API reference
Rate limits
The two limits every /v1 request passes, both counted per key; what is still counted per address; what each 429 looks like, and how to retry.
Last updated: 2026-10-11
A request to /v1 passes up to two limits, both counted per minute and both per API key: every key has its own counters, wherever it calls from. They are independent: the first one reached answers 429. A request stopped by a limit is not run, and costs nothing.
| Counted per | Value | Applies to | |
|---|---|---|---|
| The tier's limit | API key | set by the tier of the request | every /v1 endpoint, /v1/models included |
| The endpoint's limit | API key | fixed per endpoint | the /v1 endpoints in the table below |
A request without a valid key is refused with 401 before both, and counts toward neither. Neither of them promises speed. A request that passes both may still wait in a machine's queue, and the tier does not change your place in it: see Tiers.
The tier's limit
Each tier has a number of requests per minute, kept in the database. The values are in the tiers table of Pricing; the 429 message states the limit as well.
How it counts:
- One counter per key, for all
/v1endpoints together. Listing the models counts too. - The minute is the clock minute. The counter starts again at the beginning of each minute: it is not a sliding window.
- Each request is compared with the limit of its own tier. If you mix tiers on one key with
X-Siati-Tier, they share the counter.
When you exceed it:
HTTP/1.1 429 Too Many Requests
Retry-After: 23
Content-Type: application/json
{"error": {"message": "rate limit exceeded (…/min)", "type": "rate_limit_exceeded"}}
Retry-After is the number of seconds until the next minute begins. No response tells you how many requests your key has left.
The endpoint's limit
| Endpoint | Requests per minute | Counter |
|---|---|---|
POST /v1/chat/completions, /v1/completions, /v1/responses |
60 |
shared |
POST /v1/knowledge-bases/{slug}/ask |
30 |
shared |
POST /v1/embeddings, /v1/rerank |
120 |
shared |
POST /v1/audio/transcriptions |
30 |
shared |
POST /v1/audio/speech |
120 |
shared |
POST /v1/images/generations |
30 |
its own |
POST /v1/videos |
20 |
its own |
GET /v1/videos |
120 |
its own |
GET /v1/videos/{id} |
600 |
its own |
GET /v1/videos/{id}/content |
120 |
its own |
DELETE /v1/videos/{id} |
60 |
its own |
GET /v1/models, GET /v1/knowledge-bases |
not limited |
Pictures and video are served where AI Tokens has machines that make them; their limits apply wherever the endpoints answer.
How it counts:
- Per key, after the key is checked. Every key has its own counters: two people behind one address, or two projects of one account with a key each, do not share them. A request with a wrong key is refused before it is counted.
- One counter per key for the endpoints marked "shared", which each endpoint compares with its own limit. After 30 chat requests in a minute, transcriptions with the same key answer
429until the minute is over, while embeddings still have 90 to go. The endpoints marked "its own" have a counter each: asking every two seconds how a video is going does not use up the limit of chat completions. - The minute starts with the first request counted, not on the clock.
Responses of these endpoints carry two headers about this limit, for the endpoint you called and your key. They do not describe the tier's limit:
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 59
When you exceed it, the body is not the error object of the rest of the API:
HTTP/1.1 429 Too Many Requests
Retry-After: 60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1791028838
Content-Type: application/json
{"message": "Too Many Attempts."}
Retry-After is the number of seconds until the window ends; X-RateLimit-Reset is the same moment as a Unix timestamp.
What is counted per address
Both limits of /v1 count per key: two projects with a key each do not share them. What arrives with no key and no signed-in user is counted per address:
| Request | Requests per minute, per address |
|---|---|
GET a picture link, /v1/images/file/… |
120, its own counter |
| App API: signing in, registering, asking for and confirming a new password | 10 |
App API: refreshing the token, POST /auth/refresh |
60 |
One video at a time per person: a second video request while your first is queued or in progress answers 429 with code too_many_video_jobs.
What you see when a limit is reached:
| Answer | What it means | What to do |
|---|---|---|
429 with {"message": "Too Many Attempts."} and X-RateLimit-Reset |
The endpoint's limit, for your key only | Wait Retry-After seconds. If it happens often, your loop sends more than the endpoint allows: space out its requests |
429 with "type": "rate_limit_exceeded" |
The tier's limit, for your key only | Wait Retry-After seconds: the counter starts again at the next minute |
429 with "code": "too_many_video_jobs" |
A video of yours is still being made; the message names it | Wait for it to finish, then ask for the next one |
402 with "code": "insufficient_credits" |
No credits left: this is not a rate limit | Retrying does not help: write to info@daikolab.ch for more credits |
A request stopped by a limit is not run and costs nothing.
Telling them apart
| The tier's limit | The endpoint's limit | |
|---|---|---|
| Body | {"error": {"type": "rate_limit_exceeded", …}} |
{"message": "Too Many Attempts."} |
Retry-After |
seconds until the next clock minute | seconds until the window ends |
X-RateLimit-Reset |
absent | present |
In both cases, wait Retry-After seconds before you try again.
Retrying
The OpenAI SDKs already retry a failed request a few times on their own; you can change the number of attempts with max_retries in Python, maxRetries in Node.
Without an SDK, read Retry-After and add some randomness, so that many clients do not all come back in the same second:
import os
import random
import time
import requests
URL = "https://api.aitokens.ch/v1/chat/completions"
HEADERS = {"Authorization": f"Bearer {os.environ['API_KEY']}"}
def post_with_retry(body, attempts=5):
for attempt in range(attempts):
r = requests.post(URL, headers=HEADERS, json=body, timeout=120)
if r.status_code not in (429, 502, 503):
return r
if r.status_code == 503 and "model_activation_required" in r.text:
return r # retrying will not change it
wait = float(r.headers.get("Retry-After", 2 ** attempt))
time.sleep(wait + random.uniform(0, 1))
return r
r = post_with_retry({
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Hello"}],
})
print(r.status_code, r.json())
Which errors are worth retrying is in Errors.
The app API
The app API (https://my.aitokens.ch/api/v1, see Authentication) has only the endpoint limits, counted in the same way: one counter per signed-in user for the calls made with a user token, one per address for signing in and refreshing the token, each shared by the endpoints of its kind. Only the counters per address are twenty times wider from trusted networks.
| Endpoint | Requests per minute | Counted per |
|---|---|---|
POST /auth/login, /auth/register, /auth/forgot-password, /auth/password-reset/confirm |
10 |
address |
POST /auth/refresh |
60 |
address |
POST /chat/sessions/{id}/messages |
30 |
user |
POST /rag/kb |
30 |
user |
POST /rag/kb/{slug}/docs (upload) |
10 |
user |
POST /rag/kb/{slug}/chat |
30 |
user |
POST /audio/transcribe |
30 |
user |
POST /audio/synthesize |
60 |
user |
If you need more
The tier's limit changes with the tier: compare the values in Pricing. For capacity reserved to you, with nobody ahead of you in the queue, write to info@daikolab.ch.