Skip to content

API reference

Rate limits

The two limits every /v1 request passes, both counted per key; what is still counted per address; what each 429 looks like, and how to retry.

Last updated: 2026-10-11

A request to /v1 passes up to two limits, both counted per minute and both per API key: every key has its own counters, wherever it calls from. They are independent: the first one reached answers 429. A request stopped by a limit is not run, and costs nothing.

Counted per Value Applies to
The tier's limit API key set by the tier of the request every /v1 endpoint, /v1/models included
The endpoint's limit API key fixed per endpoint the /v1 endpoints in the table below

A request without a valid key is refused with 401 before both, and counts toward neither. Neither of them promises speed. A request that passes both may still wait in a machine's queue, and the tier does not change your place in it: see Tiers.

The tier's limit

Each tier has a number of requests per minute, kept in the database. The values are in the tiers table of Pricing; the 429 message states the limit as well.

How it counts:

  • One counter per key, for all /v1 endpoints together. Listing the models counts too.
  • The minute is the clock minute. The counter starts again at the beginning of each minute: it is not a sliding window.
  • Each request is compared with the limit of its own tier. If you mix tiers on one key with X-Siati-Tier, they share the counter.

When you exceed it:

http
HTTP/1.1 429 Too Many Requests
Retry-After: 23
Content-Type: application/json

{"error": {"message": "rate limit exceeded (…/min)", "type": "rate_limit_exceeded"}}

Retry-After is the number of seconds until the next minute begins. No response tells you how many requests your key has left.

The endpoint's limit

Endpoint Requests per minute Counter
POST /v1/chat/completions, /v1/completions, /v1/responses 60 shared
POST /v1/knowledge-bases/{slug}/ask 30 shared
POST /v1/embeddings, /v1/rerank 120 shared
POST /v1/audio/transcriptions 30 shared
POST /v1/audio/speech 120 shared
POST /v1/images/generations 30 its own
POST /v1/videos 20 its own
GET /v1/videos 120 its own
GET /v1/videos/{id} 600 its own
GET /v1/videos/{id}/content 120 its own
DELETE /v1/videos/{id} 60 its own
GET /v1/models, GET /v1/knowledge-bases not limited

Pictures and video are served where AI Tokens has machines that make them; their limits apply wherever the endpoints answer.

How it counts:

  • Per key, after the key is checked. Every key has its own counters: two people behind one address, or two projects of one account with a key each, do not share them. A request with a wrong key is refused before it is counted.
  • One counter per key for the endpoints marked "shared", which each endpoint compares with its own limit. After 30 chat requests in a minute, transcriptions with the same key answer 429 until the minute is over, while embeddings still have 90 to go. The endpoints marked "its own" have a counter each: asking every two seconds how a video is going does not use up the limit of chat completions.
  • The minute starts with the first request counted, not on the clock.

Responses of these endpoints carry two headers about this limit, for the endpoint you called and your key. They do not describe the tier's limit:

http
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 59

When you exceed it, the body is not the error object of the rest of the API:

http
HTTP/1.1 429 Too Many Requests
Retry-After: 60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1791028838
Content-Type: application/json

{"message": "Too Many Attempts."}

Retry-After is the number of seconds until the window ends; X-RateLimit-Reset is the same moment as a Unix timestamp.

What is counted per address

Both limits of /v1 count per key: two projects with a key each do not share them. What arrives with no key and no signed-in user is counted per address:

Request Requests per minute, per address
GET a picture link, /v1/images/file/… 120, its own counter
App API: signing in, registering, asking for and confirming a new password 10
App API: refreshing the token, POST /auth/refresh 60

One video at a time per person: a second video request while your first is queued or in progress answers 429 with code too_many_video_jobs.

What you see when a limit is reached:

Answer What it means What to do
429 with {"message": "Too Many Attempts."} and X-RateLimit-Reset The endpoint's limit, for your key only Wait Retry-After seconds. If it happens often, your loop sends more than the endpoint allows: space out its requests
429 with "type": "rate_limit_exceeded" The tier's limit, for your key only Wait Retry-After seconds: the counter starts again at the next minute
429 with "code": "too_many_video_jobs" A video of yours is still being made; the message names it Wait for it to finish, then ask for the next one
402 with "code": "insufficient_credits" No credits left: this is not a rate limit Retrying does not help: write to info@daikolab.ch for more credits

A request stopped by a limit is not run and costs nothing.

Telling them apart

The tier's limit The endpoint's limit
Body {"error": {"type": "rate_limit_exceeded", …}} {"message": "Too Many Attempts."}
Retry-After seconds until the next clock minute seconds until the window ends
X-RateLimit-Reset absent present

In both cases, wait Retry-After seconds before you try again.

Retrying

The OpenAI SDKs already retry a failed request a few times on their own; you can change the number of attempts with max_retries in Python, maxRetries in Node.

Without an SDK, read Retry-After and add some randomness, so that many clients do not all come back in the same second:

python
import os
import random
import time

import requests

URL = "https://api.aitokens.ch/v1/chat/completions"
HEADERS = {"Authorization": f"Bearer {os.environ['API_KEY']}"}


def post_with_retry(body, attempts=5):
    for attempt in range(attempts):
        r = requests.post(URL, headers=HEADERS, json=body, timeout=120)
        if r.status_code not in (429, 502, 503):
            return r
        if r.status_code == 503 and "model_activation_required" in r.text:
            return r  # retrying will not change it
        wait = float(r.headers.get("Retry-After", 2 ** attempt))
        time.sleep(wait + random.uniform(0, 1))
    return r


r = post_with_retry({
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Hello"}],
})
print(r.status_code, r.json())

Which errors are worth retrying is in Errors.

The app API

The app API (https://my.aitokens.ch/api/v1, see Authentication) has only the endpoint limits, counted in the same way: one counter per signed-in user for the calls made with a user token, one per address for signing in and refreshing the token, each shared by the endpoints of its kind. Only the counters per address are twenty times wider from trusted networks.

Endpoint Requests per minute Counted per
POST /auth/login, /auth/register, /auth/forgot-password, /auth/password-reset/confirm 10 address
POST /auth/refresh 60 address
POST /chat/sessions/{id}/messages 30 user
POST /rag/kb 30 user
POST /rag/kb/{slug}/docs (upload) 10 user
POST /rag/kb/{slug}/chat 30 user
POST /audio/transcribe 30 user
POST /audio/synthesize 60 user

If you need more

The tier's limit changes with the tier: compare the values in Pricing. For capacity reserved to you, with nobody ahead of you in the queue, write to info@daikolab.ch.

Search the docs

Type to search…