Skip to content

API reference

Automatic model selection

Send the model auto and the API picks the model for each request, keeps one model through a task's tool calls, and checks your spending limit before anything runs.

Last updated: 2026-10-11

Send "model": "auto" to /v1/chat/completions and the API chooses the model for each request. Everything else in the request stays as it is: the same SDK, the same parameters, the same answer.

bash
curl https://api.aitokens.ch/v1/chat/completions \
  -H "Authorization: Bearer $AI_TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "auto", "messages": [{"role": "user", "content": "Summarise this contract in five lines: …"}]}'

The answer's model is the model that answered, and you pay its input and output rates from the catalogue. Choosing costs nothing.

How it chooses

There are two roles: a lower-cost model for routine work (today qwen3.6-35b) and a reasoning model for work that needs it (today qwen3.8-27b). For each request, in this order:

  1. A task keeps its model. A request that carries the result of a tool call (role: "tool") goes to the model that asked for that call, for two hours. Agents do not switch model halfway through a task.
  2. Long input goes to the reasoning model, which has the longer context.
  3. Otherwise a judge decides. A small model of ours, on our machines in Switzerland, reads the last user message and estimates whether it needs reasoning: maths or logic over several steps, non-trivial code, planning, analysis with many constraints, agent work with tools. If so, the reasoning model; if not, the lower-cost one. If the judge does not answer within 3 seconds, the reasoning model. Nothing the judge reads is stored.
  4. The best available. If the chosen model has no machine available now, the next one of its ranking answers. With images, only models that read them.

What it tells you

Every answer of auto carries three headers:

header what it says
X-Auto-Model the model that answered (also in model)
X-Auto-Reason routine, reasoning, long_input, pinned (a task keeping its model) or judge_unavailable; with _fallback when the first of the ranking was not available
X-Auto-Score the judge's estimate that the request needs reasoning, 0 to 1, when the judge was asked

X-Zona says where the request was processed, as for any model.

Your spending limit

In the dashboard each key can have a monthly limit in CHF and an output allowance. Before a request runs, its highest cost is reserved: the prompt tokens plus the output allowance (your max_tokens, else the key's allowance, else 8,192), at the rates of the model that will answer. If that does not fit in what is left this month, the request stops with 402 spending_limit_reached and no model runs. auto never switches to a cheaper model to make it fit.

When the key has a limit and the request sets no max_tokens, the allowance is sent as max_tokens, so the reserve is a true upper bound. The limit holds for every model, not only auto.

Reasoning

On the reasoning model, reasoning_effort: "none" switches reasoning off for that request. See Reasoning.

Good to know

  • Choosing adds the judge's time, usually a fraction of a second, to the first token. A request with a tool result skips the judge.
  • To pin a model yourself, name it: "model": "qwen3.8-27b".
  • Questions: info@daikolab.ch.

Search the docs

Type to search…