API reference
Every parameter: honoured or refused
The complete list of what we accept on /v1/chat/completions, what reaches the engine and what we refuse with a 400, and today's exceptions: what a stream, or an Ollama engine, accepts and does not apply.
Last updated: 2026-10-11
On /v1/chat/completions a parameter can end up in two ways:
- it reaches the engine and does what its name says;
- it is refused with a 400 that names it and says what to use instead.
That is the rule, and it is the most useful one on this page: if the request comes back 200, what you wrote has been applied. Today it has exceptions, and they are all here:
- with
"stream": true, two parameters are accepted and do not do in a stream what they do without it:stream_options, andlogprobswithtop_logprobs.nabove 1 is refused. See With stream: true; - on the models served by Ollama,
tool_choiceis not applied. See Models served by Ollama.
The rule is for this endpoint. What other endpoints accept, refuse or ignore is in What works where.
Why we write it so explicitly
Because until 18 August 2026 it was not true. The gateway accepted seven
parameters and dropped them without saying so: max_completion_tokens, stop,
n, seed, logprobs, top_logprobs, parallel_tool_calls. It also
accepted an eighth, top_p, which was even in the table of this documentation.
The response came back 200 and looked right.
The worst case was n. The gateway forwarded it, the engine generated the
requested answers and billed all of them, and we returned the first one and
threw away the others: you paid for tokens you never saw. Worse than ignoring
the parameter.
The most expensive case was max_completion_tokens, the new name OpenAI
recommends for the token cap. Whoever used it had no cap at all, and found out
on the invoice.
It was reported by the team building a Swiss business application on top of this API, with the sentence that decided how the fix is made: "an error is found during development, an ignored cap is found on the invoice".
Honoured
Tested on the engine one by one on 18 August 2026, then through the public gateway.
| parameter | what it does | notes |
|---|---|---|
model |
which model | see the catalogue |
messages |
the conversation | roles system, user, assistant, tool |
temperature |
0–2 | lower is more predictable |
top_p |
0–1 | nucleus sampling |
max_tokens |
cap on generated tokens | up to 8192. On gpt-oss-120b and qwen3.8-27b the hidden reasoning counts against the cap: a low cap gives a cut or empty answer with finish_reason: "length". Without a cap, qwen3.8-27b has used 17,206 tokens for a 300-word story. |
max_completion_tokens |
the same cap, new name | use this one |
stop |
where to stop | string or list; we convert the string |
n |
how many answers | 1–8; you pay for all of them, and you see all of them. With stream, only 1: see below |
seed |
repeatability | with everything else equal, the same answer |
presence_penalty |
−2…2 | |
frequency_penalty |
−2…2 | |
logit_bias |
per-token bias | map token id → −100…100 |
logprobs |
token probabilities | appear in choices[].logprobs; not in a stream |
top_logprobs |
0–20 alternatives per token | requires logprobs: true |
parallel_tool_calls |
several calls at once, yes or no | with false they arrive one at a time |
tools / tool_choice |
function calling | see JSON and functions |
response_format |
shape of the answer | also with stream: true |
stream |
answer in chunks | see streaming; stream_options is accepted and has no effect, see below |
user |
your own reference | see below |
service_tier |
auto or default |
we have a single level; for speed use X-Siati-Tier |
reasoning_effort |
how much the model reasons | none, minimal, low, medium, high, xhigh. See Reasoning |
reasoning |
the same, in OpenRouter's form | {"enabled": false} or {"effort": "low"}. A budget in tokens (max_tokens) is refused with 422 |
Reasoning
Each engine has its own switch, and we translate the field for it:
| model | what reasoning_effort does |
|---|---|
qwen3.8-27b and the other Qwen3 models |
on or off: none switches reasoning off; any other level leaves it on, as without the field. Off, a short answer takes a few tokens instead of dozens |
gpt-oss-120b |
low, medium, high reach the engine; none and minimal give low, xhigh gives high. It cannot be switched off |
| the others | they do not reason as we serve them: the field changes nothing and costs nothing |
Reasoning tokens are output tokens: they count against max_tokens and are
billed at the output price.
The token cap, and its two names
max_tokens and max_completion_tokens are the same cap. If you send one, that
one applies. If you send both with the same number, fine. If you send both with
different numbers we answer 400: two different caps in the same request mean
that whoever wrote it believes in one of them, and guessing which is how a
customer ends up paying for our guess.
user, and where it ends up
The user field is stored with the usage record of the request. It is for
you, to separate your users, departments or customers within a single invoice.
It is a string you give us, kept as it is: we do not link it to anything and do
not use it for anything else.
It is stored the same way with stream: true. One limit today: the dashboard
does not show it yet, so ask info@daikolab.ch for an extract.
If you are looking for an X-Siati-User header, it is gone: it was documented
and never implemented. Use user, which is the OpenAI name and works with the
libraries you already have.
Models served by Ollama
The models served by vLLM apply every parameter above. Some models are served
by Ollama instead: today the qwen2.5 models, where the
catalogue has them. On those:
nabove 1,logprobs,top_logprobs,logit_biasandparallel_tool_callsget a 400 naming the parameter and saying on which models it works, not a 502 and not silence.n: 1is accepted;tool_choiceis accepted and not applied: the model decides by itself whether to call a function;response_formatis applied with and withoutstream, and{"type": "text"}gives free text;- function calls arrive in OpenAI's shape, as on the other models: Ollama writes them in its own, and we rewrite them, see What we change in responses.
With stream: true
| parameter | what happens in a stream |
|---|---|
stream_options |
include_usage: true adds a last chunk with an empty choices and the usage of the request, the counts that are billed. |
logprobs, top_logprobs |
they reach the engine, but the probabilities do not travel in the chunks. |
n above 1 |
refused with 400, code: invalid_value, param: n: a stream carries one answer, and the pieces of several would arrive mixed. Ask for one answer per stream, or call without stream to get every answer in choices. |
Everything else behaves in a stream as it does without one, response_format
and tools included, on the models served by Ollama too.
Refused, with the reason
| parameter | why |
|---|---|
metadata |
there is no stored completion to attach labels to. Use user. |
store |
completions are not stored for you to retrieve later. What the request log keeps, and for how long, is in Sovereignty. |
functions |
outdated form: use tools. |
function_call |
outdated form: use tool_choice. |
modalities, audio |
for voice there are dedicated endpoints. |
prediction |
the engine does not do it. |
web_search_options |
our models do not go out to the internet, by design. |
chat_template_kwargs |
engine-specific: use reasoning_effort instead, which we translate for each engine (Reasoning). |
Any other parameter
In a JSON body, which is what the SDKs send, a name we do not recognise gives
400 and is named. That includes typos, which
is the most useful part: tempreature used to go through, the temperature
stayed at the default and there was no way to notice it from the answer.
The answer names the field:
{
"error": {
"message": "Unrecognized request argument supplied: 'tempreature'. We refuse what we do not know instead of ignoring it: an argument accepted and dropped does not show in the answer. Accepted: …",
"type": "invalid_request_error",
"param": "tempreature"
}
}
Transcription: timestamp_granularities
This field follows the same rule. timestamp_granularities: ["word"] returns per-word
timings in words, with start, end and probability of each. Until 18 August the
field was accepted and not forwarded, so only segments came back.
Timings exist only inside response_format: "verbose_json": the other two forms
have nowhere to put them, and asking for them without verbose_json gives 400
instead of an answer without timings, which could not be told apart from audio
in which the words could not be heard.
curl https://api.aitokens.ch/v1/audio/transcriptions \
-H "Authorization: Bearer $API_KEY" \
-F file=@note.wav \
-F model=whisper-1 \
-F response_format=verbose_json \
-F "timestamp_granularities[]=word"
How we avoid repeating the mistake
The cause was not someone's distraction: request validation returned only the fields it listed, and everything else vanished without an error and without a log line. A defect like that is not found by rereading the code, because there is nothing to see.
Now the list of parameters lives in one place, and an automated test fails if a field is validated without ending up either at the engine or among the refused. We checked it by deliberately adding a forgotten parameter: the test named it.
It does not protect us from not supporting something you need. It protects us from pretending to support it, which is what costs you the most.