Skip to content

API reference

Every parameter: honoured or refused

The complete list of what we accept on /v1/chat/completions, what reaches the engine and what we refuse with a 400, and today's exceptions: what a stream, or an Ollama engine, accepts and does not apply.

Last updated: 2026-10-11

On /v1/chat/completions a parameter can end up in two ways:

  • it reaches the engine and does what its name says;
  • it is refused with a 400 that names it and says what to use instead.

That is the rule, and it is the most useful one on this page: if the request comes back 200, what you wrote has been applied. Today it has exceptions, and they are all here:

  • with "stream": true, two parameters are accepted and do not do in a stream what they do without it: stream_options, and logprobs with top_logprobs. n above 1 is refused. See With stream: true;
  • on the models served by Ollama, tool_choice is not applied. See Models served by Ollama.

The rule is for this endpoint. What other endpoints accept, refuse or ignore is in What works where.

Why we write it so explicitly

Because until 18 August 2026 it was not true. The gateway accepted seven parameters and dropped them without saying so: max_completion_tokens, stop, n, seed, logprobs, top_logprobs, parallel_tool_calls. It also accepted an eighth, top_p, which was even in the table of this documentation. The response came back 200 and looked right.

The worst case was n. The gateway forwarded it, the engine generated the requested answers and billed all of them, and we returned the first one and threw away the others: you paid for tokens you never saw. Worse than ignoring the parameter.

The most expensive case was max_completion_tokens, the new name OpenAI recommends for the token cap. Whoever used it had no cap at all, and found out on the invoice.

It was reported by the team building a Swiss business application on top of this API, with the sentence that decided how the fix is made: "an error is found during development, an ignored cap is found on the invoice".


Honoured

Tested on the engine one by one on 18 August 2026, then through the public gateway.

parameter what it does notes
model which model see the catalogue
messages the conversation roles system, user, assistant, tool
temperature 0–2 lower is more predictable
top_p 0–1 nucleus sampling
max_tokens cap on generated tokens up to 8192. On gpt-oss-120b and qwen3.8-27b the hidden reasoning counts against the cap: a low cap gives a cut or empty answer with finish_reason: "length". Without a cap, qwen3.8-27b has used 17,206 tokens for a 300-word story.
max_completion_tokens the same cap, new name use this one
stop where to stop string or list; we convert the string
n how many answers 1–8; you pay for all of them, and you see all of them. With stream, only 1: see below
seed repeatability with everything else equal, the same answer
presence_penalty −2…2
frequency_penalty −2…2
logit_bias per-token bias map token id → −100…100
logprobs token probabilities appear in choices[].logprobs; not in a stream
top_logprobs 0–20 alternatives per token requires logprobs: true
parallel_tool_calls several calls at once, yes or no with false they arrive one at a time
tools / tool_choice function calling see JSON and functions
response_format shape of the answer also with stream: true
stream answer in chunks see streaming; stream_options is accepted and has no effect, see below
user your own reference see below
service_tier auto or default we have a single level; for speed use X-Siati-Tier
reasoning_effort how much the model reasons none, minimal, low, medium, high, xhigh. See Reasoning
reasoning the same, in OpenRouter's form {"enabled": false} or {"effort": "low"}. A budget in tokens (max_tokens) is refused with 422

Reasoning

Each engine has its own switch, and we translate the field for it:

model what reasoning_effort does
qwen3.8-27b and the other Qwen3 models on or off: none switches reasoning off; any other level leaves it on, as without the field. Off, a short answer takes a few tokens instead of dozens
gpt-oss-120b low, medium, high reach the engine; none and minimal give low, xhigh gives high. It cannot be switched off
the others they do not reason as we serve them: the field changes nothing and costs nothing

Reasoning tokens are output tokens: they count against max_tokens and are billed at the output price.

The token cap, and its two names

max_tokens and max_completion_tokens are the same cap. If you send one, that one applies. If you send both with the same number, fine. If you send both with different numbers we answer 400: two different caps in the same request mean that whoever wrote it believes in one of them, and guessing which is how a customer ends up paying for our guess.

user, and where it ends up

The user field is stored with the usage record of the request. It is for you, to separate your users, departments or customers within a single invoice. It is a string you give us, kept as it is: we do not link it to anything and do not use it for anything else.

It is stored the same way with stream: true. One limit today: the dashboard does not show it yet, so ask info@daikolab.ch for an extract.

If you are looking for an X-Siati-User header, it is gone: it was documented and never implemented. Use user, which is the OpenAI name and works with the libraries you already have.

Models served by Ollama

The models served by vLLM apply every parameter above. Some models are served by Ollama instead: today the qwen2.5 models, where the catalogue has them. On those:

  • n above 1, logprobs, top_logprobs, logit_bias and parallel_tool_calls get a 400 naming the parameter and saying on which models it works, not a 502 and not silence. n: 1 is accepted;
  • tool_choice is accepted and not applied: the model decides by itself whether to call a function;
  • response_format is applied with and without stream, and {"type": "text"} gives free text;
  • function calls arrive in OpenAI's shape, as on the other models: Ollama writes them in its own, and we rewrite them, see What we change in responses.

With stream: true

parameter what happens in a stream
stream_options include_usage: true adds a last chunk with an empty choices and the usage of the request, the counts that are billed.
logprobs, top_logprobs they reach the engine, but the probabilities do not travel in the chunks.
n above 1 refused with 400, code: invalid_value, param: n: a stream carries one answer, and the pieces of several would arrive mixed. Ask for one answer per stream, or call without stream to get every answer in choices.

Everything else behaves in a stream as it does without one, response_format and tools included, on the models served by Ollama too.


Refused, with the reason

parameter why
metadata there is no stored completion to attach labels to. Use user.
store completions are not stored for you to retrieve later. What the request log keeps, and for how long, is in Sovereignty.
functions outdated form: use tools.
function_call outdated form: use tool_choice.
modalities, audio for voice there are dedicated endpoints.
prediction the engine does not do it.
web_search_options our models do not go out to the internet, by design.
chat_template_kwargs engine-specific: use reasoning_effort instead, which we translate for each engine (Reasoning).

Any other parameter

In a JSON body, which is what the SDKs send, a name we do not recognise gives 400 and is named. That includes typos, which is the most useful part: tempreature used to go through, the temperature stayed at the default and there was no way to notice it from the answer.

The answer names the field:

json
{
  "error": {
    "message": "Unrecognized request argument supplied: 'tempreature'. We refuse what we do not know instead of ignoring it: an argument accepted and dropped does not show in the answer. Accepted: …",
    "type": "invalid_request_error",
    "param": "tempreature"
  }
}

Transcription: timestamp_granularities

This field follows the same rule. timestamp_granularities: ["word"] returns per-word timings in words, with start, end and probability of each. Until 18 August the field was accepted and not forwarded, so only segments came back.

Timings exist only inside response_format: "verbose_json": the other two forms have nowhere to put them, and asking for them without verbose_json gives 400 instead of an answer without timings, which could not be told apart from audio in which the words could not be heard.

bash
curl https://api.aitokens.ch/v1/audio/transcriptions \
  -H "Authorization: Bearer $API_KEY" \
  -F file=@note.wav \
  -F model=whisper-1 \
  -F response_format=verbose_json \
  -F "timestamp_granularities[]=word"

How we avoid repeating the mistake

The cause was not someone's distraction: request validation returned only the fields it listed, and everything else vanished without an error and without a log line. A defect like that is not found by rereading the code, because there is nothing to see.

Now the list of parameters lives in one place, and an automated test fails if a field is validated without ending up either at the engine or among the refused. We checked it by deliberately adding a forgotten parameter: the test named it.

It does not protect us from not supporting something you need. It protects us from pretending to support it, which is what costs you the most.

Search the docs

Type to search…