API reference
What we change in responses
Everything the gateway changes in what the models produce, and why. Declared, not hidden.
Last updated: 2026-10-08
The gateway changes six things in what the models produce, listed here, and it marks the files it generates. Nothing else. We write them down because a service whose point is saying exactly what it does cannot rewrite answers silently.
1. The model's internal delimiters
After a tool call, some models leave in the text the markers they use to separate their internal channels:
<|channel>thought
<channel|>OK. I created the task for the customer Rossi.
Those symbols are protocol, not content: anyone showing that field on a screen
would see them. We remove them, together with the word thought when it stays
alone at the start of a line.
The list of delimiters is closed: <|channel>, <channel|>,
<|start|>, <|end|>, <|message|>, <|im_start|>, <|im_end|> and the
role markers. A generic rule on <|…|> could eat legitimate text — for example
an answer that talks about syntax — and we do not use one.
2. The word "null" instead of the null value
In the arguments of a tool call, on a field that allows null, the model
sometimes writes the four-letter string. A client that checks !== null saves
it, and the database ends up with a record whose description literally says
"null".
We turn into null only these values, and only when the field contains
them entirely, after trimming spaces:
null none nil nessuno n/a undefined (empty string)
The comparison is case-insensitive. A sentence that contains the word is not
touched: "the field was null and must be fixed" stays as it is.
It applies to the arguments of tool calls, where the field has a declared type. We do not touch the free text of the answer.
3. Transcriptions from a machine in your zones
A whisper-1 transcription can be made by a machine in your key's zones or by
the server in Switzerland (which one).
The two engines write their answers differently, and we rewrite the answer of
the machine in your zones in the shape of the server's, so that your code reads
one shape:
- segments are numbered from 1, in one sequence: the engine starts again from 0 every 30 seconds of audio;
seekis in hundredths of a second: the engine gives whole seconds;startandendstay within the length of the audio: on silence,endwent past it;wordsisnullin every segment, as on the server in Switzerland without word timings, andno_speech_probisnull, because this engine does not compute it. We do not invent one;textis the text of the segments joined and trimmed, without the space in front and the double spaces the engine leaves.
The words themselves are not touched.
4. Phrases invented on silence
On silence and noise, Whisper writes phrases nobody said: «Grazie a tutti.»,
«Grazie.», «Thank you.», «so», or punctuation alone. When everything the
machine in your zones wrote is one of these phrases, each with a low confidence
(an average log-probability under -0.3), the answer is empty text, with the
duration of the file.
Only the whole answer is judged. A recording that ends with «Grazie a tutti» keeps it, and so does what the model writes over a pause between two sentences. Answers from the server in Switzerland are passed on as they come.
5. Function calls on the models served by Ollama
Ollama writes a function call in its own shape: no id, no type, and the
arguments as a JSON object. We rewrite it in OpenAI's, as the models served by
vLLM already answer: an id (Ollama's when it sends one, otherwise call_ and
24 hexadecimal characters), "type": "function", the arguments as a JSON
string, and in a stream an index for each call. A turn that ends with calls
has finish_reason: "tool_calls", where Ollama writes stop.
The other way round, when you send a call back with its arguments as a string, as the OpenAI SDKs do on the next turn, Ollama receives them as the object it reads. The name and the arguments of a call are not changed.
6. The reasoning of models that think first
gpt-oss-120b and qwen3.8-27b reason before they answer. The reasoning is
not returned: content has only the answer. Its tokens are in
completion_tokens and are billed. On qwen3.8-27b the answer may start with
two blank lines.
Files we generate carry a mark
Pictures, video clips and voices get metadata that says they were generated by AI, for article 50 of the EU AI Act: an IPTC field and a comment in a PNG, comment and description tags in an MP4 and an MP3, a comment in a WAV. The picture, the video and the sound are not changed. Which fields, and what removes them: AI-generated media.
Silence on input
This is not a change to the answer, but it belongs to the same family and it should be said.
On audio without sound, /v1/audio/transcriptions returns empty text
without asking the model. The reason: on digital silence Whisper answers
"Grazie." or "Thank you." with no_speech_prob at 0.0000, that is with full
confidence. On a dictation path an accidental tap on the microphone would turn
into a command.
The check measures the energy of the waveform, not the model's judgement, and triggers below about -50 dBFS — below are digital silence and a closed microphone; above is any speech, even whispered into a distant microphone. It applies only to 16-bit WAV files: compressed formats always go through, because we cannot measure them without decoding and when in doubt, we transcribe. Good audio thrown away is much worse than a hallucination.
What we do not touch
The text of the answer, the numbers, the language, the punctuation, the order of JSON keys. If the model gets a total wrong, that total reaches you as it was written: add up the lines and compare, as you should with any provider.