Skip to content

API reference

What we change in responses

Everything the gateway changes in what the models produce, and why. Declared, not hidden.

Last updated: 2026-10-08

The gateway changes six things in what the models produce, listed here, and it marks the files it generates. Nothing else. We write them down because a service whose point is saying exactly what it does cannot rewrite answers silently.

1. The model's internal delimiters

After a tool call, some models leave in the text the markers they use to separate their internal channels:

text
<|channel>thought
<channel|>OK. I created the task for the customer Rossi.

Those symbols are protocol, not content: anyone showing that field on a screen would see them. We remove them, together with the word thought when it stays alone at the start of a line.

The list of delimiters is closed: <|channel>, <channel|>, <|start|>, <|end|>, <|message|>, <|im_start|>, <|im_end|> and the role markers. A generic rule on <|…|> could eat legitimate text — for example an answer that talks about syntax — and we do not use one.

2. The word "null" instead of the null value

In the arguments of a tool call, on a field that allows null, the model sometimes writes the four-letter string. A client that checks !== null saves it, and the database ends up with a record whose description literally says "null".

We turn into null only these values, and only when the field contains them entirely, after trimming spaces:

text
null    none    nil    nessuno    n/a    undefined    (empty string)

The comparison is case-insensitive. A sentence that contains the word is not touched: "the field was null and must be fixed" stays as it is.

It applies to the arguments of tool calls, where the field has a declared type. We do not touch the free text of the answer.

3. Transcriptions from a machine in your zones

A whisper-1 transcription can be made by a machine in your key's zones or by the server in Switzerland (which one). The two engines write their answers differently, and we rewrite the answer of the machine in your zones in the shape of the server's, so that your code reads one shape:

  • segments are numbered from 1, in one sequence: the engine starts again from 0 every 30 seconds of audio;
  • seek is in hundredths of a second: the engine gives whole seconds;
  • start and end stay within the length of the audio: on silence, end went past it;
  • words is null in every segment, as on the server in Switzerland without word timings, and no_speech_prob is null, because this engine does not compute it. We do not invent one;
  • text is the text of the segments joined and trimmed, without the space in front and the double spaces the engine leaves.

The words themselves are not touched.

4. Phrases invented on silence

On silence and noise, Whisper writes phrases nobody said: «Grazie a tutti.», «Grazie.», «Thank you.», «so», or punctuation alone. When everything the machine in your zones wrote is one of these phrases, each with a low confidence (an average log-probability under -0.3), the answer is empty text, with the duration of the file.

Only the whole answer is judged. A recording that ends with «Grazie a tutti» keeps it, and so does what the model writes over a pause between two sentences. Answers from the server in Switzerland are passed on as they come.

5. Function calls on the models served by Ollama

Ollama writes a function call in its own shape: no id, no type, and the arguments as a JSON object. We rewrite it in OpenAI's, as the models served by vLLM already answer: an id (Ollama's when it sends one, otherwise call_ and 24 hexadecimal characters), "type": "function", the arguments as a JSON string, and in a stream an index for each call. A turn that ends with calls has finish_reason: "tool_calls", where Ollama writes stop.

The other way round, when you send a call back with its arguments as a string, as the OpenAI SDKs do on the next turn, Ollama receives them as the object it reads. The name and the arguments of a call are not changed.

6. The reasoning of models that think first

gpt-oss-120b and qwen3.8-27b reason before they answer. The reasoning is not returned: content has only the answer. Its tokens are in completion_tokens and are billed. On qwen3.8-27b the answer may start with two blank lines.

Files we generate carry a mark

Pictures, video clips and voices get metadata that says they were generated by AI, for article 50 of the EU AI Act: an IPTC field and a comment in a PNG, comment and description tags in an MP4 and an MP3, a comment in a WAV. The picture, the video and the sound are not changed. Which fields, and what removes them: AI-generated media.

Silence on input

This is not a change to the answer, but it belongs to the same family and it should be said.

On audio without sound, /v1/audio/transcriptions returns empty text without asking the model. The reason: on digital silence Whisper answers "Grazie." or "Thank you." with no_speech_prob at 0.0000, that is with full confidence. On a dictation path an accidental tap on the microphone would turn into a command.

The check measures the energy of the waveform, not the model's judgement, and triggers below about -50 dBFS — below are digital silence and a closed microphone; above is any speech, even whispered into a distant microphone. It applies only to 16-bit WAV files: compressed formats always go through, because we cannot measure them without decoding and when in doubt, we transcribe. Good audio thrown away is much worse than a hallucination.

What we do not touch

The text of the answer, the numbers, the language, the punctuation, the order of JSON keys. If the model gets a total wrong, that total reaches you as it was written: add up the lines and compare, as you should with any provider.

Search the docs

Type to search…