API reference
Structured output and function calling
Constraining the shape of the answer with response_format, json_schema and tools: what the constraint gives you, what it does not, and the checks your code still makes.
Last updated: 2026-10-06
If you need to extract fields from a document and load them into a business system, this is the page you need. With a schema, the engine constrains the shape of the answer: a complete answer comes back as JSON with the keys and types you declared, and you do not need a tolerant parser to read it.
The schema does not make the answer right. The values can be wrong, an answer cut short is not an object at all, and what your code does with the result is your decision. The checks at the end of this page are the ones to make, in order.
Both features follow the shape of the OpenAI API, so they work with the
libraries you already use. They work on the chat models of the
catalogue; on the models served by Ollama tool_choice is not
applied, see Every parameter.
JSON mode
With response_format the engine constrains generation to JSON.
curl https://api.aitokens.ch/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Extract the fields from this invoice: ..."}],
"response_format": {"type": "json_object"}
}'
When the answer ends with finish_reason: "stop", the engine has kept it
inside the grammar, and it parses. When it ends with "length", it was cut at
the token cap: what you get is the beginning of an object, such as
{"document_number": "202, and it does not parse. Raise max_completion_tokens
or ask for fewer fields.
With a schema, if you also want the right fields
json_object constrains the answer to JSON. To constrain which keys are
there and of which type, pass the schema:
{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "…"}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "invoice",
"schema": {
"type": "object",
"properties": {
"document_number": {"type": "string"},
"date": {"type": "string"},
"total": {"type": "number"},
"vat_rate": {"type": "number"}
},
"required": ["document_number", "date", "total"]
}
}
}
}
A complete answer has the required keys, with the declared types. Validate it against your schema anyway: it costs one line, and it is what tells a complete answer from one that was cut or not constrained.
Function calling
Pass tools and the model answers with tool_calls instead of text, exactly as
with the OpenAI API.
{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Record invoice 2026-114 for 1,250 francs"}],
"tools": [{
"type": "function",
"function": {
"name": "record_invoice",
"description": "Create a draft supplier invoice",
"parameters": {
"type": "object",
"properties": {
"number": {"type": "string"},
"amount": {"type": "number"}
},
"required": ["number", "amount"]
}
}
}]
}
The response contains choices[0].message.tool_calls, with function.name and
function.arguments as a JSON string, and finish_reason is tool_calls. The
arguments are written by the model: parse and validate them like any other
answer, and let your code decide whether the call runs.
Until when it did not work
From 16 to 18 August 2026 the gateway did not forward response_format to
the engine: whoever passed it got free prose, or an object wrapped in a markdown
fence, while this page promised the opposite. Fixed on 18 August. If you wrote a
downstream validation to work around it, keep it: it is the third of the checks
below.
With stream: true the defect lasted two more days: the format was forwarded on
the normal path but not on the streaming one, so whoever asked for JSON while
streaming still got prose. That was also fixed on 18 August; response_format
now works on both paths, and since 6 October on the models served by Ollama too.
In the same round tool_choice: "required" was fixed; it used to answer 502.
The cause was not the value: a "properties": {} in the schema reached the
engine as [], and the engine rejected the grammar.
What the constraint does not give you
An answer cut short is not an object. With finish_reason: "length" the JSON
stops in the middle. In a stream, an answer that ends without [DONE], or with
finish_reason: "error", is cut short as well: see Streaming.
A refusal is not data. There is no refusal field in our answers. A model
that declines writes it where it can: as prose when the format is not applied,
or inside the fields when it is — an apology as a description, an empty
required string. Treat values that are not data as no answer.
Numbers must be recomputed. The model reads well, but a total that was read is not a total that was checked: add up the lines and compare. It holds for any model, not only ours, and it is why the result should always land in a draft to be confirmed, never straight into the books.
The schema constrains the shape, not the truth. If the date field is
required, it is in the answer; whether it is the right date is for your check to
say, not us.
enum inside tools is not enforced. In the response_format schema it
is; in the parameters of a function it is not, because there the constraint does
not go through the grammar. The model tends to answer with the word in the
language of the conversation: "urgente" instead of "urgent". Two things
work, and integrators taught them to us: putting the mapping in the property's
description ("urgente/urgentissimo=urgent, alta=high") raises adherence a lot,
and validating downstream remains necessary.
The model sometimes writes the string "null". On a field that allows
null, the four-letter word may arrive instead of the JSON value. We fix it in
tool-call arguments (what we change); in your own
parsing, treat "null", "none", "nessuno" and the empty string as absence.
On a Swiss invoice, read the QR code first. If the document has the payment part with the code, IBAN, reference, amount and creditor can be read deterministically, without a model and without margin of error. Use the model for what the code does not contain: document number, dates, lines, VAT rates. It is the right order, and it costs you less.
The checks, in order
Each check assumes the one before it has passed.
- Transport. The HTTP status is
200and the answer arrived within your own timeout. In a stream,[DONE]arrived. Keep theX-Request-Id: it is what to quote if you write to us. - Protocol.
finish_reasonisstop, ortool_callswhen you asked for a function; notlength, noterror. There is noerrorobject, and the answer is not a refusal. - Structure. The content parses as JSON, and the object validates against your schema: required keys, types, enums.
- Meaning. The values make sense for the task: totals recomputed from the lines, units and currencies, dates that exist, references that match the source (an IBAN's check digits, an invoice number), line numbers that are in the document.
- Action. What happens next is authorised by your code, not by the model: a draft that someone confirms, and a person's approval wherever money, people's data or a production system are involved.
The first three, in Python:
import json
import os
from jsonschema import validate # pip install jsonschema
from openai import OpenAI
client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"], timeout=60)
schema = {
"type": "object",
"properties": {
"document_number": {"type": "string"},
"date": {"type": "string"},
"total": {"type": "number"},
"vat_rate": {"type": "number"},
},
"required": ["document_number", "date", "total"],
}
# 1. Transport: errors and timeouts raise an exception here.
resp = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Extract the fields: INVOICE 2026-0417, 14 September 2026, VAT 8.1%, TOTAL CHF 5026.65"}],
response_format={"type": "json_schema", "json_schema": {"name": "invoice", "schema": schema}},
max_completion_tokens=300,
)
choice = resp.choices[0]
# 2. Protocol: a cut answer is not an object.
if choice.finish_reason != "stop":
raise RuntimeError(f"incomplete answer ({choice.finish_reason}), request {resp.id}")
# 3. Structure: parse, then validate against the schema.
invoice = json.loads(choice.message.content)
validate(invoice, schema)
# 4. Meaning and 5. action are yours: compare with the document, then make a draft.
print(invoice)
Search on your documents
To upload documents, ask a question and get the answer with the passage it comes from, see Knowledge bases. If your case is extracting fields from invoices you do not need it — the two features above are enough.