Skip to content

API reference

Structured output and function calling

Constraining the shape of the answer with response_format, json_schema and tools: what the constraint gives you, what it does not, and the checks your code still makes.

Last updated: 2026-10-06

If you need to extract fields from a document and load them into a business system, this is the page you need. With a schema, the engine constrains the shape of the answer: a complete answer comes back as JSON with the keys and types you declared, and you do not need a tolerant parser to read it.

The schema does not make the answer right. The values can be wrong, an answer cut short is not an object at all, and what your code does with the result is your decision. The checks at the end of this page are the ones to make, in order.

Both features follow the shape of the OpenAI API, so they work with the libraries you already use. They work on the chat models of the catalogue; on the models served by Ollama tool_choice is not applied, see Every parameter.


JSON mode

With response_format the engine constrains generation to JSON.

bash
curl https://api.aitokens.ch/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Extract the fields from this invoice: ..."}],
    "response_format": {"type": "json_object"}
  }'

When the answer ends with finish_reason: "stop", the engine has kept it inside the grammar, and it parses. When it ends with "length", it was cut at the token cap: what you get is the beginning of an object, such as {"document_number": "202, and it does not parse. Raise max_completion_tokens or ask for fewer fields.

With a schema, if you also want the right fields

json_object constrains the answer to JSON. To constrain which keys are there and of which type, pass the schema:

json
{
  "model": "qwen3.8-27b",
  "messages": [{"role": "user", "content": "…"}],
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "invoice",
      "schema": {
        "type": "object",
        "properties": {
          "document_number": {"type": "string"},
          "date":            {"type": "string"},
          "total":           {"type": "number"},
          "vat_rate":        {"type": "number"}
        },
        "required": ["document_number", "date", "total"]
      }
    }
  }
}

A complete answer has the required keys, with the declared types. Validate it against your schema anyway: it costs one line, and it is what tells a complete answer from one that was cut or not constrained.


Function calling

Pass tools and the model answers with tool_calls instead of text, exactly as with the OpenAI API.

json
{
  "model": "qwen3.8-27b",
  "messages": [{"role": "user", "content": "Record invoice 2026-114 for 1,250 francs"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "record_invoice",
      "description": "Create a draft supplier invoice",
      "parameters": {
        "type": "object",
        "properties": {
          "number": {"type": "string"},
          "amount": {"type": "number"}
        },
        "required": ["number", "amount"]
      }
    }
  }]
}

The response contains choices[0].message.tool_calls, with function.name and function.arguments as a JSON string, and finish_reason is tool_calls. The arguments are written by the model: parse and validate them like any other answer, and let your code decide whether the call runs.


Until when it did not work

From 16 to 18 August 2026 the gateway did not forward response_format to the engine: whoever passed it got free prose, or an object wrapped in a markdown fence, while this page promised the opposite. Fixed on 18 August. If you wrote a downstream validation to work around it, keep it: it is the third of the checks below.

With stream: true the defect lasted two more days: the format was forwarded on the normal path but not on the streaming one, so whoever asked for JSON while streaming still got prose. That was also fixed on 18 August; response_format now works on both paths, and since 6 October on the models served by Ollama too.

In the same round tool_choice: "required" was fixed; it used to answer 502. The cause was not the value: a "properties": {} in the schema reached the engine as [], and the engine rejected the grammar.

What the constraint does not give you

An answer cut short is not an object. With finish_reason: "length" the JSON stops in the middle. In a stream, an answer that ends without [DONE], or with finish_reason: "error", is cut short as well: see Streaming.

A refusal is not data. There is no refusal field in our answers. A model that declines writes it where it can: as prose when the format is not applied, or inside the fields when it is — an apology as a description, an empty required string. Treat values that are not data as no answer.

Numbers must be recomputed. The model reads well, but a total that was read is not a total that was checked: add up the lines and compare. It holds for any model, not only ours, and it is why the result should always land in a draft to be confirmed, never straight into the books.

The schema constrains the shape, not the truth. If the date field is required, it is in the answer; whether it is the right date is for your check to say, not us.

enum inside tools is not enforced. In the response_format schema it is; in the parameters of a function it is not, because there the constraint does not go through the grammar. The model tends to answer with the word in the language of the conversation: "urgente" instead of "urgent". Two things work, and integrators taught them to us: putting the mapping in the property's description ("urgente/urgentissimo=urgent, alta=high") raises adherence a lot, and validating downstream remains necessary.

The model sometimes writes the string "null". On a field that allows null, the four-letter word may arrive instead of the JSON value. We fix it in tool-call arguments (what we change); in your own parsing, treat "null", "none", "nessuno" and the empty string as absence.

On a Swiss invoice, read the QR code first. If the document has the payment part with the code, IBAN, reference, amount and creditor can be read deterministically, without a model and without margin of error. Use the model for what the code does not contain: document number, dates, lines, VAT rates. It is the right order, and it costs you less.

The checks, in order

Each check assumes the one before it has passed.

  1. Transport. The HTTP status is 200 and the answer arrived within your own timeout. In a stream, [DONE] arrived. Keep the X-Request-Id: it is what to quote if you write to us.
  2. Protocol. finish_reason is stop, or tool_calls when you asked for a function; not length, not error. There is no error object, and the answer is not a refusal.
  3. Structure. The content parses as JSON, and the object validates against your schema: required keys, types, enums.
  4. Meaning. The values make sense for the task: totals recomputed from the lines, units and currencies, dates that exist, references that match the source (an IBAN's check digits, an invoice number), line numbers that are in the document.
  5. Action. What happens next is authorised by your code, not by the model: a draft that someone confirms, and a person's approval wherever money, people's data or a production system are involved.

The first three, in Python:

python
import json
import os

from jsonschema import validate  # pip install jsonschema
from openai import OpenAI

client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"], timeout=60)

schema = {
    "type": "object",
    "properties": {
        "document_number": {"type": "string"},
        "date": {"type": "string"},
        "total": {"type": "number"},
        "vat_rate": {"type": "number"},
    },
    "required": ["document_number", "date", "total"],
}

# 1. Transport: errors and timeouts raise an exception here.
resp = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Extract the fields: INVOICE 2026-0417, 14 September 2026, VAT 8.1%, TOTAL CHF 5026.65"}],
    response_format={"type": "json_schema", "json_schema": {"name": "invoice", "schema": schema}},
    max_completion_tokens=300,
)
choice = resp.choices[0]

# 2. Protocol: a cut answer is not an object.
if choice.finish_reason != "stop":
    raise RuntimeError(f"incomplete answer ({choice.finish_reason}), request {resp.id}")

# 3. Structure: parse, then validate against the schema.
invoice = json.loads(choice.message.content)
validate(invoice, schema)

# 4. Meaning and 5. action are yours: compare with the document, then make a draft.
print(invoice)

Search on your documents

To upload documents, ask a question and get the answer with the passage it comes from, see Knowledge bases. If your case is extracting fields from invoices you do not need it — the two features above are enough.

Search the docs

Type to search…