Skip to content

Cookbook

Three examples, with measured numbers

Invoice extraction, document search with citations, transcription of a meeting. Runnable code; latency and cost measured on 22 September 2026, not estimated: two examples through the public API, document search on the pipeline itself.

Last updated: 2026-10-06

Every number on this page was measured on 22 September 2026, three times. Nothing is estimated. Examples 1 and 3 are real calls to the public API. Example 2 is not: document upload was out of service that day, so the document reached the pipeline another way, and its numbers measure the pipeline itself — reading, splitting, vectors, search, rerank, answer — not the calls you would make. Its section says it again.

The code can be copied and run: if your numbers differ from ours, that difference is information, and we want to know about it.

All you need is a key and your base_url:

bash
export API_KEY="sk-…"
export BASE_URL="https://api.aitokens.ch/v1"

1. Extracting the fields of an invoice

A fiduciary's case: a supplier invoice becomes a bookkeeping entry. With json_schema a complete answer is an object with those fields and those types; whether the values are right is for your checks to say, as below.

bash
curl $BASE_URL/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26b",
    "temperature": 0,
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "invoice",
        "schema": {
          "type": "object",
          "properties": {
            "document_number": {"type": "string"},
            "issue_date":      {"type": "string"},
            "due_date":        {"type": "string"},
            "supplier":        {"type": "string"},
            "net_amount":      {"type": "number"},
            "vat_rate":        {"type": "number"},
            "vat_amount":      {"type": "number"},
            "total":           {"type": "number"},
            "iban":            {"type": "string"}
          },
          "required": ["document_number","issue_date","supplier","net_amount","total"]
        }
      }
    },
    "messages": [
      {"role":"system","content":"Extract the fields from the invoice. Amounts as numbers, without apostrophes."},
      {"role":"user","content":"FATTURA N. 2026-0417\nStudio Bernasconi SA, Via Nassa 12, 6900 Lugano\nData: 14 settembre 2026 — Scadenza: 14 ottobre 2026\nImponibile CHF 4'\''650.00\nIVA 8.1% CHF 376.65\nTOTALE CHF 5'\''026.65\nIBAN CH93 0076 2011 6238 5295 7"}
    ]
  }'

What comes back:

json
{
  "document_number": "2026-0417",
  "issue_date": "14 settembre 2026",
  "due_date": "14 ottobre 2026",
  "supplier": "Studio Bernasconi SA",
  "net_amount": 4650.0,
  "vat_rate": 8.1,
  "vat_amount": 376.65,
  "total": 5026.65,
  "iban": "CH93 0076 2011 6238 5295 7"
}

Measured — three runs, identical:

model gemma-4-26b
latency 1.81 s (1.74 / 1.76 / 1.93)
tokens 263 in + 183 out
cost multiply by the model's prices in Pricing; at the reference list price of that day, CHF 0.000189 per invoice, CHF 0.19 per thousand

Fields checked one by one against the document: number, net amount, VAT amount, total and rate, all correct in all three runs.

What we do not guarantee, and should say. The schema constrains the shape, not the truth: total is always there and will be a number; whether it is the total is for your check to say. Add up the lines and compare. And on a Swiss invoice with the payment part, read the QR code first: IBAN, reference, amount and creditor come out deterministically, without a model. Use the model for what the code does not contain.


2. Querying your documents, with the source

⚠️ Not measured through the public API. The numbers of this example were taken on the pipeline itself, with the document provided without the upload, which was out of service that day (see the note at the end of this section).

Correction of 1 October 2026. Until that day this section showed three curl commands to /v1/rag/kb with the API key. Those commands answer 404. We had written them without testing them from outside, which is exactly what this page promises not to do, and we removed them instead of replacing them with others not tested. Today an API key can ask questions to a knowledge base, with POST /v1/knowledge-bases/{slug}/ask, and uploading a document needs the dashboard (https://my.aitokens.ch) or the app API: the RAG quickstart has the calls.

What comes back — on a five-article mandate contract:

Ciascuna parte può disdire il contratto con un preavviso di sei mesi per la fine di un anno civile [contratto-mandato.md: parte 0]. La disdetta deve essere comunicata per iscritto mediante raccomandata [contratto-mandato.md: parte 0].

The citation is not decoration: it is the point. An answer without a source, on a contract, cannot be used by whoever has to answer the client.

Measured — three different questions on the same contract:

indexing 0.20 s (2 chunks)
answer 0.57 s on average (0.47 / 0.51 / 0.73)
sources cited 2 per answer
pipeline bge-m3 vectors, 1024 dim → hybrid search → rerank → gemma-4-26b

The three questions — notice period, deadline for adjusting the fee, yearly amount and instalments — all got the right answer with the article cited.

⚠️ On 22 September 2026 document upload was not in service: the object storage that receives the files had a fault since 12 September. The numbers above measure the real pipeline — reading, splitting, vectors, search, rerank, answer — with the document provided another way. On 3 October 2026 the object storage was replaced; we are re-verifying the upload end to end, and this note will go, with the date. If your case is field extraction (example 1), it does not concern you.


3. Transcribing a meeting

bash
curl $BASE_URL/audio/transcriptions \
  -H "Authorization: Bearer $API_KEY" \
  -F file=@minutes.wav \
  -F model=whisper-1 \
  -F response_format=verbose_json \
  -F "timestamp_granularities[]=word"

Measured on 22.2 seconds of minutes in Italian:

latency 1.48 s on average (0.91 / 1.12 / 2.40)
speed 15× real time
cost see the note below
per-word timings 55 words with start, end and probability

⚠️ On cost, correction of 1 October 2026. This page used to say CHF 0.36 for an hour of audio, which is the configured price. Checking the usage records we found that transcription was not billed at all when these numbers were taken, and that the public price list showed a third value (CHF 0.24 per hour). So this example has no cost of that day. Since 3 October 2026 transcription is charged per second of audio, at the configured price, and the price list reads the same value.

A useful thing we did not expect to have to document: spoken numbers come out as digits.

«…un utile di 47.000 franchi. Secondo punto, la fattura numero 2026-0417 dello studio Bernasconi, da 5.026 franchi e 65, va registrata entro il 14 ottobre.»

Dictated as "quarantasettemila", "cinquemilaventisei franchi e sessantacinque", "quattordici ottobre". In minutes that is exactly what you need, and it means the text can go straight into example 1 to extract its fields.

Per-word timings are for whoever has to highlight the exact point in the audio, cut a recording or show where recognition was uncertain — the probability is per word:

json
{"start": 21.44, "end": 22.02, "word": " ottobre.", "probability": 0.9997}

How these numbers were obtained

Examples 1 and 3: three calls each, from outside, through the load balancer, with a normal key and the medium tier. Latency is the total time of the HTTP request, not the model's time alone: it includes network, authentication, accounting and response. The cost is computed with the real billing formula, which since 3 October 2026 is also the one shown on the pricing page.

Example 2 is the exception: its document did not go through the public upload, which was out of service, and its numbers were taken on the pipeline, not with calls to the public API from outside. Read its times as those of the pipeline, not of a request you make.

What the numbers are, and what they are not: three runs of one invoice, three questions on one contract and one recording show that each example works end to end and how long it takes. They do not measure how often the answers are right on other documents: for that, try it on a sample of your own.

The numbers change with the length of your document, the tier of the key and the load. Measure on your own case: that is why the code is here and not only the table.

Search the docs

Type to search…