Skip to content

Concepts

Models

How we choose what to serve, what open weights mean here, and which model fits which job.

Last updated: 2026-10-08

AI Tokens serves open-weight models on its own hardware. No closed third-party API in the middle and no opaque costs: you know which model answers, and you can check where it comes from and under which licence.

The catalogue

The list of the models you can use lives in the Models catalogue, read straight from the database. GET /v1/models lists every model AI Tokens serves, built at each call: the chat models of the catalogue, embeddings and rerank, pictures and video while a machine makes them, transcription and speech. This page does not repeat the list: a copy would go out of date the day a model is added. For every model the catalogue shows:

  • the countries where it runs. A request never leaves the zones allowed for your key, and so never leaves their countries: see Zones below;
  • whether it is ready to use or started on request. A model started on request answers 503 model_activation_required while it is not loaded, and the message names the models ready now: see Errors;
  • the images badge, on the models that accept images as input;
  • its precision, when it is served quantized;
  • a description of what it is good at, and its price in Pricing.

Some models are outside the catalogue:

  • Embeddings and rerank always use bge-m3 and bge-reranker-v2-m3, through /v1/embeddings and /v1/rerank. /v1/models lists both; on some brands the catalogue does not show them.
  • Audio: whisper-1 and whisper-1-hd for transcription, tts-1 for speech, documented in Audio API; /v1/models lists them. The voices are the voice of tts-1, listed at GET /v1/audio/voices.
  • Pictures and video, where AI Tokens has machines that make them, are in /v1/models while a machine makes them, and in their own pages of the API reference.

Where the models come from

The families you may find in the catalogue, and who publishes them:

Family Published by
Apertus 🇨🇭 Swiss AI Initiative — EPFL, ETH Zurich, CSCS
Gemma 🇺🇸 Google DeepMind
gpt-oss 🇺🇸 OpenAI
Ministral, Mistral 🇫🇷 Mistral AI
Qwen 🇨🇳 Alibaba
BGE (embeddings and rerank) 🇨🇳 BAAI

Which of them AI Tokens serves, and in which version and size, is in the catalogue.

Zones

A zone is a group of machines in one country, with a declared owner: the platform's own machines, or a partner's. Every model in the catalogue lists the countries of its zones, and every API key has the zones it may use — by default all the zones offered by AI Tokens.

  • A request is sent only to machines in the zones allowed for the key, so it never leaves the countries of those zones.
  • If no machine is available there, the answer is 503 with the zone named in the message. The request is never moved to another zone to get an answer.
  • Text answers — chat completions, completions and the Responses API — say where they were processed, in the X-Zona header, and your usage records keep the zone and the site of every request. In a stream the headers leave before a machine is chosen: if your key allows more than one zone, the header lists the possible ones, and the usage record has the exact one. Transcriptions made in a zone carry X-Zona too.

Zones apply today to text generation — /v1/chat/completions, /v1/completions and /v1/responses — to pictures and video, where AI Tokens makes them, and to whisper-1 transcriptions, where AI Tokens has a transcription machine in your zones. The answers of knowledge bases are written in the zones offered by AI Tokens, not only in those of your key. Embeddings and rerank are served by dedicated machines that do not go through zones yet; speech synthesis, and the transcriptions a machine in your zones does not take, run in Switzerland, outside the zones: see Audio. If it matters for your case, write to info@daikolab.ch and we tell you where they run for AI Tokens.

Images as input

The models with the images badge in the catalogue accept images in the messages of a chat completion. How many per request depends on the model: above its limit, the 400 states it. Sending an image to a model without the badge returns a 400 that lists the models that accept images — not a generic backend error.

Formats: PNG, JPEG, WebP, GIF. At most 8 MB and 40 megapixels per image. The request format is in Chat completions.

How we choose what to serve

Three criteria:

  1. Licence: usable commercially, without surprises.
  2. Quality: at the top of its size class when it is added.
  3. Sustainability: it must run efficiently on the capacity we have. We do not add a model if it makes the service worse for all the others.

We exclude models that contact the outside world during inference, models whose access is bound in ways we cannot verify, and models whose ecosystem collects user data.

Quantization

Large models may be served quantized: a more compact numerical representation that reduces the memory needed and increases serving capacity, at the cost of a small, measured loss of quality. The precision of each model is stated in the catalogue.

It is invisible to whoever calls the API: the model_id stays the same however the model is loaded, and if we change the quantization your code does not change.

Which one to pick

The catalogue changes more often than these pages, so the advice here is about what to look for in it, not about names.

If you need to Look for
Chat, general case qwen3.8-27b, the default of AI Tokens
Read an image, a scan or a screenshot gemma-4-26b, or another model with the images badge
Write or review code a model whose description says it is made for programming
Sums, invoices, contracts: several steps of reasoning a model whose description says it reasons before answering. Recompute the numbers anyway: see Structured output
Answers in a given language the languages named in the description of each model
An answer at any time a model ready to use, not one started on request
Search your texts by meaning bge-m3 via /v1/embeddings
Improve the order of search results bge-reranker-v2-m3 via /v1/rerank

Then try two or three of them on a sample of your own work, and keep the one that gets it right at the lower price: a description is a starting point, not a measurement. Every request names a model_id; which machine runs it, inside your zones, is up to us.

Tiers and models

The tier does not decide which model you can use: it decides which machines may serve you, the rate limit of your key and a multiplier on the price, which depends on the model. It does not change the order in a machine's queue. Details in Tiers and in the live price list.

If you need a model we do not have

Write to info@daikolab.ch. Adding an open-weight model is a mechanical operation, and we do it on request for business customers. A model trained on your own terminology is a project of its own: tell us about it.

Search the docs

Type to search…