Catalogue
Models catalogue
What we can serve you: what is already loaded, what we start on request, and the country where each model runs.
Compare
The chat models we serve now, with the values we serve: context in tokens, prompt and answer together; the longest answer is max_tokens up to 8,192 on every model. Choose what you need:
All countries 🇨🇭 Switzerland Calls tools Reads images Reasons
| Model | Runs in | Context | Tools | Images | Reasons | CHF / 1M in · out |
|---|---|---|---|---|---|---|
gemma-4-26b |
🇨🇭 | 32k | Yes | Yes | No | 0.10 · 0.30 |
qwen3.6-35b |
🇨🇭 | 32k | Yes | Yes | No | 0.10 · 0.70 |
qwen3-coder-30b |
🇨🇭 | 32k | Yes | No | No | 0.08 · 0.40 |
qwen3.8-27b |
🇨🇭 | 256k | Yes | No | Yes, can be off | 0.30 · 1.75 |
qwen2.5:32b |
🇨🇭 | 32k | Yes | No | No | 0.15 · 0.40 |
ministral-3-14b |
🇨🇭 | 32k | Yes | Yes | No | 0.10 · 0.20 |
gpt-oss-120b |
🇨🇭 | 32k | Yes | No | Yes, 3 levels | 0.10 · 0.40 |
«—» means not tried yet. Every value comes from our engines or from a request through the API, not from the publisher's card.
Ready to use
Already loaded: the first request answers without a start-up wait. There are 12. Each one lists the countries where it runs. A request never leaves the zones allowed for your key — groups of machines, each in one country — and so never leaves their countries.
apertus-70b-instruct
Apertus 70B, the open model of the Swiss AI Initiative (EPFL, ETH Zurich and CSCS), trained on text in more than a thousand languages, Swiss German and Romansh among them. Version 2509, 4-bit weights, context of 16,384 tokens.
Precision: 4 bits, w4a16 (Red Hat), version 2509 · Everything about it
CHF 0.40 input · CHF 1.50 output, per million tokens
gemma-4-26b
Gemma 4 26B A4B by Google, a mixture of experts. Reads images as well as text, such as photos of documents, scans and screenshots, and returns structured data well. Context of 32,768 tokens, 8-bit weights.
Precision: FP8, 8 bits (Red Hat) · Everything about it
CHF 0.10 input · CHF 0.30 output, per million tokens
qwen3.6-35b
Qwen3.6 35B A3B by Alibaba, a mixture of experts that uses about 3 billion parameters per token: fast and low-cost. For chat, summaries, translation and extracting fields from text. Reads images. As served it answers without reasoning first. Context of 32,768 tokens, 8-bit weights.
Precision: FP8, 8 bits (Qwen) · Everything about it
CHF 0.10 input · CHF 0.70 output, per million tokens
qwen3-coder-30b
Qwen3-Coder 30B A3B by Alibaba, trained for code: writing, explaining and changing code, and calling tools. A mixture of experts, about 3 billion parameters per token, so it answers fast. Context of 32,768 tokens, 8-bit weights.
Precision: FP8, 8 bits (Qwen) · Everything about it
CHF 0.08 input · CHF 0.40 output, per million tokens
qwen3.8-27b
Qwen3.8 27B by Alibaba, dense. It reasons before it answers, which helps with code, maths, planning and agents that call tools; the reasoning counts as output tokens, and reasoning_effort "none" switches it off. Context of 262,144 tokens. 4-bit MXFP4 weights published by AMD.
Precision: MXFP4, 4 bits (AMD Quark), with an 8-bit speculative head · Everything about it
CHF 0.30 input · CHF 1.75 output, per million tokens
qwen2.5:32b
Qwen 2.5 32B, an earlier generation kept for programs that use it. For new work, Qwen3.6 35B is faster and also reads images; it costs less for input and more for output. 4-bit weights.
Precision: 4 bits (Ollama) · Everything about it
CHF 0.15 input · CHF 0.40 output, per million tokens
ministral-3-14b
Ministral 3 14B by Mistral AI: small and fast, for short answers, classification and extraction. Reads images. Context of 32,768 tokens.
Precision: As released by Mistral AI · Everything about it
CHF 0.10 input · CHF 0.20 output, per million tokens
gpt-oss-120b
gpt-oss-120b, the open-weight model by OpenAI: a mixture of experts with about 5 billion parameters per token. It reasons before it answers, at three levels (reasoning_effort low, medium or high); the reasoning counts as output tokens. Weights as released, MXFP4.
Precision: MXFP4, 4 bits, as released by OpenAI · Everything about it
CHF 0.10 input · CHF 0.40 output, per million tokens
qwen2.5:14b
Qwen 2.5 14B, an earlier generation kept for programs that use it. For new work, Qwen3.6 35B or Ministral 3 14B. 4-bit weights.
Precision: 4 bits (Ollama) · Everything about it
CHF 0.10 input · CHF 0.25 output, per million tokens
qwen2.5:1.5b
Qwen 2.5 1.5B, a very small model for simple, repetitive tasks where the time of the answer matters most. 4-bit weights.
Precision: 4 bits (Ollama) · Everything about it
CHF 0.02 input · CHF 0.05 output, per million tokens
bge-m3
BGE-M3 embeddings: turns a text into a vector of 1,024 numbers to search by meaning rather than by words, in more than 100 languages, also mixed. Input tokens only.
CHF 0.02 per million input tokens
bge-reranker-v2-m3
BGE reranker v2 m3: reads the results of a search together with the question and puts the relevant ones first. Not charged.
Not charged
If you need a model that is not listed
Ask us. Any open-weight model that fits the capacity we have can be added: it is a mechanical operation, not a project. The list above is what we serve today, not the limit of what we can serve.
What we do not do is resell the closed models of OpenAI or Anthropic: we could not run them ourselves, so we could not make any of the commitments we make on everything else.
Write to info@daikolab.ch.
Audio
Transcription with whisper-1 and whisper-1-hd; speech with the voices
paola and riccardo in Italian, amy and ryan in
English, thorsten in German, siwis in French. OpenAI voice names are
accepted as synonyms. Details in Audio API.