Skip to content

Catalogue

Models catalogue

What we can serve you: what is already loaded, what we start on request, and the country where each model runs.

Compare

The chat models we serve now, with the values we serve: context in tokens, prompt and answer together; the longest answer is max_tokens up to 8,192 on every model. Choose what you need:

All countries 🇨🇭 Switzerland Calls tools Reads images Reasons

Model Runs in Context Tools Images Reasons CHF / 1M in · out
gemma-4-26b 🇨🇭 32k Yes Yes No 0.10 · 0.30
qwen3.6-35b 🇨🇭 32k Yes Yes No 0.10 · 0.70
ministral-3-14b 🇨🇭 32k Yes Yes No 0.10 · 0.20

«—» means not tried yet. Every value comes from our engines or from a request through the API, not from the publisher's card.

Ready to use

Already loaded: the first request answers without a start-up wait. There are 12. Each one lists the countries where it runs. A request never leaves the zones allowed for your key — groups of machines, each in one country — and so never leaves their countries.

Apertus 70B

apertus-70b-instruct

Switzerland instant

Apertus 70B, the open model of the Swiss AI Initiative (EPFL, ETH Zurich and CSCS), trained on text in more than a thousand languages, Swiss German and Romansh among them. Version 2509, 4-bit weights, context of 16,384 tokens.

Precision: 4 bits, w4a16 (Red Hat), version 2509 · Everything about it

CHF 0.40 input · CHF 1.50 output, per million tokens

Gemma 4 26B

gemma-4-26b

Switzerland images instant

Gemma 4 26B A4B by Google, a mixture of experts. Reads images as well as text, such as photos of documents, scans and screenshots, and returns structured data well. Context of 32,768 tokens, 8-bit weights.

Precision: FP8, 8 bits (Red Hat) · Everything about it

CHF 0.10 input · CHF 0.30 output, per million tokens

Qwen3.6 35B

qwen3.6-35b

Switzerland images instant

Qwen3.6 35B A3B by Alibaba, a mixture of experts that uses about 3 billion parameters per token: fast and low-cost. For chat, summaries, translation and extracting fields from text. Reads images. As served it answers without reasoning first. Context of 32,768 tokens, 8-bit weights.

Precision: FP8, 8 bits (Qwen) · Everything about it

CHF 0.10 input · CHF 0.70 output, per million tokens

Qwen3-Coder 30B

qwen3-coder-30b

Switzerland instant

Qwen3-Coder 30B A3B by Alibaba, trained for code: writing, explaining and changing code, and calling tools. A mixture of experts, about 3 billion parameters per token, so it answers fast. Context of 32,768 tokens, 8-bit weights.

Precision: FP8, 8 bits (Qwen) · Everything about it

CHF 0.08 input · CHF 0.40 output, per million tokens

Qwen3.8 27B

qwen3.8-27b

Switzerland instant

Qwen3.8 27B by Alibaba, dense. It reasons before it answers, which helps with code, maths, planning and agents that call tools; the reasoning counts as output tokens, and reasoning_effort "none" switches it off. Context of 262,144 tokens. 4-bit MXFP4 weights published by AMD.

Precision: MXFP4, 4 bits (AMD Quark), with an 8-bit speculative head · Everything about it

CHF 0.30 input · CHF 1.75 output, per million tokens

Qwen 2.5 32B

qwen2.5:32b

Switzerland instant

Qwen 2.5 32B, an earlier generation kept for programs that use it. For new work, Qwen3.6 35B is faster and also reads images; it costs less for input and more for output. 4-bit weights.

Precision: 4 bits (Ollama) · Everything about it

CHF 0.15 input · CHF 0.40 output, per million tokens

Ministral 3 14B

ministral-3-14b

Switzerland images instant

Ministral 3 14B by Mistral AI: small and fast, for short answers, classification and extraction. Reads images. Context of 32,768 tokens.

Precision: As released by Mistral AI · Everything about it

CHF 0.10 input · CHF 0.20 output, per million tokens

gpt-oss-120b

gpt-oss-120b

Switzerland instant

gpt-oss-120b, the open-weight model by OpenAI: a mixture of experts with about 5 billion parameters per token. It reasons before it answers, at three levels (reasoning_effort low, medium or high); the reasoning counts as output tokens. Weights as released, MXFP4.

Precision: MXFP4, 4 bits, as released by OpenAI · Everything about it

CHF 0.10 input · CHF 0.40 output, per million tokens

Qwen 2.5 14B

qwen2.5:14b

Switzerland instant

Qwen 2.5 14B, an earlier generation kept for programs that use it. For new work, Qwen3.6 35B or Ministral 3 14B. 4-bit weights.

Precision: 4 bits (Ollama) · Everything about it

CHF 0.10 input · CHF 0.25 output, per million tokens

Qwen 2.5 1.5B

qwen2.5:1.5b

Switzerland instant

Qwen 2.5 1.5B, a very small model for simple, repetitive tasks where the time of the answer matters most. 4-bit weights.

Precision: 4 bits (Ollama) · Everything about it

CHF 0.02 input · CHF 0.05 output, per million tokens

Switzerland instant

BGE-M3 embeddings: turns a text into a vector of 1,024 numbers to search by meaning rather than by words, in more than 100 languages, also mixed. Input tokens only.

CHF 0.02 per million input tokens

BGE Reranker v2 m3

bge-reranker-v2-m3

Switzerland instant

BGE reranker v2 m3: reads the results of a search together with the question and puts the relevant ones first. Not charged.

Not charged

If you need a model that is not listed

Ask us. Any open-weight model that fits the capacity we have can be added: it is a mechanical operation, not a project. The list above is what we serve today, not the limit of what we can serve.

What we do not do is resell the closed models of OpenAI or Anthropic: we could not run them ourselves, so we could not make any of the commitments we make on everything else.

Write to info@daikolab.ch.

Audio

Transcription with whisper-1 and whisper-1-hd; speech with the voices paola and riccardo in Italian, amy and ryan in English, thorsten in German, siwis in French. OpenAI voice names are accepted as synonyms. Details in Audio API.

Search the docs

Type to search…