Qwen 2.5 14B
qwen2.5:14b
🇨🇭 Switzerland
Qwen 2.5 14B, an earlier generation kept for programs that use it. For new work, Qwen3.6 35B or Ministral 3 14B. 4-bit weights.
At a glance
| Publisher | Alibaba (Qwen) |
|---|---|
| Architecture | Dense, 14 billion parameters |
| Licence | Apache 2.0 |
| Precision | 4 bits (Ollama) |
| Context | 32,768 tokens, prompt and answer together |
| Longest answer | max_tokens up to 8,192 |
| Reads | Text |
| Calls tools | Not tried yet |
| JSON schema | Not tried yet |
| Reasons | No, as served: it answers straight away. |
| Runs in | 🇨🇭 Switzerland |
| Status | Loaded: the first request answers without a start-up wait. Live state on Status. |
| Price | CHF 0.10 input · CHF 0.25 output, per million tokens, excluding VAT. All prices |
Use it
curl https://api.aitokens.ch/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen2.5:14b",
"messages": [{"role": "user", "content": "Summarise in one line: the meeting moves to Thursday at 10."}]
}'
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])
answer = client.chat.completions.create(
model="qwen2.5:14b",
messages=[{"role": "user", "content": "Summarise in one line: the meeting moves to Thursday at 10."}],
)
print(answer.choices[0].message.content)
More in Quickstart and Chat completions.
Limits
- Prompt and answer together stay within 32,768 tokens; beyond that the engine refuses the request with a
400that says so. - Text only: a request with an image gets a
400. - Served at 4 bits, which makes it faster and fits our machines; we write the precision of every model.
- A request goes only to the zones your key allows; if the model has no machine there now, a
503names the zones. It is never sent elsewhere.