Skip to content

Models

Qwen3.8 27B

qwen3.8-27b 🇨🇭 Switzerland

Qwen3.8 27B by Alibaba, dense. It reasons before it answers, which helps with code, maths, planning and agents that call tools; the reasoning counts as output tokens, and reasoning_effort "none" switches it off. Context of 262,144 tokens. 4-bit MXFP4 weights published by AMD.

At a glance

PublisherAlibaba (Qwen)
ArchitectureDense, 27 billion parameters
LicenceApache 2.0
PrecisionMXFP4, 4 bits (AMD Quark), with an 8-bit speculative head
Context262,144 tokens, prompt and answer together
Longest answermax_tokens up to 8,192, reasoning included
ReadsText
Calls toolsYes
JSON schemaYes
ReasonsYes, before every answer. Switch it off with reasoning_effort: "none"; the reasoning counts as output tokens.
Runs in🇨🇭 Switzerland
StatusLoaded: the first request answers without a start-up wait. Live state on Status.
PriceCHF 0.30 input · CHF 1.75 output, per million tokens, excluding VAT. All prices
In autoThe reasoning model of the automatic model selection.

Use it

curl https://api.aitokens.ch/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "reasoning_effort": "none",
    "messages": [{"role": "user", "content": "Summarise in one line: the meeting moves to Thursday at 10."}]
  }'
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])
answer = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Summarise in one line: the meeting moves to Thursday at 10."}],
)
print(answer.choices[0].message.content)

More in Quickstart and Chat completions.

Limits

  • Prompt and answer together stay within 262,144 tokens; beyond that the engine refuses the request with a 400 that says so.
  • The reasoning counts against max_tokens: a low cap can end the answer before it starts, with finish_reason: "length".
  • Text only: a request with an image gets a 400.
  • Served at 4 bits, which makes it faster and fits our machines; we write the precision of every model.
  • A request goes only to the zones your key allows; if the model has no machine there now, a 503 names the zones. It is never sent elsewhere.

Search the docs

Type to search…