Skip to content

Models

Gemma 4 26B

gemma-4-26b 🇨🇭 Switzerland

Gemma 4 26B A4B by Google, a mixture of experts. Reads images as well as text, such as photos of documents, scans and screenshots, and returns structured data well. Context of 32,768 tokens, 8-bit weights.

At a glance

PublisherGoogle
ArchitectureMixture of experts, 26 billion parameters, about 4 billion per token
LicenceApache 2.0
PrecisionFP8, 8 bits (Red Hat)
Context32,768 tokens, prompt and answer together
Longest answermax_tokens up to 8,192
ReadsText and images
Calls toolsYes
JSON schemaYes
ReasonsNo, as served: it answers straight away.
Runs in🇨🇭 Switzerland
StatusLoaded: the first request answers without a start-up wait. Live state on Status.
PriceCHF 0.10 input · CHF 0.30 output, per million tokens, excluding VAT. All prices

Use it

curl https://api.aitokens.ch/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26b",
    "messages": [{"role": "user", "content": "Summarise in one line: the meeting moves to Thursday at 10."}]
  }'
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])
answer = client.chat.completions.create(
    model="gemma-4-26b",
    messages=[{"role": "user", "content": "Summarise in one line: the meeting moves to Thursday at 10."}],
)
print(answer.choices[0].message.content)

More in Quickstart and Chat completions.

Limits

  • Prompt and answer together stay within 32,768 tokens; beyond that the engine refuses the request with a 400 that says so.
  • A request goes only to the zones your key allows; if the model has no machine there now, a 503 names the zones. It is never sent elsewhere.

Search the docs

Type to search…