Integrations
Open WebUI
Open WebUI with AI Tokens: one OpenAI connection, the chat models only, a small model for the background tasks, and tools in Native mode.
Last updated: 2026-10-11
Connection
Settings → Admin → Connections, under Manage OpenAI API Connections, ➕ Add Connection:
| Field | Value |
|---|---|
| URL | https://api.aitokens.ch/v1 |
| API Key | your key, sk-at-… |
| Prefix ID | optional, such as aitokens, to tell these models from others |
| Model IDs | the chat models you want to offer, such as qwen3.8-27b, qwen3.6-35b, gemma-4-26b |
Verify Connection reads https://api.aitokens.ch/v1/models. That list also holds
models that do not chat (bge-m3, whisper-1, tts-1 and others), and
without the Model IDs list they would all appear in the chat's menu.
On a server, the same connection can come from the environment:
OPENAI_API_BASE_URL=https://api.aitokens.ch/v1
OPENAI_API_KEY=sk-at-…
Background tasks
Open WebUI writes a title for each chat, tags and suggested follow-ups with
extra requests to a model. By default that is the model of the chat, so with
qwen3.8-27b each of them reasons and is billed as output. In
Settings → Admin → Interface, set Task Model (External) to a small
model, such as ministral-3-14b, or switch off the tasks you do not need.
Tools
Open WebUI calls tools in Native mode, its default since version 0.10.0.
Native mode needs a model with tools: qwen3.8-27b, qwen3.6-35b,
qwen3-coder-30b, gpt-oss-120b, gemma-4-26b and ministral-3-14b
returned tool calls on 11 October 2026; apertus-70b-instruct does not take
tools.
Images
For models that read images (qwen3.6-35b, gemma-4-26b,
ministral-3-14b), switch on Vision in the model's settings under
Settings → Admin → Models. Leave it off on the others.
auto
Add auto to the Model IDs and the API picks the model for
each message: the low-cost one for routine questions, the reasoning one when a
question needs it. The answer's model says which one answered.
Limits
- Advanced parameters. A parameter the API does not know, such as
top_kormin_p, makes the request fail with400and the parameter's name. Leave those at their default. - Max Tokens up to 8192. A higher value is refused with
422. - No thinking panel for
qwen3.8-27b: the API does not return its reasoning.