Skip to content

Integrations

Open WebUI

Open WebUI with AI Tokens: one OpenAI connection, the chat models only, a small model for the background tasks, and tools in Native mode.

Last updated: 2026-10-11

Connection

Settings → Admin → Connections, under Manage OpenAI API Connections, ➕ Add Connection:

Field Value
URL https://api.aitokens.ch/v1
API Key your key, sk-at-…
Prefix ID optional, such as aitokens, to tell these models from others
Model IDs the chat models you want to offer, such as qwen3.8-27b, qwen3.6-35b, gemma-4-26b

Verify Connection reads https://api.aitokens.ch/v1/models. That list also holds models that do not chat (bge-m3, whisper-1, tts-1 and others), and without the Model IDs list they would all appear in the chat's menu.

On a server, the same connection can come from the environment:

bash
OPENAI_API_BASE_URL=https://api.aitokens.ch/v1
OPENAI_API_KEY=sk-at-…

Background tasks

Open WebUI writes a title for each chat, tags and suggested follow-ups with extra requests to a model. By default that is the model of the chat, so with qwen3.8-27b each of them reasons and is billed as output. In Settings → Admin → Interface, set Task Model (External) to a small model, such as ministral-3-14b, or switch off the tasks you do not need.

Tools

Open WebUI calls tools in Native mode, its default since version 0.10.0. Native mode needs a model with tools: qwen3.8-27b, qwen3.6-35b, qwen3-coder-30b, gpt-oss-120b, gemma-4-26b and ministral-3-14b returned tool calls on 11 October 2026; apertus-70b-instruct does not take tools.

Images

For models that read images (qwen3.6-35b, gemma-4-26b, ministral-3-14b), switch on Vision in the model's settings under Settings → Admin → Models. Leave it off on the others.

auto

Add auto to the Model IDs and the API picks the model for each message: the low-cost one for routine questions, the reasoning one when a question needs it. The answer's model says which one answered.

Limits

  • Advanced parameters. A parameter the API does not know, such as top_k or min_p, makes the request fail with 400 and the parameter's name. Leave those at their default.
  • Max Tokens up to 8192. A higher value is refused with 422.
  • No thinking panel for qwen3.8-27b: the API does not return its reasoning.

Search the docs

Type to search…