Integrations
Integrations
What any OpenAI-compatible tool needs to work with AI Tokens: the base URL, a key and a model id. Which model to pick, the context window to set, and what the API does not do.
Last updated: 2026-10-11
A tool that speaks OpenAI's chat completions API works with AI Tokens. It needs three values:
| Setting | Value |
|---|---|
| Base URL | https://api.aitokens.ch/v1 |
| API key | a key from the dashboard, starting with sk-at- |
| Model | a model id from the catalogue, written as it is there: qwen3.8-27b, not Qwen3.8 27B |
Tools give the base URL different names: Base URL, API Base, api_url,
baseURL, OPENAI_API_BASE. The value is the same everywhere, and it ends in
/v1, without /chat/completions.
Where a tool asks for a provider type, choose OpenAI Compatible. Where it offers only OpenAI, choose it and change the base URL.
Tested, or from the documentation
Every page says which of the two it is.
- Tested: we ran the example on the live API on the date given, with the version of the tool named on the page, and the page shows what came back.
- From the documentation: the settings come from the tool's own documentation, read on the date given. We did not run the tool. If a setting does not work as written, tell us at info@daikolab.ch.
| Tool | What it is | Status |
|---|---|---|
| Cline | coding agent for VS Code | from the documentation |
| Continue | coding assistant for VS Code and JetBrains | from the documentation |
| Roo Code | coding agent for VS Code | from the documentation |
| Kilo Code | coding agent for VS Code and the terminal | from the documentation |
| OpenCode | coding agent for the terminal | from the documentation |
| Aider | coding assistant for the terminal | tested |
| Zed | code editor with an agent | from the documentation |
| Open WebUI | chat interface you host yourself | from the documentation |
| n8n | workflow automation | from the documentation |
| LangChain | Python framework | tested |
| LlamaIndex | Python framework | tested |
| LiteLLM | Python SDK and proxy | SDK tested, proxy from the documentation |
| Vercel AI SDK | TypeScript framework | tested |
| OpenAI Agents SDK | Python framework for agents | tested |
Which model
The catalogue lists the models AI Tokens serves today, with their prices and the country where each one runs. This table adds what a tool needs to be told, as we served it on 11 October 2026.
| Model | Context window | Tools | Images | Reasoning | Pick it for |
|---|---|---|---|---|---|
qwen3.8-27b |
262,144 | yes | no | before every answer; reasoning_effort: "none" switches it off |
coding agents, long files, long conversations |
qwen3-coder-30b |
32,768 | yes | no | no | quick code edits, questions about code |
qwen3.6-35b |
32,768 | yes | yes | no | chat, summaries, translation, extraction at a low price |
gpt-oss-120b |
32,768 | yes | no | low, medium or high; it cannot be switched off |
problems that need reasoning, with short inputs |
gemma-4-26b |
32,768 | yes | yes | no | photos of documents, scans, structured data |
ministral-3-14b |
32,768 | yes | yes | no | short answers, classification, the background tasks of chat apps |
apertus-70b-instruct |
16,384 | no | no | no | text in many languages, Swiss German and Romansh among them |
Tools is what we checked on 11 October 2026: one request with a tool on
each model, with the OpenAI SDK, and the same with stream: true on the first
four rows. Each model returned the call in OpenAI's shape. A coding agent needs
a model with tools.
Context window. Set it in the tool to the number in this table. A tool that does not know it assumes its own default, and may send a conversation longer than the model can read.
Maximum output. Set it to 8,192 tokens or less. The API refuses a larger
max_tokens with 422.
Reasoning
qwen3.8-27b and gpt-oss-120b reason before they answer. The reasoning is
not returned: tools that show a model's thinking show nothing for these two.
It is counted in usage.completion_tokens, billed as output, and it counts
against max_tokens. With a low cap the answer comes back cut or empty, with
finish_reason: "length".
On qwen3.8-27b, "reasoning_effort": "none" switches reasoning off for a
request. Asked for one sentence about the Gotthard tunnel on 11 October 2026,
it used 176 output tokens with reasoning and 47 without. Each tool page shows
how to send the field, where the tool allows it. More in
Every parameter.
auto
With "model": "auto", the API chooses the model for each
request: today qwen3.6-35b for routine work, and qwen3.8-27b for work that
needs reasoning or a long input. It suits chat interfaces and workflows, such
as Open WebUI and n8n. Where a tool asks for the context window of auto,
give it 32,768, the smaller of the two.
For coding agents, name a model. auto keeps a task on the model it started
with, so that the agent does not change model halfway through. A task that
started on qwen3.6-35b stays within its 32,768 tokens, even when the files
the agent reads grow past them.
What the API does not do
- Anthropic's Messages API.
/v1/messagesanswers404withunknown_endpoint. Tools that speak only Anthropic's format, such as Claude Code, cannot connect directly. - Fill-in-the-middle.
/v1/completionspasses the prompt to the model as a chat message, and refusessuffix. Tab completion built on fill-in-the-middle does not work here: Continue's autocomplete, Zed's edit predictions. Chat, edits and agents do. - Tools and streaming on the Responses API.
/v1/responsesanswers501tostream: true. On 11 October 2026 it also did not read a tool result sent back asfunction_call_output: the model asked for the same tool again. Tools that can use either API have to use chat completions; the pages of n8n and the OpenAI Agents SDK show the setting. - Parameters it does not know. A field that is not in
Every parameter gets a
400that names it. Some tools send fields of other engines when you fill in their settings, such astop_k,min_p,repetition_penaltyorprompt_cache_key. Leave those settings empty. - Embeddings from other models.
/v1/embeddingsalways answers withbge-m3: 1,024 dimensions, at most 32 texts per request. See Embeddings.
/v1/models lists every model AI Tokens serves, and some of them do not chat:
bge-m3, bge-reranker-v2-m3, whisper-1, whisper-1-hd, tts-1, and the
picture and video models where there are any. When a tool fills its model menu
from that list, choose a chat model.
Headers
Where a tool lets you add HTTP headers, X-Siati-Tier sets the
tier of its requests. The zone where a request ran comes
back in the X-Zona header of the answer.