Integrations
LiteLLM
LiteLLM with AI Tokens: the openai/ prefix and api_base in the Python SDK, and a model_list for the proxy.
Last updated: 2026-10-11
Python SDK
pip install litellm
export API_KEY="sk-at-…"
import os
import litellm
resp = litellm.completion(
model="openai/qwen3.8-27b",
api_base="https://api.aitokens.ch/v1",
api_key=os.environ["API_KEY"],
messages=[{"role": "user", "content": "In one sentence: why is Ticino Italian-speaking?"}],
max_tokens=1024,
)
print(resp.choices[0].message.content)
stream = litellm.completion(
model="openai/qwen3.8-27b",
api_base="https://api.aitokens.ch/v1",
api_key=os.environ["API_KEY"],
messages=[{"role": "user", "content": "Count from 1 to 5."}],
max_tokens=1024,
stream=True,
)
print("".join(chunk.choices[0].delta.content or "" for chunk in stream))
What came back on qwen3.8-27b, shortened:
Ticino was historically part of the Duchy of Milan … before being conquered by the Swiss cantons …
Usage(completion_tokens=220, prompt_tokens=21, total_tokens=241, …)
1, 2, 3, 4, 5.
The 220 output tokens of a one-sentence answer are mostly reasoning: see Reasoning.
The openai/ prefix tells LiteLLM to call an OpenAI-compatible API at
api_base; the rest is the model id of the catalogue. The
api_base ends in /v1.
Proxy
In the proxy's config.yaml, one entry per model:
model_list:
- model_name: qwen3.8-27b
litellm_params:
model: openai/qwen3.8-27b
api_base: https://api.aitokens.ch/v1
api_key: os.environ/API_KEY
- model_name: qwen3-coder-30b
litellm_params:
model: openai/qwen3-coder-30b
api_base: https://api.aitokens.ch/v1
api_key: os.environ/API_KEY
os.environ/API_KEY makes the proxy read the key from the environment
variable API_KEY when it starts.
Limits
max_tokensup to 8192. A higher value is refused with422.- Parameters the API does not know are refused with
400, and the error names them: see Every parameter. - Reasoning tokens are counted in
completion_tokens, with no separate count.