Skip to content

Integrations

LiteLLM

LiteLLM with AI Tokens: the openai/ prefix and api_base in the Python SDK, and a model_list for the proxy.

Last updated: 2026-10-11

Python SDK

bash
pip install litellm
export API_KEY="sk-at-…"
python
import os

import litellm

resp = litellm.completion(
    model="openai/qwen3.8-27b",
    api_base="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    messages=[{"role": "user", "content": "In one sentence: why is Ticino Italian-speaking?"}],
    max_tokens=1024,
)
print(resp.choices[0].message.content)

stream = litellm.completion(
    model="openai/qwen3.8-27b",
    api_base="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    messages=[{"role": "user", "content": "Count from 1 to 5."}],
    max_tokens=1024,
    stream=True,
)
print("".join(chunk.choices[0].delta.content or "" for chunk in stream))

What came back on qwen3.8-27b, shortened:

text
Ticino was historically part of the Duchy of Milan … before being conquered by the Swiss cantons …
Usage(completion_tokens=220, prompt_tokens=21, total_tokens=241, …)
1, 2, 3, 4, 5.

The 220 output tokens of a one-sentence answer are mostly reasoning: see Reasoning.

The openai/ prefix tells LiteLLM to call an OpenAI-compatible API at api_base; the rest is the model id of the catalogue. The api_base ends in /v1.

Proxy

In the proxy's config.yaml, one entry per model:

yaml
model_list:
  - model_name: qwen3.8-27b
    litellm_params:
      model: openai/qwen3.8-27b
      api_base: https://api.aitokens.ch/v1
      api_key: os.environ/API_KEY
  - model_name: qwen3-coder-30b
    litellm_params:
      model: openai/qwen3-coder-30b
      api_base: https://api.aitokens.ch/v1
      api_key: os.environ/API_KEY

os.environ/API_KEY makes the proxy read the key from the environment variable API_KEY when it starts.

Limits

  • max_tokens up to 8192. A higher value is refused with 422.
  • Parameters the API does not know are refused with 400, and the error names them: see Every parameter.
  • Reasoning tokens are counted in completion_tokens, with no separate count.

Search the docs

Type to search…