Skip to content

Cookbook

OpenAI SDK migration

Change the base URL, the key and the model name. What else is different, and what does not exist here.

Last updated: 2026-10-08

If your code uses the OpenAI SDK, three things change: the base URL, the key and the model name, which must be a model of the catalogue. The rest of your code works if it uses only what is supported here: the differences are on this page, and every endpoint, with what it honours, refuses or ignores, is in What works where.

OpenAI AI Tokens
Base URL https://api.openai.com/v1 https://api.aitokens.ch/v1
Key sk-… from OpenAI sk-… from the dashboard
Model gpt-4o, … an open-weight model from the catalogue

Chat model names are not translated. A request for gpt-4o answers 404 with model_not_found: it is never sent to another model. Embeddings are the exception, see below.

Python

python
import os

from openai import OpenAI

# Before:
# client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

client = OpenAI(
    base_url="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
)

resp = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)

JavaScript / TypeScript

typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.aitokens.ch/v1",
  apiKey: process.env.API_KEY,
});

const resp = await client.chat.completions.create({
  model: "qwen3.8-27b",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);

Which SDK calls work

SDK call Endpoint Notes
chat.completions.create /v1/chat/completions streaming, tools, response_format, images on the models that accept them
completions.create /v1/completions one prompt per request
responses.create /v1/responses no streaming, no stored conversations: see Responses API
embeddings.create /v1/embeddings bge-m3 only: see Embeddings below
audio.transcriptions.create /v1/audio/transcriptions whisper-1 or whisper-1-hd
audio.speech.create /v1/audio/speech our voices, wav or mp3: see Audio below
images.generate /v1/images/generations only where AI Tokens has machines that make pictures: see Pictures and video below
videos.create, videos.retrieve, videos.list, videos.delete, videos.download_content /v1/videos only where AI Tokens has machines that make video: see Pictures and video below
models.list /v1/models every model served: chat, embeddings, rerank, pictures and video while a machine makes them, transcription, tts-1; the voices are at /v1/audio/voices

/v1/rerank exists too, but the OpenAI SDK has no method for it: call it over HTTP, see Rerank. What each endpoint honours, refuses or accepts without effect is in What works where.

Everything else does not exist here: Assistants and threads, Batch, Files, fine-tuning, image edits and variations, video remix, edits, extensions and characters, audio translations, moderation, Realtime, vector stores. Those paths answer 404 with code: unknown_endpoint. For answers over your own documents there are knowledge bases: with the API key you list them and ask them questions, while creating them and uploading documents need a signed-in user. Or build your own on /v1/embeddings and /v1/rerank.

Models

If your code says Use Notes
a chat model: gpt-4o, gpt-4o-mini, gpt-4.1, … qwen3.8-27b, or another chat model which one for which job: Models
a reasoning model: o3, o4-mini, gpt-5, … gpt-oss-120b or qwen3.8-27b, where the catalogue has them reasoning_effort is translated for each model (Reasoning); the reasoning counts in completion_tokens
a chat model, with images in the messages a model marked images in the catalogue Images as input
text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002 bge-m3 1024 dimensions
whisper-1, gpt-4o-transcribe whisper-1 or whisper-1-hd other names are refused
tts-1, tts-1-hd the same the name is not checked: one speech engine serves every request
dall-e-2, dall-e-3, gpt-image-1 a picture model of AI Tokens, where there is one not translated: 400 model_not_found, and the message lists the picture models served
sora-2, sora-2-pro a video model of AI Tokens, where there is one not translated: 400 model_not_found, and the message lists the video models served

Chat completions: what is different

  • Unknown parameters are refused. A parameter we do not recognise gets a 400 that names it, typos included. These are refused on purpose: store, metadata, functions and function_call (use tools and tool_choice), modalities, audio, prediction, web_search_options, chat_template_kwargs (use reasoning_effort). The full list is in Every parameter.
  • Some parameters work only on some models. n, logprobs, top_logprobs, logit_bias and parallel_tool_calls are refused with a 400 on the models that cannot apply them. Some libraries send n: 1 or logprobs: false by default: if that 400 appears, drop them from the request.
  • The token cap is max_tokens or max_completion_tokens, up to 8192. If you send both with different values, the answer is 400.
  • Streaming sends the same chunks and ends with data: [DONE]. With stream_options: {"include_usage": true} a last chunk carries the usage, as with OpenAI. logprobs/top_logprobs are accepted in a stream and do not work there; n above 1 is refused with 400: see With stream: true.
  • Errors. An unknown model gives 404 model_not_found. No machine available in your zones gives 503 with type: zone_unavailable and the zone in the message: the request is not sent to another zone. The two limits of your key, the tier's and the endpoint's, give 429 with Retry-After (Rate limits). The SDKs retry 429 and 5xx twice by default (max_retries). In code, check the status and type, not the text. All the codes: Errors.

Tools (function calling)

The same shape as OpenAI: tools, tool_choice, and tool_calls in the answer.

python
tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

resp = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "What is the weather in Lugano?"}],
    tools=tools,
    tool_choice="auto",
)
print(resp.choices[0].message.tool_calls)

Send the result back as a message with role: "tool" and the tool_call_id, as with OpenAI. Schemas, structured output and their limits: Structured output.

Tier and zone from the SDK

The tier is a request header. The zone where a chat completion was processed comes back in the X-Zona header.

python
client = OpenAI(
    base_url="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    default_headers={"X-Siati-Tier": "fast"},   # every request
)

raw = client.chat.completions.with_raw_response.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Hello"}],
    extra_headers={"X-Siati-Tier": "slow"},     # this request only
)
print(raw.headers.get("X-Zona"))
resp = raw.parse()
typescript
const { data, response } = await client.chat.completions
  .create(
    { model: "qwen3.8-27b", messages: [{ role: "user", content: "Hello" }] },
    { headers: { "X-Siati-Tier": "fast" } },
  )
  .withResponse();
console.log(response.headers.get("x-zona"));

What a tier changes: Tiers. What a zone is: Zones.

Embeddings

python
emb = client.embeddings.create(
    model="bge-m3",
    input=["first text", "second text"],
)
print(len(emb.data[0].embedding))  # 1024
  • Re-embed your whole corpus. Vectors from OpenAI models and from bge-m3 are not comparable, and they have different lengths: do not mix them in one index.
  • 1024 dimensions, always. dimensions is accepted and has no effect.
  • input is a string or a list of up to 32 strings. Lists of token ids are refused.
  • encoding_format is float or base64. The Python SDK asks for base64 by default and decodes it for you.
  • model is not checked today: every request is answered by bge-m3, and the response says so in model. Write bge-m3 anyway.

Audio

Transcription takes whisper-1 or whisper-1-hd, and response_format json, verbose_json or text: srt and vtt are not available.

Speech uses our voices, listed in Audio. OpenAI voice names (alloy, echo, nova, …) are accepted and mapped to one of our voices, female or male, in the language given by language: a field of ours, Italian by default. For English text, send it:

python
speech = client.audio.speech.create(
    model="tts-1",
    voice="alloy",
    input="Hello from the other side.",
    response_format="mp3",
    extra_body={"language": "en"},
)
speech.write_to_file("hello.mp3")

The default format is wav, where OpenAI's is mp3: ask for mp3 if you need it. Only wav and mp3 exist. speed is accepted and has no effect today.

Pictures and video

Pictures (/v1/images/generations) and video (/v1/videos) are served only where AI Tokens has machines that make them. There, the API reference has a page for each, after Audio, with the models, the sizes and the prices. Where there are none, the requests answer 503 with code no_engine_available, or 400 model_not_found for a model name that is not served.

Where they exist, they differ from OpenAI in a few ways:

  • Pictures: prompt, model, size, n, response_format (b64_json or url) and seed, nothing else. quality, style, background and output_format are refused with a 400 that names them, and so is a size the model cannot make.
  • Video: a job you follow until it is completed, as with OpenAI. input_reference (starting from an image) is refused, and remix, edits, extensions and characters do not exist. One job at a time per person.
  • Both: the prompt can be checked before anything is made, and the result after; the pages of the two endpoints say which checks run. A refused prompt answers 400 prompt_refused; a refused result is not delivered and not charged.

LangChain

python
import os

from langchain_openai import ChatOpenAI, OpenAIEmbeddings

llm = ChatOpenAI(
    base_url="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="qwen3.8-27b",
)
llm.invoke("Hello")

emb = OpenAIEmbeddings(
    base_url="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="bge-m3",
    check_embedding_ctx_length=False,  # send text, not token ids
    chunk_size=32,                     # at most 32 texts per request
)

Without those two settings OpenAIEmbeddings sends token ids and up to 1000 texts per request, and both are refused.

LlamaIndex

python
import os

from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    api_base="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="qwen3.8-27b",
    is_chat_model=True,
)

OpenAILike is the LlamaIndex class for OpenAI-compatible endpoints. If you use tools through LlamaIndex, add is_function_calling_model=True.

Moving one call at a time

The SDK binds a client to one base URL, so an OpenAI client and a AI Tokens client can live side by side in the same code. Move one call, compare the answers, then move the next.

Search the docs

Type to search…