Skip to content

API reference

What works where

Every endpoint on one page: modes, models, inputs, the parameters honoured, refused or accepted without effect, the credential, the zones and what is billed.

Last updated: 2026-10-08

This page is the summary of the API reference: one row per endpoint, then what each one accepts. The pages of each endpoint have the details and the examples; when this page and another one disagree, write to info@daikolab.ch: one of them is wrong.

Three terms used below:

  • Honoured: the parameter does what its name says.
  • Refused: the request gets a 400 (or a 422) that names the field. Nothing runs and nothing is charged.
  • Accepted, no effect: the request succeeds and the field is ignored. These are the cases to know before you rely on a field.

At a glance

Endpoint Modes Credential Zones Billed per
POST /v1/chat/completions answer, or stream API key the key's; X-Zona input and output tokens, × tier
POST /v1/completions answer, or stream API key the key's; X-Zona input and output tokens, × tier
POST /v1/responses answer only API key the key's; X-Zona input and output tokens, × tier
GET /v1/models list API key — not billed
POST /v1/embeddings answer API key not yet input tokens, estimated, × tier
POST /v1/rerank answer API key not yet tokens, estimated, × tier
GET /v1/knowledge-bases, POST /v1/knowledge-bases/{slug}/ask list, answer API key; the rest of knowledge bases needs a signed-in user the brand's, not the key's the answer, at the chat model's prices
POST /v1/audio/transcriptions answer API key; signed-in user on the app API whisper-1 in the key's, where there is a machine; the rest in Switzerland seconds of audio
POST /v1/audio/speech the whole file API key; signed-in user on the app API not yet: in Switzerland characters of text
POST /v1/images/generations answer API key; signed-in user in the Studio the key's pictures delivered
/v1/videos a job you follow API key; signed-in user in the Studio the key's seconds of a completed clip

The last two exist only where AI Tokens has machines that make pictures and video; elsewhere they answer 503 no_engine_available. Prices are in Pricing, zones in Sovereignty, limits per minute in Rate limits. Every response carries X-Request-Id, the reference to quote when you write to us.

Text

POST /v1/chat/completions

  • Models: the chat models of the catalogue. Images as input only on the models with the images badge.
  • Inputs: messages with the roles system, user, assistant and tool; images as image_url parts, a data URI or a public http(s) address, PNG, JPEG, WebP or GIF, up to 8 MB each.
  • Honoured: temperature, top_p, max_tokens or max_completion_tokens (up to 8192), stop, n (1–8), seed, presence_penalty, frequency_penalty, logit_bias, logprobs, top_logprobs, parallel_tool_calls, tools, tool_choice, response_format (json_object, json_schema), user, service_tier (auto or default). Header X-Siati-Tier.
  • Refused: any name not in that list, typos included; metadata, store, functions, function_call, modalities, audio, prediction, web_search_options; n above 1 with stream: true. On the models served by Ollama, also n above 1, logprobs, top_logprobs, logit_bias and parallel_tool_calls. A value out of range gets 422.
  • Accepted, no effect: with stream: true, logprobs and top_logprobs (not in the chunks). stream_options.include_usage works: a last chunk carries the usage. On the models served by Ollama, tool_choice.
  • Details: Chat completions, Every parameter, Streaming, Structured output.

POST /v1/completions

  • Inputs: prompt, a string or a list with one string; more than one gets 400.
  • Parameters: those of chat completions, with the same rules.
  • Differences: the answer has OpenAI's text_completion shape, but with stream: true the chunks are those of chat completions, with the text in delta.content.

POST /v1/responses

  • Honoured: input (a string or a list of messages), instructions, max_output_tokens, temperature, top_p, text.format (json_object, json_schema), tools (in the Responses shape, your own functions only), tool_choice.
  • Refused: stream: true, with 501 stream_not_implemented; a value out of range, with 400.
  • Accepted, no effect: every other field, previous_response_id and store included: there is no state on the server, and a typo goes unnoticed.
  • Differences: an answer cut at max_output_tokens has status: incomplete.
  • Details: Responses API.

GET /v1/models

Every model AI Tokens serves, built at each call: the chat models of the catalogue, bge-m3 and bge-reranker-v2-m3, the picture and video models while a machine makes them, then whisper-1, whisper-1-hd and tts-1. Not listed: the voices, which are the voice of tts-1 (GET /v1/audio/voices), and tts-1-hd, accepted for OpenAI's clients and the same engine as tts-1. Every row has the same owned_by, the name of AI Tokens.

Search

POST /v1/embeddings

  • Model: always bge-m3, 1024 dimensions.
  • Honoured: input, a string or a list of up to 32 strings; encoding_format (float or base64).
  • Refused: lists of token ids, empty texts, more than 32 texts.
  • Accepted, no effect: model (the vectors always come from bge-m3), dimensions (always 1024), user, and any other field.
  • Details: Embeddings.

POST /v1/rerank

  • Model: always bge-reranker-v2-m3.
  • Honoured: query (up to 32,000 characters), documents (1 to 32), top_n, return_documents.
  • Refused: a missing query, more than 32 documents, an empty one.
  • Accepted, no effect: model, and any other field.
  • Details: Rerank.

Knowledge bases

  • With an API key: GET /v1/knowledge-bases lists yours and the shared ones that have a document ready; POST /v1/knowledge-bases/{slug}/ask honours question (up to 2,000 characters), model (a model of the catalogue, default qwen3.8-27b) and top_k (1–20). A value out of range gets 422; any other field is accepted with no effect.
  • With a signed-in user: creating knowledge bases, uploading documents (PDF, DOCX, TXT, Markdown, up to 50 MB), following and deleting them, in the dashboard or the app API.
  • Zones: the answer is written in the zones of AI Tokens; the zone limits of your key do not apply.
  • Details: Knowledge bases API.

Audio

POST /v1/audio/transcriptions

  • Models: whisper-1 and whisper-1-hd; another name gets 400.
  • Inputs: one file up to 25 MB, in WAV, MP3, M4A, MP4, Ogg, Opus, WebM, FLAC, 3GPP or AMR, checked on its content.
  • Honoured: language, prompt, response_format (json, verbose_json, text), temperature (0–1), timestamp_granularities (word, segment, with verbose_json only).
  • Refused: srt and vtt; word timings without verbose_json; a format we do not accept.
  • Accepted, no effect: any other field.
  • Where it runs: whisper-1 on a machine in your key's zones where AI Tokens has one, with X-Zona; whisper-1-hd, word timings, a prompt, audio over 10 minutes and the requests that machine cannot take run in Switzerland, outside the zones.
  • Details: Audio.

POST /v1/audio/speech

  • Voices: those of Audio, and OpenAI's voice names, mapped to ours by language.
  • Honoured: input (up to 5,000 characters), voice, response_format (wav, the default, or mp3), language (ours).
  • Refused: an unknown voice, a longer text, another format.
  • Accepted, no effect: model (one engine serves every request), speed, and any other field.
  • Output: the whole file at once, no streaming, marked as generated by AI: see AI-generated media.

Pictures and video

Where AI Tokens makes them, each has its own page after Audio in this reference.

POST /v1/images/generations

  • Model: z-image-turbo; another name gets 400 model_not_found.
  • Honoured: prompt (up to 4,000 characters), size (1024x1024, 1024x1536, 1536x1024), n (1–4), response_format (b64_json or url, a link valid 24 hours), seed.
  • Refused: every other field, OpenAI's quality, style, background and output_format included; a value the model cannot take; a prompt refused by the content check.
  • Output: PNG, marked as generated by AI. A picture refused by the check after it was made is not delivered and not charged.

/v1/videos

  • Mode: POST /v1/videos creates a job; GET /v1/videos/{id} follows it; GET /v1/videos/{id}/content downloads the MP4; GET /v1/videos lists, DELETE /v1/videos/{id} deletes. One job at a time per person.
  • Model: wan-2.2-5b-fast, the default.
  • Honoured: prompt (up to 2,000 characters), seconds (5), size (1280x704, 704x1280), seed.
  • Refused: input_reference, every other field, a value the model cannot make, a prompt refused by the content check.
  • Output: MP4, 24 frames a second, no sound, kept 7 days, marked as generated by AI.

Without code

A signed-in user reaches the same engines from the dashboard: the playground for chat, knowledge bases, and the Studio for voice, transcription, pictures and video. They count on the same credits and usage records as the API.

Not available

Assistants and threads, Batch, Files, fine-tuning, image edits and variations, video remix, edits, extensions and characters, audio translations, moderation, Realtime and vector stores. Those paths answer 404 with code: unknown_endpoint. Coming from OpenAI: OpenAI SDK migration.

Search the docs

Type to search…