Skip to content

API reference

Audio: transcription and speech

POST /v1/audio/transcriptions and POST /v1/audio/speech, in the shape of OpenAI's: what they accept, what they return, and what does not work yet.

Last updated: 2026-10-08

Two endpoints, in the shape of OpenAI's:

Endpoint Does Engine
POST /v1/audio/transcriptions audio → text Whisper large-v3-turbo (whisper-1) or Whisper large-v3 (whisper-1-hd)
POST /v1/audio/speech text → audio Piper, six voices in four languages

The OpenAI SDKs work with them once you change base_url and the key, and use the models and fields of this page: What works where and OpenAI SDK migration list the differences. The examples below use:

bash
export API_KEY="sk-…"

Transcription

POST https://api.aitokens.ch/v1/audio/transcriptions, as a multipart form.

bash
curl https://api.aitokens.ch/v1/audio/transcriptions \
  -H "Authorization: Bearer $API_KEY" \
  -F file=@message.m4a \
  -F model=whisper-1 \
  -F language=en
json
{ "text": "Hello, I wanted to ask whether we can meet tomorrow." }

Python

python
import os

from openai import OpenAI

client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])

with open("message.m4a", "rb") as f:
    result = client.audio.transcriptions.create(
        model="whisper-1",
        file=f,
        language="en",  # optional: detected from the audio if omitted
    )

print(result.text)

TypeScript (Node 20 or later, on your server: the key must not reach a browser)

typescript
// transcribe.mts. Run: API_KEY=sk-… npx tsx transcribe.mts
import { openAsBlob } from "node:fs";

const form = new FormData();
form.append("file", await openAsBlob("message.m4a"), "message.m4a");
form.append("model", "whisper-1");

const res = await fetch("https://api.aitokens.ch/v1/audio/transcriptions", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.API_KEY}` },
  body: form,
});
const body = await res.json();
if (!res.ok) throw new Error(body.error?.message ?? body.message);
console.log(body.text);

Fields

Field
file required The audio, up to 25 MB. See accepted formats.
model optional whisper-1 (default) runs Whisper large-v3-turbo; whisper-1-hd runs Whisper large-v3, in Switzerland only. Any other value gives 400.
language optional A two-letter code: it, en, de, fr… Without it, Whisper detects the language.
prompt optional Up to 1,024 characters of context: names and terms you expect, spelled the way you want them back. A request with a prompt runs in Switzerland.
response_format optional json (default), text or verbose_json.
temperature optional From 0 to 1.
timestamp_granularities[] optional word, segment or both. Only with verbose_json. Word timings run in Switzerland.

prompt, language and response_format sent empty count as absent. SDKs often send an empty variable instead of leaving the field out, and that used to make the whole transcription fail.

What comes back

With json, the default, the object shown above: {"text": "…"}.

With text, the bare text, as Content-Type: text/plain; charset=utf-8.

With verbose_json, also the language, the duration in seconds and the segments with their timings. The example shows only some of their fields. Their id starts at 1, not at 0 as at OpenAI:

json
{
  "task": "transcribe",
  "language": "en",
  "duration": 4.32,
  "text": "Hello, how are you today? I wanted to ask whether we can meet tomorrow.",
  "segments": [
    { "id": 1, "start": 0.0, "end": 2.1, "text": " Hello, how are you today?" },
    { "id": 2, "start": 2.1, "end": 4.32, "text": " I wanted to ask whether we can meet tomorrow." }
  ]
}

Per-word timings

timestamp_granularities[]=word adds words at the top level of the response, with the start, end and probability of every word. It is what you need to highlight the exact point in the audio, cut a recording, or show where recognition was uncertain.

bash
curl https://api.aitokens.ch/v1/audio/transcriptions \
  -H "Authorization: Bearer $API_KEY" \
  -F file=@note.wav \
  -F model=whisper-1 \
  -F response_format=verbose_json \
  -F "timestamp_granularities[]=word"
json
{
  "words": [
    { "start": 0.0, "end": 0.42, "word": " Hello,", "probability": 0.97 },
    { "start": 0.42, "end": 0.71, "word": " how", "probability": 0.99 }
  ]
}

Word timings also move the segments. Without them, segments follow one another with no gaps from 0:00, so the first one starts at 0:00 even when the audio opens with silence or music. With them, each segment starts at its first word. Measured on 6 October 2026, on a video with 2.5 seconds of silence before the voice: the first segment started at 0.00 s without word timings and at 2.38 s with them; after 8 seconds of music, it ran from 0.00 s to 14.82 s without them and started at 10.36 s with them. For subtitles, ask for word timings even if you use only the segments.

Word timings are made in Switzerland only: a request that asks for them runs there, outside your zones. Segment timings (timestamp_granularities[]=segment) are made wherever the request runs.

Timings exist only inside verbose_json: the other two forms have nowhere to put them. Asking for them with another format gives 400, instead of an answer without timings that you could not tell apart from audio in which no words were heard. Until 18 August 2026 the field was accepted and not forwarded.

A meeting transcribed with word timings is one of the Three examples, with measured numbers.

Which model

whisper-1 is the default and the faster of the two. If it makes mistakes on your audio — noise, phone calls, strong accents — run the same files through whisper-1-hd and compare.

On long recordings, start with whisper-1: whisper-1-hd can get stuck repeating one sentence, and leave out the speech that follows.

whisper-1-hd runs only in Switzerland, where the model is loaded on demand: the first request after a long pause can take several seconds longer.

Which machine transcribes

whisper-1 runs on the transcription machine in your key's zones, where AI Tokens has one, whenever that machine can do what you ask. These requests go straight to the server in Switzerland instead:

  • model=whisper-1-hd;
  • per-word timings, timestamp_granularities[]=word;
  • a prompt;
  • audio longer than 10 minutes, or a file larger than 25 MB;
  • 3GPP and AMR files.

If the machine in your zones does not answer, answers with an error, or refuses the file, the request goes on to the server in Switzerland and you get the transcription from there, a little later. 502 comes only if that server fails too.

The response has the same fields wherever it ran. You can tell where from the headers: X-Zona names the zone when the machine in your zones transcribed the audio, and is missing when the server in Switzerland did. In verbose_json two things differ:

  • no_speech_prob is null when the machine in your zones transcribed the audio: its engine does not compute it. The server in Switzerland gives a number. If you rely on it, ask for word timings or use whisper-1-hd: both run there.
  • Each engine cuts the audio into segments in its own way, so the same file can come back with a different number of segments, and different boundaries.

Accepted formats

The check looks at the content of the file, not at its name or at the Content-Type you declare. A renamed file does not pass; a file with the wrong Content-Type passes if what is inside is a format we accept.

These are the accepted types, as detected from the content:

text
audio/wav  audio/wave  audio/x-wav  audio/vnd.wave
audio/mpeg  audio/mp3  audio/x-mpeg
audio/mp4  audio/m4a  audio/x-m4a  video/mp4
audio/ogg  video/ogg  application/ogg  audio/opus
audio/webm  video/webm
audio/flac  audio/x-flac
audio/3gpp  video/3gpp  audio/amr

3GPP and AMR files are always transcribed in Switzerland; the others wherever the request runs.

Why some start with video/. WebM, MP4 and Ogg are containers, and the type detected from the content does not say whether there is video inside: a voice-only recording is detected as video/webm or video/mp4. Browsers record in these containers (Chrome in WebM, Safari in MP4), so refusing video/webm would mean refusing dictation from a browser. Until 18 August 2026 it did.

When a file is refused, you get 400 and the message names the type that was detected and lists the formats we accept:

json
{"error":{"message":"Unsupported audio format for 'dictation.avi': the file's content was recognised as video/x-msvideo. We accept WAV, MP3, M4A, MP4, Ogg, WebM, FLAC and 3GPP. …","type":"invalid_request_error"}}

Audio without sound

A file without sound returns empty text, with status 200, without asking the model. It is not an error: the file is valid, there is just nothing to transcribe. On digital silence Whisper answers "Thank you." with full confidence, and on a dictation path an accidental tap on the microphone would otherwise become a command. The threshold and the reasoning are in What we change in responses.

How it works today:

  • Only 16-bit PCM WAV is measured. Compressed formats always go to the model.
  • The file is read a short stretch at a time, up to its first 2,000,000 bytes of samples: about a minute of 16 kHz mono audio, 11 seconds of 44.1 kHz stereo. As soon as one stretch is louder than about -50 dBFS, the file goes to the model. A WAV that starts with a pause and has speech later is transcribed: on 6 October 2026 a recording with 2.5 seconds of digital silence at the start came back whole. Until 3 October 2026 only the beginning of the file was measured, and a recording like that came back empty.
  • A file is treated as silent only if we have read all of it: a longer WAV whose first 2,000,000 bytes are quiet goes to the model anyway.
  • Noise is not silence. On 6 October 2026, three seconds of faint noise (about -59 dBFS) came back empty; three seconds of louder noise (about -48 dBFS) went to the model, which answered «Grazie.». With verbose_json, the server in Switzerland gives every segment a no_speech_prob: 0.82 on that noise, 0.004 on a real question.
  • With verbose_json, the empty answer has duration set to 0 and language set to the one you sent, or it if you sent none.
  • With response_format=text, the empty answer is an empty body with status 200. It used to be a 500.
  • Only this endpoint does it. The app endpoint /api/v1/audio/transcribe sends every file to the model.

When the machine in your zones transcribes, a second check looks at its answer, for every format and on both endpoints: if all the model wrote is one of the phrases Whisper writes on silence and noise («Grazie a tutti.», «Grazie.», «Thank you.», «so», or punctuation alone), each with a low confidence, the answer is empty text, with the duration of the file. Only the whole answer is judged: a recording that ends with «Grazie a tutti» keeps it, and so does what the model writes over a pause between two sentences. Answers from the server in Switzerland are passed on as they come.


Speech

POST https://api.aitokens.ch/v1/audio/speech, with a JSON body. The answer is the audio file.

bash
curl https://api.aitokens.ch/v1/audio/speech \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "voice": "amy",
    "input": "Hello! This is a test of speech synthesis.",
    "response_format": "mp3"
  }' \
  --output hello.mp3

Python, with the client from the transcription example:

python
res = client.audio.speech.create(
    model="tts-1",          # the SDK requires it; the value is ignored
    voice="amy",
    input="Hello! This is a test of speech synthesis.",
    response_format="mp3",  # the default here is wav, not mp3 as at OpenAI
)
res.write_to_file("hello.mp3")

TypeScript (Node 20 or later)

typescript
// speak.mts. Run: API_KEY=sk-… npx tsx speak.mts
import { writeFile } from "node:fs/promises";

const res = await fetch("https://api.aitokens.ch/v1/audio/speech", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ voice: "amy", input: "Hello!", response_format: "mp3" }),
});
if (!res.ok) {
  const body = await res.json();
  throw new Error(body.error?.message ?? body.message);
}
await writeFile("hello.mp3", new Uint8Array(await res.arrayBuffer()));

Fields

Field
input required The text, up to 5,000 characters.
voice optional One of our voices or an OpenAI voice name. Without it, the default voice of language.
language optional Ours, not OpenAI's. A two-letter code that picks the default voice and the language OpenAI voice names are served in. Default it.
response_format optional wav (default) or mp3. Other formats give 400.
model optional Accepted and ignored, whatever its value: there is one engine.
speed optional Accepted between 0.5 and 2.0, and ignored: the audio always comes at normal speed.

Here too, a field that is not in the table is ignored, not refused.

Where the text to read is checked against the content policy — it is decided zone by zone — a text the check refuses is not read: you get 400 with code input_refused. If the check cannot answer, nothing is read either: 503 with moderation_unavailable. See Errors.

Voices

Voice Language
paola Italian female; default for it
riccardo Italian male; a lower-quality voice model (x_low)
amy English (US) female; default for en
ryan English (US) male
thorsten German (Germany) male; default for de
siwis French (France) female; default for fr

Voice names are case-sensitive. An unknown voice gives 400, and the message lists the valid ones.

OpenAI voice names work too, in upper or lower case. They carry only the character of the voice; the language comes from language:

  • alloy, nova, shimmer, coral, sage → the female voice of the language;
  • echo, onyx, fable, ash, ballad, verse → the male voice.

German has only a male voice and French only a female one: there you get that voice whatever the name asks for. So alloy with "language": "en" is amy, and alloy with no language is paola.

Other languages. There is no voice for Romansh, or for any language not in the table: with such a language, the text is read by an Italian voice.

What comes back

The audio, as Content-Type: audio/wav or audio/mpeg. The type always matches the content: if the MP3 encoder were missing you would get 502, never a WAV labelled as MP3. MP3 is several times smaller than WAV, which is the reason to ask for it on a mobile network.

There is no streaming: the whole audio is generated before the answer starts. For spoken replies, split the text into sentences and request them one at a time, so the first can play while the next is generated.

Every audio file is marked as generated by AI inside its metadata, for article 50 of the EU AI Act: a LIST/INFO comment in a WAV, ID3 tags in an MP3. Editors can drop them: where you publish a voice, say in words that it was generated. Details: AI-generated media.


Limits

  • Size: 25 MB per audio file; 5,000 characters per speech request. A file over 25 MB gets 400 with the limit in the message («The file field must not be greater than 25600 kilobytes.»); over 50 MB the message says only «The file failed to upload.». A request over 64 MB is stopped before it reaches the API: 413, with an HTML page as its body, not JSON.
  • Length: the transcription machine in your zones takes up to 10 minutes of audio per file; longer audio is transcribed in Switzerland, within the time below.
  • Time: a transcription has 300 seconds. If it takes longer, the request ends with 504, and that body too is an HTML page, not JSON. On the server in Switzerland, a recording of more than about ten minutes can take that long: send it in parts of five minutes. Parts shorter than ten minutes also stay on the machine in your zones, where there is one.
  • Per key: the per-minute limit of the request's tier, on one counter with all your other /v1 calls.
  • Per endpoint: 30 requests per minute on transcriptions, 120 on speech, also counted per key. The counter is shared with the other /v1 endpoints: once a key has made 30 requests to any of them in the same minute, its transcriptions get 429 until the minute is over. Other keys behind the same address, such as an office or a class, have their own: see Rate limits.

How both limits count, and how to tell their 429 apart, is in Rate limits.

Errors

Status type When
400 invalid_request_error Something in the request is refused: a missing or invalid field, a file over 25 MB or in a format we do not accept, word timings without verbose_json, an unknown voice, a text over 5,000 characters. The message says what.
400 invalid_request_error, with code input_refused and param input Speech, where the text to read is checked: the check refused the text, and it was not read. The message says why when it can.
401 invalid_request_error, with code invalid_api_key The key is missing, malformed, revoked or expired, or the account is disabled.
429 see below Over one of the limits.
413 none: the body is an HTML page The request is over 64 MB. The web server in front of the API stops it.
502 bad_gateway No transcription server answered (the machine in your zones, then the server in Switzerland), or the speech engine did not answer, or one answered with an error.
503 server_error, with code moderation_unavailable Speech, where the text to read is checked: the check did not answer, so the text was not read. Try again in a minute.
504 none: the body is an HTML page The transcription took more than 300 seconds. Send shorter parts.

On 429, Retry-After says how many seconds to wait. The two limits answer with different bodies: the tier's limit with the usual error object ("type": "rate_limit_exceeded"), the endpoint's limit with only {"message": "Too Many Attempts."}.

On 502, retry a few times with growing pauses. When the engines refuse a file, that also comes back as 502 today: if the same file keeps failing, the file is the likely cause.

Branch on the status and on type, not on the text of the message: the text can change, and the messages of the field checks follow the language of AI Tokens.

To report a problem, write to info@daikolab.ch with the X-Request-Id header of the response, the time of the request in UTC and the name of your key — never the key itself. The header is the reference we record the request under, and the same value as the request_id in an error body. The other statuses of the API are in Errors.

Price

Transcription is charged per second of audio, and speech per character of the text you send, at the prices in Pricing: in CHF, or in credits where AI Tokens works with credits. A request refused before it runs, a request that fails, and a silent WAV answered with empty text without asking the model cost nothing. Until 3 October 2026 transcription was not billed at all; since then the charge and the price list read the same value.

Audio calls that reach an engine appear in your usage records, in the dashboard.


A voice conversation

Transcribe what the user said, answer with a chat model, read the answer aloud:

python
import os

from openai import OpenAI

client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])

with open("question.wav", "rb") as f:
    heard = client.audio.transcriptions.create(model="whisper-1", file=f, language="en")

if not heard.text:
    raise SystemExit("Nothing was said.")  # a silent WAV gives empty text

answer = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": heard.text}],
    max_tokens=300,  # a short answer: speech takes up to 5,000 characters
).choices[0].message.content

client.audio.speech.create(
    model="tts-1",
    voice="amy",
    input=answer,
    response_format="mp3",
).write_to_file("answer.mp3")

Names and terms. If the audio contains names or technical terms, put them in prompt, spelled the way you want them back. A request with a prompt is transcribed in Switzerland:

python
with open("meeting.m4a", "rb") as f:
    client.audio.transcriptions.create(
        model="whisper-1-hd",
        file=f,
        prompt="Kowalczyk, Qdrant, rerank, embeddings",
    )

In apps that sign their users in

Apps that sign their users in use the app API at https://my.aitokens.ch/api/v1, with a user token instead of a key: see Authentication. The same engines are there, with other names:

App API Key API
POST /api/v1/audio/transcribe, file in audio POST /v1/audio/transcriptions, file in file
POST /api/v1/audio/synthesize, text in text, format in format POST /v1/audio/speech, input, response_format
GET /api/v1/audio/voices none
bash
curl https://my.aitokens.ch/api/v1/audio/transcribe \
  -H "Authorization: Bearer $USER_TOKEN" \
  -F audio=@note.m4a \
  -F with_segments=true
json
{ "text": "…", "language": "en", "duration_seconds": 3.4, "model": "whisper-1", "segments": [] }
bash
curl https://my.aitokens.ch/api/v1/audio/voices \
  -H "Authorization: Bearer $USER_TOKEN"
json
{
  "voices": [
    { "name": "paola", "language": "it", "gender": "female" },
    { "name": "riccardo", "language": "it", "gender": "male" }
  ],
  "defaults_by_language": { "it": "paola", "en": "amy", "de": "thorsten", "fr": "siwis", "rm": "paola" }
}

How they differ from the key API:

  • transcribe accepts audio, model, language, prompt and with_segments. There is no temperature, no per-word timing and no check of silent WAV files, and an empty field is refused instead of counting as absent. It runs where the key API would run it, in the zones of AI Tokens, and its answer carries X-Zona the same way.
  • synthesize accepts text, voice, language and format (wav or mp3). Without language, the default voice follows the language of the user's account. The answer adds X-Voice, the voice used, and X-Duration-Seconds, an estimate of the length.
  • Errors have the app shape, {"error": {"code": "…", "message": "…"}}: 422 validation_error with the fields, 400 invalid_voice, 502 upstream_unavailable, and where the text to read is checked 400 input_refused and 503 moderation_unavailable.
  • Limits are per signed-in user: 30 transcriptions and 60 syntheses per minute, on one counter shared with the other app endpoints that have a limit (the list). voices has none.

Search the docs

Type to search…