API reference
Audio: transcription and speech
POST /v1/audio/transcriptions and POST /v1/audio/speech, in the shape of OpenAI's: what they accept, what they return, and what does not work yet.
Last updated: 2026-10-08
Two endpoints, in the shape of OpenAI's:
| Endpoint | Does | Engine |
|---|---|---|
POST /v1/audio/transcriptions |
audio → text | Whisper large-v3-turbo (whisper-1) or Whisper large-v3 (whisper-1-hd) |
POST /v1/audio/speech |
text → audio | Piper, six voices in four languages |
The OpenAI SDKs work with them once you change base_url and the key, and use
the models and fields of this page: What works where and
OpenAI SDK migration list the differences.
The examples below use:
export API_KEY="sk-…"
Transcription
POST https://api.aitokens.ch/v1/audio/transcriptions, as a multipart form.
curl https://api.aitokens.ch/v1/audio/transcriptions \
-H "Authorization: Bearer $API_KEY" \
-F file=@message.m4a \
-F model=whisper-1 \
-F language=en
{ "text": "Hello, I wanted to ask whether we can meet tomorrow." }
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])
with open("message.m4a", "rb") as f:
result = client.audio.transcriptions.create(
model="whisper-1",
file=f,
language="en", # optional: detected from the audio if omitted
)
print(result.text)
TypeScript (Node 20 or later, on your server: the key must not reach a browser)
// transcribe.mts. Run: API_KEY=sk-… npx tsx transcribe.mts
import { openAsBlob } from "node:fs";
const form = new FormData();
form.append("file", await openAsBlob("message.m4a"), "message.m4a");
form.append("model", "whisper-1");
const res = await fetch("https://api.aitokens.ch/v1/audio/transcriptions", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.API_KEY}` },
body: form,
});
const body = await res.json();
if (!res.ok) throw new Error(body.error?.message ?? body.message);
console.log(body.text);
Fields
| Field | ||
|---|---|---|
file |
required | The audio, up to 25 MB. See accepted formats. |
model |
optional | whisper-1 (default) runs Whisper large-v3-turbo; whisper-1-hd runs Whisper large-v3, in Switzerland only. Any other value gives 400. |
language |
optional | A two-letter code: it, en, de, fr… Without it, Whisper detects the language. |
prompt |
optional | Up to 1,024 characters of context: names and terms you expect, spelled the way you want them back. A request with a prompt runs in Switzerland. |
response_format |
optional | json (default), text or verbose_json. |
temperature |
optional | From 0 to 1. |
timestamp_granularities[] |
optional | word, segment or both. Only with verbose_json. Word timings run in Switzerland. |
prompt, language and response_format sent empty count as absent. SDKs
often send an empty variable instead of leaving the field out, and that used to
make the whole transcription fail.
What comes back
With json, the default, the object shown above: {"text": "…"}.
With text, the bare text, as Content-Type: text/plain; charset=utf-8.
With verbose_json, also the language, the duration in seconds and the
segments with their timings. The example shows only some of their fields. Their
id starts at 1, not at 0 as at OpenAI:
{
"task": "transcribe",
"language": "en",
"duration": 4.32,
"text": "Hello, how are you today? I wanted to ask whether we can meet tomorrow.",
"segments": [
{ "id": 1, "start": 0.0, "end": 2.1, "text": " Hello, how are you today?" },
{ "id": 2, "start": 2.1, "end": 4.32, "text": " I wanted to ask whether we can meet tomorrow." }
]
}
Per-word timings
timestamp_granularities[]=word adds words at the top level of the response,
with the start, end and probability of every word. It is what you need to
highlight the exact point in the audio, cut a recording, or show where
recognition was uncertain.
curl https://api.aitokens.ch/v1/audio/transcriptions \
-H "Authorization: Bearer $API_KEY" \
-F file=@note.wav \
-F model=whisper-1 \
-F response_format=verbose_json \
-F "timestamp_granularities[]=word"
{
"words": [
{ "start": 0.0, "end": 0.42, "word": " Hello,", "probability": 0.97 },
{ "start": 0.42, "end": 0.71, "word": " how", "probability": 0.99 }
]
}
Word timings also move the segments. Without them, segments follow one another with no gaps from 0:00, so the first one starts at 0:00 even when the audio opens with silence or music. With them, each segment starts at its first word. Measured on 6 October 2026, on a video with 2.5 seconds of silence before the voice: the first segment started at 0.00 s without word timings and at 2.38 s with them; after 8 seconds of music, it ran from 0.00 s to 14.82 s without them and started at 10.36 s with them. For subtitles, ask for word timings even if you use only the segments.
Word timings are made in Switzerland only: a request that asks for them runs there,
outside your zones. Segment timings (timestamp_granularities[]=segment) are
made wherever the request runs.
Timings exist only inside verbose_json: the other two forms have nowhere to put
them. Asking for them with another format gives 400, instead of an answer
without timings that you could not tell apart from audio in which no words were
heard. Until 18 August 2026 the field was accepted and not forwarded.
A meeting transcribed with word timings is one of the Three examples, with measured numbers.
Which model
whisper-1 is the default and the faster of the two. If it makes mistakes on
your audio — noise, phone calls, strong accents — run the same files through
whisper-1-hd and compare.
On long recordings, start with whisper-1: whisper-1-hd can get stuck
repeating one sentence, and leave out the speech that follows.
whisper-1-hd runs only in Switzerland, where the model is loaded on demand: the
first request after a long pause can take several seconds longer.
Which machine transcribes
whisper-1 runs on the transcription machine in your key's zones, where
AI Tokens has one, whenever that machine can do what you ask. These requests go
straight to the server in Switzerland instead:
model=whisper-1-hd;- per-word timings,
timestamp_granularities[]=word; - a
prompt; - audio longer than 10 minutes, or a file larger than 25 MB;
- 3GPP and AMR files.
If the machine in your zones does not answer, answers with an error, or refuses
the file, the request goes on to the server in Switzerland and you get the
transcription from there, a little later. 502 comes only if that server fails
too.
The response has the same fields wherever it ran. You can tell where from the
headers: X-Zona names the zone when the machine in your zones transcribed the
audio, and is missing when the server in Switzerland did. In verbose_json two things differ:
no_speech_probisnullwhen the machine in your zones transcribed the audio: its engine does not compute it. The server in Switzerland gives a number. If you rely on it, ask for word timings or usewhisper-1-hd: both run there.- Each engine cuts the audio into segments in its own way, so the same file can come back with a different number of segments, and different boundaries.
Accepted formats
The check looks at the content of the file, not at its name or at the
Content-Type you declare. A renamed file does not pass; a file with the wrong
Content-Type passes if what is inside is a format we accept.
These are the accepted types, as detected from the content:
audio/wav audio/wave audio/x-wav audio/vnd.wave
audio/mpeg audio/mp3 audio/x-mpeg
audio/mp4 audio/m4a audio/x-m4a video/mp4
audio/ogg video/ogg application/ogg audio/opus
audio/webm video/webm
audio/flac audio/x-flac
audio/3gpp video/3gpp audio/amr
3GPP and AMR files are always transcribed in Switzerland; the others wherever the request runs.
Why some start with video/. WebM, MP4 and Ogg are containers, and the type
detected from the content does not say whether there is video inside: a
voice-only recording is detected as video/webm or video/mp4. Browsers record
in these containers (Chrome in WebM, Safari in MP4), so refusing video/webm
would mean refusing dictation from a browser. Until 18 August 2026 it did.
When a file is refused, you get 400 and the message names the type that
was detected and lists the formats we accept:
{"error":{"message":"Unsupported audio format for 'dictation.avi': the file's content was recognised as video/x-msvideo. We accept WAV, MP3, M4A, MP4, Ogg, WebM, FLAC and 3GPP. …","type":"invalid_request_error"}}
Audio without sound
A file without sound returns empty text, with status 200, without asking
the model. It is not an error: the file is valid, there is just nothing to
transcribe. On digital silence Whisper answers "Thank you." with full
confidence, and on a dictation path an accidental tap on the microphone would
otherwise become a command. The threshold and the reasoning are in
What we change in responses.
How it works today:
- Only 16-bit PCM WAV is measured. Compressed formats always go to the model.
- The file is read a short stretch at a time, up to its first 2,000,000 bytes of samples: about a minute of 16 kHz mono audio, 11 seconds of 44.1 kHz stereo. As soon as one stretch is louder than about -50 dBFS, the file goes to the model. A WAV that starts with a pause and has speech later is transcribed: on 6 October 2026 a recording with 2.5 seconds of digital silence at the start came back whole. Until 3 October 2026 only the beginning of the file was measured, and a recording like that came back empty.
- A file is treated as silent only if we have read all of it: a longer WAV whose first 2,000,000 bytes are quiet goes to the model anyway.
- Noise is not silence. On 6 October 2026, three seconds of faint noise (about -59 dBFS) came back empty; three seconds of louder noise (about -48 dBFS) went to the model, which answered «Grazie.». With
verbose_json, the server in Switzerland gives every segment ano_speech_prob: 0.82 on that noise, 0.004 on a real question. - With
verbose_json, the empty answer hasdurationset to0andlanguageset to the one you sent, oritif you sent none. - With
response_format=text, the empty answer is an empty body with status200. It used to be a500. - Only this endpoint does it. The app endpoint
/api/v1/audio/transcribesends every file to the model.
When the machine in your zones transcribes, a second check looks at its answer,
for every format and on both endpoints: if all the model wrote is one of the
phrases Whisper writes on silence and noise («Grazie a tutti.», «Grazie.»,
«Thank you.», «so», or punctuation alone), each with a low confidence, the
answer is empty text, with the duration of the file. Only the whole answer is
judged: a recording that ends with «Grazie a tutti» keeps it, and so does what
the model writes over a pause between two sentences. Answers from the server in
Switzerland are passed on as they come.
Speech
POST https://api.aitokens.ch/v1/audio/speech, with a JSON body. The answer is the audio
file.
curl https://api.aitokens.ch/v1/audio/speech \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"voice": "amy",
"input": "Hello! This is a test of speech synthesis.",
"response_format": "mp3"
}' \
--output hello.mp3
Python, with the client from the transcription example:
res = client.audio.speech.create(
model="tts-1", # the SDK requires it; the value is ignored
voice="amy",
input="Hello! This is a test of speech synthesis.",
response_format="mp3", # the default here is wav, not mp3 as at OpenAI
)
res.write_to_file("hello.mp3")
TypeScript (Node 20 or later)
// speak.mts. Run: API_KEY=sk-… npx tsx speak.mts
import { writeFile } from "node:fs/promises";
const res = await fetch("https://api.aitokens.ch/v1/audio/speech", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ voice: "amy", input: "Hello!", response_format: "mp3" }),
});
if (!res.ok) {
const body = await res.json();
throw new Error(body.error?.message ?? body.message);
}
await writeFile("hello.mp3", new Uint8Array(await res.arrayBuffer()));
Fields
| Field | ||
|---|---|---|
input |
required | The text, up to 5,000 characters. |
voice |
optional | One of our voices or an OpenAI voice name. Without it, the default voice of language. |
language |
optional | Ours, not OpenAI's. A two-letter code that picks the default voice and the language OpenAI voice names are served in. Default it. |
response_format |
optional | wav (default) or mp3. Other formats give 400. |
model |
optional | Accepted and ignored, whatever its value: there is one engine. |
speed |
optional | Accepted between 0.5 and 2.0, and ignored: the audio always comes at normal speed. |
Here too, a field that is not in the table is ignored, not refused.
Where the text to read is checked against the content policy — it is decided
zone by zone — a text the check refuses is not read: you get 400 with code
input_refused. If the check cannot answer, nothing is read either: 503 with
moderation_unavailable. See Errors.
Voices
| Voice | Language | |
|---|---|---|
paola |
Italian | female; default for it |
riccardo |
Italian | male; a lower-quality voice model (x_low) |
amy |
English (US) | female; default for en |
ryan |
English (US) | male |
thorsten |
German (Germany) | male; default for de |
siwis |
French (France) | female; default for fr |
Voice names are case-sensitive. An unknown voice gives 400, and the message
lists the valid ones.
OpenAI voice names work too, in upper or lower case. They carry only the
character of the voice; the language comes from language:
alloy,nova,shimmer,coral,sage→ the female voice of the language;echo,onyx,fable,ash,ballad,verse→ the male voice.
German has only a male voice and French only a female one: there you get that
voice whatever the name asks for. So alloy with "language": "en" is amy,
and alloy with no language is paola.
Other languages. There is no voice for Romansh, or for any language not in
the table: with such a language, the text is read by an Italian voice.
What comes back
The audio, as Content-Type: audio/wav or audio/mpeg. The type always matches
the content: if the MP3 encoder were missing you would get 502, never a WAV
labelled as MP3. MP3 is several times smaller than WAV, which is the reason to
ask for it on a mobile network.
There is no streaming: the whole audio is generated before the answer starts. For spoken replies, split the text into sentences and request them one at a time, so the first can play while the next is generated.
Every audio file is marked as generated by AI inside its metadata, for article
50 of the EU AI Act: a LIST/INFO comment in a WAV, ID3 tags in an MP3.
Editors can drop them: where you publish a voice, say in words that it was
generated. Details: AI-generated media.
Limits
- Size: 25 MB per audio file; 5,000 characters per speech request. A file over 25 MB gets
400with the limit in the message («The file field must not be greater than 25600 kilobytes.»); over 50 MB the message says only «The file failed to upload.». A request over 64 MB is stopped before it reaches the API:413, with an HTML page as its body, not JSON. - Length: the transcription machine in your zones takes up to 10 minutes of audio per file; longer audio is transcribed in Switzerland, within the time below.
- Time: a transcription has 300 seconds. If it takes longer, the request ends with
504, and that body too is an HTML page, not JSON. On the server in Switzerland, a recording of more than about ten minutes can take that long: send it in parts of five minutes. Parts shorter than ten minutes also stay on the machine in your zones, where there is one. - Per key: the per-minute limit of the request's tier, on one counter with all your other
/v1calls. - Per endpoint:
30requests per minute on transcriptions,120on speech, also counted per key. The counter is shared with the other/v1endpoints: once a key has made 30 requests to any of them in the same minute, its transcriptions get429until the minute is over. Other keys behind the same address, such as an office or a class, have their own: see Rate limits.
How both limits count, and how to tell their 429 apart, is in
Rate limits.
Errors
| Status | type |
When |
|---|---|---|
400 |
invalid_request_error |
Something in the request is refused: a missing or invalid field, a file over 25 MB or in a format we do not accept, word timings without verbose_json, an unknown voice, a text over 5,000 characters. The message says what. |
400 |
invalid_request_error, with code input_refused and param input |
Speech, where the text to read is checked: the check refused the text, and it was not read. The message says why when it can. |
401 |
invalid_request_error, with code invalid_api_key |
The key is missing, malformed, revoked or expired, or the account is disabled. |
429 |
see below | Over one of the limits. |
413 |
none: the body is an HTML page | The request is over 64 MB. The web server in front of the API stops it. |
502 |
bad_gateway |
No transcription server answered (the machine in your zones, then the server in Switzerland), or the speech engine did not answer, or one answered with an error. |
503 |
server_error, with code moderation_unavailable |
Speech, where the text to read is checked: the check did not answer, so the text was not read. Try again in a minute. |
504 |
none: the body is an HTML page | The transcription took more than 300 seconds. Send shorter parts. |
On 429, Retry-After says how many seconds to wait. The two limits answer
with different bodies: the tier's limit with the usual error object
("type": "rate_limit_exceeded"), the endpoint's limit with only
{"message": "Too Many Attempts."}.
On 502, retry a few times with growing pauses. When the engines refuse a file,
that also comes back as 502 today: if the same file keeps failing, the file is
the likely cause.
Branch on the status and on type, not on the text of the message: the text
can change, and the messages of the field checks follow the language of
AI Tokens.
To report a problem, write to info@daikolab.ch with the X-Request-Id header
of the response, the time of the request in UTC and the name of your key —
never the key itself. The header is the reference we record the request under,
and the same value as the request_id in an error body. The other statuses of
the API are in Errors.
Price
Transcription is charged per second of audio, and speech per character of the text you send, at the prices in Pricing: in CHF, or in credits where AI Tokens works with credits. A request refused before it runs, a request that fails, and a silent WAV answered with empty text without asking the model cost nothing. Until 3 October 2026 transcription was not billed at all; since then the charge and the price list read the same value.
Audio calls that reach an engine appear in your usage records, in the dashboard.
A voice conversation
Transcribe what the user said, answer with a chat model, read the answer aloud:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])
with open("question.wav", "rb") as f:
heard = client.audio.transcriptions.create(model="whisper-1", file=f, language="en")
if not heard.text:
raise SystemExit("Nothing was said.") # a silent WAV gives empty text
answer = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": heard.text}],
max_tokens=300, # a short answer: speech takes up to 5,000 characters
).choices[0].message.content
client.audio.speech.create(
model="tts-1",
voice="amy",
input=answer,
response_format="mp3",
).write_to_file("answer.mp3")
Names and terms. If the audio contains names or technical terms, put them in
prompt, spelled the way you want them back. A request with a prompt is
transcribed in Switzerland:
with open("meeting.m4a", "rb") as f:
client.audio.transcriptions.create(
model="whisper-1-hd",
file=f,
prompt="Kowalczyk, Qdrant, rerank, embeddings",
)
In apps that sign their users in
Apps that sign their users in use the app API at https://my.aitokens.ch/api/v1,
with a user token instead of a key: see
Authentication. The same engines are
there, with other names:
| App API | Key API |
|---|---|
POST /api/v1/audio/transcribe, file in audio |
POST /v1/audio/transcriptions, file in file |
POST /api/v1/audio/synthesize, text in text, format in format |
POST /v1/audio/speech, input, response_format |
GET /api/v1/audio/voices |
none |
curl https://my.aitokens.ch/api/v1/audio/transcribe \
-H "Authorization: Bearer $USER_TOKEN" \
-F audio=@note.m4a \
-F with_segments=true
{ "text": "…", "language": "en", "duration_seconds": 3.4, "model": "whisper-1", "segments": [] }
curl https://my.aitokens.ch/api/v1/audio/voices \
-H "Authorization: Bearer $USER_TOKEN"
{
"voices": [
{ "name": "paola", "language": "it", "gender": "female" },
{ "name": "riccardo", "language": "it", "gender": "male" }
],
"defaults_by_language": { "it": "paola", "en": "amy", "de": "thorsten", "fr": "siwis", "rm": "paola" }
}
How they differ from the key API:
- transcribe accepts
audio,model,language,promptandwith_segments. There is notemperature, no per-word timing and no check of silent WAV files, and an empty field is refused instead of counting as absent. It runs where the key API would run it, in the zones of AI Tokens, and its answer carriesX-Zonathe same way. - synthesize accepts
text,voice,languageandformat(wavormp3). Withoutlanguage, the default voice follows the language of the user's account. The answer addsX-Voice, the voice used, andX-Duration-Seconds, an estimate of the length. - Errors have the app shape,
{"error": {"code": "…", "message": "…"}}:422validation_errorwith the fields,400invalid_voice,502upstream_unavailable, and where the text to read is checked400input_refusedand503moderation_unavailable. - Limits are per signed-in user: 30 transcriptions and 60 syntheses per minute, on one counter shared with the other app endpoints that have a limit (the list).
voiceshas none.