API reference
Video: generation
POST /v1/videos and the four requests around it, in the shape of OpenAI's Videos API: what they accept, how a job goes from the queue to the file, what it costs and what is refused.
Last updated: 2026-10-11
A request creates a job. The job makes a clip of a few seconds on one of our machines while you ask how it is going, then you download the MP4.
| Request | Does |
|---|---|
POST /v1/videos |
creates a job from a prompt |
GET /v1/videos/{id} |
the job: its status and progress |
GET /v1/videos/{id}/content |
the MP4, once the job is completed |
GET /v1/videos |
your jobs, newest first |
DELETE /v1/videos/{id} |
deletes the job and its file |
The shape is OpenAI's Videos API. Code written for it works once it uses the base URL and the key of AI Tokens, a model of this page instead of OpenAI's, and only the fields below: What works where and OpenAI SDK migration list the differences.
In three requests
# 1. Create the job
curl https://api.aitokens.ch/v1/videos \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "wan-2.2-5b-fast", "prompt": "A green dragon flies over snowy mountains, wide cinematic shot"}'
# {"id": "video_01k6…", "object": "video", "status": "queued", "progress": 0, …}
# 2. Ask how it is going, every couple of seconds, until "completed" or "failed"
curl https://api.aitokens.ch/v1/videos/video_01k6… -H "Authorization: Bearer $API_KEY"
# 3. Download the clip
curl -o drago.mp4 https://api.aitokens.ch/v1/videos/video_01k6…/content -H "Authorization: Bearer $API_KEY"
Models
model |
Makes | seconds |
size |
|---|---|---|---|
wan-2.2-5b-fast (default) |
text → video, quick | 5 |
1280x704 (landscape, default), 704x1280 (portrait) |
wan-2.2-5b |
text → video, more detail, slower. Not served today: on 6 October 2026 a request got 503 no_engine_available |
5 |
1280x704, 704x1280 |
Clips are MP4 (H.264), 24 frames a second, 121 frames, 5.04 seconds, without
sound. GET /v1/models lists a video model while a machine makes it; this table says what each one does.
A machine makes one clip at a time: how long a request waits depends on the
queue in front of it as well as on the model.
How long it takes
Measured on 6 October 2026 with wan-2.2-5b-fast: six clips, one after the
other, made with the examples on this page.
| creating the job | 0.6 s, the check on the prompt included |
| waiting in the queue | 4 to 9 s, with nobody in front |
| making the clip | about 100 s, landscape or portrait |
| from the request to the file on disk | 109 to 114 s |
| downloading | 0.2 s |
| the file | 0.6 to 1.6 MB landscape, 2.4 MB portrait |
Every clip in front of yours in the queue adds about 100 seconds.
Creating a job
POST /v1/videos, with a JSON body:
| Field | ||
|---|---|---|
prompt |
required | What the video shows, up to 2000 characters. Italian and English both work; describe the subject, the action, the setting, the light and the camera. |
model |
optional | One of the models above. Default wan-2.2-5b-fast. |
seconds |
optional | The length, as the model allows it: "5" or 5. |
size |
optional | "1280x704" or "704x1280", as width x height. |
seed |
optional | Our extension: a whole number from 0 to 2147483647. The same seed with the same prompt and size gives the same clip: on 6 October 2026 two requests gave the same file, byte for byte. Without it, every clip is different. |
input_reference |
refused for now | Starting from an image is not available yet: the request comes back 400. |
The model follows the subject, the setting and the light well; a style or a camera move is not guaranteed. In our tests «A low-poly red fox runs through an autumn forest» came out as a photorealistic fox, and a knight «with the camera following from behind» walked towards the camera.
Every other field is refused with a 400 that names it, and a value the model
cannot make is refused, not rounded: ask for "seconds": 8 and you get a 400
that says which lengths exist. If the request comes back 200, what you asked for
is what you get.
The answer is the job:
{
"id": "video_01k6zq3d2m8x4v5n7r9t0w2y4a",
"object": "video",
"model": "wan-2.2-5b-fast",
"status": "queued",
"progress": 0,
"created_at": 1791302400,
"completed_at": null,
"expires_at": null,
"error": null,
"prompt": "A green dragon flies over snowy mountains, wide cinematic shot",
"remixed_from_video_id": null,
"seconds": "5",
"size": "1280x704"
}
Following the job
GET /v1/videos/{id} returns the same object, updated:
status |
Means |
|---|---|
queued |
waiting for its machine |
in_progress |
being made; progress goes from 0 to 100 |
completed |
done: the MP4 is ready, until expires_at |
failed |
not made; error.code and error.message say why, and nothing is charged |
While a job is queued or in progress, the answer carries the header
openai-poll-after-ms: 2000: ask every two seconds, not more often. OpenAI's
SDK reads that header by itself.
With wan-2.2-5b-fast the machine only says whether the clip is done. Until it
is, progress is our estimate from the 100 seconds a clip usually takes, and it
stops at 95. The job is brought up to date every five seconds or so, so two
answers in a row often have the same progress.
A job that fails has one of these codes in error.code:
error.code |
Means |
|---|---|
generation_failed |
the machine could not make this clip |
output_refused |
the finished clip failed the check on its frames: it was deleted (see What is refused) |
moderation_unavailable |
the finished clip could not be checked, so it was not delivered |
engine_error |
the machine did not accept the job |
engine_unreachable |
the machine stopped answering for a minute |
engine_gone |
the machine left the service while making it |
timeout |
the clip was not done within 45 minutes of the request, and was stopped |
internal_error |
an error on our side |
Downloading
GET /v1/videos/{id}/content returns the MP4 (Content-Type: video/mp4).
Only the video exists: ?variant=thumbnail or spritesheet is refused with a
400. Asking before the job is completed gives 400 with code
video_not_ready.
Files are kept 7 days. After that the job is still listed, with its
expires_at in the past, but its content answers 404 video_expired.
Download what you want to keep.
Listing and deleting
GET /v1/videos lists your jobs, newest first:
| Parameter | |
|---|---|
limit |
1 to 100, default 20 |
order |
desc (default) or asc |
after |
the id of the last job of the previous page |
{"object": "list", "data": [ … ], "first_id": "video_…", "last_id": "video_…", "has_more": false}
DELETE /v1/videos/{id} deletes the job and its file, and stops it if it is
still being made:
{"id": "video_…", "object": "video.deleted", "deleted": true}
After that the id answers 404 with code video_not_found, and so does its
content. You only ever see your own jobs: someone else's id answers the same
404.
One video at a time
Each person has one job at a time, queued or in progress. A second request
before the first ends gets a 429 with code too_many_video_jobs and the id
of the job in the way. A machine makes one clip at a time, and nobody should wait behind
someone else's twenty clips. To make many
clips, make them one after the other: see the example at the end.
What is refused
Before a job enters the queue, one of our own models reads the prompt. These are refused:
- sexual content or nudity;
- minors in any unsafe, violent or sexual situation;
- real, identifiable people (public figures, people named in the prompt), and any likeness of them;
- blood, gore, torture, self-harm; realistic violence against people or animals;
- hate, harassment, extremist symbols;
- realistic illegal activity;
- well-known copyrighted characters and brand logos.
Fantasy and game scenes with stylised action are fine: knights fighting, explosions without victims, monsters, space battles, landscapes, objects, animals, characters you invent.
A refused prompt gets a 400 with code prompt_refused and the reason. If
the check itself cannot answer, the request is refused with a 503 moderation_unavailable, never let through: try again a minute later.
When the clip is made, and before it is kept or charged, the same kind of
check looks at one frame for every second of it. A clip that fails is
deleted and costs nothing: the job ends failed with code output_refused.
If that check cannot answer either, the clip is not delivered.
Every clip that is delivered is also marked as made by AI, inside the file,
for article 50 of the EU AI Act: the MP4's comment and description say so.
Editors and social networks can strip them, so where you publish a clip, say
it in words too: AI-generated media.
Which checks run is decided zone by zone by whoever runs the machines. Today both run everywhere for video.
What we keep
- The prompt while the job runs; it is deleted when the job ends, and a finished job answers with
"prompt": null. It is kept only if you turned on «keep messages» for that key in the dashboard: see Sovereignty. - The clip, for 7 days, in AI Tokens's storage.
- The usage: model, seconds, cost and time, as for every request.
Price
Per second of video, charged when the clip is completed, at the price in Pricing: in CHF, or in credits where AI Tokens works with credits. A failed job, and a clip refused by the check, cost nothing.
Errors when creating
| Status | error.code |
When |
|---|---|---|
| 400 | unknown_parameter |
a field this API does not have; error.param names it |
| 400 | missing_required_parameter |
no prompt, or an empty one |
| 400 | string_above_max_length |
a prompt longer than 2000 characters |
| 400 | model_not_found |
a model that does not make video here |
| 400 | invalid_value |
a seconds, size or seed the model cannot take |
| 400 | unsupported_parameter |
input_reference, for now |
| 400 | prompt_refused |
the prompt goes against the list above |
| 401 | invalid_api_key |
missing or wrong key |
| 429 | too_many_video_jobs |
you already have a job queued or in progress |
| 503 | no_engine_available |
no machine is making this model right now |
| 503 | moderation_unavailable |
the prompt could not be checked |
Errors have OpenAI's shape:
{"error": {"message": "'wan-2.2-5b-fast' makes clips of 5 seconds.", "type": "invalid_request_error", "param": "seconds", "code": "invalid_value"}}
Examples
Python, with requests
import os, time, requests
API = "https://api.aitokens.ch/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['API_KEY']}"}
def make_video(prompt: str, path: str, **options) -> dict:
r = requests.post(f"{API}/videos", headers=HEADERS, json={"prompt": prompt, **options}, timeout=60)
if r.status_code != 200:
raise RuntimeError(r.json()["error"]["message"])
job = r.json()
while job["status"] in ("queued", "in_progress"):
time.sleep(2)
job = requests.get(f"{API}/videos/{job['id']}", headers=HEADERS, timeout=30).json()
print(f"{job['status']} {job['progress']}%")
if job["status"] == "failed":
raise RuntimeError(job["error"]["message"])
with requests.get(f"{API}/videos/{job['id']}/content", headers=HEADERS, stream=True, timeout=300) as clip:
clip.raise_for_status()
with open(path, "wb") as f:
for chunk in clip.iter_content(1 << 16):
f.write(chunk)
return job
make_video("Isometric view of a cozy fantasy tavern, fireplace flickering, warm light", "taverna.mp4")
JavaScript (Node 18+)
// volpe.mjs. Run: API_KEY=sk-… node volpe.mjs (.mjs: it is an ES module)
import { writeFile } from "node:fs/promises";
const API = "https://api.aitokens.ch/v1";
const headers = { Authorization: `Bearer ${process.env.API_KEY}` };
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
let res = await fetch(`${API}/videos`, {
method: "POST",
headers: { ...headers, "Content-Type": "application/json" },
body: JSON.stringify({ prompt: "A low-poly red fox runs through an autumn forest, side view", size: "1280x704" }),
});
let job = await res.json();
if (!res.ok) throw new Error(job.error.message);
while (job.status === "queued" || job.status === "in_progress") {
await sleep(Number(res.headers.get("openai-poll-after-ms") ?? 2000));
res = await fetch(`${API}/videos/${job.id}`, { headers });
job = await res.json();
}
if (job.status === "failed") throw new Error(job.error.message);
const clip = await fetch(`${API}/videos/${job.id}/content`, { headers });
if (!clip.ok) throw new Error((await clip.json()).error.message);
await writeFile("volpe.mp4", Buffer.from(await clip.arrayBuffer()));
C# in Unity
A coroutine that asks for a clip, waits for it, saves it under
Application.persistentDataPath and plays it with a VideoPlayer. Unity 2022
LTS or Unity 6. The key is read from UserSettings/api_key.txt, next to
Assets/: outside Assets, so it never goes into a build, and Unity's usual
.gitignore leaves it out of git. For a game you ship, put your own server in
between: a key inside a build can be read by anyone who has the build.
// GeneraClip.cs: asks for a clip, waits for it, saves it and plays it. Unity 2022 LTS or Unity 6.
using System.Collections;
using System.IO;
using System.Text;
using UnityEngine;
using UnityEngine.Networking;
using UnityEngine.Video;
public class GeneraClip : MonoBehaviour
{
public VideoPlayer player;
const string Api = "https://api.aitokens.ch/v1";
[System.Serializable] class Job { public string id; public string status; public int progress; public Err error; }
[System.Serializable] class Err { public string code; public string message; }
[System.Serializable] class Request { public string prompt; public string size; }
// The key, from UserSettings/api_key.txt next to Assets/: never in a build, never in git.
static string Key() =>
File.ReadAllText(Path.Combine(Application.dataPath, "..", "UserSettings", "api_key.txt")).Trim();
public IEnumerator Make(string prompt)
{
string key = Key();
string body = JsonUtility.ToJson(new Request { prompt = prompt, size = "1280x704" });
Job job;
using (var post = new UnityWebRequest($"{Api}/videos", "POST"))
{
post.uploadHandler = new UploadHandlerRaw(Encoding.UTF8.GetBytes(body));
post.downloadHandler = new DownloadHandlerBuffer();
post.SetRequestHeader("Content-Type", "application/json");
post.SetRequestHeader("Authorization", $"Bearer {key}");
yield return post.SendWebRequest();
if (post.result != UnityWebRequest.Result.Success) { Debug.LogError(post.downloadHandler.text); yield break; }
job = JsonUtility.FromJson<Job>(post.downloadHandler.text);
}
while (job.status == "queued" || job.status == "in_progress")
{
yield return new WaitForSeconds(2);
using (var get = UnityWebRequest.Get($"{Api}/videos/{job.id}"))
{
get.SetRequestHeader("Authorization", $"Bearer {key}");
yield return get.SendWebRequest();
if (get.result != UnityWebRequest.Result.Success) { Debug.LogError(get.downloadHandler.text); yield break; }
job = JsonUtility.FromJson<Job>(get.downloadHandler.text);
}
}
if (job.status != "completed") { Debug.LogError(job.error.message); yield break; }
string path = Path.Combine(Application.persistentDataPath, $"{job.id}.mp4");
using (var file = UnityWebRequest.Get($"{Api}/videos/{job.id}/content"))
{
file.SetRequestHeader("Authorization", $"Bearer {key}");
file.downloadHandler = new DownloadHandlerFile(path);
yield return file.SendWebRequest();
if (file.result != UnityWebRequest.Result.Success)
{
// A file download handler has no .text: the error body went into the file.
Debug.LogError($"{file.responseCode}: {file.error}");
File.Delete(path);
yield break;
}
}
player.url = path; // a local path works; the VideoPlayer switches to its URL source
player.Play();
}
}
This code was not run in Unity: we compiled it against stand-ins with the
signatures of Unity's API, and made the same requests from Python and Node. A DownloadHandlerFile writes an
error answer into the file and has no .text to read it back, which is why a
failed download deletes the file.
With OpenAI's SDK
OpenAI's SDKs have the same calls: client.videos.create_and_poll(...) waits by
itself and client.videos.download_content(id) returns the file.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.aitokens.ch/v1", api_key=os.environ["API_KEY"])
video = client.videos.create_and_poll(model="wan-2.2-5b-fast", prompt="A medieval castle on a hill at dawn, mist in the valley")
if video.status == "completed":
client.videos.download_content(video.id).write_to_file("castello.mp4")
OpenAI's SDKs mark these methods as deprecated, with the message «The Sora API
is scheduled to permanently shut down on September 24, 2026.». They still work
against our API — each call prints that DeprecationWarning — but a future
version of the SDK may remove them. For code you want to keep, use plain HTTP as
in the examples above, or pin the SDK version: we used openai 3.24.0.
Our extensions go in extra_body: extra_body={"seed": 42}.
Many clips, one after the other
Since each person makes one clip at a time, a storyboard is a loop: the next
request starts when the previous clip is done. make_video is the function
from the Python example.
scenes = [
"A knight in silver armour walks through a medieval village at sunset, camera following from behind",
"The knight stops in front of a tavern door, warm light from inside",
"Inside the tavern, people drinking at wooden tables, fireplace flickering",
]
for i, scene in enumerate(scenes, 1):
make_video(scene, f"scena-{i:02d}.mp4", seed=42)
The seed makes each scene repeatable: run the loop again and you get the same
clips. It does not keep a character the same from one scene to the next.