Skip to content

API reference

Knowledge bases API

Create knowledge bases, upload documents, follow their indexing and ask questions answered with their sources: everything with a signed-in user's token, listing and asking also with an API key.

Last updated: 2026-10-06

A knowledge base is a set of your documents: you upload them, we index them, and you ask questions that are answered from them, with the sources. How it works inside is in Knowledge bases (RAG); this page is the reference of the calls. A complete run, from the file to the answer, is in the RAG quickstart.

Who can do what

Operation Dashboard App API, user token API key
List the knowledge bases you can use yes GET /rag/kb: yours and the shared ones GET /v1/knowledge-bases: yours and the shared ones, only those with at least one document ready
Ask a question yes POST /rag/kb/{slug}/chat: yours and the shared ones POST /v1/knowledge-bases/{slug}/ask: yours and the shared ones
Create a knowledge base yes POST /rag/kb no
Upload a document yes POST /rag/kb/{slug}/docs: yours only no
Follow the indexing yes GET /rag/kb/{slug}/docs and /docs/{id}: yours only no
Delete a document yes DELETE /rag/kb/{slug}/docs/{id}: yours only no
Delete a knowledge base yes no no
Rename a knowledge base no no no
Share a knowledge base with everybody no: an administrator of AI Tokens does it no no

A key and a token of the same account see the same knowledge bases. What a key can do on the rest of the API is in What works where.

With an API key

Two calls work with your API key, on your own knowledge bases and on the ones shared with everybody on AI Tokens (for example a product manual, loaded once by an administrator).

bash
curl https://api.aitokens.ch/v1/knowledge-bases \
  -H "Authorization: Bearer $API_KEY"
json
{"object": "list", "data": [{"slug": "course-notes-1a2b3c", "name": "Course notes", "shared": true, "owned": false, "documents": 1}]}
bash
curl https://api.aitokens.ch/v1/knowledge-bases/course-notes-1a2b3c/ask \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"question": "What is the principle of least privilege?"}'
json
{
  "object": "knowledge_base.answer",
  "knowledge_base": "course-notes-1a2b3c",
  "model": "qwen3.8-27b",
  "answer": "…",
  "sources": [{"document": "notes.pdf", "part": 12, "score": 0.83, "text": "…"}],
  "usage": {"prompt_tokens": 1450, "completion_tokens": 160}
}
  • question is required, up to 2,000 characters; model (one of the catalogue, default qwen3.8-27b) and top_k (1 to 20, default 5) are optional. A value outside these gets 422 with the field in errors. The tier is the request's: your key's default, or the X-Siati-Tier header.
  • You pay for your question, also on a shared knowledge base. Unknown or not visible knowledge base: 404 with code knowledge_base_not_found. Documents that could not be searched: 503 with code retrieval_unavailable, and nothing to pay; see Errors.
  • When no document is ready yet, or the search finds nothing, answer is one of the fixed messages of Ask a question, in English, with empty sources and usage at 0: the model is not called and nothing is charged.
  • The list shows only knowledge bases with at least one document ready; documents counts those.
  • The answer is written on machines in the zones offered by AI Tokens. The zone limits of your key do not apply to this call today.
  • ask allows 30 requests a minute per key; the list has only the limit of your key's tier. See Rate limits.
  • A shared knowledge base can be queried by everybody and changed only by whoever created it.

Authentication

The same token as for chat sessions: sign in with the email and the password of the account.

bash
export BASE="https://my.aitokens.ch/api/v1"
export TOKEN=$(curl -s $BASE/auth/login \
  -H "Content-Type: application/json" \
  -d '{"email": "you@example.com", "password": "…"}' | jq -r .access_token)

The account needs a verified email: otherwise every call below answers 403 with code email_unverified. The token lasts 24 hours by default; Signing in explains how to renew it.

Endpoints

All paths start from https://my.aitokens.ch/api/v1.

Method Path What it does Limit per minute
GET /rag/kb Your knowledge bases —
POST /rag/kb Create a knowledge base 30
GET /rag/kb/{slug}/docs The documents of a knowledge base —
POST /rag/kb/{slug}/docs Upload a document 10
GET /rag/kb/{slug}/docs/{id} One document —
DELETE /rag/kb/{slug}/docs/{id} Delete a document —
POST /rag/kb/{slug}/chat Ask a question 30

The limits are counted per signed-in user, on one counter shared with the other limited calls of the app API made with a user token, chat sessions included. Over the limit you get 429, and Retry-After says how long to wait.

A knowledge base cannot be renamed, and it can be deleted only from the dashboard.

Create a knowledge base

bash
curl -s $BASE/rag/kb \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"name": "Contracts 2026", "description": "Client contracts signed in 2026"}'

The answer is 201:

json
{ "id": "…", "slug": "contracts-2026-3f9a1c", "name": "Contracts 2026" }

name is required, up to 200 characters; description is optional, up to 1000. The slug — the name made safe for a URL, plus six random hexadecimal characters — identifies the knowledge base in every other call.

List your knowledge bases

GET /rag/kb, the most recently changed first:

json
{
  "knowledge_bases": [
    {
      "id": "…",
      "slug": "contracts-2026-3f9a1c",
      "name": "Contracts 2026",
      "description": "Client contracts signed in 2026",
      "docs_count": 3,
      "chunks_count": 41,
      "created_at": "2026-10-03T08:00:00+00:00"
    }
  ]
}

docs_count counts the documents that are ready; chunks_count the passages indexed.

Upload a document

bash
curl -s $BASE/rag/kb/contracts-2026-3f9a1c/docs \
  -H "Authorization: Bearer $TOKEN" \
  -F file=@./contract.pdf

The answer is 202:

json
{ "id": "…", "status": "pending", "original_filename": "contract.pdf", "size_bytes": 423812 }
  • Formats: PDF, DOCX, TXT and Markdown. The type is checked on the content of the file, not on its name.
  • Size: up to 50 MB. One file per call; the dashboard takes several at once.
  • 202 means accepted, not usable. Indexing happens later, in a queue. Follow the status of the document until it is ready — or failed.

Follow the indexing

GET /rag/kb/{slug}/docs lists the documents, the most recent first:

json
{
  "documents": [
    {
      "id": "…",
      "original_filename": "contract.pdf",
      "mime_type": "application/pdf",
      "size_bytes": 423812,
      "status": "ready",
      "error": null,
      "chunks_count": 12,
      "ingested_at": "2026-10-03T08:01:10+00:00"
    }
  ]
}

GET /rag/kb/{slug}/docs/{id} returns a single document, with created_at as well.

status goes pending → parsing → chunking → embedding → ready, or stops at failed with the reason in error, as a technical message. A failed document is not retried: delete it, fix the cause, upload it again.

Delete a document

DELETE /rag/kb/{slug}/docs/{id} removes the document's vectors, its passages and the file:

json
{ "deleted": true, "id": "…", "removed_chunks": 12 }

Ask a question

bash
curl -s $BASE/rag/kb/contracts-2026-3f9a1c/chat \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "question": "What is the notice period?",
    "model": "qwen3.8-27b",
    "tier": "medium",
    "top_k": 5
  }'
Field Required Default Notes
question yes Up to 2000 characters.
model no qwen3.8-27b The chat model that writes the answer: a model_id from the catalogue. Another name gets 422.
tier no medium slow, medium, fast or ludicrous: which machines may write the answer. See Tiers.
top_k no 5 How many passages reach the model, from 1 to 20.

The answer, with example values:

json
{
  "answer": "The notice period is six months, to the end of a calendar year [contract.pdf — part 0].",
  "sources": [
    {
      "score": 0.97,
      "dense_score": 0.5,
      "text": "Either party may end the contract with six months of notice, to the end of a calendar year…",
      "document_filename": "contract.pdf",
      "chunk_idx": 0
    }
  ],
  "model": "qwen3.8-27b",
  "tier": "medium",
  "prompt_tokens": 1141,
  "completion_tokens": 97
}
  • sources are the passages given to the model, the best first. score is the score of the reranker, from 0 to 1. dense_score is the score of the first search: another scale, useful to compare passages with each other. If the reranker does not answer, the passages keep the order of the first search and score is the first search's score.
  • chunk_idx is the position of the passage in its document, counting from 0. The answer cites passages as [file name — part N], where N is chunk_idx and the word for part follows the language of AI Tokens. The model writes the citations: check them against sources.
  • The answer is at most 800 tokens long, at temperature 0.3. Neither can be changed, and neither can the instruction: what the model is told is in Knowledge bases (RAG).
  • If no document of the knowledge base is ready yet, the model is not called. answer is then a fixed message, in English on every brand, sources is empty, the tokens are 0, and kb_status says why: processing, all_failed or empty.
  • If the search runs and finds no passage, the model is not called either: answer is The uploaded documents contain nothing about this question., sources is empty, the tokens are 0, and kb_status is no_match. Nothing is charged.

Errors

Status When
401 The token is missing, wrong or expired, or the account is disabled
403 The email of the account is not verified. Code email_unverified
404 The knowledge base or the document does not exist, or it is not yours. The body is {"message": …}, and the message names it
422 Validation failed: name or question missing, a model that is not in the catalogue, a file too large or of a type not accepted, top_k outside 1–20. The body lists the fields in errors
429 Too many requests. Retry-After says how long to wait
500 The answer could not be generated: no machine in the zones of AI Tokens can serve the model right now, or the machines that could did not answer
503 The documents could not be searched: the vector database or the embeddings did not answer, or answered with an error. Code retrieval_unavailable, also with an API key. No answer is written and nothing is charged

Where it runs, and what is kept

The answer is generated only on machines in the zones offered by AI Tokens, as for chat sessions, and is never moved to another zone. Reading the files, the vectors, the search and the rerank run on dedicated machines that do not go through zones yet. If it matters for your case, write to info@daikolab.ch and we tell you where they run for AI Tokens.

Files, passages and vectors are kept until you delete them. Deleting a document removes all three, and so does deleting a knowledge base from the dashboard, for every document in it.

Search the docs

Type to search…