Skip to content

Concepts

Knowledge bases (RAG)

Answers from your own documents, with their sources: what happens to a file, how passages are found, what the model is told, and what does not work yet.

Last updated: 2026-10-06

A chat model knows only what was in its training data. With RAG — retrieval-augmented generation — it answers from your documents instead:

  1. Your documents are split into passages, and every passage becomes a vector.
  2. A question becomes a vector too, and the passages closest to it are found.
  3. The best passages go to the model with the question. The model answers from them, and cites them.

Here a set of documents you query together is called a knowledge base.

Where to use it

  • In the dashboard, at https://my.aitokens.ch/dashboard/rag: create a knowledge base, upload files, ask questions. The answer shows its sources. Nothing to install.
  • Through the app API, at https://my.aitokens.ch/api/v1/rag/…, with the token of a signed-in user: everything the dashboard does except deleting a whole knowledge base. A complete run is the RAG quickstart.
  • With an API key, at https://api.aitokens.ch/v1/knowledge-bases: list the knowledge bases you can use and ask them questions. Creating one, uploading and deleting documents need the dashboard or the app API.

Which operation works with which credential is in the table of the Knowledge bases API. To build your own pipeline with a key, the pieces are /v1/embeddings and /v1/rerank.

What happens to a document

Step What happens
Reading PDF: first Docling, which keeps tables and the reading order of columns, and runs OCR on scanned pages. If Docling fails or returns nothing, the text is taken with pdftotext, which does not read scanned pages. DOCX: the text of the main body; headers, footers and notes are not read, and tables become plain text. TXT and Markdown: as they are.
Splitting All spacing, line breaks included, becomes a single space. The text is then cut into passages of at most 2,048 characters, overlapping by 256: at the end of a sentence when possible, otherwise between two words.
Vectors Each passage gets two vectors. A dense one from bge-m3, 1024 dimensions, for meaning. A keyword one: how many times each word appears, leaving out words shorter than 3 characters and the most common Italian, English, German and French words.
Storage The vectors go to a Qdrant collection, one per knowledge base, which weighs each keyword by how rare it is in that knowledge base. The file stays in storage, and the text of the passages is also kept in the database.

Indexing runs in a queue: the upload answers at once, and the document goes from pending to ready — or failed — later.

What happens to a question

  1. The question gets the same two vectors.
  2. Hybrid search. Qdrant takes the 60 passages closest in meaning and the 60 best by keywords, merges the two rankings with Reciprocal Rank Fusion, and keeps the best 30.
  3. Rerank. bge-reranker-v2-m3 reads the question together with each of these passages and scores them. The best top_k stay: 5 unless you ask otherwise, up to 20.
  4. Answer. The passages and the question go to the chat model — qwen3.8-27b unless you choose another — on machines in the zones offered by AI Tokens.

The keyword half is what finds exact words — a name, a product code, an article number — that the dense vector tends to blur into a general idea. The rerank is what puts the passage that really answers first.

If the hybrid search fails, the search runs again by meaning only. If it cannot run at all — the embeddings or the vector database do not answer, or answer with an error — no answer is written: the API answers 503 with code retrieval_unavailable (see Knowledge bases API), the dashboard says the documents could not be searched, and the question is not charged. If the search runs and finds no passage, the model is not asked either: the answer says that the uploaded documents contain nothing about the question.

What the model is told

A fixed instruction, written in Italian, tells the model to answer only from the passages it is given, to say so when the answer is not there, to cite its sources as [file name — part N], and to answer in the language of the question. The passages follow, each under that same label: its file name and its position in the document (chunk_idx, counting from 0). Their scores are not sent to the model. The word for part follows the language of AI Tokens: part in English, parte in Italian.

The answer is at most 800 tokens long, at temperature 0.3. You cannot change the instruction: if you need your own, build the pipeline with /v1/embeddings and /v1/rerank.

The instruction makes answers from outside the documents less likely; it does not rule them out. That is why every answer comes with the passages it was given: check them.

Measured

On 22 September 2026 we measured the pipeline on a five-article mandate contract, two passages long. Indexing took 0.20 s; the answers took 0.57 s on average over three questions, each citing two sources, and all three were right. The answers were written by gemma-4-26b.

The document reached the pipeline without the upload, which was out of service that day because the object storage that receives the files had failed. The object storage was replaced on 3 October 2026. Code and method are in Three examples, with measured numbers.

What does not work yet

  • No settings per knowledge base. Passage size and overlap are the same for every knowledge base and cannot be changed, in the dashboard or in the API. Neither can the number of candidates (30), the instruction, or the length of the answer.
  • Renaming a knowledge base is not possible. Deleting one is possible only from the dashboard, and deletes the files, passages and vectors of all its documents.
  • A failed document is not retried. Only ready documents are searched, so the passages it indexed before failing are never used; they are kept until you delete it.
  • Scanned PDFs depend on Docling. On the pdftotext fallback they come out empty, and the document fails with No text found in the document after parsing.
  • The status page checks the search only. Service status shows whether the vector database, the embeddings and the reranker answer (without the reranker the passages keep the first search's order). It does not tell you whether upload, reading or indexing work: if a document stays pending or ends failed, write to info@daikolab.ch.
  • Old reasons stay as they were written. The error of a failed document is written once, when it fails, and kept. It is in English now; a document that failed before can still show its reason in Italian, such as Documento vuoto dopo parsing.

Where it runs

The answer is generated only on machines in the zones offered by AI Tokens, and is never moved to another zone. Reading the files, the vectors, the search and the rerank run on dedicated machines that do not go through zones yet. If it matters for your case, write to info@daikolab.ch and we tell you where they run for AI Tokens.

When it is the right tool

Good fits:

  • Questions on documents that change slowly: contracts, manuals, regulations, an internal wiki.
  • Finding by meaning and by exact words — codes, names — at the same time.
  • Answers someone has to be able to check: each one comes with its sources.

Poor fits:

  • Live data: give the model a tool to call instead, see OpenAI SDK migration.
  • Changing how the model writes: that is a matter of instructions, or of another model.
  • A few short documents: they may fit in the prompt of an ordinary chat request.

Search the docs

Type to search…