Skip to content

Cookbook

Streaming in TypeScript

A chat completion read as it arrives, in TypeScript: a parser that tells a complete answer from a cut one, an error and a broken stream, a Next.js route that keeps the key on the server, a React component, and the same with the OpenAI SDK.

Last updated: 2026-10-06

A parser in plain TypeScript, a Next.js route that keeps your key on the server, and a React component that shows the answer as it arrives. The format itself is described in Streaming.

What the code has to handle

See the stream first:

bash
curl -N https://api.aitokens.ch/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "stream": true,
    "messages": [{"role": "user", "content": "Count from 1 to 5."}]
  }'

-N turns off curl's buffering. What arrives, shortened:

text
data: {"id":"req_…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"req_…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"content":"1, 2, 3"},"finish_reason":null}]}

data: {"id":"req_…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Six facts shape the code below:

  1. Events. Each one is a data: line followed by an empty line. The text is in choices[0].delta.content; the stream ends with data: [DONE].

  2. No event arrives until a machine starts answering. While the request waits for one, the stream is silent: do not put a short timeout on the first event.

  3. How it ended. The last chunk before [DONE] carries finish_reason:

    • stop: the answer is complete;
    • length: the answer was cut at the token cap, max_tokens or max_completion_tokens. The stream is complete, the answer is not;
    • tool_calls: the model asks you to run a function. The pieces of each call arrive in delta.tool_calls;
    • error: the answer was interrupted.

    A [DONE] without a finish_reason, or with a value not in this list, means you cannot tell how it ended: treat the answer as incomplete.

  4. Errors come in two forms. What is wrong before the stream starts (a parameter, the key, the model, a rate limit, no machine in your zones) is an ordinary JSON response with an error status. Once the stream has started the status is already 200, and an error arrives as an event, with no [DONE] after it:

    text
    data: {"error":{"message":"Upstream backend unavailable. Retry shortly.","type":"bad_gateway","request_id":"req_…"}}
    
  5. Broken streams. A stream that ends without [DONE], because the connection closed or broke, is an incomplete answer.

  6. Token counts. Ask for them with stream_options: { include_usage: true }: the last chunk before [DONE] has an empty choices and the usage. Read it with if (chunk.usage) { … }, before you look at chunk.choices[0], which that chunk does not have.

Every chunk carries the id of the request, and the response the same value in X-Request-Id: if something goes wrong, quote it to info@daikolab.ch.

A parser

Plain TypeScript, for Node 20 or later and for browsers. It calls onText for every piece of text, and resolves with how the stream ended: one of five outcomes, each with the text received so far. It does not throw for anything the server does; it throws only when you abort the request yourself, so that an abort is never mistaken for an answer.

typescript
// lib/read-chat-stream.ts
export type ToolCall = {
  id: string;
  type: "function";
  function: { name: string; arguments: string };
};

export type StreamOutcome =
  | { kind: "stop"; text: string } // complete
  | { kind: "length"; text: string } // cut at the token cap: incomplete
  | { kind: "tool_calls"; text: string; toolCalls: ToolCall[] } // run them, then ask again
  | { kind: "error"; text: string; message: string; status?: number; requestId?: string }
  | { kind: "incomplete"; text: string; reason: string }; // no [DONE], or no clear end

export async function readChatStream(
  res: Response,
  onText: (text: string) => void = () => {},
): Promise<StreamOutcome> {
  const requestId = res.headers.get("x-request-id") ?? undefined;

  if (!res.ok || !res.body) {
    // An error before the stream: ordinary JSON. The 429 of an endpoint's
    // limit has only a top-level `message`, the other errors an `error` object.
    const body = await res.json().catch(() => null);
    return {
      kind: "error",
      text: "",
      message: body?.error?.message ?? body?.message ?? `HTTP ${res.status}`,
      status: res.status,
      requestId: body?.error?.request_id ?? requestId,
    };
  }

  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = "";
  let text = "";
  let finishReason: string | null = null;
  const calls: ToolCall[] = [];

  for (;;) {
    let read: ReadableStreamReadResult<Uint8Array>;
    try {
      read = await reader.read();
    } catch (e) {
      if ((e as Error).name === "AbortError") throw e; // you stopped it
      return { kind: "incomplete", text, reason: `connection lost: ${(e as Error).message}` };
    }
    const { value, done } = read;
    buffer += done ? decoder.decode() + "\n" : decoder.decode(value, { stream: true });

    const lines = buffer.split("\n");
    buffer = lines.pop() ?? ""; // a line without its end waits for the next read

    for (const line of lines) {
      if (!line.startsWith("data:")) continue; // empty lines between events
      const data = line.slice(5).trim();

      if (data === "[DONE]") return ending(finishReason, text, calls, requestId);

      const event = JSON.parse(data);
      if (event.error) {
        return {
          kind: "error",
          text,
          message: event.error.message,
          requestId: event.error.request_id ?? requestId,
        };
      }

      const choice = event.choices?.[0];
      if (choice?.delta?.content) {
        text += choice.delta.content;
        onText(choice.delta.content);
      }
      // A function call arrives in pieces: the name first, the arguments in fragments.
      for (const piece of choice?.delta?.tool_calls ?? []) {
        const call = (calls[piece.index ?? 0] ??= { id: "", type: "function", function: { name: "", arguments: "" } });
        if (piece.id) call.id = piece.id;
        if (piece.function?.name) call.function.name = piece.function.name;
        if (piece.function?.arguments) call.function.arguments += piece.function.arguments;
      }
      if (choice?.finish_reason) finishReason = choice.finish_reason;
    }

    if (done) return { kind: "incomplete", text, reason: "the stream ended without [DONE]" };
  }
}

function ending(finishReason: string | null, text: string, calls: ToolCall[], requestId?: string): StreamOutcome {
  switch (finishReason) {
    case "stop":
      return { kind: "stop", text };
    case "length":
      return { kind: "length", text };
    case "tool_calls":
      return { kind: "tool_calls", text, toolCalls: calls.filter(Boolean) };
    case "error":
      return { kind: "error", text, message: "The answer was interrupted (finish_reason: error).", requestId };
    case null:
      return { kind: "incomplete", text, reason: "[DONE] arrived without a finish_reason" };
    default:
      return { kind: "incomplete", text, reason: `unexpected finish_reason: ${finishReason}` };
  }
}

To try it from a terminal, put the function and these lines in one file, main.mts, and run API_KEY=sk-… npx tsx main.mts:

typescript
const res = await fetch("https://api.aitokens.ch/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "qwen3.8-27b",
    stream: true,
    max_completion_tokens: 400,
    messages: [{ role: "user", content: "Count from 1 to 5." }],
  }),
  signal: AbortSignal.timeout(120_000), // your own time limit
});

const outcome = await readChatStream(res, (text) => process.stdout.write(text));
switch (outcome.kind) {
  case "stop":
    console.log();
    break;
  case "length":
    console.error("\n[cut at the token cap: raise max_completion_tokens or ask for less]");
    break;
  case "tool_calls":
    console.error("\n[the model asks for functions]", outcome.toolCalls);
    break;
  case "error":
    console.error(`\n[error${outcome.status ? ` ${outcome.status}` : ""}: ${outcome.message}, request ${outcome.requestId}]`);
    process.exitCode = 1;
    break;
  case "incomplete":
    console.error(`\n[incomplete answer: ${outcome.reason}]`);
    process.exitCode = 1;
    break;
}

tool_calls comes back only if you sent tools. Run the functions, add their results to the conversation as messages with role: "tool", and ask again: see Structured output and function calling.

Next.js: the key stays on the server

The browser must never see your key. A route handler calls the API with it and passes the stream on unchanged:

typescript
// app/api/chat/route.ts
export async function POST(req: Request) {
  const { messages } = await req.json();

  const upstream = await fetch("https://api.aitokens.ch/v1/chat/completions", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ model: "qwen3.8-27b", stream: true, messages }),
    signal: req.signal, // abort this call when the browser's request is aborted
  });

  if (!upstream.ok || !upstream.body) {
    // An error before the stream: pass it on as it is, Retry-After included.
    const headers = new Headers({
      "Content-Type": upstream.headers.get("Content-Type") ?? "application/json",
    });
    const retryAfter = upstream.headers.get("Retry-After");
    if (retryAfter) headers.set("Retry-After", retryAfter);
    return new Response(await upstream.text(), { status: upstream.status, headers });
  }

  return new Response(upstream.body, {
    headers: {
      "Content-Type": "text/event-stream",
      "Cache-Control": "no-cache, no-transform",
      "X-Accel-Buffering": "no",
    },
  });
}

Cache-Control: no-cache, no-transform and X-Accel-Buffering: no are the headers the API itself sends, so that proxies pass the stream on as it comes instead of holding it back. Keep them if a proxy such as nginx sits in front of your app.

React: the answer as it arrives

The parser above in lib/read-chat-stream.ts, and a component that uses it. Every outcome but stop leaves a note under the text received so far:

tsx
"use client";
// app/chat.tsx

import { useEffect, useRef, useState } from "react";
import { readChatStream } from "@/lib/read-chat-stream";

export function Chat() {
  const [answer, setAnswer] = useState("");
  const [note, setNote] = useState("");
  const [busy, setBusy] = useState(false);
  const current = useRef<AbortController | null>(null);

  // Leaving the page stops the answer.
  useEffect(() => () => current.current?.abort(), []);

  async function ask(question: string) {
    current.current?.abort(); // a new question stops the previous answer
    const controller = new AbortController();
    current.current = controller;

    setAnswer("");
    setNote("");
    setBusy(true);
    try {
      const res = await fetch("/api/chat", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ messages: [{ role: "user", content: question }] }),
        signal: controller.signal,
      });
      const outcome = await readChatStream(res, (text) => setAnswer((a) => a + text));
      switch (outcome.kind) {
        case "stop":
          break;
        case "length":
          setNote("The answer was cut short: it is incomplete.");
          break;
        case "tool_calls":
          setNote("The model asked for a function this page does not run.");
          break;
        case "error":
          setNote(`Error: ${outcome.message}`);
          break;
        case "incomplete":
          setNote(`Incomplete answer: ${outcome.reason}`);
          break;
      }
    } catch (e) {
      // Network errors, and aborts: an abort started by a new question says nothing.
      if (!controller.signal.aborted) setNote(`Incomplete answer: ${(e as Error).message}`);
    } finally {
      if (current.current === controller) setBusy(false);
    }
  }

  return (
    <div>
      <button onClick={() => ask("Count from 1 to 10.")} disabled={busy}>
        Ask
      </button>
      <pre>{answer}</pre>
      {note && <p role="alert">{note}</p>}
    </div>
  );
}

Stopping

A stream is stopped by closing the connection; it cannot be resumed, and a new request starts from the beginning. That is why the component aborts its request when a new question comes or the page goes away: the gateway checks, at every piece of text, whether the connection is still open, and stops relaying the answer when it is not.

With the OpenAI SDK

If you use openai-node, the same rules apply. The SDK throws an APIError when an error event arrives; a stream that stops early simply ends, with no finish_reason. So check how it ended:

typescript
// sdk.mts. npm install openai, then run: API_KEY=sk-… npx tsx sdk.mts
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.aitokens.ch/v1", apiKey: process.env.API_KEY });

const stream = await client.chat.completions.create({
  model: "qwen3.8-27b",
  stream: true,
  messages: [{ role: "user", content: "Count from 1 to 5." }],
});

let finishReason: string | null = null;
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  finishReason = chunk.choices[0]?.finish_reason ?? finishReason;
}

if (finishReason === "length") console.error("\n[cut at the token cap: the answer is incomplete]");
else if (finishReason !== "stop") throw new Error(`Incomplete answer (finish_reason: ${finishReason}).`);

With tools, a finish_reason of tool_calls is a third good ending: the SDK's client.chat.completions.stream(…) helper collects the pieces of the calls for you.

Search the docs

Type to search…