Skip to content

Integrations

LangChain

LangChain in Python with AI Tokens: ChatOpenAI with the base URL, tools, reasoning switched off, and OpenAIEmbeddings with the two settings bge-m3 needs.

Last updated: 2026-10-11

bash
pip install langchain-openai
export API_KEY="sk-at-…"

Chat and tools

python
import os

from langchain_core.tools import tool
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    base_url="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="qwen3.8-27b",
    max_tokens=2048,  # room for the reasoning of the models that reason
)

print(llm.invoke("Name the four national languages of Switzerland, comma-separated.").content)


@tool
def get_weather(city: str) -> str:
    """Current weather for a city."""
    return "14 °C, light rain"


msg = llm.bind_tools([get_weather]).invoke("What is the weather in Lugano?")
print(msg.tool_calls)

What came back on qwen3.8-27b, shortened:

text
German, French, Italian, Romansh
[{'name': 'get_weather', 'args': {'city': 'Lugano'}, 'id': 'chatcmpl-tool-…', 'type': 'tool_call'}]

The first answer used 53 output tokens, most of them reasoning.

Reasoning off

On qwen3.8-27b, reasoning_effort="none" switches reasoning off:

python
quick = ChatOpenAI(
    base_url="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="qwen3.8-27b",
    max_tokens=2048,
    reasoning_effort="none",
)

The same question then used 9 output tokens instead of 53, with the same answer.

Embeddings

python
from langchain_openai import OpenAIEmbeddings

emb = OpenAIEmbeddings(
    base_url="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="bge-m3",
    check_embedding_ctx_length=False,  # send text, not token ids
    chunk_size=32,                     # at most 32 texts per request
)
vectors = emb.embed_documents(["Lugano is in Ticino.", "Chur is in Graubünden."])
print(len(vectors), len(vectors[0]))  # 2 1024

Without those two settings OpenAIEmbeddings sends token ids and up to 1,000 texts per request, and the API refuses both.

Limits

  • Chat completions only. ChatOpenAI moves to the Responses API when you pass use_responses_api=True or OpenAI's built-in tools. AI Tokens's Responses API has no streaming and, on 11 October 2026, did not read tool results sent back to it: keep to chat completions, the default.
  • max_tokens up to 8192. A higher value is refused with 422.
  • usage_metadata has no separate count of reasoning tokens: they are inside output_tokens.

Search the docs

Type to search…