Integrations
LangChain
LangChain in Python with AI Tokens: ChatOpenAI with the base URL, tools, reasoning switched off, and OpenAIEmbeddings with the two settings bge-m3 needs.
Last updated: 2026-10-11
pip install langchain-openai
export API_KEY="sk-at-…"
Chat and tools
import os
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
base_url="https://api.aitokens.ch/v1",
api_key=os.environ["API_KEY"],
model="qwen3.8-27b",
max_tokens=2048, # room for the reasoning of the models that reason
)
print(llm.invoke("Name the four national languages of Switzerland, comma-separated.").content)
@tool
def get_weather(city: str) -> str:
"""Current weather for a city."""
return "14 °C, light rain"
msg = llm.bind_tools([get_weather]).invoke("What is the weather in Lugano?")
print(msg.tool_calls)
What came back on qwen3.8-27b, shortened:
German, French, Italian, Romansh
[{'name': 'get_weather', 'args': {'city': 'Lugano'}, 'id': 'chatcmpl-tool-…', 'type': 'tool_call'}]
The first answer used 53 output tokens, most of them reasoning.
Reasoning off
On qwen3.8-27b, reasoning_effort="none" switches reasoning off:
quick = ChatOpenAI(
base_url="https://api.aitokens.ch/v1",
api_key=os.environ["API_KEY"],
model="qwen3.8-27b",
max_tokens=2048,
reasoning_effort="none",
)
The same question then used 9 output tokens instead of 53, with the same answer.
Embeddings
from langchain_openai import OpenAIEmbeddings
emb = OpenAIEmbeddings(
base_url="https://api.aitokens.ch/v1",
api_key=os.environ["API_KEY"],
model="bge-m3",
check_embedding_ctx_length=False, # send text, not token ids
chunk_size=32, # at most 32 texts per request
)
vectors = emb.embed_documents(["Lugano is in Ticino.", "Chur is in Graubünden."])
print(len(vectors), len(vectors[0])) # 2 1024
Without those two settings OpenAIEmbeddings sends token ids and up to 1,000
texts per request, and the API refuses both.
Limits
- Chat completions only.
ChatOpenAImoves to the Responses API when you passuse_responses_api=Trueor OpenAI's built-in tools. AI Tokens's Responses API has no streaming and, on 11 October 2026, did not read tool results sent back to it: keep to chat completions, the default. max_tokensup to 8192. A higher value is refused with422.usage_metadatahas no separate count of reasoning tokens: they are insideoutput_tokens.