Skip to content

Integrations

LlamaIndex

LlamaIndex in Python with AI Tokens: OpenAILike with the base URL, the context window and function calling switched on.

Last updated: 2026-10-11

bash
pip install llama-index-llms-openai-like
export API_KEY="sk-at-…"

OpenAILike is LlamaIndex's class for OpenAI-compatible APIs. The OpenAI class checks model names against OpenAI's list and does not know AI Tokens's.

Chat and tools

python
import os

from llama_index.core.llms import ChatMessage
from llama_index.core.tools import FunctionTool
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    api_base="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="qwen3.8-27b",
    is_chat_model=True,              # use /v1/chat/completions
    is_function_calling_model=True,  # let agents and tools use it
    context_window=262144,           # the model's, from the table below
    max_tokens=2048,
)

resp = llm.chat([ChatMessage(role="user", content="In one sentence: what is the Rhaetian Railway?")])
print(resp.message.content)


def get_weather(city: str) -> str:
    """Current weather for a city."""
    return "14 °C, light rain"


tool = FunctionTool.from_defaults(fn=get_weather)
resp = llm.chat_with_tools([tool], user_msg="What is the weather in Lugano?")
print(llm.get_tool_calls_from_response(resp, error_on_no_tool_call=False))

What came back on qwen3.8-27b, shortened:

text
The Rhaetian Railway (Rhätische Bahn, RhB) is a Swiss narrow-gauge (1,000 mm) railway network …
[ToolSelection(tool_id='chatcmpl-tool-…', tool_name='get_weather', tool_kwargs={'city': 'Lugano'})]

The three settings

Setting Why
is_chat_model=True without it, OpenAILike calls /v1/completions, which passes the prompt as one chat message
is_function_calling_model=True without it, LlamaIndex's agents treat the model as one that cannot call functions
context_window LlamaIndex fits retrieved text and chat history into it; the default, 3,900 tokens, is far below what the models here read

Context windows: 262,144 for qwen3.8-27b, 32,768 for qwen3.6-35b, qwen3-coder-30b, gemma-4-26b and ministral-3-14b, 16,384 for apertus-70b-instruct. All of them in Which model.

Reasoning off

On qwen3.8-27b, pass the field in additional_kwargs:

python
llm = OpenAILike(
    api_base="https://api.aitokens.ch/v1",
    api_key=os.environ["API_KEY"],
    model="qwen3.8-27b",
    is_chat_model=True,
    context_window=262144,
    additional_kwargs={"reasoning_effort": "none"},
)

Asked for the four national languages of Switzerland, this answered German, French, Italian, Romansh with 9 output tokens.

Limits

  • max_tokens up to 8192. A higher value is refused with 422.
  • Embeddings: /v1/embeddings answers with bge-m3 and takes at most 32 texts per request. See Embeddings.

Search the docs

Type to search…