Chat completions
The same API key that calls /v1/videos,
/v1/images, and the speech endpoints also
serves an OpenAI-compatible POST /v1/chat/completions.
Setting model:"auto" lets our router pick
the cheapest text model that clears the quality bar for your prompt; pinning an explicit model id skips
the router entirely.
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/chat/completions",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "auto",
"mode": "balanced",
"messages": [{"role": "user", "content": "Write a Python web scraper"}],
},
).json()
mode is one of
cost ·
balanced ·
quality — it only affects requests routed via
"auto" and defaults to your key's configured mode.
Full routing behavior — how "auto" scores a
prompt and picks a model — is documented in Models & routing.
Use the OpenAI SDK
No new client library needed. Point any existing OpenAI SDK at our
base_url and swap the key —
everything else (streaming, tool calls, response shape) works unchanged. The OpenAI SDK has no
first-class video/image-generation methods against a third-party base URL, so use its low-level
client.post/client.get,
or plain requests/httpx,
for video and
image generation instead.
# pip install openai
from openai import OpenAI
client = OpenAI(
base_url="https://videorouter.sh/api/v1",
api_key="llmr_sk_live_...",
)
# model:"auto" routes; a slug (e.g. "anthropic/claude-opus-4.8") pins it
resp = client.chat.completions.create(
model="auto",
messages=[{"role": "user",
"content": "Summarize this contract..."}],
)
print(resp.choices[0].message.content)
Same idea in any OpenAI-compatible client (Node's openai package,
LangChain, LlamaIndex, etc.) — set the base URL and key, nothing else changes.
Streaming
Video/image generation has no streaming mode; poll /v1/videos/{id}
instead. Chat completions support the standard stream: true
flag for a standard SSE stream of chat.completion.chunk events,
ending in data: [DONE] — identical shape to the OpenAI API.
stream = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Count to 5"}],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
Failover only happens before the first token — once bytes are streaming, a mid-stream upstream error surfaces as an SSE error event carrying the request's trace id instead of silently switching providers. Non-streaming requests can fail over across the full response. Chunk shape, usage accounting on a stream, time-to-first-token, and cancellation: Streaming.