Discord
此页面暂无中文版本 — 以下显示英文内容。 查看英文版

Chat completions

The same API key that calls /v1/videos, /v1/images, and the speech endpoints also serves an OpenAI-compatible POST /v1/chat/completions. Setting model:"auto" lets our router pick the cheapest text model that clears the quality bar for your prompt; pinning an explicit model id skips the router entirely.

import requests

resp = requests.post(
    "https://videorouter.sh/api/v1/chat/completions",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "auto",
        "mode": "balanced",
        "messages": [{"role": "user", "content": "Write a Python web scraper"}],
    },
).json()

mode is one of cost · balanced · quality — it only affects requests routed via "auto" and defaults to your key's configured mode. Full routing behavior — how "auto" scores a prompt and picks a model — is documented in Models & routing.

Use the OpenAI SDK

No new client library needed. Point any existing OpenAI SDK at our base_url and swap the key — everything else (streaming, tool calls, response shape) works unchanged. The OpenAI SDK has no first-class video/image-generation methods against a third-party base URL, so use its low-level client.post/client.get, or plain requests/httpx, for video and image generation instead.

# pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="https://videorouter.sh/api/v1",
    api_key="llmr_sk_live_...",
)

# model:"auto" routes; a slug (e.g. "anthropic/claude-opus-4.8") pins it
resp = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user",
               "content": "Summarize this contract..."}],
)

print(resp.choices[0].message.content)

Same idea in any OpenAI-compatible client (Node's openai package, LangChain, LlamaIndex, etc.) — set the base URL and key, nothing else changes.

Streaming

Video/image generation has no streaming mode; poll /v1/videos/{id} instead. Chat completions support the standard stream: true flag for a standard SSE stream of chat.completion.chunk events, ending in data: [DONE] — identical shape to the OpenAI API.

stream = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Count to 5"}],
    stream=True,
)
for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

Failover only happens before the first token — once bytes are streaming, a mid-stream upstream error surfaces as an SSE error event carrying the request's trace id instead of silently switching providers. Non-streaming requests can fail over across the full response. Chunk shape, usage accounting on a stream, time-to-first-token, and cancellation: Streaming.