Chat completions
Every route below is for sending model requests, and requires
Authorization: Bearer llmr_sk_<env>_… —
see Authentication. Full request/response detail: Models & routing, Streaming, Tool calling.
POST/v1/chat/completions
Parameters
messagesarrayRequired. Standard OpenAI messages array.modelstringOptional, default "auto". A slug from GET /v1/models pins it; also accepts {author}/{model} ids.streambooleanOptional, default false. SSE stream of chat.completion.chunk events.modestringOptional, default your key's configured mode (else "balanced"). One of cost/balanced/quality — only affects model:"auto" routing.modalitiesarrayOptional. Pass ["image","text"] to get a generated image back on message.images — see image output in chat completions. Pass ["audio","text"] against a gpt-audio model to get spoken audio back on message.audio — see audio output in chat completions.audioobjectRequired when modalities includes "audio". {voice, format} — e.g. {"voice": "alloy", "format": "wav"}.service_tierstringOptional. "flex" for ~50% off at a slower, blocking latency (not compatible with stream). "hour" is rejected — see Async Jobs instead.providerobjectOptional. OpenRouter-compatible routing preferences (only/ignore/order/sort/max_price/quantizations/zdr/data_collection) — see Provider selection.routeobjectOptional. Routing hints for the "auto" path — mode and cost_quality_tradeoff (0–10).Example response (non-streaming)
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1731430000,
"model": "anthropic/claude-opus-4.8",
"choices": [{
"index": 0,
"finish_reason": "stop",
"message": { "role": "assistant", "content": "...", "tool_calls": null, "images": [] }
}],
"usage": { "prompt_tokens": 12, "completion_tokens": 34, "total_tokens": 46, "cost": 0.00041 }
}
Streaming (stream: true) returns a standard SSE stream of chat.completion.chunk events ending in data: [DONE] — see Streaming. Image generation is also reachable here — pass modalities: ["image", "text"] to get a generated image back on message.images, billed the same rate as /v1/images. Spoken audio output works the same way — pass modalities: ["audio", "text"] to a gpt-audio model to get a reply back on message.audio, streaming not supported — see Text-to-speech.
POST/v1/route
Explains the routing decision model:"auto" would make for a given prompt — no upstream model call, no charge.
Parameters
messagesarrayEither this or query is required.querystringShorthand for a single user message — wrapped as [{"role":"user","content":query}].modestringOptional, same default/behavior as on /v1/chat/completions.toolsarrayOptional — factored into the prompt-complexity estimate.Example response
{
"profile": { "difficulty": "medium", "...": "..." },
"profile_source": "complexity_router:medium",
"model": "anthropic/claude-haiku-4.5",
"provider": "anthropic",
"provider_model_id": "claude-haiku-4-5-20251001",
"base_url": null,
"fallbacks": ["openai/gpt-4o-mini", "..."],
"mode": "balanced",
"cost_quality_tradeoff": null,
"reason": "cheapest model clearing the quality bar for this prompt's estimated difficulty",
"relaxed": false
}
GET/v1/models
No parameters. Restricted to your key's model allowlist, if one is set.
Example response
{
"object": "list",
"data": [
{ "id": "auto", "object": "model", "owned_by": "llmrouter" },
{ "id": "anthropic/claude-opus-4.8", "object": "model", "owned_by": "anthropic" },
{ "...": "..." }
]
}