Model fallbacks
Model fallback is automatic failover between
models: the top-level models array
lets you automatically try other models if the primary model's providers are down, rate-limited, or
refuse to reply due to content moderation.
Chat completions
Retry, then fall back
The default. When a model has more than one registered provider, the gateway picks one at random for
each request, skewed heavily toward the cheaper ones — a $1/M-token provider is
about 9× more likely to get picked than a $3/M one (details, and how to fail fast instead:
Provider selection).
On failure, it gets one retry, which — since the failed provider is cooled down for 30s — usually lands
on a different one if the model has multiple. Across models, pass a top-level
models array — the same
parameter OpenRouter uses:
curl https://videorouter.sh/api/v1/chat/completions \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "anthropic/claude-opus-4.8",
"models": ["anthropic/claude-sonnet-5", "openai/gpt-4o"],
"messages": [{"role": "user", "content": "Hello!"}]
}'
claude-opus-4.8 gets its own
retry first, then claude-sonnet-5,
then gpt-4o — first one to
succeed wins. Two caps stack: at most 3 distinct models get tried per request
(LLMROUTER_MAX_ATTEMPTS), out of
up to 6 computed into the candidate list
(LLMROUTER_MAX_CANDIDATES, primary +
auto-routing's own fallback ranking + your models array).
With model:"auto", that ranking is
a real list the matcher returns (e.g. gpt-4o-mini →
claude-3-5-haiku → …), not just the
one model you'd get with a pin.
Turning it off: allow_model_fallback
Top-level, sibling to models —
not nested under provider.
Defaults to true (everything
above). Set it false and
models is never even
consulted, no matter how many entries it has — a hard pin to the one model you named, full stop. This
is the flag to reach for if you want
provider: {"allow_fallbacks": false}'s
old behavior back: that field now only disables provider-level retry within one named model (see
Provider selection) —
it doesn't touch this array.
curl https://videorouter.sh/api/v1/chat/completions \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "anthropic/claude-opus-4.8",
"models": ["anthropic/claude-sonnet-5"],
"allow_model_fallback": false,
"messages": [{"role": "user", "content": "Hello!"}]
}'
Streaming caveat
Fallback only happens before the first token. Once bytes are streaming, a mid-stream
upstream error can't be transparently swapped for another model — it surfaces as an SSE
error event carrying the trace id
instead. Full event shape: Streaming.
Seeing it happen
Response headers x-llmrouter-fallbacks (the
ordered candidate list) and x-llmrouter-primary-model
(present only when a fallback actually served) show whether this happened. Full trace —
fallback_used,
fallbacks_tried,
fallback_latency_ms — at
/logs.
Video generation
Cross-model fallback: models[]
Above only ever swaps hosts for the same checkpoint. To fall back to a genuinely different
model if model's entire host
group is exhausted, pass a top-level models
array — same parameter and priority-order semantics as chat's. Each named model gets its own
cheapest-host walk first; only once that model's hosts are exhausted does the request move to
the next name in the list:
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0/fal",
"models": ["bytedance/seedance-2.5/fal"],
"prompt": "a paper airplane gliding over a city"
}'
Up to 5 entries from models
are considered (plus the primary model,
so 6 distinct models total) — unlike chat's separate attempts-vs-candidates split above, there's no
extra cap on how many actually get tried: the walk continues through every host on every considered
model, in order, until one succeeds. provider: {"allow_fallbacks": false}
doesn't disable this array — it restricts each named model (this one and every backup) to a single
provider attempt, chat parity (Provider selection),
never a same-model retry — but the walk still moves on to the next name in the list if that one
attempt fails. To disable the array itself, set top-level
allow_model_fallback: false
instead — see above.
Full guide: Video generation.
Image generation
Same split as video: pinning/preferring/automatically picking which host serves one named model is Provider selection. Below is cross-model fallback.
Cross-model fallback: models[]
Same as video: a top-level models
array falls back to a genuinely different model once model's
own hosts are exhausted — each named model gets its own cheapest-host walk first, in priority order:
curl https://videorouter.sh/api/v1/images \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "openai/gpt-image-1/openai",
"models": ["qwen-image-3.0"],
"prompt": "a paper airplane gliding over a city"
}'
Same cap as video: up to 5 entries from
models (6 distinct models
total including the primary), no separate attempts-vs-candidates split — every host on every
considered model gets tried, in order, until one succeeds. Same
allow_fallbacks: false
caveat as video: it restricts each named model to a single provider attempt, not this array. Same
allow_model_fallback: false
to disable the array itself — see above.
Full guide: Image generation.
What triggers a fallback
The same status set, everywhere — chat, video, and image all check the identical list. Only transient / capacity upstream statuses fail over — a real client-side mistake (bad request, auth) does not, since retrying it elsewhere wouldn't help:
| Upstream status | Meaning |
|---|---|
| 408, 425 | timeout / too-early |
| 409 | conflict |
| 429 | provider-side rate limit |
| 500, 502, 503, 504 | provider server error / unavailable |