Provider selection
If you don't want to bother with provider selection
Skip this whole page. Send just model
and prompt — no
provider field at all — and we route
to the cheapest host currently healthy for that model automatically. If it rejects the job, we retry once on
the same host, then walk every other host serving that model, cheapest first, until one accepts. Everything
below only matters once you want to pin, exclude, or hedge across hosts yourself.
# no "provider" field at all — routes to the cheapest healthy host
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city"
}'
Fields at a glance
The provider object can contain
the following fields:
| Field | Type | Default | Description |
|---|---|---|---|
| only | string[] | — | Allow-list of provider slugs to restrict candidates to — any
size, not just one (e.g. ["fal", "atlascloud"]). Chat, video,
image. Learn more |
| ignore | string[] | — | Block-list of provider slugs to exclude. Chat, video, image. Learn more |
| order | string[] | — | Priority sequence, not a restriction — the first listed slug
with a healthy/priceable candidate wins, e.g.
["together", "deepinfra"]. Chat, video, image.
Learn more |
| sort | string | — | "price",
"latency", "reliability", or
"queue"
(video/image only — chat instead accepts "throughput", not
"reliability"/"queue") —
deterministic pick instead of the default order. Chat, video, image.
Learn more |
| policy | string | — | Named shorthand for a common
sort choice: "lowest_cost",
"fast_finish", "most_reliable", or
"fast_start".
An explicit sort on the same request always wins over what
policy implies. Video, image only.
Learn more |
| allow_fallbacks | boolean | true |
false disables retries within
the named model's own provider candidates — zero retries of any kind there, not just
cross-candidate fallback. Scoped to provider, not
models[]: chat, video, and image all still cascade to
models[] on a retryable error even with this set — each named
model just gets exactly one provider attempt instead of its normal fallback pool. Chat, video,
image. Learn more |
| max_price | object | — | {prompt, completion} ceiling in
$/1M tokens — deployments above it are excluded. Chat only.
Learn more |
| require_parameters | boolean | false |
Drop candidates that can't satisfy params your request
actually carries (e.g. tools, JSON mode). Chat only.
Learn more |
| failover.on_timeout_sec | number | — | Separate failover object, not
part of provider. Hedges to the next-cheapest untried host if
the job/response hasn't finished by this deadline, without cancelling the original — both
attempts billed if they land. Video, image only.
Learn more |
| preferences | object | — | Nested field —
provider: {"preferences": {max_p95_latency_ms,
min_success_rate}} — hard-filters hosts by measured 24h performance. On video/image
this filters hosts of the one model you named; on chat it instead ranks which model
"auto" picks on /v1/route — same
field shape, different layer. Chat, video, image.
Learn more |
Video generation
A caller almost always names one model — the provider it actually runs on is what this section
controls. POST /v1/videos shares
the same provider field names
as chat below (only/ignore/order/allow_fallbacks/sort),
but the default selection rule is different: deterministic cheapest-first instead of chat's price-weighted
random pick — no max_price/require_parameters here either, and
sort takes "price"/"latency"/"reliability"/"queue" instead of chat's "throughput".
Video/image also accept a named policy shorthand for the four — see
below.
Pin to an allowed host set — fail fast
only isn't limited to naming a
single host — it's an allow-list, so provider: {"only": [...], "allow_fallbacks": false}
can list several trusted hosts at once. With allow_fallbacks: false,
only the cheapest surviving candidate from that list is ever attempted — a rejected submission returns
as an error immediately, with no same-host retry and none of the other hosts in the list tried either.
Same semantics as chat's allow_fallbacks: false below:
zero retries of any kind, not just cross-candidate fallback — "fail fast" means one attempt total for
this named model, even when the allow-list has more than one entry.
models[] is a separate knob and
still gets consulted if this one host fails — see
Model fallbacks.
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city",
"provider": {"only": ["fal", "atlascloud"], "allow_fallbacks": false}
}'
Automatic provider selection
The default is free and always on, no parameter needed: a rejected submission retries once on the same
host, then walks every other host confirmed to serve that exact checkpoint cheapest first,
until one accepts the job. "Cheapest first" is just provider.sort
defaulting to "price" — set it
explicitly to reorder the exact same walk by something other than price, without changing any of the
retry/fallback mechanics above. This only covers the submission call itself, before a job id exists — a
job that's already accepted and then stalls or fails needs the opt-in hedge below instead. Narrow which
hosts are eligible without disabling fallback via
provider.only/ignore
(no allow_fallbacks: false) —
the walk still runs, just over that narrowed set:
# no "provider" field at all — same as provider: {"sort": "price"}, the default
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city"
}'
provider.sort also accepts
"latency" (re-ranks the walk by
each host's measured p50 generation time instead of price),
"reliability" (re-ranks by
measured 24h success rate, highest first), and
"queue" (re-ranks by measured
avg queue time — how long a host makes you wait before generation even starts, distinct from
"latency"'s total generation
time) — same field as chat's own sort below,
four values instead of chat's price/latency/throughput (media has no separate throughput signal, but
does have measured reliability/queue signals chat's own
sort doesn't expose yet).
Queue-transition timestamps come from each provider's own poll response, not our own wall-clock
measurement the way generation time and success/failure are — treat
"queue" as a softer signal than
the other two. Any host with no measured data for the chosen sort yet sorts to the very back, after
every host that has data — never rewarded or punished for a signal we don't actually have:
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city",
"provider": {"sort": "reliability"}
}'
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city",
"provider": {"sort": "queue"}
}'
provider.order — same field,
same semantics as chat's below — is a priority
sequence, not a re-sort: listed hosts jump to the front in that exact order, everything else keeps
falling back behind them in whatever sort
already put them in:
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city",
"provider": {"order": ["wavespeed", "replicate"]}
}'
Response body tells you which host actually took it:
provider,
fallback_used,
fallbacks_tried. To prefer one
host over the rest without disabling fallback, append it to the model id instead of writing a
provider object:
"model": "minimax/h3/fal" tries
Fal first, still falls back to the rest of the pool if Fal itself errors.
{
"id": "video_...",
"status": "processing",
"provider": "wavespeed",
"fallback_used": true,
"fallbacks_tried": 1
}
Named policies: lowest_cost / fast_finish / most_reliable / fast_start
provider.policy is a named
shorthand over the same sort
values above, for when you'd rather say what you want than know which field controls it:
| policy: "lowest_cost" | same as today's default — cheapest first, no re-ranking |
| policy: "fast_finish" | same as sort: "latency" — lowest measured p50 generation time first |
| policy: "most_reliable" | same as sort: "reliability" — highest measured 24h success rate first |
| policy: "fast_start" | same as sort: "queue" — shortest measured avg queue time first |
An explicit sort on the same
request always overrides what policy
would have implied — policy only
fills in sort when you leave it
unset, and it composes normally with only/ignore/order:
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "minimax/h3",
"prompt": "a paper airplane gliding over a city",
"provider": {"policy": "most_reliable"}
}'
Latency/reliability filter
provider.preferences — nested
inside the same provider
object as everything else on this page — hard-filters the candidate pool by each host's actual
measured performance over the last 24h:
max_p95_latency_ms drops any
host whose measured p95 latency exceeds it, min_success_rate
(0–1) drops any host below it. A host with too few samples in that window to trust yet is never
penalized by either — same "don't punish missing data" rule as the
latency/reliability sort above. This is a
hard filter (drop hosts below a threshold), not a ranking — combine it with
sort: "reliability"/policy: "most_reliable"
if you want the survivors ordered by success rate too, not just filtered. Like
only/max_price,
an over-restrictive threshold degrades to "ignore this filter" rather than failing the request outright.
Chat has the same nested field, but it only ranks which model
model: "auto" picks on
POST /v1/route — a different
layer from this, which filters hosts of the one model you already named:
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city",
"provider": {"preferences": {"max_p95_latency_ms": 60000, "min_success_rate": 0.9}}
}'
Opt-in: wait, then hedge to a second attempt
For a job that's already been accepted. POST /v1/videos
is async and can take minutes, so retrying on your own end just leaves the first one running too. Pass
failover.on_timeout_sec: if your
job hasn't finished by that deadline, the platform resubmits it to the next-cheapest untried host —
without cancelling the original. Both attempts get billed — a second
chance at finishing sooner, not a free retry.
curl https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "minimax/h3/fal",
"prompt": "a paper airplane gliding over a city",
"failover": {"on_timeout_sec": 60, "max_attempts": 2}
}'
max_attempts defaults to 2, capped
at 3. A background sweep checks for overdue jobs every 30s. The creation response tags
hedge_armed: true once one's
armed. Poll either job id —
GET /v1/videos/{id} transparently
returns whichever finishes first, tagged
pricing_mode: "timeout_failover_duplicate_billed"
once a hedge exists. Cross-model fallback (a genuinely different checkpoint, not just a different
host): Model fallbacks. Full guide:
Video generation.
Image generation
POST /v1/images shares
video's provider/failover behavior byte for
byte — same fields (only/ignore/order/sort/policy/allow_fallbacks/preferences/failover.on_timeout_sec),
same defaults, same shared fallback-pool code, same
provider/fallback_used/fallbacks_tried
response fields, same model-id suffix trick. See Video generation
above for the full explanation and examples — swap /v1/videos
for /v1/images and any video model for an image one.
The one real difference: image generation is synchronous, so there's no job id to poll. A timeout hedge races candidates — the leading and the next-cheapest untried host — inside the single HTTP request/response instead of across separate polls; whichever finishes successfully first is returned, and anything still running keeps going in the background, billed for real if it lands. Cross-model fallback: Model fallbacks. Full guide: Image generation.
Chat completions
Default behavior (no provider)
With no provider object at all, deployment
choice among a model's healthy providers follows the same rule OpenRouter documents: deployments with a
significant outage in the last ~30 seconds are deprioritized first; if a prior request from your session or
key already landed on a deployment (sticky routing, for provider-side prompt-cache hit rate), that one is
preferred; otherwise a deployment is drawn at random weighted by the inverse square of price —
a $1/M provider is ~9× more likely to be picked than a $3/M one. Setting sort
or order disables both the weighting and
the stickiness in favor of your explicit choice. This also means every call already gets automatic failover
across a model's own healthy deployments for free — if one errors or times out before the first token, the
gateway retries the next one automatically:
# no "provider" field at all — this already retries across
# meta-llama/llama-3.3-70b-instruct's own healthy provider deployments
resp = client.chat.completions.create(
model="meta-llama/llama-3.3-70b-instruct",
messages=[{"role": "user", "content": "Summarize this contract..."}],
)
The provider object
curl https://videorouter.sh/api/v1/chat/completions \
-H "Authorization: Bearer llmr_sk_live_..." \
-H "content-type: application/json" \
-d '{
"model": "meta-llama/llama-3.3-70b-instruct",
"provider": {
"only": ["together", "deepinfra"],
"max_price": {"prompt": 0.5, "completion": 0.5},
"require_parameters": true
},
"messages": [{"role": "user", "content": "Hello!"}]
}'
| Field | Default | Effect |
|---|---|---|
| allow_fallbacks | true | false restricts the call to a single deployment with no within-model retry |
| order | unset | ordered list of provider names — first match among healthy deployments wins, pinned |
| only / ignore | unset | allow-list / block-list of provider names |
| max_price | unset | {prompt, completion} ceiling in $/1M tokens — deployments above it are excluded |
| require_parameters | false | drop model candidates that can't satisfy params your request actually carries (e.g. tools, JSON mode) — checked before a model is even selected |
| sort | unset | see below |
| preferences | unset | {max_p95_latency_ms, min_success_rate} — ranks which model "auto" picks on POST /v1/route by measured performance; see Video generation for the same field's different role there |
only/ignore/max_price
never prune a model down to zero servable deployments — an over-restrictive filter degrades to "ignore that
one preference" rather than making an otherwise-healthy model fail your request outright.
sort & order
Setting either one opts out of the default weighted/sticky behavior above in favor of a deterministic pick:
| order: ["together", "deepinfra"] | first name in the list with a healthy deployment wins, every time |
| sort: "price" | always the cheapest healthy deployment |
| sort: "latency" / "throughput" | lowest measured average latency among deployments we have data for; falls back to cheapest on a cold-start deployment with no samples yet |
Filters never hard-fail
A too-narrow only/max_price
combination that would exclude every deployment of a model is treated as "ignore this one filter" instead of
a hard failure — we'd rather serve your request than enforce a preference into an outage.
order and sort
still only ever choose among deployments that are actually healthy right now.
Outage cooldown
A deployment that fails repeatedly is cooled down and excluded from selection for a window before being retried — the same "no significant outage in the last ~30s" rule used in the default weighting above applies here too, so a struggling provider stops receiving new traffic well before your request would otherwise time out against it.