Discord
此页面暂无中文版本 — 以下显示英文内容。 查看英文版

Provider selection

If you don't want to bother with provider selection

Skip this whole page. Send just model and prompt — no provider field at all — and we route to the cheapest host currently healthy for that model automatically. If it rejects the job, we retry once on the same host, then walk every other host serving that model, cheapest first, until one accepts. Everything below only matters once you want to pin, exclude, or hedge across hosts yourself.

# no "provider" field at all — routes to the cheapest healthy host
curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "a paper airplane gliding over a city"
  }'

Fields at a glance

The provider object can contain the following fields:

Field Type Default Description
only string[] — Allow-list of provider slugs to restrict candidates to — any size, not just one (e.g. ["fal", "atlascloud"]). Chat, video, image. Learn more
ignore string[] — Block-list of provider slugs to exclude. Chat, video, image. Learn more
order string[] — Priority sequence, not a restriction — the first listed slug with a healthy/priceable candidate wins, e.g. ["together", "deepinfra"]. Chat, video, image. Learn more
sort string — "price", "latency", "reliability", or "queue" (video/image only — chat instead accepts "throughput", not "reliability"/"queue") — deterministic pick instead of the default order. Chat, video, image. Learn more
policy string — Named shorthand for a common sort choice: "lowest_cost", "fast_finish", "most_reliable", or "fast_start". An explicit sort on the same request always wins over what policy implies. Video, image only. Learn more
allow_fallbacks boolean true false disables retries within the named model's own provider candidates — zero retries of any kind there, not just cross-candidate fallback. Scoped to provider, not models[]: chat, video, and image all still cascade to models[] on a retryable error even with this set — each named model just gets exactly one provider attempt instead of its normal fallback pool. Chat, video, image. Learn more
max_price object — {prompt, completion} ceiling in $/1M tokens — deployments above it are excluded. Chat only. Learn more
require_parameters boolean false Drop candidates that can't satisfy params your request actually carries (e.g. tools, JSON mode). Chat only. Learn more
failover.on_timeout_sec number — Separate failover object, not part of provider. Hedges to the next-cheapest untried host if the job/response hasn't finished by this deadline, without cancelling the original — both attempts billed if they land. Video, image only. Learn more
preferences object — Nested field — provider: {"preferences": {max_p95_latency_ms, min_success_rate}} — hard-filters hosts by measured 24h performance. On video/image this filters hosts of the one model you named; on chat it instead ranks which model "auto" picks on /v1/route — same field shape, different layer. Chat, video, image. Learn more

Video generation

A caller almost always names one model — the provider it actually runs on is what this section controls. POST /v1/videos shares the same provider field names as chat below (only/ignore/order/allow_fallbacks/sort), but the default selection rule is different: deterministic cheapest-first instead of chat's price-weighted random pick — no max_price/require_parameters here either, and sort takes "price"/"latency"/"reliability"/"queue" instead of chat's "throughput". Video/image also accept a named policy shorthand for the four — see below.

Pin to an allowed host set — fail fast

only isn't limited to naming a single host — it's an allow-list, so provider: {"only": [...], "allow_fallbacks": false} can list several trusted hosts at once. With allow_fallbacks: false, only the cheapest surviving candidate from that list is ever attempted — a rejected submission returns as an error immediately, with no same-host retry and none of the other hosts in the list tried either. Same semantics as chat's allow_fallbacks: false below: zero retries of any kind, not just cross-candidate fallback — "fail fast" means one attempt total for this named model, even when the allow-list has more than one entry. models[] is a separate knob and still gets consulted if this one host fails — see Model fallbacks.

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "a paper airplane gliding over a city",
    "provider": {"only": ["fal", "atlascloud"], "allow_fallbacks": false}
  }'

Automatic provider selection

The default is free and always on, no parameter needed: a rejected submission retries once on the same host, then walks every other host confirmed to serve that exact checkpoint cheapest first, until one accepts the job. "Cheapest first" is just provider.sort defaulting to "price" — set it explicitly to reorder the exact same walk by something other than price, without changing any of the retry/fallback mechanics above. This only covers the submission call itself, before a job id exists — a job that's already accepted and then stalls or fails needs the opt-in hedge below instead. Narrow which hosts are eligible without disabling fallback via provider.only/ignore (no allow_fallbacks: false) — the walk still runs, just over that narrowed set:

# no "provider" field at all — same as provider: {"sort": "price"}, the default
curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "a paper airplane gliding over a city"
  }'

provider.sort also accepts "latency" (re-ranks the walk by each host's measured p50 generation time instead of price), "reliability" (re-ranks by measured 24h success rate, highest first), and "queue" (re-ranks by measured avg queue time — how long a host makes you wait before generation even starts, distinct from "latency"'s total generation time) — same field as chat's own sort below, four values instead of chat's price/latency/throughput (media has no separate throughput signal, but does have measured reliability/queue signals chat's own sort doesn't expose yet). Queue-transition timestamps come from each provider's own poll response, not our own wall-clock measurement the way generation time and success/failure are — treat "queue" as a softer signal than the other two. Any host with no measured data for the chosen sort yet sorts to the very back, after every host that has data — never rewarded or punished for a signal we don't actually have:

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "a paper airplane gliding over a city",
    "provider": {"sort": "reliability"}
  }'
curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "a paper airplane gliding over a city",
    "provider": {"sort": "queue"}
  }'

provider.order — same field, same semantics as chat's below — is a priority sequence, not a re-sort: listed hosts jump to the front in that exact order, everything else keeps falling back behind them in whatever sort already put them in:

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "a paper airplane gliding over a city",
    "provider": {"order": ["wavespeed", "replicate"]}
  }'

Response body tells you which host actually took it: provider, fallback_used, fallbacks_tried. To prefer one host over the rest without disabling fallback, append it to the model id instead of writing a provider object: "model": "minimax/h3/fal" tries Fal first, still falls back to the rest of the pool if Fal itself errors.

{
  "id": "video_...",
  "status": "processing",
  "provider": "wavespeed",
  "fallback_used": true,
  "fallbacks_tried": 1
}

Named policies: lowest_cost / fast_finish / most_reliable / fast_start

provider.policy is a named shorthand over the same sort values above, for when you'd rather say what you want than know which field controls it:

policy: "lowest_cost"same as today's default — cheapest first, no re-ranking
policy: "fast_finish"same as sort: "latency" — lowest measured p50 generation time first
policy: "most_reliable"same as sort: "reliability" — highest measured 24h success rate first
policy: "fast_start"same as sort: "queue" — shortest measured avg queue time first

An explicit sort on the same request always overrides what policy would have implied — policy only fills in sort when you leave it unset, and it composes normally with only/ignore/order:

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "minimax/h3",
    "prompt": "a paper airplane gliding over a city",
    "provider": {"policy": "most_reliable"}
  }'

Latency/reliability filter

provider.preferences — nested inside the same provider object as everything else on this page — hard-filters the candidate pool by each host's actual measured performance over the last 24h: max_p95_latency_ms drops any host whose measured p95 latency exceeds it, min_success_rate (0–1) drops any host below it. A host with too few samples in that window to trust yet is never penalized by either — same "don't punish missing data" rule as the latency/reliability sort above. This is a hard filter (drop hosts below a threshold), not a ranking — combine it with sort: "reliability"/policy: "most_reliable" if you want the survivors ordered by success rate too, not just filtered. Like only/max_price, an over-restrictive threshold degrades to "ignore this filter" rather than failing the request outright. Chat has the same nested field, but it only ranks which model model: "auto" picks on POST /v1/route — a different layer from this, which filters hosts of the one model you already named:

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "a paper airplane gliding over a city",
    "provider": {"preferences": {"max_p95_latency_ms": 60000, "min_success_rate": 0.9}}
  }'

Opt-in: wait, then hedge to a second attempt

For a job that's already been accepted. POST /v1/videos is async and can take minutes, so retrying on your own end just leaves the first one running too. Pass failover.on_timeout_sec: if your job hasn't finished by that deadline, the platform resubmits it to the next-cheapest untried host — without cancelling the original. Both attempts get billed — a second chance at finishing sooner, not a free retry.

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "minimax/h3/fal",
    "prompt": "a paper airplane gliding over a city",
    "failover": {"on_timeout_sec": 60, "max_attempts": 2}
  }'

max_attempts defaults to 2, capped at 3. A background sweep checks for overdue jobs every 30s. The creation response tags hedge_armed: true once one's armed. Poll either job id — GET /v1/videos/{id} transparently returns whichever finishes first, tagged pricing_mode: "timeout_failover_duplicate_billed" once a hedge exists. Cross-model fallback (a genuinely different checkpoint, not just a different host): Model fallbacks. Full guide: Video generation.

Image generation

POST /v1/images shares video's provider/failover behavior byte for byte — same fields (only/ignore/order/sort/policy/allow_fallbacks/preferences/failover.on_timeout_sec), same defaults, same shared fallback-pool code, same provider/fallback_used/fallbacks_tried response fields, same model-id suffix trick. See Video generation above for the full explanation and examples — swap /v1/videos for /v1/images and any video model for an image one.

The one real difference: image generation is synchronous, so there's no job id to poll. A timeout hedge races candidates — the leading and the next-cheapest untried host — inside the single HTTP request/response instead of across separate polls; whichever finishes successfully first is returned, and anything still running keeps going in the background, billed for real if it lands. Cross-model fallback: Model fallbacks. Full guide: Image generation.

Chat completions

Default behavior (no provider)

With no provider object at all, deployment choice among a model's healthy providers follows the same rule OpenRouter documents: deployments with a significant outage in the last ~30 seconds are deprioritized first; if a prior request from your session or key already landed on a deployment (sticky routing, for provider-side prompt-cache hit rate), that one is preferred; otherwise a deployment is drawn at random weighted by the inverse square of price — a $1/M provider is ~9× more likely to be picked than a $3/M one. Setting sort or order disables both the weighting and the stickiness in favor of your explicit choice. This also means every call already gets automatic failover across a model's own healthy deployments for free — if one errors or times out before the first token, the gateway retries the next one automatically:

# no "provider" field at all — this already retries across
# meta-llama/llama-3.3-70b-instruct's own healthy provider deployments
resp = client.chat.completions.create(
    model="meta-llama/llama-3.3-70b-instruct",
    messages=[{"role": "user", "content": "Summarize this contract..."}],
)

The provider object

curl https://videorouter.sh/api/v1/chat/completions \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "content-type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b-instruct",
    "provider": {
      "only": ["together", "deepinfra"],
      "max_price": {"prompt": 0.5, "completion": 0.5},
      "require_parameters": true
    },
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Field Default Effect
allow_fallbackstruefalse restricts the call to a single deployment with no within-model retry
orderunsetordered list of provider names — first match among healthy deployments wins, pinned
only / ignoreunsetallow-list / block-list of provider names
max_priceunset{prompt, completion} ceiling in $/1M tokens — deployments above it are excluded
require_parametersfalsedrop model candidates that can't satisfy params your request actually carries (e.g. tools, JSON mode) — checked before a model is even selected
sortunsetsee below
preferencesunset{max_p95_latency_ms, min_success_rate} — ranks which model "auto" picks on POST /v1/route by measured performance; see Video generation for the same field's different role there

only/ignore/max_price never prune a model down to zero servable deployments — an over-restrictive filter degrades to "ignore that one preference" rather than making an otherwise-healthy model fail your request outright.

sort & order

Setting either one opts out of the default weighted/sticky behavior above in favor of a deterministic pick:

order: ["together", "deepinfra"]first name in the list with a healthy deployment wins, every time
sort: "price"always the cheapest healthy deployment
sort: "latency" / "throughput"lowest measured average latency among deployments we have data for; falls back to cheapest on a cold-start deployment with no samples yet

Filters never hard-fail

A too-narrow only/max_price combination that would exclude every deployment of a model is treated as "ignore this one filter" instead of a hard failure — we'd rather serve your request than enforce a preference into an outage. order and sort still only ever choose among deployments that are actually healthy right now.

Outage cooldown

A deployment that fails repeatedly is cooled down and excluded from selection for a window before being retried — the same "no significant outage in the last ~30s" rule used in the default weighting above applies here too, so a struggling provider stops receiving new traffic well before your request would otherwise time out against it.