How to Call the Wan 3.0 API (Python + curl, Base vs Prime)
Alibaba's Wan 3.0 has one of the widest host rosters of any video model VideoRouter routes — eleven separate providers for the base checkpoint, plus a handful more for the Prime variant. That width is exactly why this model benefits from a dedicated integration guide: with this many hosts, understanding the provider-suffix pattern and knowing which hosts actually carry which checkpoint matters more here than it does for a two- or three-host model.
Text-to-video
import time
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "alibaba/wan-3.0",
"prompt": "a paper boat floating down a rain-slicked street at night",
"resolution": "480p",
"aspect_ratio": "16:9",
"duration_secs": 4,
},
)
job = resp.json()
while job["status"] not in ("completed", "failed"):
time.sleep(5)
job = requests.get(
f"https://videorouter.sh/api/v1/videos/{job['id']}",
headers={"Authorization": "Bearer llmr_sk_live_..."},
).json()
if job["status"] == "completed":
video_url = job["data"][0]["url"]
else:
raise RuntimeError(job["error"])
Equivalent curl:
job=$(curl -s https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_..." \
-H "Content-Type: application/json" \
-d '{"model": "alibaba/wan-3.0", "prompt": "a paper boat floating down a rain-slicked street at night", "resolution": "480p", "duration_secs": 4}')
id=$(echo "$job" | jq -r .id)
status=$(echo "$job" | jq -r .status)
while [ "$status" != "completed" ] && [ "$status" != "failed" ]; do
sleep 5
job=$(curl -s "https://videorouter.sh/api/v1/videos/$id" -H "Authorization: Bearer llmr_sk_...")
status=$(echo "$job" | jq -r .status)
done
echo "$job" | jq -r 'if .status == "completed" then .data[0].url else .error end'
Switching between base Wan 3.0 and Prime
Prime is a separate checkpoint, not a resolution tier — swap the model field to move between them:
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "alibaba/wan-3.0-prime", # or "alibaba/wan-3.0" for the base checkpoint
"prompt": "a paper boat floating down a rain-slicked street at night",
"resolution": "480p",
"duration_secs": 4,
},
).json()
Prime costs roughly double base Wan 3.0 at its own cheapest host — see Wan 3.0 vs Wan 3.0 Prime: Which Should You Call? before defaulting a new integration to it.
Pinning to a specific provider
With eleven hosts carrying base Wan 3.0 and a nearly 4.25x spread between the cheapest and priciest of them, pinning explicitly is a bigger lever here than for most models in this series:
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "alibaba/wan-3.0/replicate", # currently the cheapest verified host
"prompt": "a paper boat floating down a rain-slicked street at night",
"duration_secs": 4,
},
).json()
Confirmed suffixes for base Wan 3.0: replicate, atlascloud, machgen, pika, fal, alibaba (direct), wavespeed, and openrouter. Not every host carries Prime — notably, Replicate and MachGen (two of the cheapest base-model hosts) don't list Prime at all, so a pipeline that pins to replicate for base Wan 3.0 and expects the same suffix to work for alibaba/wan-3.0-prime will get a 400 instead. See Wan 3.0 API Pricing: Every Provider Compared for the full per-host table this maps to.
Handling errors correctly
Every error follows the same envelope shape as the OpenAI API — {"error": {"message", "type", "code"}}:
| Status | type / code | What it means |
|---|---|---|
| 400 | invalid_request_error |
Missing prompt, or a provider suffix that doesn't carry the requested checkpoint (e.g. alibaba/wan-3.0-prime/replicate) |
| 401 | invalid_api_key |
Key missing, malformed, revoked, or expired |
| 402 | spend_cap_exceeded / insufficient_credits |
Monthly cap hit, or prepaid balance ≤ $0 |
| 403 | model_not_allowed |
The requested Wan checkpoint isn't in this key's model_allowlist |
| 429 | rpm_limit / tpm_limit |
Rate limit exceeded — Retry-After tells you how long to wait |
| 500 / 502 / 503 / 504 | upstream_error |
Every candidate host in the fallback chain failed — not billed |
import time
import requests
API_BASE = "https://videorouter.sh/api/v1"
API_KEY = "llmr_sk_live_..."
def generate_wan_video(prompt, prime=False, provider=None, **kwargs):
model = "alibaba/wan-3.0" + ("-prime" if prime else "") + (f"/{provider}" if provider else "")
headers = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
payload = {"model": model, "prompt": prompt, **kwargs}
for attempt in range(3):
resp = requests.post(f"{API_BASE}/videos", headers=headers, json=payload)
if resp.status_code == 429:
time.sleep(int(resp.headers.get("Retry-After", 5)))
continue
resp.raise_for_status()
job = resp.json()
break
else:
raise RuntimeError("exceeded retry budget on job creation")
deadline = time.time() + 600
while job["status"] not in ("completed", "failed"):
if time.time() > deadline:
raise TimeoutError(f"job {job['id']} did not complete within 10 minutes")
time.sleep(5)
job = requests.get(f"{API_BASE}/videos/{job['id']}", headers=headers).json()
if job["status"] == "failed":
raise RuntimeError(f"generation failed: {job['error']}")
return job["data"][0]["url"]
# usage
url = generate_wan_video("a paper boat floating down a rain-slicked street at night", provider="replicate", resolution="480p")
Given eleven hosts carry base Wan 3.0, a genuine upstream_error on an unpinned request is a stronger signal than the same error on a two- or three-host model — either something is affecting several independent providers at once, or (far more likely) you've pinned to a single host and that one specifically is down.
Common mistakes
Pinning to a suffix that doesn't carry Prime. As noted above, replicate and machgen are base-model-only — the equivalent Prime request needs a different suffix (pika, atlascloud, fal, alibaba, or wavespeed).
Assuming "direct from Alibaba" is the cheap option. It isn't — Alibaba's own direct listing ties with Fal and WaveSpeedAI at the median price, not the floor. Replicate currently undercuts it by more than 2x. See Wan 3.0 API Pricing.
Not re-checking the host ranking per resolution tier. Wan 3.0's tier multiplier (roughly 2x per step from 480p → 720p → 1080p) is consistent across hosts, so the cheapest host at 480p stays cheapest at 1080p for this model — but that's not guaranteed to hold for every model, so verify rather than assume when integrating a different one.
See /docs/video-generation for the complete request reference.
Frequently asked questions
Can I get Wan 3.0 output back synchronously? No — generation is async and job-based across every host, the same as every other video model VideoRouter routes. Poll on a 5-10 second interval as shown above.
Does pinning change VideoRouter's own fee, or just the underlying host's rate? Only the underlying host's rate — pinning determines which of the eleven hosts' prices applies to your request; it doesn't change VideoRouter's routing fee structure.
What's the fastest way to check which host served an unpinned request? The completed job object carries provider attribution — check the exact field name in /docs/api-reference-video.
Is there a Wan 2.7 equivalent of this guide? Wan 2.7 is a separate, earlier-generation model with its own narrower host list, not covered here — check VideoRouter's Wan 2.7 model page for its own pricing and hosts rather than assuming this guide's suffixes carry over.
How do I A/B test base Wan 3.0 against Prime fairly? Run both through a host that carries both (Fal or Alibaba direct are the simplest same-host comparison) rather than comparing each at its own cheapest host — see Wan 3.0 vs Wan 3.0 Prime for why isolating the checkpoint as the only variable matters.