How to Call the Wan 3.0 API (Python + curl, Base vs Prime)

Alibaba's Wan 3.0 has one of the widest host rosters of any video model VideoRouter routes — eleven separate providers for the base checkpoint, plus a handful more for the Prime variant. That width is exactly why this model benefits from a dedicated integration guide: with this many hosts, understanding the provider-suffix pattern and knowing which hosts actually carry which checkpoint matters more here than it does for a two- or three-host model.

Text-to-video

import time
import requests

resp = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "alibaba/wan-3.0",
        "prompt": "a paper boat floating down a rain-slicked street at night",
        "resolution": "480p",
        "aspect_ratio": "16:9",
        "duration_secs": 4,
    },
)
job = resp.json()

while job["status"] not in ("completed", "failed"):
    time.sleep(5)
    job = requests.get(
        f"https://videorouter.sh/api/v1/videos/{job['id']}",
        headers={"Authorization": "Bearer llmr_sk_live_..."},
    ).json()

if job["status"] == "completed":
    video_url = job["data"][0]["url"]
else:
    raise RuntimeError(job["error"])

Equivalent curl:

job=$(curl -s https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_..." \
  -H "Content-Type: application/json" \
  -d '{"model": "alibaba/wan-3.0", "prompt": "a paper boat floating down a rain-slicked street at night", "resolution": "480p", "duration_secs": 4}')
id=$(echo "$job" | jq -r .id)

status=$(echo "$job" | jq -r .status)
while [ "$status" != "completed" ] && [ "$status" != "failed" ]; do
  sleep 5
  job=$(curl -s "https://videorouter.sh/api/v1/videos/$id" -H "Authorization: Bearer llmr_sk_...")
  status=$(echo "$job" | jq -r .status)
done

echo "$job" | jq -r 'if .status == "completed" then .data[0].url else .error end'

Switching between base Wan 3.0 and Prime

Prime is a separate checkpoint, not a resolution tier — swap the model field to move between them:

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "alibaba/wan-3.0-prime",  # or "alibaba/wan-3.0" for the base checkpoint
        "prompt": "a paper boat floating down a rain-slicked street at night",
        "resolution": "480p",
        "duration_secs": 4,
    },
).json()

Prime costs roughly double base Wan 3.0 at its own cheapest host — see Wan 3.0 vs Wan 3.0 Prime: Which Should You Call? before defaulting a new integration to it.

Pinning to a specific provider

With eleven hosts carrying base Wan 3.0 and a nearly 4.25x spread between the cheapest and priciest of them, pinning explicitly is a bigger lever here than for most models in this series:

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "alibaba/wan-3.0/replicate",  # currently the cheapest verified host
        "prompt": "a paper boat floating down a rain-slicked street at night",
        "duration_secs": 4,
    },
).json()

Confirmed suffixes for base Wan 3.0: replicate, atlascloud, machgen, pika, fal, alibaba (direct), wavespeed, and openrouter. Not every host carries Prime — notably, Replicate and MachGen (two of the cheapest base-model hosts) don't list Prime at all, so a pipeline that pins to replicate for base Wan 3.0 and expects the same suffix to work for alibaba/wan-3.0-prime will get a 400 instead. See Wan 3.0 API Pricing: Every Provider Compared for the full per-host table this maps to.

Handling errors correctly

Every error follows the same envelope shape as the OpenAI API — {"error": {"message", "type", "code"}}:

Status type / code What it means
400 invalid_request_error Missing prompt, or a provider suffix that doesn't carry the requested checkpoint (e.g. alibaba/wan-3.0-prime/replicate)
401 invalid_api_key Key missing, malformed, revoked, or expired
402 spend_cap_exceeded / insufficient_credits Monthly cap hit, or prepaid balance ≤ $0
403 model_not_allowed The requested Wan checkpoint isn't in this key's model_allowlist
429 rpm_limit / tpm_limit Rate limit exceeded — Retry-After tells you how long to wait
500 / 502 / 503 / 504 upstream_error Every candidate host in the fallback chain failed — not billed
import time
import requests

API_BASE = "https://videorouter.sh/api/v1"
API_KEY = "llmr_sk_live_..."

def generate_wan_video(prompt, prime=False, provider=None, **kwargs):
    model = "alibaba/wan-3.0" + ("-prime" if prime else "") + (f"/{provider}" if provider else "")
    headers = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
    payload = {"model": model, "prompt": prompt, **kwargs}

    for attempt in range(3):
        resp = requests.post(f"{API_BASE}/videos", headers=headers, json=payload)
        if resp.status_code == 429:
            time.sleep(int(resp.headers.get("Retry-After", 5)))
            continue
        resp.raise_for_status()
        job = resp.json()
        break
    else:
        raise RuntimeError("exceeded retry budget on job creation")

    deadline = time.time() + 600
    while job["status"] not in ("completed", "failed"):
        if time.time() > deadline:
            raise TimeoutError(f"job {job['id']} did not complete within 10 minutes")
        time.sleep(5)
        job = requests.get(f"{API_BASE}/videos/{job['id']}", headers=headers).json()

    if job["status"] == "failed":
        raise RuntimeError(f"generation failed: {job['error']}")
    return job["data"][0]["url"]

# usage
url = generate_wan_video("a paper boat floating down a rain-slicked street at night", provider="replicate", resolution="480p")

Given eleven hosts carry base Wan 3.0, a genuine upstream_error on an unpinned request is a stronger signal than the same error on a two- or three-host model — either something is affecting several independent providers at once, or (far more likely) you've pinned to a single host and that one specifically is down.

Common mistakes

Pinning to a suffix that doesn't carry Prime. As noted above, replicate and machgen are base-model-only — the equivalent Prime request needs a different suffix (pika, atlascloud, fal, alibaba, or wavespeed).

Assuming "direct from Alibaba" is the cheap option. It isn't — Alibaba's own direct listing ties with Fal and WaveSpeedAI at the median price, not the floor. Replicate currently undercuts it by more than 2x. See Wan 3.0 API Pricing.

Not re-checking the host ranking per resolution tier. Wan 3.0's tier multiplier (roughly 2x per step from 480p → 720p → 1080p) is consistent across hosts, so the cheapest host at 480p stays cheapest at 1080p for this model — but that's not guaranteed to hold for every model, so verify rather than assume when integrating a different one.

See /docs/video-generation for the complete request reference.

Frequently asked questions

Can I get Wan 3.0 output back synchronously? No — generation is async and job-based across every host, the same as every other video model VideoRouter routes. Poll on a 5-10 second interval as shown above.

Does pinning change VideoRouter's own fee, or just the underlying host's rate? Only the underlying host's rate — pinning determines which of the eleven hosts' prices applies to your request; it doesn't change VideoRouter's routing fee structure.

What's the fastest way to check which host served an unpinned request? The completed job object carries provider attribution — check the exact field name in /docs/api-reference-video.

Is there a Wan 2.7 equivalent of this guide? Wan 2.7 is a separate, earlier-generation model with its own narrower host list, not covered here — check VideoRouter's Wan 2.7 model page for its own pricing and hosts rather than assuming this guide's suffixes carry over.

How do I A/B test base Wan 3.0 against Prime fairly? Run both through a host that carries both (Fal or Alibaba direct are the simplest same-host comparison) rather than comparing each at its own cheapest host — see Wan 3.0 vs Wan 3.0 Prime for why isolating the checkpoint as the only variable matters.