How to Call the Veo 3.1 API (Python + curl, Flagship/Fast/Lite)

Google's Veo 3.1 ships as three separate tiers — the flagship, Fast, and Lite — served through VideoRouter's same /v1/videos endpoint every other video model uses. The request shape doesn't change between tiers or hosts; only the model field and, optionally, a provider suffix do. That makes Veo one of the simpler integrations in this series to get right once you understand the tier naming, which is also the part most likely to trip up a first integration.

Text-to-video

import time
import requests

resp = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "google/veo-3.1",
        "prompt": "a lighthouse beam sweeping across a foggy harbor",
        "resolution": "1080p",
        "aspect_ratio": "16:9",
        "duration_secs": 8,
    },
)
job = resp.json()

while job["status"] not in ("completed", "failed"):
    time.sleep(5)
    job = requests.get(
        f"https://videorouter.sh/api/v1/videos/{job['id']}",
        headers={"Authorization": "Bearer llmr_sk_live_..."},
    ).json()

if job["status"] == "completed":
    video_url = job["data"][0]["url"]
else:
    raise RuntimeError(job["error"])

Equivalent curl:

job=$(curl -s https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_..." \
  -H "Content-Type: application/json" \
  -d '{"model": "google/veo-3.1", "prompt": "a lighthouse beam sweeping across a foggy harbor", "resolution": "1080p", "duration_secs": 8}')
id=$(echo "$job" | jq -r .id)

status=$(echo "$job" | jq -r .status)
while [ "$status" != "completed" ] && [ "$status" != "failed" ]; do
  sleep 5
  job=$(curl -s "https://videorouter.sh/api/v1/videos/$id" -H "Authorization: Bearer llmr_sk_...")
  status=$(echo "$job" | jq -r .status)
done

echo "$job" | jq -r 'if .status == "completed" then .data[0].url else .error end'

Switching between flagship, Fast, and Lite

All three tiers share the identical request shape — only the model field changes:

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "google/veo-3.1-lite",  # or "google/veo-3.1-fast", or "google/veo-3.1"
        "prompt": "a lighthouse beam sweeping across a foggy harbor",
        "resolution": "720p",
        "duration_secs": 6,
    },
).json()

That means a tiered pipeline — generating drafts on Lite, previews on Fast, and finals on the flagship — is a routing decision inside your own application, not three separate integrations. See Veo 3.1 vs Veo 3.1 Fast vs Veo 3.1 Lite: Pricing and When to Use Each for the full price gap between tiers (up to 6.7x between Lite and the flagship) and when each is the right call.

Pinning to a specific provider

Veo 3.1 has one of the widest host rosters in this series — seven providers at the flagship tier alone — with a real price split worth knowing before you pick a default:

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "google/veo-3.1/replicate",  # cheapest verified host for the flagship tier
        "prompt": "a lighthouse beam sweeping across a foggy harbor",
        "duration_secs": 8,
    },
).json()

Confirmed suffixes for the flagship tier: google (direct), pika, deepinfra, machgen, replicate, wavespeed, and openrouter. Three of these — pika, replicate, openrouter — currently price the flagship at $0.20/sec; the other four sit at $0.40/sec, exactly double, for the identical model. See Veo 3.1 API Pricing Comparison Across Providers for the full breakdown, including the counterintuitive part: calling Google directly (the google suffix) is 2x the price of Replicate or Pika, not the cheap option.

Handling errors correctly

Every error follows the same envelope shape as the OpenAI API — {"error": {"message", "type", "code"}} — regardless of tier or host:

Status type / code What it means
400 invalid_request_error Missing prompt, or a model id that doesn't exist (e.g. a typo like google/veo3.1)
401 invalid_api_key Key missing, malformed, revoked, or expired
402 spend_cap_exceeded / insufficient_credits Monthly cap hit, or prepaid balance ≤ $0
403 model_not_allowed The requested Veo tier isn't in this key's model_allowlist
429 rpm_limit / tpm_limit Rate limit exceeded — Retry-After tells you how long to wait
500 / 502 / 503 / 504 upstream_error Every candidate host failed — not billed
import time
import requests

API_BASE = "https://videorouter.sh/api/v1"
API_KEY = "llmr_sk_live_..."

def generate_veo_video(prompt, tier="flagship", provider=None, **kwargs):
    tier_suffix = {"flagship": "", "fast": "-fast", "lite": "-lite"}[tier]
    model = f"google/veo-3.1{tier_suffix}" + (f"/{provider}" if provider else "")
    headers = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
    payload = {"model": model, "prompt": prompt, **kwargs}

    for attempt in range(3):
        resp = requests.post(f"{API_BASE}/videos", headers=headers, json=payload)
        if resp.status_code == 429:
            time.sleep(int(resp.headers.get("Retry-After", 5)))
            continue
        resp.raise_for_status()
        job = resp.json()
        break
    else:
        raise RuntimeError("exceeded retry budget on job creation")

    deadline = time.time() + 600
    while job["status"] not in ("completed", "failed"):
        if time.time() > deadline:
            raise TimeoutError(f"job {job['id']} did not complete within 10 minutes")
        time.sleep(5)
        job = requests.get(f"{API_BASE}/videos/{job['id']}", headers=headers).json()

    if job["status"] == "failed":
        raise RuntimeError(f"generation failed: {job['error']}")
    return job["data"][0]["url"]

# usage
url = generate_veo_video("a lighthouse beam sweeping across a foggy harbor", tier="fast", provider="pika")

Common mistakes

Assuming Lite is available at every host the flagship is. Lite currently ships through only three hosts (Pika, OpenRouter, WaveSpeedAI) versus seven for the flagship — pinning to a provider suffix that doesn't carry Lite (e.g. google/veo-3.1-lite/machgen) returns a 400 rather than silently falling back to a host that does.

Assuming "direct from Google" means cheapest. It's the opposite for Veo 3.1 — the google suffix prices at $0.40/sec, double what replicate, pika, or openrouter charge for the identical model. Default to one of the cheaper suffixes unless you have a specific compliance or SLA reason to prefer Google's own infrastructure.

Not re-checking pricing at the 4K/2160p tier before assuming the standard-tier cheapest host stays cheapest. MachGen's 2160p rate ($0.60/sec) breaks from the pattern that holds at 720p/1080p — see Veo 3.1 API Pricing Comparison Across Providers for the tier-by-tier detail.

See /docs/video-generation for the complete request reference.

Frequently asked questions

Can I get Veo output back synchronously, without polling? No — like every video model VideoRouter routes, Veo generation is async and job-based because it takes longer than a single HTTP request should stay open for. Poll on a 5-10 second interval as shown above.

Does the aspect_ratio field work identically across all three Veo tiers? Treat it as tier-independent unless you observe otherwise — check /docs/video-generation for the current accepted values, since supported aspect ratios are a parameter-level detail more likely to change than pricing.

What's the fastest way to check which host actually served a request I didn't pin? The completed job object carries provider attribution — check the exact field name in /docs/api-reference-video rather than assuming based on VideoRouter's routing bias alone.

Is there a Veo 3.1 image-to-video mode? Check /docs/video-generation for the current parameter shape — this is a capability detail worth verifying against live docs rather than a pricing-focused guide like this one.

Can I mix tiers within the same user session, like generating a Lite draft then a flagship final? Yes — since all three tiers share the identical request shape, that's a routing decision in your own code, not three integrations. See the worked pipeline example in Veo 3.1 vs Veo 3.1 Fast vs Veo 3.1 Lite.