Discord

Multimodal

Async, job-based — create a job, then poll it. Billed per requested second at creation time; polling for status is free. There's no OpenRouter backstop for this endpoint: each model is served direct by the provider that owns it. Three generation modes share this one endpoint — plain text-to-video (below), image-to-video (animate a starting frame), and reference-to-video (guide generation from multiple images, videos, and audio clips at once) — which one you get depends only on which model you pass.

Text-to-video

import time
import requests

# create
resp = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "minimax/h3/fal",
        "prompt": "a paper airplane gliding over a city",
        "duration_secs": 8,
    },
)
job = resp.json()

# poll until it leaves the queue ("queued" -> "in_progress" -> "completed"/"failed")
while job["status"] not in ("completed", "failed"):
    time.sleep(5)
    job = requests.get(
        f"https://videorouter.sh/api/v1/videos/{job['id']}",
        headers={"Authorization": "Bearer llmr_sk_live_..."},
    ).json()

if job["status"] == "completed":
    video_url = job["data"][0]["url"]  # the finished video
else:
    raise RuntimeError(job["error"])

The OpenAI SDK doesn't have first-class video methods yet — use its low-level client.post/client.get, or plain requests, against the same base_url.

model is "<creator>/<model>/<host>" — "minimax/h3/fal" above is a soft preference for Fal, not a hard pin: it still falls back across every other host confirmed to serve that checkpoint if Fal itself errors. Drop the host segment (just "minimax/h3") to route automatically to whichever host is cheapest right now instead. For a hard pin (no fallback at all), use "provider": {"only": ["fal"], "allow_fallbacks": false}. Full mechanics: Provider selection.

Resolution & aspect ratio

Pass resolution (a tier the model supports, e.g. "720p") and aspect_ratio (one of 16:9/9:16/1:1/4:3/3:4/21:9) to request a size together — same shape as /v1/images. Each model only supports some combinations; an unrecognized or unsupported one never 400s — it's ignored and the model's default resolution/orientation is used instead, same as omitting both. Explicit height/width pixel integers are still accepted too, and override resolution/aspect_ratio when given — but for most callers resolution/aspect_ratio is simpler than computing exact pixel dimensions by hand.

import requests

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "minimax/h3/fal",
        "prompt": "a paper airplane gliding over a city",
        "resolution": "720p",
        "aspect_ratio": "16:9",
        "duration_secs": 8,
    },
).json()

For veo-3.1/veo-3.1-fast resolution/aspect_ratio select one of a small set of confirmed tiers — the DeepInfra and SiliconFlow rows ignore both. For veo-3.1-fast, the tier also changes the price.

Image-to-video

Pass start_image_url — a public https:// URL or an inline data:image/...;base64,... URI — to animate a starting frame. On most models this is the same public model id you'd use for text-to-video — omit start_image_url and it's plain text-to-video, pass it and the model animates that frame instead: Fal (fal/seedance-2.0, fal/h3), Atlas Cloud (e.g. atlascloud/h3), Replicate (replicate/h3), WaveSpeedAI (e.g. wavespeed/h3), and MachGen (machgen/minimax-h3). A few models have no text-to-video mode at all and require start_image_url (400 without it): fal/happy-horse, novita/wan2.6-i2v, machgen/vidu-q3-pro-fast, and machgen/grok-imagine-video-1.5. Every other model rejects start_image_url with a 400.

import requests

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "bytedance/seedance-2.0/fal",
        "prompt": "the subject turns and smiles, gentle camera push-in",
        "start_image_url": "https://example.com/photo.jpg",
        "duration_secs": 5,
        "resolution": "720p",
        "aspect_ratio": "16:9",
    },
).json()

fal/happy-horse's prompt is optional — omit it and the image is animated with no text guidance at all. end_image_url (last-frame/keyframe generation) is accepted by no model yet — passing it is rejected with a 400 rather than silently ignored.

Uploading a file

Every image/video/audio field above — start_image_url and every input_references / input_video_references / input_audio_references entry below — takes either a public https:// URL or an inline data:image/...;base64,... URI, so if you already have the file hosted somewhere, or you're fine base64-encoding it inline, you don't need anything below this point. If you'd rather not do either — the file is only local, or base64 would inflate a large image/video past a reasonable request size — POST /v1/uploads stages it in storage for you and hands back a URL to pass into the field instead.

import requests

# 1. upload the file, get back a URL
with open("photo.jpg", "rb") as f:
    upload = requests.post(
        "https://videorouter.sh/api/v1/uploads",
        headers={"Authorization": "Bearer llmr_sk_live_..."},
        files={"file": f},
    ).json()
# -> {"url": "https://...", "expires_in": 1800}

# 2. use that URL right away as start_image_url (or an input_references entry)
job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "bytedance/seedance-2.0/fal",
        "prompt": "the subject turns and smiles, gentle camera push-in",
        "start_image_url": upload["url"],
    },
).json()

file is a multipart/form-data field — any image/*, video/*, or audio/* content type, up to 50 MB. The returned URL is presigned and expires in 30 minutes — plenty of time to pass it straight into your very next /v1/videos or /v1/images call, but it's scratch space for that one call, not somewhere to keep a file — upload again for each new job.

Reference-to-video

Distinct from image-to-video above — not just a bigger version of it. Some models (bytedance/seedance-2.0, minimax/h3/fal, minimax/h3/atlas-cloud) can composite multiple reference files of three different kinds in one call: up to 9 images, 3 videos, and 3 audio clips (12 combined) — same model id as text/image-to-video above, mode picked automatically by which fields you send, no separate "-reference" id to look up. Two new request fields extend the same {"type", "<kind>_url": {"url"}} convention input_references already uses, just for video and audio:

import requests

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "bytedance/seedance-2.0/fal",
        "prompt": "the cat from the reference image walks across the scene",
        "input_references": [
            {"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}}
        ],
        "input_video_references": [
            {"type": "video_url", "video_url": {"url": "https://example.com/motion-ref.mp4"}}
        ],
        "input_audio_references": [
            {"type": "audio_url", "audio_url": {"url": "https://example.com/ambience.mp3"}}
        ],
        "duration_secs": 5,
    },
).json()

All three arrays are optional individually, but at least one reference across all three is required — a call with none is rejected with a 400 (pass a plain prompt with no references to fal/seedance-2.0 instead for text-to-video). Exceeding any per-kind cap, or 12 combined, is also rejected with a 400.

MachGen's reference-to-video rows (machgen/vidu-q3, machgen/minimax-h3-reference) are narrower — images only, up to 7, no input_video_references/input_audio_references support at all. Pass input_references the same way, just omit the other two arrays.

Video editing

Different from reference-to-video above — this modifies an existing clip instead of guiding a new generation. Pass the clip as input_video_url (a single URL, not an array) alongside a prompt describing the edit — it's mutually exclusive with start_image_url/input_references/input_video_references in the same call.

import requests

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "xai/grok-imagine-video-edit/atlas-cloud",
        "prompt": "Adjust the video style to an American comic book style.",
        "input_video_url": "https://example.com/source-clip.mp4",
    },
).json()

Currently wired: grok-imagine-video-edit (xAI Grok Imagine, billed per second of the input clip — output length always matches input, capped at 8.7s), replicate/kling-v3-omni-video (send input_video_url instead of input_video_references to switch this model from reference/style mode into true editing mode), runway/aleph-2 (Runway's Aleph 2.0, input capped at 30s), and pika/pikadditions (flat $0.03/request, object-insertion-style edits), and flux-3-edit-video (Black Forest Labs' FLUX Video Edit, direct — $0.03 per second of output, which always matches the input length; also available via Fal as fal/flux-3-edit-video). To continue a clip instead of modifying it, see Video extension.

Billing

The full cost is billed once, at job creation, from the requested duration_secs — GET /v1/videos/{id} status polling has no incremental provider cost and is unbilled. Like every other mode, this endpoint carries the flat 2% platform fee on top of provider cost — see Pricing & billing.

Exception: SiliconFlow's two siliconflow/-prefixed models bill a flat $0.29 per clip regardless of any duration_secs you request — SiliconFlow's own API ignores duration and always renders a fixed ~5s clip, so any duration_secs in the request body for these two models is ignored on our side too.

Fallback & failover

A rejected submission automatically fails over to the next-cheapest host still serving the same checkpoint — free, on by default, no parameter needed. For a job that's already been accepted and then stalls or fails mid-generation, pass failover.on_timeout_sec to hedge to a second attempt (both attempts get billed). Full details, triggers, and request examples: Provider selection. Falling back to a genuinely different model instead of another host of the same one: Model fallbacks.

Not yet supported

Video editing is wired for 5 models only (see Video editing above), and video extension for one model. Video as chat input isn't supported either.

Image-to-video is wired for Fal, Atlas Cloud, Replicate, WaveSpeedAI, Novita, and MachGen. Reference-to-video (multiple image/video/audio inputs at once) is wired for Fal, Atlas Cloud, and MachGen. Kling and Luma each have their own image-to-video capability upstream, but it isn't wired into this endpoint yet.