Multimodal
Async, job-based — create a job, then poll it. Billed per requested second at creation time; polling for
status is free. There's no OpenRouter backstop for this endpoint: each model is served direct by the
provider that owns it. Three generation modes share this one endpoint — plain
text-to-video (below),
image-to-video (animate a starting
frame), and reference-to-video
(guide generation from multiple images, videos, and audio clips at once) — which one you get depends
only on which model you pass.
Text-to-video
import time
import requests
# create
resp = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "minimax/h3/fal",
"prompt": "a paper airplane gliding over a city",
"duration_secs": 8,
},
)
job = resp.json()
# poll until it leaves the queue ("queued" -> "in_progress" -> "completed"/"failed")
while job["status"] not in ("completed", "failed"):
time.sleep(5)
job = requests.get(
f"https://videorouter.sh/api/v1/videos/{job['id']}",
headers={"Authorization": "Bearer llmr_sk_live_..."},
).json()
if job["status"] == "completed":
video_url = job["data"][0]["url"] # the finished video
else:
raise RuntimeError(job["error"])
The OpenAI SDK doesn't have first-class video methods yet — use its low-level
client.post/client.get,
or plain requests, against the same
base_url.
model is
"<creator>/<model>/<host>" —
"minimax/h3/fal" above is a
soft preference for Fal, not a hard pin: it still falls back across every other host
confirmed to serve that checkpoint if Fal itself errors. Drop the host segment (just
"minimax/h3") to route automatically
to whichever host is cheapest right now instead. For a hard pin (no fallback at all), use
"provider": {"only": ["fal"], "allow_fallbacks": false}.
Full mechanics: Provider selection.
Resolution & aspect ratio
Pass resolution (a tier the
model supports, e.g. "720p") and
aspect_ratio (one of
16:9/9:16/1:1/4:3/3:4/21:9)
to request a size together — same shape as
/v1/images. Each model only
supports some combinations; an unrecognized or unsupported one never 400s — it's ignored and the
model's default resolution/orientation is used instead, same as omitting both. Explicit
height/width
pixel integers are still accepted too, and override
resolution/aspect_ratio
when given — but for most callers resolution/aspect_ratio
is simpler than computing exact pixel dimensions by hand.
import requests
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "minimax/h3/fal",
"prompt": "a paper airplane gliding over a city",
"resolution": "720p",
"aspect_ratio": "16:9",
"duration_secs": 8,
},
).json()
For veo-3.1/veo-3.1-fast
resolution/aspect_ratio
select one of a small set of confirmed tiers — the DeepInfra and SiliconFlow rows ignore both. For
veo-3.1-fast, the tier also changes the
price.
Image-to-video
Pass start_image_url — a public
https:// URL or an inline
data:image/...;base64,... URI — to
animate a starting frame. On most models this is the same public model id you'd use for
text-to-video — omit start_image_url and
it's plain text-to-video, pass it and the model animates that frame instead: Fal
(fal/seedance-2.0,
fal/h3), Atlas Cloud (e.g.
atlascloud/h3), Replicate
(replicate/h3), WaveSpeedAI (e.g.
wavespeed/h3), and MachGen
(machgen/minimax-h3). A few models have
no text-to-video mode at all and require start_image_url
(400 without it): fal/happy-horse,
novita/wan2.6-i2v,
machgen/vidu-q3-pro-fast, and
machgen/grok-imagine-video-1.5. Every
other model rejects start_image_url with
a 400.
import requests
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "bytedance/seedance-2.0/fal",
"prompt": "the subject turns and smiles, gentle camera push-in",
"start_image_url": "https://example.com/photo.jpg",
"duration_secs": 5,
"resolution": "720p",
"aspect_ratio": "16:9",
},
).json()
fal/happy-horse's
prompt is optional — omit it and the
image is animated with no text guidance at all.
end_image_url (last-frame/keyframe
generation) is accepted by no model yet — passing it is rejected with a 400 rather than silently
ignored.
Uploading a file
Every image/video/audio field above — start_image_url
and every input_references /
input_video_references /
input_audio_references entry below —
takes either a public https:// URL or an
inline data:image/...;base64,... URI, so
if you already have the file hosted somewhere, or you're fine base64-encoding it inline, you don't need
anything below this point. If you'd rather not do either — the file is only local, or base64 would
inflate a large image/video past a reasonable request size — POST /v1/uploads
stages it in storage for you and hands back a URL to pass into the field instead.
import requests
# 1. upload the file, get back a URL
with open("photo.jpg", "rb") as f:
upload = requests.post(
"https://videorouter.sh/api/v1/uploads",
headers={"Authorization": "Bearer llmr_sk_live_..."},
files={"file": f},
).json()
# -> {"url": "https://...", "expires_in": 1800}
# 2. use that URL right away as start_image_url (or an input_references entry)
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "bytedance/seedance-2.0/fal",
"prompt": "the subject turns and smiles, gentle camera push-in",
"start_image_url": upload["url"],
},
).json()
file is a
multipart/form-data field — any
image/*,
video/*, or
audio/* content type, up to 50 MB.
The returned URL is presigned and expires in 30 minutes — plenty of time to pass it straight into your
very next /v1/videos or
/v1/images call, but it's scratch
space for that one call, not somewhere to keep a file — upload again for each new job.
Reference-to-video
Distinct from image-to-video above — not just a bigger version of it. Some models
(bytedance/seedance-2.0,
minimax/h3/fal,
minimax/h3/atlas-cloud) can
composite multiple reference files of three different kinds in one
call: up to 9 images, 3 videos, and 3 audio clips (12 combined) — same model id as
text/image-to-video above, mode picked automatically by which fields you send, no separate
"-reference" id to look up. Two new request fields extend the same
{"type", "<kind>_url": {"url"}}
convention input_references already
uses, just for video and audio:
import requests
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "bytedance/seedance-2.0/fal",
"prompt": "the cat from the reference image walks across the scene",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}}
],
"input_video_references": [
{"type": "video_url", "video_url": {"url": "https://example.com/motion-ref.mp4"}}
],
"input_audio_references": [
{"type": "audio_url", "audio_url": {"url": "https://example.com/ambience.mp3"}}
],
"duration_secs": 5,
},
).json()
All three arrays are optional individually, but at least one reference across all
three is required — a call with none is rejected with a 400 (pass a plain
prompt with no references to
fal/seedance-2.0 instead for
text-to-video). Exceeding any per-kind cap, or 12 combined, is also rejected with a 400.
MachGen's reference-to-video rows (machgen/vidu-q3,
machgen/minimax-h3-reference) are
narrower — images only, up to 7, no input_video_references/input_audio_references
support at all. Pass input_references
the same way, just omit the other two arrays.
Video editing
Different from reference-to-video above — this modifies an existing clip instead of
guiding a new generation. Pass the clip as input_video_url
(a single URL, not an array) alongside a prompt
describing the edit — it's mutually exclusive with
start_image_url/input_references/input_video_references
in the same call.
import requests
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "xai/grok-imagine-video-edit/atlas-cloud",
"prompt": "Adjust the video style to an American comic book style.",
"input_video_url": "https://example.com/source-clip.mp4",
},
).json()
Currently wired: grok-imagine-video-edit
(xAI Grok Imagine, billed per second of the input clip — output length always matches input,
capped at 8.7s), replicate/kling-v3-omni-video
(send input_video_url instead of
input_video_references to switch this
model from reference/style mode into true editing mode),
runway/aleph-2 (Runway's Aleph 2.0,
input capped at 30s), and pika/pikadditions
(flat $0.03/request, object-insertion-style edits), and
flux-3-edit-video (Black Forest Labs'
FLUX Video Edit, direct — $0.03 per second of output, which always matches the input length; also
available via Fal as fal/flux-3-edit-video).
To continue a clip instead of modifying it, see Video extension.
Billing
The full cost is billed once, at job creation, from the requested duration_secs —
GET /v1/videos/{id} status polling has
no incremental provider cost and is unbilled. Like every other mode, this endpoint carries the flat
2% platform fee on top of provider cost — see Pricing & billing.
Exception: SiliconFlow's two siliconflow/-prefixed
models bill a flat $0.29 per clip regardless of any duration_secs
you request — SiliconFlow's own API ignores duration and always renders a fixed ~5s clip, so any
duration_secs in the request body for
these two models is ignored on our side too.
Fallback & failover
A rejected submission automatically fails over to the next-cheapest host still serving the same
checkpoint — free, on by default, no parameter needed. For a job that's already been accepted and then
stalls or fails mid-generation, pass failover.on_timeout_sec to
hedge to a second attempt (both attempts get billed). Full details, triggers, and request examples:
Provider selection.
Falling back to a genuinely different model instead of another host of the same one:
Model fallbacks.
Not yet supported
Video editing is wired for 5 models only (see Video editing above), and video extension for one model. Video as chat input isn't supported either.
Image-to-video is wired for Fal, Atlas Cloud, Replicate, WaveSpeedAI, Novita, and MachGen. Reference-to-video (multiple image/video/audio inputs at once) is wired for Fal, Atlas Cloud, and MachGen. Kling and Luma each have their own image-to-video capability upstream, but it isn't wired into this endpoint yet.