Multimodal
Lip sync
A dedicated POST /v1/lipsync endpoint
— not one more /v1/videos request
shape. A driving audio track re-syncs an existing face, in one of two modes depending on the model you
pick: image mode animates a still photo (image_url
+ audio_url), video mode
dubs an existing clip's mouth to new speech (video_url
+ audio_url). There's no
prompt field at all, and no
duration_secs either — output length
always follows the real length of your audio (and, in video mode, your source video too).
This is deliberately kept off /v1/videos'
own start_image_url/input_video_url/input_audio_url
fields — those already mean something else there (starting frame, video editing, LTX's own audio-to-video),
and combining them for lip sync would be one field-name accident away from a reference-to-video call
(input_references/input_audio_references)
that generates something new instead of preserving your exact source face/scene and just re-timing the mouth.
Image mode: animate a still photo
import requests, time
job = requests.post(
"https://videorouter.sh/api/v1/lipsync",
headers={"Authorization": "Bearer llmr_sk_live_..."},
json={
"model": "fal/h3-max-lipsync",
"image_url": "https://example.com/face.jpg",
"audio_url": "https://example.com/speech.mp3",
},
).json()
# -> {"id": "fal:...", "status": "queued", ...}
status = requests.get(
f"https://videorouter.sh/api/v1/lipsync/{job['id']}",
headers={"Authorization": "Bearer llmr_sk_live_..."},
).json()
Video mode: dub an existing clip
import requests, time
job = requests.post(
"https://videorouter.sh/api/v1/lipsync",
headers={"Authorization": "Bearer llmr_sk_live_..."},
json={
"model": "fal/sync-lipsync-v3",
"video_url": "https://example.com/clip.mp4",
"audio_url": "https://example.com/new-speech.mp3",
},
).json()
# -> {"id": "fal:...", "status": "queued", ...}
status = requests.get(
f"https://videorouter.sh/api/v1/lipsync/{job['id']}",
headers={"Authorization": "Bearer llmr_sk_live_..."},
).json()
Real price is computed from your file's own probed duration — for video-mode models, that's
max(video_secs, audio_secs), never
the shorter of the two, since some models can stretch the shorter input to match the longer one. That cost
already includes our 2% platform fee.
Which models
10 models across 2 providers — fal.ai and Atlas Cloud.
Image mode (image_url + audio_url)
fal/h3-max-lipsyncfal.ai (MiniMax)$0.05–$0.32/s by resolutionfal/sync-lipsync-v3-imagefal.ai (Sync Labs)$0.1333/satlascloud/infinitetalkAtlas Cloud (MeiGen-AI)$0.03/s (480p), $0.06/s (720p)Video mode (video_url + audio_url)
fal/sync-lipsync-v3fal.ai (Sync Labs sync-3)$0.1333/sfal/sync-lipsync-react-1fal.ai (Sync Labs)$0.1667/s, max 15s in/outfal/sync-lipsync-v2-profal.ai (Sync Labs lipsync-2-pro)$0.0833/sfal/sync-lipsync-v2fal.ai (Sync Labs lipsync-2)$0.05/sfal/veed-lipsync-v2fal.ai (VEED)$0.07/sfal/heygen-v3-lipsync-precisionfal.ai (HeyGen)$0.10/sfal/heygen-v3-lipsync-speedfal.ai (HeyGen)$0.05/sAtlas Cloud rows bill through Atlas Cloud's reserve-then-settle flow (same as every other Atlas Cloud video model) — the price shown is a pre-call estimate; the real charge settles from Atlas Cloud's own completed-job invoice. Every model here also shows up under the Lip Sync filter on the models catalog.
Not yet supported
Avatar/digital-human models that need a voice_id
or avatar selection (e.g. HeyGen's own Avatar-IV/Digital Twin) aren't onboarded — this endpoint has no
avatar/voice-selection concept, only a real image or video you supply yourself. Kling LipSync and Hedra
aren't wired in yet either. Self-hosted open-weight options (LatentSync, MuseTalk) would need their own GPU
deployment, not a simple resale row, and aren't on this platform today.