Discord

Multimodal

Lip sync

A dedicated POST /v1/lipsync endpoint — not one more /v1/videos request shape. A driving audio track re-syncs an existing face, in one of two modes depending on the model you pick: image mode animates a still photo (image_url + audio_url), video mode dubs an existing clip's mouth to new speech (video_url + audio_url). There's no prompt field at all, and no duration_secs either — output length always follows the real length of your audio (and, in video mode, your source video too).

This is deliberately kept off /v1/videos' own start_image_url/input_video_url/input_audio_url fields — those already mean something else there (starting frame, video editing, LTX's own audio-to-video), and combining them for lip sync would be one field-name accident away from a reference-to-video call (input_references/input_audio_references) that generates something new instead of preserving your exact source face/scene and just re-timing the mouth.

Image mode: animate a still photo

import requests, time

job = requests.post(
    "https://videorouter.sh/api/v1/lipsync",
    headers={"Authorization": "Bearer llmr_sk_live_..."},
    json={
        "model": "fal/h3-max-lipsync",
        "image_url": "https://example.com/face.jpg",
        "audio_url": "https://example.com/speech.mp3",
    },
).json()
# -> {"id": "fal:...", "status": "queued", ...}

status = requests.get(
    f"https://videorouter.sh/api/v1/lipsync/{job['id']}",
    headers={"Authorization": "Bearer llmr_sk_live_..."},
).json()

Video mode: dub an existing clip

import requests, time

job = requests.post(
    "https://videorouter.sh/api/v1/lipsync",
    headers={"Authorization": "Bearer llmr_sk_live_..."},
    json={
        "model": "fal/sync-lipsync-v3",
        "video_url": "https://example.com/clip.mp4",
        "audio_url": "https://example.com/new-speech.mp3",
    },
).json()
# -> {"id": "fal:...", "status": "queued", ...}

status = requests.get(
    f"https://videorouter.sh/api/v1/lipsync/{job['id']}",
    headers={"Authorization": "Bearer llmr_sk_live_..."},
).json()

Real price is computed from your file's own probed duration — for video-mode models, that's max(video_secs, audio_secs), never the shorter of the two, since some models can stretch the shorter input to match the longer one. That cost already includes our 2% platform fee.

Which models

10 models across 2 providers — fal.ai and Atlas Cloud.

Image mode (image_url + audio_url)

ModelProviderPrice
fal/h3-max-lipsyncfal.ai (MiniMax)$0.05–$0.32/s by resolution
fal/sync-lipsync-v3-imagefal.ai (Sync Labs)$0.1333/s
atlascloud/infinitetalkAtlas Cloud (MeiGen-AI)$0.03/s (480p), $0.06/s (720p)

Video mode (video_url + audio_url)

ModelProviderPrice
fal/sync-lipsync-v3fal.ai (Sync Labs sync-3)$0.1333/s
fal/sync-lipsync-react-1fal.ai (Sync Labs)$0.1667/s, max 15s in/out
fal/sync-lipsync-v2-profal.ai (Sync Labs lipsync-2-pro)$0.0833/s
fal/sync-lipsync-v2fal.ai (Sync Labs lipsync-2)$0.05/s
fal/veed-lipsync-v2fal.ai (VEED)$0.07/s
fal/heygen-v3-lipsync-precisionfal.ai (HeyGen)$0.10/s
fal/heygen-v3-lipsync-speedfal.ai (HeyGen)$0.05/s

Atlas Cloud rows bill through Atlas Cloud's reserve-then-settle flow (same as every other Atlas Cloud video model) — the price shown is a pre-call estimate; the real charge settles from Atlas Cloud's own completed-job invoice. Every model here also shows up under the Lip Sync filter on the models catalog.

Not yet supported

Avatar/digital-human models that need a voice_id or avatar selection (e.g. HeyGen's own Avatar-IV/Digital Twin) aren't onboarded — this endpoint has no avatar/voice-selection concept, only a real image or video you supply yourself. Kling LipSync and Hedra aren't wired in yet either. Self-hosted open-weight options (LatentSync, MuseTalk) would need their own GPU deployment, not a simple resale row, and aren't on this platform today.