Discord

Multimodal

Music generation

POST /v1/audio/music — text-to-music across 5 providers: Google Lyria, Mureka, ElevenLabs, StepFun, and Suno Chirp (via Atlas Cloud). Returns raw audio bytes in one response (no job/poll, no streaming). Most models are billed flat per generated song, regardless of input length or output duration — ElevenLabs is the one exception, billed per minute of requested track length (see Which models below).

import requests

resp = requests.post(
    "https://videorouter.sh/api/v1/audio/music",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "lyria-3.5",
        "prompt": "An upbeat acoustic guitar folk song about road trips, 2 minutes, with vocals",
        "response_format": "mp3",
    },
)
with open("song.mp3", "wb") as f:
    f.write(resp.content)

Which models

ModelProviderLengthPrice
lyria-3.5GoogleFull song (verses, choruses, bridges) — a couple of minutes, controllable via prompt$0.08 / song
lyria-3-clip-previewGoogleShort clip/loop — fixed 30 seconds$0.04 / song
lyria-3-pro-previewGoogleFull song, prompt-controlled$0.08 / song
mureka-v9MurekaFull song, prompt/lyrics-controlled$0.45 / song
mureka-v8MurekaFull song, prompt/lyrics-controlled$0.45 / song
elevenlabs-music-v2.5ElevenLabsSet via duration in seconds — defaults to 30s if omitted$0.15 / minute
stepaudio-3-musicStepFunFull song, prompt-controlledFree (preview)
atlascloud/suno-chirp-v6Suno (via Atlas Cloud)Full song, prompt-controlled$0.132 / song
atlascloud/suno-chirp-v6-wildSuno (via Atlas Cloud)Full song, prompt-controlled, more adventurous style adherence$0.132 / song
atlascloud/suno-chirp-v6-miniSuno (via Atlas Cloud)Full song, prompt-controlled, lighter/faster tier$0.132 / song

Every model except elevenlabs-music-v2.5 is billed flat per generated track, regardless of how long the output ends up — there's no per-second or per-character component the way text-to-speech is billed. ElevenLabs is billed $0.15 per minute of the duration you request.

Controlling genre, vocals & length

Three optional request fields work across the non-Lyria providers: lyrics (Mureka, StepFun — your own lyrics text, section tags like [Verse]/[Chorus] supported), instrumental (boolean, every provider except the Lyria rows), and duration (seconds — ElevenLabs only; every other provider's length is either fixed or purely prompt-controlled). The three Google Lyria models take none of these — every one of genre/length/vocals is controlled purely through the natural-language prompt string:

To control (Lyria)Say in the prompt
Instrumental only"Instrumental only, no vocals."
Song structureEmbed [Verse]/[Chorus]/[Bridge] tags directly in the prompt.
Timed sectionsTimestamp syntax, e.g. [0:00 - 0:10] Intro: ....
Approximate length"...a 2-minute song..." — lyria-3.5/lyria-3-pro-preview only; lyria-3-clip-preview is always a fixed 30 seconds.

Response

response_format accepts mp3 (default) or wav — the WAV toggle only takes effect for the Google Lyria rows; every other provider always returns MP3 regardless of what's requested. The response body is the raw audio file — same "binary bytes, not JSON" shape as text-to-speech.

Not yet supported

No image-conditioned generation (Lyria's API also accepts an image alongside the text prompt for mood/style grounding — not exposed here yet), no streaming playback (the full track is generated and returned in one response, even for the providers whose upstream API is async internally), no reference-audio-conditioned generation (Mureka's own voice-cloning/reference-track parameters aren't exposed yet), and real Suno V5/V5.5 aren't available at all — atlascloud/suno-chirp-v6* is a different, newer Suno model generation, not those specific leaderboard versions.