Multimodal
Music generation
POST /v1/audio/music —
text-to-music across 5 providers: Google Lyria, Mureka, ElevenLabs, StepFun, and Suno Chirp
(via Atlas Cloud). Returns raw audio bytes in one response (no job/poll, no streaming). Most
models are billed flat per generated song, regardless of input length or output
duration — ElevenLabs is the one exception, billed per minute of requested track length (see
Which models below).
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/audio/music",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "lyria-3.5",
"prompt": "An upbeat acoustic guitar folk song about road trips, 2 minutes, with vocals",
"response_format": "mp3",
},
)
with open("song.mp3", "wb") as f:
f.write(resp.content)
Which models
lyria-3.5GoogleFull song (verses, choruses, bridges) — a couple of minutes, controllable via prompt$0.08 / songlyria-3-clip-previewGoogleShort clip/loop — fixed 30 seconds$0.04 / songlyria-3-pro-previewGoogleFull song, prompt-controlled$0.08 / songmureka-v9MurekaFull song, prompt/lyrics-controlled$0.45 / songmureka-v8MurekaFull song, prompt/lyrics-controlled$0.45 / songelevenlabs-music-v2.5ElevenLabsSet via duration in seconds — defaults to 30s if omitted$0.15 / minutestepaudio-3-musicStepFunFull song, prompt-controlledFree (preview)atlascloud/suno-chirp-v6Suno (via Atlas Cloud)Full song, prompt-controlled$0.132 / songatlascloud/suno-chirp-v6-wildSuno (via Atlas Cloud)Full song, prompt-controlled, more adventurous style adherence$0.132 / songatlascloud/suno-chirp-v6-miniSuno (via Atlas Cloud)Full song, prompt-controlled, lighter/faster tier$0.132 / song
Every model except elevenlabs-music-v2.5 is billed flat per
generated track, regardless of how long the output ends up — there's no per-second or
per-character component the way text-to-speech
is billed. ElevenLabs is billed $0.15 per minute of the
duration you request.
Controlling genre, vocals & length
Three optional request fields work across the non-Lyria providers:
lyrics (Mureka, StepFun —
your own lyrics text, section tags like [Verse]/[Chorus] supported),
instrumental (boolean, every
provider except the Lyria rows), and
duration (seconds — ElevenLabs
only; every other provider's length is either fixed or purely prompt-controlled). The three Google
Lyria models take none of these — every one of genre/length/vocals is controlled purely through the
natural-language prompt string:
[Verse]/[Chorus]/[Bridge] tags directly in the prompt.[0:00 - 0:10] Intro: ....lyria-3.5/lyria-3-pro-preview only; lyria-3-clip-preview is always a fixed 30 seconds.Response
response_format accepts
mp3 (default) or
wav — the WAV toggle only
takes effect for the Google Lyria rows; every other provider always returns MP3 regardless of what's
requested. The response body is the raw audio file — same "binary bytes, not JSON" shape as
text-to-speech.
Not yet supported
No image-conditioned generation (Lyria's API also accepts an image alongside the text prompt for
mood/style grounding — not exposed here yet), no streaming playback (the full track is generated
and returned in one response, even for the providers whose upstream API is async internally), no
reference-audio-conditioned generation (Mureka's own voice-cloning/reference-track parameters
aren't exposed yet), and real Suno V5/V5.5 aren't available at all — atlascloud/suno-chirp-v6*
is a different, newer Suno model generation, not those specific leaderboard versions.