Audio
POST/v1/audio/speech
Parameters
inputstringRequired. Text to speak.modelstringOptional, default "tts-1" (also the "auto" resolution).voicestringOptional for OpenAI models (defaults "alloy") and the Gemini TTS family (defaults "Kore", one of 30 Gemini prebuilt names) — required for MiniMax/ElevenLabs models, which use their own voice-id vocabulary (400 if omitted with no default). Voice vocabularies aren't interchangeable across providers.response_formatstringOptional, default "mp3". Also opus/aac/flac/wav/pcm.speednumberOptional, forwarded as-is when present.Response
Raw audio bytes in the requested response_format, Content-Type set accordingly (e.g. audio/mpeg for mp3) — not a JSON envelope.
Billed per input character, except gpt-4o-mini-tts and the Gemini TTS family (gemini-3.1-flash-tts-preview, gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts), which bill on real per-call token usage instead — see Text-to-speech.
POST/v1/audio/transcriptions
multipart/form-data, not JSON.
Parameters
filefileRequired. The audio file.modelstring (form)Optional, default "whisper-1" (also the "auto" resolution).response_formatstring (form)Optional, default "json". Also text (raw body) and verbose_json (full object incl. duration/segments). srt/vtt aren't supported yet.Example response (response_format: "json", the default)
{ "text": "..." }
Billed on the real provider-reported audio duration (never a size-based guess when the provider supplies one) — see Speech-to-text. Exception: gemini-3.5-transcribe is billed per-token (input audio + output text tokens), not per-minute.
POST/v1/audio/music
Text-to-music across 5 providers (Google Lyria, Mureka, ElevenLabs, StepFun, Suno Chirp via Atlas Cloud). No job/poll — returns audio in one response, same shape as /v1/audio/speech.
Parameters
promptstringRequired. Describes the song. For the Lyria rows, genre/mood/vocals/structure/length are all controlled here, not via separate parameters. See Music generation.modelstringOptional, default "lyria-3.5" (also the "auto" resolution). Also lyria-3-clip-preview, lyria-3-pro-preview, mureka-v9, mureka-v8, elevenlabs-music-v2.5, stepaudio-3-music, atlascloud/suno-chirp-v6 (and -wild/-mini).lyricsstringOptional. Used by Mureka and StepFun only — your own lyrics text (section tags like [Verse] supported). Ignored by every other provider.instrumentalbooleanOptional. Requests no vocals — supported by every provider except the Lyria rows (use the prompt string for those instead).durationnumberOptional, seconds. ElevenLabs only (defaults to 30s if omitted) — this is what its per-minute price is computed from. Ignored by every other provider.response_formatstringOptional, default "mp3". Also wav — only takes effect for the Lyria rows; every other provider always returns MP3.Response
Raw audio bytes in the requested response_format, Content-Type set accordingly — not a JSON envelope.
Billed flat per generated song for every model except elevenlabs-music-v2.5 (billed per minute of requested duration) — see Music generation for the full price table.