Discord
此页面暂无中文版本 — 以下显示英文内容。 查看英文版

← API reference

Audio

POST/v1/audio/speech

Parameters

ParameterTypeDescription
inputstringRequired. Text to speak.
modelstringOptional, default "tts-1" (also the "auto" resolution).
voicestringOptional for OpenAI models (defaults "alloy") and the Gemini TTS family (defaults "Kore", one of 30 Gemini prebuilt names) — required for MiniMax/ElevenLabs models, which use their own voice-id vocabulary (400 if omitted with no default). Voice vocabularies aren't interchangeable across providers.
response_formatstringOptional, default "mp3". Also opus/aac/flac/wav/pcm.
speednumberOptional, forwarded as-is when present.

Response

Raw audio bytes in the requested response_format, Content-Type set accordingly (e.g. audio/mpeg for mp3) — not a JSON envelope.

Billed per input character, except gpt-4o-mini-tts and the Gemini TTS family (gemini-3.1-flash-tts-preview, gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts), which bill on real per-call token usage instead — see Text-to-speech.

POST/v1/audio/transcriptions

multipart/form-data, not JSON.

Parameters

ParameterTypeDescription
filefileRequired. The audio file.
modelstring (form)Optional, default "whisper-1" (also the "auto" resolution).
response_formatstring (form)Optional, default "json". Also text (raw body) and verbose_json (full object incl. duration/segments). srt/vtt aren't supported yet.

Example response (response_format: "json", the default)

{ "text": "..." }

Billed on the real provider-reported audio duration (never a size-based guess when the provider supplies one) — see Speech-to-text. Exception: gemini-3.5-transcribe is billed per-token (input audio + output text tokens), not per-minute.

POST/v1/audio/music

Text-to-music across 5 providers (Google Lyria, Mureka, ElevenLabs, StepFun, Suno Chirp via Atlas Cloud). No job/poll — returns audio in one response, same shape as /v1/audio/speech.

Parameters

ParameterTypeDescription
promptstringRequired. Describes the song. For the Lyria rows, genre/mood/vocals/structure/length are all controlled here, not via separate parameters. See Music generation.
modelstringOptional, default "lyria-3.5" (also the "auto" resolution). Also lyria-3-clip-preview, lyria-3-pro-preview, mureka-v9, mureka-v8, elevenlabs-music-v2.5, stepaudio-3-music, atlascloud/suno-chirp-v6 (and -wild/-mini).
lyricsstringOptional. Used by Mureka and StepFun only — your own lyrics text (section tags like [Verse] supported). Ignored by every other provider.
instrumentalbooleanOptional. Requests no vocals — supported by every provider except the Lyria rows (use the prompt string for those instead).
durationnumberOptional, seconds. ElevenLabs only (defaults to 30s if omitted) — this is what its per-minute price is computed from. Ignored by every other provider.
response_formatstringOptional, default "mp3". Also wav — only takes effect for the Lyria rows; every other provider always returns MP3.

Response

Raw audio bytes in the requested response_format, Content-Type set accordingly — not a JSON envelope.

Billed flat per generated song for every model except elevenlabs-music-v2.5 (billed per minute of requested duration) — see Music generation for the full price table.