Discord
此页面暂无中文版本 — 以下显示英文内容。 查看英文版

Multimodal

Text-to-speech

Same shape as OpenAI's audio.speech — POST /v1/audio/speech. Returns raw audio bytes in one response (no streaming playback yet). Most models bill per input character; gpt-4o-mini-tts and the Gemini TTS family below bill on real per-call token usage instead (input text tokens + output audio tokens, at very different rates) — see each model's page for its exact rate.

import requests

resp = requests.post(
    "https://videorouter.sh/api/v1/audio/speech",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "tts-1",
        "voice": "alloy",
        "input": "Your order has shipped.",
        "response_format": "mp3",
    },
)
with open("out.mp3", "wb") as f:
    f.write(resp.content)

Which models and voices

ModelProviderVoice id
tts-1OpenAIe.g. alloy (default)
tts-1-hdOpenAIOpenAI voice ids
gpt-4o-mini-ttsOpenAIOpenAI voice ids
speech-02-hdMiniMaxMiniMax voice ids — required, no default
speech-02-turboMiniMaxMiniMax voice ids — required, no default
eleven_multilingual_v2ElevenLabsElevenLabs voice ids — required, no default
eleven_v3ElevenLabsElevenLabs voice ids — required, no default
gemini-3.1-flash-tts-previewGoogleGemini prebuilt voice names, e.g. Kore (default)
gemini-3.8-flash-ttsGoogleGemini prebuilt voice names, e.g. Kore (default)
gemini-3.8-flash-lite-ttsGoogleGemini prebuilt voice names, e.g. Kore (default)

voice only defaults to "alloy" for OpenAI models and "Kore" for the Gemini TTS family — MiniMax and ElevenLabs use their own voice-id vocabulary and a request against those models must set voice explicitly. Gemini has 30 prebuilt voice names (no custom/cloned voices); an OpenAI voice id like alloy will 400 against a Gemini model. response_format accepts mp3 (default), opus, aac, flac, wav, and pcm.

Audio output in chat completions

A second way to get spoken audio back: pass modalities: ["audio", "text"] to POST /v1/chat/completions against one of OpenAI's gpt-audio family models — the model hears/speaks as part of the conversation, rather than just converting a fixed string. The reply comes back on message.audio.data (base64, matching audio.format), with message.audio.transcript alongside it. Streaming is not supported for this mode.

from openai import OpenAI

client = OpenAI(base_url="https://videorouter.sh/api/v1", api_key="llmr_sk_live_...")

resp = client.chat.completions.create(
    model="gpt-audio-mini",
    messages=[{"role": "user", "content": "Say a friendly hello."}],
    modalities=["audio", "text"],
    audio={"voice": "alloy", "format": "wav"},
)
message = resp.choices[0].message
print(message.audio.transcript)
print(f"audio bytes (base64): {len(message.audio.data)}")

Three models today, all OpenAI — gpt-audio, gpt-audio-1.5, and the cheaper gpt-audio-mini. Billed on real per-call usage, split by text vs. audio tokens on both the input and output side (audio tokens cost substantially more than text tokens) — see each model's page for its exact rates.

Not yet supported

No real-time/streamed audio output outside chat completions — a plain /v1/audio/speech call generates and returns the full file in one response. For full-duplex voice conversation, see Realtime. For transcription (audio → text) see Speech-to-text.