Multimodal
Text-to-speech
Same shape as OpenAI's audio.speech —
POST /v1/audio/speech. Returns raw audio
bytes in one response (no streaming playback yet). Most models bill per input character;
gpt-4o-mini-tts and the Gemini TTS family below
bill on real per-call token usage instead (input text tokens + output audio tokens, at very different rates) —
see each model's page for its exact rate.
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/audio/speech",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "tts-1",
"voice": "alloy",
"input": "Your order has shipped.",
"response_format": "mp3",
},
)
with open("out.mp3", "wb") as f:
f.write(resp.content)
Which models and voices
tts-1OpenAIe.g. alloy (default)tts-1-hdOpenAIOpenAI voice idsgpt-4o-mini-ttsOpenAIOpenAI voice idsspeech-02-hdMiniMaxMiniMax voice ids — required, no defaultspeech-02-turboMiniMaxMiniMax voice ids — required, no defaulteleven_multilingual_v2ElevenLabsElevenLabs voice ids — required, no defaulteleven_v3ElevenLabsElevenLabs voice ids — required, no defaultgemini-3.1-flash-tts-previewGoogleGemini prebuilt voice names, e.g. Kore (default)gemini-3.8-flash-ttsGoogleGemini prebuilt voice names, e.g. Kore (default)gemini-3.8-flash-lite-ttsGoogleGemini prebuilt voice names, e.g. Kore (default)
voice only defaults to
"alloy" for OpenAI models and
"Kore" for the Gemini TTS family — MiniMax and
ElevenLabs use their own voice-id vocabulary and a request against those models must set
voice explicitly. Gemini has 30 prebuilt
voice names (no custom/cloned voices); an OpenAI voice id like alloy
will 400 against a Gemini model.
response_format accepts
mp3 (default),
opus,
aac,
flac,
wav, and
pcm.
Audio output in chat completions
A second way to get spoken audio back: pass modalities: ["audio", "text"] to
POST /v1/chat/completions against one of
OpenAI's gpt-audio family models — the model
hears/speaks as part of the conversation, rather than just converting a fixed string. The reply comes back on
message.audio.data (base64, matching
audio.format), with
message.audio.transcript alongside it.
Streaming is not supported for this mode.
from openai import OpenAI
client = OpenAI(base_url="https://videorouter.sh/api/v1", api_key="llmr_sk_live_...")
resp = client.chat.completions.create(
model="gpt-audio-mini",
messages=[{"role": "user", "content": "Say a friendly hello."}],
modalities=["audio", "text"],
audio={"voice": "alloy", "format": "wav"},
)
message = resp.choices[0].message
print(message.audio.transcript)
print(f"audio bytes (base64): {len(message.audio.data)}")
Three models today, all OpenAI — gpt-audio,
gpt-audio-1.5, and the cheaper
gpt-audio-mini. Billed on real per-call
usage, split by text vs. audio tokens on both the input and output side (audio tokens cost substantially more
than text tokens) — see each model's page for its exact rates.
Not yet supported
No real-time/streamed audio output outside chat completions — a plain /v1/audio/speech call
generates and returns the full file in one response. For full-duplex voice conversation, see
Realtime. For transcription (audio → text) see
Speech-to-text.