Multimodal
Speech-to-text
Multipart upload, same shape as OpenAI's
audio.transcriptions —
POST /v1/audio/transcriptions.
Billed per minute of audio — the real provider-reported duration, not a size-based estimate.
import requests
with open("call.wav", "rb") as f:
resp = requests.post(
"https://videorouter.sh/api/v1/audio/transcriptions",
headers={"Authorization": "Bearer llmr_sk_live_..."},
files={"file": f},
data={"model": "whisper-1"},
)
print(resp.json()["text"])
Which models
whisper-1OpenAI (default)gpt-4o-transcribeOpenAIgpt-4o-mini-transcribeOpenAIscribe_v1ElevenLabsgemini-3.5-transcribeGoogle
gemini-3.5-transcribe is billed
differently from every other model in this table — genuinely per-token (input audio tokens + output
text tokens), not per-minute — see Pricing & billing.
Response format
response_format accepts
json (default,
{"text": ...}),
text (raw body), and
verbose_json (full provider object
including duration and
segments). We always request
verbose_json from the upstream provider
regardless of what you asked for, purely to get a real duration to bill on, then reshape the response to match.
Not yet supported
srt and
vtt response formats aren't built —
they need real reformatting, not just a passthrough of the provider response.