Discord

Multimodal

Speech-to-text

Multipart upload, same shape as OpenAI's audio.transcriptions — POST /v1/audio/transcriptions. Billed per minute of audio — the real provider-reported duration, not a size-based estimate.

import requests

with open("call.wav", "rb") as f:
    resp = requests.post(
        "https://videorouter.sh/api/v1/audio/transcriptions",
        headers={"Authorization": "Bearer llmr_sk_live_..."},
        files={"file": f},
        data={"model": "whisper-1"},
    )
print(resp.json()["text"])

Which models

ModelProvider
whisper-1OpenAI (default)
gpt-4o-transcribeOpenAI
gpt-4o-mini-transcribeOpenAI
scribe_v1ElevenLabs
gemini-3.5-transcribeGoogle

gemini-3.5-transcribe is billed differently from every other model in this table — genuinely per-token (input audio tokens + output text tokens), not per-minute — see Pricing & billing.

Response format

response_format accepts json (default, {"text": ...}), text (raw body), and verbose_json (full provider object including duration and segments). We always request verbose_json from the upstream provider regardless of what you asked for, purely to get a real duration to bill on, then reshape the response to match.

Not yet supported

srt and vtt response formats aren't built — they need real reformatting, not just a passthrough of the provider response.