Realtime
One WebSocket endpoint serving every full-duplex speech-to-speech model on this platform — GPT Realtime
(OpenAI direct and Azure AI Foundry), Grok Voice (xAI), Gemini Live (Google), and GPT Live 1's own
dedicated protocol (see below). Switching models is a
?model= change on the connection URL,
not a different endpoint — but each model still speaks its own upstream vendor's native event protocol once
connected; this is a routing convenience, not a protocol translator. Full guide with context on why:
Realtime guide →
WS/v1/realtime
Connect with Authorization: Bearer llmr_sk_live_... as a
request header on the WebSocket handshake (same key you use everywhere else). The connection is gated once,
at connect time (scope, RPM/TPM, monthly spend cap, credit balance) — a rejected connection closes
immediately with a 4400-range close code and a plain-text reason
(4400 unsupported model/provider,
4401 invalid key,
4403 missing scope,
4402 spend cap / no credit,
4429 rate limited,
4502 upstream provider not
configured/reachable).
Model & provider selection
?model= picks the model (defaults to
gpt-realtime-2 if omitted);
?provider= picks which real
inference provider serves it, for the models with more than one — omit it to use that model's default.
An unsupported model or an unsupported model/provider pair both close with
4400.
gpt-live-1openaiopenai — its own protocol (session.start/session.closed), not this endpoint's usual session.update/response.done shape.gemini-3.8-livegooglegoogle — own protocol (setup/serverContent), auth via ?key= not a header.gemini-3.8-live-extended-thinkinggooglegoogle — own protocol (setup/serverContent), auth via ?key= not a header.gpt-realtime-2openaiopenai — default if ?model= is omitted too.gpt-realtime-2.1openai, azureazuregpt-realtime-2.1-miniopenai, azureazuregpt-realtime-translateopenaiopenai — billed per session-minute, not per token.gpt-realtime-whisperopenaiopenai — billed per session-minute, not per token.grok-voice-think-fast-2.0xaixai — billed per session-minute, not per token.Event protocol
Every frame forwards verbatim in both directions — this is a thin proxy, not a reimplementation of any
vendor's protocol, so any client written against the real API works here unmodified with just a
base_url/key change. Which
protocol that is depends entirely on the chosen model, not on this endpoint. Three distinct shapes:
1. GPT Realtime & Grok Voice — OpenAI's GA Realtime shape
gpt-realtime-2/-2.1/-2.1-mini,
the narrower single-purpose gpt-realtime-translate
(streaming speech-to-speech translation) and gpt-realtime-whisper
(streaming speech-to-text), and grok-voice-think-fast-2.0 —
xAI built Grok Voice compatible with this exact vocabulary (live-verified 2026-09-15).
session.updatesession.instructions, .voice, .modalities, .turn_detection (or null), .toolsinput_audio_buffer.append / .commit / .clearbase64 PCM16 audio; manual commit/clear when turn_detection is offconversation.item.create / .delete / .truncateitem.role, item.content[]response.create / .canceloptional per-response .modalities/.instructions overridesession.created / .updatedpushed on connect and after every session.updateinput_audio_buffer.speech_started / .speech_stopped / .committedserver VAD lifecycleconversation.item.created / .donehistory item added/finalizedresponse.created, .output_item.added/.done, .content_part.added/.doneresponse lifecycle scaffoldingresponse.output_audio.delta/.done, .output_audio_transcript.delta/.done, .output_text.delta/.donestreamed output contentresponse.function_call_arguments.delta/.donestreamed tool-call argumentsresponse.doneterminal — carries response.usage (GPT Realtime bills from this; Grok Voice, gpt-realtime-translate, and gpt-realtime-whisper never populate it, all three billed by session duration instead)pingGrok Voice only — keepalive, harmless to ignoreerrorerror.message/.code2. Gemini Live — Google's own shape
gemini-3.8-live /
gemini-3.8-live-extended-thinking
— client speaks FIRST (setup), not server-first like the shape above. Auth
to Google is a ?key= query param internally, not a header — irrelevant to
you, you always authenticate to this platform the same way regardless of model.
setupfirst message only — .model ("models/gemini-3.8-live"), .generationConfig.responseModalities (must be ["AUDIO"] — ["TEXT"] 1007s, live-confirmed), .systemInstruction, .toolsclientContent.turns[], .turnComplete: truerealtimeInput.audio/.video/.text, .activityStart/.activityEnd, .audioStreamEndtoolResponse.functionResponses[]setupCompleteack, empty objectsessionResumptionUpdatereconnect handle, pushed unpromptedserverContent.modelTurn.parts[], .generationComplete, .turnComplete, .interrupted, .inputTranscription/.outputTranscriptiontoolCall / toolCallCancellationfunction-call request / cancellationgoAway.timeLeft before server disconnectsusageMetadatasibling key on ANY message above, not standalone — cumulative for the session; this platform keeps only the latest snapshot and bills off it once at teardown3. GPT Live 1 — its own shape
Client speaks first with session.start
— session.model is
force-rewritten to "gpt-live-1"
server-side regardless of what's sent, since this is the only model billed this way.
session.startfirst message; optional session.delegation.responses (OpenAI's Responses API request shape, passed through untouched — this platform only reads .model, for billing)session.input_audio.appendbase64-encoded audio chunkssession.closerequest closesession.startedconfirms the session, echoes config (incl. delegation) back verbatim — live-verified byte-for-byte 2026-09-15session.output_audio.deltabase64-encoded audio chunks backsession.usage.updatedrunning usage.seconds snapshot mid-sessionsession.closedauthoritative usage.seconds — billed as max(that, wall-clock elapsed), never trusted alone (a real invoice under-counted ~1.8% otherwise)session.update, transcript deltas, delegated response.*, error — passes through to/from OpenAI unmodified
delegation.responses fields
OpenAI's own docs list: model, instructions,
reasoning.effort, text.verbosity,
max_output_tokens, parallel_tool_calls — every
field forwards untouched, no validation on our side. If model doesn't
match this platform's registry, its token usage is silently never billed (voice-minutes still are).
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=gpt-live-1",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
await ws.send(json.dumps({
"type": "session.start",
"session": {
"instructions": "Be concise.",
"audio": {"format": {"type": "audio/pcm", "rate": 24000}, "output": {"voice": "marin"}},
"delegation": {"type": "responses", "responses": {"model": "gpt-5.6-luna"}}, # optional
},
}))
print(json.loads(await ws.recv())) # -> {"type": "session.started", ...}
await ws.send(json.dumps({"type": "session.close"}))
print(json.loads(await ws.recv())) # -> {"type": "session.closed", "usage": {"seconds": ...}, ...}
asyncio.run(main())
Example (Python) — gpt-realtime-2
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=gpt-realtime-2",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
print(json.loads(await ws.recv())) # -> {"type": "session.created", ...}
await ws.send(json.dumps({
"type": "conversation.item.create",
"item": {"type": "message", "role": "user",
"content": [{"type": "input_text", "text": "Say OK."}]},
}))
await ws.send(json.dumps({"type": "response.create"}))
async for message in ws:
event = json.loads(message)
if event["type"] == "response.done":
print(event["response"]["usage"])
break
asyncio.run(main())
Swap the URL's ?model= and send
that model's own event shapes per the tables above to reach any other model on this same endpoint —
Grok Voice below is a drop-in swap since it shares the shape above; Gemini Live needs its own client
logic since the shape is genuinely different.
Example (Python) — grok-voice-think-fast-2.0
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=grok-voice-think-fast-2.0&provider=xai",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
async for message in ws:
event = json.loads(message)
if event["type"] == "session.created":
await ws.send(json.dumps({
"type": "session.update",
"session": {"voice": "eve", "instructions": "Be concise.", "turn_detection": None},
}))
elif event["type"] == "session.updated":
await ws.send(json.dumps({
"type": "conversation.item.create",
"item": {"type": "message", "role": "user",
"content": [{"type": "input_text", "text": "Say OK."}]},
}))
await ws.send(json.dumps({"type": "response.create"}))
elif event["type"] == "response.done":
print(event["response"]) # usage is always {} — billed by session duration, not tokens
break
asyncio.run(main())
Example (Python) — gemini-3.8-live
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=gemini-3.8-live&provider=google",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
await ws.send(json.dumps({
"setup": {"model": "models/gemini-3.8-live",
"generationConfig": {"responseModalities": ["AUDIO"]}}, # TEXT is rejected
}))
sent = False
async for message in ws:
event = json.loads(message)
if "setupComplete" in event and not sent:
await ws.send(json.dumps({
"clientContent": {"turns": [{"role": "user", "parts": [{"text": "Say OK."}]}],
"turnComplete": True},
}))
sent = True
if event.get("serverContent", {}).get("turnComplete"):
print(event.get("usageMetadata")) # real token usage, per modality
break
asyncio.run(main())
Billing
Two shapes, depending on the model — plus the flat 2% platform fee on top either way, same margin as every other endpoint, see /pricing:
- Per token — GPT Realtime and Gemini Live: billed from the real response's usage (three separately-priced modalities — text/audio/image — each with its own input/cached-input/output rate; unaccounted tokens bill at the most expensive known rate rather than being dropped).
- Per session-minute — Grok Voice and GPT Live 1: billed off wall-clock session duration rather than a token count (neither exposes an in-band usage object this platform trusts). GPT Live 1 additionally bills any delegated backend model's own token usage separately — see the GPT Live 1 section above.