Overview
Realtime
One full-duplex voice-and-text WebSocket, WS /v1/realtime,
serving four model families — GPT Realtime (OpenAI direct
and Azure AI Foundry), Grok Voice (xAI),
Gemini Live (Google), and
GPT Live 1 (OpenAI, also reachable on
its own dedicated endpoint). Pick the model with ?model=
on the connection URL — switching models is just that query-param change, not a different endpoint.
What does NOT change is the event vocabulary you speak once connected. Three distinct shapes:
GPT Realtime and Grok Voice share one OpenAI-shaped protocol
(session.update,
response.create, streamed
response.* events); Gemini Live
speaks its own
(setup,
clientContent,
serverContent); and GPT Live 1
speaks yet another
(session.start,
session.closed). Full event
tables and parameter details for each are below.
This is a routing convenience, not a protocol translator: every frame forwards verbatim to/from the real
upstream, so any client already built against one of these vendors' own SDKs works here unmodified with
just a base_url/key change.
Models & providers
Select a model with ?model= on the
connection URL — an unrecognized model closes the socket immediately with code
4400, no silent fallback.
gpt-realtime-2.1 and
gpt-realtime-2.1-mini are each
available from two independent inference providers — OpenAI directly, and this platform's own Azure AI
Foundry deployment. Pick one explicitly with
?provider=openai or
?provider=azure; omit it to use
that model's default (an unsupported model/provider pair is also a
4400). Every model below shares
this one endpoint — switching between them (including
gpt-live-1) is just a
?model= change — though each
model still speaks its own upstream vendor's wire protocol once connected, so a client still needs to
send the right event shapes for whichever one it picks; this is a routing convenience, not a protocol
translator.
gpt-live-1openaiopenai — its own protocol (session.start/session.closed), not this endpoint's usual session.update/response.done shape.gemini-3.8-livegooglegoogle — own protocol (setup/serverContent), auth via ?key= not a header.gemini-3.8-live-extended-thinkinggooglegoogle — own protocol (setup/serverContent), auth via ?key= not a header.gpt-realtime-2openaiopenai — default if ?model= is omitted too.gpt-realtime-2.1openai, azureazuregpt-realtime-2.1-miniopenai, azureazuregpt-realtime-translateopenaiopenai — billed per session-minute, not per token.gpt-realtime-whisperopenaiopenai — billed per session-minute, not per token.grok-voice-think-fast-2.0xaixai — billed per session-minute, not per token.GPT Realtime (OpenAI & Azure)
gpt-realtime-2,
gpt-realtime-2.1, and
gpt-realtime-2.1-mini speak
OpenAI's own GA Realtime API protocol verbatim — any client built against the real API works here
unmodified with just a base_url/key
change. Configure the session with session.update,
stream input with input_audio_buffer.append,
and trigger a reply with response.create.
Two narrower single-purpose variants speak this exact same protocol —
gpt-realtime-translate
(streaming speech-to-speech translation, 70+ input languages → 13 output languages) and
gpt-realtime-whisper
(streaming speech-to-text) — so everything below (events, example) applies to them unchanged, just
swap the ?model= value.
The one real difference is billing: neither returns a response.usage
block, so both are billed per session-minute rather than per token — see the
Models & providers table above.
Client → server events
session.updateConfigure the session — session.instructions, .voice, .modalities (["text"]/["audio"]), .turn_detection (server VAD config, or null to disable and commit manually), .tools. Can be sent any time, not just first.input_audio_buffer.appendBase64 PCM16 audio chunk in .audio.input_audio_buffer.commit / .clearManually end the input turn (when turn_detection is disabled) / discard buffered audio.conversation.item.createAdd a message to history — item.role, item.content[] (input_text/input_audio/function_call_output parts).conversation.item.delete / .truncateRemove a history item, or cut an assistant audio item short (barge-in).response.createAsk the model to generate a response now. Optional response.modalities/.instructions override the session default for this one response.response.cancelStop an in-flight response (e.g. on user barge-in).Server → client events
session.created / .updatedPushed immediately on connect, and again after every session.update.input_audio_buffer.speech_started / .speech_stopped / .committedServer-side VAD lifecycle, only fires when turn_detection is enabled.conversation.item.created / .doneEchoes a history item (yours or the model's) as it's added/finalized.response.createdA response started generating.response.output_item.added / .doneOne output item (a message) started/finished within the response.response.content_part.added / .doneOne content part (text or audio) within an output item.response.output_audio.delta / .doneStreamed base64 PCM16 audio chunks, then completion.response.output_audio_transcript.delta / .doneStreamed transcript of the audio being generated.response.output_text.delta / .doneStreamed text, for text-modality responses.response.function_call_arguments.delta / .doneStreamed tool-call arguments, if the model invokes a tool.response.doneTerminal event — carries the authoritative response.usage block this platform bills from (see Billing).errorMalformed event or an upstream failure — carries error.message/.code.
Live-verified 2026-09-15 against real OpenAI (?provider=openai)
and real Azure AI Foundry (?provider=azure).
Example
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=gpt-realtime-2.1",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
print(json.loads(await ws.recv())) # -> {"type": "session.created", ...}
await ws.send(json.dumps({"type": "session.update",
"session": {"type": "realtime", "instructions": "You are a helpful assistant."}}))
# ... input_audio_buffer.append chunks, then response.create ...
await ws.send(json.dumps({"type": "response.create"}))
print(json.loads(await ws.recv())) # -> response.* stream, ending in response.done + usage
asyncio.run(main())
Grok Voice (xAI)
grok-voice-think-fast-2.0 speaks the
same event vocabulary as GPT Realtime above (xAI built theirs to be compatible) — session.update,
conversation.item.create,
response.create all work exactly the
same way. Two real differences, both live-verified 2026-09-15:
- An extra
pingkeepalive event arrives unprompted (harmless to ignore) — not part of OpenAI's own vocabulary. response.done'susageblock is always{}— confirmed on a real call, not documentation. This platform bills Grok Voice by wall-clock session duration instead (see Billing), never from that empty object.
Example (Python)
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=grok-voice-think-fast-2.0&provider=xai",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
async for message in ws:
event = json.loads(message)
if event["type"] == "session.created":
await ws.send(json.dumps({
"type": "session.update",
"session": {"voice": "eve", "instructions": "Be concise.", "turn_detection": None},
}))
elif event["type"] == "session.updated":
await ws.send(json.dumps({
"type": "conversation.item.create",
"item": {"type": "message", "role": "user",
"content": [{"type": "input_text", "text": "Say OK."}]},
}))
await ws.send(json.dumps({"type": "response.create"}))
elif event["type"] == "response.done":
print(event["response"]) # usage is always {} — billed by duration, not tokens
break
asyncio.run(main())
26 built-in voices (default eve), 20
languages, PCM/µ-law/A-law/Opus audio formats — see
xAI's own reference for the full parameter list on session.update; every field passes through unmodified.
Gemini Live (Google)
gemini-3.8-live and
gemini-3.8-live-extended-thinking
speak Google's own protocol — genuinely different from everything above, not OpenAI-shaped at all. The
client sends setup FIRST
(server-first handshakes like GPT Realtime's don't apply here), then
clientContent to send a complete
turn.
Client → server events
setupFirst message only. setup.model ("models/gemini-3.8-live"), .generationConfig.responseModalities — must be ["AUDIO"], ["TEXT"] is rejected by Google with a 1007 close for this model (live-confirmed) — .systemInstruction, .tools.clientContent.turns[] ({role, parts:[{text}]}), .turnComplete: true to signal end-of-turn and trigger a response.realtimeInputContinuous streaming input — .audio/.video/.text, .activityStart/.activityEnd for manual VAD, .audioStreamEnd.toolResponse.functionResponses[] replying to a toolCall.Server → client events
setupCompleteAck for setup — empty object.sessionResumptionUpdateA reconnect handle, pushed proactively (unprompted) after setup.serverContent.modelTurn.parts[] (streamed audio/text), .generationComplete, .turnComplete, .interrupted, .inputTranscription/.outputTranscription.toolCall / toolCallCancellationModel wants to call a function / cancels a prior call.goAwayServer will disconnect soon — carries .timeLeft.usageMetadataNOT a standalone event — rides alongside any of the messages above as a sibling key. This platform keeps only the LATEST snapshot seen (it's cumulative for the session) and bills off that once at teardown — see Billing.Example (Python)
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=gemini-3.8-live&provider=google",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
await ws.send(json.dumps({
"setup": {"model": "models/gemini-3.8-live",
"generationConfig": {"responseModalities": ["AUDIO"]}}, # TEXT is rejected
}))
sent = False
async for message in ws:
event = json.loads(message)
if "setupComplete" in event and not sent:
await ws.send(json.dumps({
"clientContent": {"turns": [{"role": "user", "parts": [{"text": "Say OK."}]}],
"turnComplete": True},
}))
sent = True
if event.get("serverContent", {}).get("turnComplete"):
print(event.get("usageMetadata")) # real token usage, per modality
break
asyncio.run(main())
Auth to Google happens internally via a ?key=
query param on our side of the connection, not a header — irrelevant to you, you always authenticate to
this platform with Authorization: Bearer
regardless of model. Input audio is 16-bit PCM at 16kHz, output at 24kHz.
GPT Live 1
Reachable here via ?model=gpt-live-1
— the one model on this page that speaks a genuinely different protocol: the client
sends session.start first, and
session.model is force-rewritten to
"gpt-live-1" server-side regardless
of what's sent, since this is the only model billed this way.
Session lifecycle
session.startclient → server, first message. Optionally set session.delegation.responses.model to hand reasoning/tool calls to a separate backend chat model (see below).session.startedserver → client, confirms the session is live — echoes the (possibly delegation-bearing) session config back.session.input_audio.appendclient → server, base64-encoded audio chunks.session.output_audio.deltaserver → client, base64-encoded audio chunks back.session.usage.updatedserver → client, a running usage.seconds snapshot mid-session.session.close / session.closedclient requests close / server confirms — session.closed carries the authoritative usage.seconds for the session, billed as max(that, wall-clock elapsed) (see Billing).
Every other event — session.update,
transcript deltas, delegated response.*
events, error — passes through to/from
OpenAI unmodified; this is a thin proxy, not a reimplementation of the protocol.
The delegation parameter
session.delegation.responses takes
OpenAI's own Responses API request shape, passed through untouched — this platform only reads
model (for billing); every other
field forwards to OpenAI exactly as sent, with no validation on our side. Live-verified 2026-09-15 —
OpenAI's real server echoes every field back byte-for-byte in session.started:
{
"type": "session.start",
"session": {
"delegation": {
"type": "responses",
"responses": {
"model": "gpt-5.6-luna",
"instructions": "You are the backend reasoning model for a realtime voice partner.",
"parallel_tool_calls": false,
"reasoning": {"effort": "none"},
"text": {"verbosity": "low"},
"max_output_tokens": 120
}
}
}
}
If model doesn't match a row in this
platform's model registry, its token usage is simply never billed (voice-minutes still are) — not an
error, a silent $0 for that piece. Omit delegation entirely and GPT Live 1
handles the whole conversation itself, no backend model involved.
Example (Python)
import asyncio, json, websockets
async def main():
async with websockets.connect(
"wss://videorouter.sh/v1/realtime?model=gpt-live-1",
additional_headers={"Authorization": "Bearer llmr_sk_live_..."},
) as ws:
await ws.send(json.dumps({
"type": "session.start",
"session": {
"instructions": "Be concise.",
"audio": {"format": {"type": "audio/pcm", "rate": 24000}, "output": {"voice": "marin"}},
# Optional — omit for GPT Live 1 to reason for itself:
"delegation": {"type": "responses", "responses": {"model": "gpt-5.6-luna"}},
},
}))
print(json.loads(await ws.recv())) # -> {"type": "session.started", ...}
# ... send session.input_audio.append chunks, read session.output_audio.delta back ...
await ws.send(json.dumps({"type": "session.close"}))
print(json.loads(await ws.recv())) # -> {"type": "session.closed", "usage": {"seconds": ...}, ...}
asyncio.run(main())
Live-verified 2026-09-15 against real OpenAI, including a real delegation
payload — everything above is what actually happened on the wire, not documentation paraphrased from
OpenAI's own docs.
Auth & gating
Connect with Authorization: Bearer llmr_sk_live_...
as a request header on the WebSocket handshake — the same key you use everywhere else. The connection is
gated once, at connect time (scope, RPM/TPM, monthly spend cap, credit balance) — a rejected connection
closes immediately with a 4400-range close code and a short reason
(4400 unsupported model,
4401 invalid key,
4403 missing scope,
4402 spend cap / no credit,
4429 rate limited). A very long
session can't be re-gated mid-flight, the same limitation an ordinary streaming chat completion already has.
Billing
Two shapes, depending on the model — plus the flat 2% platform fee either way, see Pricing & billing:
- GPT Realtime, Gemini Live — per token. Split across three separately-priced
modalities (text, audio, image), each with its own input/cached-input/output rate. GPT Realtime reads
this from every
response.doneevent'sresponse.usageblock, summed for the whole session. Gemini Live reads the LATESTusageMetadatasnapshot seen on any message (it's cumulative, not incremental) and bills that once at teardown. A response still in flight when the connection drops isn't billed for; one that already finished stays billed regardless of what happens afterward. - Grok Voice, GPT Live 1 — per session-minute. Billed off wall-clock session duration,
not a token count — neither model exposes an in-band usage object this platform trusts (Grok Voice's
response.done.usageis always empty, confirmed live). GPT Live 1 additionally bills anydelegationbackend model's own token usage separately — see the GPT Live 1 section above.