API reference
Every route below is for sending model requests, and requires
Authorization: Bearer llmr_sk_<env>_… —
see Authentication.
Chat completions
| Endpoint | Description |
|---|---|
| POST /v1/chat/completions | OpenAI-compatible chat completions; model:"auto" routes, a slug pins — streaming, tools, automatic fallback |
| POST /v1/route | routing decision only — no upstream call, no charge |
| GET /v1/models | catalog + the auto pseudo-model, restricted to your key's allowlist |
Full request/response detail: Models & routing, Streaming, Tool calling. Full parameters & example responses: Chat completions API reference →
Image generation is also reachable here — pass modalities: ["image", "text"] to get a generated image back on message.images, billed the same rate as /v1/images below. See image output in chat completions.
Spoken audio output too — pass modalities: ["audio", "text"] to a gpt-audio model to get a reply back on message.audio. See audio output in chat completions.
Embeddings
A provider passthrough endpoint kept for drop-in SDK compatibility — not part of the chat/routing surface.
| Endpoint | Description |
|---|---|
| POST /v1/embeddings | OpenAI-compatible embeddings, billed the same as chat (provider cost + 2%) |
Full parameters & example responses: Embeddings API reference →
Images
| Endpoint | Description |
|---|---|
| POST /v1/images | unified text-to-image generation and image-to-image editing (pass input_references) — one endpoint for both, billed per image |
| GET /v1/images/models | live list of every valid model slug for POST /v1/images — bare ids only, no price/capability breakdown (see /models for that); unbilled |
Full guide: image generation → Full parameters & example responses: Images API reference →
Uploads
A helper for any caller — images, video, lip sync, face swap — that has raw file bytes rather than an
already-public URL to pass as a reference (input_references,
start_image_url,
input_video_url, etc.). A base64
data: URI works directly on those
fields too and skips this call entirely — use it for small files; use this endpoint when that would blow
past a request-body-size limit.
| Endpoint | Description |
|---|---|
| POST /v1/uploads | multipart file upload (image ≤10 MB, audio ≤20 MB, video ≤50 MB); returns a presigned url valid for 30 minutes — stage it, then immediately reference it in the same session's generate call; unbilled, not durable storage |
Audio
| Endpoint | Description |
|---|---|
| POST /v1/audio/speech | text-to-speech; returns raw audio, billed per input character (per-token for gpt-4o-mini-tts and the Gemini TTS family) |
| POST /v1/audio/transcriptions | speech-to-text, billed per minute of real audio duration (per-token for Gemini 3.5 Transcribe) |
| POST /v1/audio/music | text-to-music (Google Lyria); returns raw audio, billed flat per generated song |
Full guide: text-to-speech → · Full guide: speech-to-text → Full parameters & example responses: Audio API reference →
Realtime
Full-duplex voice, one WebSocket for GPT Realtime (OpenAI/Azure), Grok Voice (xAI), Gemini Live (Google),
and GPT Live 1 — a ?model= change,
not a different endpoint.
| Endpoint | Description |
|---|---|
| WS /v1/realtime | open a WebSocket, pick a model with ?model=, exchange that model's own event frames — billed per token (GPT Realtime/Gemini Live) or per session-minute (Grok Voice, GPT Realtime Translate/Whisper, GPT Live 1) |
Full guide: Realtime guide → Full parameters & example responses: Realtime API reference →
Video
| Endpoint | Description |
|---|---|
| POST /v1/videos | submit a video generation job — text-to-video, or image-to-video (pass start_image_url)/reference-to-video (pass input_references/input_video_references/input_audio_references) on supporting models; async, returns a video id immediately; billed per requested second at creation time |
| GET /v1/videos/models | live list of every valid model slug for POST /v1/videos — bare ids only, no price/capability breakdown (see /models for that); unbilled |
| GET /v1/videos/{id} | poll status; the finished video URL is included once status is completed |
Full guide: video generation → Full parameters & example responses: Video API reference → — also see Motion control → for Kling v3.0 Motion Control's character-image + reference-motion-video shape.
Lip sync
| Endpoint | Description |
|---|---|
| POST /v1/lipsync | a SEPARATE endpoint from /v1/videos — re-sync a face to a driving audio track, either animating a still (pass image_url + audio_url) or dubbing an existing clip's mouth to new speech (pass video_url + audio_url); async, returns a job id immediately; billed once at creation time from the driving audio's (or the longer of audio/video's) real duration |
| GET /v1/lipsync/{id} | poll status — identical job-id scheme and response shape as GET /v1/videos/{id} |
Full guide: lip sync → Full parameters & example responses: Lip sync API reference →
Face swap
| Endpoint | Description |
|---|---|
| POST /v1/face-swap | a SEPARATE endpoint from /v1/videos — replace the face throughout an existing video (video_url) with a source face photo (face_image_url); requires source_face_consent: true; async, returns a job id immediately; billed reserve-then-settle from the provider's own completed-job invoice |
| GET /v1/face-swap/{id} | poll status — identical job-id scheme and response shape as GET /v1/videos/{id} |
Full guide: face swap → Full parameters & example responses: Face swap API reference →
3D generation
| Endpoint | Description |
|---|---|
| POST /v1/meshes | submit a text-to-3d (prompt), image-to-3d (image_url), or multi-image-to-3d (image_urls) job — async, returns a job id immediately; billed flat at creation time from the model + option choice |
| GET /v1/meshes/{id} | poll status; the finished mesh (GLB + whichever of FBX/OBJ/USDZ/textures the model produces) is included once status is completed |
Full guide: 3D generation → Full parameters & example responses: 3D generation API reference →
Async Jobs
Submit a single request that's OK to wait 1 or 6 hours for a deeper discount — same body as
/v1/chat/completions, wrapped.
| Endpoint | Description |
|---|---|
| POST /v1/jobs | submit with completion_window: "1h" or "6h"; returns a job id, or push results to a webhook_url |
| GET /v1/jobs/{id} | poll status and read the result once completed |
Full guide: Async Jobs → Full parameters & example responses: Async Jobs API reference →
Batch
For bulk work on the full 24-hour window — the OpenAI-compatible Files + Batches shape, so any OpenAI Batch
SDK client works with a one-line base_url change.
| Endpoint | Description |
|---|---|
| POST /v1/files | upload a JSONL file of requests, purpose "batch" |
| GET /v1/files | list files |
| GET /v1/files/{id} | file metadata |
| GET /v1/files/{id}/content | download raw file content — input request lines, or the output/error JSONL once a batch completes |
| POST /v1/batches | create a batch from input_file_id, or an inline requests array for small jobs |
| GET /v1/batches | list batches |
| GET /v1/batches/{id} | status, progress counts, and output/error file ids once complete |
| POST /v1/batches/{id}/cancel | cancel a batch any time before it completes |
Full guide: Batch API → Full parameters & example responses: Batch API reference →