Multimodal
Image generation
One endpoint for both text-to-image and image-to-image —
POST /v1/images, matching
OpenRouter's own /api/v1/images contract
exactly (confirmed against their real OpenAPI spec — there's no separate "generate" vs. "edit" endpoint on
their side either). Image-to-image happens on this same endpoint via
input_references, not a separate
multipart upload. Billed per image, with a real
usage block — including
cost — inside the JSON response body,
not just a response header. That cost already includes our 2% platform fee — image and
video generation carry a higher fee than chat and other modes; see
Pricing & billing.
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/images",
headers={"Authorization": "Bearer llmr_sk_live_..."},
json={
"model": "openai/gpt-image-1/openai",
"prompt": "a red panda astronaut floating in space",
"aspect_ratio": "16:9",
"resolution": "2K",
},
).json()
print(resp["data"][0]["b64_json"][:50], resp["usage"]["cost"])
Image-to-image: pass one or more reference images via input_references (URL
or base64 data URL, same two shapes as image input)
instead of uploading a file:
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/images",
headers={"Authorization": "Bearer llmr_sk_live_..."},
json={
"model": "gemini-2.5-flash-image",
"prompt": "make this scene look like a watercolor painting",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
],
},
).json()
Every model listed under Which models below,
plus the chat-native image models
(gemini-2.5-flash-image,
gemini-3-pro-image-preview,
gemini-3.1-flash-image-preview,
gemini-3.1-flash-lite-image) work here —
one catalog, one endpoint, whichever mechanism each model actually uses under the hood. Other parameters:
size (an explicit
"WxH" string overrides
resolution/aspect_ratio when
given), quality,
output_format,
background, and
output_compression (all OpenAI-models-only —
ignored elsewhere, same as OpenRouter's own "ignored by providers lacking X control" behavior),
n for multiple images per call.
stream: true is not supported —
rejected with a 400, not silently ignored. Non-square aspect_ratio on
dall-e-2 is clamped to square (it only has
one size tier per axis); DeepInfra/SiliconFlow models don't honor
resolution/aspect_ratio yet.
Which models
Fifteen models across three providers, served direct (no cross-provider failover for this endpoint):
gpt-image-1OpenAIdefault for model:"auto", flat per-(size,quality) pricinggpt-image-1-miniOpenAIflat per-(size,quality) pricinggpt-image-1.5OpenAIflat per-(size,quality) pricingdall-e-3OpenAI—dall-e-2OpenAI—gpt-image-2OpenAIgeneration + edits, priced per tokengpt-image-2.5-flareOpenAIlatest OpenAI model; faster default, priced per tokengpt-image-2.5-sunburstOpenAIlatest OpenAI model; extra precision, slower, priced per tokenimagen-4.0Google—imagen-4.0-fastGoogle—imagen-4.0-ultraGoogle—qwen-image-2.0Qwen—qwen-image-2.0-proQwen—qwen-image-3.0Qwen—qwen-image-3.0-proQwenhigher output resolutions cost more
size and
quality are only forwarded for OpenAI
models — Imagen and Qwen don't take them the same way and ignore/reject them.
Other providers
A handful of open-weight and independent models, served direct through Fal or WaveSpeedAI:
fal/ideogram-4.0Falbilled per output megapixelfal/ideogram-4.0-qualityFalIdeogram's QUALITY rendering mode; billed per output megapixelfal/flux-2-devFalFLUX.2 [dev]; billed per output megapixelflux-3-imageBlack Forest Labs (direct)FLUX 3 Image; flat per resolution tier: 768sq $0.041, 1k $0.048, 2k $0.10, 4k $0.607. Text-to-image and editing (up to 10 input_references) at the same pricefal/muse-imageFalMeta's Muse Image, via Meta Model API on Falfal/krea-2-medium-turboFal—fal/cosmos3-super-text2imageFalNVIDIA's open-weights Cosmos 3 Superwavespeed/p-image-ideogram-highWaveSpeedAIPruna AI × Ideogram P-Image, "high" thinking level
size/resolution/aspect_ratio are
honored (and change the price) for fal/ideogram-4.0,
fal/ideogram-4.0-quality, and
fal/flux-2-dev only — the others bill (and generate) at a fixed size regardless of what's requested. flux-3-image is different again: resolution (512/1K/2K/4K → 768sq/1k/2k/4k) picks the price tier and aspect_ratio is honored.
Image output in chat completions
A second way to get an image back: pass modalities: ["image", "text"] to
POST /v1/chat/completions itself — the
same wire contract OpenRouter uses. The generated image comes back on the assistant message as
message.images[0].image_url.url (a base64
data URL), alongside whatever text the model also returned.
from openai import OpenAI
client = OpenAI(base_url="https://videorouter.sh/api/v1", api_key="llmr_sk_live_...")
resp = client.chat.completions.create(
model="gemini-2.5-flash-image",
messages=[{"role": "user", "content": "Generate a beautiful sunset over mountains"}],
modalities=["image", "text"],
)
message = resp.choices[0].message
for image in (message.images or []):
image_url = image["image_url"]["url"] # base64 data URL
print(f"Generated image: {image_url[:50]}...")
This works with an uploaded image too — combine an image_url
content part (see Image understanding) in the
same request with modalities: ["image", "text"] to
send an image in and get an edited/re-imagined image back through the same endpoint.
Four models today, all Google:
gemini-2.5-flash-imagereturns PNGgemini-3-pro-image-previewreturns JPEG; highest quality, highest costgemini-3.1-flash-image-previewreturns JPEGgemini-3.1-flash-lite-imagecheapest of the four, returns JPEG
model: "auto" is not supported for this —
pass the model explicitly. Billed per request off the response's own
usage.completion_tokens_details split: image
tokens (and, on gemini-3-pro-image-preview,
reasoning tokens) bill at a materially higher per-token rate than plain text tokens (this is the provider's own
pricing, not a platform markup), so cost varies with image size/complexity rather than being a flat per-call
price like /v1/images above. This path
still carries our own 2% platform fee on top of that provider cost, same as
/v1/images — image generation
reached through chat completions is billed the same rate as the dedicated endpoint.
A vague, templated prompt (e.g. "a tiny icon of a red circle") can trigger Gemini's own recitation safety
filter and come back with no image at all — a descriptive, original prompt avoids this.
Grok, DeepSeek, and Meta's muse-spark-1.1 don't
support this yet — each was tested directly, not assumed: Grok and OpenAI's gpt-5.1/gpt-5.2 reject
the modalities parameter outright,
DeepSeek silently ignores it and returns text only, and Muse Spark's own API rejects
"image" as an invalid modality
value (it supports image input, not output). We'll add a model here as soon as one of these — or
another provider — has a real, verified image-output path.
Fallback & failover
For a model with more than one registered host, a rejected submission automatically fails over to the
next-cheapest one still serving the same checkpoint — free, on by default, no parameter needed. Pin an
exact host instead with provider: {"only": [...], "allow_fallbacks": false}.
Full details, triggers, and request examples: Provider selection.
Falling back to a genuinely different model instead of another host of the same one:
Model fallbacks.
Not yet supported
Generating multiple unprompted variations of an existing image (no edit instruction, just "more like this")
has no equivalent here. Mask-based inpainting via input_references —
restricting an edit to one region of the image — isn't supported either, only whole-image edits.
stream: true is rejected outright (see
above). Chat-native image output is scoped to four Google models for now.