Discord
此页面暂无中文版本 — 以下显示英文内容。 查看英文版

← Chat completions

PDF, audio & video input

Beyond image input (see Image understanding), chat completions also accepts a type:"file" content part (PDF/document), OpenAI's own type:"input_audio", and a video content part. None of these are gated by capability the way images are (see capability-aware auto-routing) — support depends entirely on whether the model you route to actually converts the shape into something its real upstream API accepts. Don't build against it without checking the table below first.

Provider PDF Audio Video
Google Gemini / VertexReal supportReal supportReal support
Anthropic (direct or Vertex)Real supportNot supportedNot supported
Anthropic via BedrockReal supportNot supportedReal support
MistralPartial — file_id onlyNot supportedNot supported
OpenAINative shape, forwarded as-isNative shape, forwarded as-isNot supported
DeepSeek (native deepseek-* slugs)Blocked (400)Blocked (400)Blocked (400)
xAI, and every OpenAI-compatible passthrough model (resale providers, Qwen, GLM, MiniMax, etc.)Untested — raw passthroughUntested — raw passthroughUntested — raw passthrough

"Real support" means our dispatch layer actually converts the content part into that provider's own native shape (e.g. Anthropic's document block, confirmed by reading the conversion code) — not that we tested every provider live. "Untested — raw passthrough" means the OpenAI-shaped block is forwarded exactly as you sent it; whether it works depends entirely on that specific backend accepting OpenAI's own shape. DeepSeek is the one confirmed-bad case: see below.

PDF / document input

Pass a type:"file" content part with a base64 file_data data URL — same shape OpenAI itself uses:

resp = client.chat.completions.create(
    model="anthropic/claude-opus-4.8",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Summarize this contract"},
            {"type": "file", "file": {"filename": "contract.pdf", "file_data": f"data:application/pdf;base64,{b64}"}},
        ],
    }],
)

On Gemini/Vertex, unsupported MIME types are rejected with a clean error rather than silently accepted — the conversion enforces an allow-list. The OpenRouter-style top-level plugins array (used there to pick a PDF parser engine, e.g. mistral-ocr) is silently dropped before the request reaches a provider on this platform — sending it doesn't error, it just does nothing.

Audio input

Only Gemini/Vertex has real conversion for OpenAI's type:"input_audio" content part today. Everywhere else it's either dropped unhandled or, on native DeepSeek, would have been silently stripped (now blocked instead — see below). If you need a transcript rather than raw audio understanding, speech-to-text works across a wider model set and is the safer default.

Video input (understanding, not generation)

This is a video file as chat input — for generating a video, see Video generation instead. Real support exists only on Gemini/Vertex (which also accepts optional video_metadata: fps, start_offset, end_offset) and Anthropic reached specifically via Bedrock — not direct Anthropic, not Anthropic via Vertex, which both lack it.

DeepSeek is blocked, not silently degraded

Native DeepSeek's own request builder collapses any non-text content part — image, file, or audio — down to plain text before the request is even sent, with no error and no indication anything was dropped. Rather than let that happen invisibly, sending an image/file/audio content part to a native deepseek-* model (explicit pin or via model:"auto") is rejected outright with 400 unsupported_content_type. This only applies to native DeepSeek deployments — DeepSeek models resold via DeepInfra/DigitalOcean route through a different, unaffected code path.

What to do on an unsupported combination

PDF: convert pages to images and send them through image understanding, or extract the text yourself and inline it as a plain text content part. Audio: transcribe it first via speech-to-text and send the transcript as text. Video: extract representative frames as images and send those instead.