How to Call the MiniMax H3 API (Python + curl, Provider Pinning Explained)
MiniMax H3 is served through the same /v1/videos endpoint as every other video model VideoRouter routes — async, job-based, billed per requested second at creation time. What makes H3 worth its own integration guide isn't the request shape (it's identical to every other model's) — it's that H3 has one of the widest provider rosters we track (eleven hosts), which makes explicit provider pinning genuinely useful here in a way it isn't for a model with only two or three hosts.
Text-to-video
import time
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "minimax/h3",
"prompt": "a paper airplane gliding over a city",
"duration_secs": 8,
},
)
job = resp.json()
while job["status"] not in ("completed", "failed"):
time.sleep(5)
job = requests.get(
f"https://videorouter.sh/api/v1/videos/{job['id']}",
headers={"Authorization": "Bearer llmr_sk_live_..."},
).json()
if job["status"] == "completed":
video_url = job["data"][0]["url"]
else:
raise RuntimeError(job["error"])
Equivalent curl:
job=$(curl -s https://videorouter.sh/api/v1/videos \
-H "Authorization: Bearer llmr_sk_..." \
-H "Content-Type: application/json" \
-d '{"model": "minimax/h3", "prompt": "a paper airplane gliding over a city", "duration_secs": 8}')
id=$(echo "$job" | jq -r .id)
status=$(echo "$job" | jq -r .status)
while [ "$status" != "completed" ] && [ "$status" != "failed" ]; do
sleep 5
job=$(curl -s "https://videorouter.sh/api/v1/videos/$id" -H "Authorization: Bearer llmr_sk_...")
status=$(echo "$job" | jq -r .status)
done
echo "$job" | jq -r 'if .status == "completed" then .data[0].url else .error end'
Image-to-video
H3 supports animating a starting frame on most of its hosts (Fal, Atlas Cloud, Replicate, WaveSpeedAI, and MachGen all support start_image_url for H3) — pass it and the prompt guides motion from that frame instead of generating from scratch:
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "minimax/h3",
"prompt": "the subject turns and smiles, gentle camera push-in",
"start_image_url": "https://example.com/photo.jpg",
"duration_secs": 5,
"resolution": "768p",
},
).json()
Why provider pinning matters more for H3 than most models
Most video models in VideoRouter's catalog have two to five hosts. H3 has eleven, spanning a 2.5x price range at the 768p tier alone (see MiniMax H3 API Pricing: Every Provider Compared for the full table). With that much spread, letting automatic routing pick a deployment for you is a reasonable default, but pinning explicitly to the cheapest verified host is a bigger, more predictable saving here than for a model where every host is within 20% of every other.
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "minimax/h3/machgen",
"prompt": "a paper airplane gliding over a city",
"duration_secs": 8,
},
).json()
The tradeoff: pinning to one host means no automatic failover to a second one if MachGen specifically has an outage. A middle-ground pattern that gets most of the price benefit while keeping some failover: pin to an explicit only allow-list of your two or three cheapest verified hosts rather than a single one, so routing still fails over, but only among hosts you've already confirmed are competitively priced — see /docs/provider-selection for the provider.only field's exact shape.
Handling errors correctly
VideoRouter's error envelope is consistent across every endpoint, video included — {"error": {"message", "type", "code"}}, the same shape as the OpenAI API. The codes that actually come up calling H3:
| Status | type / code | What it means |
|---|---|---|
| 400 | invalid_request_error |
Missing prompt, or a model id not in the catalog |
| 401 | invalid_api_key |
Key missing, malformed, revoked, or expired |
| 402 | spend_cap_exceeded / insufficient_credits |
Monthly cap hit, or prepaid balance ≤ $0 |
| 403 | model_not_allowed |
H3 isn't in this key's model_allowlist |
| 429 | rpm_limit / tpm_limit |
Rate limit exceeded — Retry-After header tells you how long to wait |
| 500 / 502 / 503 / 504 | upstream_error |
Every candidate host failed — not billed |
Given H3 has eleven hosts to fail over across, a genuine upstream_error on this specific model is a stronger signal than it would be on a two-host model — either something is affecting several independent providers at once, or (more likely) you've pinned to a single host via the provider suffix and that one host specifically is down, which is exactly the failover you give up when pinning:
import time
import requests
API_BASE = "https://videorouter.sh/api/v1"
API_KEY = "llmr_sk_live_..."
def generate_h3_video(prompt, provider=None, **kwargs):
model = "minimax/h3" + (f"/{provider}" if provider else "")
headers = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
payload = {"model": model, "prompt": prompt, **kwargs}
for attempt in range(3):
resp = requests.post(f"{API_BASE}/videos", headers=headers, json=payload)
if resp.status_code == 429:
time.sleep(int(resp.headers.get("Retry-After", 5)))
continue
resp.raise_for_status()
job = resp.json()
break
else:
raise RuntimeError("exceeded retry budget on job creation")
deadline = time.time() + 600
while job["status"] not in ("completed", "failed"):
if time.time() > deadline:
raise TimeoutError(f"job {job['id']} did not complete within 10 minutes")
time.sleep(5)
job = requests.get(f"{API_BASE}/videos/{job['id']}", headers=headers).json()
if job["status"] == "failed":
raise RuntimeError(f"generation failed: {job['error']}")
return job["data"][0]["url"]
If you're pinning to MachGen specifically for price reasons and want a fallback if it's unavailable, catch the upstream_error and retry once with no provider suffix (letting default routing pick any healthy host) rather than surfacing the failure straight to a user — you'll pay a higher per-second rate on the fallback attempt, but that's a better tradeoff than a hard failure for most products.
Running a batch of prompts
A common pattern for H3 specifically — given how cheap its lower tiers are relative to most video models — is generating several prompt variations for the same concept and picking the best result, rather than generating one clip per concept and hoping the first attempt is usable. Since job creation returns immediately (a 202-style response with an id, not a blocking call), you can submit a whole batch before polling any of them:
import time
import requests
API_BASE = "https://videorouter.sh/api/v1"
API_KEY = "llmr_sk_live_..."
headers = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
def submit(prompt, **kwargs):
resp = requests.post(f"{API_BASE}/videos", headers=headers,
json={"model": "minimax/h3", "prompt": prompt, **kwargs})
resp.raise_for_status()
return resp.json()["id"]
def poll_all(job_ids, interval=5, timeout=600):
results = {}
deadline = time.time() + timeout
remaining = set(job_ids)
while remaining and time.time() < deadline:
time.sleep(interval)
for job_id in list(remaining):
job = requests.get(f"{API_BASE}/videos/{job_id}", headers=headers).json()
if job["status"] in ("completed", "failed"):
results[job_id] = job
remaining.remove(job_id)
return results
prompts = [
"a paper airplane gliding over a city, golden hour",
"a paper airplane gliding over a city, blue hour, neon reflections",
"a paper airplane gliding over a city, storm clouds gathering",
]
job_ids = [submit(p, duration_secs=5, resolution="480p") for p in prompts]
results = poll_all(job_ids)
Submitting all three jobs up front rather than one-at-a-time-with-a-poll-loop-between-each cuts your total wall-clock time roughly to the slowest single job's completion time, instead of the sum of all three — a meaningful difference once you're generating more than a couple of variations per concept. At H3's 480p rate ($0.035-$0.05/sec depending on host — see MiniMax H3 API Pricing), generating three 5-second variants to pick the best one costs under a dollar total, which is cheap enough to make this a reasonable default rather than a special-case optimization.
Common mistakes
Assuming every host supports every tier. H3's 480p and 1440p/2K tiers aren't available at every one of the eleven hosts — check the specific tier against MiniMax H3 API Pricing before assuming a pinned host will accept a given resolution value; an unsupported tier is silently ignored rather than rejected, so you can end up billed for a different resolution than you requested without an error telling you so.
Confusing WaveSpeedAI's two H3 listings. WaveSpeedAI carries both a native H3 listing and a resale listing at a meaningfully higher price — pinning without checking which one you're targeting can land you on the pricier of the two. See the pricing article linked above for both rates side by side.
Polling too aggressively. Status polling is free, but hammering the status endpoint every second instead of every 5-10 seconds adds no value — job completion times for an 8-second H3 clip are typically in the tens of seconds to low minutes, not sub-second, so a tighter poll interval only adds request overhead without getting you the result any faster.
See /docs/video-generation for the complete request reference and How to Call the Seedance API for the equivalent walkthrough on ByteDance's model family.
Frequently asked questions
Why does H3 have so many more hosts than most video models? MiniMax has licensed or opened H3 up to a wider set of third-party hosts than most model creators do for their flagship checkpoints — that's a business decision on MiniMax's side, not something VideoRouter controls, but it's the direct reason H3's price comparison table (see MiniMax H3 API Pricing) has eleven rows where most models in this series have three to seven.
Is the minimax/h3 model id case-sensitive? Treat it as case-sensitive and use it exactly as shown — model ids in general aren't guaranteed to normalize case, and an unrecognized model id returns a 400 rather than silently falling back to a close match.
Can I request multiple H3 variants (H3, H3 Max, H3 Max Turbo) through the same code path? Yes — swap the model field between minimax/h3, minimax/h3-max, and minimax/h3-max-turbo; the request shape is identical across all three, only the price and quality/speed tradeoff changes.
What's the fastest way to verify which host actually served a given job? The completed job object includes provider/host attribution in its response — check the specific field name against /docs/api-reference-video, since relying on which host you think you pinned to (versus which one actually served the request, if you didn't pin) can otherwise lead to debugging the wrong host when a quality or latency issue comes up.