Same AI Video Model, Different API Prices: Why?
Here's a question worth asking before you wire a video model into your product: is the price you're looking at actually the price, or just one price among several for the same underlying model? We pulled real per-second pricing for two popular video models — Alibaba's Wan 3.0 Prime and ByteDance's Seedance 2.5 — across every host we track that serves them, to see how big the gap actually is between the cheapest and most expensive way to call the identical checkpoint. The results are worth walking through in detail, because they're not what most people assume.
We're treating this less like a general explainer and more like an audit: pull the numbers, lay them side by side, and see what they actually say, rather than starting from a theory about GPU costs or provider margin and working backward to confirm it. Both models below are checkpoints available from more than one host, which is what makes the comparison possible in the first place — plenty of newer or more obscure models are still effectively single-sourced, and for those there's no arbitrage to find yet. But for models with real multi-host availability, the numbers tell a fairly blunt story.
Case study 1: Wan 3.0 Prime at 480p
Wan 3.0 Prime is Alibaba's video generation model, and it's available through Alibaba's own official API (DashScope) as well as through a handful of third-party resellers. The intuitive assumption is that the model creator's own official API should be the cheapest option — after all, they're not paying anyone else's margin. That assumption turns out to be wrong.
Here's the real per-second pricing at 480p, across every host in our registry that serves this model, as of September 2026:
| Provider | Price per second (480p) | Notes |
|---|---|---|
| Pika (resale, via fal.ai) | $0.051 | Cheapest tested |
| Atlas Cloud | $0.0612 | |
| DashScope (Alibaba, official) | $0.068 | Model creator's own API |
| Fal.ai | $0.068 | Matches official rate |
| WaveSpeed | $0.075 | |
| OpenRouter | $0.17 | Real invoiced rate, live-tested |

A few things jump out here. First, Alibaba's own official DashScope pricing — $0.068/sec — is not the cheapest option available. A Pika resale listing, served through fal.ai, undercuts the official price by about 25%, landing at $0.051/sec. That's not a typo or an error in either direction; it's simply how wholesale reselling works in this market. A reseller that negotiates favorable bulk terms with the underlying infrastructure, or runs a more efficient serving setup, can pass most of that saving through to customers and still turn a profit, ending up cheaper than the model creator's own retail-facing price. "Official" and "cheapest" are two different questions, and conflating them is an easy way to overpay.
Second, and more striking: we ran a real, live-invoiced test job against OpenRouter's Wan 3.0 Prime endpoint, and the actual billed rate came out to $0.17 per second — a 2-second test job was invoiced $0.34 total. That's roughly 2.5x OpenRouter's own advertised catalog price for this model, and more than 3x what the cheapest tested host (the Pika resale listing) charges for the identical checkpoint. We're flagging this as a genuinely useful, differentiated data point from our own live testing, not as an accusation — pricing pages can drift out of sync with what actually gets billed, and this is exactly the kind of discrepancy that only shows up when someone runs a real job and checks the invoice line-by-line rather than trusting the rate card.
Put together, that means the realistic spread for Wan 3.0 Prime at 480p, from the cheapest host we tested to what actually gets billed on the most expensive one, is north of 3x for output from the exact same model.
It's also worth being precise about what "Pika" means as a line item in that table, since it's easy to misread. Pika Labs is the creator of its own consumer-facing video models, but it isn't a first-party API vendor in the way Alibaba is for Wan — the Wan 3.0 Prime access attributed to Pika here comes through a developer-facing API surface that resells access to multiple third-party models, Wan included, alongside Pika's own. That's a useful thing to understand about the market generally: several of the names that show up in pricing comparisons like this one are themselves small aggregators reselling other companies' models, not single-model vendors, and reseller economics — buying wholesale, marking up less than the notional "retail" rate — is a large part of why a resold listing can beat the model creator's own official price.
To make the dollar stakes concrete: at 480p, the gap between Pika's $0.051/sec and OpenRouter's real invoiced $0.17/sec is $0.119 for every second of video generated. A product generating 5,000 seconds of Wan 3.0 Prime output in a month — not an unusual volume for a video feature with even modest usage — would pay roughly $255 on the cheap end and $850 on the expensive end for identical output. That's not a rounding error in a startup's infrastructure budget; it's the kind of gap that shows up as a line item someone eventually asks about in a spend review.
Case study 2: Seedance 2.5 at 720p
Seedance 2.5 is ByteDance's video model, and the pricing landscape here is a cleaner illustration of the same underlying pattern, without the added wrinkle of an advertised-versus-invoiced gap. Every number below is a currently quoted rate, per second of output at 720p:
| Provider | Price per second (720p) |
|---|---|
| MachGen | $0.19 |
| OpenRouter | $0.231 |
| Atlas Cloud | $0.300 (discounted from $0.376) |
| WaveSpeed | $0.36 |
| Fal.ai | $0.473 |
MachGen sits at the bottom, $0.19/sec. Fal.ai sits at the top, $0.473/sec. That's a 2.5x spread for identical output from the same checkpoint — no ambiguity about official versus resold, no advertised-versus-invoiced discrepancy, just five providers charging materially different amounts for the same model.
Interestingly, the ranking here doesn't map cleanly onto any single obvious explanation like "the aggregator is always cheapest" or "the specialist media platform is always cheapest." MachGen, a resale aggregator, comes in lowest. Fal.ai, a platform that specializes specifically in serving diffusion and generative media models, comes in highest. Atlas Cloud, a broad pay-as-you-go aggregator similar in shape to MachGen, lands in the middle rather than at the bottom. That's consistent with the idea that the price you see reflects each individual provider's own infrastructure efficiency and margin decisions for this specific model, not some fixed rule about which category of platform is cheapest in general.
Note, too, that Atlas Cloud's $0.300 figure is itself a discounted rate, down from an official $0.376 list price — a 20% promotional cut. That's a reminder that even within a single provider's own pricing, the number on the page today isn't necessarily the number that was there last month or will be there next month. A comparison snapshot is exactly that: a snapshot, not a permanent ranking.
Prices — and availability — are not guaranteed to hold
Both case studies above are about price differences between hosts serving the same model right now. It's worth widening the lens slightly, because the same underlying lesson — don't assume today's setup is permanent — applies to model availability too, and there's a concrete, current example of why that matters. OpenAI notified developers in March 2026 that its Videos API and the entire Sora 2 model family (sora-2, sora-2-pro, and dated snapshots) would be removed from the API on September 24, 2026, with no replacement model listed in its own deprecation table. Anyone who built directly on OpenAI's native Sora 2 API has to find a new host for that workload within days of this article being written. Third-party resellers — Pika, Replicate, DeepInfra, and MachGen among them — still serve Sora 2 Pro as of now, since they call OpenAI's backend under their own business agreements, but that continuity is worth treating as "true as of today," not as a guarantee.
The connective thread between that and the Wan 3.0 Prime and Seedance 2.5 numbers above is the same: whatever provider and price you've settled on for a given model is a snapshot of current market conditions, not a fixed fact. Models get deprecated. Hosts change pricing. Discounts expire or get introduced. A reseller renegotiates its wholesale rate and passes the change through. None of that shows up unless you're actually watching for it, and "I checked this once when I integrated it" isn't the same as "I know this is still true."
What both cases actually show
Line the two case studies up side by side and the pattern is consistent even though the specific numbers and ordering differ:
| Cheapest host | Most expensive tested | Spread | |
|---|---|---|---|
| Wan 3.0 Prime (480p) | Pika resale — $0.051/sec | OpenRouter — $0.17/sec (real invoiced) | >3x |
| Seedance 2.5 (720p) | MachGen — $0.19/sec | Fal.ai — $0.473/sec | 2.5x |
Two takeaways worth sitting with. The first is that "same model" and "same price" are not the same statement, and treating them as interchangeable is an easy way to leave real money on the table. Neither of these is an edge case or a cherry-picked outlier — they're two popular, actively used video models, and the spread on both is large enough that it would show up clearly on an invoice for any team running meaningful generation volume. A team generating a few thousand seconds of Seedance 2.5 output a month is looking at a difference of hundreds to low thousands of dollars depending purely on which host they picked, for literally identical output.
The second takeaway is less about the specific dollar amounts and more about process: the "official" or "first listed" price for a model is not a reliable signal of what the best available price actually is. It can be beaten by resellers with better wholesale terms, and — as the OpenRouter case shows — even an advertised catalog price can diverge meaningfully from what a live job actually bills. That means checking isn't something you do once when you first pick a model and then stop thinking about. Prices shift as hosts add discounts, adjust rates, or start serving a model for the first time, and a provider's ranking today isn't guaranteed to hold next quarter. For anyone building on these models at real volume, treating price comparison as a recurring check rather than a one-time decision is the difference between paying what you have to and paying considerably more than you have to for the exact same output.
Doing that check by hand — across five or six providers, for every model you use, on a recurring basis — is tedious enough that most teams simply don't do it consistently, which is exactly how these spreads persist unnoticed. VideoRouter (videorouter.sh) is built around exactly this problem: a live per-model price comparison table on every model's page, pulled from real provider pricing rather than static assumptions, so the question "which host is actually cheapest for this model right now" has a direct answer instead of requiring a fresh round of manual research every time you check.