AI Video API Pricing Falls Below $0.10/s as Wan3.0 and Grok Imagine Video 1.5 Both Enter Arena This Week
Two AI video generation models reached Arena.ai’s Text-to-Video leaderboard in the same week: Alibaba’s Wan3.0 on September 4 and xAI’s Grok Imagine Video 1.5 on September 5. Together they establish the current pricing floor for commercially deployed video APIs — and reveal a clear split in how the major labs are positioning for different use cases.
The Pricing Comparison
| Model | 480P | 720P | 1080P | Max Duration |
|---|---|---|---|---|
| Wan3.0 | $0.05/s | $0.10/s | $0.20/s | 30 seconds |
| Grok Imagine Video 1.5 | $0.08/s | $0.14/s | $0.25/s | 15 seconds |
At 480P, Wan3.0 is 37.5% cheaper. At 720P, the gap is 29%. At 1080P, Grok costs 25% more per second but produces output at half the maximum duration Wan3.0 supports.
Neither is the cheapest deployed video API. India’s Varya reached $0.005/s earlier this year on open-weight infrastructure — a full order of magnitude below both. But Wan3.0 and Grok Imagine Video 1.5 are now the two highest-visibility commercial APIs from top-tier labs with formal Arena ranking, which anchors them as the reference points builders compare against.
Capability Differences
The pricing gap reflects genuinely different product designs.
Wan3.0 targets volume and modality breadth. Its 30-second ceiling triples what Wan2.7 shipped, and its document-to-video input mode — accepting PDFs, spreadsheets, presentations, and web links directly as structured generation input — is the first of its kind among competitive APIs. Up to 20 reference assets per request (10 images, 5 video clips, 5 audio clips). Native audio in the same pass. Lower per-second cost reflects a focus on high-volume, document-heavy enterprise workflows.
Grok Imagine Video 1.5 targets quality and creative control. The July 31 update added native text-to-video, 1080p, and up to seven simultaneous visual reference anchors per generation. Voice references for face and audio consistency. Single-pass synchronized audio — dialogue, ambience, and sound effects generated with the video, described by xAI as speech that “lands on the action.” Arena.ai’s Image-to-Video leaderboard recorded 1460 ELO points for the model, placing it ahead of FLUX 3 Video.
Grok also uses a different architecture for text inputs: text-to-video routes through text-to-image → image-to-video under the hood on some modes (returning only the final video), while reference modes use direct multi-modal conditioning. Wan3.0 routes via a single unified model regardless of input type.
Duration Is the Structural Constraint
The 30-second vs 15-second ceiling is not a minor spec difference. At 30fps, 30 seconds is 900 frames. A 15-second clip is 450 frames. For narrative content, enterprise training videos, or any use case requiring sustained visual storytelling, halving the duration means halving what can be communicated in a single API call.
Wan3.0’s extension capability partially addresses this: generate a first segment, then extend from the last frame without full regeneration. Grok Imagine Video 1.5 also supports video extension. In practice both models can produce longer sequences through chained calls, but Wan3.0 reaches 30 seconds natively; Grok needs at least two calls to match it.
Arena Context
Both models now appear on Arena.ai’s Text-to-Video leaderboard. ELO scores take two to three weeks to stabilise after placement. Wan3.0’s Text-to-Video ranking will be visible by late September. Grok Imagine Video 1.5’s 1460 ELO reading — from the Image-to-Video arena — gives an initial reference point, but cross-arena comparisons are indirect.
ByteDance’s Seedance 2.5 and Google’s Veo 3.x have led Arena.ai’s video categories for most of 2026. Grok at 1460 points sits behind Gemini Omni Flash in the recorded Image-to-Video snapshot. Whether either new entrant competes with the current leaderboard leaders will depend on human evaluator preference across actual generation tasks, not spec sheets.