GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Wan3.0 Enters Text-to-Video Arena With Document Input and 30-Second Generation at $0.05/s

Alibaba’s Wan3.0 was added to Arena.ai’s Text-to-Video leaderboard on September 4, 2026. The listing opens a preference-based ranking for a model that ships with two capabilities that distinguish it from the existing field: native document ingestion and a 30-second single-pass generation ceiling.

Document Input Is the Feature That Matters Most

Every competitive video generation API accepts text prompts, reference images, and reference video clips. Wan3.0 adds a fourth category: documents and web pages parsed directly as generation input.

Supported file types include DOCX, XLSX, PPTX, PDF, TXT, Markdown, and Apple iWork formats (key, pages, numbers). Maximum one file or link per request, capped at 100MB and 50 pages. Wan3.0 reads the document content and uses it alongside the prompt to generate video — a product PPT becomes a brand film, a training deck becomes video courseware, a spreadsheet becomes animated charts.

Web links are also accepted, with the same mechanism: the model fetches and parses the page content, then generates from the structured material.

The API treats this through a unified wan3.0-video model name. No model switching required — the system automatically routes based on the type field in the media array (file or link).

30 Seconds: Three Times Wan2.7’s Ceiling

Wan2.7 capped single-generation output at 10 seconds. Wan3.0 raises that to 30 seconds, with a smart duration mode (set duration: -1) that instructs the model to recommend an appropriate length based on prompt complexity and input material. Short product shots stay short; complex narratives get room.

Video extension is also native: generate a first segment, then extend the storyline from the last frame without full regeneration.

Multimodal Reference Surface: Up to 20 Assets

Wan3.0 accepts up to 20 reference assets per request across images (max 10), video clips (max 5, total ≤15 seconds), and audio clips (max 5, total ≤15 seconds), combined in any mix. In the prompt, use “Image 1”, “Video 1”, “Audio 1” notation to refer to specific assets in order.

Document or web link inputs are mutually exclusive with first_frame/last_frame frame-control inputs — these modes cannot be combined in a single request.

Pricing and Models

ResolutionPrice per Second
480P$0.05
720P$0.10
1080P$0.20

Two model variants: wan3.0-video (standard) and wan3.0-video-prime (high-speed, same capabilities). Both available on Alibaba Cloud Model Studio. Default output resolution is 1080P. Aspect ratio defaults to adaptive — the model infers appropriate framing from input content.

Native audio is on by default: dialogue, background music, and sound effects are generated in the same pass as video. Audio generation does not affect pricing.

Arena Context

The Text-to-Video Arena is separate from Arena.ai’s Video Edit Arena, where Wan3.0 entered on August 27. Preference-based ELO scores on the Text-to-Video leaderboard take two to three weeks to stabilise after initial placement. Arena.ai’s existing Video Edit rankings for Wan3.0 are not directly comparable.

ByteDance’s Seedance 2.5 has led Arena.ai’s video rankings for multiple months across categories. Wan3.0’s ranked position in Text-to-Video will be visible by late September.

What Document-to-Video Means for the API Market

AI video APIs have competed on duration, resolution, and reference control since mid-2025. Document input shifts the competitive axis: instead of treating video generation as a media production tool, it repositions the API as a document rendering layer — a way to express information that exists as structured text into a format that plays back. For enterprise use cases — brand films from brand guidelines, training videos from training decks — this is a materially different value proposition than what any current competitor offers.