GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Qwen3.8 Max 0902: Alibaba Post-Trains 2.4T MoE for Agentic Coding at $2 Input

Alibaba shipped Qwen3.8 Max 0902 on September 3, the latest date-stamped snapshot of its 2.4-trillion-parameter mixture-of-experts flagship. The notable change from the preview release in July is not architecture — that remains 2.4T MoE with a 1M-token context — but post-training focus.

What the Snapshot Changes

The 0902 snapshot is described as post-trained specifically for:

  • Multi-step software projects and agentic coding pipelines
  • Multi-tool orchestration across extended sessions
  • Long-horizon task execution
  • Chart reasoning and document parsing
  • Multimodal understanding over long documents and extended video sequences

The previous July preview was positioned as a capability benchmark (AA Intelligence Index placement, SWE-bench comparison). The September snapshot signals Qwen is now treating the model as a production-grade coding agent rather than a general intelligence demonstration.

Pricing and Positioning

MetricValue
Input$2.00 / 1M tokens
Output$6.00 / 1M tokens
Context1M tokens
ProviderAlibaba Cloud International
ReasoningOn by default
Input modalitiesText, image, video

At $2 input, Qwen3.8 Max 0902 sits at roughly 12.5x the cost of Qwen3.8 Flash ($0.16 input) and significantly below what frontier Anthropic and OpenAI models charge for comparable context windows. The blended cost ($3.50 at a 25/75 input/output mix) makes it accessible for agentic workloads where output tokens dominate.

Video Input at Scale

The model accepts text, image, and video input. The post-training notes emphasize “multimodal understanding over long documents and extended video” specifically — not single-frame analysis but video understanding across an extended context window. At 1M tokens, that translates to roughly 10-12 hours of video at standard token densities, though practical limits depend on resolution and sampling rate.

This is the segment of the market where Alibaba appears to be competing directly with Google’s Gemini 3.x Flash family on value-per-capability rather than raw benchmark ranking.

Tool Calling and Structured Outputs

The snapshot explicitly supports tool calling, structured outputs, and configurable reasoning effort. Configurable reasoning is relevant for cost control in production: customers can dial down thinking token usage for simpler tasks and reserve full reasoning for complex ones.

The single-provider setup (Alibaba Cloud International only, no multi-provider routing on OpenRouter) means uptime risk is concentrated. Current uptime shows 100% but that reflects a short post-launch window. For production agentic pipelines, the single-provider constraint is a consideration.