GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
← Back to feed

ERNIE 5.1 Preview Hits Arena #13 Globally With 1/3 the Parameters of ERNIE 5.0 — and Leads Worldwide in Legal

Baidu’s ERNIE 5.1 Preview landed on LMArena’s April 30 rankings at 1476 ELO — 13th globally on the Text leaderboard and first among all Chinese models. That places it ahead of models from DeepSeek, Qwen, Kimi, MiMo, and GLM on the platform that determines rankings through blind head-to-head comparisons.

The category results are the more striking data point. On Legal and Government, ERNIE 5.1 Preview ranks first globally, above every model from OpenAI, Anthropic, and Google. On Mathematics it sits ninth. On Business, Management, and Financial Operations it ranks fourth. On Software and IT Services it ranks seventh. Its weakest relative performance is in creative and general text, where the Western frontier models have an advantage.

Architecture: Compression as the Strategy

ERNIE 5.1 Preview is not a standalone model. It inherits the pre-training foundation of ERNIE 5.0 — a 2.4 trillion parameter native multimodal model Baidu released in late 2025 — and then applies aggressive compression: total parameters are cut to approximately one-third of the 5.0 base, and active parameters are cut to approximately one-half.

The pre-training cost is reported at around 6% of what comparable models at the same scale typically require. That figure is plausible given that ERNIE 5.1 is a distilled and post-trained derivative of a much larger base, not a full training run from scratch. The cost advantage compounds: Baidu gets competitive quality at a fraction of the compute bill, while ERNIE 5.0’s depth provides a richer foundation than a lean model trained from zero.

Post-training methodology uses two disclosed techniques: decoupled fully-asynchronous reinforcement learning, which Baidu says stabilizes training at scale by decoupling the policy update from rollout generation; and what the company calls “scaled agentic post-training,” the details of which have not been published.

The Legal and Government category ranking is the result that warrants attention. Legal reasoning requires precise instruction-following, sensitivity to jurisdiction-specific language, and tolerance for nuanced ambiguity — all properties that tend to favor models with strong general language modeling rather than raw benchmark optimization. That ERNIE 5.1 leads this category globally is either a genuine capability signal or an artifact of heavy Chinese legal corpus exposure in training. The benchmark is human preference, not automated grading, which makes gaming harder.

A full public release is expected at Create 2026, Baidu’s AI Developer Conference scheduled for May 2026. The preview is currently available for enterprise users and developers through Baidu’s Qianfan model platform.

Context

ERNIE 5.1 is Baidu’s answer to a specific competitive pressure: Chinese models have been competitive on general intelligence benchmarks for several months, but the LMArena rankings have been dominated by Western models in preference-based head-to-head evaluation. Reaching #1 in a specific professional category changes the commercial narrative. Baidu’s stated customers for ERNIE are primarily Chinese enterprises and government entities, where the Legal and Government ranking is the one that matters commercially.

The efficiency claims, if they hold under independent scrutiny, are also significant. A model that extracts competitive performance at 6% of pre-training cost represents a different approach to the scaling question — not more compute, but better reuse of existing compute.