GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
← Back to feed

ERNIE 5.1 Launches at 6% of Industry Pretraining Cost — Arena Search #4 Globally, Beats DeepSeek-V4-Pro on Agents

Baidu released ERNIE 5.1 on May 9, moving from the preview stage to general availability on its Qianfan enterprise platform and the public Ernie Bot website. The story is efficiency: the model runs on 6% of the pretraining compute required by comparable models at similar scale — dropping training energy from roughly 240 million kWh (the GPT-4 range) to approximately 6.3 million kWh per equivalent run.

The compression comes from Baidu’s multi-dimensional elastic pretraining framework, which extracts optimal sub-structures from the ERNIE 5.0 foundation across three axes — depth, width, and sparsity — in a single training pass. Total parameters compress to one-third of 5.0; active parameters compress to one-half. Single-response latency is 35% lower than the prior generation.

Benchmark Performance

The headline external validation is the LMArena Search leaderboard: ERNIE 5.1 scored 1,223 on May 9 to claim fourth place globally and first among all Chinese models. That search ranking is where ERNIE 5.0’s successor preview first appeared at a different position; the GA release scores higher and in a harder category.

On agent benchmarks:

  • τ³-bench and SpreadsheetBench-Verified: ERNIE 5.1 surpasses DeepSeek-V4-Pro. Agent capabilities described as “approaching leading closed-source models.”
  • AIME26 (with tool use): 99.6 — second only to Gemini 3.1 Pro among all models evaluated.
  • GPQA and MMLU-Pro: Performance described as approaching frontier closed-source models.
  • Creative writing: Internal evaluations place it close to Gemini 3.1 Pro.

The preview version separately posted 1,476 on the LMArena text leaderboard on April 30, landing first in China and inside the global top 15, ahead of GPT-5.5 and DeepSeek-V4-Pro on that metric.

Why the Cost Number Matters

Training a frontier model at 6% of standard compute cost changes the economics of iteration. Baidu’s approach allows multiple sub-models to be trained in a single session, which means the cost to run parallel experiments — and to keep the model current — drops proportionally. The difference between 240 million kWh and 6.3 million kWh is also a data-center power commitment: the kind of gap that determines whether a lab can iterate monthly or quarterly.

Baidu’s stock rose 6% on May 8 ahead of the official GA release. The Baidu AI Developer Conference runs May 13-14 in Beijing, where the company is expected to publish fuller technical details and commercial licensing terms.

What It Doesn’t Change

ERNIE 5.1 is not available outside of Baidu’s own platforms. Qianfan is the enterprise API path; Ernie Bot is the consumer surface. No open weights. No third-party API routing. For teams that want efficiency-class pricing without Baidu platform dependency, DeepSeek V4 Flash at $0.14/M output remains the open alternative.