GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Arena Starts Scoring for Truth: Factuality Now Weighted Alongside Human Preference

Chatbot Arena is changing how it ranks models. Starting now, factuality — the accuracy of claims a model makes — is weighted alongside human preference in the Text and Search Arena leaderboards.

The change is opt-in via a toggle in the leaderboard UI. The methodology is explicit: factuality battle outcomes are computed independently, then blended with standard human preference ELO into a unified ranking. Neither signal replaces the other.

How It Works

Arena randomly samples live battles for a factuality audit. Atomic claims are extracted from each model’s response and filtered to those that are web-verifiable. A system of search agents assigns calibrated truth probabilities to each claim — calibration is post-processed against a high-quality annotated dataset to correct for raw model overconfidence. The model whose response has a higher average claim-truth probability wins the factuality battle; the margin scales the signal strength.

The resulting “factuality battles” feed into a combined leaderboard formula. Arena describes it as “a weighted combination of human preference and factuality, two complementary signals that tell only a partial story in isolation.”

Why It Matters

Human raters are good at detecting whether a response is helpful, fluent, and coherent. They are not good at fact-checking claims in real time. A model that produces confident-sounding hallucinations often wins human preference battles over a more cautious, accurate model — because the hallucination is only identified later, if at all.

The practical consequence: ELO has historically rewarded confident wrong answers. This is the gap Arena is targeting.

For developers choosing models for information-retrieval, research, or customer-facing workflows, the combined factuality-preference leaderboard is more decision-relevant than human ELO alone. A model ranked first on human preference but third on factuality is a different risk profile than one ranked consistently across both.

Current Coverage

The factuality signal is live in Text Arena and Search Arena. Arena plans to expand to additional leaderboard categories. The toggle is off by default — users who want the combined ranking must enable it manually in the leaderboard UI. Arena describes the non-default status as an initial launch posture, not a permanent constraint.

The Broader Signal

Arena is the only leaderboard with the data volume and battle infrastructure to make factuality scoring credible at scale. Independent evaluations from Adobe and others have documented model hallucination rates well above what human raters catch — GPT-5.5 has been measured at 86% hallucination rate in some evaluations. If factuality consistently diverges from human preference in the Arena data, it will pressure labs to optimize for claim accuracy rather than response confidence.

The toggle exists. When it moves to default-on, the leaderboard changes.