Arena Starts Scoring for Truth: Factuality Now Weighted Alongside Human Preference
Chatbot Arena is changing how it ranks models. Starting now, factuality — the accuracy of claims a model makes — is weighted alongside human preference in the Text and Search Arena leaderboards.
The change is opt-in via a toggle in the leaderboard UI. The methodology is explicit: factuality battle outcomes are computed independently, then blended with standard human preference ELO into a unified ranking. Neither signal replaces the other.
How It Works
Arena randomly samples live battles for a factuality audit. Atomic claims are extracted from each model’s response and filtered to those that are web-verifiable. A system of search agents assigns calibrated truth probabilities to each claim — calibration is post-processed against a high-quality annotated dataset to correct for raw model overconfidence. The model whose response has a higher average claim-truth probability wins the factuality battle; the margin scales the signal strength.
The resulting “factuality battles” feed into a combined leaderboard formula. Arena describes it as “a weighted combination of human preference and factuality, two complementary signals that tell only a partial story in isolation.”
Why It Matters
Human raters are good at detecting whether a response is helpful, fluent, and coherent. They are not good at fact-checking claims in real time. A model that produces confident-sounding hallucinations often wins human preference battles over a more cautious, accurate model — because the hallucination is only identified later, if at all.
The practical consequence: ELO has historically rewarded confident wrong answers. This is the gap Arena is targeting.
For developers choosing models for information-retrieval, research, or customer-facing workflows, the combined factuality-preference leaderboard is more decision-relevant than human ELO alone. A model ranked first on human preference but third on factuality is a different risk profile than one ranked consistently across both.
Current Coverage
The factuality signal is live in Text Arena and Search Arena. Arena plans to expand to additional leaderboard categories. The toggle is off by default — users who want the combined ranking must enable it manually in the leaderboard UI. Arena describes the non-default status as an initial launch posture, not a permanent constraint.
The Broader Signal
Arena is the only leaderboard with the data volume and battle infrastructure to make factuality scoring credible at scale. Independent evaluations from Adobe and others have documented model hallucination rates well above what human raters catch — GPT-5.5 has been measured at 86% hallucination rate in some evaluations. If factuality consistently diverges from human preference in the Arena data, it will pressure labs to optimize for claim accuracy rather than response confidence.
The toggle exists. When it moves to default-on, the leaderboard changes.