GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

AI Medical Chatbots Hit 95% in Lab Tests — Then Drop to 35% When Real Patients Talk to Them

Controlled evaluations of AI medical chatbots consistently report accuracy around 95% — which is why they’re being deployed at scale as a first-line health information layer. A BBC investigation published this week, drawing on converging research, puts a number on what happens when the same systems interact with actual patients: accuracy drops to approximately 35%.

The gap — 60 percentage points between lab performance and real-world performance — is not a model failure. It’s a measurement failure. Lab benchmarks present clean, complete, well-structured clinical descriptions. Real users give partial, distracted, meandering symptom reports where the relevant detail arrives in the wrong order, or doesn’t arrive at all.

The Mechanics of the Collapse

The problem is conversational — not computational. AI health systems are evaluated on cases constructed by medical professionals who know what information the model needs. Real patients don’t know what information the model needs. They describe their subjective experience, skip the objective details that would matter to a clinician, and use colloquial language that maps imprecisely to diagnostic categories.

In this environment, the stakes of phrasing asymmetry become acute. A wording difference of a few words — not a factual difference, just a phrasing difference — can produce advice recommendations that sit on opposite ends of the clinical action spectrum: “rest at home and monitor symptoms” versus “seek emergency care.”

A parallel evaluation published by Modern Healthcare, covering five major AI chatbot platforms, found 50% of AI medical advice was “problematic” — a broader category that includes incomplete, misleading, or inappropriately confident responses, not just factually wrong ones. A separate PublicTechnology analysis of the same five platforms found Grok performed worst.

The Deployment Problem

Roughly 25% of US adults used an AI tool for health information in the 30 days prior to an AP-NORC poll conducted in late 2025. That figure is increasing, and the use pattern matters: people are turning to AI chatbots for health questions at the point when they’re trying to decide whether to seek care — precisely the decision threshold where 35% accuracy and advice-flipping phrasing sensitivity are most consequential.

These systems were not designed for this use case in their current form. They were designed to answer health questions, and they’ve been deployed at scale because they answer most questions correctly in testing. The gap between “most questions correctly in testing” and “35% of real conversations” is not disclosed to users at the point of use.

What This Looks Like in Practice

A user describing chest tightness after exercise with shortness of breath might frame it as “I’ve been feeling a bit off after walking” — omitting the specifics that would trigger the right clinical response path. The same model that correctly identifies cardiac risk when given structured input generates a different risk assessment for the same underlying condition described informally.

The model hasn’t changed. The input distribution has.

The Accountability Gap

No current regulatory framework requires health AI systems to publish real-world accuracy figures alongside controlled benchmark performance. The FDA has guidance on AI medical devices, and the EU AI Act classifies high-risk AI applications including health — but accuracy disclosure norms for conversational health tools haven’t been formalised in either jurisdiction.

The practical effect: a chatbot can report 95% clinical accuracy in marketing materials and product documentation while operating at 35% accuracy with its actual user base. The 60-point gap is not visible to the users in that 25%-of-adults cohort making health decisions based on the answers they receive.

Key Numbers

  • Controlled test accuracy: ~95%
  • Real-user conversation accuracy: ~35%
  • Accuracy gap: ~60 percentage points
  • Problematic AI medical advice rate (Modern Healthcare): 50% across 5 platforms
  • US adults using AI for health info (AP-NORC, late 2025): ~25% in past 30 days
  • Worst performer (PublicTechnology evaluation): Grok
  • Clinical stakes of phrasing error: advice flip between home rest and emergency care