GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Ontario Audit: 12 of 20 Government-Approved AI Medical Scribes Got the Drug Wrong

Ontario’s Auditor General Shelley Spence released a special report on May 13 covering AI use across the provincial public service. The findings on medical AI scribes are the kind that make a procurement process look like it was designed to fail.

The province approved 20 vendors to supply AI transcription software to family doctors and other healthcare workers. Before approval, Supply Ontario ran two standardised test conversations between a simulated doctor and patient through each system. Here is what the tests found:

  • 12 of 20 systems captured a different drug than what the doctor prescribed
  • 9 of 20 systems fabricated information — referring patients for therapy, ordering blood tests, or suggesting clinical steps that were never mentioned in the recording
  • 17 of 20 systems missed key details about the patients’ mental health issues in at least one of the two tests

Eleven of the 20 approved vendors did not submit third-party audit reports, ISO 27001 certifications, or threat risk assessments during the procurement process. Five skipped threat risk assessments and privacy impact assessments. All were approved anyway.

The procurement weighting

Spence’s report details how Supply Ontario scored vendors:

CriterionWeight
Domestic presence in Ontario30%
Data privacy/legal controls23%
System security controls11%
Accuracy of medical notes4%

A vendor could score zero on accuracy, security controls, and bias mitigation and still clear the minimum threshold for approval. The auditor notes this explicitly.

Scale of deployment

The province says the systems save doctors five to seven hours of administrative work per week on average. Provincial Minister Stephen Crawford insisted the errors occurred during the testing phase, not live use, and that doctors review all AI-generated notes before acting on them. Auditor General Spence pushed back: she visited her own doctor during the audit and noticed they were using one of the approved AI scribes in a live appointment.

A privacy breach in September 2024, traced to an unapproved AI scribe in use at a hospital, preceded the province’s formal April 2025 rollout of the approved vendor list.

Spence made 10 recommendations. The province accepted nine. The one it rejected: increasing the weight given to security and privacy controls in future AI procurement scoring, which Supply Ontario said was “appropriate” as currently configured.

The accepted recommendations focus on mandatory doctor attestation that notes were reviewed, annual third-party audits in vendor contracts, and better risk assessment processes for procurement.

Context

Ontario is not the only jurisdiction running into this problem. A broader pattern has emerged across healthcare AI deployments: systems that perform well in controlled evaluations degrade substantially when variables include real patients, regional accents, medical terminology variation, and conversational context. The province approved 2024-era systems. Critics of the report, including some healthcare AI consultants, note that current-generation scribes have improved. The auditor’s position: the 2025 rollout used the results from those evaluations, and the evaluation process itself was the problem regardless of when the testing took place.

The structural issue — low accuracy weighting in procurement, vendors approved without security documentation — is a policy failure that will repeat unless the weighting changes.