GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

146,932 Hallucinated Citations Entered the Scientific Record in 2025. Now Venues Are Fighting Back.

A paper posted to arXiv this week (2605.07723) presents the most comprehensive audit of AI-induced citation contamination in the scientific literature to date. The numbers are not reassuring.

Researchers audited 111 million references across 2.5 million papers spanning arXiv, bioRxiv, SSRN, and PubMed Central, using a reference verification pipeline to identify citations that do not correspond to any real publication. Their conservative estimate: 146,932 hallucinated citations entered the scientific record in 2025 alone, projected from monthly rates reaching 3,353 on arXiv, 478 on bioRxiv, 767 on SSRN, and 8,140 on PubMed Central as of August 2025.

The rise begins sharply in mid-2024, roughly 18 months after widespread LLM adoption. Despite concurrent improvements in reasoning models and retrieval-augmented generation, the hallucination rate shows no sign of plateauing through the study’s data cutoff in late 2025.

What gets through

The existing safeguards are largely failing. arXiv moderation — 240 volunteer reviewers — does catch a disproportionate share of papers with hallucinated references: rejected manuscripts reach a 2.2% hallucination rate versus 0.5% for accepted ones. But the volume of submissions has far outpaced what moderation can handle. The paper estimates that 78.8% of non-existent citations pass arXiv moderation and appear on the platform.

When papers are traced from preprint to journal publication, 85.3% of the hallucinations that survived arXiv’s moderation also survived peer review and appear in the final published record.

Venue-level enforcement arriving

Conferences are now moving from documentation to enforcement:

  • ICLR 2026 assembled a desk-reject queue of more than 600 submissions flagged for fabricated references. An independent multi-agent detection framework evaluated 647 of those submissions and flagged 796 citations as hallucinated at 98.6% recall.
  • ICML 2026 and ACM CCS 2026 have announced citation-verification requirements for the 2026 cycle.
  • NeurIPS 2025 documented more than 100 fabricated citations across 53 accepted papers — references that passed peer review by three or more reviewers.

A separate study (2601.18724) found nearly 300 papers with at least one fabricated citation across ACL, NAACL, and EMNLP conferences in 2024 and 2025, with half of them appearing at EMNLP 2025 alone.

Who gets blamed

Hallucinated references are not distributed randomly. The arXiv paper finds they are concentrated in manuscripts with linguistic signatures of AI-assisted writing, in fields with rapid AI uptake, and among small and early-career author teams. They also disproportionately assign credit to already-prominent and male scholars — LLMs fabricating plausible-sounding citations tend to invent names that match patterns in their training data.

Thomas Dietterich, who chairs arXiv’s computer science section, posted a thread on May 14 reaffirming the platform’s code of conduct: authors take full responsibility for all content in their papers, irrespective of how it was generated. Hallucinated citations are the author’s problem, not a mitigation arXiv will provide.

The verification gap

A February 2026 paper surveying 97 researchers found that 41.5% copy-paste BibTeX without checking validity, 44.4% take no action when they encounter suspicious references, and 80% of peer reviewers never suspect fabricated citations during review. The infrastructure assumption — that citations are real unless proven otherwise — is being systematically exploited by LLM writing assistance.

Tools for automated citation verification exist and are improving. The harder problem is changing the incentive structure: reviewers are already overwhelmed, and hallucinated citations in submitted papers are indistinguishable from real ones without active verification.

The trajectory is clear. 146,932 in 2025, with no plateau. The venues enforcing desk-rejects are addressing the symptom. The underlying mechanism — LLMs that produce fluent, plausible-looking citations at high rates — is still in every writing tool researchers are using.