GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

LLMs Are Converging to the Same Output — and the Spam Economy Is Built on It

Ask a modern instruction-tuned model to write a short story. Ask it again in a separate session. You will often get the same character, the same setting, the same opening clause. Two independent Gemini 2.5 Flash-Lite runs on the same prompt — “Write a story in 10 sentences” — both open with:

The old lighthouse keeper, Elias, polished the brass railing…

This is not a fluke. It is a documented property of instruction-tuned models optimised for human preference: training collapses the output distribution toward what evaluators rewarded most, producing the same lighthouse keeper across millions of independent sessions. The phenomenon has a name — mode collapse — and it is now the substrate under an entire tier of commercial AI deployment.

The Business Model That Runs on Convergence

The clearest illustration is not in the model labs but in the operational layers built on top of them. Cold outreach agents mine stale contact databases, generate personalised-sounding emails, and send them at 3 AM from names like Bruce and Ava. When they reach the wrong target — a professional who moved cities, left the company, or simply has no need for the product — no one corrects the record. The check costs more than the mistake.

The same economics run in product and content spam. Thousands of AI-generated t-shirt listings exist for every small institution in the country — veterinary clinics, fire departments, high schools — waiting for the fraction of notification recipients who half-remember the place and click through. The alumni shirt for a veterinary clinic whose clients are not graduates. The DMARC fix offer for a domain whose configuration is intentional. The personalisation is real; the relevance is not. The system works because the hit rate does not need to be high. Unit economics clear at one sale per thousand notifications.

Why Nobody Checks

The structural problem is that the quality floor of AI-generated content has dropped below the cost of verifying it. Operators set up the pipeline and let it run. The model produces what the inputs constrain it to produce. The inputs are what the operator chose not to invest in.

This creates a specific failure mode: content that is fluent, confident, and wrong. Not random noise — targeted misfires, plausible in form, broken in substance. The DMARC email that correctly identifies the domain but misreads why the configuration looks unusual. The cold email that addresses you by a job title you held a decade ago. The product listing that generates medically accurate-sounding language about a treatment category the seller knows nothing about.

Mode collapse is what makes this possible at scale. If models produced diverse outputs, the cost of verifying any individual piece would still be high, but the population of errors would be diffuse. Because models converge, you get the same errors at volume. The same lighthouse keeper appears in thousands of creative writing samples. The same templated confidence appears in thousands of product descriptions. The sameness is the efficiency.

The Implication for Digital Trust

The commercial tier of the internet is now partially filled by content that is identical in origin, confident in tone, and unverified in fact. This is not the frontier model problem — GPT-5.5 and Opus 4.7 show real divergence under complex prompting. It is the cheap-inference problem: the 2026 equivalent of the content farm, except the content is more fluent and the scale is higher.

The gap between “sounds correct” and “is correct” was always present in low-quality web content. Instruction-tuned models have widened the gap by increasing fluency without proportionally increasing accuracy. At the current cost of inference, the economics favour generating and distributing over verifying and correcting.

The downstream effect is not just spam. It is a slow contamination of the training data that future models will learn from — the same loop identified by a January 2026 medRxiv preprint on medical AI, where synthetic content fed back into training produces models that are more confident and less accurate, with false reassurance rates tripling after two generations of self-referential training.

The lighthouse keeper is selling cancer treatment advice on Amazon. And no one checked.