MIT, Oxford, CMU Study: 10 Minutes of AI Assistance Measurably Reduces Independent Problem-Solving Persistence
Researchers from MIT, Oxford, Carnegie Mellon, and several other institutions have published findings that cut directly against the standard case for AI productivity adoption: GPT-5-based assistance improved task performance while available, then measurably reduced independent performance shortly after.
The paper — “AI Assistance Reduces Persistence and Hurts Independent Performance” (arXiv:2604.04721) — ran three experiments across approximately 1,200 participants working on math and reading tasks. Participants given direct AI answers completed early problems faster. Within roughly 10 minutes of the AI being removed, those same participants solved fewer problems, stalled more often, and quit sooner than the control group.
The Mechanism: Persistence, Not Accuracy
The researchers make a precise distinction between two types of performance degradation: knowing the wrong answer (accuracy deficit) and stopping effort sooner (persistence deficit). AI-assisted participants showed the second type. They remained capable of attempting problems but gave up earlier — a pattern consistent with reduced tolerance for sustained difficulty rather than reduced knowledge.
That finding points to a specific mechanism. Skill development, the paper argues drawing on educational research, requires repeated contact with difficulty — not exposure to correct answers. AI tools optimised for friction removal compress away exactly that contact. Users get the answer; they do not build the habit of holding a hard problem in mind and pushing through confusion.
The Hint Mode Finding
The most policy-relevant result in the paper is the hint mode comparison. Participants who used the AI as a hint system — receiving partial guidance rather than complete answers — showed substantially smaller persistence drops compared to participants who received direct answers.
That isolates the harm source. The problem is not AI exposure or AI-assisted work per se. It is the specific pattern of replacing effortful thinking with answer retrieval. A tool configured to scaffold rather than complete produces a measurably different cognitive outcome in the short term.
Why This Study Is Harder to Dismiss
AI-and-cognition research has a credibility problem: most of it relies on self-reported data or single-session surveys without proper controls. This study controls for prior ability, randomises AI access, and tests the persistence dimension directly using task-completion timing and drop-off rates rather than asking participants how they felt about the AI.
Three separate experiments across different task types — math and reading — showing consistent effects adds generalisability that single-study AI cognition claims typically lack.
Key Numbers
- Participants: ~1,200 across 3 experiments
- Task types: Math and reading comprehension
- Recovery window: Performance gap measurable within ~10 minutes post-AI removal
- Key variable: Answer mode vs. hint mode (hint mode shows substantially smaller persistence loss)
- arXiv: 2604.04721