GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

3.2 Million Math Records: AI Lifts Completion Speed 23%, Drops Proctored Retention 25%

A study of 3.2 million ALEKS math learning records spanning 10 years has isolated a clean trade-off in AI-assisted learning: students complete problems faster, then answer fewer of them correctly when AI is removed.

The paper — “Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build” (arxiv.org/abs/2605.21629) — uses an elegant natural experiment that separates AI’s effect from confounding variables like platform changes or general student improvement.

The Design

ALEKS, a widely deployed adaptive math platform, uses two structurally different problem types:

Word problems — text-based, pasteable into any chatbot. Students can extract the problem, receive a step-by-step solution, and submit the answer without having done the mathematical reasoning.

Graph problems — require visual manipulation inside the platform. They cannot be meaningfully handed to a language model without losing the core challenge.

After ChatGPT became publicly available, researchers tracked changes in time spent on each type. Students spent significantly less time on word problems. Time on graph problems was unchanged. The divergence disappeared on proctored sessions, ruling out platform updates or general capability gains as explanations.

The 25% Retention Finding

The core result: on proctored retention questions administered after the learning session, students became 25% less likely to correctly answer AI-friendly word problem items compared to pre-ChatGPT cohorts. AI-resistant graph problems showed no equivalent degradation.

The proposed mechanism is straightforward. Math builds durable knowledge through cognitive friction: choosing a representation, testing an approach, encountering failure, and correcting it. When a chatbot supplies the path, the student submits the correct answer without having done the work that encodes the knowledge. Completion rates improve; retention does not.

Who Is Affected

The effect was concentrated in high school and college students. Younger students — who are less likely to strategically deploy external tools — showed smaller or no change. The pattern is consistent with older students actively offloading AI-delegable tasks while younger students either lack the tools or the inclination to use them.

The Companion Finding on Perceived Efficiency

A separate study (arxiv.org/abs/2605.22687), drawing on 2,691 participants across MIT, Stanford, New York University, and Princeton, quantified the expectations gap on the time side. Students expected AI to save an average of 55.7 seconds per problem. The measured saving was 7.5 seconds.

The completion speed benefit is real but minor. The retention cost — 25% on proctored tests — is not.

The two findings compound: students consistently overestimate how much time AI saves, and they consistently underestimate what they lose by not doing the cognitive work. After using AI on just two tasks, participants in the efficiency study became more likely to reach for it again, even when their own unaided performance would have been faster and produced a durable skill improvement.

What This Means for AI in Education

The ALEKS result does not say AI makes students worse at math. It says AI assistance on tasks that could be done independently — and that build foundational skills through practice — substitutes completion for learning. The students are not failing; they are finishing. The gap shows up later, in controlled conditions where the scaffold is gone.

The graph problem finding is the control that makes this interpretation credible. It is not that ALEKS got harder, or that student populations shifted. The problems that resisted AI handoff showed no degradation. Only the problems that were easy to delegate did.

The policy implication for educators is not to ban AI tools, but to recognize that completion metrics and retention metrics have decoupled. A student who finishes assignments faster on AI-amenable problem types may be producing worse learning outcomes on the same material.