The most striking number in a new study of 26,811 Chinese secondary students is not the 18% jump in homework scores. It is the 20% drop in exam scores that followed within six months, and the 18% to 24% decline in college entrance exam results measured over two years. The working paper, published through the Centre for Economic Policy Research by David Strömberg of Stockholm University and Victor Lei and Yanhui Wu of the University of Hong Kong, tracked students in grades seven through twelve across a county in central China over 30 months, comparing AI users against a control group.
The divergence is the story. Homework completion time fell from an average of 64 minutes to 45 minutes per assignment. Around 80% of students used generative AI platforms such as Doubao and DeepSeek. And roughly 80% of AI users displayed what researchers call homework outsourcing: finishing assignments unusually quickly while still scoring highly. Those outsourcers drove almost all of the learning losses, in both the short and long term.
The transfer problem, restated for the AI era
The researchers put it plainly: “For students, completing these tasks efficiently is not the goal; learning from them is.” Their conclusion is that generative AI, “which is likely to become a prevalent technology for education, has a substantial negative impact on student learning.”
That finding should worry anyone building AI tools for classrooms, and anyone betting that AI-assisted education will automatically improve outcomes. The mechanism at work is old. Neuroscientist Jared Cooney Horvath, who testified to the U.S. Senate Committee on Commerce, Science, and Transportation on this topic, traces it back to 1924, when Ohio State University psychology professor Sidney Pressey invented the “teaching machine.” Students answered questions the machine presented, but could not generalize the knowledge when tested off the device. B.F. Skinner built a more advanced version three decades later, with the same result. Both psychologists abandoned the project. In a letter to Skinner, Pressey conceded that students had not mastered the subject matter; they had mastered the machine.
“The reason they all quit was the transfer problem,” Horvath said. “They found that kids would be very good so long as they were using the tool, but as soon as they went off the tool, they couldn’t do it anymore.”
The Chinese study is the transfer problem at population scale. Students who used AI to outsource homework got better at using the tool, not at the underlying material. When the tool was removed, the learning deficit showed up in closed-book exams.
Not all AI use is equal
The study’s most useful distinction is between AI as tutor and AI as substitute. Students who used AI but took as long on their homework as non-users experienced comparatively small learning losses. The damage was concentrated among those who finished unusually quickly. That is the signature of outsourcing: the tool did the thinking, and the student did not.
Losses were largest in social science subjects, followed by STEM and languages. They were most pronounced among younger students, high achievers, and boys. In China, the effect extended to national examinations: the zhongkao, the high school entrance exam, recorded a 24% decline among affected AI users, while the gaokao, the university entrance exam, fell by 18%.
A separate study by Zara Contractor and Germán Reyes of Middlebury College found the opposite pattern: undergraduates using a chatbot to learn an unfamiliar topic performed better on tests, with the advantage persisting a week later. The difference between the two studies is not the technology. It is the use pattern. The chatbot learners in the Middlebury study used AI as a study aid. The Chinese outsourcers used it as a way to finish faster.
The incentive structure is the problem
The study’s authors are clear that the issue is not simply student laziness. The incentives to use AI to skip learning are “overwhelmingly powerful,” as the research puts it. Homework is graded. Exams are graded. But homework is graded immediately and exams are graded months or years later. Students optimize for the immediate signal.
That mismatch is structural. Schools grade the output, not the process. AI makes the output cheap to produce. When the output is all that is measured, students rationally choose the cheap path. The exam score drop is the delayed cost of that rational choice, borne by the student years later, when no one is watching.
Jacob Shelley, an associate professor of health law at Western University, told Fortune in May that he was convinced his students cheated on a final exam using AI: 8% got a perfect score on the multiple choice section, only to struggle on the essay portion, submitting answers with content not in the curriculum. “The results were anomalous,” Shelley said. “That just never happened in 20 years of teaching.” Yet he does not blame the students. “AI is going to replace them, at least a lot of them, and they know that, and we’re pretending that it won’t,” he said. “I think they see through it.”
Almost 90% of graduates from the class of 2026 are worried AI or automation could replace entry-level jobs, according to job search platform Monster. Students are not cheating because they are lazy. They are cheating because they believe the skills being taught in school will not matter in a labor market where AI does the entry-level work.
What this means for AI builders
The study is a warning to the AI education industry, not just to schools. The market for AI tutoring tools is growing on the assumption that more AI in the classroom means better learning. This data says the assumption is conditional. AI that reduces friction improves homework metrics and degrades learning. AI that adds friction, that forces the student to engage with the material, preserves learning.
The researchers’ own framing points to the design implication. Tools that individualize learning by generating answers do not produce the friction necessary for learning, Horvath argues. “The tools experts use to make their lives easier are not the tools children should use to learn how to become experts,” he said. “When you use offloading tools that experts use to make their lives easier as a novice, as a student, you don’t learn the skill. You simply learn dependency.”
There is a design path forward. An AI tutor that asks questions, withholds answers, and forces retrieval practice would look very different from the chatbot that generates a complete essay in seconds. The Middlebury study suggests such tools can work. The Chinese study suggests the market is currently shipping the wrong kind.
The measurement problem
The deeper issue is that homework scores are the wrong metric, and AI has made them worse. The 18% homework gain is not a productivity gain in any meaningful sense. It is a measurement artifact: students are being graded on how well they can prompt a model, not on what they know.
The study’s authors call this a “dissonance in AI productivity versus actual productivity gains.” The homework scores said students were learning more. The exams said they were learning less. One of the two measurements was lying, and it was the one that schools use to assign grades.
Adoption data suggests this is not a China-specific problem. Recent polling puts AI usage among undergraduates at 94% in Britain and 93% in Germany. A CollegeBoard survey of more than 1,000 U.S. high schoolers found 84% report using the technology for homework. The Chinese study is the first large-scale longitudinal evidence of what that adoption does to learning. The numbers are consistent with the anecdotal reports from teachers like Shelley.
The study leaves open the question of what to do about it. Banning AI in classrooms is one option, but the adoption rates suggest enforcement would be difficult. Teaching students to use AI as a tutor rather than a substitute is another, but it requires schools to change how they grade, and how they define learning.
What is clear is that the optimistic case for AI in education, the case that says more capable models will automatically produce better students, is not supported by this data. The models are more capable. The students who outsource to them are learning less. The gap between what the tool can do and what the student can do is the entire problem, and the current generation of AI tools is widening it.
The study’s final implication for builders is uncomfortable: the most valuable AI education tools may be the ones that deliberately make themselves less useful, that refuse to generate the answer, that force the student to struggle. The teaching machine failed because it made learning too easy. The same failure mode is now shipping at scale, and the exam scores are the proof.