Executive Summary
Researchers gave roughly a thousand Turkish high-school students access to GPT-4 during mathematics practice, then took it away and tested them. During practice, the students with the assistant scored 48 percent higher than those without it. On the unaided exam that followed, the same students scored 17 percent lower than classmates who had never touched the tool.
A third group used the identical model, reconfigured to give hints rather than answers. Those students gained even more during practice — 127 percent — and lost nothing on the exam. Same technology, same students, same curriculum. The only variable was how the tool had been instructed to behave.
The question this report poses. Ninety-four percent of UK undergraduates now use generative AI for assessed work, and 18 percent of teachers have received any formal guidance on it. If the difference between a tool that teaches and a tool that damages learning is a configuration choice, who is currently making that choice — and on what evidence?
This report sets out what the controlled research shows, what schools are facing operationally, and where the evidence runs the other way. The judgment is left to the reader.
The Experiment That Isolated the Mechanism
Most claims about AI and learning are correlational, and therefore weak: students who use AI differ from students who do not in a dozen ways at once. The study published in PNAS by Hamsa Bastani, Osbert Bastani and colleagues avoids that problem by randomising. Roughly 1,000 students across about 50 classrooms, in grades 9 to 11, were assigned to three arms across four 90-minute sessions covering approximately 15 percent of the year’s mathematics curriculum.
The arms were: a control group with standard practice materials; a GPT Base group given an interface essentially identical to ChatGPT; and a GPT Tutor group given the same GPT-4 model with a modified system prompt containing the worked solution and the errors students commonly make, instructed to supply incremental hints and never the full answer.
The results, expressed against the control group’s mean, are unusually clean. During practice, GPT Base students performed 48 percent better and GPT Tutor students 127 percent better. On the subsequent exam, taken without any assistance, GPT Base students performed 17 percent worse than students who had never had access to the tool at all — a statistically significant decline. GPT Tutor students showed no significant difference from control.
The finding in one sentence: unguarded generative AI did not merely fail to help — it left students measurably worse off than if they had never used it, while simultaneously making their practice work look far better. The authors describe students using the model as a crutch.
Why This Is the Dark Side, Specifically
The danger identified here is not that AI produces bad work. It is the opposite, and it is worse: the visible output improves while the underlying capability degrades. Practice scores, homework quality, apparent engagement — the signals a school actually collects — all move in the reassuring direction. The damage surfaces only later, under conditions where the tool is absent.
This is a measurement failure before it is a pedagogical one. A school tracking assignment completion and homework grades through an AI-saturated year would see improvement. It would be measuring the tool, not the student. The structure will be familiar to readers of this publication: it is the same shape we found in enterprise AI spending, where activity metrics rose while nobody could demonstrate the underlying return, and in workforce disengagement, where the costs migrate to places no accounting system records. A metric that improves while the thing it proxies deteriorates is the most expensive kind of number.
Separate work at the MIT Media Lab, led by Nataliya Kosmyna, offers a possible neurological correlate. Fifty-four participants wrote essays under three conditions — unaided, with a search engine, and with an LLM — while EEG measured brain activity. Connectivity scaled down with the level of external support: unaided writers showed the strongest and widest-ranging networks, search users an intermediate level, LLM users the weakest. Self-reported ownership of the finished essay was lowest in the LLM group. The authors call the effect “cognitive debt.”
That study should be handled carefully, and this report does so below. It is a preprint with 54 participants, it measures a proxy for cognitive effort rather than learning itself, and it has attracted published methodological criticism. It is suggestive, not decisive.
What Schools Are Actually Facing
Adoption is effectively total. The 2026 HEPI/Kortext survey of 1,054 UK undergraduates found 94 percent using generative AI for assessed work and 95 percent using it in some form. Comparable rich-world figures run in the same range. Any policy premised on students not using these tools is describing a world that no longer exists.
The behaviour schools most want to stop is the one growing fastest. The share of students placing AI-generated text directly into submitted work has moved from 3 percent in 2024, to 8 percent in 2025, to 12 percent in 2026.
Teachers have been left to improvise. Gallup and the Walton Family Foundation found that six in ten US teachers use AI in their work and three in ten use it weekly — but only 18 percent have received any formal guidance from administrators on how it should be used. Weekly users report saving close to six hours a week, a genuine benefit against a backdrop of burnout. The staffroom has acquired a productivity tool and no policy.
Detection does not work, and it fails unevenly. Stanford researchers tested seven widely used AI detectors against TOEFL essays written by human non-native English speakers: 61 percent were flagged as AI-generated, 97 percent were flagged by at least one detector, and 19 percent were unanimously misclassified by all seven. The tools mistake the plainer sentence construction of a second-language writer for machine output. Vanderbilt, Cornell, Pittsburgh and Iowa have disabled such tools, citing reliability and equity concerns.
Put those four numbers together and the operational position is stark: near-universal student use, a rising share of outright substitution, almost no institutional guidance, and an enforcement mechanism that disproportionately accuses immigrant and international students. Whatever the right answer is, policing is not available as a strategy.
The Evidence That Runs the Other Way
A report that stopped there would be incomplete. The strongest counter-evidence is a World Bank randomised trial in Benin City, Nigeria, where about 800 senior secondary students attended after-school English sessions using Microsoft Copilot, led by teachers, over six weeks.
The effect on English, the pre-specified outcome of interest, was 0.23 standard deviations — a result that outperformed roughly 80 percent of educational interventions in a comparison database of randomised trials in developing countries. Gains were strongest among girls, including at a single-sex school with lower baseline scores.
Two features of that trial matter for the argument. It was teacher-led and structured — closer in design to the GPT Tutor arm than to a student alone with a chatbot. And its benefits skewed toward students with stronger prior attainment and higher socioeconomic background, likely reflecting greater familiarity with digital tools. Even where AI clearly helps, it may widen gaps while lifting averages.
| Study | Design | Result |
|---|---|---|
| Bastani et al., PNAS — Turkey | Unguarded chatbot during maths practice | −17% on the unaided exam |
| Bastani et al., PNAS — Turkey | Same model, hints-only tutor configuration | No harm; +127% during practice |
| World Bank — Nigeria | Teacher-led, structured after-school sessions | +0.23 SD in English |
| MIT Media Lab (preprint) | EEG during essay writing, n=54 | Weakest brain connectivity with LLM |
Read as a set, these studies do not contradict one another. They converge on a single variable: whether the tool is configured and supervised so that the student still does the cognitive work, or so that it does the work for them. Where the model hands over answers, learning falls. Where it withholds them — by prompt design, by teacher presence, or both — learning holds or improves.
The Question, Put Directly
If the harm is a function of configuration rather than of the technology itself, then someone is doing the configuring. In practice, three parties are.
The vendors, who ship consumer assistants optimised to be maximally helpful — which, in a homework context, means maximally answer-giving. The GPT Base arm was not a defective product. It was a well-designed product being used for a purpose its defaults are wrong for.
The schools, 82 percent of which have given teachers no formal guidance at all, and which are therefore configuring by default rather than by decision.
The students, 94 percent of whom are using these tools, and most of whom have never been shown the distinction between a model that hints and a model that answers — a distinction worth, in the only controlled test available, the entire difference between learning and not learning.
So: is the right question whether AI stops children from learning? Or is it why the version proven not to harm learning is the one almost nobody is using?
What This Analysis Does Not Establish
One subject, one country, four sessions. The Turkish trial covers high-school mathematics across roughly 15 percent of a curriculum. Mathematics is unusually well suited to a hint-based tutor because problems have determinate answers and well-catalogued error patterns. Whether the same design transfers to essay writing, history or foreign languages is untested.
Short horizons. Every trial cited here measures outcomes within weeks. None can say whether the exam deficit persists, or whether students who learn alongside AI develop different but equally valuable capabilities over years.
The MIT study is not evidence of harm to learning. It measures brain connectivity during a writing task, with 54 participants, in preprint form, and has attracted published criticism. Reduced neural effort during a task is not the same thing as reduced learning from it.
Falling test scores are not attributable to AI. Declining attainment among the lowest-performing students in US national assessments began more than a decade before ChatGPT existed. Any account linking recent score movements to generative AI is unsupported by the timeline, and this report makes no such claim.
Survey populations differ. The 94 percent figure describes UK undergraduates; the teacher-guidance figures describe US school staff; the Turkish trial covers secondary students. They are assembled here to describe a landscape, not a single population.
What to Watch
Whether the hints-only configuration reaches ordinary classrooms. The GPT Tutor arm required a bespoke system prompt containing worked solutions and anticipated errors for every problem. That is a curriculum-authoring task, not a settings toggle. Whether vendors ship it as a default — and whether schools adopt it — is the most consequential open question here.
Replication outside mathematics. A well-powered trial of guarded versus unguarded AI in writing instruction would settle far more than another round of opinion surveys.
Assessment redesign. Sixty-five percent of students already report that assessment has changed significantly in response to AI. Since detection has failed, the practical response is assessment robust to assistance — oral defence, supervised in-class work, process portfolios. Watch whether institutions build that, or keep buying detectors.
Conclusion. The best controlled evidence available says that generative AI, in the form students actually encounter it, made them 48 percent better at practice and 17 percent worse at the thing practice was for. It also says that the harm vanished when the same model was told to withhold answers. The dark side of AI in education is not a machine that produces nonsense; it is a machine that produces excellent work while the person operating it learns less than they would have alone — and every metric a school routinely collects will report that as success. Whether that amounts to AI stopping children from learning, or to adults failing to configure it, is the question this report leaves with the reader.
Sources
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS, 2025. https://www.pnas.org/doi/10.1073/pnas.2422633122
Kosmyna, N. et al. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. MIT Media Lab, arXiv:2506.08872 (preprint). https://arxiv.org/abs/2506.08872
World Bank Education Global Department, with the Mastercard Foundation. From Chalkboards to Chatbots: generative AI and learning outcomes in Nigeria. https://blogs.worldbank.org/en/developmenttalk/addressing-the-learning-crisis-with-generative-ai–lessons-from-
Higher Education Policy Institute and Kortext. Student Generative AI Survey 2026 (n=1,054 UK undergraduates), with the 2024 and 2025 editions for the trend. https://www.hepi.ac.uk/reports/student-generative-ai-survey-2026/
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. and Zou, J. GPT detectors are biased against non-native English writers, Stanford University. Coverage: https://themarkup.org/machine-learning/2023/08/14/ai-detection-tools-falsely-accuse-international-students-of-cheating
Gallup and the Walton Family Foundation. Most Teachers Receive No Formal Guidance on AI Use. https://news.gallup.com/poll/710534/teachers-receive-no-formal-guidance.aspx
The Economist. Does AI stop children from learning? Graphic detail, 18 August 2026 — the article that prompted this analysis.








