How can AI enable real learning?
What the research actually says, kept current.
How can AI actually help someone learn, and not just finish faster?
For a few months I have been following that question, and I wanted to step back from the individual headlines to draw the bigger map. This is that attempt: the through-line across the research, and the difference between an AI that answers you and one that makes you better at what you do.
Two honest caveats. First, I am not an education researcher: this is my map of what the research currently shows, meant to provoke a better conversation, definitely not to close one. Second, it began as a snapshot of June 2026, and the field moves fast: the models, the tools and the studies keep shifting, and some of this will date. So the map now lives: when new evidence arrives it sharpens a beat in a small note right where the claim stands, or earns a new beat of its own, and every move is logged in how this map has changed. You will find the sources at the end, or read the original carousel as a PDF.
The problem
AI can improve performance and reduce learning.
In a randomised study of about 1,000 high school maths students in Türkiye, students using a standard ChatGPT scored 48% higher than classmates without it while they had it. When access was taken away, they scored 17% lower than students who never had it. A 2025 paper in Nature Reviews Psychology makes the underlying point: performance gains are not the same as learning.
Source: Bastani et al. (2025), PNAS; Yan, Greiff, Lodge & Gašević (2025), Nature Reviews Psychology.
Last sharpened Jul 2026
Better output is not the same as learning.
Why it happens
When AI does the thinking, the learning does not happen.
Most AI tools were built for work, not to optimise learning. At work, the goal is to finish the task with the least effort. But in learning, that effort is the point: it is what builds the capability. The task gets finished, but the understanding never develops. Researchers call this “metacognitive laziness”: the learner stops planning, monitoring and self-evaluating, because the AI always has an answer.
Source: Khosravi et al. (2026); Fan et al. (2025), BJET.
If the AI carries the effort, it also carries off the learning.
Why effort is the point
The struggle is not an obstacle to learning. It is the learning.
Learning scientists Robert and Elizabeth Bjork call these “desirable difficulties”: things like recalling an answer from memory, spacing your practice out, or trying a problem before you see the solution. They make learning feel harder now, but they make it last. When AI removes that effort, it can quietly remove the learning the effort produced. A 2026 review of 67 studies names the same mechanism, epistemic friction: without it, “AI-generated fluency can bypass the reflective struggle central to deep learning.”
Source: Bjork & Bjork; Li, Cui & Hagedorn (2026), Computers and Education: AI.
Protect the difficulty that does the teaching.
It starts with the person
When intelligence is plentiful, volition is valuable.
If the struggle is the learning, the first thing it depends on is the person doing it. Writing in The Atlantic, David Brooks argues that what will set people apart in an AI-saturated world is not how smart they are but their relationship to mental effort. His frame is need for cognition, the trait named and measured by John Cacioppo and colleagues, who reviewed more than 100 studies of it: at one pole, people who read dense books and play hard games for pleasure; at the other, the cognitive miser, who avoids effortful thought where possible. It correlates with intelligence, Brooks notes, but is not the same as it.
The evidence he marshals is largely about what happens when effort is offloaded. An MIT Media Lab team led by Nataliya Kosmyna measured brain connectivity dropping by up to 55% during ChatGPT use; Michael Gerlich (SBS Swiss Business School) found a significant negative correlation between frequent AI use and critical thinking; a Carnegie Mellon team led by Grace Liu found that after about ten minutes of AI-assisted problem solving, people who then lost the tool did worse than those who never had it. A study of endoscopists found precancerous-lesion detection fell from 28.4% to 22.4% once AI was withdrawn. From this he sketches three responses: the Productive Passenger, who lets AI think and stops noticing; the Reluctant Optimizer, who means to resist and gets pulled in; and the Mental Marathoner, who keeps the hard parts hard. His worry is a cognitive polarization between the two ends. His practical rule: ask AI for thinkers, not thinking, and treat it as a brilliant librarian, not an oracle.
Source: David Brooks (2026), The People Who Will Thrive in the AI Age, The Atlantic; on need for cognition, Cacioppo & Petty (1982), Journal of Personality and Social Psychology.
Last sharpened Jul 2026
The first variable is the learner, not the tool.
With what's already in their head
When answers are plentiful, background knowledge is the bottleneck.
Volition is not the only person-side variable. Stephen Fitzpatrick, a history teacher of thirty years, argues that the skill gating all of this is reading: AI output is fluent, confident and endless, and reading it with the skepticism it demands takes the vocabulary and background knowledge that only years of reading build. “The bottleneck to effective student AI use is mostly a problem of reading”, not writing, which gets the attention because writing is what we grade. Deep research makes the point sharper: finding sources is no longer the hard part; reading the report, and its sources, carefully is.
Marie Dollé, writing in French about the summer’s “brainmaxxing” apps, takes the same point one level deeper. Against the comforting idea that machines can hold the knowledge while we keep the judgment, she argues that judgment cannot form in a vacuum: “on ne problématise pas sans repères”, you cannot frame a problem without reference points already in your head. And she names a limit no better model fixes: what matters is often only knowable later, so no tool can pre-sort “the essential” for us. Consulting a fact on demand is not the same as having it in mind.
Daniel Susskind, an economist who has spent fifteen years studying AI and work, gives the same claim a mechanism. AI systems are what Geoffrey Hinton calls “idiot savants”: impressive on hard problems, wrong on easy ones, and nothing in the output tells you which one you are reading. So the basics are not a safe harbour, they are an instrument. Use AI critically rather than blindly, he argues, “keeping our basics sharp”, precisely so you can tell the savant from the idiot. Which is why he wants literacy and numeracy taught intensely even where AI already does them better, and calls it a no-regrets investment: one that pays off whatever the future turns out to be. The backdrop is not reassuring. Since 2009, literacy and numeracy have been falling among young people worldwide, according to the OECD’s PISA programme, and among adults too.
Source: Stephen Fitzpatrick (2026), It's the Reading, Stupid, Fitzy's History; Marie Dollé (2026), La star de l'été…, mariedolle.substack.com; Daniel Susskind (2026), The Guardian.
Last sharpened Sep 2026
You can only check a machine against what you already hold.
Then the design
The same technology can help or harm. The design decides which.
In a Harvard physics study, a purpose-built AI tutor designed with proper scaffolding beat in-person active learning by 0.73 to 1.3 standard deviations, two to three times the usual bar for a substantial effect in education research. In the Türkiye study, the standard ChatGPT left students worse off once it was removed, while a guardrailed tutor version avoided that loss entirely. Same models, opposite outcomes, depending on how they were designed.
Source: Kestin et al. (2025), Scientific Reports; Bastani et al. (2025), PNAS.
The result is set by the design, not by the model.
And by the human behind it
Believing a human is paying attention changes how hard we try.
In a controlled study in a university creative-coding course, students received identical AI-generated feedback on their work. Those told it came from a human teaching assistant ran their code more, wrote more code, and spent more time on later work. They rated the feedback equally helpful either way. The effect on effort was large (d = 0.88 to 1.56). The content was the same.
Source: Morris & Maes (2026), Same Feedback, Different Source.
Same words land differently when we believe a human wrote them.
What only a human does
AI can help with the content. It rarely touches the rest.
Education does three things at once: it builds knowledge and skills (qualification), it helps you find your place among others (socialization), and it helps you become someone who thinks independently and takes responsibility (subjectification). AI tools mostly reach the first. They rarely, if ever, address the other two, and those are where a teacher does their deepest work.
Source: Gert Biesta; Wayne Holmes (2026).
AI can teach the content. A human helps you become someone.
Three ways AI can show up
An LLM, a tutor, and a learning companion are not the same thing.
An LLM answers your question. Faster work, less learning.
An AI tutor asks questions back, no matter what you actually need. Often frustration and drop-out.
An AI learning companion (Dr Philippa Hardman calls it a “study mate”) remembers where you got stuck and pushes you towards the thinking you avoid. Capability that lasts.
Source: Khosravi et al. (2026); Dr Philippa Hardman.
Aim for a learning companion, not an answer machine.
What education is for now
AI does not shorten what there is to learn. It lengthens it.
Holmes (UNESCO) asks it directly: if generative AI is this powerful, do we still need to learn? His answer is yes, and the list grows: on top of what we wish to learn, we now need to learn AI’s profound limitations, its impacts on human rights, social justice and the environment, and “perhaps most importantly, [to] learn how to think… critically.” The World Economic Forum (WEF) keeps the balance: rote memorisation may matter less, but “the process of mastering knowledge continues to develop broader capabilities” such as grit, curiosity, communication and critical thinking, and assessment must evolve to capture them.
Source: Wayne Holmes (2026), UNESCO Courier; WEF (2026).
The emphasis moves from having answers to judging them.
So how do we make it help?
The challenge is also systemic. It runs through the tool, the classroom, and the system around them.
The tool: the guardrailed tutor and the learning companion are instructional design built into software: scaffolding, answers withheld, help that fades as you grow.
The classroom: co-design tools with teachers rather than deploying them on teachers, and set tasks that make learners compare, justify and revise what the AI produces. In the 67-study review, that scaffolding is what separated gains from cognitive offloading.
The system: AI only helps where the conditions are ready. As the WEF puts it, “learning outcomes will not be determined by technology itself, but by the conditions in which it is deployed”, and isolated fixes across policy, pedagogy and technology are unlikely to be sufficient.
In August 2026, the system layer got its first institution-scale test case. MIT’s committee on AI in teaching and learning refused patching (“this is not a moment for patches and duct tape”) and called for all three moves in concert: AI-aware redesign of every subject, an in-person social component in every subject, and permanent adaptation structures (a standing committee, per-school AI leads, a pilot fund), on a platform deliberately tied to no single AI vendor. One survey number shows what is at stake: MIT undergraduates report feeling more replaceable than capable (40% versus 34%).
Susskind adds the one move here with a precedent. In 1982 a British government report by Wilfred Cockcroft answered the electronic calculator not by banning it but by splitting mathematics in two: part of the time learning to work with a calculator, the rest learning to cope without one, and, the decisive part, examining both. That split is now the standard almost everywhere. He proposes the same structure for AI, in every subject, and calls it teach both, test both. His reason for putting the weight on assessment rather than detection is the part worth keeping: a teacher cannot know whether a student used AI alone in their bedroom, but nobody forgets sitting an exam they have only half prepared for.
Source: Bastani; Khosravi; OECD (2026); Li/Cui/Hagedorn; WEF; MIT AI committee (2026); Susskind (2026).
Last sharpened Sep 2026
Real progress needs all three moving.
- Protect the difficulty that does the teaching.
- The result is set by the design, not by the model.
- AI can teach the content. A human helps you become someone.
- Aim for a learning companion, not an answer machine.
- The differentiator is your relationship to mental effort.
One test cuts through it all: ask who is doing the thinking, you or the AI. Keep yours alive, and use AI to learn, not just to finish.
Are you here for the output, or to get better at what you do?
A map of the research, drawn over time and kept current as new sources arrive.
Download the original carousel (PDF)- Bastani et al. (2025), Generative AI without guardrails can harm learning, PNAS
- Khosravi et al. (2026), Building AI Companions that Prioritise Learning over Performance
- Kestin et al. (2025), AI tutoring outperforms in-class active learning, Scientific Reports
- Morris & Maes (2026), Same Feedback, Different Source
- Yan, Greiff, Lodge & Gašević (2025), Nature Reviews Psychology
- Li, Cui & Hagedorn (2026), Computers and Education: AI
- Fan et al. (2025), Beware of metacognitive laziness, British Journal of Educational Technology
- Inara Scott (2026), The AI Cognitive Pyramid, SSRN
- Robert & Elizabeth Bjork, desirable difficulties
- Gert Biesta, three functions of education
- OECD, Digital Education Outlook 2026
- Wayne Holmes (2026), Learning to think in the AI era, UNESCO Courier
- World Economic Forum, Shaping the Future of Learning (2026)
- Dr Philippa Hardman, From AI Tutors to AI Study Mates
- David Brooks (2026), The People Who Will Thrive in the AI Age, The Atlantic
- Cacioppo & Petty (1982), The need for cognition, Journal of Personality and Social Psychology, 42(1), 116-131
- Stephen Fitzpatrick (2026), It's the Reading, Stupid, Fitzy's History
- Marie Dollé (2026), La star de l'été…, mariedolle.substack.com
- MIT Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training (2026), Final Report
- Daniel Susskind (2026), I'm a father of three who studies the impact of artificial intelligence: this is what parents need to know about AI, The Guardian
This is the provenance ledger: every source that has moved the map since it was drawn, in the order it arrived. It never renders in the main reading flow above. It is the only place a date appears on this page. A JSON feed is derived from the same list.
- New section
Founded section 5, “When answers are plentiful, background knowledge is the bottleneck.”
Stephen Fitzpatrick brings reading as the gating skill; Marie Dollé brings the repères argument (judgment runs on reference points already in the head, and no tool can pre-sort what will matter); Daniel Susskind supplies the mechanism, that AI’s confident errors are undetectable from the output, so the basics are the instrument you check the machine with, plus the OECD finding that literacy and numeracy have been falling worldwide since 2009. Honest note: two practitioner essays and a newspaper essay by an economist, not studies. Their shared claim is exactly what someone should now measure.
Stephen Fitzpatrick (2026), It's the Reading, Stupid, Fitzy's History; Marie Dollé (2026), La star de l'été…, mariedolle.substack.com; Daniel Susskind (2026), The Guardian
- Sharpens
Attaches to section 11, “The challenge is also systemic. It runs through the tool, the classroom, and the system around them.”
Extends the final section. MIT’s AI committee report (August 2026) gives the system layer its first institution-scale case: all three moves called for in concert, on a vendor-neutral platform. Honest note: a governance report, not an outcome study; its recommendations were unfunded at publication.
- Sharpens
Attaches to section 11, “The challenge is also systemic. It runs through the tool, the classroom, and the system around them.”
Extends the final section. Daniel Susskind revives a forgotten precedent: the 1982 Cockcroft report answered the electronic calculator by splitting maths into with-calculator and without-calculator halves and examining both, a structure now standard almost everywhere. He proposes the same for AI, in every subject: teach both, test both. It is the only move on this map with a track record. Honest note, and it cuts against the grain of this whole page: the same author writes that AI already provides a level of tailored instruction he and the teachers he has watched have struggled to achieve, and quotes a student saying nobody had ever paid such pure attention to her thinking. That is professional testimony rather than a trial, and it describes the quality of the teaching experience, not what learners retain once the tool is removed, which is where the studies here find the damage. Recorded because citing an author only where he agrees with the map would be the wrong kind of map.
- New section
Founded section 4, “When intelligence is plentiful, volition is valuable.”
David Brooks argues that who thrives in an AI-saturated world depends less on the tool and more on a person’s relationship to mental effort. He builds on need for cognition, the psychology trait named and measured by Cacioppo and Petty (1982), and names three responses to AI, the Productive Passenger, the Reluctant Optimizer and the Mental Marathoner, warning of a coming cognitive polarization as some people use AI to think more and others use it to think less. This founded the beat “It starts with the person,” now placed among the factors that set the outcome.
- Sharpensconceptualpartly on map
Attaches to section 1, “AI can improve performance and reduce learning.”
A peer-reviewed Perspective in Nature Medicine splits “lost learning” by career stage: deskilling (losing abilities you already had), never-skilling (never building foundational reasoning in the first place, because AI substitutes for the effort that would have built it), and mis-skilling (internalising AI’s errors as correct knowledge). It sharpens the section above: AI-assisted performance during training can create what the authors call false proficiency, apparent competence that does not persist when the support is withdrawn. Two honest notes: the authors are explicit that this is a conceptual risk model, not yet a measured effect in medical training, and the evidence they assemble includes studies already on this map.
Ke et al. (2026), Nature Medicine · via Sam Illingworth