
A living map of the research on AI and learning
What Do We Want to Remain Capable Of?
- Does AI make me think less?
- What do we still need to learn in the age of AI?
- What is its environmental cost?
For several months, I have been digging into this subject. Its effects are systemic, but the research remains largely compartmentalised: each study lights up one piece of the puzzle without always showing the whole. So I wanted to go beyond the announcements, set these studies side by side and see what picture emerges when they are read together.
At first, my question was simple: how can we use AI without losing our ability to think?
As I kept reading, this question began to feel too narrow. Thinking for ourselves does not mean thinking alone. We also learn with others, within institutions, communities and environments that shape the way we think. And our use of AI takes place in a world of limited resources.
So the question shifted. It is no longer only about what AI lets us do, but about what we want to remain able to do without it, with it, and with one another.
That is the question running through this overview:
What do we want to remain capable of?
Two caveats before we begin. I am not an education researcher. This text offers a cross-cutting reading of the available evidence, intended to contribute to the discussion.
Moreover, this work began with a review of the field in June 2026. Since then, models, tools and research have continued to evolve. Some conclusions will therefore need to be clarified, qualified or revised.
This overview will be updated regularly. Depending on their significance, new publications will either be added as notes to the relevant sections or lead to new sections. Every change will be recorded in “Changes over time”.
Show the table of contentsHide the table of contents
I · Why effort matters
The problem
AI can improve performance and reduce learning.
In a randomised study of about 1,000 high school maths students in Türkiye, students using a standard ChatGPT scored 48% higher than classmates without it while they had it. When access was taken away, they scored 17% lower than students who never had it. A 2025 paper in Nature Reviews Psychology makes the underlying point: performance gains are not the same as learning.
Better output is not the same as learning.
Why it happens
When the effort of thinking goes, part of the learning goes with it.
Most AI tools were built for work, not to optimise learning. At work, the goal is to finish the task with the least effort. But in learning, that effort is the point: it is what builds the capability. The task gets finished, but the understanding never develops. Researchers call this “metacognitive laziness”: the learner stops planning, monitoring and self-evaluating, because the AI always has an answer.
If the AI carries the effort, it also carries off the learning.
Why effort is the point
The struggle is not an obstacle to learning. It is the learning.
Learning scientists Robert and Elizabeth Bjork call these “desirable difficulties”: things like recalling an answer from memory, spacing your practice out, or trying a problem before you see the solution. They make learning feel harder now, but they make it last. When AI removes that effort, it can quietly remove the learning the effort produced. A 2026 review of 67 studies names the same mechanism, epistemic friction: without it, “AI-generated fluency can bypass the reflective struggle central to deep learning.”
Protect the difficulty that does the teaching.
II · What decides the outcome
It starts with the person
When intelligence is plentiful, volition is valuable.
If the struggle is the learning, the first thing it depends on is the person doing it. Writing in The Atlantic, David Brooks argues that what will set people apart in an AI-saturated world is not how smart they are but their relationship to mental effort. His frame is need for cognition, the trait named and measured by John Cacioppo and colleagues, who reviewed more than 100 studies of it: at one pole, people who read dense books and play hard games for pleasure; at the other, the cognitive miser, who avoids effortful thought where possible. It correlates with intelligence, Brooks notes, but is not the same as it.
The evidence he marshals is largely about what happens when effort is offloaded. An MIT Media Lab team led by Nataliya Kosmyna measured brain connectivity dropping by up to 55% during ChatGPT use; Michael Gerlich (SBS Swiss Business School) found a significant negative correlation between frequent AI use and critical thinking; a Carnegie Mellon team led by Grace Liu found that after about ten minutes of AI-assisted problem solving, people who then lost the tool did worse than those who never had it. A study of endoscopists found precancerous-lesion detection fell from 28.4% to 22.4% once AI was withdrawn. From this he sketches three responses: the Productive Passenger, who lets AI think, gaining productivity and losing capacity; the Reluctant Optimizer, who means to resist and gets pulled in; and the Mental Marathoner, who keeps the hard parts hard. His worry is a cognitive polarisation between the two ends. His practical rule: ask AI for thinkers, not thinking, and treat it as a brilliant librarian, not an oracle.
The first variable is the learner, not the tool.
With what's already in their head
When answers are plentiful, background knowledge is the bottleneck.
Volition is not the only person-side variable. Stephen Fitzpatrick, a history teacher who has worked with students for over thirty years, argues that the most important skill here is reading: AI output is fluent, confident and endless, and reading it with the scepticism it demands takes the vocabulary and background knowledge that only years of reading build. “The bottleneck to effective student AI use is mostly a problem of reading”, not writing, which gets the attention because writing is what we grade. Deep research makes the point sharper: finding sources is no longer the hard part; reading the report, and its sources, carefully is.
Marie Dollé, writing in French about the summer’s “brainmaxxing” apps, takes the same point one level deeper. Against the comforting idea that machines can hold the knowledge while we keep the judgement, she argues that judgement cannot form in a vacuum: “on ne problématise pas sans repères”, you cannot frame a problem without reference points already in your head. And she names a limit for any technology that claims to sort the world for us: we often only know what mattered afterwards, so no tool can pre-sort “the essential” for us. Consulting a fact on demand is not the same as having it in mind.
Daniel Susskind, an economist who has spent fifteen years studying AI and work, gives the same claim a mechanism. AI systems are what Geoffrey Hinton calls “idiot savants”: impressive on hard problems, wrong on easy ones, and nothing in the output tells you which one you are reading. So the basics are not a safe harbour, they are an instrument. Use AI critically rather than blindly, he argues, “keeping our basics sharp”, precisely so you can tell the savant from the idiot. Which is why he wants literacy and numeracy taught intensely even where AI already does them better, and calls it a no-regrets investment: one that pays off whatever the future turns out to be. The backdrop is not reassuring. Since 2009, literacy and numeracy have been falling among young people worldwide, according to the OECD’s PISA programme, and among adults too.
You can only check a machine against what you already hold.
And by what we weren't looking for
Background knowledge helps us judge. It cannot tell us what will matter tomorrow.
Part of what shapes us comes from encounters whose importance we could not have measured in advance: a book, an idea, a person, an unexpected experience.
Marie Dollé sees here a fundamental limit of any technology that would sort the world on our behalf: it can learn what matters to us today, but how could it know what will matter tomorrow, when that also depends on events that have not yet happened and on who we will have become?
Research already documents this risk: writing with a language model makes texts more alike (Doshi & Hauser, 2024), and an assistant can nudge our opinions (Jakesch et al., 2023), or simply confirm what we already thought (Sharma et al., 2023).
Learning, then, is not only choosing better from what we already know. It is also staying open to what can shift our bearings.
Preserve our capacity to be surprised, to wander off course, to come out changed.
Then the design
The same technology can help or harm. The design decides which.
In a Harvard physics study, a purpose-built AI tutor designed with proper scaffolding beat in-person active learning by 0.73 to 1.3 standard deviations, which the authors describe as a large effect. In the Türkiye study, the standard ChatGPT left students worse off once it was removed, while a guardrailed tutor version essentially avoided that loss. The same model could harm learning or spare it, depending on how the tool was designed.
The result is set by the design, not by the model.
And by the human behind it
Believing a human is paying attention changes how hard we try.
In a controlled study in a university creative-coding course, students received identical AI-generated feedback on their work. Those told it came from a human teaching assistant ran their code more, wrote more code, and spent more time on later work. They rated the feedback equally helpful either way. The effect on effort was large (d = 0.88 to 1.56). The content was the same.
Same words land differently when we believe a human wrote them.
III · What people bring
What only a human does
AI can help with the content. It rarely touches the rest.
Education does three things at once: it builds knowledge and skills (qualification), it helps you find your place among others (socialisation), and it helps you become someone who thinks independently and takes responsibility (subjectification). AI tools mostly reach the first. They rarely, if ever, address the other two, and those are where a teacher does their deepest work.
AI can teach the content. A human helps you become someone.
Three ways AI can show up
An LLM, a tutor, and a learning companion are not the same thing.
An LLM answers your question. Faster work, less learning.
An AI tutor asks questions back, no matter what you actually need. Often frustration and drop-out.
An AI learning companion (Dr Philippa Hardman calls it a “study mate”) remembers where you got stuck and pushes you towards the thinking you avoid. Capability that lasts.
Aim for a learning companion, not an answer machine.
What education is for now
AI does not shorten what there is to learn. It lengthens it.
Holmes (UNESCO) asks it directly: if generative AI is this powerful, do we still need to learn? His answer is yes, and the list grows: on top of what we wish to learn, we now need to learn AI’s profound limitations, its impacts on human rights, social justice and the environment, and “perhaps most importantly, [to] learn how to think… critically.” The World Economic Forum (WEF) keeps the balance: rote memorisation may matter less, but “the process of mastering knowledge continues to develop broader capabilities” such as grit, curiosity, communication and critical thinking, and assessment must evolve to capture them.
The emphasis moves from having answers to judging them.
IV · Wider stakes
In a world of finite resources
Every answer has an energy cost. The task decides how large.
Learning with AI also happens somewhere physical: on servers, drawing electricity. A team from Hugging Face and Carnegie Mellon (Luccioni, Jernite & Strubell) measured it task by task: 88 models, a thousand requests per dataset, each run ten times. Sorting a text into categories used about 0.002 kWh per thousand requests. On average, generating text used over ten times more; generating images over sixty times more again. The least efficient image model used around half a smartphone charge for every picture.
The larger finding is about generality. A model built for one well-defined task used far less energy than a general-purpose model doing the same job: in the authors’ tests, from a few times less to about thirty times less. Their overall verdict is stronger still: “multi-purpose, generative architectures are orders of magnitude more expensive than task-specific systems for a variety of tasks, even when controlling for the number of model parameters.” For well-defined tasks such as web search, they see no convincing evidence that general models are needed, given how much energy they use.
So discernment has a physical side. Choosing the right tool for the task, a search, a classifier, a smaller model, or no AI at all, is also a way of taking the environment into account.
Use the smallest tool that does the job.
Ourselves, others, and the planet
Three levels of stakes. At each level, a relationship to attend to.
This map keeps meeting three stakes: what AI does to our thinking, to how we live and learn with others, and to the world it draws on. They are usually studied apart. Learning scientists measure thinking, ethicists debate fairness and power, engineers count energy. The declarations that try to hold them together, from the IEEE’s design guidelines to the Montreal and Toronto declarations, tend to do it around a single aim: human well-being.
In 2019 a group of scholars, artists, technologists and knowledge keepers, most of them Indigenous, from Aotearoa, Australia, North America and the Pacific answered those declarations with a position paper of their own, published with the Canadian Institute for Advanced Research and edited by Jason Edward Lewis (Cherokee, Hawaiian and Samoan), who holds a research chair in computational media at Concordia University in Montreal. Their objection is precise: “none of these efforts challenge the fundamental anthropocentrism of Western science and technology.” Their alternative is relational. Instead of asking only how a technology serves us, they ask what relationships it creates, and what each side owes. In their words, “while the developers might assume they are building a product or tool, they are actually building a relationship to which they should attend.”
Read this way, the three stakes become three relationships. With what we know: the historian Noelani Arista (Kanaka Maoli) asks whether “computer memory” will “replace experts and elders as repositories of knowledge”, and calls for institutions that train knowledge keepers “fluent in language and trained in computer science.” With others: for Caroline Running Wolf and Noelani Arista, research and applied projects need to be built “collaboratively with, not on behalf of, and certainly not without the community”. More broadly, the paper asks communities to build their own values into these technologies, since “the alternative is to have our worlds designed for us.” With the planet: the artist Suzanne Kite (Oglála Lakȟóta) asks “What is being offered to the Earth when we extract these mined materials?”, the question the previous section began to answer in kilowatt hours.
This is a lens, not evidence, and it asks for care. The paper does not claim a unified voice (“Indigenous ways of knowing are rooted in distinct, sovereign territories”), and Suzanne Kite writes that her essay “does not attempt to speak for all Lakota”. The paper does not forbid others from learning from it, but many of its participants “expressed concern about issues of appropriation and misuse of traditional knowledge”, so it is credited here as a way of seeing the question, not borrowed as a method. It was written in January 2020, before ChatGPT, and gives a central place to systems that communities control or build themselves. And as a position paper, it opens a question rather than settling one.
Ask what each relationship owes, and to whom.
V · What to do
So how do we make it help?
The challenge is also systemic. It runs through the tool, the classroom, and the system around them.
The tool: the guardrailed tutor and the learning companion are instructional design built into software: scaffolding, answers withheld, help that fades as you grow.
The classroom: co-design tools with teachers rather than deploying them on teachers, and set tasks that make learners compare, justify and revise what the AI produces. In the 67-study review, that scaffolding is what separated gains from cognitive offloading.
The system: AI only helps where the conditions are ready. As the WEF puts it, “learning outcomes will not be determined by technology itself, but by the conditions in which it is deployed”, and isolated fixes across policy, pedagogy and technology are unlikely to be sufficient.
In August 2026, the system layer got its first institution-scale test case. MIT’s committee on AI in teaching and learning refused patching (“this is not a moment for patches and duct tape”) and called for all three moves in concert: AI-aware redesign of every subject, an in-person social component in every subject, and permanent adaptation structures (a standing committee, per-school AI leads, a pilot fund), on a platform deliberately tied to no single AI vendor. One survey number shows what is at stake: MIT undergraduates report feeling more replaceable than capable (40% versus 34%).
Susskind adds the one move here with a precedent. In 1982 a British government report by Wilfred Cockcroft answered the electronic calculator not by banning it but by splitting mathematics in two: part of the time learning to work with a calculator, the rest learning to cope without one, and, the decisive part, examining both. That split is now the standard almost everywhere. He proposes the same structure for AI, in every subject, and calls it teach both, test both. His reason for putting the weight on assessment rather than detection is the part worth keeping: a teacher cannot know whether a student used AI alone in their bedroom, but nobody forgets sitting an exam they have only half prepared for.
Real progress needs all three moving.
What to take away
- Preserve the difficulties that help us grow.
- Make learning design a priority in the development of AI tools.
- Recognise that human relationships shape us far beyond the transmission of knowledge.
- Choose tools that strengthen our capabilities, rather than simply producing answers for us.
- Cultivate a taste for intellectual effort.
- Also seek, with and without AI, what challenges or shifts our ideas.
Wayne Holmes sums it up: learn AI’s limitations, “take into account its broader impacts on human rights, social justice and the environment, and, perhaps most importantly, learn how to think… critically”. In other words: a cognitive challenge, a societal one, an environmental one.
Three reasons to aim for what MIT describes as the highest-quality education: “of humans, by humans, in support of human flourishing, and for the betterment of humankind.”

What do we want to remain capable of?
About this map
My real ask: what am I missing? Drop the studies, voices, or counter-evidence that would make this map fuller, including anything newer that shifts it. I know a map is not the territory, so I’d love to hear from people on the ground. And it works: several of the first entries in how this map has changed came from exactly this kind of pointer.
Full disclosure, and a small irony: I used AI to build this.
It all started in conversations with colleagues and partners at the Learning Planet Institute. Faced with research that is plentiful but scattered across cognitive, social and planetary stakes, I wanted to try an experiment: could AI help me keep an overview of the subject while letting it evolve as I read?
I then read most of the sources cited (a few only through their summaries) and used AI to sort them, set them side by side and bring out convergences, tensions and the main lines of force. Without that help, it would have been hard for me to build such a broad view of the subject.
The result taught me as much through its content as through the way it was produced. The process remains very time-consuming: rereading, checking, qualifying and rewriting sometimes take as long as working without AI, if not longer. More than once, I also noticed it pushing towards smoother, more clear-cut wording, at the risk of flattening some nuances.
This experiment brings me back, on my own scale, to the question that closes this map: what do I want to remain capable of?
Of building an overview that makes sense, of course. But also of keeping a way of working that sustains the very capacities this map examines: searching, comparing, checking, doubting, rephrasing and building my own reasoning. In other words, using AI without giving up a form of cognitive hygiene. One matters to me as much as the other.
Where this comes from
A map of the research, drawn over time and kept current as new sources arrive.
Download this map as a PDF- Bastani et al. (2025), Generative AI without guardrails can harm learning, PNAS ↩ 1, 7, 14
- Cacioppo & Petty (1982), The need for cognition, Journal of Personality and Social Psychology, 42(1), 116-131 ↩ 4
- Daniel Susskind (2026), I'm a father of three who studies the impact of artificial intelligence: this is what parents need to know about AI, The Guardian ↩ 5, 14
- David Brooks (2026), The People Who Will Thrive in the AI Age, The Atlantic ↩ 4
- Doshi & Hauser (2024), Generative AI enhances individual creativity but reduces the collective diversity of novel content, Science Advances, 10(28), eadn5290 ↩ 6
- Dr Philippa Hardman, From AI Tutors to AI Study Mates ↩ 10
- Fan et al. (2025), Beware of metacognitive laziness, British Journal of Educational Technology ↩ 2
- Gert Biesta, three functions of education ↩ 9
- Jakesch, Bhat, Buschek, Zalmanson & Naaman (2023), Co-Writing with Opinionated Language Models Affects Users' Views, Proceedings of CHI '23 ↩ 6
- Kestin et al. (2025), AI tutoring outperforms in-class active learning, Scientific Reports ↩ 7
- Khosravi et al. (2026), Building AI Companions that Prioritise Learning over Performance ↩ 2, 10, 14
- Lewis, ed. (2020), Indigenous Protocol and Artificial Intelligence Position Paper, Initiative for Indigenous Futures and CIFAR ↩ 13
- Li, Cui & Hagedorn (2026), Computers and Education: AI ↩ 3, 14
- Luccioni, Jernite & Strubell (2024), Power Hungry Processing: Watts Driving the Cost of AI Deployment?, ACM FAccT ↩ 12
- Marie Dollé (2026), La star de l'été…, mariedolle.substack.com ↩ 5, 6
- MIT Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training (2026), Final Report ↩ 14
- Morris & Maes (2026), Same Feedback, Different Source ↩ 8
- OECD, Digital Education Outlook 2026 ↩ 14
- Robert & Elizabeth Bjork, desirable difficulties ↩ 3
- Sharma et al. (2023), Towards Understanding Sycophancy in Language Models, ICLR 2024 ↩ 6
- Stephen Fitzpatrick (2026), It's the Reading, Stupid, Fitzy's History ↩ 5
- Wayne Holmes (2026), Learning to think in the AI era, UNESCO Courier ↩ 9, 11
- World Economic Forum, Shaping the Future of Learning (2026) ↩ 11, 14
- Yan, Greiff, Lodge & Gašević (2025), Nature Reviews Psychology ↩ 1
This is the provenance ledger: every source that has moved the map since it was drawn, in the order it arrived. It never renders in the main reading flow above. It is the only place a date appears on this page. A JSON feed is derived from the same list.
- Sharpensconceptualnew to map
Attaches to section 5, “When answers are plentiful, background knowledge is the bottleneck.”
SharpensBearman, Tai, Dawson, Boud & Ajjawi (2024), Assessment & Evaluation in Higher Education
There is a name for what this section describes, and a long research tradition behind it: evaluative judgement, the capability to judge the quality of work, your own and other people’s. Margaret Bearman and colleagues at Deakin University argue that generative AI makes it urgent, and that building it needs no new tool: the familiar practices of self-assessment, peer review, feedback, rubrics and examples of good work, turned on what AI produces and on how we use it. Honest note: a conceptual paper; the measurement this section asks for is still missing.
Bearman, Tai, Dawson, Boud & Ajjawi (2024), Assessment & Evaluation in Higher Education
- New sectionmeasurednew to map
Founded section 12, “Every answer has an energy cost. The task decides how large.”
The closing names an environmental stake; this section gives it evidence. Luccioni, Jernite and Strubell measured the energy of 88 models task by task: a spread of more than 1,450 times between the cheapest and costliest tasks, and general-purpose models far costlier than task-specific ones on the same job. Honest note: open models available in 2023, the largest general-purpose one at 11 billion parameters, measured on one cloud region; energy and carbon only, with no water or manufacturing; independent measurements of today’s much larger tools remain scarce.
- New sectionconceptualnew to map
Founded section 13, “Three levels of stakes. At each level, a relationship to attend to.”
The closing names three stakes, cognitive, societal and environmental; this section holds them together through the Indigenous Protocol and Artificial Intelligence position paper, a reply by scholars, artists, technologists and knowledge keepers, most of them Indigenous, to AI ethics declarations that aim only at human well-being. Honest note: a position paper, not a study; written in January 2020, before ChatGPT, giving a central place to systems that communities control or build themselves; it does not claim a unified voice, so it is credited as a lens, not borrowed as a method.
- Conclusion revised
Closing: the three stakes (cognitive, societal, environmental) are now explicitly grounded in Wayne Holmes (UNESCO Courier).
Wayne Holmes (2026), Learning to think in the AI era, UNESCO Courier
- Conclusion revised
The closing question shifts the problem from assisted performance to durable capability: one question that holds at the individual, collective and societal scale. The MIT quote is restored in full (“and for the betterment of humankind”).
- New section
On serendipity and curation: background knowledge lets us judge but cannot pre-sort what will matter tomorrow. Preserving our capacity to be surprised, displaced and transformed becomes part of the answer. Via Marie Dollé.
Marie Dollé (2026), La star de l'été…, mariedolle.substack.com
- Sharpensmeasurednew to map
Sources added: the serendipity section is now grounded in research on homogenisation (Doshi & Hauser), opinion influence (Jakesch et al.) and sycophancy (Sharma et al.).
The closing takeaway now names the mechanism: seek, with and without AI, what challenges or shifts our ideas.
Doshi & Hauser (2024), Science Advances; Jakesch et al. (2023), CHI ’23; Sharma et al. (2023), ICLR 2024
- New section
Founded section 5, “When answers are plentiful, background knowledge is the bottleneck.”
Stephen Fitzpatrick brings reading as the bottleneck; Marie Dollé brings the repères argument (judgement runs on reference points already in the head, and no tool can pre-sort what will matter); Daniel Susskind supplies the mechanism, that AI’s confident errors are undetectable from the output, so the basics are the instrument you check the machine with, plus the OECD finding that literacy and numeracy have been falling worldwide since 2009. Honest note: two practitioner essays and a newspaper essay by an economist, not studies. Their shared claim is exactly what someone should now measure.
Stephen Fitzpatrick (2026), It's the Reading, Stupid, Fitzy's History; Marie Dollé (2026), La star de l'été…, mariedolle.substack.com; Daniel Susskind (2026), The Guardian
- Sharpens
Extends the final section. MIT’s AI committee report (August 2026) gives the system layer its first institution-scale case: all three moves called for in concert, on a vendor-neutral platform. Honest note: a governance report, not an outcome study.
- Sharpens
Extends the final section. Daniel Susskind revives a forgotten precedent: the 1982 Cockcroft report answered the electronic calculator by splitting maths into with-calculator and without-calculator halves and examining both, a structure now standard almost everywhere. He proposes the same for AI, in every subject: teach both, test both. It is the only move on this map with a track record. Honest note, and it cuts against the grain of this whole page: the same author writes that AI already provides a level of tailored instruction he and the teachers he has watched have struggled to achieve, and quotes a student, reported in The New Yorker, saying nobody had ever paid such pure attention to their thinking. That is professional testimony rather than a trial, and it describes the quality of the teaching experience, not what learners retain once the tool is removed, which is where the studies here find the damage. Recorded because citing an author only where he agrees with the map would be the wrong kind of map.
- New section
Founded section 4, “When intelligence is plentiful, volition is valuable.”
David Brooks argues that who thrives in an AI-saturated world depends less on the tool and more on a person’s relationship to mental effort. He builds on need for cognition, the psychology trait named and measured by Cacioppo and Petty (1982), and names three responses to AI, the Productive Passenger, the Reluctant Optimizer and the Mental Marathoner, warning of a coming cognitive polarisation as some people use AI to think more and others use it to think less. This founded the beat “It starts with the person,” now placed among the factors that set the outcome.
- Sharpensconceptualpartly on map
Attaches to section 1, “AI can improve performance and reduce learning.”
SharpensKe et al. (2026), Nature Medicine
A peer-reviewed Perspective in Nature Medicine splits “lost learning” by career stage: deskilling (losing abilities you already had), never-skilling (never building foundational reasoning in the first place, because AI substitutes for the effort that would have built it), and mis-skilling (internalising AI’s errors as correct knowledge). It sharpens the section above: AI-assisted performance during training can create what the authors call false proficiency, apparent competence that does not persist when the support is withdrawn. Two honest notes: the authors are explicit that this is a conceptual risk model, not yet a measured effect in medical training, and the evidence they assemble includes studies already on this map.
Ke et al. (2026), Nature Medicine · via Sam Illingworth