INTELLIGENCE REPORT SERIES SEPTEMBER 2026 OPEN ACCESS

SERIES: SCIENCE & TECHNOLOGY

Cognitive Offloading (2026) — Skill Fell 21% Without AI

After three months of AI-assisted colonoscopy, 19 experienced doctors detected 21% fewer adenomas without it. Forty years of offloading evidence, audited.

Reading Time42 min
Word Count8,208
Published2 September 2026
Evidence Tier Key → ✓ Established Fact ◈ Strong Evidence ⚖ Contested ✕ Misinformation ? Unknown
Contents
42 MIN READ
EN FR ES PT DE RU AR JP ZH

After three months of AI-assisted colonoscopy, 19 experienced doctors detected 21% fewer adenomas without it. Forty years of offloading evidence, audited.

01

The Withdrawal Test
Nineteen experienced doctors, three months of assistance, and a six-point hole where a skill had been

Between September 2021 and March 2022, four Polish endoscopy centres introduced artificial intelligence into routine colonoscopy. When the same 19 doctors later worked without it, their adenoma detection rate had fallen from 28.4% to 22.4%✓ Established [1] — a relative decline of 21% in the defining diagnostic skill of the specialty✓ Established [2]. Nothing had been removed from them except practice. The question this report audits is when that trade happens, when it does not, and what forty years of controlled studies on calculators, pens and satellite navigation already establish about the difference.

The evidence has a specific answer, and it is older than the current argument. Offloading a cognitive task to a machine degrades the underlying capacity when the machine removes the retrieval or generation step the person would otherwise perform, and leaves that capacity intact — sometimes improving it — when the machine removes friction around a step the person still performs◈ Strong Evidence [12]. The Polish colonoscopy study is the cleanest natural experiment yet published on the first case. Nineteen endoscopists, each with more than 2,000 procedures behind them, were exposed to AI-assisted detection over three months at four centres taking part in a cancer-prevention trial; their unassisted detection rate then fell by six percentage points✓ Established [1].

The population running this experiment on itself is now most of a generation. The third annual survey of United Kingdom undergraduates, published in March 2026 from 1,054 full-time responses, found 95% using generative AI in at least one way and 94% using it for assessed work✓ Established [20]. The proportion pasting AI-generated text directly into submitted work reached 12%, against 8% in 2025 and 3% in 2024✓ Established [20]. Institutional provision tripled over the same two years, from 9% of students reporting university-supplied tools in 2024 to 38% in 2026✓ Established [20]. Adoption is no longer a student behaviour to be policed; it is infrastructure being purchased.

Behavioural data at scale shows what the purchase changed. A ten-year panel of 3.2 million interactions on the ALEKS adaptive mathematics platform compared topics that transcribe easily into a chatbot prompt against graph-based topics that do not. Study time on the AI-susceptible topics fell 2.8% per quarter among university students after the release of ChatGPT, a cumulative 26.9% over eleven quarters, with high-school students at 31.3% and fifth-graders showing no detectable change◈ Strong Evidence [21]. On randomly assigned, proctored retention items the odds of a correct answer fell by a cumulative 25%◈ Strong Evidence [21]. Faster completion, less durable knowledge — measured on the same students, in the same curriculum.

22.4%
Unassisted adenoma detection after three months of AI exposure, down from 28.4%
The Lancet Gastroenterology and Hepatology, 2025 · ✓ Established
95%
United Kingdom undergraduates using generative AI in at least one way
HEPI and Kortext Survey, 2026 · ✓ Established
26.9%
Cumulative fall in university study time on AI-susceptible mathematics topics
ALEKS panel, eleven quarters, 2026 · ◈ Strong Evidence
666
Participants in the first large survey linking AI use to weaker critical thinking
Societies 15:6, 2025 · ✓ Established

The aggregate skill series were already falling before any of this. The OECD Survey of Adult Skills, which tested about 160,000 people aged 16 to 65 across 31 countries and reported in December 2024, found literacy proficiency improved over the preceding decade in only two countries, Finland and Denmark, and was stable or declining everywhere else✓ Established [18]. Numeracy did better, improving in eight countries✓ Established [18]. The declines were largest among the least educated, widening the gap within each country✓ Established [19]. School-age data points the same way: mean OECD mathematics performance fell about 15 points between 2018 and 2022 and reading about 10✓ Established [34].

None of that is evidence about artificial intelligence, and treating it as such would be the first error available. The adult skills decline predates ChatGPT by a decade; the reversal of the Flynn effect measured in roughly 394,000 United States adults runs from 2006 to 2018, before large language models existed, and it did not touch every domain — verbal, matrix and letter-number reasoning fell while three-dimensional rotation rose✓ Established [22][33]. What these series establish is a baseline, not a cause. They describe a population whose measured cognitive performance was already drifting when a general-purpose offloading device arrived.

✓ Established Three months of routine AI assistance measurably reduced experienced endoscopists' unassisted detection rate

The design is observational but unusually clean: the same 19 operators, the same four centres, the same procedure, before and after. Adenoma detection at colonoscopies performed without AI fell from 28.4% — 226 of 795 procedures — to 22.4%, or 145 of 648✓ Established [1]. AI-assisted procedures in the same period ran at 25.3%, 186 of 734✓ Established [2]. Every operator had performed more than 2,000 colonoscopies before the study began, which removes the obvious confound of inexperience✓ Established [1]. The result was published in The Lancet Gastroenterology and Hepatology on 12 August 2025✓ Established [2].

This report proceeds in the order the evidence was produced. Section two sets out why minds delegate at all, and why the judgement governing that delegation is systematically unreliable. Section three covers navigation, the one domain where the structural cost has been imaged directly. Section four returns to the controlled studies on calculators and note-taking that already answered the general question. Sections five and six audit the large language model literature and the professional deskilling record. Section seven compares what five governments have actually done. Section eight states what the evidence will and will not carry.

02

Why Minds Delegate
Offloading is a rational strategy resting on a self-assessment that is frequently wrong

Cognitive offloading is not a modern pathology. It is the ordinary use of physical action to reduce internal cognitive demand — tilting the head to read a rotated image, writing an appointment down, setting an alarm✓ Established [12]. The decision to offload is governed by a metacognitive judgement about one's own capacity, and that judgement is measurably fallible in both directions◈ Strong Evidence [12]. The consequential question is not whether people offload. It is what the act of offloading does to the ability being assessed.

The framework that organises this literature was set out by Evan Risko and Sam Gilbert in Trends in Cognitive Sciences in 2016, and it has three moving parts. The propensity to offload is influenced by the internal cognitive demand the task would otherwise impose and by a metacognitive evaluation of one's own ability; the act of offloading then feeds back into that evaluation; and offloading has direct effects on cognitive capability◈ Strong Evidence [12]. The middle term is where the trouble sits. A person who reliably delegates a task stops receiving the evidence about their own competence that would tell them whether delegation is still warranted.

The most cited empirical anchor is also the most damaged. In 2011 Betsy Sparrow, Jenny Liu and Daniel Wegner reported four studies indicating that when people expect continued access to information, they recall the information itself less well and recall where to find it better✓ Established [11]. The finding entered general usage as the Google effect, or digital amnesia. It also entered the replication crisis: high-powered replications failed to reproduce the most-cited component, the computer-priming Stroop experiment, and later synthesis suggests the overall effect is real but smaller and more context-dependent than the original paper implied⚖ Contested [11].

That correction matters more than it is usually allowed to. The strong form of the Google effect — that search engines have hollowed out semantic memory — is not established, and this report does not assert it. The defensible form is narrower and older: memory is transactive. People have always distributed what they know across other people, and the internet extended an existing mechanism to a non-human partner✓ Established [11]. A transactive system is efficient exactly as long as the partner is available and accurate. Its failure mode is not forgetting. It is the inability to detect that the partner has gone wrong.

They will cease to exercise memory because they rely on that which is written, calling things to remembrance no longer from within themselves, but by means of external marks.

— Socrates, in Plato's Phaedrus, c. 370 BCE

The objection that this argument is 2,400 years old and has been wrong every time is worth taking seriously, because it is half right. Writing did externalise memory, and Socrates was correct about the mechanism while wrong about the consequence✓ Established [32]. Literate societies did not become less capable; they became capable of different and larger things, because writing did not remove a step from thinking so much as add a permanent substrate to it. The historical record therefore supports neither panic nor complacency. It supports a question about which specific step a given tool removes.

That question can be made concrete. A calculator removes the execution of an algorithm the student has already been taught to select. A satellite navigation device removes the construction of the spatial representation, not merely its execution. A spellchecker removes the retrieval of an orthographic form after the word has been chosen. A large language model, asked for an essay, removes the selection of the argument, the retrieval of the evidence, the ordering of the claims and the generation of the sentences — every step except the prompt and the judgement of the output. These are not the same intervention, and there is no reason to expect them to produce the same result.

The metacognitive layer is what makes the difference hard to notice from the inside. The Microsoft Research and Carnegie Mellon survey of 319 knowledge workers, reporting 936 real-world uses at the 2025 Conference on Human Factors in Computing Systems, found that higher confidence in the AI system was associated with less critical thinking applied to its output, while higher confidence in one's own skill was associated with more✓ Established [6]. Competence and the perception of competence come apart under assistance. The person best placed to notice a degraded skill is the person whose degraded skill is doing the noticing.

03

The Navigation Evidence
One domain where the structural cost has been imaged, measured longitudinally, and shown to reverse

Spatial navigation is the only cognitive function for which the offloading argument has direct neuroanatomical evidence on both sides of the ledger. Building a cognitive map enlarges the posterior hippocampus✓ Established [9]; ceasing to build one is associated with a steeper decline in hippocampal-dependent spatial memory over three years◈ Strong Evidence [8]. The London taxi trade supplied the first half of that finding by accident, and satellite navigation supplied the second.

The Knowledge, the licensing examination for London taxi drivers, requires the memorisation of roughly 25,000 streets and the ability to construct a route between any two points without aid. In 2000 Eleanor Maguire and colleagues reported in the Proceedings of the National Academy of Sciences that the posterior hippocampi of licensed London taxi drivers were significantly larger than those of controls, and that hippocampal volume correlated with time spent driving — positively in the posterior region, negatively in the anterior✓ Established [9]. Structure followed use, in adults, in a brain region long assumed to be fixed after development.

The follow-up work is the part that carries the argument. The enlargement is not permanent and not free. Reviews of the taxi-driver literature record that the posterior hippocampal advantage recedes with disuse after drivers retire, and that elderly drivers still working retained an advantage over those who had stopped◈ Strong Evidence [10]. The trade-off is equally documented: the same drivers performed less well than controls on some tasks involving the acquisition of new visuospatial information◈ Strong Evidence [10]. Expertise of this kind is a lease, not a purchase, and the rent is paid in continued retrieval.

The obvious modern question — what happens when the retrieval stops — was answered by Louisa Dahmani and Veronique Bohbot in Scientific Reports in April 2020. They assessed lifetime satellite navigation experience in 50 regular drivers alongside several facets of spatial memory: strategy use, cognitive mapping and landmark encoding. Greater lifetime use was associated with worse spatial memory during self-guided navigation, that is, when the device was unavailable✓ Established [8]. The cross-sectional result on its own is ambiguous, and the authors knew it.

✓ Established Longitudinal data indicate that satellite navigation use precedes spatial memory decline rather than following it

Thirteen of the original participants were retested three years after the first session. Greater device use over the intervening period was associated with a steeper decline in hippocampal-dependent spatial memory◈ Strong Evidence [8]. The reverse-causation account — that people with a poor sense of direction adopt the device more heavily — was tested and not supported: heavier users did not report adopting it because they navigated badly◈ Strong Evidence [8]. The direction of the arrow is the whole finding. Without it, the cross-sectional correlation would carry no more weight than the observation that people who use dictionaries spell less confidently.

The effect is not confined to a laboratory task. Landmark-dependent navigation strategy, assessed across more than 37,000 participants, declines across the human lifespan, and the balance between landmark-based and map-based strategy is exactly what turn-by-turn guidance removes✓ Established [10]. A driver following spoken instructions is executing a sequence, not maintaining a representation. The behaviour looks identical from outside the car. Inside it, one of the two operations that built the taxi drivers' hippocampi is no longer being performed.

Two cautions belong here. The satellite navigation sample is small — 50 participants cross-sectionally, 13 longitudinally — and a three-year follow-up of thirteen people is a signal, not a settlement⚖ Contested [8]. And no study has shown that reduced spatial memory in this population produces any functional harm in ordinary life, because ordinary life now includes the device. The finding is that the internal capacity atrophies, not that the person navigates worse. Those are different claims, and the second one has not been demonstrated.

What navigation contributes to the general argument is a mechanism with a measured substrate. A tool that substitutes for the construction of an internal representation is associated with the decay of that representation and of the tissue that maintains it; the decay tracks the quantity of substitution; and it reverses in the other direction when the representation is rebuilt✓ Established [9][10]. That is a considerably stronger claim than anything the general offloading literature can currently support, and it is available in exactly one domain. The rest of this report is about how far it generalises.

04

What the Controlled Studies Already Settled
Calculators, the pen-versus-laptop war and the desirable difficulties literature answered the general question before it was asked

The AI debate is being conducted as though there were no prior evidence on tools that perform part of a cognitive task. There is a great deal, it is unusually well controlled, and it converges: the harm is not in the tool but in whether the tool removes the step at which learning happens◈ Strong Evidence [13][17]. A meta-analysis of 79 calculator studies published in 1986 found improvement at every grade level except one — and the exception is the most informative result in the set✓ Established [13].

Ray Hembree and Donald Dessart integrated 79 research reports on hand-held calculators for the Journal for Research in Mathematics Education in 1986. At every grade except the fourth, calculator use alongside conventional instruction improved the average student's pencil-and-paper skills, in both exercises and problem-solving✓ Established [13]. Attitude and mathematical self-concept improved across all grades and ability levels✓ Established [13]. Later work on graphing calculators reached compatible conclusions, with effects strongest where the device was integrated into instruction for exploration and conceptual work rather than issued for drill and checking◈ Strong Evidence [14].

The fourth grade is the exception that specifies the rule. Sustained calculator use at that level appeared to hinder the development of basic skills in average students✓ Established [13]. Fourth grade is where multi-digit arithmetic procedures are being consolidated — where the retrieval and execution are themselves the object of learning rather than the overhead surrounding it. Deploy the tool before the procedure is automatic and the tool occupies the slot in which automaticity would have formed. Deploy it afterwards and it frees capacity for the next thing. Same device, opposite sign, and the variable is the developmental position of the learner.

The note-taking literature shows the opposite failure mode, and it is a caution against the argument this report is making. The widely-cited 2014 finding that longhand note-taking beats laptop note-taking on conceptual questions — because the slower medium forces generative processing — did not survive direct replication. Kayla Morehead, John Dunlosky and Katherine Rawson found no consistent differences between longhand, laptop, electronic writer and, notably, a group that took no notes at all⚖ Contested [15]. A large direct replication published in Psychological Science reached the same conclusion, with mini meta-analyses returning small and non-significant effects favouring longhand⚖ Contested [16].

c. 370 BCE
The argument is stated for the first time — In the Phaedrus, Socrates warns that writing will produce forgetfulness because learners will rely on external marks rather than internal recollection✓ Established [32]. He was right about the mechanism and wrong about the outcome.
1986
The calculator question is settled by meta-analysis — Hembree and Dessart integrate 79 studies: calculator use improves pencil-and-paper skill at every grade except the fourth, where sustained use hinders basic skill development✓ Established [13].
2000
Navigation expertise is shown to reshape adult brain structure — Maguire and colleagues report enlarged posterior hippocampi in licensed London taxi drivers, with volume correlating with years of driving✓ Established [9].
2006
The integration condition is identified — Meta-analysis of graphing calculators finds gains concentrated where the device is built into instruction for conceptual exploration rather than issued for computation and checking◈ Strong Evidence [14].
2011
The Google effect is named — Sparrow, Liu and Wegner report that expected future access to information reduces recall of the information and improves recall of where to find it✓ Established [11].
2016
Offloading gets a framework — Risko and Gilbert formalise cognitive offloading as a metacognitively governed strategy whose evaluations are themselves affected by the offloading◈ Strong Evidence [12].
2019
The pen-versus-laptop result fails to replicate — Morehead, Dunlosky and Rawson find no consistent advantage for longhand over laptop notes, and no reliable advantage over taking no notes at all⚖ Contested [15].
2020
Satellite navigation is linked to spatial memory decline — Dahmani and Bohbot report worse self-guided spatial memory in heavy device users, and a steeper three-year decline among those who used it more◈ Strong Evidence [8].
2021
A large direct replication confirms the null — A multi-laboratory replication in Psychological Science finds small, non-significant effects favouring longhand, and advises against discarding the laptop⚖ Contested [16].
2024
The adult skills baseline is published — The OECD reports that literacy improved in only Finland and Denmark over a decade across 31 countries, with the sharpest declines among the least educated✓ Established [18][19].

Two nulls and one strong positive is not a contradiction; it is a specification. What the calculator result and the note-taking result have in common is that neither tool removed the generative step in the cases where it did no harm. A calculator issued after the arithmetic procedure is automatic leaves the selection of the procedure with the student. A laptop used for notes still requires the student to decide what is worth recording. Where the tool did harm — fourth-grade arithmetic — it occupied the exact operation the curriculum existed to install.

The learning-science literature names the missing variable directly. Robert and Elizabeth Bjork's desirable difficulties framework holds that conditions which depress performance during learning — spacing, interleaving, self-generation, retrieval practice — improve long-term retention and transfer◈ Strong Evidence [17]. The act of retrieving information is a more powerful event, in its effect on later accessibility, than the act of restudying it◈ Strong Evidence [17]. Any tool that substitutes for retrieval is therefore removing the operation that builds durable knowledge, whatever it does to performance on the day.

The Fourth-Grade Rule

The single most useful number in the entire offloading literature is the exception in a forty-year-old meta-analysis. Calculators helped at every grade except the one where the procedure being automated was itself the thing being learned✓ Established [13]. That is not a fact about calculators. It is a rule for tool deployment: assistance is safe on operations the user has already internalised and corrosive on operations they have not. Every AI-in-education policy currently being written is, whether or not it knows it, a bet on where each student sits relative to that line.

The desirable difficulties framework also comes with its own limit, and honesty requires stating it. The effects are not universal: they attenuate or reverse when working memory is already loaded by high element-interactivity material⚖ Contested [17]. Difficulty is desirable only where the learner has the capacity to meet it. That caveat cuts both ways in the present argument. It weakens any blanket claim that friction is good, and it strengthens the narrower claim that the value of a tool depends entirely on which step it takes and from whom.

05

The Language Model Evidence
Four independent designs, three years of data, and one consistent direction of travel

The literature on large language models and cognition is barely three years old and already contains four methodologically unrelated studies pointing the same way: electroencephalography in a laboratory, a survey of 666 adults, a survey of 319 knowledge workers, and a controlled learning experiment on essay revision◈ Strong Evidence [3][5][6][7]. None is decisive alone. What makes the set worth reading is that each fails in a different direction, and they agree anyway.

The most publicised is also the weakest as evidence and the most useful as a demonstration. A team at the MIT Media Lab divided 54 participants drawn from five Boston-area universities into three groups — one writing essays with a large language model, one with a conventional search engine, one unaided — and recorded brain activity throughout with a 32-electrode array sampling at 500 hertz✓ Established [3]. Participants had 20 minutes and chose from philosophical prompts taken from standardised admission tests. The unaided group showed the strongest and most distributed neural connectivity, the search group intermediate, the model group the weakest✓ Established [3].

The finding that matters in that study is not the connectivity measure, which is difficult to interpret and was not peer-reviewed at release⚖ Contested [4]. It is the behavioural residue. Participants who wrote with the model recalled their own essays less well and reported a weaker sense of ownership over them✓ Established [3]. The authors named the pattern cognitive debt: shallow encoding at the time of production, with the deficit becoming visible only when the text has to be retrieved or defended. That is precisely the signature the desirable difficulties literature predicts for a tool that removes the generative step◈ Strong Evidence [17].

12%
United Kingdom undergraduates pasting AI-generated text directly into assessed work
HEPI and Kortext Survey, 2026 · ✓ Established
25%
Cumulative fall in the odds of a correct answer on proctored retention items
ALEKS behavioural panel, 2026 · ◈ Strong Evidence
39.8%
Student AI prompts falling in the highest category of Bloom's taxonomy, Create
Anthropic Education Report, 2025 · ✓ Established
31.3%
Fall in high-school study time on AI-susceptible topics since 2022
ALEKS behavioural panel, 2026 · ◈ Strong Evidence

The survey evidence is larger and blunter. Michael Gerlich's mixed-method study of 666 participants, published in Societies in January 2025, found a significant negative association between frequent AI tool use and critical thinking scores, statistically mediated by cognitive offloading✓ Established [5]. Younger participants showed higher dependence and lower scores; higher educational attainment buffered the association✓ Established [5]. This is correlational, self-selected and self-reported, and it cannot separate the possibility that people with weaker critical thinking adopt the tools more heavily⚖ Contested [5]. It is a hypothesis with 666 data points attached, not a demonstration.

The controlled learning experiment is the one that isolates the variable. Yizhou Fan and colleagues, writing in the British Journal of Educational Technology in 2025, compared learners supported by ChatGPT, by human experts, by a writing analytics tool, and by nothing. The ChatGPT group produced the largest improvement in essay scores. Their knowledge gain and knowledge transfer were not significantly better than the other groups✓ Established [7]. The authors called the mechanism metacognitive laziness: learners followed the model's feedback to complete the task efficiently rather than engaging in the evaluation and monitoring that produce transferable understanding◈ Strong Evidence [7].

◈ Strong Evidence Generative assistance reliably improves the artefact while leaving the learner's transferable knowledge unchanged

This is the most replicated result in the young literature and the one with the clearest policy consequence. Essay scores rose in the ChatGPT condition; knowledge gain and transfer did not differ significantly from the unsupported condition✓ Established [7]. Completion time on AI-susceptible mathematics topics fell by a cumulative 26.9% among university students, while the odds of a correct proctored retention answer fell 25%◈ Strong Evidence [21]. Any assessment regime that measures the artefact rather than the retained capacity will therefore record improvement in precisely the population that is learning less◈ Strong Evidence [7][21].

What students actually delegate is the part of the work the curriculum values most. Anthropic's analysis of 574,740 academic conversations over an 18-day window, classified against Bloom's taxonomy, found 39.8% of prompts in the highest category, Create, and 30.2% in Analyze, against 10.9% for Apply and 10.0% for Understand✓ Established [23]. The delegation is inverted relative to the usual assumption: the routine is retained and the higher-order work is exported. The sample skewed heavily towards early adopters and computer science, which made up 38.6% of users against 5.4% of the United States student population⚖ Contested [23].

The Assessment Is Measuring the Wrong Object

Institutions are responding to generative AI by redesigning assessment, and 65% of United Kingdom students report that assessment has already changed significantly✓ Established [20]. Almost all of that redesign is aimed at detecting or preventing AI-produced text. The evidence says the exposure is elsewhere: the tool improves the submitted artefact and does not improve retained knowledge✓ Established [7][21]. An institution that successfully authenticates authorship, and continues to grade the artefact, will have solved an integrity problem while leaving the learning problem exactly where it was.

The commercial trajectory sets the timescale for any policy response. The market for artificial intelligence in education is estimated at about 10.6 billion dollars in 2026, up from roughly 7.5 billion in 2025, and is forecast to reach about 42.5 billion by 2030◈ Strong Evidence [35]. Institutional provision of these tools to United Kingdom students rose from 9% to 38% in two years✓ Established [20]. Procurement is running several years ahead of the evaluation literature, which is a familiar pattern in education technology and has not previously ended well⚖ Contested [26].

06

Deskilling Outside the Classroom
Medicine and aviation ran this experiment first, with instrumentation and consequences

The strongest evidence that machine assistance degrades expert skill does not come from education at all. It comes from two professions that instrument their own performance continuously: gastroenterology, where detection rates are audited procedure by procedure, and commercial aviation, where the loss of manual handling skill has been a named safety concern for two decades✓ Established [1][31]. In both, the degradation was found in practitioners who were already expert.

The Polish study is worth restating precisely because its design is so unglamorous. It was observational, not randomised. Nineteen endoscopists at four centres, each with more than 2,000 prior colonoscopies, were tracked across the introduction of AI-assisted detection at their institutions between September 2021 and March 2022✓ Established [1]. Adenoma detection at procedures performed without assistance fell from 28.4% to 22.4%; assisted procedures ran at 25.3%✓ Established [2]. The comparison is within-operator and within-centre, which removes most of the alternative explanations that usually swallow a result of this size.

The mechanism proposed is attentional rather than mnemonic. An operator working alongside a system that flags candidate lesions gradually reallocates visual search from exhaustive scanning to confirmation of the system's output. The skill that decays is not knowledge of what an adenoma looks like; it is the sustained, self-directed search that finds one nobody has pointed at. That is the same structural change the aviation literature has been describing since the 1990s: the transition from operator to monitor, and the specific vigilance costs that come with it◈ Strong Evidence [31].

Our results are concerning given the adoption of AI in medicine is rapidly spreading. We urgently need more research.

— Marcin Romanczyk, Academy of Silesia, in The Lancet Gastroenterology and Hepatology, August 2025

Aviation is the longest-running natural experiment available. Manual flying skills decay towards the edges of tolerable performance without relatively frequent practice, with airspeed control among the first capacities to degrade◈ Strong Evidence [31]. The Federal Aviation Administration's position, arrived at after a series of loss-of-control investigations, is that automation requires more training rather than less, that manual handling skill remains paramount for safety, and that pilots should hand-fly for substantial portions of revenue flights in suitable conditions✓ Established [31]. That is a regulator instructing professionals to deliberately forgo an available tool in order to preserve the capacity it replaces.

ExposureSeverityAssessment
Erosion of unassisted expert performance
Critical
Demonstrated in the one profession that audits itself procedure by procedure: a six-point fall in unassisted adenoma detection after three months of exposure, among operators with more than 2,000 procedures each✓ Established [1]. The exposure is greatest wherever assistance is continuous, performance is not separately measured without it, and the skill involves sustained self-directed search◈ Strong Evidence [31].
Loss of manual capability in automated systems
High
Manual flying skill decays rapidly without frequent practice, and insufficient manual experience has been a contributing factor in loss-of-control investigations◈ Strong Evidence [31]. The regulatory remedy — deliberate hand-flying during normal operations — is an explicit admission that the tool must be periodically declined for the capacity to survive✓ Established [31].
Practitioners who never acquire the baseline
High
Every deskilling result to date concerns experts who had the skill first. Nobody has yet measured a cohort that trained with continuous assistance from the outset⚖ Contested [1]. The fourth-grade calculator exception is the closest available analogue, and it points the wrong way: sustained assistance during the formation of a procedure impeded the procedure✓ Established [13].
Confidence decoupled from competence
Medium
Higher confidence in the AI system predicts less critical scrutiny of its output, while higher confidence in one's own skill predicts more✓ Established [6]. Because offloading also degrades the metacognitive signal used to judge one's own ability, the population least able to detect the deficit is the population most exposed to it◈ Strong Evidence [12].
Institutional measurement blindness
Medium
Almost no institution measures unassisted performance once assistance is deployed. The Polish result exists only because unassisted colonoscopies continued to be performed and recorded within a trial✓ Established [1]. Where the counterfactual is not collected, degradation is not merely undetected — it is undetectable◈ Strong Evidence [21].

The generalisation has a hard limit, and it should be stated before anyone extends it. Every deskilling result so far concerns practitioners who acquired the skill before the tool arrived. There is no published measurement of a cohort trained under continuous assistance from the beginning, in any profession⚖ Contested [1]. It is possible that such a cohort develops a different and equally effective competence organised around the tool. It is also possible that they never acquire the baseline against which the tool's errors are detectable. Both are hypotheses. Neither has been tested, and the second is the one that would matter.

The asymmetry of consequences is what makes this worth regulating in advance of the evidence. In education, a degraded capacity produces a graduate who performs adequately with tools and poorly without them, in a world that mostly supplies the tools. In medicine and aviation, the tool's failure and the human's degraded backup capacity are correlated events: the operator is called upon precisely when the system has stopped working◈ Strong Evidence [31]. Redundancy that decays in proportion to its disuse is not redundancy. It is a deferred single point of failure.

The Counterfactual Is Never Collected

The Polish finding is visible only because a research protocol required colonoscopies to continue being performed without assistance while assistance was available✓ Established [1]. Almost no deployment outside a trial preserves that comparison. Hospitals do not periodically withdraw diagnostic support to audit unaided performance; schools do not assess without devices to measure what the devices are holding up; firms do not benchmark staff against their own pre-deployment output. The single cheapest intervention available in every one of these settings is to keep measuring the unassisted case◈ Strong Evidence [21].

What medicine and aviation contribute is proof of principle at professional scale. Skill decays when the operation that maintained it is performed by something else, the decay is measurable within months rather than decades, and it occurs in populations whose expertise is not in doubt✓ Established [1][31]. Neither field concluded that the tool should be abandoned — assisted colonoscopy detects more adenomas than unassisted, which is why it was adopted✓ Established [2]. Both concluded that the human capacity has to be separately maintained, deliberately, against the grain of the tool's convenience.

07

Five Governments, Five Opposite Bets
China mandates, Estonia distributes, South Korea reverses, and almost nobody has designed an evaluation

National responses to machine-assisted cognition have diverged more sharply than the underlying evidence justifies, because the evidence arrived after most of the decisions. China has made artificial intelligence a compulsory school subject from the age of six while barring primary pupils from using generative systems on their own✓ Established [27]. Estonia is issuing frontier models to every upper-secondary student✓ Established [28]. South Korea legislated its flagship AI textbook programme out of existence in August 2025✓ Established [26].

China's Ministry of Education issued two guidelines in May 2025 that convert AI instruction into compulsory infrastructure. The curriculum begins at age six with a floor of at least eight hours a year of age-appropriate material, progresses through data and coding in the fourth grade to algorithms and intelligent agents by the fifth, and reaches machine-learning logic and misinformation detection in junior high✓ Established [27]. Beijing made the subject mandatory for every elementary and middle-school pupil✓ Established [27]. The same guidelines prohibit primary-school children from using generative systems independently and bar teachers from substituting them for core instruction✓ Established [27].

That combination is the most internally consistent policy any government has adopted. It treats the technology as an object of study rather than a substitute for study, and it places the prohibition exactly where the fourth-grade calculator evidence would place it — during the formation of the procedures being automated◈ Strong Evidence [13]. Whether the eight-hour floor is meaningful and whether the prohibition survives contact with 190 million pupils are separate questions, and neither has been evaluated. What is notable is that the structure of the policy matches the structure of the evidence, apparently by reasoning rather than by accident.

Estonia has taken the opposite bet with equal deliberateness. The AI Leap programme, launched on 1 September 2025, provides free access to leading AI applications and accompanying skills training to 20,000 students in the tenth and eleventh grades and 3,000 teachers, with a planned second wave adding 38,000 students and 2,000 teachers✓ Established [28]. The programme was negotiated directly with frontier model developers and explicitly frames itself as the successor to the Tiger Leap initiative that put computers into Estonian schools three decades ago✓ Established [28]. It targets an age band well past the fourth-grade danger zone.

2017
France restricts phones in schools — France becomes an early mover on classroom device restriction, a policy since adopted in some form by 106 countries and 39 United States states✓ Established [36].
July 2023
Japan issues provisional guidance — Japan's education ministry publishes provisional guidelines on text-based generative AI in primary and secondary schools, warning that pupils may accept model output as fact✓ Established [30].
January 2024
The Netherlands bans devices in secondary classrooms — A national agreement rather than a statute takes effect in secondary schools, extended to primary and special education after the summer✓ Established [25].
December 2024
Japan revises to version 2.0 — The revised guidance centres children's developmental stage, information literacy and teacher training, and restricts the use of pupils' sensitive data✓ Established [30].
December 2024
The adult skills baseline lands — The OECD reports literacy declining or stagnating across 31 countries over a decade, with the widest deterioration among the least educated✓ Established [18][19].
February 2025
The European Union makes AI literacy an obligation — Article 4 of the AI Act becomes applicable, requiring providers and deployers to ensure sufficient AI literacy among staff and affected users, proportionate to role and competence✓ Established [29].
February 2025
Denmark legislates a phone-free school day — The Danish parliament agrees a legal requirement for mobile-free primary and lower-secondary schools, with a standard policy template supplied centrally✓ Established [24].
May 2025
China makes AI compulsory from age six — Two Ministry of Education guidelines set a floor of eight hours a year from the first grade and bar independent generative AI use by primary pupils✓ Established [27].
August 2025
South Korea reverses — The National Assembly strips AI digital textbooks of legal textbook status, reclassifying them as supplementary materials; adoption falls from 37% to 19% within a semester✓ Established [26].
September 2025
Estonia distributes frontier models to schools — AI Leap begins with 20,000 upper-secondary students and 3,000 teachers, negotiated directly with leading model developers✓ Established [28].

South Korea supplies the cautionary case, and it is a procurement failure rather than a pedagogical one. AI digital textbooks were piloted in the first semester of 2025 for English and mathematics in the third and fourth grades and for English, mathematics and computer science in secondary schools✓ Established [26]. Facing sustained opposition from teachers and parents, the ministry moved from a national mandate to voluntary school-by-school adoption; on 4 August 2025 the National Assembly narrowed the statutory definition of a textbook to printed books and electronic books, excluding intelligent learning-support software✓ Established [26]. Adoption fell from 37% to 19% in a single semester✓ Established [26].

The device-restriction wave is a different policy addressing a different mechanism, and its evidence base is thin but not empty. Restrictions now exist in 106 countries and 39 United States states✓ Established [36]. In the Netherlands, a survey of 317 secondary schools after the national ban found three-quarters reporting improved concentration, close to two-thirds an improved social climate, and one-third better academic performance◈ Strong Evidence [25]. A 2025 study found higher grades where teachers collected phones, with the largest gains among lower-achieving first-year and non-STEM students◈ Strong Evidence [25]. UNESCO's assessment remains that outcomes are mixed and that restriction does not remove the need to teach navigation of digital environments⚖ Contested [24].

✓ Established The European Union has made AI literacy a legal obligation without defining what competence it protects

Article 4 of the AI Act became applicable on 2 February 2025, requiring every provider and deployer to take measures ensuring a sufficient level of AI literacy among staff and other affected persons, proportionate to their role, competence and the systems involved✓ Established [29]. The obligation is real and general. What it does not specify is which human capacities the literacy is meant to preserve — a gap that matters because the deskilling evidence concerns the erosion of domain skill, not of understanding about the tool◈ Strong Evidence [1]. Knowing how a model works does not maintain the search behaviour it replaces.

The common failure across all five approaches is evaluative rather than ideological. None of these programmes was launched with a pre-registered design capable of detecting the effect that matters — unassisted capability, measured over years, against a comparable cohort. The OECD's next round of international student assessment, with results due in September 2026 from more than 760,000 students in 91 countries, will for the first time include evidence on how students use AI for schoolwork◈ Strong Evidence [34]. That is the only large instrument currently pointed at the question, and it is observational.

08

What the Evidence Will Carry
The tool is not the variable; the removed step is

Two propositions survive the audit. Offloading is an ancient, rational and largely benign strategy that has repeatedly been mistaken for decline◈ Strong Evidence [12][32]. And tools that substitute for the retrieval or generation step measurably degrade the capacity they substitute for, in adults, within months, at professional levels of expertise✓ Established [1][8]. These are not competing claims. They describe different tools, and the entire practical question is which of the two a given deployment resembles.

The case that offloading is adaptive rests on the strongest kind of evidence available in this field, which is repeated historical falsification of the alarm. Writing did not degrade civilisation✓ Established [32]. Calculators improved mathematical attainment at every grade but one and improved attitudes at all of them✓ Established [13]. The pen-versus-laptop result, deployed for a decade as proof that digital media impair encoding, failed direct replication twice⚖ Contested [15][16]. The Google effect's most-cited experiment did not replicate⚖ Contested [11]. A field with this record of overclaiming has earned scepticism about its newest claim.

The case that offloading is corrosive rests on a smaller body of work with better identification. The satellite navigation result has a longitudinal component and a tested alternative explanation◈ Strong Evidence [8]. The colonoscopy result is within-operator, within-centre and concerns experts✓ Established [1]. The aviation position is a regulator's conclusion after accident investigation rather than a laboratory finding✓ Established [31]. The Fan experiment separates artefact quality from knowledge transfer under controlled conditions✓ Established [7]. None of these is a survey of self-reported thinking, which is what most of the alarming AI literature currently consists of.

The reconciliation is the fourth-grade rule generalised. Assistance applied to an operation the user has already automated is close to free, and often better than free, because it releases capacity for the next level of the task✓ Established [13][14]. Assistance applied to an operation the user has not yet automated occupies the slot where automaticity would have formed✓ Established [13]. Assistance applied to an operation the user automated long ago, and then continuously, allows that automaticity to lapse✓ Established [1][8]. Same tool, three populations, three signs.

The case that offloading is adaptive

The alarm has a two-thousand-year record of being wrong
Socrates predicted that writing would destroy memory; literate societies became more capable, not less✓ Established [32]. Every subsequent iteration of the argument has been made about a technology now regarded as obviously beneficial.
The controlled evidence on tools is mostly positive
Seventy-nine calculator studies found improved pencil-and-paper skill at every grade except the fourth, and improved attitude and self-concept at all of them✓ Established [13].
The headline psychological findings did not replicate
The most-cited component of the Google effect failed high-powered replication⚖ Contested [11], and the longhand note-taking advantage failed two direct replications⚖ Contested [15][16].
Offloading is metacognitively governed, not automatic
People offload more when internal demand is high and less when they judge their own capacity sufficient — a strategy, not a reflex◈ Strong Evidence [12].
Aggregate skill declines predate the technology
Adult literacy was already stagnating across the OECD over a decade✓ Established [18], and the reversal of the Flynn effect was measured from 2006 to 2018, before large language models existed✓ Established [22].

The case that offloading is corrosive

Experts deskilled within three months
Nineteen endoscopists with more than 2,000 procedures each lost six percentage points of unassisted detection after routine AI exposure✓ Established [1].
The structural cost has been measured longitudinally
Heavier satellite navigation use predicted a steeper three-year decline in hippocampal-dependent spatial memory, with the reverse-causation account tested and unsupported◈ Strong Evidence [8].
The artefact improves while the knowledge does not
Essay scores rose in the ChatGPT condition; knowledge gain and transfer did not differ significantly from the unsupported condition✓ Established [7].
Behaviour has already changed at population scale
Study time on AI-susceptible mathematics topics fell a cumulative 26.9% among university students, with a 25% fall in the odds of a correct proctored retention answer◈ Strong Evidence [21].
The people affected cannot detect it
Confidence in the system predicts less scrutiny of its output✓ Established [6], and offloading degrades the metacognitive signal used to judge one's own competence◈ Strong Evidence [12].

Three propositions in current circulation are not supported and should be retired. That generative AI is measurably lowering intelligence: the population-level cognitive series that are falling began falling before the technology existed✓ Established [22][18]. That search engines have hollowed out memory: the strong form of the Google effect did not survive replication⚖ Contested [11]. And that friction is inherently valuable in learning: desirable difficulty effects attenuate or reverse when working memory is already loaded⚖ Contested [17]. Each of these is a real finding stretched past what its design can bear.

What follows practically is narrow and unglamorous. Withhold assistance during the formation of the procedures being automated, which is where the only clear negative result in the calculator literature sits✓ Established [13]. Maintain unassisted performance deliberately in any setting where the human is the backup, which is what the aviation regulator concluded and what the colonoscopy finding implies✓ Established [31][1]. And, above all, keep measuring the unassisted case, because it is the only observation that makes the effect visible and it is the first one every deployment stops collecting◈ Strong Evidence [21].

The Real Finding

Cognitive offloading is not a single phenomenon with a single effect, which is why forty years of studies appear to contradict each other◈ Strong Evidence [12]. The variable that reconciles them is not the sophistication of the tool but the position of the removed step relative to the user's existing competence. Remove a step the user has automated and you free capacity✓ Established [13]. Remove a step the user is in the middle of automating and you prevent the automation✓ Established [13]. Remove a step the user automated years ago, continuously, and you allow it to lapse✓ Established [1][8]. Large language models are the first widely deployed tool that can occupy all three positions at once, in the same person, on the same afternoon.

The measurable questions for the next five years are already specified. Whether a cohort trained under continuous assistance from the outset acquires the baseline capacity at all, which no study has yet examined⚖ Contested [1]. Whether the retention decline visible in the mathematics panel survives designs that are not quasi-experimental⚖ Contested [21]. Whether the international student assessment due in September 2026 detects anything in the first cohort to sit it with these tools in general use◈ Strong Evidence [34]. And whether any institution, anywhere, chooses to keep collecting the unassisted measurement once it is no longer required to. On present evidence, that last one is the constraint that binds.

SRC

Primary Sources

All factual claims in this report are sourced to specific, verifiable publications. Projections are clearly distinguished from empirical findings.

Cite This Report

APA
OsakaWire Intelligence. (2026, September 2). Cognitive Offloading (2026) — Skill Fell 21% Without AI. Retrieved from https://osakawire.com/en/cognitive-offloading-minds-that-stop-remembering/
CHICAGO
OsakaWire Intelligence. "Cognitive Offloading (2026) — Skill Fell 21% Without AI." OsakaWire. September 2, 2026. https://osakawire.com/en/cognitive-offloading-minds-that-stop-remembering/
PLAIN
"Cognitive Offloading (2026) — Skill Fell 21% Without AI" — OsakaWire Intelligence, 2 September 2026. osakawire.com/en/cognitive-offloading-minds-that-stop-remembering/

Embed This Report

<blockquote class="ow-embed" cite="https://osakawire.com/en/cognitive-offloading-minds-that-stop-remembering/" data-lang="en">
  <p>After three months of AI-assisted colonoscopy, 19 experienced doctors detected 21% fewer adenomas without it. Forty years of offloading evidence, audited.</p>
  <footer>— <cite><a href="https://osakawire.com/en/cognitive-offloading-minds-that-stop-remembering/">OsakaWire Intelligence · Cognitive Offloading (2026) — Skill Fell 21% Without AI</a></cite></footer>
</blockquote>
<script async src="https://osakawire.com/embed.js"></script>