Research

The research on AI and cognition

What the studies actually found, what they do not establish, and where the evidence is thinner than the headlines. Every entry carries its own caveat, because a summary that drops the sample size is how a modest result becomes a certainty.

Written by Suman Debnath, creator of IMPRINTLast updated 5 September 2026
How to read this page. The evidence that AI delegation reduces cognitive engagement during a task is reasonably good. The evidence that it causes lasting capability loss is not — it is inferred from skill-decay research that predates AI. IMPRINT is built on the first claim and is agnostic about the second. Anyone telling you the science is settled in either direction has not read it.

Direct measurement of AI-assisted cognition

Studies that instrumented people while they worked with an AI assistant, rather than surveying them afterwards. These are the closest thing to direct evidence, and also the smallest.

Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task

Kosmyna et al., MIT Media Lab · 2025 · arXiv preprint

What it found. Fifty-four participants wrote essays across three conditions — with an LLM, with a search engine, or unaided — under EEG. The LLM group showed the weakest and least distributed neural connectivity of the three, and performed worst at the neural, linguistic and scoring levels across sessions. Asked to quote from essays they had just submitted, LLM-group participants frequently could not.

What it does not show. Fifty-four participants, with only eighteen completing the fourth crossover session. One task type, one model, and a preprint rather than a peer-reviewed publication. It shows reduced engagement during assisted work; it does not show lasting cognitive damage.

Why it is here: The study that introduced 'cognitive debt', and the one most often cited in this conversation.

Cognitive offloading and critical thinking

Survey and task-based work on whether frequent AI use tracks with weaker critical thinking, and whether offloading is the mechanism.

Cognitive offloading, critical thinking and attitudes towards artificial intelligence in the era of ChatGPT

Comparative study in young adults · 2025 · Peer-reviewed; indexed on PubMed

What it found. Reports a negative association between AI-assisted task performance and measures of critical thinking in young adults, consistent with cognitive offloading as the mediating mechanism rather than AI exposure as such.

What it does not show. Correlational. People who already prefer to offload may simply reach for AI more often, which the design cannot separate from AI causing the offloading.

Why it is here: Compares AI-assisted against manual task performance directly, rather than relying on self-report alone.

AI tools in society: impacts on cognitive offloading and the future of critical thinking

Gerlich, M. · 2025 · Societies · link unverified

What it found. Across 666 participants, frequent AI tool use correlated negatively with critical thinking scores, with cognitive offloading mediating the relationship. Younger participants showed higher AI dependence and lower critical thinking scores than older ones.

What it does not show. Correlational and largely self-reported, so it cannot establish direction. The age effect is confounded with familiarity and with baseline differences in how the two groups were educated. Listed without a link because the canonical URL was not verified — cite the journal record directly.

Why it is here: The largest sample commonly cited in this area, and the source of the 666-participant figure quoted throughout the debate.

Mechanism and theory

Work on why delegation would degrade capability at all — the learning-science and human-factors literature that predates generative AI.

From Co-Design to Metacognitive Laziness: Evaluating Generative AI in Vocational Education

Multiple authors · 2025 · arXiv preprint

What it found. Describes 'metacognitive laziness': learners working with generative AI accepted output with less evaluation over time, and the capacity to judge output quality degraded alongside the habit of judging it.

What it does not show. Educational setting with a specific cohort. A preprint, and the construct is newer than the evidence base supporting it.

Why it is here: Names and examines the pattern where fluent output suppresses the impulse to evaluate it.

A Review of the Negative Effects of Digital Technology on Cognition

Review · 2026 · arXiv preprint

What it found. Surveys evidence on digital technology and cognitive function, covering attention, memory and self-regulation, and finds effects that are real but generally smaller and more context-dependent than popular accounts suggest.

What it does not show. A review rather than new evidence, and it inherits the heterogeneity of what it reviews. Effect sizes across this literature are frequently small.

Why it is here: Places the AI conversation inside the longer literature on technology and attention, which is a useful corrective to treating 2025 as year zero.

Measuring reliance itself

Attempts to quantify how much a person is depending on a system — the problem IMPRINT is also trying to solve, approached differently.

Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

Multiple authors · 2026 · arXiv preprint

What it found. Proposes measuring AI reliance by comparing a completed workflow against a counterfactual in which the assistance was unavailable, producing a reliance measure grounded in task outcome rather than in self-report or in surface features of the output.

What it does not show. A proposed method rather than a validated instrument, and counterfactual workflows are expensive to construct. IMPRINT's lexical approach is cheaper and correspondingly weaker.

Why it is here: The closest published work to what IMPRINT's Drift Score attempts, and a more rigorous approach to the same question.

The Cognitive Divergence: AI Context Windows, Human Attention Decline, and the Delegation Feedback Loop

Multiple authors · 2026 · arXiv preprint

What it found. Argues that expanding model context capability and declining sustained human attention form a reinforcing loop, in which each increment of delegation makes the next one more attractive and less noticeable.

What it does not show. Largely theoretical. The feedback loop is argued rather than measured, and the attention-decline evidence it draws on is contested.

Why it is here: Frames delegation as a feedback loop rather than a one-off choice, which is the dynamic IMPRINT's recurring calibration assumes.

Corrections welcome

Studies are listed with a link only where the canonical URL has been checked. Where it has not, the entry is marked and the finding is described without one — a wrong citation on a page like this costs more than a missing one.

If something here is misdescribed, out of date, or missing a paper that belongs, say so. How IMPRINT applies this research is set out on the methodology page, along with what its own measurements cannot do.