Are We Outsourcing Our Thinking to AI? What the Research Actually Shows

I’ve been using an AI assistant for everything for two years now, and at some point I noticed a small, unsettling thing: when the model handed me a perfect answer without me struggling first, I felt great and retained nothing. So I went down the research rabbit hole, expecting a clean answer, and instead kept finding studies that said both things. AI-assisted students got 48% more math problems right but scored 17% worse on understanding the concepts. Physicians got faster with AI, and then got worse without it.

So, are we outsourcing our thinking to AI? The honest answer is conditional: the offloading mechanism is real and measurable, several viral studies are much weaker than the headlines suggest, and usage structure, not the technology, decides whether you come out sharper or softer.

Key Takeaways

A Polish natural experiment (Budzy? et al. 2025, Lancet Gastroenterology & Hepatology) found physicians’ unassisted polyp-detection rate fell from 28.4% to 22.4% within three months of AI becoming available.

The same tools help or harm depending on sequencing: Akgun and Toker’s 2024 study found pretesting before AI use improved retention, while prolonged AI exposure led to memory decline.

Gerlich’s 2025 three-condition study (150 participants) found structured ChatGPT prompts after independent effort produced the highest expert-rated essays, which is the basis of the attempt-first protocol.

What cognitive offloading actually is, and why AI is different

Cognitive offloading is the use of external tools to reduce mental effort, with people constantly weighing task goals against tool capability. That’s the definition from Evan Risko and Sam Gilbert’s 2016 paper in Trends in Cognitive Sciences (volume 20, pages 676-688, if you want the receipts), and it’s the repo everything else in this article forks from.The bulleted version:

  • You have a goal (get to the restaurant, remember the groceries).
  • A tool can carry part of the load (GPS, a list on your phone).
  • Your brain runs a quick cost-benefit check and offloads, or doesn’t.

The check works fine until the tool isn’t there. GPS is the clean example: Miola et al. in 2024 and Dahmani and Bohbot in 2020 both linked GPS reliance to a weaker sense of direction, and offloading only helps while the stored information stays accessible. Dead phone in a strange city is the failure mode. A grocery list you left at home is the low-stakes version.

Here’s what makes AI categorically different from writing, calculators, and GPS: those offloaded discrete tasks. AI offloads judgment, critical thinking, and creativity, which is the cognitive work people actually build identities around. It’s the difference between outsourcing your errands and outsourcing your decisions.

The evidence that skills actually decay

Yes, measurable skill decline shows up under passive AI reliance, and the most convincing evidence comes from experts, not students.Okay, this one is wild: Budzy? et al. published a natural experiment in Lancet Gastroenterology & Hepatology 10(10), 2025, tracking Polish physicians after AI-assisted colonoscopy software was introduced. Within three months, their unassisted polyp-detection rate fell from 28.4% to 22.4%. Six points in a quarter. These are experienced doctors whose unaided performance slipped once a better tool was always in the room.

The student side comes from the University of Pennsylvania study of Turkish high schoolers (reported by Jill Barshay at The Hechinger Report): AI-assisted students answered 48% more math problems correctly but scored 17% lower on a concept-understanding test, and the students who practiced with ChatGPT did worse on exams. The reading that stuck with me: it trains the syntax, not the semantics. There’s also Lee et al. 2025, a Microsoft and Carnegie Mellon study presented at CHI with 319 knowledge workers, which found that higher confidence in generative AI correlated with less critical thinking during use. That one’s self-report survey data, so calibrate accordingly.

So, is AI making us dumber?Depends on the conditions. Natural experiments with experts carry more weight than homework anecdotes because the stakes and the measurement are real, and the same concern pattern shows up in radiology, surgery, law, and finance. It’s not proof that every profession decays; it’s proof the mechanism exists.

False mastery: why AI-assisted work feels like learning

Relying on AI does weaken problem-solving when the use is passive and unstructured, and the reason it feels productive is the cruelest part.The Lodge report on AI in education names the mechanism: unstructured AI use produces output fluency, which creates an illusion of competence the researchers call metacognitive laziness, and it bypasses the generation effect, the well-documented quirk where things you generate yourself encode into memory and things you merely receive don’t.

Dan Levy, a Harvard Kennedy School lecturer, puts the underlying rule simply: no learning occurs unless the brain is actively making meaning. AI working with a student does more than AI doing the work for them. Bloom’s taxonomy maps this neatly: AI happily handles recall and synthesis while the higher-order judgment work quietly goes undone.

One question diagnoses it: can you reproduce and defend the reasoning without the tool? The copy-pasted Stack Overflow answer that compiled but taught nothing is the classic version. Fluency in production paired with vagueness on demand is the tell, and it’s common enough among heavy AI users that researchers have a name for it.

Memory and retention: sequencing matters more than dosage

No, AI isn’t uniformly bad for memory, except when exposure replaces retrieval first. That condition is the whole story, and someone actually tested it.Akgun and Toker, in 2024, ran a two-group study with 73 information science undergraduates at a Pennsylvania university, and the design is the interesting part: one group pretested, meaning they tried recalling material before touching the AI tool. The pretesters showed better retention and engagement. Students who got prolonged AI exposure first showed memory decline.

Student quizzing themselves with flashcards before opening an AI tool, showing sequencing for memory retention
Akgun and Toker’s finding in one workflow: recall first, AI second, and retention goes the right way.

Same tool, different order, different outcome.Sequence, not dosage alone, is the moderating variable.

The broader concern here has a name, digital amnesia, and the GPS findings are the everyday anchor: memory offloading works only while the external store stays reachable, and the underlying skill quietly weakens when you never exercise it. It’s skipping leg day because the elevator works. The reps matter, especially for people still building the skill.

Two more data points round out the picture. Bai, Liu, and Su in 2023 found that ChatGPT enhances personalized learning, but excessive reliance reduces cognitive engagement and long-term retention. And Ododo et al. in 2024 surveyed 206 vocational education students in Akwa Ibom State, Nigeria, and the students themselves perceived AI as a threat to retention and critical thinking, with passive acceptance of AI answers flagged as the default-mode trap.

(Curious wrinkle: male students were more concerned than female peers. Perception data, not proof of harm, but worth surfacing.) Çela et al. 2024 and Grinschgl and Neubauer 2022 fill in the mechanism: AI cuts opportunities for active recall and problem-solving, which is what drives engagement and retention down. Caveat honestly: 73 undergraduates in one discipline is a benchmark spec, not a verdict on the general population.

When offloading helps vs. harms, from the evidence:

ConditionHelps retentionHarms retention
Pretest/recall before AI exposureYes (Akgun and Toker 2024)No
Prolonged AI exposure replacing retrievalNoYes (memory decline in same study)
Moderate, personalized ChatGPT useYes (Bai, Liu, and Su 2023)No
Excessive relianceNoYes (lower engagement, worse long-term retention)

Creativity: amplified divergence, induced fixation

The decisive variable in AI creativity is who controls the ideation loop.Habib et al. 2024 ran a mixed-methods experiment with undergrads in a Creative Thinking and Problem-Solving class at a Southeastern US university, using ChatGPT-3 on the Alternative Uses Task. AI-supported students scored better on divergent-thinking fluency, flexibility, and elaboration, which is a genuine win, but they also showed cognitive fixation and lower creative confidence. Anyone who’s anchored to the first suggestion a model offers, on a naming task or a debugging one, knows exactly what that fixation feels like.

Lu et al. 2025, an MIT field experiment with 250 Chinese tech consulting employees, found the workplace half: employees with ChatGPT access produced work rated more creative, and the effect was strongest among people high in metacognitive skill.Real workers, real output, real ratings.

There’s a ceiling, though, and expert consensus across Macnamara, Gerlich, and Farahany lands on it: AI recombines existing knowledge but doesn’t make original leaps. The analogy from APA Monitor is too good not to steal: AI can improve a candle, but it can’t invent the light bulb. Good remix artist, not the originator.

What’s overblown: auditing the viral studies

The most-shared AI-brain study is also the weakest evidence in this whole space, and that should recalibrate how you read every headline like it. The MIT Media Lab EEG study (Kosmyna et al., 2025, arXiv, doi 10.48550/arXiv.2506.08872) found weaker neural connectivity in LLM-assisted essay writers versus search-engine or no-tool groups. Cool preprint, worth reading. It’s also small, and it’s a preprint that hasn’t been peer-reviewed.

Suggestive, low-confidence. Anyone telling you it proves brain harm is reading past the methods.

Then there’s the claimed 15% Stanford-linked AI-tutoring test-score gain. I went looking for the source, and it traces to a commercial blog (ideta.io), not a paper.Flagged as low-confidence and unverified, and it loses badly to the Penn study’s paired 48%-more-correct / 17%-worse-understanding finding, which at least measures both sides. Dede’s observation is worth keeping in your back pocket too: 95% of AI-in-education coverage frames the tech as “do things better.” The hype runs in both directions, panic and boosterism.

Longer precedents exist for tools shaping cognition: longhand versus keyboard note-taking, predictive text, calculators, GPS. Some of the concern is legitimate without being catastrophic, and Lu’s field experiment showing real creativity gains is exactly why blanket panic doesn’t survive contact with the data.

A framework for when AI helps vs. harms thinking

Structured AI use is the answer: it helps when it reduces busywork without replacing effortful thinking, and it harms under three predictable conditions. Gerlich’s 2025 study is the proof point: in his three-condition experiment (150 participants, essay on democracy), the structured-prompt ChatGPT condition scored highest on expert-rated essays, while AI use reduced critical thinking when it was mediated by passive offloading.

Scale balancing AI assistance against human judgment, representing the framework for when AI helps thinking
Same tool, opposite outcomes, the framework says the deciding variable is whether the tool carries busywork or judgment.

Three lenses, each in one breath:

  • Cognitive Load Theory (Schnotz and Kürschner, 2007): AI cuts extraneous load, the friction of the task, but over-reliance starves germane load, the effortful processing where deep learning actually happens.
  • Bloom’s Taxonomy (Shaikh et al., 2021): AI can automate recall, the bottom of the pyramid, while judgment and analysis atrophy if it always supplies the answer.
  • Self-Determination Theory (Ryan and Deci, 2000): AI builds competence through customization, but excessive dependence compromises autonomy, and relatedness still needs actual humans.

The unified rule as a checklist you can run on any AI task: Would I have attempted this myself first?Is the tool carrying busywork or the judgment? Does the stored answer stay substitutable for my skill, or is the skill the point? If AI enters before thinking, you’ve offloaded the thinking. If it enters after, you’ve offloaded the typing.

It’s not just willpower: incentives, equity, and children

Whether AI upskills or de-skills workers depends on organizational incentives, psychological safety, and usage patterns rather than the technology, a conclusion drawn by researchers including Brooke Macnamara, PhD, Nita Farahany, Mindy Shoss, and Microsoft’s David Evans, and Microsoft’s own reporting shows both outcomes happening in the wild. The incentive bug is that speed-optimized workplaces reward output volume, so unvetted AI drafts (“AI slop”) flow because slowing to verify is personally costly. Equity inverts the democratization story: strong prior knowledge and metacognition predict who benefits, while disadvantaged students are more susceptible to harmful offloading, widening gaps. And on children’s development, the Brookings “cognitive stunting” hypothesis is explicitly low-confidence: kids may shortcut skill development rather than offload mastered skills, with proposed measurement through the NIH Toolbox (ages 7-85+), the CDC’s “Learn the Signs.

Act Early.” program (ages 0-5), and the ABCD Study (10,000+ adolescents). Meanwhile, an 84-study review found narrow, teacher-mediated AI within strong instructional design enriches learning, and the FairAIED work (Chinta et al., 2024) flags that AI trained on unbalanced datasets can perpetuate grading inequalities, with transparency, diversified multilingual datasets, and fairness-aware algorithms as the countermeasures.

What actually keeps your thinking sharp: the attempt-first protocol

The attempt-first protocol, attempt the task yourself, then bring in AI, is what keeps thinking sharp during daily AI use. That’s not vibes; it’s the winning condition in Gerlich 2025: 150 participants, three conditions, essay on democracy, and structured ChatGPT prompts issued after independent effort produced the highest expert-rated essays, while passive offloading mediated the critical-thinking decline.Solve first, then query. The protocol, bench-tested against the evidence:

  • Attempt first, then query. Gerlich’s structured-prompt condition won. The workflow difference matters: drafting your own attempt or question before opening the chat window means you can spot AI errors, whereas starting from the AI answer tends to mean accepting it wholesale.

    I’ve done the latter. It compiles and teaches nothing.
  • Pretest before exposure. Akgun and Toker’s sequencing rule: test yourself first, then offload. Quiz yourself on the material before asking the model to explain it.
  • Restate AI answers in your own words.

    Reflection points work for adult learners too, not just classrooms. If you can’t restate it, you didn’t learn it.
  • Sort your tasks. Mutlu Cukurova’s task analysis separates completion-only tasks (offload freely) from essential-learning tasks (protect them). Transactional interactions, quick in, answer out, no thinking in between, skew toward cognitive atrophy.
  • Keep AI in a subsidiary role. Socratic partner, cognitive mirror, verification partner rather than answer oracle, per the Brookings task force.

    It’s a design philosophy you can apply to your own prompts, and it’s the same pattern the evidence favors in schools: teacher augmentation beats straight AI-tutor deployment, and New South Wales’ Department of Education trains teachers on AI plus Load Reduction Instruction, with critical-thinking gains reported.
  • Periodic AI-free solving sessions. Rest days for the brain, and peer discussion counts toward the same reps.

Where the habits come from and what they protect:

Person working through a problem on paper before consulting AI, demonstrating the attempt-first protocol
The attempt-first protocol in practice: earn your own attempt before the chat window gets a turn.
HabitEvidence behind itWhat it protects
Attempt first, then queryGerlich 2025: structured prompts after independent effort wonCritical thinking, error-spotting
Pretest before AI exposureAkgun and Toker 2024: pretesting improved retentionMemory encoding
Restate answers in your own wordsReflection points in the education literatureConcept understanding vs. syntax
Sort completion-only vs. essential-learning tasksCukurova’s task analysisHigher-order judgment
AI-free solving sessionsRest-day framing from the mitigation researchActive recall reps

Two honesty notes. The principles-over-rote principle matters: teach the periodic table and long division as reasoning, not drilling, because understanding transfers and rote doesn’t. And the evidence suggests these habits preserve thinking; it doesn’t guarantee it.Nothing here is a certificate against decay.

One more reason not to worry about humans being obsoleted: human cognition still brings things models don’t. Grotzer and Damasio’s work on somatic markers and analogical reasoning describes the gut checks and lived-experience pattern-matching AI can’t do, and there’s a delightful data point where kindergarteners outperformed a purely Bayesian model at a learning task. Tiny humans beat the model.Also, researchers are actually studying clinician evaluation of AI outputs with NSF funding, because checking the checker is its own skill.

And sometimes offloading is just a choice: Nita Farahany offloads her email to AI to buy family time. That’s offloading as a decision, not a trap.

The verdict

We’re outsourcing our thinking only where AI enters before thinking, in passive, unstructured, incentive-driven use, and not where it enters after.Shiri Melumad’s 2025 study in PNAS Nexus 4(10) is a good caution here: AI-tool users searched less and wrote shorter, less trustworthy garden advice rated as less helpful, which is what defaulting to the chat window over the open web quietly costs you. Many long-term cognitive effects remain unanswered, and the American Psychological Association (APA) is actively examining them. AI is an amplifier or an inhibitor. The usage structure decides which one you get.

Frequently Asked Questions

Is AI making us dumber?

It depends on the conditions. The offloading mechanism is real and measurable — doctors’ unassisted polyp-detection rates fell within months of AI-assisted colonoscopy software arriving — but the same tools help or harm depending on usage structure. Structured use after independent effort produced the best outcomes in controlled studies, while passive, unstructured reliance mediates the declines.

Is AI bad for your brain and memory?

Not uniformly — sequencing matters more than dosage. In Akgun and Toker’s 2024 study, students who tested their recall before using an AI tool showed better retention, while students with prolonged AI exposure first showed memory decline. The GPS analogy applies: memory offloading works only while the external store stays reachable, and the underlying skill weakens when you never exercise it.

When does using AI help learning instead of hurting it?

When AI enters after effort rather than before it. The winning pattern in Gerlich’s 2025 experiment was structured ChatGPT prompts issued after independent effort, and pretesting your recall before AI exposure improved retention in Akgun and Toker’s study. Moderate, personalized use helps; excessive reliance that replaces retrieval starves the effortful processing where deep learning happens.

How is AI affecting children’s cognitive development?

Honest answer: nobody knows yet with confidence. The Brookings ‘cognitive stunting’ hypothesis — that kids may shortcut skill development rather than offload already-mastered skills — is explicitly low-confidence, and the American Psychological Association is still examining long-term effects. An 84-study review did find that narrow, teacher-mediated AI within strong instructional design enriches learning, while equity concerns remain since disadvantaged students are more susceptible to harmful offloading.

Is outsourcing a dying concept?

No, but it’s shifting scope. Writing, calculators, and GPS offloaded discrete tasks like errands and navigation; AI outsources judgment, critical thinking, and creativity — the decisions themselves. The practical line is whether the skill is the point: offload completion-only tasks freely, protect essential-learning tasks, and keep AI in a subsidiary role like a Socratic partner rather than an answer oracle.

Leave a Comment