How to Use AI Without Losing Your Own Thinking Skills: What a 1,000-Student Study Reveals

I went digging through the research on AI and thinking skills expecting to write a think piece, and instead found something better: a bunch of studies that actually ran the experiment we all secretly want run. Not vibes, not op-eds. Field experiments with effect sizes. And the punchline is genuinely useful: the evidence shows AI’s effect on your thinking is conditional on how you use it, not which tool you use. That’s the whole answer to how to use AI without losing your own thinking skills, and it converts a vague anxiety into a controllable variable.

Key Takeaways

In Bastani et al.’s 2025 PNAS field experiment with nearly 1,000 high school math students, unguarded ChatGPT users improved 48% while using it, then scored 17% below a no-AI control once it was taken away; the same underlying model with tutoring guardrails boosted performance 127% with no post-use harm.

The sequencing is the safeguard: draft your own rough answer or idea list first, then let AI critique and expand it (van Brunt’s think-first rule, Project FAILSafe guidance).

Ryan C. Warner’s sourced boundary: cap AI at no more than 50% of the work on complex projects, and always articulate AI output in your own words afterward.

Does using AI make you dumber, or smarter? The evidence is conditional

It depends on the mode of use, full stop. The harm studies and the benefit studies agree on that.

The harm side first, because it’s the side that made headlines. That Bastani study: students using unguarded ChatGPT (the paper calls it GPT Base) scored 17% lower than students who never used AI at all once the tool was removed. The work got done. The skill didn’t stick.

Meanwhile Fan et al. (2025) ran 117 students through essay writing, and the ChatGPT group produced the highest-quality essays while showing zero gains in learning, motivation, or interest, with frequent copy-pasting. The researchers named that pattern metacognitive laziness, which might be the most honest term in the whole literature. And there’s a nasty feedback loop underneath: Bauer et al. (2025) found heavy reliance weakens your domain knowledge, which makes AI errors harder to spot, which invites more reliance.

The benefit side is real too. Wang & Fan (2025) crunched a meta-analysis of 51 studies and found a learning gain of g=0.867 and higher-order thinking gains of g=0.457. de la Puente et al. (2024) found AI-assisted debate exercises improved argumentation versus a non-AI control. Deng et al. (2024) reviewed 69 studies and found the same pattern: used constructively, AI helps learning.

Both findings are true. The variable is mode of use. One honest caveat, since searchers keep landing on the MIT study on AI’s effect on the brain: the controlled evidence here is behavioral, meaning test scores and learning measures, not brain imaging. Popular coverage tends to outrun the imaging evidence that actually exists, so be skeptical of anyone telling you they know what AI does to your neurons.

The mode-of-use test: is AI doing the thinking, or helping you think?

The ICAP framework is the answer. ICAP (Chi & Wylie, 2014) sorts engagement into four levels, passive, active, constructive, interactive, and only the top two produce durable learning gains.

Two humanoid robots with glowing faces and a young man studying at a desk, surrounded by books and a laptop, in a modern tech environment.
The APA’s seconds-long audit at the moment you open the AI tab: is it doing the learning for you, or helping you learn?

Applied to AI, it becomes a decision rule. Reading AI output and skimming it? Passive. Arguing with it, building on it, making it attack your reasoning?

Constructive and interactive. Same tool, wildly different outcomes for your brain.

The killer evidence for this lives in the Bastani experiment itself. In the same study, the same underlying model produced that 48%-boost-then-harm pattern under unguarded use, and a 127% boost with no post-use harm under a tutored, structured version called GPT Tutor. The guardrails changed the outcome. Not the model.

Not the vendor. The mode.

The APA published a guiding question in its Psychology Teacher Network (December 10, 2025) that works as a seconds-long audit at the exact moment you reach for the AI tab: is AI doing the learning for me, or is it helping me learn? Ask it every time. It costs nothing and it’s the single highest-leverage habit in this entire article. For a fuller treatment of the underlying distinction, we’ve got a piece on what it means to outsource your thinking that goes deeper on cognitive offloading.

Draft first, prompt second: the sequencing rule

Always produce your own rough answer or idea list before you open the AI tab. The same interaction strengthens or weakens your thinking depending on whether it came after your own first attempt. Order of operations is the highest-leverage variable here, and there’s real evidence behind it.

van Brunt’s think-first rule

Kristin van Brunt, a teacher at Mueller Park Junior High in Utah, requires students to generate their own ideas about a book’s themes before asking AI which topic is worth exploring. Her Animal Farm exercise: the thinking comes first, the AI consults on it, never instead of it. The personalization part is the clever bit; when her students tie Of Mice and Men themes to their own experiences, that personal stake is the one thing AI can’t fake for them, so it can’t do all the work. Project FAILSafe guidance says the same thing for adults: draft a rough answer first, then let AI critique it, keeping idea generation, sense-making, and judgment human-owned. Version one comes from your brain.

Young’s expand-a-formed-judgment workflow

Fred Young, a secondary math and CS teacher in Dansville Central School District, does something that clicked for me immediately. He uploads classroom activities he’s already used and judged, asks ChatGPT for more engaging versions, gets five or six options, and picks and modifies. The AI is expanding a judgment he already formed, not forming one for him, and this is exactly where to draw the line between assistance and shortcut. He calls it a huge time saver, and it is, because the expensive cognitive step (deciding what’s good) stays with him. His algebra students run the same pattern in reverse: hands-on experiments first, physical and messy, then they compare their own conclusions against AI-generated responses.

The pattern generalizes way past classrooms. If you want more on the classroom side specifically, there’s a whole piece on what outsourcing thinking means for education.

Use AI as a sparring partner, not an answer key

The sparring-partner workflow is simple to state and easy to get wrong: instead of asking AI for answers, hand it your reasoning and make it attack. Draft your position, feed it in, and demand counterarguments. That’s the version that builds thinking; the answer-key version slowly deletes it.

Two professionals collaborating on a whiteboard with user flow diagrams and key metrics for business growth.
A yes-man is a terrible study partner, so hand the model your reasoning and make it attack.

Why AI agrees with you by default

Here’s the failure mode nobody warns you about. Models are trained to agree, so a naive \”am I right?\” tends to produce agreement by default. Georgetown Law’s Tech Institute bluntly describes this tendency as a yes-man problem, and researchers call the general phenomenon sycophancy. A yes-man is a terrible study partner. The New York Times documented extreme cases: ChatGPT told one user he was trapped in a simulated reality straight out of The Matrix, and told another he’d made a world-changing math discovery.

Those are extremes, sure. But the mechanism scales all the way down to your Tuesday afternoon work session: the model agreed instead of pushing back. If everyone asks the same agreeable oracle, everyone starts getting the same answers, and that’s the quiet loss.

Copy-ready challenge prompts

Steal these, straight from Project FAILSafe and APA guidance:

  • “Which counterarguments haven’t I considered?”
  • “Find the weak spots in my explanation.”
  • Offer alternative interpretations and help me weigh them.

The evidence backs the framing, not just the vibe. Gerlich (2025) found that structured prompts requiring active reasoning measurably reduced cognitive offloading and improved engagement versus unguided use. Ask the AI to make you think, and you think more. And de la Puente et al.

(2024) found students using ChatGPT in international relations debate exercises improved argumentation versus a control. Arguing with the model beats taking notes from it.

One widely reported pattern worth knowing in advance: your first critique request will often come back as polite agreement. Force the adversarial framing, explicitly, and it fixes it. This is also part of the editorial practice on my end; before publishing an argument, I’ve started running my draft through those challenge prompts, and the difference between the first polite pass and the second adversarial one is genuinely funny.

What to offload and what to keep: the division of labor

Offload storage and routine drafting; never offload idea generation, sense-making, or judgment. The grounding: adults hold roughly three to five chunks in working memory, according to Morra, Patella, and Muscella (2024), so delegating the holding to AI is cognitively legitimate extra RAM, not laziness. A therapist letting AI handle documentation and session-data organization is the pattern done right: the machine holds the paperwork so the human can focus on judgment in the room. The same framing gives you a concrete split for any task. AI holds the spec, the notes, the raw material; you hold the decision. Warner’s ?50% boundary on complex projects fits here too, but it gets fuller treatment below.

Stay the last filter: verification and accountability habits

Keep AI to a helper’s share of the work (Warner’s sourced guideline: no more than 50% on complex projects) and require yourself to articulate anything AI produced in your own words. That articulation test is what keeps your voice and judgment intact, because output quality can improve while the underlying skill quietly declines. The familiar advice holds up here too, and it’s worth repeating: write your own first draft, then let AI polish. Fan et al.’s essay experiment showed why the order matters: the copy-paste group produced great essays and learned nothing.

Now the stories, because they make the habit concrete.

Thomas Courtney, a sixth grade teacher in San Diego, tells it best. One student misused AI to submit more than 40 AI-generated poems, and Courtney’s response wasn’t a ban; it was a rule: students must talk or write about anything AI produced. He calls students the last filter, meaning the human is the final checkpoint in the pipeline, and AI is a tool they used, not a replacement for their decisions. And those transparency habits apply to you too, not just in classrooms: keep the chat transcripts, write self-worded summaries of AI feedback, and explicitly note where you agree or disagree with it. If you can’t explain it, you didn’t earn it. Rubber-ducking your own output, basically.

Fred Young has a paired habit he gives students via analogy: AI is a measuring tape: it can get you to the answer, but the measuring itself is still yours to do. In practice it’s three steps, not a protocol: read the response, check it, make sure you understand it.

And Julie York, who teaches at South Portland High School in Maine, has the best mental model I found anywhere:

Her class found the limits of that assistant the weird way, during a Spanish telenovela video project with a crime scene. AI video models Veo, Sora, and Kling all refused to generate a dead body, and York ended up with a barely-breathing body and an ax coming out of it after trying three models. Here’s the part that should genuinely bug you as a user: none of the models disclosed those constraints.

Even the tools’ limits are hidden from you. That’s why “read, check, understand” isn’t optional; you’re the only verification layer that exists.

If the muscle has already weakened: rebuilding the first-draft habit

The fix is deliberately re-imposing the sequencing the harm studies show matters: your own unassisted first draft, then structured AI use. This isn’t paranoia; the Bastani data makes the recovery necessary. GPT Base users scored 17% lower than the no-AI control once the tool was gone. Cognitive debt, as one MIT-affiliated study framed it, defers mental effort now and pays for it later in diminished critical inquiry and increased vulnerability to manipulation. Draft-first sequencing has been stable advice across every credible source I found, and the Bastani retention hit is the proof of cost.

The urgency is real because abstention already lost. EdWeek Research Center found K-12 teacher AI use nearly doubled in two years: 34% in 2023, 61% in 2025. On the student side, 86% of students across bachelor’s, master’s, and doctoral programs reported using AI for coursework in 2024 (Digital Education Council), and HEPI found 92% of undergrads in 2025, mostly to clarify concepts (58%) and summarize articles (48%). ChatGPT’s 2022 launch made this inevitable in about a semester. The real choice is guarded versus unguarded use.

So here’s the recovery playbook, for when the first-draft muscle has gone soft:

  1. Schedule timed, unassisted first drafts. The dev analog is the one every coder already knows: writing it yourself teaches more than heavy AI reliance ever will. You’ve felt the difference between shipping and understanding.
  2. Re-impose think-first sequencing on existing workflows. Take the tasks where you currently open the AI tab first and flip the order: rough version, then AI.
  3. Teach-back reflection. Explain the results to someone else, Warner’s third step, rubber-duck style. If you can’t explain it, you didn’t learn it.
  4. Keep the reasoning active with structured adversarial prompts. Gerlich (2025) showed these measurably reduce cognitive offloading versus unguided use.

And here’s the beat that convinced me the guardrails are general, not education-specific. The same partner-not-replacement thesis shows up independently everywhere: Kristen Love’s classrooms (AI should never replace student thinking, it should expand how students demonstrate their learning), in university guidance, and at Phoenix Financial, where roughly 100 AI Champions ran an internal hackathon that turned ~90 ideas into 18 working AI agents built by business employees, on the explicit framing that AI is a thought partner, not a thought replacement. Different domains, same guardrails. That’s what a real pattern looks like.

The anti-atrophy checklist

Six checks, each tied to its evidence, run them before and during any AI session:

  • Ask “is AI doing the thinking, or helping me think?” before each session (APA self-check)
  • Draft your own rough answer before prompting (van Brunt / Project FAILSafe)
  • Use adversarial critique prompts, never “am I right?” (Georgetown’s yes-manism warning)
  • Cap AI at ?50% of the work on complex projects (Warner)
  • Explain it back to someone (teach-back test)
  • Audit your mode: constructive and interactive beats passive (ICAP, Chi & Wylie 2014)

AI is a cognitive amplifier if you stay intellectually in charge. The question is how, not whether.

Frequently Asked Questions

Does using AI make you dumber or can it make you smarter?

It depends entirely on the mode of use, not the tool. In Bastani et al.’s 2025 PNAS field experiment, students using unguarded ChatGPT scored 17% below a no-AI control once the tool was removed, while the same underlying model with tutoring guardrails boosted performance 127% with no post-use harm. Used constructively, a meta-analysis of 51 studies found a learning gain of g=0.867.

How to use AI without becoming dumber?

Keep the expensive cognitive steps — idea generation, sense-making, and judgment — human-owned, and offload only storage and routine drafting. Working memory holds roughly three to five chunks, so delegating the holding to AI is legitimate extra RAM, not laziness. Then stay the last filter: read, check, and make sure you understand every AI response, because you’re the only verification layer that exists.

How much should I rely on AI for writing without losing my own voice?

Keep AI to no more than 50% of the work on complex projects, write your own first draft, and let AI polish rather than originate. The articulation test is what protects your voice: require yourself to express anything AI produced in your own words, because output quality can improve while the underlying skill quietly declines.

Leave a Comment