THE ARTICLE · 14 MIN
Every tool that does some of our thinking for us brings the same worry: are we doing less of it ourselves? This page looks at what critical thinking is, whether it can be taught, what employers say about it, and what the research on AI does and does not show. Then it turns what holds up into habits you can try.
What critical thinking means
As an educational goal, the idea goes back to the American philosopher John Dewey (1910), who more commonly called it “reflective thinking”. He defined it as “active, persistent and careful consideration of any belief or supposed form of knowledge in the light of the grounds that support it, and the further conclusions to which it tends.”
Between 1988 and 1989 a panel of 46 scholars, educators and others, brought together at the request of a committee of the American Philosophical Association, agreed a longer definition, published in 1990. Critical thinking, they wrote, is “purposeful, self-regulatory judgment which results in interpretation, analysis, evaluation, and inference, as well as explanation of the evidential, conceptual, methodological, criteriological, or contextual considerations upon which that judgment is based.” Only one of the panel declined to be listed as supporting the final document.
The philosopher Robert Ennis puts it more briefly, in his 2011 wording: “Critical thinking is reasonable and reflective thinking focused on deciding what to believe or do.”
The definitions do not fully agree. Some limit critical thinking to forming a judgment; others, in the words of the Stanford Encyclopedia of Philosophy, “allow for actions as well as beliefs as the end point”. That matters for what follows, because the studies people quote do not all measure the same thing. One of the AI studies below defines critical thinking through Bloom’s taxonomy of learning goals, and its authors say plainly that “this definition of critical thinking is not uncontested.”
Can it be taught?
On average, a little. A 2008 review by Philip Abrami and colleagues pooled 117 studies with 20,698 participants and found an average effect of 0.341 standard deviations, with results that varied a great deal (“the distribution was highly heterogeneous”). The authors concluded that improvement “cannot be a matter of implicit expectation” and that “educators must take steps to make CT objectives explicit in courses”.
A 2015 review by the same group, drawing on 341 effect sizes from experimental and quasi-experimental studies, found a similar average of 0.30 and concluded that there are effective strategies “at all educational levels and across all disciplinary areas.” Three things stood out: “the opportunity for dialogue, the exposure of students to authentic or situated problems and examples, and mentoring”.
Our reading: an effect of about a third of a standard deviation is a real but modest average gain, not a transformation. The same review found that teaching critical thinking on its own and inside subject lessons together did better than either alone, but, as the Stanford Encyclopedia notes, that difference “was not statistically significant; that is, it might have arisen by chance”, and most of the studies lacked long-term follow-up.
A 2016 review by Huber and Kuncel found that critical thinking skills and dispositions “improve substantially over a normal college experience”, but that curriculum-wide programmes to improve them “do not necessarily produce incremental long-term gains.”
The harder question is whether the skill travels from one subject to another. The cognitive scientist Daniel Willingham argued in 2007 that critical thinking “is not a set of skills that can be deployed at any time, in any context” and that it is “very much dependent on domain knowledge and practice.” He did not say it never carries over: “such transfer does occur”, he wrote, but when and why is complex. The Stanford Encyclopedia describes the evidence for one popular answer, teaching critical thinking across many subjects with explicit attention to the skills they share, as “scanty”.
Our reading: teaching helps on average, and knowing a subject well is what lets the habit work in that subject.
What employers say
The line “critical thinking is the No. 1 skill employers want” is usually credited to the World Economic Forum’s Future of Jobs reports, and the Forum’s own summary of its 2020 report said that critical thinking and problem-solving “top the list”. These are surveys of employers, what companies say they expect, not measurements of what workers can do.
- In the October 2020 report, “critical thinking and analysis” comes first among the skill groups employers saw rising in importance; the report says it and problem-solving “have stayed at the top of the agenda”. In the same report’s list of the top 15 individual skills for 2025, it ranked fourth, behind analytical thinking and innovation, active learning, and complex problem-solving.
- The May 2023 report, based on 803 companies, says “analytical thinking and creative thinking remain the most important skills for workers in 2023.”
- The January 2025 report, based on over 1,000 employers surveyed in late 2024, says “analytical thinking remains the most sought-after core skill among employers, with seven out of 10 companies considering it as essential in 2025.”
We searched the text of the 2023 and 2025 reports and did not find the phrase “critical thinking” in either. So the claim fits one 2020 chart of skill groups; in the ranked lists of individual skills, a close cousin, analytical thinking, comes first.
A second statistic often travels with these reports: “65% of children entering primary school today will ultimately end up working in completely new job types that don’t yet exist.” The Forum’s 2016 report gives it as “one popular estimate”, and its endnote gives a web address for Shift Happens, not a study. We could not find a study behind it. An author who used the figure in a 2011 book wrote in 2017 that it “didn’t originate with me”, that she had traced it to an Australian website that no longer existed, and that she now regards it as “a pretty vague figure”.
AI and thinking: what the studies found
Our guide to reading a study explains why a survey, a preprint and an experiment answer different questions.
A survey of knowledge workers
Researchers at Microsoft Research and Carnegie Mellon University surveyed 319 knowledge workers, who shared 936 examples of using generative AI at work. The paper, presented at the CHI 2025 conference, found that “higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking”, and that AI shifted critical thinking “toward information verification, response integration, and task stewardship.”
The paper’s own title calls the reductions in effort “self-reported”. Its authors note that participants “occasionally conflated reduced effort in using GenAI with reduced effort in critical thinking with GenAI”, and that the sample “was biased towards younger, more technologically skilled participants”. It shows an association in what people reported, not a cause.
A survey of 666 people
A 2025 paper in the journal Societies surveyed 666 people in the UK, recruited online through social media, interviewed 50 of them, and found “a significant negative correlation between frequent AI tool usage and critical thinking abilities, mediated by increased cognitive offloading.” Critical thinking was measured partly by self-report. The author lists “reliance on self-reported measures and the potential for sample bias” among the limits, and says experiments “could offer causal evidence”. A correction in September 2025 replaced a duplicated table; the author states that the scientific conclusions are unaffected.
The brain-scan preprint
The study behind many “ChatGPT and your brain” stories is a preprint from the MIT Media Lab, first posted in June 2025. In its words, “a total of 54 participants took part in sessions 1-3, with 18 completing session 4”: they wrote essays with ChatGPT, with a search engine or with no tools while their brain activity was recorded by EEG. The authors report that “LLM users displayed the weakest connectivity.”
As of 1 October 2026 its arXiv record shows a preprint, last revised on 31 December 2025, with no journal publication listed. Other researchers have posted a constructive comment raising “the limited sample size”, reproducibility and the EEG methods. The authors’ own FAQ answers the question of whether it shows that LLMs make us “dumber” with “No!” and asks people not to use words like “brain rot”.
The strongest experiment
The clearest causal evidence among these studies comes from a field experiment with nearly 1,000 students at a high school in Turkey, published in PNAS in 2025. While students had GPT-4 during practice, their grades rose (“48% improvement in grades for GPT Base and 127% for GPT Tutor”). When access was taken away, the students who had used the plain tool did worse than those who never had access (“17% reduction in grades for GPT Base”).
The tutor version, whose instructions asked it “to avoid giving them the answer and instead guide them in a step-by-step fashion”, largely avoided the drop. The authors’ explanation is that, without guardrails, students used GPT-4 “as a ‘crutch’ during practice problem sessions”. The study measured maths learning over a few sessions, not critical thinking in general, and the authors say they “focus on short-term outcomes”.
Experiments that point the other way
When the tool is built for learning, results can be positive. In a randomised trial in a Harvard undergraduate physics course with 194 students, each taking one lesson with an AI tutor and one in class, a custom tutor built on the course’s own teaching methods let students “learn significantly more in less time” than an in-class active-learning lesson (Scientific Reports, 2025).
Two 2026 studies add detail, one a preprint and one a CHI 2026 conference paper. In the preprint, access to a general AI tool raised undergraduates’ immediate test scores by 0.27 standard deviations, and the gains persisted a week later; essay gains a week later were larger among students who used AI “to explain concepts rather than generate text”. In the conference paper, an experiment with 393 people, having an AI model from the start “improved performance under time pressure but impaired it with sufficient time”.
What 15-year-olds told PISA
The OECD’s PISA 2025 results, published on 8 September 2026, add a large survey: “Students who say they did not use AI to draft text for writing assignments outperformed those who say they did.” Among frequent users, students who used AI to help them learn and were taught to assess the quality of AI-generated information “tended to perform better than those who received no such guidance.” These are associations, not proof that AI lowered anyone’s scores. The full results of PISA’s new Learning in the Digital World test are due in May 2027.
Our reading: surveys find that people who lean on AI more report, or score, less critical thinking, but they cannot say which causes which. The clearest experiment among them found that using AI to skip the work during practice left students worse off on their own, and that a tool built to make students do the thinking largely avoided that.
Over-trusting machines is an old problem
Researchers have a name for leaning too hard on a confident machine: automation bias, “the tendency to over-rely on automation”. A 2012 systematic review screened 13,821 papers and kept 74. It found that trust and confidence, workload, task complexity and time pressure all played a part, and that training and “emphasizing user accountability” helped.
In a 2023 experiment, 27 radiologists read mammograms with suggestions from what they were told came from an AI system; in 12 of 40 cases the suggestion was deliberately wrong. Inexperienced readers rated 79.7% of mammograms correctly when the suggestion was right and 19.8% when it was wrong; very experienced readers fell from 82.3% to 45.5%. The authors concluded that radiologists at every level of experience “are prone to automation bias when being supported by an AI-based system.”
Experience softened the pull without removing it. Why AI makes things up covers the errors chatbots make.
Checking what you read: what helps, and how much
The first approach is lateral reading: leaving a page to see what other sources say about it, the core of the SIFT method. In a 2022 field study in one urban school district, high-school students who had six 50-minute lessons (271 students, compared with 228 in regular classes, in a matched rather than randomised design) “grew significantly in their ability to judge the credibility of digital content”. As our SIFT page explains, a lasting benefit measured against a comparison group has not been established.
The second is prebunking, or inoculation: short games or videos that show people a manipulation technique before they meet it. In a 2022 study in Science Advances, five short videos improved recognition of techniques such as false dichotomies across six randomised experiments with 6,464 people. In a field test on YouTube the effect was smaller: recognition rose “by about 5% on average”, measured by a single test question within 24 hours of the advert. Effects also fade. In a 2021 study of the Bad News game, the benefit lasted at least three months when people were tested at regular intervals, but without regular testing it was “no longer significant” after two months.
Whether these lessons help people tell reliable from unreliable news, and not just name a trick, is disputed. A 2023 reanalysis of five studies found that two inoculation games “did not improve discrimination” but made people answer “false” more often to everything. A 2026 reanalysis of 33 experiments with 37,025 people found that games and videos “consistently improve discrimination between reliable and unreliable news, without increasing response bias.” Two of that reanalysis’s five authors also co-wrote the 2022 YouTube study; the 2023 critique was written by two researchers who were not among the authors of either. Our logical fallacies page weighs the two and leans, cautiously, toward the larger reanalysis, because it covers far more experiments. Whether any benefit survives a realistic feed has not been shown: in a 2025 study in PNAS Nexus, mixing real posts into a simulated feed “appeared to nullify the effect” of a lesson on emotional manipulation.
How to use it
- Doing the thinking first on practice tasks. In the maths experiment, students who used a plain AI tool to get answers during practice did worse on their own later. A 2026 conference paper found that, with enough time, people who started without the AI model did better than those who had it from the start; under time pressure the pattern reversed.
- Asking for hints and explanations rather than answers. The tutor that guided students step by step largely avoided the drop, and in a 2026 preprint students who used AI to explain concepts showed larger gains a week later.
- Checking AI output against a source. Automation bias catches experts as well as beginners. In PISA, frequent AI users who used it to help them learn and had been taught to assess its output tended to do better, though that is an association. Our page on spotting a fake quote or an invented source covers the checks.
- Leaving the page to check who is behind it. Lateral reading helped students judge credibility in the classroom study; whether the habit lasts has not been established.
- Learning the subject. Willingham’s point is that critical thinking depends on knowing the subject you are thinking about.
- Repeating the lesson. In the Bad News study, the effect held when people were tested regularly and faded when they were not.
Related: Think Again · Cognitive biases checked
Sources
- Stanford Encyclopedia of Philosophy, “Critical Thinking” (David Hitchcock), §§1, 11 and 12.
- P. A. Facione, Critical Thinking: A Statement of Expert Consensus for Purposes of Educational Assessment and Instruction (“The Delphi Report”), 1990, ERIC ED315423.
- R. H. Ennis, “The Nature of Critical Thinking”, University of Illinois, revised May 2011.
- P. C. Abrami et al., “Instructional Interventions Affecting Critical Thinking Skills and Dispositions”, Review of Educational Research 78(4), 2008.
- P. C. Abrami et al., “Strategies for Teaching Students to Think Critically”, Review of Educational Research 85(2), 2015.
- C. R. Huber and N. R. Kuncel, “Does College Teach Critical Thinking? A Meta-Analysis”, Review of Educational Research 86(2), 2016.
- D. T. Willingham, “Critical Thinking: Why Is It So Hard to Teach?”, American Educator, Summer 2007.
- World Economic Forum, The Future of Jobs (2016), The Future of Jobs Report (2020, 2023, 2025).
- C. N. Davidson, “How can you prepare for a career in an era of automation? BBC interview”, 1 June 2017.
- H.-P. Lee et al., “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers”, CHI ‘25.
- M. Gerlich, “AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking”, Societies 15(1), 2025, and correction, Societies 15(9), 2025.
- N. Kosmyna et al., “Your Brain on ChatGPT”, arXiv:2506.08872 (preprint), and the authors’ FAQ; M. Stanković et al., “Comment on: Your Brain on ChatGPT”, arXiv:2601.00856 (preprint).
- H. Bastani et al., “Generative AI without guardrails can harm learning: Evidence from high school mathematics”, PNAS 122(26), 2025.
- G. Kestin et al., “AI tutoring outperforms in-class active learning”, Scientific Reports 15, 2025.
- Z. Contractor and G. Reyes, “Experimental Evidence on the Learning Impact of Generative AI”, arXiv:2607.08849 (preprint, 2026).
- J. Zhi, H. Kumar and M. Lee, “Investigating the Effects of LLM Use on Critical Thinking Under Time Constraints: Access Timing and Time Availability”, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (April 2026); also arXiv:2603.08849.
- OECD, “PISA 2025: Students’ reading and mathematics performance declined sharply across the OECD”, press release, 8 September 2026.
- K. Goddard, A. Roudsari and J. C. Wyatt, “Automation bias: a systematic review of frequency, effect mediators, and mitigators”, JAMIA 19(1), 2012.
- T. Dratsch et al., “Automation Bias in Mammography”, Radiology 307(4), 2023.
- S. Wineburg et al., “Lateral reading on the open Internet: A district-wide field study in high school government classes”, Journal of Educational Psychology 114(5), 2022.
- J. Roozenbeek et al., “Psychological inoculation improves resilience against misinformation on social media”, Science Advances 8(34), 2022.
- R. Maertens et al., “Long-term effectiveness of inoculation against misinformation”, Journal of Experimental Psychology: Applied 27(1), 2021.
- A. Modirrousta-Galian and P. A. Higham, “Gamified inoculation interventions do not improve discrimination between true and fake news”, Journal of Experimental Psychology: General, 2023.
- A. Simchon et al., “A Signal Detection Theory Meta-Analysis of Psychological Inoculation Against Misinformation”, Current Opinion in Psychology 67, 2026.
- S. Y. N. Wang et al., “Limited effectiveness of psychological inoculation against misinformation in a social media feed”, PNAS Nexus 4(6), 2025.
Checked October 2026. What we read: the Stanford Encyclopedia entry; the Delphi Report, Ennis’s outline and Willingham’s article in full; the four World Economic Forum reports (searched as text); the full texts of the Lee, Gerlich, Bastani, Kestin and Roozenbeek papers and the Gerlich correction; the OECD’s PISA 2025 press release; and the abstracts of the other papers and preprints, with the Kosmyna authors’ FAQ. What we could not read: the full texts of the Abrami and Huber & Kuncel reviews (so their effect sizes are taken from the abstracts), the full PISA 2025 report, the full texts of the two 2026 studies, and the BBC programme on the 65% figure. If you can show any of this wrong, with a source, we want to see it.
- critical thinking
- ai
- learning
- misinformation
- research
