Kali Trzesniewski · Research program

The research program · 01

Correct the record

The most talked-about studies cannot support their headlines — and the counter-evidence is not being read.

The field is moving fast and obscuring what we are learning. Since 2023, studies of AI and learning have appeared faster than the field can absorb them: thousands of papers, dozens of meta-analyses, and already meta-analyses of the meta-analyses, all built on a few years of work. Speed is not the problem. The problem is what the speed hides. The summaries pool studies that differ in the tool, the task, the outcome, and in what “using AI” even meant, and some of what gets pooled is not about the tools being argued over at all: chatbot syntheses without a single Generative AI (GenAI) study, “generative AI” categories that exclude ChatGPT. An average of things that are not alike comes out precise and settles nothing. One of the most prominent of these summaries, a Nature-family meta-analysis reporting a large ChatGPT learning effect (Wang & Fan, 2025, Humanities and Social Sciences Communications), collected about three hundred citations and was retracted. Retraction corrected the record. It did not correct the story.

Four headlines, and the designs behind them. Evaluating four studies that made big headlines reveals the gap between the headline that traveled and the conclusions the design could support.

The headline that traveled Why the design cannot say that
“AI makes students metacognitively lazy”Fan et al., 2025 · 1,018 citations The coding scheme filed every minute of AI conversation into a category where metacognition cannot be counted, the dialogue logs that could have shown it were collected and never analyzed, and the authors' own limitations say the construct was never measured. What remains is click-patterns from one hour of revising one essay, 117 students.
“AI is eroding your brain”Kosmyna et al., 2025 · preprint · 1,367 citations The claim is about lasting change; the design measures brain activity during tool use. No learning, memory, or transfer outcome exists anywhere in the study, and “accumulation” rests on a sub-study of eighteen people the authors call preliminary. The group that wrote unaided first, then used AI, showed the opposite pattern, and the authors have objected to the brain-rot coverage.
“AI harms learning”Bastani et al., 2025 · 314 citations The best-executed of the wave — preregistered, randomized, an unassisted exam. It compared a bare chatbot with a tutor-style version built with guardrails, but “without guardrails” was one of five simultaneous differences between those conditions, so nothing in the result attributes anything to guardrails. The harm itself: a few points on next-day near-copies of practiced problems from the bare version, none from the tutored one.
“Brief AI use makes people give up”Liu et al., 2026 · preprint · 13 citations in under four months The design removed the motivation persistence requires: adults doing arithmetic drills for a flat fee, told in writing that pay did not depend on correctness. Its own cleanest experiment found no effect on giving up. What replicates — people handed free answers practiced less and solved fewer problems minutes later — is a fact about how the tool was built, not about what people are like.

Citation counts (roughly, how many other papers have picked a study up) as of July 2026.

Four headlines, one shared failure. These studies are not measuring what they believe they are measuring — bundled treatments, constructs named but never measured, performance with the tool scored as learning without it — every one a design that cannot answer its question. That is not a storytelling problem; it starts before the first participant is recruited — and only standards reach that far upstream.

The attention runs against the evidence. “Coach not crutch” (Lira, Rogers, Goldstein, Ungar & Duckworth, 2025, preprint) measured skill the way the studies above are faulted for not measuring it: unassisted, a day later, with real incentives, across three preregistered experiments and roughly 7,200 adults. Held to the same standard as the table, it is the only study here whose design can answer its question — it isolated the mechanism, and the mechanism corrects its title: seeing one high-quality worked example carried the benefit, not the coaching. It has been public since February 2025. It has been cited three times. The counter-evidence is not losing an argument; it is not being read.

Where the record actually stands. There is little rigorous evidence that GenAI improves durable, independent learning — and little that it harms it — because durable, independent learning is what almost no one measures. The one design above that did measure it found the benefit in the worked example, not the tool. The widest lens shows the same gap at scale: a preregistered synthesis of 45 meta-analyses (40 years of AI-in-education research, roughly 197,000 learners) finds a medium average benefit on the outcomes those literatures do measure, and no significant differences in effect across types of AI, with GenAI and ChatGPT running descriptively smaller, not larger, than the broader AI tools and chatbots that came before; reported effects grow with publication year while nearly half the meta-analyses are underpowered (Emslander, Lindner, Eitel, Kasneci, & Bardach, 2026, preprint). The harm story is on the same footing: three universities' administrative records have now been used to test whether grades rose in GenAI-susceptible courses after ChatGPT arrived, and the answer flips with the design: grades rose at the first two (Hausman, Rigbi, & Weisburd, 2025, working paper; Chirikov, 2026, working paper); the third, with the tightest design of the three (susceptibility anchored to pre-GenAI syllabi and validated against human coders, course fixed effects, 138,000 students), finds no effect (Zumel Dumlao et al., 2026, preprint). None of the three is peer-reviewed yet, and all three measure availability and grades: none measures how students used the tool, or what they learned. That rules out both prevailing stories, AI as miracle tutor and AI as cognitive decay, as evidence-based positions. It rules out use-amount, how much rather than how, as a goal in either direction: neither story's evidence ties amount to durable learning. And it rules out waiting for more of the same designs to settle the question.

What this feeds. A narrative: that AI happens to people — the active variable in every design. The question becomes what AI does to them, never what they could learn to do with it. We teach reading; we do not run studies of learning with books against learning without. The research designs, the headlines, and the policies all inherit the applied-to frame.

Where this leads if the record stands

Into how people see themselves and each other. The stories are already inside students' heads, and they are measurable there — the same internalization my colleagues and I documented when a generation absorbed the Generation Me story, the claim that young people were uniquely narcissistic and entitled (Trzesniewski & Donnellan, 2014). In my current college sample, about one in three students agrees there is not much point teaching students to use AI carefully. Nearly half agree that people who use chatbots heavily will end up letting the chatbot do most of the thinking, no matter how good their intentions — the inevitability belief. About one in five say their ability to resist letting AI do the thinking is something basic about them that they cannot change (Trzesniewski, Gripshover, & Master, unpublished data). Those are beliefs about whether trying is worthwhile. Separate from them, about three in ten report actually finding it difficult to keep AI from taking over their thinking, reports that track behavior under pressure rather than the inevitability belief. Whether the narrative causes the beliefs, or the struggle does, decides what fixes them.

Into policy. Bans, detection regimes, and locked-down systems are the story's natural policies — and they push use underground rather than out of existence: unsupervised, unmentored, spread across whatever model quality a family can reach. The students with coaches at home lose least. In my adolescent sample, the climate of AI-cheating accusations already lands hardest on lower-income students — nearly twice as likely to be accused of AI use they did not commit (44% vs 25%; see The measures). A record corrected too late is not neutral; it is regressive. And there is no social vaccine here, in either direction: no one-shot exposure, and no one-shot ban, turns every student into an effective user. The work is learning how to help everyone.

The next step: keep auditing, extract the lessons, fix it upstream. The evaluations continue, study by study, because that is where the standards come from: every failure in the table is a recurring one, and a failure that recurs can be written into a checklist and prevented where it starts, in how studies are designed before they run and how they are reviewed before they spread. The failures cluster in a seam between research communities — the AI-and-computing community on one side, education and psychology on the other — where each community's reviewing reflexes cover only half the pitfalls; standards and review training built for that seam are how the record stops needing correction one paper at a time.

The task ahead

Clarify the primary questions and constructs. What exactly the field is trying to learn about knowledge, skills, and use, stated before the comparison is chosen. The table's recurring failure — constructs named but never measured — starts here: nothing was defined well enough to measure. A commentary in preparation at Educational Researcher (Trzesniewski, Gripshover, & Master) makes that argument; the definitions live and grow with the science (see The measures).

Test whether the narrative itself is doing the damage. If the stories cause the defeatism, that is a self-fulfilling effect a correction can undo. If the beliefs instead reflect genuine struggle, that is a skill gap education can close. The two need separating, because they call for opposite responses. College pilot data already suggest they are separable (see The measures). And beneath both sits a missing piece: the field has no shared definition of effective use, so people have little way to judge when they — or anyone else — are using AI well, and in a climate where nearly all use is treated as cheating, use itself reads as the offense. The design follows: estimate what narrative exposure moves; find the predictors of stronger and weaker beliefs about controlling one's offloading; test whether a clear definition of effective use sharpens how people evaluate their own and others' use; then build to move the beliefs. The bet is mine and it is testable: teaching effective use should raise those control beliefs — and a null there means the belief is not the lever and the premise needs revising.

Systemize the standards. Measurement and research-design standards exist, but they sit siloed by field, and none have been systematically applied to GenAI-and-learning research. The work is consolidation: extract what transfers from the standards psychology, medicine, and education already have — JARS-Quant, CONSORT-AI, PRISMA 2020, What Works Clearinghouse review practice — add what GenAI needs (e.g., model and version, interaction logs, assisted performance separated from independent learning), and win the support of the experts and organizations that make standards stick. One standard, accessible across disciplines: an AI researcher with no experimental training has something to reach for when designing a study, and a psychology reviewer with no AI experience knows what to look for when reviewing one.

What would narrow this critique. Studies comparing learning with and without AI that state the question they answer, add taught-use conditions, measure delayed independent learning, and still produce stable, general answers. The critique is a claim about designs, and designs can prove it wrong.

How I use GenAI to support this work. The barrier here is volume: more studies than any reader can appraise. I asked Claude how to create a self-updating literature knowledge system. Through its guidance, I learned to move from using Claude as a chatbot to directing it through Claude Code — a version of Claude that works from my written instructions and can use the files and tools on my computer. This changed a lot about what I can accomplish now with GenAI.

The tools I built: I now have a knowledge base I can draw from for looking up that article I read that one time — the one with the person from ________ university and the anecdote about bread. Or that blog I read six months ago and remember a few keywords from. And I can use it in the flow of writing and analyses to look up references.

It has since grown to include Skills — written instructions Claude saves and follows for a recurring task — that I have trained Claude in (very similar to training a new student): conducting literature searches, verifying citations, and evaluating design and measurement against the research questions and conclusions. It does this by launching Claude Agents — separate copies of Claude, each given one focused job — trained to wear different reviewer hats and focus on those areas of expertise (e.g., review this as an expert in experimental design would, or as a motivation researcher would). By focusing an Agent on one body of knowledge, it does a much more thorough evaluation than when it is trying to evaluate 10 characteristics and knowledge bases at the same time.

A struggle I faced: the GenAI models inherited a bias the scientific field carries — a randomized trial is gold standard and always gets an A grade. By training the specialized Agents, I was able to teach Claude the Skills to evaluate based on fit with the research question, sample, and design instead of leading with “it is a randomized trial.”

References