ChatGPT Raised Student Scores by 0.86—Causal Training Drove Originality

OpenAI’s August 27, 2026 account of a randomized Bocconi University experiment found two complementary effects among more than 1,000 first-year undergraduates: ChatGPT improved evaluated quality, coherence and similarity to expert recommendations, while causal-reasoning training produced a wider range of ideas. Students who received both interventions retained benefits associated with each.
Bocconi’s August 27 report on the 1,053-participant study states that ChatGPT access raised the five-point evaluation by about 0.86 points from an estimated control score of 2.09. The result does not show that AI improved critical thinking generally: evaluator scores and idea originality were separate outcomes, and the latter was driven primarily by the reasoning exercise.
Four groups separated AI access from reasoning training

The preregistered 2×2 experiment assigned class sections within degree programs to one of four conditions: neither intervention, causal-reasoning training, ChatGPT Edu access using GPT-4o, or both. The CEPR discussion-paper record identifies the participants as first-year economics, management and finance students and the common assignment as a real-world merchandising business problem.
Students in the reasoning condition completed a structured exercise about causal links, underlying mechanisms and the circumstances in which an explanation could fail. The intervention was separate from AI instruction. This design allowed the researchers to estimate the effect of ChatGPT access, the effect of causal training and whether combining them changed either result.
Every participant then prepared a consultant-style recommendation for increasing alumni awareness and use of university merchandise. That matters for interpreting the findings: this was a bounded business assignment with established marketing objectives, not an examination of general academic achievement or a test of what students could do later without AI.
The outcome matrix shows why quality and originality diverged

- Evaluator score: Trained human graders assessed how well each recommendation addressed the specified awareness-and-usage objectives. ChatGPT produced the clear gain on this conventional measure; causal training did not.
- Expert similarity: Researchers separately compared student submissions with recommendations prepared by domain experts. AI-assisted work moved closer to those expert responses, a result distinct from the human grading score.
- Coherence: ChatGPT improved the logical connection between claims and recommendations. The combined condition also strengthened coherent logic.
- Number of ideas: AI access increased how many central proposals appeared in a submission. Causal training also increased idea count, but this was not its main contribution.
- Idea diversity: Causal training increased the variety of ideas within submissions and made students’ central proposals more distinct from those of their peers. ChatGPT alone increased idea volume without clearly widening the collective solution space.
The distinction between idea count and idea diversity is central. A response can contain several proposals while remaining close to conventional recommendations shared across the class. Conversely, a submission can introduce a less common approach without satisfying a rubric designed around standard answers.
“Originality” in this experiment therefore refers to measured variation and distance among students’ ideas. It was not a separate expert judgment that every unusual proposal was creative, workable or commercially valuable.
Higher evaluations did not demonstrate broader thinking
Random assignment supports a causal interpretation within the conditions of the experiment. ChatGPT improved performance against the specified rubric, and AI-assisted responses were more coherent, included more ideas and resembled expert recommendations more closely. Coherence and idea count explained about half of the estimated ChatGPT advantage, leaving part of the score gain associated with other measured or unmeasured features of the submissions.
Causal training changed a different layer of the work. Students exposed to it used more mechanism-based reasoning, considered why a proposal might work and produced less conventional combinations of ideas. Those changes did not translate into higher evaluations under the study’s standard marketing criteria.
This is why the experiment does not support a simple verdict that ChatGPT either helped or harmed original thinking. AI improved polished, expert-like performance without being the main source of between-student diversity. The reasoning intervention expanded the range of solutions without receiving a comparable reward from the evaluation rubric.
The scoring system is part of the finding, not merely a neutral measurement device. When assessment criteria reward established objectives and recognizable solutions, work that moves farther from the typical answer space may receive no additional credit even when it demonstrates a broader search for alternatives.
The combined intervention preserved breadth, within narrow limits

Students receiving both interventions showed the quality-related advantages associated with ChatGPT while preserving the idea variety associated with causal training. Their rubric scores and idea counts were similar to those of the AI-only group, while their diversity resembled that of students who completed the reasoning exercise. The combined treatment also strengthened some measured features of causal and coherent logic.
The experiment does not establish that ChatGPT improves learning, critical thinking or originality across subjects. It shows that access to GPT-4o did not erase the immediate effect of a causal-reasoning exercise during one structured business assignment. The researchers evaluated submitted recommendations, not later unaided performance, conceptual mastery or long-term retention.
Generalization is also constrained by the task: a short marketing problem with explicit objectives and conventional performance criteria. Less structured questions, unfamiliar domains, longer projects or grading systems that explicitly reward novelty could produce a different balance between polished answers and diverse thinking.
As of the August 27 release, the supported conclusion is deliberately narrow. ChatGPT raised conventional evaluations in this randomized experiment, while causal-reasoning training was the primary driver of measured originality; whether those complementary effects persist across other assignments or over time remains unanswered.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.