ChatGPT Improved Student Answers—but Causal Training Produced Rarer Ideas

On August 27, 2026, OpenAI and Bocconi University published findings from a randomized classroom experiment in which ChatGPT improved the evaluated quality and coherence of students’ marketing recommendations, while causal-reasoning training produced ideas that were less typical of their peers’ work. OpenAI’s account of the experiment describes the effects as complementary, but they were changes in different properties of one written assignment—not interchangeable evidence of learning or critical-thinking growth.
Bocconi’s August 27 report identifies 1,053 first-year economics, finance and management students and four conditions: ChatGPT Edu with GPT-4o, causal-reasoning training, both interventions, or neither. Every student completed the same task—recommending ways to increase alumni awareness and use of the university’s merchandise store.
Four groups separated AI access from reasoning training

The experiment used a preregistered 2×2 design, assigning class periods within academic programs rather than assigning each student independently. This let the researchers estimate the effects of being offered ChatGPT access, receiving causal training and receiving both together.
The causal intervention was a structured game about chains of cause and effect. Students were prompted to identify mechanisms, construct coherent causal explanations and specify conditions under which a claim might prove false. Comparison students played a placebo version with the same scenarios but without reasoning instructions or feedback; the training therefore did not teach students how to use ChatGPT.
The research paper’s experimental design allocated 249 students to control, 256 to causal training, 197 to GPT access and 351 to both; participants then had 45 minutes to write no more than 180 words, after which 20 master’s students rated the submissions, three raters assessed each response, and the texts were separately compared with recommendations from three domain experts.
The outcomes capture different kinds of performance

The study did not use a single definition of success. Human ratings assessed how well a recommendation addressed the predefined marketing goals, while separate text analyses examined coherence, causal structure, idea count, semantic diversity and similarity to expert responses.
- Evaluated answer quality: ChatGPT access increased the awareness-and-usage score by an estimated 0.862 points from an estimated control score of 2.09 on a five-point scale. This was performance on the assigned rubric, not a general measure of academic ability.
- Coherence: GPT access produced clearer logical structure. The combined condition added a further coherence effect, suggesting that the model could help express reasoning introduced by the causal intervention.
- Expertise-gap reduction: GPT-assisted responses were semantically closer to recommendations supplied by the three domain experts. That means the completed texts resembled expert output on this problem; it does not show that students acquired the experts’ knowledge.
- Causal reasoning: Training increased explicit mechanism identification and falsification logic. GPT alone did not generate clear gains on those two measures, although it strengthened them when paired with the training.
- Idea volume: GPT produced the largest increase in the number of core ideas. Causal training also increased idea count, but by less.
- Idea uniqueness: Causal training produced the strongest increase in semantic diversity within individual answers and was the only intervention with a clear positive effect on diversity between students’ responses. This measured distance from the sample’s typical ideas—the evidence behind the title’s “rarer ideas”—rather than verified novelty in the wider world.
The combined condition retained both patterns
Students offered both interventions kept the diversity associated with causal training while also receiving the main output benefits associated with ChatGPT. Adding GPT did not erase the causal intervention’s effect on diversity, and the combined condition strengthened some measures of coherent logic, mechanisms and falsification.
It was not an across-the-board victory for the combined group. Its conventional evaluation result was largely attributable to GPT access, with no statistically meaningful additional scoring effect from the interaction. GPT alone also did not increase between-student diversity, so more ideas and more varied ideas remained separate findings.
The experiment measured output, not durable learning

The direct evidence concerns written products created during one class exercise. The study did not test later recall, unaided performance, transfer to another subject or critical-thinking ability months after the intervention. It therefore cannot establish that ChatGPT taught marketing knowledge or that the short causal exercise produced lasting intellectual development.
The assignment also favored a language model. Awareness and usage are widely documented marketing concepts, the problem and rubric were tightly specified, and the required product was a short recommendation. The causal estimate for ChatGPT applies to being offered GPT-4o under these conditions, not to every model, discipline or assessment format.
“Originality” needs the same restraint. The analysis measured semantic distance among ideas within a response and between participants’ responses. An uncommon proposal can be distinctive without being feasible, useful or correct, none of which was established by the diversity measures.
The rubric rewarded alignment more readily than novelty
Causal training changed students’ reasoning without raising their conventional evaluation scores. The rubric rewarded recommendations aligned with the familiar goals of awareness and store use; it did not separately reward explicit mechanisms, falsifiable claims or distance from the common solution space. In the paper’s follow-up analysis, mechanism identification, falsifiability and between-response diversity were negatively associated with the rubric score.
The immediate implication for teaching and workplace assessment is about what a score represents. A polished, expert-like response can demonstrate strong task performance while revealing little about retention or independent expertise. Conversely, a more unusual answer can score poorly when the rubric values conformity to established criteria. The experiment supports measuring answer quality, causal reasoning and originality separately; evidence about lasting learning and transfer still requires other tasks and longer follow-up.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.