Back to News
InnovationAI Understanding briefing

Bocconi experiment finds ChatGPT and critical-thinking training improve different aspects of student work

A randomized experiment involving more than 1,000 Bocconi University students found that ChatGPT access improved the quality and coherence of business recommendations, while causal-reasoning training produced more original ideas.

By 5 min read
Primary-source image accompanying Bocconi experiment finds ChatGPT and critical-thinking training improve different aspects of student work
The short version

A randomized experiment involving more than 1,000 Bocconi University students found that ChatGPT access improved the quality and coherence of business recommendations, while causal-reasoning training produced more original ideas.

What happened

Researchers at Bocconi University, working with OpenAI Economic Research, randomly assigned more than 1,000 first-year students to receive ChatGPT access, causal-reasoning training, both, or neither. Students completed a marketing case for the university merchandise store. According to the source, ChatGPT access raised rubric scores by almost one point on a five-point scale, while critical-thinking training increased idea variety and originality.

OpenAI’s account, dated August 27, 2026, describes a randomized experiment conducted by researchers at Bocconi University in collaboration with OpenAI Economic Research. More than 1,000 first-year undergraduate students worked on a real-world business case involving marketing recommendations for Bocconi’s merchandise store. Students were assigned by class period to one of four groups: access to ChatGPT using GPT-4o, training in causal reasoning, both interventions, or neither. The source presents the design as a way to separate the effects of ChatGPT access, critical-thinking instruction, and their combination.

Student submissions were assessed by trained human graders using a five-point rubric. The rubric measured how well recommendations addressed two stated marketing goals: increasing awareness and increasing use of the university store. The source says students with ChatGPT access scored almost one point higher on that scale. Automated text analysis also found that their submissions contained more ideas, followed clearer logic, and were more similar to recommendations written by three experts. The source characterizes this as helping less-experienced students produce work that appeared more professional.

The causal-reasoning intervention had a different measured effect. The exercise was unrelated to AI and taught concepts through a game, examples, questions, and feedback. Students who completed it explained more clearly why their proposals might work and when they might fail, but did not score higher on the study’s main rubric. Automated analysis found that their ideas covered a wider range and were more distinct from those produced by peers. Students who received both interventions showed the combination of effects across the study’s measures, including stronger logical coherence, more ideas, and more evidence of questioning assumptions. The source does not state the experiment’s duration, detailed statistical results, participant demographics, or how students used ChatGPT during the assignment.

Read the source: openai.com

Why it matters

The experiment suggests that polished AI-assisted work and independent reasoning are not interchangeable outcomes. ChatGPT access improved performance on the study’s conventional grading measure, while causal-reasoning training helped students generate more distinctive ideas and explain when recommendations might succeed or fail. The findings also question whether grading only final answers can show what students understand.

The central implication is that a high-quality final answer may reflect more than subject knowledge. In this experiment, ChatGPT access improved the features that a conventional rubric rewarded: coherence, idea count, similarity to expert recommendations, and performance against specified marketing goals. That could make AI-assisted work look stronger even when the assessment does not reveal how much of the reasoning came from the student’s own prior expertise. The source says students still had to decide what to ask ChatGPT, evaluate its responses, and select material for their submissions, but it does not quantify those activities.

The study also separates originality from polish. Causal-reasoning training did not increase scores on the stated marketing rubric, yet it was associated with a broader and more distinctive set of ideas and clearer explanations of possible success or failure. That result matters for educators because an assignment can reward a well-structured conventional answer while missing whether a student considered multiple approaches or produced an idea that peers did not. The finding is presented as evidence that critical thinking can affect dimensions of performance that ordinary grading may not capture.

The combined group showed why the source describes the interventions as complementary rather than as alternatives. Students who received both displayed the wider idea variety associated with the exercise and the stronger rubric performance and idea counts associated with ChatGPT access. The source reports that this group also showed stronger logical coherence and more evidence of looking for explanations and questioning assumptions. These findings could inform assessment design, but they do not establish that ChatGPT improves learning in every context or that causal-reasoning instruction always increases originality. The experiment measured work on one business case, not long-term learning, retention, or independent performance without AI.

What to watch next

The study’s conclusions need to be interpreted within its setting: first-year students at one university completing one business case, with assignment to groups organized by class period. The source does not provide the paper’s full methods, participant demographics, training duration, ChatGPT-use records, or evidence from other subjects. Further research should test whether the pattern holds across disciplines and whether redesigned assessments can measure reasoning and originality more directly.

A key question is whether the same pattern appears outside a university merchandise-marketing case. The source gives no results for mathematics, science, writing, professional training, or other kinds of assignments. Researchers and educators would need comparable experiments across subjects, age groups, institutions, and levels of prior expertise before treating the findings as broadly generalizable. It is also unknown whether the benefits attributed to ChatGPT depend on GPT-4o specifically, on the way students were instructed to use it, or on the structure of the task.

Assessment changes are another practical issue. If students can submit polished, expert-like answers with AI assistance, educators may place more weight on drafts, oral explanations, source evaluation, reasoning records, or novel applications. Those approaches could make student understanding more visible, but the source does not test any redesigned assessment. It also does not report whether students’ final work contained factual errors, whether graders knew which students had ChatGPT access, or how automated measures of originality were validated against human judgments.

The study’s assignment structure deserves attention in future reporting and replication. Students were randomly assigned by class period, and the source does not describe how many classes were involved, how the groups were balanced, or whether classroom conditions differed. The source also leaves open how much time students spent with ChatGPT, what prompts they used, and how independently they completed the work. Follow-up research could examine longer-term effects, performance on unaided tasks, and whether training students to question AI outputs changes the quality and originality of their work. Until those questions are answered, the results support a targeted conclusion: in this experiment, AI assistance and causal-reasoning instruction improved different aspects of one student assignment.

Related guides & quizzes

ChatGPT & LLMsAI EthicsAI TrainingPrompt EngineeringTest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?