The skills that earn top grades are the ones AI can fake best

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 30, 2026 4 min read
The skills that earn top grades are the ones AI can fake best

A study at Bocconi University found that GPT-4o significantly boosted grades on a business assignment for over 1,000 freshmen.

Experiment design and results

In November 2025, researchers split 13 sections of an introductory management course into four groups. One group received no intervention. A second group took a short lesson on causal reasoning. A third group had access to GPT-4o. The fourth group received both the lesson and the AI tool.

Students had to write marketing recommendations for the university merchandise shop in up to 180 words. The lesson covered coherent causal logic, falsifiability, and how a proposed action might lead to a desired outcome.

Students with GPT-4o scored nearly a full point higher on a 1-to-5 scale than those in the control group. Their answers contained about two more ideas on average. The text was more logically coherent and matched the recommendations of three subject-matter experts more closely. Even after researchers controlled for argumentation quality, number of ideas, idea diversity, and text properties, a measurable GPT-4o advantage remained. The authors attribute this to higher content quality, not greater student knowledge.

The causal reasoning lesson did not raise traditional scores. Work from those students actually scored slightly worse on average. However, they more often explained why a proposed action should work and under what conditions it might fail. They also generated more diverse ideas that diverged from what their peers wrote.

Combining the lesson with GPT-4o did not add any further boost to traditional scores. On causal reasoning markers, the two approaches complemented each other. The greater idea diversity from the lesson group held up.

The grading rubric favored conventional answers

More ideas, more coherent arguments, and greater idea diversity within a single answer all correlated with higher scores. But stronger falsifiability, more detailed explanations of how proposed actions would work, and greater divergence from other students’ ideas correlated with lower scores.

That does not mean originality was penalized across the board. For this assignment, the traditional score mainly rewarded well-structured answers that stayed within the expected solution space. The authors conclude that diversity and originality need to be explicitly built into grading criteria if they are supposed to count.

Current grading systems measure polish, structure, and completeness, but not learning and understanding. That makes it easy to use AI as a cheating tool.

What the study does not show

There was no follow-up test where students had to demonstrate what they actually understood or retained without ChatGPT. The authors explicitly acknowledge that it is unclear whether the GPT advantage came from knowledge students actually gained or simply reflected better output with AI assistance. The study shows improved graded performance on this assignment, not improved learning.

The experiment only looked at freshmen at a single university working on a narrow marketing task. Randomization happened across 13 class sections rather than individually among all 1,053 students.

The main performance score came from human graders who did not know which group each text belonged to. Several additional measures like causal reasoning and idea diversity were evaluated using models from OpenAI and Anthropic.

OpenAI provides the technology being studied and was involved in the research. Several authors work at OpenAI or were employed there during the study.

Other studies ask what sticks after the AI goes away

Research so far clearly suggests that what matters is not whether students use AI but whether it supports their own thinking or replaces it. When it replaces it, students suffer.

A study covering more than 500,000 US college grades found that top grades after ChatGPT’s launch increased most in writing- and programming-heavy courses with a large homework component. Controlled experiments found that after AI was taken away, participants performed worse than a control group. The drop was steepest among those who had mainly used GPT to get direct answers.

A 30-month study of more than 26,000 students in China showed a similar pattern over a longer period. Homework got better and faster with AI, but exam scores dropped. On later entrance exams, results were 18 to 24 percent lower over the long term. Students who spent roughly the same amount of time on homework as non-users despite having AI access did not show comparable declines.

Despite this growing body of evidence, OpenAI frames the results mainly as a challenge for grading systems: if AI can produce polished, expert-like work, the final product alone says less about what students actually understand. The authors argue for rewarding students for producing work that reflects originality, reasoning, and consideration of multiple approaches, not just the most conventional or polished answers. That may be true, but changing the grading system likely means changing the entire education system along with it.

Scroll to Top