A recent study on Chinese students shows that they generally use Artificial Intelligence (AI) to assist in learning. While students using AI showed significant improvement in their homework grades, their performance in exams noticeably declined.
According to a report from The New York Post on Wednesday, a research paper titled “Learning Penalties from Generative AI: Evidence from Chinese Secondary Education” was published on June 23 on the Social Science Research Network platform.
The study was conducted collaboratively by David Stromberg from Stockholm University in Sweden, Victor Lei and Wu Yanhui from the University of Hong Kong.
The research team tracked a total of 26,811 middle and high school students aged 12 to 18 in a county in central China over a span of 30 months.
The data indicates that students using AI for writing assignments showed an 18% increase in homework grades six months later compared to the baseline average. The time required to complete each assignment also decreased from 64 minutes to approximately 45 minutes, a reduction of about 30%. However, the same group of students saw a 20% decline in closed-book exam scores, a decrease that gradually accumulated over approximately six months and stabilized.
Approximately 80% of the interviewed students reported using Chinese domestic AI models, including Doubao and DeepSeek, while the remaining 20% served as a control group.
As of June 2025, Doubao was the most widely used AI tool at 47%, followed by DeepSeek at 36%. The most concentrated subject for AI use was mathematics at 66%, followed by English at 55%.
The phenomenon observed by the authors is coined as “outsourcing assignments,” revealing that many students are not actually writing their own assignments but rather directly copying answers provided by AI.
The study found that students who completed assignments particularly quickly had scores that closely matched the AI’s accuracy on similar topics, indirectly affirming that the high scores on these assignments are likely not calculated by the students themselves but rather the result of AI assistance.
Furthermore, the research discovered an interesting aspect: when comparing AI users and non-users who spent similar amounts of time on assignments, their exam scores did not exhibit significant differences.
This demonstrates that the issue lies not in the quality of AI tools themselves but in whether students, due to using AI, are neglecting the time that should be allocated for critical thinking and practice.
In other words, as long as students are willing to invest time in thoughtful consideration and practice, the use of AI has minimal impact; the real disparity arises from individuals who slack off and devote less time.
The study also revealed that the impact of artificial intelligence on major exams such as the “Zhongkao” and “Gaokao” (national college entrance exams) is much slower compared to regular tests. On a general scale, the difference in academic performance between students using AI and those who do not is only around 5% to 6%.
However, with a prolonged timeframe in consideration, the scenario changes. For students who have been using AI for assignments for over two years, the gap in scores for major exams widened to 24% for “Zhongkao” and 18% for “Gaokao.”
The research team explained that this discrepancy arises because entrance exams test knowledge accumulated over multiple years. Even if the recent reliance on AI leads to laziness, the foundation of solid learning from earlier remains, making the impact less immediately discernible.
However, monthly and mid-term exams generally evaluate recently taught content, and the use of AI immediately affects performance levels more directly.
The study also found discrepancies in the impacts across various subjects: students in liberal arts (such as politics, history, and geography) experienced the most severe decline, averaging a 27% decrease, while those in science subjects saw an average decline of 22%.
Of note, students who originally had higher grades experienced a more significant decline, surpassing even those with average grades—a contrast to previous workplace studies where AI was found to offer the most assistance to individuals with weaker abilities.
The authors of the paper cautioned that if parents and teachers solely focus on metrics like assignment scores or exam results, they may overlook underlying issues—achieving good grades does not necessarily equate to genuine understanding.
They suggested that instead of fixating on grades, it is essential to pay attention to how much time and effort children actually devote to their studies to detect problems sooner.
However, the study also mentioned a contrasting case from Middlebury College in Vermont, where an experiment was conducted with students learning a completely unfamiliar topic. The results indicated that with the right approach and cautious use of AI for learning assistance, students scored better on tests compared to those not using AI, and this advantage persisted even during a retest a week later.
The researchers emphasized that the issue does not lie with AI itself but in “how it is applied.” The key factor is whether students can resist the urge to immediately seek answers from AI when faced with problems, opting instead to think independently first and utilize AI for checking or reinforcing—this way, AI can serve as a learning aide rather than a shortcut for laziness.
