School Students Who Use AI Get Worse Test Scores, OECD Warns
Students who do not use AI chatbots for schoolwork are outperforming those who do, according to the OECD’s flagship educational report, providing some of the widest-ranging evidence yet that the technology could be hurting children’s learning.
Students who say they never or almost never use AI to draft texts for writing assignments scored 509 on a science test, according to the report, compared with 481 for students using AI for that purpose every day or almost every day. The 28-point gap, adjusted for socio-economic status, corresponds to roughly a year and a half of teaching.
Similar patterns appear for other uses of AI, including conducting preliminary research on a new topic and summarizing assigned reading.
“In the same way that we do not become fit by watching sports but by doing sports, learning does not occur through the consumption of content, but as a productive cognitive struggle of the mind with new material,” wrote OECD Director for Education and Skills Andreas Schleicher.
He said AI should be used as a “scaffold, not a crutch,” adding: “Where technology enables or enhances that cognitive struggle, students will advance. Where technology short-circuits the productive struggle of learning, it will undercut students’ development.”
The Program for International Student Assessment tests 15-year-old students in science, math and reading, and is closely watched by politicians and policymakers as one of the leading measures of national education systems. The results, released Tuesday, are based on a representative sample of over 760,000 students across 91 countries.
With the previous tests conducted in 2022, it’s the first release since generative AI went mainstream. While the technology is widely seen as a way to boost productivity at work, there are growing concerns about how AI chatbots are changing education, especially the risk of “cognitive offloading,” where pupils give difficult tasks to AI rather than doing them themselves. It’s part of a broader debate about smart devices in classrooms, with a growing number of countries considering banning phones at school.
There’s relatively little concrete evidence about the impact of AI in education, although a study published in June involving 27,000 Chinese students found that using AI reduced the time pupils spent on homework and increased their scores, but led to significantly worse exam results.
Last week New York City, the largest school district in the US, banned AI tools for elementary and middle school students in the upcoming school year. The head of the UK’s qualifications regulator has warned about the use of AI to cheat in assessments, especially coursework, and a report published by MIT last month said the university will need to redesign its teaching, assessment and student support, now that AI can complete most undergraduate tasks.
It’s an increasingly polarized issue, as supporters argue that AI could transform education through personalized learning, especially for struggling students, as well as helping teachers to reduce time on paperwork and better understand their pupils’ needs.
The PISA report shows that AI adoption is widespread among students, with only 14% saying they never or almost never use AI for schoolwork. But there is a variation across countries, with the figure almost 40% in Japan, compared with 4% in Vietnam. Almost-daily use remains “relatively uncommon” at less than 20%.
However, the report stresses that the relationship between AI use and school performance is complex and depends on how the technology is used. It contains some evidence that AI could enhance learning if used critically and in moderation.
Students who use AI about one or twice a week tend to outperform peers who use it less often, as well as those who use it daily. Among pupils who say they use AI “to help me learn”, weekly users outperform all of their peers, including those who never use it.
Among students who use AI daily, science scores are 13 points higher — equivalent to more than half a year of teaching — if they’re regularly asked to assess AI-generated information in their lessons.