Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

A collage of a computer screen showing red corrections and X marks next to a stack of scientific reports containing charts.

AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpublished NeurIPS papers rated the results as "Reject." According to the study, conducted with Princeton and the UK AI Security Institute, frontier models can handle the full research engineering process but fall short on research judgment, creative problem-solving, and the ability to abandon failed approaches.

The article Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach appeared first on The Decoder.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论