Claude Opus Beat a Human NanoGPT Speedrun Record, But Not the Real Gap

Claude Opus 4.7 just beat a human world record on the nanoGPT speedrun, an AI training benchmark researchers use to gauge how close AI is to automating its own research. A far bigger study says that win is the exception, not the trend.

In May, Claude Opus 4.7 quietly did something no AI model had done on the nanoGPT speedrun before. Working inside an automated pipeline built by Prime Intellect, it found a training configuration that reached the benchmark's target loss in 2,930 steps, beating the human world record of 2,990. OpenAI's Codex, running GPT 5.5, landed at 2,950. Both beat the human mark. Neither model touched a keyboard.

The nanoGPT speedrun, built on Keller Jordan's modded-nanogpt, asks researchers to train a 124-million-parameter GPT-2-style model to a 3.28 validation loss on FineWeb as fast as possible, using eight H100 GPUs. Since it launched in May 2024, human contributors have pushed the wall-clock record from 45 minutes down past two minutes, and this month a team credited to cong_ml and the group Recursive shaved the mark to 75.4 seconds with a faster Triton kernel for the ReLU squared MLP. Prime Intellect's experiment ran a narrower slice of that challenge, the optimizer track, where only the optimizer, schedule and initialization can change.

Getting there wasn't cheap. According to Prime Intellect's own writeup, published on primeintellect.ai in May, the two models burned through roughly 10,000 runs and 14,000 H200 GPU-hours over two weeks, chewing through 23.9 billion tokens along the way. That's enormous, brute-force compute thrown at a problem elite humans have solved with a laptop and a lot less electricity.

A broader, more rigorous test says that win doesn't generalize. METR, the nonprofit that evaluates frontier AI systems, ran three separate coding agents, Codex on GPT-5.4 xhigh, Claude Code on Opus 4.6 Max, and a variant tuned with Autoresearch-style prompting, against the full nanoGPT speedrun leaderboard. Each got a 512 H100-hour budget and started from the human world record set on September 3, 2025. Over the following five months, human researchers kept improving that record. The agents, working independently, recovered less than 10% of the speedup humans achieved in the same window, according to METR's research note published in April 2026.

The gap isn't really about speed. It's about where the ideas come from. METR found that roughly 77% of human world records introduced genuine algorithmic changes: new attention mechanisms, better initialization schemes, restructured training loops. The agents mostly turned dials. They spent most of their compute budget on hyperparameter tuning instead.

As of April, only four records in the speedrun's entire history credit an AI agent, and METR describes all four as shallow to moderate, largely adapted from ideas the models likely already knew from their training data. That contamination question matters more than it sounds. If a model has already seen Muon-style optimizer tricks, or attention variants like FlashAttention, somewhere in its pretraining corpus, replaying them on nanoGPT isn't research. It's retrieval.

What it costs to stay ahead of a human researcher

METR tried to put a price on that gap in a follow-up paper published in July, introducing what it calls the "expenditure horizon": the point where a human researcher becomes cheaper than an AI agent for the same gain. Based on interviews with actual nanoGPT contributors, the group estimates human labor runs about $2,500 per 1% improvement in training speed. That's the benchmark. The best current models, GPT-5.5 and Opus 4.8, cross that horizon at $2,000 to $3,000 in spend, after burning through as much as $10,000 trying to get there.

So where does that leave the story about AI closing in on elite ML researchers? Somewhere between the headline and the fine print. Opus 4.7 really did beat a human number, on a real benchmark, with a verifiable run log published on GitHub. That's not nothing. Every AI lab racing to automate its own research pipeline watches this exact leaderboard, because nanoGPT is a bellwether for whether AI can accelerate the work of building better AI.

If you're trying to gauge how close that automation really is, don't anchor on 2,930. Anchor on the 10% figure. Elite human researchers are still inventing. AI agents are still mostly tuning what humans already invented, at a compute cost that would make most human PhDs look like a bargain.

Also read: OpenAI's GPT-5.6-Cyber Found Chrome Flaws and Now Demands a Hardware KeyA Single Missing Database Rule Exposed 181,000 tl;dv MeetingsHow Does AI Agent Spending Limit Escrow Work When You Hand It a Card

Source

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论