Measuring Human Performance on ARC-AGI-3
Measuring Human Performance on ARC-AGI-3
AGI is here when a system can learn like a human.
However there is still a gap between what humans can learn and what AI can learn. ARC Prize Foundation exists to understand this gap. The ARC-AGI benchmarks are our tools for measuring it.
Using these tools requires understanding human performance. How do real people, our only proof point of general intelligence, actually learn and solve novel problems?
Today we’re releasing the human dataset for ARC-AGI-3 - a controlled study of 458 participants and the most exhaustive human testing study in the ARC-AGI series to date.
We do not yet have AGI. This dataset is the receipt.
ARC-AGI-3 Public Demo environments: AR25, LF52, SB26
ARC-AGI-3
ARC-AGI-3 is a series of 135 abstract reasoning environments. Play them yourself.
The test taker, whether human or AI, is not given instructions on how to play. They must explore, infer the rules that govern the environment, and come up with a strategy on their own.
A key design constraint of ARC-AGI-3 is that every environment must be solvable by humans with no prior training. To ensure this, we conducted the largest formal study on ARC-AGI human performance ever done.
To dive deeper into ARC-AGI-3, check out the benchmark, play the environments, or watch our launch video.
Human Baselines on ARC-AGI-3
To gather human baselines we tested members of the general population. Participants included various levels of education, income, job sectors, and ages. We did not control for one particular demographic.
We held weekly, in-person, focus groups in a San Francisco-based testing center. No references to ARC Prize Foundation or AI testing were made at any time.
Each testing session lasted 90 minutes. Every participant received a base payment of ~$130, plus an extra $5 for each environment they successfully solved.
Tests were held under "first-run" conditions. This means every participant only saw…