OpenAI O3 breakthrough high score on ARC-AGI-PUB

OpenAI o3 Breakthrough High Score on ARC-AGI-Pub

OpenAI has released a new version of o3. Read our analysis to learn how it differs from the preview below.

Updated (April 16, 2025): OpenAI has officially released o3. OpenAI has confirmed that this version is not the same as the one we tested in this original post. See more information on this. We will publish updated results for released o3 shortly.

OpenAI's new o3 system - trained on the ARC-AGI-1 Public Training set - has scored a breakthrough 75.7% on the Semi-Private Evaluation set at our stated public leaderboard $10k compute limit. A high-compute (172x) o3 configuration scored 87.5%.

o Series Performance

This is a surprising and important step-function increase in AI capabilities, showing novel task adaptation ability never seen before in the GPT-family models. For context, ARC-AGI-1 took 4 years to go from 0% with GPT-3 in 2020 to 5% in 2024 with GPT-4o. All intuition about AI capabilities will need to get updated for o3.

The mission of ARC Prize goes beyond our first benchmark: to be a North Star towards AGI. And we're excited to be working with the OpenAI team and others next year to continue to design next-gen, enduring AGI benchmarks.

ARC-AGI-2 (same format - verified easy for humans, harder for AI) will launch alongside ARC Prize 2025. We're committed to running the Grand Prize competition until a high-efficiency, open-source solution scoring 85% is created.

Read on for the full testing report.

OpenAI o3 ARC-AGI Results

Update 12/20/2024: ARC Prize presented o3's performance results in person with OpenAI's Sam Altman (CEO) and Mark Chen (SVP Research) during the final "12 Days of OpenAI" event. Watch the recording here.

We tested o3 against two ARC-AGI datasets:

  • *Semi-Private Eval**: 100 private tasks used to assess overfitting
  • *Public Eval**: 400 public tasks

At OpenAI's direction, we tested at two levels of compute with variable sample sizes: 6…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论