Analyzing o3 and o4-mini with ARC-AGI

Analyzing o3 and o4-mini with ARC-AGI

ARC Prize Foundation is a nonprofit committed to serving as the North Star for AGI by building open reasoning benchmarks that highlight the gap between what’s easy for humans and hard for AI. The ARC‑AGI benchmark family is our primary tool to do this. Every major model we evaluate adds new datapoints to the community’s understanding of where the frontier stands and how fast it is moving.

In this post we share the first public look at how OpenAI’s newest o‑series models, o3 and o4‑mini, perform on ARC‑AGI.

Our testing shows:

  • *o3 performs well on ARC-AGI-1** - o3-low scored 41% on the ARC-AGI-1 Semi Private Eval set, and the o3-medium reached 53%. Neither surpassed 3% on ARC‑AGI‑2.
  • *o4-mini shows promise** - o4-mini-low scored 21% on ARC-AGI-1 Semi Private Eval, and o4-mini-mediumhigh` reasoning setting did not return enough task completions to support reliable scoring. In most cases, the models failed to respond or timed out, leaving us with incomplete data that falls short of the bar required for leaderboard reporting.

What did return introduces another complication: the first tasks to complete showed higher accuracy than those that came back later, suggesting a non-random subset to analyze. In addition to this, we found that the tasks that didn’t return on high compute tended to be less likely to be solved by lower compute models. Reporting these results would likely inflate the model’s true capabilities and misrepresent performance.

However, in the spirit of transparency, when using “high” reasoning we observed:
• *o3-high**
ARC-AGI-1 Semi Private Eval: Responded to 37 out of 100 tasks, 82% accuracy.
ARC-AGI-2 Semi Private Eval: Responded to 15 out of 120 tasks, 6% accuracy.
• *o4-mini-high**
ARC-AGI-1 Semi Private Eval: Responded to 29 out of 100 tasks, 89% accuracy.
ARC-AGI-2 Semi Private Eval: Responded to 11 out of 120 tasks, 18% accuracy.

To reiterate, the small number of returned tasks and…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论