ARC Prize

RSS: https://arcprize.org/feed.xml
ARC 奖基金会(ARC Prize Foundation)是一个非营利组织,通过基准测试和奖金推动开源通用人工智能(AGI)研究。

DeepSeek V4 Flash 0731 推理基准成绩发布

DeepSeek 发布 V4 Flash 0731 模型,在 ARC-AGI 推理基准上取得新成绩。最高档位 ARC-AGI-1 得分 89.0%,ARC-AGI-2 得分 61.4%,单任务成本仅 0.02-0.04 美元。模型提供 Max/High/Low 三档推理强度,成本与性能呈线性关系。
评论点赞收藏8 天前

ARC-AGI 排行榜:从被动流体智力到实时交互适应

ARC Prize 发布 ARC-AGI 排行榜,涵盖 ARC-AGI 1/2/3 三个版本。ARC-AGI-3 转向评估 AI 在新型交互环境中的即时适应能力。榜单可视化了任务成本与性能的效率关系,区分了推理系统趋势线、基础 LLM 单次推理及 Kaggle 竞赛方案。仅展示运行成本低于 1 万美元的系统,部分结果为预览或估算。
评论点赞收藏22 天前

Inkling:ARC Prize 榜单上得分最高的开源权重模型

Thinking Machines 发布开源模型 Inkling,在 ARC-AGI-1 上得分 79.5%,ARC-AGI-2 上得分 36.5%,刷新了开放权重模型在 ARC Prize 榜单上的最高纪录。该模型单任务推理成本仅约 0.3 至 0.64 美元,且尚未在 ARC-AGI-3 上进行评估。
评论点赞收藏29 天前

GPT-5.6系列:ARC-AGI测试结果

OpenAI发布GPT-5.6系列在ARC-AGI基准测试的成绩。Sol模型在最大推理力度下,于ARC-AGI-3公共测试集平均得分13.33%,并在ft09任务中以87%正确率成为首个赢得该任务的模型。文章详细列出了Sol、Terra、Luna三款模型在不同推理层级(Max至Low)下的具体得分与通过率数据。
评论点赞收藏37 天前

ARC Prize 2026:ARC-AGI-3 第一阶段里程碑大奖揭晓

ARC Prize 2026 公布 ARC-AGI-3 基准测试第一阶段里程碑奖($37.5K)及前三名开源方案。冠军 Tufa Labs 的 "The Duck" 采用小型开源 LLM,通过实时 REPL 编写运行 Python 代码解决交互式推理游戏,利用图像、ASCII 和分割工具灵活感知状态,并通过消息淘汰机制实现无限上下文下的持续游玩。该方案展示了让模型驱动轻量级 Harness 处理长视界规划与探索的新思路。
评论点赞收藏41 天前

Analyzing GPT-5.5 & Opus 4.7 with ARC-AGI-3

本文通过 ARC-AGI-3 基准测试,对比分析了 GPT-5.5 与 Opus 4.7 的推理过程,发现三大共性失败模式:模型能感知局部操作效果却无法构建正确世界模型;错误地将新环境类比为已知游戏(如 Tetris、Sokoban);即使解决单个关卡也未能真正理解机制并迁移到下一关。研究还指出两者压缩策略差异——Opus 倾向快速形成错误理论,GPT-5.5 则难以有效压缩信息。该分析工具已开源,旨在推动更透明的 AI 能力评估。
评论点赞收藏107 天前

Measuring Human Performance on ARC-AGI-3

Measuring Human Performance on ARC-AGI-3 AGI is here when a system can learn like a human. However there is still a gap between what humans can learn and what AI can learn. ARC Prize Foundation exists...
评论点赞收藏123 天前

Announcing ARC-AGI-3

Announcing ARC-AGI-3 A New Challenge for Frontier Agentic Intelligence Today we're excited to announce the release of ARC-AGI-3, a series of hundreds of interactive environments and thousands of game-...
评论点赞收藏144 天前

Arc Prize Verified

Announcing ARC Prize Verified Today we're announcing **ARC Prize Verified**, a program to increase the rigor of evaluating frontier systems on the ARC-AGI benchmark. In addition to certified score ver...
评论点赞收藏282 天前

Hidden drivers of HRM's performance on ARC-AGI

The Hidden Drivers of HRM's Performance on ARC-AGI We scored on hidden tasks, ran ablations, and found that performance comes from an unexpected source On June 8, 2025, the Hierarchical Reasoning Mode...
评论点赞收藏312 天前

Arc-AGI-3 Preview: 30-day learnings

ARC-AGI-3 Preview: 30-day learnings Highlighting the gap between humans and AI with Interactive Benchmarks ARC-AGI-3, the first Interactive Reasoning Benchmark by ARC Prize Foundation On July 17, we r...
评论点赞收藏361 天前

Analyzing o3 and o4-mini with ARC-AGI

Analyzing o3 and o4-mini with ARC-AGI ARC Prize Foundation is a nonprofit committed to serving as the **North Star for AGI** by building open reasoning benchmarks that highlight the gap between what’s...
评论点赞收藏480 天前

登录芦苇

登录后关注作者、收藏内容和参与讨论。