METR 发布了一份关于智能体能力评估指标的综述,构建了基于“支出-得分”曲线的指标分类体系。文章详细梳理了不同场景下的衡量标准,但未提供具体推荐或深层测量讨论。对于关注 AI 对齐、模型评估方法论的研究者和工程师而言,这是一份梳理现有评估框架的重要参考资料,有助于理解如何量化智能体与人类在资源消耗下的表现差异。
We reviewed the “Risks from automated R&D” section of Anthropic’s February 2026 Risk Report, producing two corresponding review documents: our original review and our updated review. We recommend that...
In collaboration with Anthropic, a METR staff member (David Rein) recently spent three weeks red-teaming a subset of Anthropic’s internal agent monitoring and security systems, many of which are descr...
As METR’s time horizon task suite saturates, the results are becoming more sensitive to analysis choices. One example of this was the recent update to fix a modelling mistake with regularization, whic...
Introduction METR aims to keep the public informed about the capabilities of and risks posed by AI — by some metrics the fastest-moving technology in history, and one that could speed up further as AI...