Principles for Embedded Evaluations
This is a linkpost for https://www.apolloresearch.ai/blog/principles-for-embedded-evaluations
Blog post
In September 2026, leaders of frontier AI companies called for pacing the frontier of AI development, with embedded evaluators as the first step. Frontier AI companies committed to giving outside evaluators employee-like access to their training, evaluation, and deployment, and some have since published principles for third-party assessments.We are very excited about this development. Public third-party assessments of how frontier AI is developed are urgently needed, and embedded evaluations are a good first step. But their impact will depend heavily on implementation. If evaluators lack necessary access, if they are given too few resources, or if their findings carry little weight in actual decisions about frontier development, embedded evaluations may not amount to much.This post sets out our current thinking on core principles for embedded evaluations that assess loss of control risks from scheming, i.e., AI models covertly subverting their developers in pursuit of unintended goals. We first describe at a high level what effective embedded evaluations should achieve. We then propose a concrete design based on verifying or falsifying developers' safety claims, which we hope developers and evaluators will adopt. We see these principles as a minimal starting point rather than a complete framework. Many of them build on established practice for independent auditing in other high-stakes industries, adapted to the specific challenges of frontier AI.
Twitter Thread
1/
Frontier AI companies have committed to giving outside evaluators employee-like access to training, evaluation and deployment. We're very excited about this. But its impact depends heavily on implementation.
Today we're sharing our principles for embedded evaluations. 🧵
2/
If evaluators lack access or resources, or if their findings carry little weight in real decisions, embedded evaluations may not amount to much. In the worst case, they give a false sense of security.
3/
We think an embedded evaluation must meet four criteria:
1. It reduces risk from frontier AI
2. It informs the public
3. It gives both parties incentives that favor safety
4. It is fair to the developer and the evaluator
4/
Reducing risk means catching problems early. Warning signs like models subverting their training or evading monitoring could appear long before release. So evaluators need ongoing access to training data, rollouts, checkpoints and internal deployment.
5/
Informing the public means publishing by default and stating clearly what access the evaluator had. Developers may redact sensitive details, and the report states what kind of information was removed and whether it mattered. The evaluator's conclusion can never be redacted.
6/
Good incentives and fairness mean agreeing on scope up front, keeping the evaluator independent, and protecting both sides' confidential IP. We also believe voluntary commitments are unlikely to be enough, so embedded evaluations should eventually be required by law.
7/
Our concrete proposal is claim-based assessment. Instead of an overall judgment of the developer, evaluators assess specific claims fixed in advance, such as "Models never attempted to disable or evade their monitoring during internal deployment."
8/
Claims come from two sources.
Standard claims apply to every developer and cover properties any safe developer should be able to demonstrate.
Recommended claims are specific to one developer and drawn from its commitments, safety case, past incidents or what evaluators observe.
9/
Each claim goes through three steps:
1. Review Evidence. The developer shares existing evidence.
2. Gather Evidence. The evaluator looks for new counterexamples.
3. Assess Evidence Adequacy. The evaluator tests whether the developer's methods would have found one if it existed.
10/
Not finding a counterexample is not enough to establish a claim as true. We also need confidence that the developer's methods could find one.
For a monitoring claim, the evaluator might insert realistic evasion attempts and check how many the monitors flag.
11/
Each claim ends with one of five verdicts:
A. Verified
B. Unrefuted
C. Self-reported
D. Evaluator-found
E. Blocked
12/
The verdicts are ranked to reward openness. Insufficient access gives the worst verdict. A problem the developer reports itself scores better than one the evaluator finds.Beyond that, the aim is that the easiest path to a better verdict is to actually improve safety.
13/
Public reports come out on a fixed schedule. They list every claim and verdict, as well as the access the evaluator asked for and received.
If the developer disagrees, it can add a statement, but it cannot veto the verdict.
14/
We see this as a starting point and expect the methodology to evolve. We welcome feedback from developers, other evaluators and policymakers, and we hope the first embedded evaluation agreements set a high bar for those that follow.Full post: https://www.apolloresearch.ai/blog/principles-for-embedded-evaluations
评论
?
参与讨论