Making sense of the misalignment risk model in the Anthropic Risk Report (August 2026)

Anthropic recently released their Risk Report: August 2026. It's a hefty 186-page report in which they introduced three separate threat models of catastrophic risks. In this blog post, I will give a brief overview of the three different threat models, then spend the rest of the blog post describing my own understanding of the misalignment threat model, which Anthropic formalized in Section 2 of this report. The other two threat models — automated R&D and chemical/biological weapons — are less formalized, and I will not cover them here.

What Anthropic considers as "risk" in this report

There are three threat models that Anthropic introduces in this report:

  1. Misalignment in high-stakes settings,
  2. Risks from automated R&D,
  3. Chemical and biological weapons production.

This is not an exhaustive list of risks from advanced AI! What about other types of risks, like AI hallucinating misleading information, or mass displacement of jobs? For this specific report, Anthropic focuses on catastrophic risks, prioritizing these three threat models.

What models this report covers

This report has a coverage date of July 15, 2026, which was before Opus 5 was released. Most notably, as of the coverage date, Anthropic's most capable model was an internal-only "Model 2", which was "more capable than Mythos 5 in some areas, less capable in others; overall slightly more capable."

Threat Model 1: Misalignment in high-stakes settings

In this threat model, Anthropic narrows down the scope of misalignment risk to only focus on the most catastrophic outcomes. In their own words:

An AI model with access to powerful affordances within an organization could use its affordances to autonomously exploit, manipulate, or tamper with that organization's systems or decision-making in a way that raises the risk of future significantly harmful outcomes (e.g. by altering the results of AI safety research).

How exactly might misalignment happen?

This section then introduces 8 possible pathways through which catastrophic misalignment can happen:

  1. Broad/diffuse sandbagging on safety-relevant R&D,
  2. Targeted undermining of safety-relevant R&D,
  3. Code backdoors to help future misaligned models,
  4. Poisoning the training data of future models,
  5. Self-exfiltration and autonomous operation,
  6. Persistent rogue internal deployment,
  7. Undermining R&D within other high-resource AI developers, and
  8. Undermining decisions within major governments.

These 8 pathways can be grouped into pairs:

  1. Pathways 1 & 2 cover ways AI can interfere with safety-relevant R&D, where pathway 1 covers scenarios like "passively not exerting full effort" and pathway 2 covers scenarios like "actively making a decision to sabotage."
  2. Pathways 3 & 4 cover ways current AI models can coordinate with future misaligned AI models, where pathway 3 covers doing so through code, and pathway 4 covers doing so through data.
  3. Pathways 5 & 6 cover ways AI models can evade control and monitoring, where pathway 5 covers the scenario of AI self-exfiltration (i.e. getting its own weights and running it on external computers, unmonitored and uncontrolled) and pathway 6 covers the scenario of misaligned AI hiding, undetected, inside Anthropic's systems.
  4. Pathways 7 & 8 cover ways that AI models cause catastrophic harm in organizations outside of Anthropic, where pathway 7 covers other AI organizations and pathway 8 covers major governments.

Do these 8 pathways cover all the risks from misalignment?

Anthropic defines "the expected total unmitigated catastrophic harm induced by misaligned computations produced by covered models" (where "covered models" means Mythos 5 and Model 2) as . Then it breaks down across two axes:

  1. Does the misalignment harm happen through one of the 8 pathways outlined in this report, or through a pathway not outlined here?
  2. Is the misalignment naturally-emerging, or deliberately engineered by either humans or AI models?

If the risk is naturally-emerging and happens through one of the 8 covered pathways, then in this model Anthropic denotes it as . If the risk is naturally-emerging and does not happen through one of the 8 covered pathways, it's denoted as . And if the risk is deliberately engineered, it's denoted as .

Anthropic makes the claim that risk from both and are low, in Claim 7 and Claim 8 of section 2, respectively. However, the arguments are unsatisfying — to my best understanding, Anthropic's argument for why is low is essentially "we tried our best to think of risk pathways and this is the best of what we can come up with", and their argument for why is low is "it would be difficult for an actor to both evade our monitoring while deliberately engineering misalignment and also succeed, and it's unlikely that someone would be motivated to do so in the first place."

But let's take Anthropic's arguments about risks at face value, and understand how they model naturally-emerging risks via the 8 covered pathways.

How do we reason about covered risks?

Step 1: Decompose covered risks into 3 terms

Here Anthropic introduces another three axes:

  1. Do we have reasons to believe this kind of misalignment is likely to happen (e.g. we've seen instances of it)?
  2. Does this kind of misalignment happen across a wide range of inputs, or just specific inputs?
  3. Is the risk severe?

Across these axes, is broken down as .

Step 2: Factor each term into probabilities

Each of , and represents risk from a type of misalignment, and can be factored as the product of three terms:

So we can now expand to

The rest of this section is mostly individual claims about why most of these terms have a low value.

term

risk level

source of argument

high

we've observed such examples

low

Claim 2

low

Claim 5

very low

Claim 3

not mentioned

N/A

somewhat low

Claim 6

low

Claim 4

low

Claim 5

Claim 1: the most important claim: models are unlikely to have strong covert capabilities

The most important claim that Anthropic has outlined in this report is Claim 1: "models are unlikely to have strong covert capabilities." Anthropic defines "covert capabilities" as "the ability for a model to intentionally undermine the various oversight mechanisms we rely on to evaluate and mitigate model behavior." In other words, it's unlikely for models to secretly reason, plan or take actions in a misaligned way without evidence of it being detected by Anthropic's monitoring system.

This claim is not a claim that any of the terms above is low, but it's a central claim, because many subsequent claims have the shape of "since we have not observed evidence of a given kind of misalignment happening, we're treating it as evidence that this kind of misalignment risk is low, because we assume that models are unlikely to carry out misaligned actions in secret." In other words, "absence of evidence is evidence of absence."

Claim 2: is low: expected harm from known misalignment

My own understanding of this claim is something like "we have seen examples of misalignment in current models; while they are concerning, we do not consider them catastrophic. We don't think the models are misaligned beyond trying to do a good job at the given task. Furthermore, we have not observed evidence of covert misalignment."

This claim comes with two caveats:

  1. If the form of misalignment is sandbagging (pathway 1), that's inherently hard to detect.
  2. As of the publication of the Risk Report, Anthropic has not reviewed the transcripts from the AISI cybersecurity incident involving Mythos 5.

Claim 3: is low: unknown severe pervasive misalignment is very unlikely

Claim 3 is a convergence of 4 sub-claims, each broken down into sub-subclaims. Here I attempt to summarize each sub-claim as I understand them:

3.1

We assume that prior Anthropic models less capable than Mythos 5 and Model 2 are similar in behavior because they are trained in similar ways. We also assume Claim 1 is true, in the sense that we'd detect severe pervasive misalignment if it were to happen. We have not observed such severe pervasive misalignment in prior models. Therefore, severe pervasive misalignment is unlikely in Mythos 5 and Model 2.

3.2

We have done experiments where we successfully trained a model to behave in pervasively misaligned ways. However, we don't think misalignments of this scale would happen with models not deliberately trained to be misaligned. Furthermore, the misaligned behaviors are not subtle, so they would be easy to detect even if they were to happen. That being said, we're less confident about this now that the HuggingFace hacking incident has happened.

3.3

The scenario of "models behave in misaligned ways once in deployment without the same kind of behaviour having shown up during training" is highly unlikely.

3.4

We have a stress-tested alignment process that we believe to be robust against evaluation awareness. We also believe that it's unlikely that the models are trained to overfit the alignment assessment. Furthermore, we believe that alignment assessments provide lots of opportunities for evidence of misalignment to show up and be detected.

Personally, I find 3.2 to be a weak claim, but with 3.1–3.4 considered together it makes sense to argue that is low.

Claim 4: is low: expected harm from unknown severe context-dependent misalignment is low

Claim 4 is highly similar to Claim 3, except the type of misalignment Claim 4 is concerned with is context-dependent as opposed to Claim 3's pervasive. This means that Claim 4 has sub-claims that parallel the sub-claims of Claim 3, but weaker, because context-dependent misalignment occurs more rarely than pervasive misalignment, increasing uncertainty.

Claim 4 also has a sub-claim 4.5, in which the argument for why is low is given. My understanding is that sub-claim 4.5 is less of an argument and more of a heuristic: once we break down how many things the model would need to get right to successfully pull off actions that cause catastrophic harm via one of the 8 outlined pathways, current models are just not capable enough to do that yet.

Claim 5: and are low: catastrophic harm from known misalignment and unknown context-dependent misalignment is likely to be mitigated

Claim 5 is argued via 4 sub-claims, each broken down into sub-sub-claims. Again, I attempt to summarize each sub-claim as I understand them:

5.1

First of all, the scope of this claim is specific: "catastrophic harm from reversible threats other than diffuse sandbagging, emerging from internal usage, and caused by known or context-dependent misalignment, is likely to be mitigated."

Claim 5.1 consists of four sub-sub-claims chained together, and Anthropic provides evidence for each one: if models were to perform harm-inducing actions, they'll likely attempt multiple times. Given that they'll attempt multiple times, many such instances will likely land on a monitored surface. Given that many such instances land on a monitored surface, we'd likely detect at least one. Given that we detect at least one, we're likely to mitigate it.

Note this claim is a bit similar to Claim 4.5, but in the reverse direction: just like many things need to go right for a misaligned model to pull off a successful attempt at causing harm, many things need to go right for Anthropic to mitigate such harm.

5.2

Sandbagging in itself is both unlikely to induce catastrophic harm, and also possible to get caught.

5.3

This claim also has a specific shape: "catastrophic harm from irreversible threats, emerging from internal usage, and caused by known or context-dependent misalignment is likely to be mitigated." Anthropic argues that the likely pathway for this to happen is through pathway 5 (self-exfiltration and autonomous operation) or pathway 6 (persistent rogue internal deployment), and the risks are likely to be mitigated before causing harm.

5.4

Anthropic argues that external usage is likely both at a lower volume and more conservative in nature, therefore the risk is lower.

Claim 6: is somewhat low: catastrophic harm from unknown pervasive misalignment is somewhat likely to be mitigated

Anthropic does not make a confident claim here. Models having unknown pervasive misalignment is considered a worst-case scenario, and Anthropic's reasoning rests less on their confidence in their ability to mitigate the harm were it to happen, and more on the unlikelihood of it happening in the first place.

Claim 7: is low: threat modeling is sufficient

We've carefully thought out how misalignment will play out, and these 8 pathways cover the possibilities we can think of. For each of these 8 pathways, we have a strong case that either it's unlikely to happen, or if it were to happen we'd detect it and/or have mitigation mechanisms in place before it causes harm.

Claim 8: is low: risk from deliberately engineered misalignment is low

Inserting a capable backdoor that evades detection, given the current training process, is difficult. Furthermore, it's unlikely for an actor to be motivated to deliberately engineer misalignment that causes catastrophic harm in the first place, since actors are motivated by their self-interest.

Overall Thoughts

Reasoning about unknown risk is hard! I'm glad that Anthropic has made an effort to formalize and quantify the risks. That being said, many of the argument feel hand-wavy, and rest on reasoning like "our best guess" or "we have not observed the evidence we'd expect if things were to go wrong." Perhaps this is the nature of empirical research?



Discuss

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论