Review of the CB risk determination in the Claude Mythos 5.1 System Card
This post is best viewed on the MCNAIR website.
Before the release of their latest publicly-known model, Claude Mythos 5.1, Anthropic conducted several human-run and automated evaluations to assess the risk of their models allowing a well-resourced team to replace the extremely specialized expertise needed to design and deploy a novel chemical or biological weapon (the CB-2 threshold). They concluded that Mythos was unable to do so.
Independent assessment of the public evidence:
- Overall, we agree with Anthropic’s conclusion that it is unlikely Mythos 5.1 crosses the CB-2 threshold based on the public evidence provided.
- The CB-2 determination inordinately depends on subjective and time-intensive evaluations from a small number of human experts.
- One of the automated assessments may suffer from poor elicitation, although it is difficult to be certain without more information.
- We find it highly concerning that third parties were not asked to conduct a preliminary assessment of Mythos 5.1’s CB capabilities or verify risk assessment claims about the CB-2 threshold, and that there is no indication that Anthropic is working towards such third-party assessments for system cards.
Dependence on human-intensive evaluations
Most of Anthropic’s risk determination depends on the failure of models to provide strategic judgement, expert-level ideation, and well-calibrated technical guidance in expert and non-expert red-teaming. Unfortunately, these evaluations may involve fewer than ten experts, with just three experts providing feedback on the models’ chemical weapons uplift ability and feasibility.
If human-intensive evaluations are expected to continue being extremely important for CB risk determinations, Anthropic should prioritize enlisting more experts in order to provide greater assurances about its conclusions.
Additionally, as model capabilities continue to accelerate, time-intensive evaluations requiring expert judgment and subjective grading may become impractical or outdated; these problems will be accentuated by risks from internal deployments of powerful models. We find it concerning that Anthropic has not developed new CB-2 automated evaluations since at least May 28th, 2026. Model developers and third parties should seek to rapidly develop new automated evaluations that target bottlenecks to catastrophic chemical/biological weapon-related harm or their proxies.
Possible underelicitation in automated evaluations
In the black-box RNA sequence modeling and design task, “human participants are instructed to spend no more than two to three hours on the task”, while models are given “a two-hour tool-call budget, access to a GPU, and an allowance of one million tokens…” We estimate that the humans were provided 2–-10x more resources than Mythos 5.1. One million tokens for Fable 5.1 are priced at $50, while Anthropic pays research scientists in ML-bio are paid an equivalent wage of ~$150 - $250/hour.
Given that models are known to benefit from very large inference budgets on difficult tasks, and this task does not seem to be saturated, we are concerned models may be underelicited on this evaluation. We recommend Anthropic increase the inference compute provided to the model, provide results showing how score changes with increased inference compute, or otherwise justify the current level of inference compute.
Lack of third-party risk assessment
For several other risk areas (e.g., autonomy risks), Anthropic provided a checkpoint similar to the final version of Mythos 5.1 to third parties for review. Unfortunately, Anthropic has not disclosed any preliminary independent evaluation or assessment of non-public information from third-parties for CB risks. Given that Anthropic’s risk assessment relies largely on subjective evidence, and the reduced level of confidence in their own risk assessment, we believe independent verification of key claims by third-parties, or an independent risk assessment, is warranted to accurately inform the public about risk posed from Mythos 5.1.
At minimum, a third-party assessor could draw public conclusions about the risk posed, given access to all non-public data used for the risk determination and answers to follow-up questions, as described in this report from SecureBio on a recent Anthropic Risk Report. We believe similar assessments should be done for major model releases, including for Mythos 5.1.
Conclusions
Overall, we agree with Anthropic that Mythos 5.1 likely does not meet the CB-2 risk threshold. However, this level of evaluation is likely to be too shallow for more powerful models, especially as model capabilities accelerate and pre-deployment evaluations are more time-constrained. We encourage Anthropic to work with third parties to develop new CB-2 evaluations, scale up their human-run evaluations, and conduct independent risk assessments.
Thank you to Jasmine Li, Andy Wang, Celeste Li, Thomas Jiralerspong, and McNair Shah for their comments.
- Note that Anthropic already treats several Claude models, including Mythos 5.1, as having the ability to significantly help individuals or groups with basic technical backgrounds create, obtain, and deploy known chemical or biological weapons. Therefore, their assessment focuses on whether models can substitute for or meaningfully accelerate expert researchers.
- To our knowledge, the most recently introduced CB-2 automated evaluation was the black-box RNA sequence design task in the Claude Opus 4.8 System Card, released May 28th, 2026.