内容讨论 图片 1

We previously described some of the changes we’ve made to our alignment and security efforts following these incidents here:

Anthropic (@AnthropicAI)
We’re sharing an update on our alignment and security efforts.

In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.

In a new post, we describe:

1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards

2. An update on our alignment assessment

3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them

4. How we hardened our security practices earlier this year to prepare for Mythos-class models

Read more: anthropic.com/news/improving…
Link
Improving our alignment and security practices (anthropic.com/news/improving-alignment-security-efforts) On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. We are conducting an in-depth analysis of both incidents, and planning to work with... anthropic.com
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论