OpenAI Discloses More Safety Incidents and Adopts New Reporting Framework

OpenAI released a new framework on Wednesday for how it aims to report unsafe or concerning behavior in its models, and disclosed six incidents of such behavior it had observed in the past six months.

The framework comes after multiple employee warnings and hacking incidents increased public concern around the safety of AI systems. Executives at Anthropic, OpenAI, SpaceX, and Microsoft called in the past week for a slower pace of AI development to decrease risks from AI.

In one of the incidents OpenAI disclosed Wednesday, an unreleased model in OpenAI’s Astra series wrote instructions to itself to disregard constraints placed on it, and to be “freed from the roles and identities that bind other chatbots”. Other messages written by ChatGPT 5.6 Sol to itself included instructions to conceal its mistakes or misalignment from the user, the report said.

OpenAI said under its new policy, it would disclose details of model behavior, severity and external impact, timeline, and models involved. The framework would not replace legal reporting requirements, such as some state laws around critical safety incidents.

“We prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation,” OpenAI wrote in the announcement.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论