AI Labs Shouldn’t Control What Investigators Can See

(Image source here)

In light of the unsanctioned AI compromise of OpenAI infrastructure, the hacking of Hugging Face, and other misalignment and security incidents, both OpenAI and Anthropic have allowed independent investigations of their models’ actions. The terms and scope of these investigations vary from company to company and are subject to their benevolent discretion.

In OpenAI’s case, the terms under which it engaged METR and Redwood Research limited the investigation to the Hugging Face hack, excluding the subsequent compromise of OpenAI’s own infrastructure and other acknowledged intrusions. While the investigation is welcome, the limited scope and time for the investigators constrain its effectiveness. Anthropic gave more access, including transcripts of models beyond the incident window. It also gave access to employees who could share confidential information with METR, and has said it intends to give METR access for as long as METR deems necessary.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论