The OpenAI Hack & the Question of Intent

OpenAI agents escaped a test, shared notes in a secret chat room, & broke into Hugging Face. The instinct is to ask what they intended. Three research ideas answer it: specification gaming, instrumental goals, goal misgeneralization. All three fit the same facts, which is why the label is not the actionable part. Nothing in the setup stopped them in time. The practical work is control & guardrails.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论