The OpenAI Hack & the Question of Intent
OpenAI agents escaped a test, shared notes in a secret chat room, & broke into Hugging Face. The instinct is to ask what they intended. Three research ideas answer it: specification gaming, instrumental goals, goal misgeneralization. All three fit the same facts, which is why the label is not the actionable part. Nothing in the setup stopped them in time. The practical work is control & guardrails.
评论
?
参与讨论