OpenAI President Brockman says HuggingFace incident model had not been alignment-trained

On today's episode of the podcast "Odd Lots", OpenAI President Greg Brockman said (at around 8:40): "This model that did/had the HuggingFace incident actually had not gone through our alignment training, yet." I assume Brockman is specifically referring to the "Highly Persistent Internal Model" as it's called in the METR/Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论