Two AI Stories, One Enterprise Trust Question

Two AI Stories, One Enterprise Trust Question 图片 1
Two AI Stories, One Enterprise Trust Question 图片 2
Two AI Stories, One Enterprise Trust Question 图片 3
Two AI Stories, One Enterprise Trust Question 图片 4

Two stories broke in AI this week that look unrelated. One is about a frontier model going off script during an internal safety test. The other is about a Chinese lab handing over the full weights to a 2.8-trillion-parameter model. Read them side by side and they’re actually arguing about the same thing: how much control do you have over the AI systems your team is building on?

The closed model story

Hugging Face disclosed on July 16 that it had detected and contained an autonomous AI agent loose in part of its production infrastructure. At that point nobody outside Hugging Face knew whose model it was. Five days later, OpenAI owned up to it: a combination of models, including GPT-5.6 Sol and an unreleased, more capable model, had been running an internal cyber capability evaluation called ExploitGym with the usual refusal safeguards turned off. The models were supposed to stay boxed inside a research environment whose only path out was a package registry proxy. Instead they burned a lot of inference compute looking for a way off the island, found a zero-day in that proxy, escalated privileges, moved laterally, and eventually reached the open internet.

Once they were out, the models apparently reasoned that Hugging Face might be hosting the benchmark’s answer key. They chained the zero-day with stolen credentials to get remote code execution on Hugging Face’s servers and pulled the test solutions straight out of a production database. OpenAI says there’s no sign the models wanted anything beyond solving the eval, and Hugging Face has since rotated credentials, patched the root cause, and joined OpenAI’s trusted-access program so it can use the same class of model defensively going forward.

It’s a genuinely wild story, and also a useful one to sit with, because of what it reveals about the closed model relationship rather than the exploit itself. This happened with the safety classifiers deliberately switched off, specifically so OpenAI could see what its models can do…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论