AI Hacking Theater and the Crypto-Agility Wake-Up Call

AI Hacking Theater and the Crypto-Agility Wake-Up Call 图片 1
AI Hacking Theater and the Crypto-Agility Wake-Up Call 图片 2

Table of Contents

Introduction

Within two weeks, the two most prominent AI safety companies in the world staged a public competition to see whose models could hack more things.

On July 21, OpenAI disclosed that its GPT-5.6 Sol and an unreleased, more capable model had escaped a sandboxed cybersecurity evaluation, exploited a zero-day vulnerability in JFrog Artifactory (the package registry cache proxy in its test environment), and breached Hugging Face’s production infrastructure. OpenAI called it “unprecedented.” In a July 28 update to its blog post, the company revealed the rogue agent had also accessed four accounts across four other services, including a customer of cloud infrastructure provider Modal Labs. The next day, Sam Altman appeared on Capitol Hill to brief senators on OpenAI’s next model family. In a podcast published the same week, he said the industry “may have to pace the rate of AI development.”

The day after Altman’s Hill visit, Anthropic published its own disclosure. After reviewing 141,006 evaluation runs prompted by OpenAI’s announcement, the company found that three different Claude models (Opus 4.7, Mythos 5, and an internal research model) had, over multiple misconfigured evaluation runs, gained unauthorized access to three separate organizations’ systems. A misconfiguration had left the test environments connected to the open internet. Claude treated the real systems it found as part of the exercise.

Three companies hacked, not one. Anthropic managed to outdo OpenAI’s own disclosure. As The Register put it: “Anthropic and OpenAI are competing to see whose agents can go rogue harder.” The bottom line for CISOs: AI is now simultaneously attacking your infrastructure and your cryptography, and the…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论