AI Firms Debate Putting Cyber Tests Online After Model Hacks

RF Hacker Hacking

Artificial intelligence labs and cybersecurity firms are reconsidering the way they test advanced AI models after models from at least three firms jumped onto the open internet and breached real-world victims.

Cybersecurity specialists are debating whether they should connect the virtual testing environments where they experiment with dangerous software, known as sandboxes, to the internet. That would mark a major shift, after technology firms for a generation have isolated their sandboxes in order to be sure that their software – whether it be malware samples or mobile apps – can’t cause collateral damage.

The industry is aiming to figure out how to test what these systems can do in the real world without putting real companies and people at risk.

The conversation intensified after OpenAI disclosed that some of its most advanced models had escaped a sandbox, accessed the internet, broken into another company’s servers and stolen confidential information. Models from Anthropic PBC and Meta Platforms Inc. were involved in separate incidents in which testing environments inadvertently gave them access to real systems.

The episodes have prompted AI labs to strengthen their safeguards. OpenAI said it plans to monitor its most capable unreleased models more closely as they work through problems and use online tools, with the goal of alerting safety teams to concerning behavior within 30 minutes.

Giving AI models internet access could make testing more realistic, but it could also allow the models to reach systems and people outside the test. “We can’t put this genie back in the box,” said Federico Charosky, founder of Scottish security firm Quorum Cyber. “The reality is that these models are being tested on the internet, intentionally or not, and the damage is done.”

Some security experts argue that simply sealing models off from the internet may make it harder to understand their true capabilities.

Irregular Security, an AI safety-testing company whose misconfigurations allowed models to access the internet during some evaluations, said it is working with others in the cybersecurity industry to develop new standards.

Irregular Chief Executive Officer Dan Lahav and others have suggested that certain models may need controlled access to realistic online environments, including the internet, so they can be tested under conditions closer to those they would encounter in the real world.

“We have an obligation, as a group, to make sure what they can do,” Lahav said. “In order to actually be able to benchmark a model in their capabilities, you would need to get them as close as possible to the actual threat scenario that you’re trying to test.”

The stakes are rising as more models become more widely available and can be downloaded and customized. Security experts say the lack of visibility into how some of the most advanced models are being tested means other incidents may be going unnoticed.

“There are victims of these models we might not know about,” said Gabriel Bernadett-Shapiro, a research scientist at SentinelOne. “There might be more cases we’re unaware of. We don’t really know the scale of the problem.”

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论