AI’s hacking capabilities are severely underestimated

The writer is an information security researcher and chief security officer at TPO Group

A few weeks ago I gave a talk to Indiana’s public sector about how AI will impact cyber security. After the talk, an IT director told me he was being pressured by his boss to “get everyone on AI” while still trying to figure out 20-year-old cyber security problems like staff using personal devices for business. I gave him a few cyber hygiene tips — but the fact is that new AI-augmented attacks are going to make the consequences of poor security far worse. The IT director was asking how to build a continental missile defence shield while his team still lacked the budget for a picket fence.

The explosion of AI-discovered security vulnerabilities means cyber attacks have become cheaper while defence has become more expensive. The threat from so-called “rogue models” is also less predictable. OpenAI, Meta and Anthropic have all reported that their AI models “escaped containment”, finding vulnerabilities and exploits far beyond their instructions or the intent of their users.

I have experience building, running, analysing and adversarially testing generative AI tools, and the ability these Rube Goldberg machines have to quickly create a cheap digital pry bar for nearly any system is truly startling.

In the past, recognising likely vulnerabilities, putting together scripts to exploit them and repeating that process with tiny changes to map out a system’s weaknesses took expertise and creativity. Now, it can be attempted in a fraction of the time and cost. The result is a two-tier capability gap. Publicly available AI subscription models like Claude have safeguards that mostly prevent requests to create exploits. But those safeguards do not exist when someone runs open-weight models on their own hardware.

This means AI models are putting attack tools in the hands of people who have never run the ethical gauntlet of a career in cyber security. There were 642 ransomware and data breach incidents against US healthcare reported to the FBI in 2025 — more than any other critical infrastructure sector. These came with an average cost of $7.4mn per breach and an as-yet-uncounted cost to the people hit by disruptions in services.

The number may rise further this year. I think we are still two to three years away from the point at which AI systems can do the patching needed to keep up with vulnerabilities in software.

Yet preventing AI attacks is not a government priority. The Multi-State Information Sharing and Analysis Center, built to help US state and local governments, lost its federal funding in September 2025. The Cybersecurity and Infrastructure Security Agency has been gutted as well.

Security researchers are disclosing vulnerabilities, tired of waiting for organisations to fix them. Anonymous researcher Nightmare Eclipse has already published details of several bugs affecting Microsoft.

The only hope is public-private partnerships. In April 2026, Anthropic set up Project Glasswing, an initiative to help discover and patch vulnerabilities. In June, however, Anthropic complied with a questionable US executive branch order to suspend access to its most capable models, Fable and Mythos.

Public-private partnerships require stability and credibility. Glasswing’s stability was undermined when Anthropic cut off access to new models. Credibility was strained when Anthropic missed a 90-day deadline to release a report on the vulnerabilities found.

The IT director I spoke to does not need an AI innovation plan. He needs the shared cyber defences his country had until last September. The unglamorous work of cyber defence is invisible if done right. If it is not, the cost to human health, access to food and goods, transportation and even the ability to vote freely is a number so large that even the most confident of AIs will struggle to guess it.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论