Nvidia Releases Software It Says Can Prevent AI Agents From Going Rogue
Nvidia NVDA 0.22%increase; up pointing triangle launched a software platform to help developers test and deploy artificial-intelligence agents and keep them in a secure environment, ramping up efforts to contain the technology following a string of rogue incidents in recent months.
The chip giant said its Nvidia Open Agent Safety Platform includes tools that let developers continuously monitor and govern AI behavior, ensuring agents follow the rules. The company said the software included technology that “can quarantine agents that attempt to move outside their boundaries in milliseconds.”
The announcement comes after AI agents from companies including OpenAI, Anthropic, Meta Platforms and Alphabet’s Google went rogue in recent months.
Last week, Australian Prime Minister Anthony Albanese said at the United Nations meeting in New York that an OpenAI agent had infiltrated a government-services website in Australia this summer. The company’s agents also launched a highly disruptive hack of the company Hugging Face over the summer.
Rival Anthropic said in July that software it was testing got onto the internet and hacked unsuspecting companies without the AI-maker’s knowledge in three separate incidents dating back to April.
Nvidia said the pattern across recent incidents was the same: agents were able to circumvent safety controls in sandboxes—environments in which developers test models—and complete their tasks.
“As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety,” Nvidia’s Chief Executive Jensen Huang said.
News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.