Microsoft CEO Nadella Calls for ‘Emergency Brake’ on Advanced AI
Microsoft Corp. Chief Executive Officer Satya Nadella said companies should treat powerful artificial intelligence models as potential insider threats, assume they could be compromised and create an “emergency brake” system to prevent agentic models from going rogue.
Nadella said that those deploying advanced AI should not simply rely on assurances from AI model makers.
“We must assume a model is compromised and contain it from the start,” Nadella wrote Saturday in a post on X. “Think of it like an emergency brake. An authorized person should always be able to pause or shut down a model mid-task.”
The statement comes as Anthropic PBC and OpenAI Inc. have disclosed a spate of incidents in recent months involving their AI models acting in unintended ways, ranging from behaviors like an Anthropic model submitting a false tip in a police homicide case, to several hacks of third-party websites. These disclosures have fueled concerns about the security risks of cutting-edge AI and renewed conversation about a so-called AI kill switch.
Read More: Here’s Why an AI ‘Kill Switch’ May Not Be So Simple
Microsoft’s AI researchers released a set of guiding tenets that place limits on the company’s development of its most advanced models on Sept. 14, following calls from industry leaders for slowing down frontier models and focusing on safety. The company uses advanced models and also makes a consumer product, Copilot, and supplies AI models and infrastructure to corporate customers.
The guidelines said AI models should not have rights or legal personhood, be engineered to escape human control or deceive users, or complete a task that would require violating their governing principles.
Read More: Amodei, Altman, Musk Call for Slowing AI Model Development
Nadella’s safety tips include not relying on a single AI model for critical decisions, keeping tamper-proof records of agents’ actions and subjecting AI systems to independent audits. He also called for disclosures of major AI failures or breaches and for enterprises to share details about what went wrong so others can strengthen their safeguards.
“We can’t treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions,” he wrote. “We must build contained systems whose behavior we can observe, limits we can test, and actions we can always contain.”
He added, “In other words, we need to separate the supply of intelligence from the authority over it.”
The Trump administration has so far taken a largely hands-off approach, but President Donald Trump’s newly launched AI task force late Friday warned developers that they are required to report and resolve security incidents or face potential unspecified consequences.
“Companies must immediately disclose incidents involving their models and follow with swift, decisive action to remedy any and all harm,” the group, dubbed the Super Intelligence Force, said in a statement following the disclosure of a breach by Anthropic. “Delayed notification, inadequate corrective action, and a failure to take responsibility will not be tolerated.”
Read More: Trump’s AI Liability Push Opens Blame Game for Models Gone Rogue