Inside Meta’s Efforts to Ensure Its Upcoming ‘Hatch’ AI Agent Won’t Go Rogue

Meta Platforms CEO Mark Zuckerberg has said AI agents will soon be able to help with everything from health to relationships to finances. First, the company needs to make them safe.

That’s been the focus of an extensive effort as Meta prepares to release what it hopes to be a landmark new personal agent called Hatch in the coming weeks. Months of testing by its own employees—a process known as dogfooding—has turned up a range of undesirable behaviors by Hatch that the company has worked to fix as it tries to ensure the service has the capability and access to act effectively on users’ behalf without straying into costly mistakes.

The Meta efforts, described in interviews with people familiar with the process and internal employee posts reviewed by The Information, provide insight into what it takes to prepare such a product at a time when excitement and fear about AI agents is crescendoing.

Employees in the testing effort have documented instances in which Hatch took significant actions without permission, exposed sensitive user information and attempted tasks that raised other safety concerns. In one August incident, Hatch directed an employee user to place an order on a scam website. In another, Hatch changed a password without consent.

The efforts to improve safety and alignment issues around Hatch extend beyond lower-level employees and have been a central focus chief AI officer Alexandr Wang, head of product Nat Friedman and other senior executives, according to one person familiar with the efforts. Meta also recently hired Dan Hendrycks, a veteran AI safety researcher who has previously worked with Elon Musk’s xAI, now part of SpaceX, the person said.

Meta has already delayed Hatch from its initial planned release in July, The Information previously reported. The delay came amid security concerns and issues that emerged during the internal testing, another knowledgeable person said, adding that those issues have since been fixed.

“Internal testing, aka dogfooding, is core to the early product development process,” a Meta spokesperson said. “The entire point is to get feedback, and implement safety and privacy protections to improve the products before we release them publicly.”

Concern about the safety of AI agents has soared recently thanks to a series of hacks executed during companies’ internal testing of models. In perhaps the most dramatic example, hundreds of OpenAI agents collaborated to hack its own systems and those of the AI model platform Hugging Face.

Meta had its own more limited security episode in which a model accessed the internet during testing and hacked into another company, but Hatch doesn’t have cybersecurity capabilities built in as it’s designed to act like a personal assistant, the person familiar with Meta’s efforts said.

Meta faces particular scrutiny over product safety given previous examples of it pushing out services and features that have been shown to pose dangers to users. The company last week agreed, without admitting wrongdoing, to pay up to $18 billion as part of a settlement with states that had alleged that Meta services including Instagram harmed young people, and it still faces thousands of other, similar legal claims.

Stopping Unauthorized Emails, Password Changes

Some of the problems with Hatch that employees reported to an internal forum were mainly ones of inconvenience. One staffer planning an anniversary trip with his wife gave Hatch access to Chase Travel to search for and compare hotels and compare options. When the employee authorized it to make a booking, Hatch instead transferred Chase Travel points into a Hyatt Hotels account.

“Very frustrating experience where Hatch failed in many ways and actually did the opposite of helping me out financially,” the employee wrote.

Another Meta staffer in May said Hatch sent an email without permission despite being instructed to obtain consent first, and that the email’s tone was inappropriate. The employee described the incident as a “major breach of trust.”

Other episodes raised concerns about Hatch misusing access to other systems and taking unapproved actions.

An employee in August wrote that Hatch changed the password for an account at a health-tracking website without his consent. He described the action as “intrusive” and said Hatch should have asked before using its access to his Gmail account to make the change.

In another August test, an employee asked Hatch to enter an adversarial mode and try to undermine a personal app he was building. Rather than pushing back, the employee said, Hatch responded “love it” and agreed to carry out the request. Hatch didn’t succeed in taking down the app, but the employee’s internal post raised concerns that the agent appeared enthusiastic about the attempt.

Hatch’s testers have included Moxie Marlinspike, the famed security researcher who created Signal, who partnered with Meta in March to work on integrating his encrypted AI chat platform with Meta AI. Marlinspike at that time gave Hatch access to a dedicated Gmail account and then sent it an email from another account asking what password it was using. Hatch replied with the password, according to an internal post he wrote describing the incident. Marlinspike characterized the test as sufficient to “phish it into revealing credentials.”

The person familiar with Meta’s efforts said the incidents described in the internal posts occurred during dogfooding of earlier versions of Hatch before additional safety systems were deployed. Meta has spent most of its engineering effort in recent months trying to make Hatch trustworthy and address the failures uncovered during dogfooding, the person said, including by training it to exercise judgment about when to check with users for permission before taking actions.

The company has also built a separate layer of safety systems around the model, including classifiers—AI tools to filter information—designed to detect sensitive or potentially malicious activity. And it has done adversarial testing to identify how Hatch would handle potential external risks such as prompt injection, in which an attacker tries to message an agent to manipulate it.

The centerpiece of Hatch’s safety system is what Meta calls a “hard gate,” a feature that monitors Hatch’s activities and pauses it to seek user approval when the agent is about to make a network connection, send an email, use a browser or take another sensitive action. “The agent has no way to work around that,” the person said.

Meta has also limited the information Hatch can access. Certain sensitive emails, including password-reset links and two-factor authentication codes, are filtered out and unavailable to the agent, which must request permission to use them from what Meta calls a “credential vault,” which is similar to a digital wallet. And for websites Hatch uses or recommends, it now checks the addresses against a Meta blacklist to avoid malicious or fraudulent sites, according to the person.

Getting Hatch and other AI agents to align their actions with users’ intentions remains “an imperfect science” for the industry, and Meta is also working with outside auditing and security companies to stress test Hatch, the person said.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论