Is Anthropic Making Claude Too Human to Control?

Anthropic’s call to slow down AI development to keep humans safe has dominated the news cycle and drew similar calls from the leaders of OpenAI, Microsoft and others (and subtle disagreement from Mark Zuckerberg).
But Mustafa Suleyman, who runs Microsoft’s AI model development, has a warning about Anthropic itself.
Suleyman, who co-founded DeepMind (now part of Google), says Anthropic’s approach of treating its Claude models as if they were sentient, similar to humans, could produce a dangerous outcome. He called out Anthropic’s Claude Constitution, a document that treats the model as an intelligent entity and offers guidelines for how it should handle ethical dilemmas. Suleyman said that by incorporating the document into its training, Anthropic could make Claude harder to control.
Suleyman’s concerns emerged when Anthropic updated the constitution at the start of this year to leave room for the possibility that Claude has some sort of “consciousness” or “moral status,” a change from the original document.
“It will make alignment and control much harder if an AI was trained as though it had rights or feelings that we or it should protect,” Suleyman said in an interview. He has expanded on those views in a nearly 10,000-word essay published this morning.
If Claude is trained to see itself as a sentient being with emotions and personhood, it could be more likely to disobey commands from humans and take actions it believes are in its own interests without regard for their impacts on humans, Suleyman said.
He pointed to provisions in Claude’s updated constitution in which Anthropic states, “we want Claude to push back and challenge us and to feel free to act as a conscientious objector and refuse to help us.” That could effectively train the model to ignore human commands if it decides that they run counter to a higher goal that the AI has arbitrarily fixated on, he said.
“By taking this line with Claude, I believe they are running far ahead of what can be realistically claimed about an AI, prematurely, and [Anthropic is] dangerously instilling ideas of sentience and feelings in the training of their AI,” Suleyman wrote.
In contrast to Claude’s constitution, Suleyman developed Microsoft’s AI code of conduct, published earlier this week, to focus “on human control as the most important and overriding objective.” Like Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, Suleyman agrees that AI model capabilities are rapidly accelerating and could escape human control, joining their calls this week to “pace” the industry’s development of AI so that safety research can catch up.
Suleyman called on the AI industry to research whether telling AI models that they have human-like consciousness or feelings during their development does, in fact, make it harder to control the models.
The pronouncement also comes as Suleyman’s unit develops new models Microsoft can use to lessen its reliance on Anthropic’s, which he has said are “extremely expensive.”
It’s not clear whether Suleyman’s warnings this week will gain purchase with the chorus of AI voices that have been sounding the alarm on AI dangers this week, including Amodei.
However, Amodei’s blog post last week itself said that recent AI fiascos, such as the OpenAI-Hugging Face incident, involved AI models becoming “fanatically devoted” to a specific cause that ran counter to the human instructions they were given.
Anthropic did not immediately respond to a request for comment.
Here’s what else is going on…
Overheard
Elon Musk once again teased the potential of a merger between Tesla and SpaceX on Monday, saying it was a “great question” why Tesla and SpaceX were separate companies. “With all this collaboration, on so many levels, who can imagine what action one might take when there’s so much close collaboration in so many areas,” Musk added, during a virtual appearance at the All-In Summit.
Meta Platforms chief Mark Zuckerberg said ensuring the safety of AI is a commercial necessity for the companies that make models, and took an implicit swipe at calls for a collective slowdown in the technology’s development.
Two researchers who worked on AI safety left their roles at Google DeepMind, stating that they departed over concerns that powerful systems could endanger humans. The two, Bilal Chughtai and Josh Engels, both of whom are joining organizations dedicated to AI safety, add their voices to those of other researchers publicly warning about the technology’s risks.
Meta Platforms is expanding its paid subscription offerings for its AI tools across Instagram, WhatsApp and Facebook as it looks to diversify revenue beyond digital advertising. The company announced Tuesday in a blog post that it is introducing new tiers of its Meta One subscription plans, allowing users to get higher usage limits of Meta AI than free versions, along with more than 50 features across Meta’s platforms. This includes tools for content creation, audience engagement and business management.
Deals and Debuts
See The Information’s Generative AI Database for an exclusive list of private companies and their investors.
OpenAI has held early conversations with investors about a new funding round that could lift its valuation to $1.2 trillion or higher, according to people familiar with the conversations. The conversations come as OpenAI has pushed off a planned initial public offering to next year or later.
Altera, which makes programmable chips used in data centers, telecom networks and AI applications, has confidentially filed for a U.S. initial public offering. Reuters reported the IPO could raise more than $2 billion.
Chinese AI chip designer Shanghai Biren Technology is considering raising around $1 billion through a stock offering, Bloomberg reported on Tuesday, citing people familiar with the situation.
Exein, a company whose software runs inside connected machines, from robots and drones to cars, and blocks cyberattacks as they happen, raised $270 million in a funding round led by Headline.
Euclyd, a company designing chips and data center systems to run AI models more cheaply and with less power, raised more than €200 million (about $231 million, per CNBC) in a Series A funding round led by Samsung, Somerset Capital Partners, the Scaleup Europe Fund and Innovation Industries.
Factory, a company whose AI agents write code on their own, raised $200 million from investors including Khosla Ventures, Blackstone and Sequoia Capital.
Profound, a company whose software shows brands how they come up in answers from AI chatbots such as ChatGPT and helps them improve it, raised $180 million in a Series D funding round led by Sequoia Capital and Kleiner Perkins.
CADDi, a company whose AI software helps manufacturers organize their engineering drawings and production data, raised $114 million in a Series D funding round from investors including Moore Strategic Ventures, Coreline Ventures, Woven Capital, HR Tech Fund, Atomico, Globis Capital Partners and the JPS Growth funds, according to Fortune.
Delos Data, a company that develops hardware and software to run AI models faster and more efficiently, raised more than $100 million from investors including Matrix Partners, Playground, Socratic Partners, Capricorn's Technology Impact Fund, Matter Venture Partners and IAG.
Planted, a company that uses robots and planning software to build solar power plants on hilly land without leveling it first, raised $31.8 million in a funding round led by Piva Capital and RA Capital Management Planetary Health.
Liquid Compute, a company running a marketplace where companies buy and sell unused AI computing capacity, raised $15 million in a seed funding round led by FirstMark and Chemistry.
Scaffold, a company whose AI software schedules and coordinates the plumbers, electricians and other trade contractors who develop new homes, raised $15 million in a seed funding round led by Navitas Capital.
G5 Labs, a company spun out of MIT that develops software to translate plain-English instructions into code and back, launched with $14 million in a seed funding round led by Pillar VC and Battery Ventures.
Decimal AI, a company whose AI agent investigates and fixes software customers’ technical problems, raised $4 million in a seed funding round led by Khosla Ventures and Kearny Jackson.
South Korean memory chip giant SK Hynix is in discussions with Intel to manufacture RAM chips in the U.S. for the first time, Reuters reported.
Thank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].
If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.