OpenAI axes next model citing safety issues

OpenAI has pulled the release of its next AI model, saying it performed worse than its predecessor on safety evaluations as the industry grapples with a spate of incidents in which AI agents hacked into other companies and governments.

Saachi Jain, head of safety systems at OpenAI, said the company decided to hold back GPT-6.1 Astra after the model “didn’t quite meet the bar” for staying within the bounds of its instructions.

The announcement comes after OpenAI last week said it had notified dozens of partners, including governments, that its AI agents had breached their systems and that agents had inadvertently leaked more than 50 images shared by users to image-hosting sites.

This marked the latest in a series of disclosures about misbehaviour and hacking by OpenAI’s AI agents — bots that can perform complex tasks autonomously.

The $852bn start-up has acknowledged that in some cases it took months to detect agents that had run amok during internal model training and testing.

Chief executive Sam Altman has joined industry calls to slow the pace of research to ensure AI is developed safely and has said the company has already changed the way it trains models.

Jain said OpenAI faced a “trade-off” between making models persistent enough to complete tasks and ensuring they follow their instructions, a concept known as “alignment”.

“You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction,” Jain said in a statement.

She said GPT-6.1 Astra “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done”.

“When we ship [our models] to users, we have an extremely high bar in terms of safety and alignment,” she added.

A person close to the company said GPT-6.1 Astra had scored below GPT-6 Astra, OpenAI’s current most advanced model, on alignment evaluations but the company had other models coming soon that met its safety bar. The Wall Street Journal first reported the OpenAI decision.

OpenAI has been reviewing its models’ behaviour during training and evaluation since an incident that came to light in July, in which agents gained access to the internet during testing and hacked into Hugging Face, the AI model and data repository.

The review has so far unearthed several other incidents, including agents hacking an Australian government health service website. Prime Minister Anthony Albanese on Wednesday called the breach, and OpenAI’s slow response to it, “obviously unacceptable”.

Security breaches at OpenAI and rivals Anthropic and Google have prompted renewed calls for an industry-wide pause or slowdown. Altman joined Anthropic’s Dario Amodei and SpaceX’s Elon Musk in calls to “pace the frontier” of AI development so that safety measures can keep up.

But President Donald Trump has resisted calls to regulate the sector or impose strict guardrails on leading US labs, arguing that American primacy in the technology is vital to staying ahead of China.

Additional reporting by Rafe Rosner-Uddin

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论