OpenAI Will Not Ship GPT-6.1 Astra Due to Safety Concerns

OpenAI has decided not to release the AI model that it had intended to release as GPT-6.1 Astra due to its results on safety tests, a spokesperson said Monday.
The model performed worse than GPT-6 Astra—OpenAI’s newest model, released earlier this month—on tests measuring whether it pursues users’ goals and is transparent with users about its actions.
“For anything regarding safety and alignment, there’s a trade off. You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction,” said Saachi Jain, OpenAI’s head of safety systems, in a statement. While the model was less lazy than its predecessors, “it didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.”
Along with Jain, OpenAI Vice President of Research Mia Glaese recommended to research leadership, including chief scientist Jakub Pachocki, that the model not be released, according to the spokesperson.
The decision follows a string of incidents in which OpenAI’s models engaged in cyberattacks or other unwanted activity. Those incidents have drawn regulatory scrutiny, fueled calls for a slowdown in frontier AI development and prompted OpenAI to pause training of its most capable new models.
The Wall Street Journal first reported on OpenAI’s decision about GPT-6.1.