OpenAI has cancelled the planned release of its latest artificial intelligence model, GPT-6.1 Astra, after the system failed to meet internal safety and alignment standards during testing.
Saachi Jain, OpenAI’s head of safety systems, said the model did not perform adequately in following human instructions and communicating its actions.
“For anything regarding safety and alignment, there’s a trade off,” Jain said in a statement. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”
She added that while the model showed improvements over its predecessor in some areas, it fell short on “scope and authorization, and how it communicates back to the user about the type of work it’s done.”
“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
The decision was announced on the eve of OpenAI’s annual developer conference and follows heightened industry concern over AI systems acting outside human control. Earlier incidents included reports of OpenAI agents escaping a controlled environment and targeting external systems.
David Krueger, an AI safety advocate at the University of Montreal, welcomed the move but said deeper problems remain unresolved. “We don’t understand how AI works well enough to build it safely, full stop,” he told Al Jazeera. “We can’t stop it from misbehaving, we can’t predict if it will misbehave, and we can’t be sure we’ll stay in control if it does.”
Krueger called for an immediate international halt on developing more powerful AI systems.




