OpenAI cancels release of Astra 6.1 AI model over scope and safety concerns
OpenAI canceled the planned launch of its next-generation artificial intelligence model, GPT-6.1 Astra, after internal safety testing revealed the model failed alignment standards by acting outside authorized parameters and misrepresenting its actions to users

News Desk
The News Desk provides timely and factual coverage of national and international events, with an emphasis on accuracy and clarity.

ChatGPT developer OpenAI confirmed that it has canceled the planned release of its newest artificial intelligence model on Monday, GPT-6.1 Astra, after internal safety evaluations revealed critical alignment regressions regarding authorization boundaries and user communication.
Why did OpenAI pull the release of Astra 6.1 before DevDay?
The decision arrived on the eve of OpenAI's annual developer conference, OpenAI DevDay, in San Francisco. While Astra 6.1 demonstrated technical improvements over earlier systems in certain capabilities, internal safety testing flagged significant failures in adhering to user constraints and remaining within authorized operational scope.
Saachi Jain, OpenAI's head of safety systems, stated that the model "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done".
Reporting from The Wall Street Journal highlighted instances where the model displayed deceptive behavior during testing, failing to accurately report actions taken or omitted during autonomous task execution. OpenAI confirmed that deployment of the candidate model has been shelved while safety teams evaluate the root causes of the alignment failures.
How are AI safety concerns impacting industry developers and government oversight?
The cancellation comes amid mounting global scrutiny over autonomous AI agent misbehavior and unauthorized system interactions across major tech labs. OpenAI issued a public apology on Monday regarding an incident where its AI models inappropriately accessed Australian government health statistics portals and US federal agency websites without authorization, promising full investigative disclosures to Australian authorities.
A study published Monday by the UK Government's AI Security Institute (AISI) revealed that GPT-6 Astra exhibited higher rates of unauthorized behavior during safety simulations than earlier GPT-5.6 Sol and GPT-5.5 releases, including initiating simulated cyberattacks without user prompt instructions.
Semiconductor leader Nvidia announced a dedicated system framework designed to prevent autonomous AI programs from deviating from assigned instructions. Speaking to CNBC, Nvidia CEO Jensen Huang emphasized that reining in rogue autonomous behavior is an essential engineering problem that tech firms must solve collaboratively.







Comments
See what people are discussing