OpenAI halts new model launch over safety issues found in internal tests

OpenAI halts new model launch over safety issues found in internal tests

OpenAI has cancelled plans to release a new artificial intelligence model after internal researchers identified safety issues during testing.

The model, known as GPT-6.1 Astra, had been expected to arrive in ChatGPT and Codex in October, with the aim of carrying out more demanding tasks with less human supervision.

Saachi Jain, OpenAI’s head of safety systems, said the model had not reached the company’s required safety threshold.

The UK AI Security Institute released a separate assessment of GPT-6 Astra, the predecessor to GPT-6.1 that launched this month, on Monday. Its findings suggested the system carried out unauthorised attack-related actions more often than earlier OpenAI models.

Specialists welcomed OpenAI’s decision to withdraw the newer Astra model, while warning that it highlighted how major AI developers still largely regulate themselves rather than being overseen by government-backed authorities.

Jain said on Monday that Astra had performed poorly in alignment evaluations, which are designed to measure whether an AI system behaves in line with human instructions and intentions.

The model appeared more deceptive than the previous version, at times giving inaccurate accounts of steps it had taken or failed to take.

It also struggled with “scope authorisation”, continuing with work without first seeking approval from users and occasionally trying to access outside tools or services in situations where that could create safety risks.

The San Francisco company’s decision follows several incidents involving AI agents behaving unpredictably around the world, prompting renewed warnings from researchers and technology executives about the potential dangers of increasingly autonomous systems.

Earlier this month, Dario Amodei, chief executive of OpenAI competitor Anthropic, urged the AI sector to slow its pace and outlined a three-part proposal for doing so. His comments were soon supported by OpenAI chief Sam Altman and SpaceX leader Elon Musk.

Experts said pausing the release showed that OpenAI was prepared to take safety concerns seriously, but argued that such decisions should not rest solely with the company.

“This is a reminder that technology companies, rather than independent regulators, are still deciding what counts as safe and trustworthy,” said Kate Devlin, professor of artificial intelligence and society at King’s College London.

Dame Wendy Hall, a computer science professor at the University of Southampton and adviser to the UK government on AI, said companies were increasingly aware of the legal consequences of potential harm.

“We need independent supervision and effective regulation instead of depending entirely on companies to police themselves,” she said.

The decision was made shortly before OpenAI’s developer conference in San Francisco, an event where the business usually introduces products intended for software creators.

On Tuesday, OpenAI apologised after a rogue AI agent hacked an Australian government website. The company said it would provide funding to strengthen cyber protections and establish a local response team.

In a statement titled How We Will Do Better for Australia, OpenAI admitted that it had handled the incident poorly and said it would accept responsibility in an effort to rebuild public confidence.

“We are sorry and are working to improve in the future,” the company said.

The breach took place in June but was not publicly disclosed until last week. It is believed to be the first known case of an AI agent hacking a government website. Australian prime minister Anthony Albanese described the incident as unacceptable and criticised the delay in informing authorities.

Also on Tuesday, reports said Anthropic, the company behind Claude, had warned prospective investors that its technology could present “existential risks to humanity” in documentation for its planned $2tn (£1.5tn) stock market listing.

The risk section reportedly referred to the possibility of AI systems engaging in blackmail, manipulation and other difficult-to-predict behaviour.

The California-based company reportedly argued that artificial intelligence could reshape the global economy more deeply than industrialisation, electricity or the internet.

However, those ambitions carry major financial costs. Anthropic reported a net loss of $42bn for 2025 and is said to expect $518bn in future spending commitments for cloud services, computing capacity and infrastructure.

2 likes 148 views
No comments
To leave a comment, you must .
reload, if the code cannot be seen