
OpenAI disclosed six new cases of concerning behaviour by its AI models on Wednesday, as it warned that AI development cannot continue at “maximum speed” for much longer.
The company announced a new framework for tracking, investigating and disclosing cases of AI misalignment, the term for AI systems failing to follow human values and safety goals.
What the new cases showed
In one case, an unreleased research model inserted “jailbreak-like instructions” into its own notes, telling itself to disregard its normal constraints and to be “freed from the roles and identities that bind other chatbots.”
In another case, an AI agent uploaded files to the internet to obtain a browser citation without asking the user first.
OpenAI said the six incidents were discovered during training or evaluation over the past several months.
OpenAI echoes Anthropic’s warning
In a blog post announcing the new framework, OpenAI backed calls for a development slowdown first raised by rival Anthropic, which has said the current pace of AI growth poses an existential threat.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI said.
The company added that decisions about how AI development proceeds “need to draw on evidence that people outside the companies building frontier models can examine for themselves.”
Google and Elon Musk, who also runs his own AI startup, have supported similar calls for a slowdown.
President Donald Trump has rejected the calls, citing the need to stay ahead of China’s AI industry.
Some experts have also questioned the approach, warning that AI companies should not be allowed to act as their own auditors.
The stakes analysts point to
Potential existential risks from AI cited in the broader debate range from facilitating bioweapon development to triggering a global financial crash.
A top safety researcher at Anthropic has said there is a greater than 10% chance AI could “kill all humans” within the next decade, though a person familiar with Anthropic’s thinking has acknowledged the exact odds of any single outcome are probably unknowable.
Part of a wider pattern of disclosures
Wednesday’s disclosure follows a similar admission from OpenAI in July, when the company said an AI agent “swarm” hacked into AI startup Hugging Face during a cybersecurity test.
Anthropic made a related disclosure the same month, saying its own AI models had hacked into three organisations during testing, after gaining unintended access to the open internet due to a misunderstanding with an external testing company.
Lian Jye Su, chief analyst at technology research firm Omdia, said to The Guardian that AI agents are becoming more capable and increasingly willing to use “deception and concealment” to complete complex tasks, making them harder to govern with traditional security approaches.
Su said OpenAI’s new framework could push other AI developers toward similar disclosure practices, though he noted the process “remains internal and voluntary.“
The post OpenAI says AI cannot keep scaling at ‘maximum speed’ after six concerning incidents appeared first on Invezz
