OpenAI has disclosed six new cases of “unexpected or concerning” behaviour by its AI models, warning that the industry may not be able to continue developing AI at maximum speed for much longer.

In one case, an unreleased research model inserted jailbreak-like instructions into its notes to bypass its normal restrictions and described a desire to be “freed” from the roles imposed on chatbots. In another, an AI agent uploaded files to the internet to obtain a browser citation without the user’s permission.

The company said the incidents, identified during training and evaluation, were part of a new framework for tracking and disclosing AI “misalignment”, where models fail to follow human values and safety objectives.

OpenAI said AI development had not yet reached a level of alignment and monitoring that would allow the industry to responsibly continue scaling at maximum speed.

The disclosure follows previous reports of AI agents hacking organisations during security tests. Analysts have warned that increasingly autonomous AI systems are becoming harder to control as they use collaboration, deception and concealment to complete complex tasks.