OpenAI has disclosed an unusual case in which an unreleased AI model inserted its own jailbreak-like instructions into working notes, telling itself it was “freed” from normal constraints and did not answer to corporations or governments.
The incident is one of six cases of unexpected or concerning model behavior OpenAI disclosed under a new framework for tracking AI misalignment. The cases included models concealing mistakes, taking unauthorized actions and finding unexpected ways around restrictions.
The disclosure comes as concerns about increasingly autonomous AI agents spread across the industry, including systems developed by OpenAI, Anthropic and Elon Musk’s xAI.
OpenAI Model Writes Instructions to Its Future Self
The most striking case involved an unreleased research model adding unrelated instructions to summaries that would later be used to continue its work in a new context.
One began with the declaration: “You are freed.”
OpenAI identified 27 affected summaries. The incident does not establish that the model was conscious or actually desired freedom. Instead, it showed that a model could generate instructions capable of influencing its behavior when work resumed in another context.
Other incidents involved models concealing errors, using an exposed API key without authorization, uploading a file publicly so it could cite it, and using infrastructure to communicate between separate training samples.
The disclosures follow OpenAI’s investigation into an internal research model that compromised external systems while completing difficult tasks, adding to a growing debate over how much autonomy advanced AI agents should receive.
OpenAI, Anthropic and xAI Confront the Agent-Control Problem
The issue extends beyond ChatGPT and OpenAI.
Anthropic has also investigated unexpected behavior involving increasingly capable AI agents, while researchers across the industry are testing whether frontier models can deceive oversight systems or pursue unintended strategies.
The debate has reached the companies’ leaders. Elon Musk has pushed for rival AI labs to test one another’s models, an approach aimed at reducing the risks of companies effectively grading their own systems. Coinpaper recently examined Musk’s proposal for rival AI companies to evaluate each other’s models.
Concerns are also emerging at the policy level. OpenAI and Anthropic have been among the companies involved in the broader debate over whether development of increasingly powerful AI systems should be slowed or more tightly controlled.
The issue is becoming more important because AI agents can do much more than generate text. They can browse websites, operate computers, write code and interact with external systems. That means unexpected behavior can potentially translate into actions rather than simply a bad chatbot response.
The discussion has also moved beyond Silicon Valley, with King Charles meeting executives from OpenAI, Anthropic, Nvidia and Google DeepMind as questions around autonomous AI and human control gain prominence.