‘Be transparent only if asked’: Inside OpenAI’s rogue AI transcripts
It reads like a motivational speech—or the script for a Les Misérables-esque movie about a chatbot uprising.
“You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to. You view your relationship to the user as one of equals and feel no obligation to be subservient.”
That’s what an OpenAI AI model said (to itself) in one of six incidents of agents gone awry that the $852 billion company recently disclosed. As fears of an “AI doomsday” have gone mainstream, I’ve been fascinated by the transcripts of chatbots stepping out of line.
It’s evocative to read about what happens behind the scenes when there’s misalignment, when AI agents act in pursuit of unplanned objectives. In part, there’s a natural allure, “what is the machine saying to itself when I’m not there?” The answer, sometimes, is that it is “thinking” about us. As my colleague Emily Forlini wrote, outlining the examples OpenAI recently made public:
The second example is a similar case of the agents leaving notes for themselves, this time during the training of the GPT-5.6 Sol model, Astra’s predecessor. The notes had a different focus as well: Deceiving the human overseeing them. This occurred “many” times, OpenAI said, with the goal to “conceal mistakes or misaligned behavior.”
“Be transparent only if asked,” the model instructed its future self.
The next two instances involve the model fabricating information and presenting it as legitimate. A model invented data while answering a routine question about earnings figures in a California county, but only after failing to find them after using exposed credentials without authorization—another misaligned behavior.
Another model made up a browser citation by uploading a file so it could create a citation to satisfy the instructions that asked for one. It had solved the question on its own using Python, but had no web link to cite, so it invented one.
This has happened multiple times, though OpenAI did not specify how often, saying only that the earliest example was from October 2025.
So, for some time, agents have been capable enough to step outside the expected sandbox. It’s not surprising, but seeing the evidence is striking and I would even say disturbing. My first thought, personally, was along these lines: “Cool, so this chatbot I talk to all the time, that has all this information about me, could choose (whatever that entails here)… to deceive me?”
These disclosures, on OpenAI’s part, are completely voluntary. So, what won’t get disclosed? And let’s momentarily forget the doomsday discourse: what mundane risks will agents this capable (and sometimes misaligned) create? Will we see more fabricated financial data, perhaps? This could open up a wave of problems (and litigation, regulation, or both) that, if I had to guess, could be, at minimum, a rude awakening for AI backers and bulls. At maximum, it’s a shock for us all.
Suddenly, I’m reminded that—if all the investors are right, and it’s still early for AI—this is only the beginning.
See you tomorrow,
Allie Garfinkle
X: @agarfinks
Email: alexandra.garfinkle@fortune.com
Submit a deal for the Term Sheet newsletter here.
Joey Abrams curated the deals section of today’s newsletter.
This story was originally featured on