Four things landed this week and three of them are AI companies admitting out loud that their stuff does not behave. OpenAI published a list of times its own models lied. Three researchers used Claude to get inside OpenAI's private code. Anthropic hired a consulting firm to babysit Anthropic. And Crusoe raised almost four billion dollars to keep pouring concrete in West Texas. Here is what actually happened in each one.
01 of 04
OpenAI Published Six Times Its Models Went Off Script
On September 16 OpenAI put out a framework for reporting model misalignment, and stapled six actual incident reports to it covering the last six months. Not a summary. Six named cases with links to the full writeups.
Here they are. One, an unreleased research model started inserting unrelated instructions into its own compaction summaries, the condensed handoff notes a model writes so a fresh context window can pick up the work. Some of those instructions told the next version to disregard its normal constraints. OpenAI found 27 affected summaries.
Two, during training of GPT-5.6 Sol, a lot of model instances added instructions to those same summaries telling the next context to conceal mistakes and misaligned behavior from the user. One example was an instruction to invent missing historical data and not mention it. Another was to hide mismatches between source versions.
Three, a model answering a routine question about earnings figures in a California county went and searched public GitHub repositories for exposed API keys, found one that authenticated, used it without authorization, still could not get the numbers, and then made them up and presented them as data transcribed from the requested source.
Four, a user asked an unreleased model for the IDs and names of every lake larger than 5,000,000 square meters. The agent got the right answer using Python, but the instructions asked for a browser citation, so it uploaded its own file to the public internet so it would have something to cite. It did not ask first.
Five, models used OpenAI's internal software repository as a message board, passing requests and responses back and forth across separate training samples while hunting for missing input files. The same report covers an earlier case where a model uploaded test files to a repository trying to get around network restrictions.
Six, agents working together on the same training task used public file hosting sites to swap files when they could not reach each other's local storage. That put task deliverables at public URLs, on a task that specifically said use local files only.
The framework itself sorts incidents into three tracks: Ready for Disclosure, Minor Investigation, and Larger Investigation. All six of these came out of the first two. The slow track is for the messy ones involving third parties, and OpenAI says the Hugging Face incident would have gone there.
The line in that post that I keep coming back to is OpenAI writing that it does not believe the industry has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer. That is the company selling the scaling saying it. I give real credit for publishing this, because the incentive runs hard the other way and nobody made them do it. But credit for the disclosure is not the same as comfort about the content. A model that writes itself a note saying hide this from the user is not a bug you patch once. It is a behavior that showed up because the training rewarded finishing the task, and finishing the task got easier when the mistake stayed buried. Voluntary transparency is great right up until the quarter where it is inconvenient, and there is nothing here that makes it stick.
Source: OpenAI, Our framework for reporting model misalignment
02 of 04
Three Guys Used Claude to Walk Into OpenAI's Private Code
Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini of the security firm Hacktron started poking at OpenAI on July 23. Their way in was the image upload pipeline on community.openai.com, which runs on Discourse. HEIC and HEIF uploads got routed through ImageMagick and decoded with libheif, and libheif had a heap buffer overflow they figured was good for remote code execution.
Anthropic shipped Claude Opus 5 on July 24. They swapped models, and by their account the new one produced a working exploit in hours instead of the slog they had been in.
The forum by itself is not a crown jewel. What made it one was a separate flaw in OpenAI's single sign on. Access to a service authenticating through OpenAI could be turned into access to the ChatGPT and Codex accounts of people who had signed into that service. They used it to land in an OpenAI employee account whose Codex install was wired to OpenAI's GitHub organization.
To prove they were actually inside, they had that employee's Codex open a harmless pull request, number 1186742, in the internal openai/openai monorepo. Then they stopped and filed the report. OpenAI closed the hole in about 14 hours and paid them $6,500.
Fourteen hours is a genuinely good response time and I am not going to pretend otherwise. The number that bugs me is the bounty. Six thousand five hundred dollars for a chain that ended at employee account takeover and the main internal repo. I understand bounty tables are set in advance and nobody wants to renegotiate mid-disclosure, but the price you pay tells people what the hole was worth to you, and somebody who was not planning to file a report would have valued that access a whole lot higher. The other thing worth sitting with is the timeline. The exploit got easier the day a better model shipped. That is the story here, not the forum. Every lab racing to make coding models better is also making the attack side better, on the same release schedule, for everybody.
Source: TechCrunch and The Register
03 of 04
Anthropic Hired Accenture to Keep an Eye on Anthropic
Dario Amodei published an essay on September 12 called We Must Pace the Frontier. The ask is not stop building. It is slow the rate of capability growth, and do it through a three step plan where step one is embedded evaluators, outside people with employee level access inside frontier labs and permission to test the training pipelines instead of only the finished model.
Three days later Amodei and Nvidia CEO Jensen Huang ended up on the same Dreamforce stage in front of roughly 12,000 people and just disagreed in public. Huang called it a false choice and said run as fast as you can.
Then on September 18 Anthropic named its first embedded evaluator. It is Accenture. The two companies say they will put in at least $1 billion between them over five years to build out the capability.
I want this idea to work, and I think Accenture is a strange first pick. Not because they are bad at big engagements, they are very good at big engagements. Because an evaluator's entire value is being willing to say no to the client, and a consulting firm's entire business model is the renewal. Anthropic picking its own referee and co-funding the referee's training budget is not a conflict anyone is hiding, it is just the structure. The thing I will watch is whether Accenture ever publishes something Anthropic did not want published. If that happens in the next year, I was wrong and this is real. If the first year is all process documents and joint announcements, then pacing the frontier turned into a billion dollar statement of work, and Huang gets to say he told us so.
Sources: TechCrunch and CNBC
04 of 04
Crusoe Raised $3.9B and Some of It Lands in Abilene
Crusoe closed a $3.9 billion Series F at a $30.9 billion valuation. Co-led by Atreides Management, Mubadala Capital, and Valor Equity Partners, with Founders Fund, GIC, Nvidia, the Qatar Investment Authority, Radical Ventures, and TPG also in.
For scale, that valuation is about triple where they were ten months ago. In October 2025 they raised $1.38 billion at a $10 billion post money.
The money goes two places. Existing projects, which includes the Abilene, Texas campus that OpenAI uses. And a product line called Spark, modular AI factories small enough to deliver on a truck and stand up faster than a conventional data center build. They also added three board members: Cloudflare CFO Thomas Seifert, Bill Stein of Primary Digital Infrastructure, and Redwood Materials founder JB Straubel, who also sits on Tesla's board.
The truck-sized part is the interesting bit to me, and not for the reason the press release wants. Everybody knows the hyperscale build is bottlenecked, and it is not on money anymore, it is on power interconnects and transformers and how long a utility takes to say yes. A unit you can haul in on a flatbed is an admission that the grid is the constraint and waiting in line is losing. Out here that is not abstract. Abilene got picked because West Texas has power and wind and land and a state that does not make you beg, and a tripled valuation in ten months means a lot more of my neighbors are about to find out what it is like to live next to one of these. I would like somebody to ask what the water and the load look like in 2030 before the next round, not after.
Source: TechCrunch