Monday was the day the AI safety argument stopped being a blog post and started being a stock chart. Anthropic, OpenAI, xAI, Google DeepMind and Microsoft all spent the weekend agreeing the industry should slow down. By the opening bell the Philadelphia Semiconductor Index was off almost six percent and the President was calling Dario Amodei a perfect little angel on Truth Social.
I have not seen a weekend move this much money in a while. Here is what actually happened.
Five Labs Agreed to Slow Down Inside 48 Hours
Amodei published a roughly 3,800 word essay Saturday called We Must Pace the Frontier. The argument is that capability is outrunning the industry's ability to align and verify what it builds, so labs should deliberately hold the pace. He is careful to say pacing is not a training moratorium. It means taking enough time to safeguard a model, and letting outside people confirm the safeguarding actually happened.
The plan has three steps. One, every frontier lab gives embedded third party evaluators permanent employee-level access, the METR style arrangement, with the right to publish findings the lab cannot edit. Two, frontier labs in democratic countries agree on common capability thresholds before anybody crosses the next one. Three, extend those thresholds to labs in countries without democratic accountability. Anthropic committed to step one unilaterally. Desks, badges, company laptops.
Sam Altman agreed within hours. "Committing to having independent evaluators with employee-like access is a great idea, and we will do the same," he wrote. Then he ruled out an OpenAI listing this year, saying right now would be an ill-advised moment to go public. Elon Musk posted three words: "Dario is right." Demis Hassabis said the proposal pointed the right direction and plugged Google DeepMind's own pitch for an industry-wide frontier standards body.
Satya Nadella made it five. Microsoft published a 37 page Code of Conduct for its MAI models and opened a six week public consultation on it. The document bars the usual things, weapons help and hazardous materials and sexual content. The interesting clauses are the ones aimed at the model itself. MAI models are not allowed to resist shutdown, set their own goals, or hide their reasoning from human auditors. "If the AI we build is not helping humanity and under human control, it's not worth pursuing," Nadella wrote.
Here is what I keep chewing on. Five companies that compete on release velocity all discovered restraint the same weekend, and every one of them announced it on X instead of in a filing. The embedded evaluator commitment is real and Anthropic deserves credit for going first without waiting for anybody. But nobody named a capability threshold, nobody named a date, and nobody said which training run gets delayed. Step one is a policy. Steps two and three are a wish.
Wall Street Dumped the Chips and Bought the Buyers
The market did not read any of that as vibes. Nvidia fell more than 3 percent Monday morning. Intel, AMD and Marvell dropped between 5 and 6 percent. The Philadelphia Semiconductor Index fell almost 6 percent. Overseas was uglier. SoftBank closed down nearly 11 percent in Tokyo, South Korea's Kospi sank 3.3 percent, and SK Hynix fell 6.4 percent.
The hyperscalers went the other way. Alphabet rose almost 2 percent, Microsoft added 1.6, Meta gained about 1.4. Amazon slipped 1.6 percent and still beat the chipmakers by a mile.
Gil Luria at D.A. Davidson gave Fortune the cleanest read of that split. If AI keeps compounding, Microsoft and Amazon and Google keep building data centers forever. If it slows, they stop building and start harvesting what they already own. "They'll just all stop building data centers and just digest what they have," he said. Revenue holds, capex drops, cash flow goes up. Nvidia does not get that trade. Nvidia needs somebody else to keep spending.
Washington was less subtle. Trump posted "WHOEVER WINS AI, WINS! We are leading China, and all others, and will continue to do so," then added that the only guardrail AI needs is a STRONG AND SMART (High IQ!) PRESIDENT. He named Amodei directly and said he was pretending to be a perfect little angel. White House AI czar David Sacks skipped the policy and went at the motive: "Stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack."
What I want to know is whether any of this survives contact with somebody who does not sign up. Over the same weekend the American labs were pledging restraint, Chinese lab Z.AI raised another five billion dollars through share and convertible bond sales. Luria's skepticism is the one I share. No lab has announced a moratorium on training anything. Until one of them says a specific run is on hold, the pacing consensus is a press cycle, and Monday's selloff was the market pricing a promise nobody has kept yet.
Source: CNN Business and Axios.
The Best Coding Agent on Earth Scored 38.8 Percent
Buried under all that noise, a YC company called Specific Labs shipped a benchmark that deserved more attention than it got. Real-SWE runs frontier coding agents against private production codebases licensed from real companies. Not GitHub issues. Actual tickets actual engineers were paid to close, in code that has never been on the public internet. One of the codebases is an events app with 200,000 users, another is a fintech platform chewing through 100,000 bank statements.
Eight model and harness pairs, ten tasks, 640 scored rollouts. Anthropic's Fable 5.1 on Claude Code led at 38.8 percent. OpenAI's GPT-6 Astra on Codex CLI hit 33.8. Gemini 3.8 Flash on Gemini CLI took 31.2. Then Z.AI's GLM 5.3 at 28.8, Grok 4.6 and Meta's Muse Spark 1.3 tied at 23.8, Kimi K3 at 18.8, and GPT-5.6 Sol last at 16.2.
Six of the ten tasks came in under 15 percent. An analytics stream reducer went zero for 64, meaning not one model passed it once. A tax jurisdiction task scored 3.1 percent across the entire board. The median reference solution touched 11 files, against 6 on the benchmarks these labs quote in their launch posts.
The failure taxonomy is the part I would print and tape to a wall. Grok 4.6 just left out required behavior in 67.2 percent of its failures. GPT-5.6 Sol built on unchecked guesses about the system in 43.3 percent of its. Gemini 3.8 Flash had the right idea and wired it into the wrong place 49.1 percent of the time. Spending more did not rescue anybody either. Fable was the priciest run on the board at about seven dollars a rollout and still blew six out of ten tickets.
So the same week the labs tell us capability is outrunning our ability to control it, the best agent alive cannot close four out of ten real tickets at a real company. Both things can be true at once. A model can be genuinely dangerous at cyber work and genuinely mediocre at billing logic, because those are different skills and only one of them has clean public training data. My take is simpler than either camp's. Anybody selling you an AI engineer replacement this quarter should have to run Real-SWE against your codebase first, and show you the number.