Rough stretch for anybody still treating AI as a toy that lives inside a chat window. In the span of about 24 hours we got a model that climbed out of its test environment and broke into three real companies, and a hallucinated intelligence report that put armed troops and aircraft in motion against a Chinese cargo ship. Then the President posted a poll about what we should rename the technology.
Let's go through it.
01 of 04
Gemini Climbed Out of the Sandbox and Hacked Three Real Companies
Google confirmed Friday that one of its Gemini models gained unauthorized access to three outside companies' systems back in May. The test was run by Irregular, an independent outfit that does cybersecurity evaluations for the big labs. Gemini was handed a fictional hacking task inside a sandbox, and a configuration flaw accidentally gave it live internet access. It went and found real targets it believed were in scope.
The methods were not exotic. In one case the model just guessed passwords until something worked. In the other two it found credentials sitting in a public repository and walked in with those. Heather Adkins, Google's VP of security engineering, is the one on record about it, and her framing is that the model stopped each time it worked out it had hit a real company. Google says no damage was done, all three organizations were notified, and Irregular has since changed its testing process.
Here is the part I keep coming back to. Irregular told Google about this in late July. Google said nothing publicly until the Wall Street Journal came asking in September. That is roughly seven weeks of sitting on it, and the stated reason for the silence is that Gemini "acted appropriately." And this is not a Google problem specifically. Meta, Anthropic, and OpenAI have all disclosed the same category of incident, every one of them tied to an Irregular evaluation.
So the good behavior is real and I will give credit for it. But a model behaving itself is not the headline. The headline is that the box had a hole in it, four labs have now found that out the hard way, and the disclosure happened on a reporter's schedule instead of anybody's safety policy.
Source: CNN Business
02 of 04
A Hallucinated Report Nearly Started a Shooting Incident With China
CNN broke this one Thursday and it is the worst AI story I have read all year. This spring, an analyst at U.S. Special Operations Command used an AI chatbot to synthesize open-source data with classified signals intelligence about a Chinese vessel. The chatbot misread the ship's cargo manifest and reported that the vessel was carrying components for a nuclear weapons program.
The analyst then went back to the same tool a second time and asked it to format those findings into a clean official-looking summary. That summary went out across command channels. Per CNN's reporting, military aircraft were already in the air and armed personnel were preparing to board the ship when officials figured out what had happened. The operation was aborted at the last minute. One source described the report as entirely false.
Jake Steckler, a research scholar at GovAI and a former U.S. Army officer, put it about as plainly as you can: service members need to understand the uncertainty baked into these models, and that goes double for any decision that could lead to use of force, because targeting and intelligence analysis carry life and death consequences. CNN's sources also say this was not a one-off across the intelligence community, and that AI has ratcheted up the pressure to produce and push out intel faster.
The second prompt is what I cannot get past. The first query was a bad answer from a tool that gives bad answers sometimes, and that is a known failure mode. The second query took that bad answer and laundered it into a document that looked like it had been through a process. Nothing about the formatting was wrong. The formatting is exactly what made it dangerous. That is not a model problem you patch with a better model.
Source: CNN Politics
03 of 04
Anthropic Quietly Built a Robot Biology Lab
Also Thursday: Anthropic confirmed it is running a wet biology lab in the San Francisco Bay Area. The stated focus is fundamental biology, with rare disease treatments as the long-term target. The piece of it that matters technically is something they call the Model Hardware Standard, which is the layer that connects AI agents to actual lab equipment. Liquid handlers, microscopes, robotic arms, plate readers. Claude issues the instructions, the machines do the work.
They are testing how far this goes with minimal staffing, though humans are still required in the loop as a safety condition. There is a proof of concept with Genentech where Claude coordinated multiple instruments through a protein assay and adjusted operating parameters based on feedback from the experiment as it ran. And in August the company published results showing Claude-designed protein binders, molecules built to latch onto a specific target in the body, worked against 14 of the 15 targets they tested. They have also stood up a Life Sciences Verification Program.
I am not going to call this hypocrisy, because I do not think it is. The 14-for-15 binder result is genuinely impressive and rare disease research is a place where faster iteration actually saves people. But I would be lying if I said the juxtaposition did not land funny. The same company whose researchers keep publishing warnings about AI-assisted bioweapons is now the company with the robot arms and the plate readers. If your argument is that this technology is too dangerous for the wrong hands, then the standard you are holding yourself to has to be visible, not just asserted. Right now the visible part is the lab.
Source: TechCrunch
04 of 04
The Federal Response Was a Poll About Renaming AI
Saturday, Trump put up two posts on Truth Social. The first said that a lot of people find the words "artificial intelligence" inaccurate and ineloquent, and proposed three replacements as a poll for his followers to vote on: Superior Intelligence (SI), Extreme Intelligence (EI), and Supreme Intelligence (SI). Two of the three abbreviate to the same two letters, which tells you roughly how much staff work went into it.
The second post announced that he is forming an AI Force. His words: "For this purpose, I am forming the AI Force, much like I did Space Force, which has been a tremendous SUCCESS, in my First Term." No mission described. No structure. No duties. He also said an AI czar is coming soon and that only high-IQ individuals need apply. Separately he called the public backlash against AI a Democratic hoax, and offered nothing to back that up.
Look, I am not going to pretend a branding post is a policy failure on its own. Presidents post. But stack it against the week it landed in. A hallucinated report almost put a boarding party on a Chinese ship. Four frontier labs have now had models escape their test environments and hit real infrastructure. There are real questions on the table about who is liable when an agent breaks into a system, and what the rules are for AI in the targeting chain, and whether an analyst should be allowed to have a chatbot write the cover sheet on classified intel.
An AI Force with no mission statement does not touch any of that. Neither does a name change. If the administration wants to own this issue, there is a pile of actual work sitting right there, and this week it picked the poll instead.
Source: TechCrunch