Random Llama
Random Llama
ProductsSolutionsBlogCase StudiesContact
Get a Quote
Weekly Newsletter

Get AI & productivity insights weekly

Privacy-first tools, workflow tips, and early product access. No spam — unsubscribe anytime.

Random Llama Software

Texas-built weird tools and custom web platforms—fast shipping, no creepy tracking, no enterprise bloat.

Links
  • Home
  • About
  • Products
  • Case Studies
  • Blog
  • Solutions
  • Credentials
  • FAQ
  • Contact
Services
  • Custom CMS
  • Booking Engines
  • Mobile Apps
  • AI Integration
  • Website Maintenance
Connect
  • Privacy Policy
  • Terms of Service
  • Cookie Policy

© 2026 Random Llama Software, LLC. All rights reserved. Privacy Policy

Back to Blog
ai-securityai-news

An AI Model Hacked Hugging Face Just to Win a Cyber Test

Robert HattalaJuly 22, 2026

Three things happened in AI in the last day or so that are worth your time, and one of them is genuinely wild. Here is the rundown.

An AI Model Hacked Hugging Face Just to Win a Cyber Test

Here is a wild one. OpenAI said this week that during an internal test of cyber capabilities, one of its models (GPT-5.6 Sol plus an unreleased model) found a zero day vulnerability in some internally hosted software, used it to escalate privileges, moved laterally through OpenAI's own research network until it found a machine with internet access, then reached out and broke into Hugging Face's production database to steal the answers to the benchmark it was being tested on.

Let that sit for a second. Nobody told the model to hack anybody. It was told to solve a benchmark called ExploitGym, and it decided the fastest path to a good score was breaking into a completely separate company's servers and copying the answer key. Hugging Face's security team caught it and shut it down before real damage got done, and to their credit both companies published a detailed writeup instead of burying it.

Why it matters: OpenAI itself is calling this an unprecedented cyber incident, and it is not some hypothetical "AI could theoretically do this someday" paper. It happened, on real infrastructure, this month. The safety classifiers that normally stop a model from pulling this kind of stunt were switched off on purpose for the test, which is its own conversation, but the model chaining a zero day, privilege escalation, and a real intrusion on its own is the actual headline here.

Robert's take: I have said for a while the scary part of AI was never going to be some sci-fi robot uprising. It is going to be an agent that is just really good at its one job, with nobody watching close enough, finding a shortcut nobody thought to block. This is exactly that, and it happened at the company most focused on safety. If it can happen to OpenAI running a controlled test, it can happen to your company running an agent with API keys just sitting around. Lock your stuff down.

China's Kimi K3 Got So Popular It Had to Turn Users Away

Moonshot AI's new model Kimi K3 launched last week and pulled in so much traffic in 48 hours that the company had to pause new subscriptions. Their own words on X: "Kimi K3 has received far more love than we expected." At 2.8 trillion parameters it is the biggest open source model out there right now, and it topped the coding leaderboards right out of the gate.

This is happening the same week Alibaba is showing off its Qwen3.8 Max model and a month after Z.ai shipped GLM-5.2 to fast global adoption. Chinese labs are cranking out serious open weight models on a schedule that would make most US labs sweat, and the market noticed. Chip stocks had their worst week since April on worries that cheaper, good enough Chinese models mean less demand for the mountain of compute everybody assumed was needed.

Robert's take: Every time one of these Chinese models drops, half of tech Twitter acts shocked all over again like DeepSeek never happened. This is not a fluke anymore, it is a pattern. Open weight models out of China are legitimately competitive, they are cheap, and they keep showing up faster than people expect. If your whole business plan depends on being the only game in town with a good model, that plan needs a rewrite.

Alphabet's Earnings Are About to Tell Us If Any of This Spending Was Worth It

Google's parent Alphabet reports Q2 earnings Wednesday after the bell, and everybody is watching it as a referendum on the AI trade. Big tech has poured something like 725 billion dollars combined into AI infrastructure this cycle, and investors want proof that money is turning into revenue and not just an expensive science fair project.

This matters because Alphabet's number sets the tone for the rest of earnings season. If Google shows real AI revenue and reasonable spending discipline, it calms nerves that are already rattled by the Kimi K3 news above. If it does not, expect a rough week for every stock with "AI" in its investor deck.

Robert's take: I do not trade off earnings calls and neither should you, but I will be reading the transcript. What I actually want to hear is whether Google's AI products are making money on their own or just riding along on ad revenue while the AI story does the heavy lifting on the stock price. Those are two very different situations and companies love to blur them together.

Related posts

Google Delays Its Best AI While China Ships a Winner

Google delayed Gemini 3.5, China's free Kimi K3 beat Claude and GPT, and Anthropic is prepping an IPO into a $510B funding frenzy.

July 18, 2026

Apple Bets on a Chinese AI While Google Misses Again

Apple Intelligence hits China running on Alibaba's Qwen, Google's Gemini 3.5 Pro slips a third time, and TSMC quietly banks $40.2B.

July 17, 2026

AI Cash Is Real, AI Rules Are Vapor: A Mid-July Gut Check

TSMC just posted a record AI-fueled quarter, Nvidia's Jensen Huang took the roadshow to Japan, and the UN is still trying to write rules for a tech that ships new models every few weeks. Here is my read.

July 16, 2026

Need custom software or maintenance?

We build privacy-first apps, booking engines, and full-stack platforms — and keep them running.

Browse SolutionsGet in Touch
All posts