Three things landed this week that matter if you're using AI for anything besides asking it to write your emails. Here's the rundown.
A Free Chinese Model Just Wiped Out $314 Billion in AI Value
Moonshot AI dropped the weights for Kimi K3 this week. Free download, 594 gigabytes, modified MIT license, do whatever you want with it. No API key, no waiting list, no subscription.
The market did not shrug this one off. Reports put the hit to OpenAI and Anthropic's pre-IPO paper valuations at $314 billion combined, with Anthropic taking the bigger chunk at $232 billion and OpenAI down $82 billion.
Why it matters: for years the story was that frontier labs had a moat because training a top tier model costs a fortune and only a few outfits could afford it. Kimi K3 is more proof that moat keeps getting shallower every few months. If a model good enough to spook investors can just show up on a Tuesday with no strings attached, the pricing power everybody assumed these labs had starts looking shaky.
Robert's take: I don't think this kills OpenAI or Anthropic, they've got distribution, enterprise contracts, and cash nobody else has. But paper valuations built on "nobody else can do this" were always fragile, and this is the market finally pricing that in. If you're building on top of any single model provider right now, this is your reminder to keep your architecture swappable. Don't marry a model.
AMD and Cerebras Team Up to Come After Nvidia's Inference Crown
AMD and Cerebras announced they're pairing AMD's Helios rack scale systems with Cerebras' Wafer Scale Engine chips to split the work of AI inference. Helios handles the heavy lifting of processing prompts and big context windows, and the Cerebras chip takes over the memory hungry job of generating tokens fast.
The companies are claiming up to 5x better tokens per second per watt than doing it the normal way, and AMD's Lisa Su says Helios can push 30 percent more inference tokens per dollar than Nvidia's upcoming Vera Rubin rack. The joint setup shows up on Cerebras Cloud later this year.
Why it matters: inference, not training, is where the real money gets spent now that everybody and their dog is running agents around the clock. Every dollar and watt saved per token is a dollar somebody doesn't have to charge you. Nvidia has owned this space for years basically unchallenged.
Robert's take: I'll believe the 5x number when independent folks run it, vendor benchmarks always look great on the vendor's own slide deck. But the fact that AMD and Cerebras are teaming up at all tells you something, nobody wants to fight Nvidia alone anymore. Good for us either way, more competition on inference cost means cheaper agents for the rest of us building on top.
Washington Wants 30 Days to Look at Your New Model Before You Ship It
OpenAI, Anthropic, and Google are reportedly close to a voluntary deal with the federal government that would give agencies up to 30 days to review a new frontier model for national security risk before it goes public. An announcement is expected before August 1.
Why it matters: this is the first real sign of a formal pre-release review process for the biggest labs, even if it's voluntary for now. Voluntary today has a habit of turning into mandatory tomorrow once there's a process in place to point to.
Robert's take: I get the national security angle, I really do. But a 30 day government review window is going to move slower than the pace this industry ships at, and the labs that agree to it are betting that looking responsible is worth the friction. Watch what happens to release cadence over the next few quarters. If the big three start shipping less often right around the same time, you'll know why.