Remember Fable? Mythos?
Those were the models Anthropic rolled out earlier this spring, then rolled back, then finally released recently. OpenAI is similarly rolling out GPT 5.6, the latest attempt to monetize agentic models while they still can. Both companies are in trouble, and they are acting like it. Why? Because a lot of their edge relied on the fact that locally-hosted models could not compete with their cloud-only ones. But they’ve unknowingly priced themselves into a corner.
Market Mistiming
AI providers were hoping to cash in on the usual stickiness cycle – provide something at an artificially low cost early on (tokens) so people will adopt, then crank up the price when they can no longer live without it. It worked for Uber, Airbnb, AWS… and to an extent, it worked for Anthropic and OpenAI! While tokens were cheap, customers were consuming enormous amounts of tokens. Unfortunately for them, what they were building was not sticky, they were just burning through tokens without making anything substantial. And now it’s time to pay up – token prices MUST be raised for OpenAI and Anthropic to get the valuation they want for their IPO.
Enter GLM 5.2

Unfortunately for them, as they crank up the price, the gap between open-source LLMs and proprietary commercial models has shrunk dramatically. Open source model options have democratized the technology in a hurry, and become very attractive to those users who were shocked by the sudden price increases. Some companies were already exploring self-hosted alternatives, which is also bad news for vendors who rely on API usage to fuel their profitability.
Z.ai picked the worst possible moment for Anthropic to prove the point. Barely a day after Fable and Mythos got yanked over export controls, Z.ai quietly rolled out GLM 5.2 to its Coding Plan members on a Saturday — an unusual release schedule that reads less like an accident and more like a company that knows exactly how to time a marketing win. The open weights and official benchmarks followed three days later, and what followed that was the kind of grassroots hype cycle you don’t get from a press release: developers running it themselves, comparing it to DeepSeek R1’s moment, and finding it holds up.
The numbers back it up. On Terminal-Bench 2.1, GLM 5.2 jumped from 63.5 to 81.0 — within a few points of Opus 4.8’s 85.0. It’s an MoE model, roughly 750B total parameters but only ~40B active per token, which keeps self-hosting realistic instead of theoretical. A 1M-token context window. MIT licensed, no regional restrictions, no usage caps. And reportedly around a sixth of the price of comparable frontier models.
None of this means GLM 5.2 is better than what Anthropic or OpenAI have shipped — on the hardest long-horizon tasks it still trails Opus 4.8 by double digits. But “still trails the best closed model” isn’t the bar anymore. The bar is “good enough, self-hostable, and not subject to export controls that can vanish your access overnight.” That’s the trap. When your pricing power depended on being irreplaceable, and someone releases a model that’s 90% as good for a sixth of the price with no strings attached, you don’t get to just raise prices and wait for the IPO. You have to explain why anyone should still pay you.
What That Means for You
If you don’t know where to start, you’re definitely not alone! The self-hosted landscape has moved fast and most of the guidance out there is either vendor marketing or deep in the weeds technically. If you’ve been putting off exploring self-hosted or open-weight models because the pricing math didn’t make sense before, that math just changed. This is a good window to figure out what running your own models actually looks like for your organization, before everyone else does the same math and the specialized help gets harder to book.
I can help! luisaherrmann.com