After Kimi K3, Z.ai and MiniMax Have About Ten Weeks to Respond
Moonshot’s Kimi K3 wiped billions off rivals’ valuations on July 17. Now MiniMax’s 2.7T M3 Pro and Z.ai’s next move must land before Q3 ends.
On July 16–17, 2026, Moonshot AI released Kimi K3 — quietly, with no keynote, no press conference, just a model page going live overnight — and by the next morning, the Chinese AI leaderboard had been reshuffled. K3 is the first open-source model to reach 2.8 trillion parameters and, as of launch, Moonshot’s most capable flagship to date. Moonshot says K3 still sits behind Anthropic’s Claude Fable 5 and OpenAI’s GPT 5.6 Sol on overall performance, but it outperformed every other model in the company’s evaluation suite — including Claude Opus 4.8 and GPT 5.5 — across coding and agentic benchmarks. For rivals Z.ai and MiniMax, that’s a problem that needs solving before summer ends.
Z.ai saw its shares drop approximately 27% following the announcement. MiniMax fell roughly 16%. The market is not subtle: when a competitor ships the world’s largest open-weight model at competitive pricing, investors do not wait around for a rebuttal press release. Both companies now face a compressed timeline to answer K3 — either with bigger models, smarter architecture, or a multimodal play that changes the scoring criteria entirely.
What K3 Actually Is — and Why It Rattled the Room
K3 is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals, with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning.
K3 uses Stable LatentMoE, activating just 16 of its 896 experts per token — meaning only a fraction of those 2.8 trillion parameters fire on any given request. That keeps inference costs manageable and sets a template for how you scale to near-3T without requiring a small power grid to run it. Together, the architectural changes yield roughly 2.5x better overall scaling efficiency than Kimi K2.
On the pricing side, API pricing is $0.30 per million cache-hit input tokens, $3 per million on cache misses, and $15 per million output tokens. That positions K3 well below what US labs charge for comparable-tier models — and it comes with open weights dropping on July 27, which means developers can self-host it entirely. As VentureBeat noted, the release was timed to land just ahead of the 2026 World Artificial Intelligence Conference in Shanghai — a deliberate statement in front of the largest annual audience in Chinese tech.
MiniMax’s Clock Is Running
MiniMax’s situation is the more precarious of the two. Its market cap had already dropped from HK$410 billion in March 2026 to HK$109 billion by early July 2026 — a decline of over 72% — before K3 even launched. The company has been losing ground in China’s model rankings and now faces the additional pressure of being lapped on parameter scale by a direct rival.
The answer it’s preparing is M3 Pro. MiniMax is building a 2.7-trillion-parameter model and plans to open-source it as early as Q3, according to The Information. MiniMax will also launch H3, its frontier-level multimodal video generation model, later this month — a separate bet that the competition isn’t just about parameter counts and coding benchmarks. If M3 Pro lands in Q3 as planned, it would clock in just 100 billion parameters shy of K3. Not a win on raw scale, but close enough to make the contest about architecture quality, real-world benchmarks, and price.
The problem for MiniMax is that its current flagship, M3, did not exactly cover itself in glory. While official claims put M3’s programming ability at 59% on SWE-Bench Pro — exceeding GPT-5.5 — the intelligence index of independent evaluator Artificial Analysis ranked it ninth among mainstream models, and on Chatbot Arena, which focuses on real user preferences, it ranked outside the top forty or fifty. M3 Pro needs to clear that credibility gap, not just the parameter gap.
Z.ai: Architecture Over Brute Force
Z.ai — formerly Zhipu AI — took a different approach with GLM-5.2, released on June 13, 2026. On industry-standard third-party benchmarks, GLM-5.2 performs above most open-source flagship models, even DeepSeek V4, and scores near or above closed-weights rivals GPT-5.5 and Claude Opus 4.8. The model is a 753-billion-parameter open-weights LLM engineered specifically to dominate long-horizon autonomous coding and engineering tasks.
On SWE-bench Pro, GLM-5.2 scored 62.1, beating GPT-5.5’s 58.6 — and VentureBeat reports it does this for roughly one-sixth the cost. The architectural key is IndexShare: an optimization to sparse attention that reduces per-token compute by 2.9x at the full 1-million-token context length. This is Z.ai’s answer to scale — not 2.8 trillion parameters, but 753 billion parameters that cost far less to run and still beat models that cost six times more per query.
Now K3 has arrived and rendered GLM-5.2’s “top open-weight coding model” title provisional at best. Arena ranked K3 first in its Frontend Code evaluation at 1,679 points, ahead of Claude Fable 5, in blind developer testing — the exact category where GLM-5.2 had made its name. According to Latent Space’s AI News, Z.ai had previously forecast an “Open Fable by EOY” — suggesting the company’s roadmap targets matching closed-source frontier quality by end of year. K3 just moved that goalpost further down the field.
The Open-Weight Bet and Its Uncomfortable Economics
All three companies — Moonshot, Z.ai, MiniMax — are racing on the same track: open-weight models that developers worldwide can download, fine-tune, and run without paying a US-based API for the privilege. US labs spend billions to train frontier models and then charge a premium. Cheap, capable open weights undercut that logic entirely.
That’s a winning strategy for developer adoption. Whether it’s a winning strategy for the balance sheet is a different question. Bank of America analyst Alex Liu said in a note that “K3 raises the capability ceiling for China AI models, shifting the burden of proof to other independent AI labs.” The burden of proof now sits squarely with Z.ai and MiniMax — and they have roughly ten weeks of Q3 left to produce it. Meanwhile, Moonshot is seeking as much as $2 billion in a new funding round that would value the startup at $30 billion, giving it the runway to keep shipping at this pace regardless of who responds.
What’s Next
The immediate calendar is clear: K3’s full model weights will be released by July 27, 2026. Once that happens, independent benchmarkers get to stress-test every claim Moonshot has made — and the community will decide whether K3 is as good as advertised or another self-reported number waiting to be deflated. For MiniMax, the new model could be released as early as Q3, which means the window is August or September. Miss that window and M3 Pro becomes a 2027 story nobody will care about. For Z.ai, the play is more nuanced: compete on efficiency and price per useful token rather than raw parameter count, and accelerate toward whatever comes after GLM-5.2 before K3’s open weights give every developer a free alternative to the entire GLM family. China’s AI parameter race has been running at sprint pace all year. K3 just turned it into a dead heat with a countdown clock.





