Qwen3.8-Max-Preview Arrives With a Bold Claim and Zero Proof
Alibaba launched Qwen3.8-Max-Preview at WAIC Shanghai, claiming it ranks second only to Claude Fable 5—but without benchmarks, a model card, or independent scores to back it up.
Alibaba previewed its new flagship AI model, Qwen3.8, at the World Artificial Intelligence Conference in Shanghai on Sunday, July 19, claiming the system ranks second only to Anthropic’s Claude Fable 5 among the world’s frontier models. That is an extraordinary claim. What is equally extraordinary is that Alibaba made it without publishing a benchmark table, a model card, an activated-parameter count, or a shred of independent verification. The model is live. The evidence is not.
Alibaba stock rose as much as 5.4% on Monday following the announcement. Investors apparently decided that the company’s word about its own product was good enough. Whether that confidence survives contact with actual third-party testing is a different question.
What Alibaba Actually Confirmed
Alibaba said the model contains 2.4 trillion parameters and is the first Qwen model with more than one trillion parameters to support multimodal tasks, including text, images, video, and document understanding. The model uses a sparse Mixture-of-Experts (MoE) architecture, meaning only a fraction of those 2.4 trillion parameters activate for any given task, which keeps inference practical at massive scale. Crucially, Alibaba has not disclosed how many parameters activate per token — the number that actually determines speed and cost. Without that figure, the 2.4 trillion headline is essentially marketing shorthand.
The preview is available through Alibaba’s Token Plan, Qoder, and QoderWork platforms at 10% of the standard price during the trial period. Alibaba also said it plans to release the model’s open weights “soon,” though it has not provided a timeline. As of July 20, there is no model card on Hugging Face, so the weights cannot be downloaded yet. “Soon” is doing a lot of heavy lifting here.
The Gap Between the Claim and the Reality
Unlike the launch of Qwen3.7-Max in May, the company has not published a model card, activated parameter count, benchmark scores, or detailed technical specifications beyond the total parameter count. It also did not provide task-level comparisons with Qwen3.7-Max, describing the new model only as “continuously evolving.” As a result, there is currently no public benchmark data available to compare its performance with other AI models.
That omission matters more given what came before. Alibaba’s previous flagship, Qwen3.7-Max, shipped in May with a full set of published results, including a score of 56.6 on the Artificial Analysis Intelligence Index. That score placed it fifth overall, making it the highest-ranked Chinese model on the leaderboard — respectable, but nowhere near Fable 5. No outside leaderboard has ranked Qwen3.8 yet, and the most recent Qwen model that was scored, Qwen3.7-Max, sits far from the top of LMArena’s list, where Fable 5 currently holds first place. Alibaba is asking buyers to accept a jump from the middle of the pack to second in the world on its word alone.
The ranking is the company’s own characterization of internal testing. Independent leaderboards including Artificial Analysis, Arena.AI, and Hugging Face’s Open LLM Leaderboard show no Qwen3.8-Max entry. Meanwhile, the public Artificial Analysis Intelligence Index snapshot ranks Claude Fable 5 first at 59.9%, ahead of GPT-5.6 Sol (58.9%) and Kimi K3 (57.1%) among 166 tested models. Qwen3.8 does not appear on that list at all.
The Context: Racing Against Kimi K3
The timing of this announcement is not subtle. Kimi K3, launched on July 16 by Beijing Moonshot AI, features a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window and native vision capabilities. Unlike Qwen3.8, Kimi K3 shipped with benchmark results. Moonshot AI released benchmark results with Kimi K3, and those results put it ahead of Fable 5 and OpenAI’s GPT-5.6 Sol on some tests, including Program Bench and SWE Marathon. The demand was so strong that Kimi.ai announced a pause on new subscriptions, saying that two days of surging usage had strained its GPU resources to near capacity. To protect existing subscribers, it prioritized available compute for current members while expanding its infrastructure and gradually reopening new spots in batches.
Alibaba walked onto the same stage three days later with a bigger story and less documentation. The timing is the story as much as the model.
The Bigger Picture: Alibaba’s $53 Billion Bet
Strip away the benchmark controversy and there is a serious infrastructure play underneath. Alibaba Group announced plans to invest at least RMB 380 billion ($53 billion) over the next three years to advance its cloud computing and AI infrastructure, an investment that exceeds Alibaba’s total AI and cloud spending over the past decade. That spend now has a flagship model attached to it, whether or not that model’s ranking holds up under scrutiny.
Alibaba also picked up a significant distribution win just days before WAIC. The Cyberspace Administration of China approved Apple’s AI services on July 15, with Alibaba’s Qwen handling the underlying model across iOS, iPadOS, macOS, and visionOS for users in mainland China. Potentially hundreds of millions of Apple devices in China running on Qwen gives Alibaba an inference volume opportunity that sits entirely outside any leaderboard debate.
What’s Actually Worth Watching
The Qwen3.8 story has two plausible endings. Either Alibaba publishes benchmark data soon and the claim turns out to be roughly accurate — in which case this is a solid model announced with unusual secrecy. Or the data arrives and the gap between “second only to Fable 5” and reality turns out to be significant, in which case this launch will be remembered as an example of marketing outrunning engineering. Every performance line so far comes from Alibaba’s own internal evals, not Artificial Analysis or LMArena. Until that changes, the rank is an assertion, not a result. Five things worth watching: an official benchmark table, the activated-parameter count, a Hugging Face model card with a license, standard API pricing, and an independent evaluation from a platform with no stake in the outcome. What is missing is everything a buyer would actually need — Alibaba’s Qwen team publishes a detailed launch post with full benchmark tables for each flagship, as it did for Qwen3.7 in May and Qwen3.6-Max-Preview in April, and no equivalent post exists for 3.8. That absence, for a company with Alibaba’s track record of thorough releases, is the most telling data point available right now.





