📊 Full opportunity report: Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba’s Qwen3.8-Max, previously only teased, has now published detailed benchmark results, confirming it as the second most capable model after Fable 5. Open weights are set to ship soon, marking a significant milestone in open AI models.
Alibaba has officially published detailed benchmark results for its Qwen3.8-Max model, confirming it as the second most capable AI model after Fable 5. The company disclosed the model’s parameters, performance scores, and plans to release open weights next week, marking a significant development in the open AI landscape.
On August 3, Alibaba released the full benchmark table for Qwen3.8-Max, revealing it has 2.4 trillion total parameters and approximately 95 billion active parameters per query. The model is built on the Qwen3.5 architecture, utilizing sparse mixture-of-experts technology, and supports multimodal input — including text, images, and video — with text output.
The benchmark results place Qwen3.8-Max as second only to GPT-5.6 Sol on the Terminal-Bench 2.1, with a score of 86.6, surpassing Claude Opus 4.8 and Claude Fable 5.6, which scored 84.6. Only GPT-5.6 Sol, with a score of 88.8, outperforms it at maximum effort. Other benchmarks, like PaperBench, rank Qwen3.8-Max at the top with a score of 93.0, highlighting its strong performance in various tasks.
Alibaba also demonstrated the model’s capabilities in long-horizon reasoning and agentic execution, outperforming previous versions significantly, especially in software engineering benchmarks. The open weights are scheduled for release next week, though the model remains a multi-node artifact due to its size, limiting direct self-hosting options.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Impact of Benchmark Results and Open Model Release
This development confirms Alibaba’s position as a major player in AI model capabilities, with Qwen3.8-Max ranking just below GPT-5.6 Sol in performance. The full benchmark disclosure and upcoming open weights signal a shift toward more accessible, high-capacity models, potentially influencing AI deployment and competition. The release of the 27B checkpoint also caters to local deployment needs, enabling broader use in practical applications, especially in environments with limited hardware resources.
These advancements could accelerate AI research, foster more open competition, and challenge existing market leaders, especially if the open weights prove effective for real-world tasks.

From Weights to Wisdom: The Complete Guide to Running and Adapting Opensource AI Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba’s AI Model Development
Alibaba has been gradually building its AI capabilities, culminating in the stealth preview of Qwen3.8-Max in July, which was initially revealed through community detection and a brief teaser during the World AI Conference in Shanghai. Prior to this, the company had released smaller models like Kimi K3 and maintained a strategy of staged disclosures, emphasizing both performance and openness.
The model’s announcement follows a competitive landscape where models like Meta’s Llama, OpenAI’s GPT series, and other large-scale models have set benchmarks. Alibaba's approach of combining high performance with open weights aims to carve out a distinct position in this crowded field.
Previous Alibaba models have focused on multimodal capabilities and agentic reasoning, with incremental improvements culminating in this latest benchmark reveal, which underscores the company's focus on scaling and practical deployment.
"We are committed to advancing AI capabilities and providing open access to our models, starting with next week's release of open weights."
— Alibaba spokesperson
Unconfirmed Aspects of Open Weights and Deployment
It remains unclear what the licensing terms will be for the open weights, and whether they will be fully open-source under permissive licenses like Apache 2.0. Additionally, the practical performance of the 27B checkpoint in real-world applications has yet to be demonstrated, as it is designed for local deployment on high-memory hardware and may not match the flagship’s performance in all tasks.
Further, the long-term stability of agentic improvements and the impact of potential licensing restrictions are still uncertain.
Upcoming Release and Practical Deployment of Open Weights
The open weights for Qwen3.8-Max are scheduled to be released next week, enabling developers and researchers to test and deploy the model locally. The 27B checkpoint will be particularly relevant for those seeking high-performance, single-machine solutions.
Alibaba will likely publish detailed licensing terms and documentation alongside the open weights, clarifying usage rights and restrictions. Monitoring how the community adopts these models will be critical in assessing their impact on the AI ecosystem.
Key Questions
What are the main performance advantages of Qwen3.8-Max?
Qwen3.8-Max demonstrates strong benchmark scores across multiple tasks, especially in multimodal and agentic reasoning benchmarks, placing it just below GPT-5.6 Sol and outperforming many competitors in specific areas like long-horizon reasoning.
When will the open weights for Qwen3.8-Max be available?
The open weights are scheduled to be released next week. Exact date and licensing details are yet to be announced by Alibaba.
Can I run Qwen3.8-Max on my own hardware?
While the full 2.4 trillion parameter model is a multi-node artifact, a 27B checkpoint will be released for local deployment on high-memory machines, making it accessible for many users.
How does Qwen3.8-Max compare to other models like GPT-4 or Claude?
In benchmark tests, Qwen3.8-Max scores highly, second only to GPT-5.6 Sol, and surpasses Claude Fable 5 in several tasks. However, its real-world performance and deployment flexibility remain to be fully evaluated.
Source: ThorstenMeyerAI.com