SemiAnalysis tested Qwen3.8-Max-Preview, a 2.4T-parameter model, per a tweet from @SemiAnalysis_ on February 2026. The independent evaluation offers a rare third-party check on Alibaba's latest flagship.
Key facts
- Qwen3.8-Max-Preview has 2.4T parameters.
- SemiAnalysis tested the model, per tweet.
- No benchmark scores disclosed in tweet.
- Model likely uses MoE architecture.
- Tweet posted February 2026.
SemiAnalysis, the independent semiconductor and AI research firm, said it tested Qwen3.8-Max-Preview, a model with 2.4 trillion parameters. According to @SemiAnalysis_, the test was announced in a brief tweet on February 2026, but no benchmark scores, methodology, or compute details were included in the post.
The 2.4T parameter count places Qwen3.8-Max-Preview among the largest open-weight models publicly known, rivaling the scale of dense models like GPT-4-class systems but with a MoE (mixture-of-experts) architecture typical of Qwen's recent releases. The "Preview" designation suggests Alibaba is soliciting external feedback before a stable release, and SemiAnalysis's independent test could surface issues that internal evals miss.
Why this test matters
The tweet is thin on specifics—no SWE-Bench, MMLU, or latency numbers were shared—but the act of testing is itself notable. Alibaba has not published a technical report for Qwen3.8-Max-Preview, so independent evaluations like SemiAnalysis's are the primary source of ground truth for researchers deciding whether to adopt the model. The 2.4T parameter count, if accurate, would make it one of the largest open-weight models to date, eclipsing Qwen2.5-Max's reported 1T+ scale.
The lack of disclosed results is a limitation. SemiAnalysis has a track record of rigorous hardware and model analysis, so their test likely includes practical metrics like inference throughput and cost per token, but the tweet alone does not confirm that. Readers should wait for a fuller report or follow-up posts.
What this means for the ecosystem
For AI engineers, an independent test of Qwen3.8-Max-Preview matters because it reduces reliance on vendor benchmarks, which can be cherry-picked. If SemiAnalysis's test reveals strong performance on code or reasoning tasks, it could accelerate adoption in production environments where Qwen models are already popular due to their permissive licensing. Conversely, if the test flags weaknesses—high latency, poor long-context handling—it would temper expectations.
The 2.4T parameter figure also raises questions about training cost and inference efficiency. At that scale, even with MoE sparsity, serving the model requires substantial GPU memory, likely multiple H100 or MI300X nodes. SemiAnalysis, which tracks data-center supply chains, is well-positioned to contextualize those requirements, but the tweet does not address them.
Key Takeaways
- SemiAnalysis tested Qwen3.8-Max-Preview, a 2.4T-param model, per a tweet.
- No results disclosed, but independent eval is notable.
What to watch
![]()
Watch for SemiAnalysis to publish a full report or follow-up tweet with benchmark scores and compute details. Also track Alibaba's official release of Qwen3.8-Max, which may include a technical paper and API pricing—both would update the model's viability for production use.









