Ant Group released Ling-3.0-flash on July 22, a 124B-parameter MoE with 5.1B activated that beats 1T-Ring-2.6 in 11 of 12 benchmarks. The model ties DeepSeek V4 Flash for the top average score across 34 evaluation dimensions.
Key facts
- 124B total parameters, 5.1B activated per token
- 15 first-place and 19 second-place results across 34 dimensions
- Beat 1T-Ring-2.6 on 11 of 12 shared benchmarks
- Tied DeepSeek V4 Flash for top average score
- Released July 22, 2026, same day as Huang's open-source comments
Ant Group released Ling-3.0-flash, a mixture-of-experts execution model with 124B total parameters and only 5.1B activated per token. According to Pandaily, the model achieved 15 first-place and 19 second-place results across 34 evaluation dimensions, beating 1T-Ring-2.6 — a model with roughly 8× the parameters — in 11 of 12 shared benchmarks. It tied DeepSeek V4 Flash for the highest average score.
Key Takeaways
- Ant's 124B-param Ling-3.0-flash with 5.1B activated beats 1T-Ring-2.6 in 11 of 12 benchmarks, tying DeepSeek V4 Flash.
- Sparse activation economics are the story.
Sparse Activation Economics

The 5.1B activated parameter count is the critical number here. At inference, Ling-3.0-flash requires only about 4% of its total parameter count to be loaded into memory for each token, a design choice that directly attacks the memory-bandwidth bottleneck that dominates serving costs. The 124B total parameter count suggests a dense training run with a wide expert fan-out, consistent with the MoE architecture Ant has been developing internally.
Open-Source Context
The release lands the same day Jensen Huang publicly championed open-source models, arguing they expand the addressable market for Nvidia's hardware. Huang's comments on July 22 noted that DeepSeek and Kimi open models boost Nvidia sales — a dynamic that directly applies to Ling-3.0-flash, which Ant has positioned as an execution-focused model for production workloads rather than a research showcase.
The benchmark results are notable for what they don't show: the source does not disclose inference latency, throughput, or cost-per-token figures, which are the metrics that matter for an "execution" model. The company also did not specify the training compute budget or the exact evaluation methodology for the 34 dimensions, leaving room for benchmark-gaming skepticism that has followed Chinese model releases this year.
Competitive Positioning
Ling-3.0-flash's tie with DeepSeek V4 Flash is significant given DeepSeek's recent momentum — the company is in early talks for a funding round at a $71 billion pre-money valuation, and its V4 line has become the reference point for Chinese open-weight models. Ant matching that average score with a smaller activated-parameter count suggests the efficiency race is tightening, not just the raw capability race. The 11-of-12 win against 1T-Ring-2.6, a model with roughly 8× the parameters, reinforces that sparse activation is becoming the dominant serving strategy for production deployments.
What to watch
Watch for Ant's disclosure of inference cost-per-token and latency benchmarks, which the release omitted. If Ling-3.0-flash delivers sub-10ms time-to-first-token at competitive pricing, it will pressure DeepSeek and Qwen on serving economics. Also monitor whether Ant open-weights the model, which would signal a direct challenge to DeepSeek's distribution strategy.
Source: pandaily.com








