Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Data chart comparing AI model benchmark scores with Ant Group logo and model names Ling-3.0-flash and 1T-Ring-2.6
Open SourceBreakthroughScore: 90

Ant Ling-3.0-flash Beats 1T-Ring-2.6 on 11 of 12 Benchmarks

Ant's 124B-param Ling-3.0-flash with 5.1B activated beats 1T-Ring-2.6 in 11 of 12 benchmarks, tying DeepSeek V4 Flash. Sparse activation economics are the story.

·1d ago·3 min read··5 views·AI-Generated·Report error
Share:
Source: pandaily.comvia pandailySingle Source
How does Ant Group's Ling-3.0-flash execution model perform against larger models like 1T-Ring-2.6?

Ant Group's Ling-3.0-flash, a 124B-parameter MoE with 5.1B activated, scored 15 first-place and 19 second-place results across 34 evaluation dimensions, beating 1T-Ring-2.6 in 11 of 12 benchmarks while using 12% of the parameters. It ties DeepSeek V4 Flash for the top average score.

TL;DR

124B total params, only 5.1B activated · 15 firsts, 19 seconds across 34 dimensions · Ties DeepSeek V4 Flash on average score

Ant Group released Ling-3.0-flash on July 22, a 124B-parameter MoE with 5.1B activated that beats 1T-Ring-2.6 in 11 of 12 benchmarks. The model ties DeepSeek V4 Flash for the top average score across 34 evaluation dimensions.

Key facts

  • 124B total parameters, 5.1B activated per token
  • 15 first-place and 19 second-place results across 34 dimensions
  • Beat 1T-Ring-2.6 on 11 of 12 shared benchmarks
  • Tied DeepSeek V4 Flash for top average score
  • Released July 22, 2026, same day as Huang's open-source comments

Ant Group released Ling-3.0-flash, a mixture-of-experts execution model with 124B total parameters and only 5.1B activated per token. According to Pandaily, the model achieved 15 first-place and 19 second-place results across 34 evaluation dimensions, beating 1T-Ring-2.6 — a model with roughly 8× the parameters — in 11 of 12 shared benchmarks. It tied DeepSeek V4 Flash for the highest average score.

Key Takeaways

  • Ant's 124B-param Ling-3.0-flash with 5.1B activated beats 1T-Ring-2.6 in 11 of 12 benchmarks, tying DeepSeek V4 Flash.
  • Sparse activation economics are the story.

Sparse Activation Economics

Ant Ling-3.0-Flash 124B-A5B... new fast model for one Spark ...

The 5.1B activated parameter count is the critical number here. At inference, Ling-3.0-flash requires only about 4% of its total parameter count to be loaded into memory for each token, a design choice that directly attacks the memory-bandwidth bottleneck that dominates serving costs. The 124B total parameter count suggests a dense training run with a wide expert fan-out, consistent with the MoE architecture Ant has been developing internally.

Open-Source Context

The release lands the same day Jensen Huang publicly championed open-source models, arguing they expand the addressable market for Nvidia's hardware. Huang's comments on July 22 noted that DeepSeek and Kimi open models boost Nvidia sales — a dynamic that directly applies to Ling-3.0-flash, which Ant has positioned as an execution-focused model for production workloads rather than a research showcase.

The benchmark results are notable for what they don't show: the source does not disclose inference latency, throughput, or cost-per-token figures, which are the metrics that matter for an "execution" model. The company also did not specify the training compute budget or the exact evaluation methodology for the 34 dimensions, leaving room for benchmark-gaming skepticism that has followed Chinese model releases this year.

Competitive Positioning

Ling-3.0-flash's tie with DeepSeek V4 Flash is significant given DeepSeek's recent momentum — the company is in early talks for a funding round at a $71 billion pre-money valuation, and its V4 line has become the reference point for Chinese open-weight models. Ant matching that average score with a smaller activated-parameter count suggests the efficiency race is tightening, not just the raw capability race. The 11-of-12 win against 1T-Ring-2.6, a model with roughly 8× the parameters, reinforces that sparse activation is becoming the dominant serving strategy for production deployments.

What to watch

Watch for Ant's disclosure of inference cost-per-token and latency benchmarks, which the release omitted. If Ling-3.0-flash delivers sub-10ms time-to-first-token at competitive pricing, it will pressure DeepSeek and Qwen on serving economics. Also monitor whether Ant open-weights the model, which would signal a direct challenge to DeepSeek's distribution strategy.


Source: pandaily.com


Sources cited in this article

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 1 verified source, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The headline numbers — 11 of 12 wins against an 8× larger model — obscure the more interesting structural claim: Ant is competing on serving efficiency, not raw capability. A 5.1B activated parameter count means the model fits in a single H100's 80GB HBM with room for KV cache, enabling single-node inference that larger MoEs cannot achieve without tensor parallelism. This is the same economic logic that drove DeepSeek's V3 design, and Ant appears to be pushing it further. The timing with Jensen Huang's open-source comments is not coincidental. Huang's argument — that open models expand the Nvidia total addressable market by driving demand for inference hardware — directly benefits Ant, which needs Nvidia GPUs despite China's domestic chip push. Ant is effectively signaling to both the open-source community and Nvidia that it can match frontier efficiency without matching frontier parameter counts. The missing serving metrics are the real story. An execution model that doesn't publish latency or cost data is incomplete, and the 34-dimensional evaluation framework is opaque. The benchmark-gaming risk in Chinese model releases this year is real, and readers should treat the 11-of-12 headline as directional until third-party replication emerges.
Compare side-by-side
Ling-3.0-flash vs 1T-Ring-2.6
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Open Source

View all