Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

A sleek AI chip with glowing circuitry sits on a dark workbench, surrounded by performance graphs and data charts…

DeepSeek V4 Pro 1.5T Beats Nemotron3 Ultra on Agentic Tasks

DeepSeek V4 Pro 0813 (1.5T) and Flash 0731 beat Nemotron3 Ultra on agentic tasks per SemiAnalysis, with Flash using 4.2x fewer active params. No benchmarks disclosed.

·11h ago·3 min read··17 views·AI-Generated·Report error
Share:
How does DeepSeek V4 Pro compare to Nemotron3 Ultra on agentic benchmarks?

DeepSeek V4 Pro 0813, a 1.5T-parameter model, massively beats Nemotron3 Ultra on agentic tasks, according to SemiAnalysis. DeepSeek V4 Flash 0731 also beats Nemotron3 Ultra with 4.2x fewer active parameters and close to 2x fewer total parameters.

TL;DR

DeepSeek V4 Pro 0813 1.5T massively beats Nemotron3 Ultra. · Flash 0731 also wins with 4.2x fewer active params. · SemiAnalysis says committee-based frontier dev fails.

DeepSeek V4 Pro 0813 (1.5T) and V4 Flash 0731 massively beat Nemotron3 Ultra on agentic tasks, per @SemiAnalysis_. The Flash model does so with 4.2x fewer active parameters.

Key facts

  • DeepSeek V4 Pro 0813 is a 1.5T-parameter model.
  • V4 Flash 0731 uses 4.2x fewer active parameters than Nemotron3 Ultra.
  • V4 Flash has close to 2x fewer total parameters than Nemotron3 Ultra.
  • SemiAnalysis claims both models 'massively beat' Nemotron3 Ultra on agentic tasks.
  • No benchmark numbers or methodology were disclosed in the tweet.

Key Takeaways

  • DeepSeek V4 Pro 0813 (1.5T) and Flash 0731 beat Nemotron3 Ultra on agentic tasks per SemiAnalysis, with Flash using 4.2x fewer active params.
  • No benchmarks disclosed.

The Claim

DeepSeek V4 Pro 0813, a 1.5T-parameter model, massively beats Nemotron3 Ultra on agentic tasks, according to a tweet from @SemiAnalysis_. The same thread claims DeepSeek V4 Flash 0731 also beats Nemotron3 Ultra massively, while having 4.2x fewer active parameters and close to 2x fewer total parameters. SemiAnalysis, a respected compute-focused research firm, is not in the habit of handing out praise; their endorsement carries weight in the AI infrastructure community.

The Missing Numbers

The tweet provides no benchmark numbers, ablation details, or evaluation methodology. No mention of specific agentic benchmarks like SWE-bench, GAIA, or Terminal-Bench. No inference cost figures. No context window specs. The source is a congrats thread, not a technical report. As of now, DeepSeek's official blog and arXiv have not published V4 Pro 0813 details — the company did not disclose the figure for active parameters in the tweet, and the total parameter count for the 1.5T model is implied by the name.

Why This Matters

SemiAnalysis argues that committee-based model frontier development does not work, despite brilliant people working on Nemotron. This is a direct shot at NVIDIA's model-building strategy, which relies on large internal teams and extensive RLHF pipelines. DeepSeek, by contrast, is known for aggressive efficiency — their V3 training ran on 2,048 H800 GPUs for about $5.6M, as previously reported. If V4 Flash maintains that efficiency at scale, the gap between frontier labs and efficient challengers is narrowing faster than most enterprise buyers expect.

The Efficiency Angle

The 4.2x active-parameter advantage is the real story. Active parameters determine inference FLOPs and cost. If Flash delivers comparable agentic performance with 4.2x fewer active params, the cost-per-task ratio flips in DeepSeek's favor. For enterprises running agent loops at scale, that delta translates directly into GPU hours and cloud bills. Nemotron3 Ultra, presumably a dense or MoE model, would need to justify its larger footprint with meaningfully better accuracy — which the tweet suggests it does not.

Caveats

This is a single tweet, not a peer-reviewed benchmark. No code, no weights, no evaluation harness. The "massively beats" phrasing is qualitative. SemiAnalysis has been bullish on DeepSeek before, but their track record on compute forecasts is solid. Until independent evals surface, treat this as a strong signal, not a settled fact.

What to watch

Watch for DeepSeek's official V4 technical report or arXiv paper, which should disclose active parameter counts, benchmark scores, and training compute. Also track independent evaluations on agentic benchmarks like SWE-bench or GAIA from third-party labs. If NVIDIA responds with a Nemotron3 Ultra update, that signals real competitive pressure.

Sources cited in this article

  1. SemiAnalysis
Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 1 verified source, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

SemiAnalysis's claim is notable for its source: a firm that tracks GPU supply chains and training costs, not a benchmark lab. Their endorsement suggests the performance delta is real enough to be visible in their compute models. But the absence of numbers is a red flag — 'massively beats' without a benchmark table is marketing language, not engineering evidence. Comparing to prior art: DeepSeek V3 already demonstrated that a $5.6M training run could rival models costing ten times more. V4 Flash's 4.2x active-parameter advantage extends that efficiency story into the inference regime. If agentic tasks are compute-bound at inference time, the cost per successful task could be an order of magnitude lower than Nemotron3 Ultra. That's the structural shift worth watching. The committee-based development critique is a broader argument: NVIDIA's model team may be too consensus-driven to iterate quickly. DeepSeek's smaller, more autonomous team has a track record of shipping bold architectural choices. But one tweet doesn't settle that debate — we need to see the model card and independent evals before concluding NVIDIA's approach is broken.
This story is part of
Hugging Face Becomes the Neutral Ground Where Google and Anthropic's Agent Protocol War Converges
As Claude Code's MCP dominance threatens Google Cloud, Hugging Face's unique position as partner to both players creates an unexpected convergence zone
Compare side-by-side
DeepSeek V4 Pro 0813 vs DeepSeek V4 Flash 0731
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all