DeepSeek V4 Pro 0813 (1.5T) and V4 Flash 0731 massively beat Nemotron3 Ultra on agentic tasks, per @SemiAnalysis_. The Flash model does so with 4.2x fewer active parameters.
Key facts
- DeepSeek V4 Pro 0813 is a 1.5T-parameter model.
- V4 Flash 0731 uses 4.2x fewer active parameters than Nemotron3 Ultra.
- V4 Flash has close to 2x fewer total parameters than Nemotron3 Ultra.
- SemiAnalysis claims both models 'massively beat' Nemotron3 Ultra on agentic tasks.
- No benchmark numbers or methodology were disclosed in the tweet.
Key Takeaways
- DeepSeek V4 Pro 0813 (1.5T) and Flash 0731 beat Nemotron3 Ultra on agentic tasks per SemiAnalysis, with Flash using 4.2x fewer active params.
- No benchmarks disclosed.
The Claim
DeepSeek V4 Pro 0813, a 1.5T-parameter model, massively beats Nemotron3 Ultra on agentic tasks, according to a tweet from @SemiAnalysis_. The same thread claims DeepSeek V4 Flash 0731 also beats Nemotron3 Ultra massively, while having 4.2x fewer active parameters and close to 2x fewer total parameters. SemiAnalysis, a respected compute-focused research firm, is not in the habit of handing out praise; their endorsement carries weight in the AI infrastructure community.
The Missing Numbers
The tweet provides no benchmark numbers, ablation details, or evaluation methodology. No mention of specific agentic benchmarks like SWE-bench, GAIA, or Terminal-Bench. No inference cost figures. No context window specs. The source is a congrats thread, not a technical report. As of now, DeepSeek's official blog and arXiv have not published V4 Pro 0813 details — the company did not disclose the figure for active parameters in the tweet, and the total parameter count for the 1.5T model is implied by the name.
Why This Matters
SemiAnalysis argues that committee-based model frontier development does not work, despite brilliant people working on Nemotron. This is a direct shot at NVIDIA's model-building strategy, which relies on large internal teams and extensive RLHF pipelines. DeepSeek, by contrast, is known for aggressive efficiency — their V3 training ran on 2,048 H800 GPUs for about $5.6M, as previously reported. If V4 Flash maintains that efficiency at scale, the gap between frontier labs and efficient challengers is narrowing faster than most enterprise buyers expect.
The Efficiency Angle
The 4.2x active-parameter advantage is the real story. Active parameters determine inference FLOPs and cost. If Flash delivers comparable agentic performance with 4.2x fewer active params, the cost-per-task ratio flips in DeepSeek's favor. For enterprises running agent loops at scale, that delta translates directly into GPU hours and cloud bills. Nemotron3 Ultra, presumably a dense or MoE model, would need to justify its larger footprint with meaningfully better accuracy — which the tweet suggests it does not.
Caveats
This is a single tweet, not a peer-reviewed benchmark. No code, no weights, no evaluation harness. The "massively beats" phrasing is qualitative. SemiAnalysis has been bullish on DeepSeek before, but their track record on compute forecasts is solid. Until independent evals surface, treat this as a strong signal, not a settled fact.
What to watch
Watch for DeepSeek's official V4 technical report or arXiv paper, which should disclose active parameter counts, benchmark scores, and training compute. Also track independent evaluations on agentic benchmarks like SWE-bench or GAIA from third-party labs. If NVIDIA responds with a Nemotron3 Ultra update, that signals real competitive pressure.




