DeepSeek open-sourced V4-Flash-0731 on July 31, scoring 82.7 on VulcanBench and matching Claude Opus-4.8. The 304B-parameter model undercuts V4-Pro preview pricing by 65%, at $0.14 per million input tokens.
Key facts
- Released July 31, 2026 as open-source
- 304B parameters, 45% of V4's 671B
- VulcanBench score: 82.7, matches Opus-4.8
- Pricing: $0.14/$0.28 per M tokens
- Hugging Face trending: #2 within 24h
DeepSeek released V4-Flash-0731 as an open-source model on July 31, 2026, a 304B-parameter lightweight variant that the company says outperforms its own V4-Pro preview. According to Pandaily, the model scored 82.7 on VulcanBench, topping the leaderboard and matching Anthropic's Claude Opus-4.8 in the top performance tier.
Pricing and positioning
The model is priced at $0.14 per million input tokens and $0.28 per million output tokens — roughly a third of V4-Pro preview's rates. This aggressive pricing puts DeepSeek's flagship-tier performance at commodity cost, directly targeting the coding-agent market where Claude Code and its Opus 4.6 backend currently dominate. On Hugging Face, the release reached the second spot on the trending models list within 24 hours.
The 304B parameter count is notable: it's roughly 45% of V4's 671B total, yet the company claims parity on VulcanBench's top tier. That benchmark, which tests agentic coding and tool-use scenarios, is the same one where Claude Code with Opus 4.8 scores 78.9% on Terminal-Bench 2.1, per Anthropic's public figures.
What the benchmark gap means
The VulcanBench 82.7 score is a single number, and DeepSeek did not disclose the full evaluation methodology or variance across runs. The company also didn't specify hardware requirements for self-hosting the 304B model — a meaningful omission for enterprises weighing on-prem deployment against API use. The open-source release includes weights and inference code, but no fine-tuning scripts or training data, per the Pandaily report.
This is the third time DeepSeek has reset the price-performance curve — following V3 in December 2024 and R1 in January 2025. Each release forced Western labs to respond on cost. Anthropic's Opus 5, shipped August 1 at half the price of Opus 4.6, suggests the pressure is already being felt.
What to watch
Watch for independent replication of the 82.7 VulcanBench score by third-party evaluators, and whether Anthropic responds with a price cut on Opus 4.8 or accelerates Opus 5 GA. Also track enterprise adoption of self-hosted V4-Flash — DeepSeek hasn't disclosed hardware specs, which will determine real-world deployment costs.
Source: pandaily.com
[Updated 03 Aug via scmp_tech]
DeepSeek is now recruiting open-source developers to beta test its upcoming ‘harness’ software, which converts LLMs into AI agents, according to a Saturday post by Cui Tianyi, who leads the project [per SCMP]. This signals a strategic push into agentic AI, expanding beyond the V4-Flash release. The harness could integrate with the new model to compete directly with coding agents like Claude Code, potentially reshaping the agentic tooling landscape.









