Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Line chart comparing AI model intelligence scores, with DeepSeek V4 Flash 0731 at 50 and GPT-5.6 Luna at 51
Open SourceBreakthroughScore: 100

DeepSeek V4 Flash 0731 Hits 50 on Intelligence Index at $0.14/M Tokens

DeepSeek V4 Flash 0731 scores 50 on Intelligence Index, one point behind GPT-5.6 Luna at ~60% lower cost. 304B params, $0.14/M input pricing.

·1d ago·4 min read··32 views·AI-Generated·Report error
Share:
Source: simonwillison.netvia simon_willison, the_decoder, openai_blog, reddit_anthropic, @kimmonismusMulti-Source
How does DeepSeek V4 Flash 0731 compare on the Artificial Analysis Intelligence Index and what does it cost?

DeepSeek's V4 Flash 0731 update, a 304B-parameter model, scored 50 on the Artificial Analysis Intelligence Index — one point behind OpenAI's GPT-5.6 Luna at roughly 60% lower cost per task. It's priced at $0.14/M input and $0.27/M output tokens, ranking ahead of MiniMax's 428B-parameter M3.

TL;DR

304B-param model scores 50 on Artificial Analysis Intelligence Index · Matches GPT-5.6 Luna at roughly 60% lower cost per task · $0.14/M input, $0.27/M output pricing undercuts MiniMax M3

DeepSeek's V4 Flash 0731 update jumped ten points to 50 on the Artificial Analysis Intelligence Index, per The Decoder. The 304B-parameter model lands one point behind OpenAI's GPT-5.6 Luna at roughly 60% lower cost per task.

Key facts

  • Intelligence Index score: 50, up 10 points from prior version
  • Pricing: $0.14/M input, $0.27/M output tokens
  • 304B parameters, 167GB on Hugging Face
  • Ranks ahead of MiniMax M3 (428B params)
  • One point behind GPT-5.6 Luna at ~60% lower cost

DeepSeek shipped a major revision to its budget model line on July 31, 2026. The V4 Flash 0731 update, a 304B-parameter model weighing 167GB on Hugging Face, scores 50 on the Artificial Analysis Intelligence Index — a ten-point jump from the prior version per The Decoder. That puts it one point behind OpenAI's GPT-5.6 Luna, at roughly 60% lower cost per task.

The pricing is the headline. At $0.14 per million input tokens and $0.27 per million output tokens, the model undercuts nearly every frontier competitor according to Simon Willison. Artificial Analysis ranks it ahead of MiniMax's M3, a substantially larger 428B-parameter model. On the Intelligence Index vs. Cost per Task chart, V4 Flash 0731 sits alone in the "most attractive quadrant," where the Pareto line jumps sharply upward — roughly $0.028 per task at an intelligence score of 50. Models of similar intelligence like MiniMax-M3, Kimi K3 (low), and GLM-5.1 cost ten times more.

The reasoning-effort gap

Simon Willison's hands-on test exposed a real quality cliff. Using the default reasoning level via OpenRouter, the model produced a "disappointing pelican" — a mangled bicycle with floating frame tubes and disconnected handlebars. With -o reasoning_effort high, the output improved dramatically: a coherent pelican gripping handlebars, one orange foot on the pedal, a small fish tucked in its beak pouch. The delta between default and high reasoning effort is a practical consideration for agentic workloads, not a theoretical one.

The company touts "substantially enhanced agentic capabilities" in the release notes per the Hugging Face model card. Hacker News commenters confirm the model is a daily driver for coding work — one user reports using it with reasonix or pi "all day long" for pennies. But the pelican test suggests default settings may not deliver the advertised agentic quality; users need to explicitly raise reasoning effort.

The value-per-intelligence case

The 60% cost advantage over GPT-5.6 Luna is the structural story here. DeepSeek continues to compress frontier-adjacent intelligence into a smaller, cheaper package — the same play that made V3 and R1 notable in late 2024 and early 2025. The company is simultaneously building gigawatt-scale data centers in Inner Mongolia and developing custom inference ASICs to cut GPU dependency [as previously reported]. The Flash line is the commercial front of that infrastructure bet.

What the source material does not disclose: training compute, architecture details beyond parameter count, or context window specifications. The company released no technical report alongside the model card. The 304B parameter count and 167GB size are the only hard numbers available.

Key Takeaways

  • deepseek-v4-flash-0731" class="entity-chip">DeepSeek V4 Flash 0731 scores 50 on Intelligence Index, one point behind GPT-5.6 Luna at ~60% lower cost.
  • 304B params, $0.14/M input pricing.

What to watch

Watch for DeepSeek's DeepSeek Harness release — the model card mentions it was evaluated with the 'minimal mode' of this unreleased agent framework. If the harness ships and lifts agentic benchmark scores, expect the Flash line to pressure OpenAI and Anthropic pricing further. Also track the $71B pre-money funding round reported in July.

Scatter plot from Artificial Analysis titled with axes


Source: simonwillison.net

[Updated 01 Aug via simon_willison]

OpenAI responded to DeepSeek's pricing pressure with an aggressive counter-move: GPT-5.6 Luna received an 80% price cut, and GPT-5.6 Terra dropped 20%, per OpenAI's announcement per OpenAI. The company credits GPT-5.6 Sol for the efficiency gains, using it to optimize load balancing and even rewrite production kernels via Codex to reduce GPU idle time. This directly challenges DeepSeek's value-per-intelligence advantage, narrowing the cost gap that made V4 Flash 0731 the standout budget model.


Sources cited in this article

  1. The Decoder. The
  2. The Decoder
  3. Task
  4. OpenAI's
  5. OpenAI
Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 6 verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

DeepSeek's Flash 0731 update is the clearest evidence yet that the company's cost-per-intelligence curve is flattening the frontier. The ten-point Intelligence Index jump in a single revision — without a parameter count increase disclosed — suggests the gains come from training efficiency or inference-time reasoning improvements rather than raw scale. The 304B-parameter model beating MiniMax's 428B model is the kind of inversion that makes the parameter-count arms race look increasingly irrelevant. The reasoning-effort cliff is the more interesting finding. Simon Willison's pelican test shows a dramatic quality gap between default and high reasoning settings. That's not a bug — it's a pricing strategy. DeepSeek can advertise cheap tokens while pushing users toward higher-effort settings that consume more compute. The HN commenter's experience coding all day for pennies suggests the default setting is fine for code generation, but the gap matters for agents that need spatial reasoning or multi-step planning. The structural threat to OpenAI is the cost-per-task curve. At $0.028 per task, V4 Flash 0731 sits an order of magnitude below GPT-5.6 Sol and Grok 4.5 at $0.4 to $3 per task. If DeepSeek's custom ASICs and gigawatt-scale data centers materialize, that gap widens. The funding round at $71B pre-money values DeepSeek as a serious infrastructure player, not just a model lab.
This story is part of
Hugging Face Becomes the Neutral Ground Where Google and Anthropic's Agent Protocol War Converges
As Claude Code's MCP dominance threatens Google Cloud, Hugging Face's unique position as partner to both players creates an unexpected convergence zone
Compare side-by-side
OpenAI vs DeepSeek
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Open Source

View all