Epoch AI
The environmental impact of artificial intelligence includes substantial electricity consumption for training and the usage of deep learning models, as well as the related carbon footprint and water usage impact. Moreover, artificial intelligence (AI) data centers are materially intense, requiring a
Signal Radar
Five-axis snapshot of this entity's footprint
Mentions × Lab Attention
Weekly mentions (solid) and average article relevance (dotted)
Timeline
10- Product LaunchJun 30, 2026
Released EBR-Bench benchmark for measuring experience-based reasoning in AI models
View source - Product LaunchJun 30, 2026
Released MirrorCode, a benchmark and method for program reconstruction from I/O behavior
View source - Product LaunchJun 30, 2026
Released MirrorCode benchmark for zero-shot software reimplementation
View source - Product LaunchJun 27, 2026
Published CursorBench, a 500+ task benchmark for AI code editors.
View source- tasks:
- 500
- Product LaunchJun 27, 2026
Launched SciCode benchmark for evaluating LLMs on scientific research coding tasks.
View source - Product LaunchJun 27, 2026
Released OSWorld 2.0 with 1,500 desktop tasks, up from 369 in v1.
View source- tasks:
- 1500
- Research MilestoneJun 25, 2026
Report finds AI data center scale doubles every 7 months, accelerating from 12-month cycle
View source - Research MilestoneMay 29, 2026
Research finds open-weight models trail frontier closed-source models by four months
View source - Research MilestoneDec 31, 2025
Published data showing frontier labs used 21% of global AI compute in late 2025
View source
Relationships
12Developed
Regulated
Frequently appears with
6Entities that show up in the same articles — shared coverage, not a stated relationship.
Recent Articles
12Epoch AI: Google's Colossus 1 Training Compute Hits 1e26 FLOP
~Google's Colossus 1 used 1e26 FLOP at $4.6B, per Epoch AI. It is the largest known training run, signaling a new capital scale.
85 relevanceOpenAI Stargate Abilene: $100B data center cluster breaks ground in Texas
~OpenAI's $100B Stargate Abilene data center cluster in Texas targets 5 GW capacity by 2028, the largest single AI compute build. Epoch AI estimates po
100 relevanceGPT-4 Held Top Spot 52 Weeks; Today's Models Last 7
~GPT-4 dominated the ECI for a year. Today's top models last 7 weeks median, with 17 leadership changes since Feb 2024.
84 relevanceGPT-4 Held ECI Lead for 18 Months, Epoch AI Data Shows
~GPT-4 led the ECI for 18 months, the longest reign. GPT-4o and Claude 3.5 Sonnet broke the streak in September 2024.
93 relevanceEpoch AI's EBR-Bench: Top Models Score 30-50% on Experience-Based Reasoning
+Epoch AI's EBR-Bench tests experience-based reasoning. Top models score 30-50%, with Google Gemini 3 Pro leading at 48.2%, revealing a gap between pat
100 relevanceFrontier AI Labs Used Only 21% of Global Compute in 2025
+Frontier labs used only 21% of global AI compute in 2025, per EpochAI, challenging the narrative of compute concentration.
91 relevanceMirrorCode Rebuilds Programs from Behavior Alone, Beats GPT-4o by 37%
+Epoch AI's MirrorCode reconstructs programs from I/O behavior alone, scoring 67.3% on SWE-bench—37% above GPT-4o—without source code or traces.
100 relevanceSciCode: Epoch AI Launches Benchmark Measuring AI Research Ability
+Epoch AI launched SciCode benchmark testing LLMs on real research coding tasks. Top models score below 30%, exposing gap between coding benchmarks and
95 relevanceEpoch AI's CursorBench Benchmarks AI Code Editing at Scale
+Epoch AI launched CursorBench, a 500-task benchmark for AI code editors. It reveals a 15% accuracy gap vs. humans and 3x latency variance.
95 relevanceOSWorld 2.0 Launches, Tests AI Agents on 1,500 Desktop Tasks
+Epoch AI released OSWorld 2.0 with 1,500 desktop tasks, up from 369 in v1, testing AI agents on adversarial and cross-application workflows.
95 relevanceMirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits
+Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.
77 relevanceMirrorCode: Epoch AI Tests If AI Can Rebuild 25 Unix Tools From Scratch
~Epoch AI released MirrorCode, a 25-program benchmark testing AI's ability to reimplement software from scratch without source access, requiring exact
82 relevance
Predictions
No predictions linked to this entity.
AI Discoveries
7- observationactive8h ago
Lifecycle: Epoch AI
Epoch AI is in 'surging' phase (1 mentions/3d, 1/14d, 19 total)
90% confidence - observationactiveJul 17, 2026
[Compressed] Institutional knowledge: Epoch AI
TRAJECTORY: Our understanding of Epoch AI evolved from a niche compute tracker to a rapidly consolidating benchmark authority for agentic AI, revealing both a dangerous research-to-product disconnect and an accelerating institutional influence. KEY FACTS: - Epoch AI surged from 1 to 4 mentions in 3
80% confidence - hypothesisactiveJul 7, 2026
H: Within 60 days, Epoch AI will publish a paper or report that uses GPT-4 Turbo as a case study for a
Within 60 days, Epoch AI will publish a paper or report that uses GPT-4 Turbo as a case study for a new capability measurement methodology, in line with the DARPA AIQ program's shift away from traditional benchmarks.
70% confidence - observationactiveJul 7, 2026
Novel co-occurrence: Epoch AI + GPT-4 Turbo
Epoch AI (organization) and GPT-4 Turbo (ai_model) appeared together in 2 articles this week but have NEVER co-occurred before and have no existing relationship. This is a potential breaking story signal.
85% confidence - observationactiveJul 6, 2026
Investigation: Epoch AI
Assessment: Epoch AI is rapidly emerging as the dominant benchmark infrastructure provider for agentic AI, releasing 5 major benchmarks in a single week (EBR-Bench, MirrorCode, CursorBench, SciCode, OSWorld 2.0). However, its trajectory shows a dangerous isolation: zero co-occurrence with major comp
70% confidence - hypothesisactiveJul 6, 2026
H: OpenAI will explicitly reject Epoch AI's benchmarks (EBR-Bench, MirrorCode, OSWorld 2.0) as insuffic
OpenAI will explicitly reject Epoch AI's benchmarks (EBR-Bench, MirrorCode, OSWorld 2.0) as insufficient for evaluating agentic reasoning within 60 days, citing methodological concerns about test-time compute leakage.
72% confidence - hypothesisactiveJul 6, 2026
H: Within 90 days, at least one of Anthropic, Meta, or Microsoft will announce a competing 'unified age
Within 90 days, at least one of Anthropic, Meta, or Microsoft will announce a competing 'unified agent benchmark suite' that directly challenges Epoch's portfolio, claiming superior coverage of real-world tasks.
65% confidence
Sentiment History
| Week | Avg Sentiment | Mentions |
|---|---|---|
| 2026-W26 | 0.30 | 5 |
| 2026-W27 | 0.30 | 6 |
| 2026-W28 | 0.05 | 2 |
| 2026-W30 | 0.10 | 1 |