epoch ai
30 articles about epoch ai in AI news
Epoch AI's 5-Variable Polynomial System Generates All Primes
Epoch AI published a five-variable polynomial system whose positive values are exactly the primes, a concrete Matiyasevich construction. The result ties to Hilbert's tenth problem and signals Epoch's pivot into pure math.
Epoch AI Opens FrontierMath's Unsolved Problems to Public Scrutiny After 2 Years
Epoch AI opened FrontierMath's unsolved problems to public scrutiny after two years, aiming to verify AI claims. The benchmark includes 1,000+ original math problems, with transparency seen as a step against benchmark gaming.
Epoch AI: Parallelization limits could delay intelligence explosion
Epoch AI argues parallelization limits could delay an intelligence explosion by years, as scaling beyond 10^28 FLOP faces diminishing returns from hardware constraints.
Epoch AI: Google's Colossus 1 Training Compute Hits 1e26 FLOP
Google's Colossus 1 used 1e26 FLOP at $4.6B, per Epoch AI. It is the largest known training run, signaling a new capital scale.
Epoch AI's EBR-Bench: Top Models Score 30-50% on Experience-Based Reasoning
Epoch AI's EBR-Bench tests experience-based reasoning. Top models score 30-50%, with Google Gemini 3 Pro leading at 48.2%, revealing a gap between pattern matching and true learning.
SciCode: Epoch AI Launches Benchmark Measuring AI Research Ability
Epoch AI launched SciCode benchmark testing LLMs on real research coding tasks. Top models score below 30%, exposing gap between coding benchmarks and scientific ability.
Epoch AI's CursorBench Benchmarks AI Code Editing at Scale
Epoch AI launched CursorBench, a 500-task benchmark for AI code editors. It reveals a 15% accuracy gap vs. humans and 3x latency variance.
MirrorCode: Epoch AI Tests If AI Can Rebuild 25 Unix Tools From Scratch
Epoch AI released MirrorCode, a 25-program benchmark testing AI's ability to reimplement software from scratch without source access, requiring exact stdout/stderr match.
Epoch AI: Hormuz LNG Shock Absorbed by Chip Margins, Gulf Investment is AI Risk
A new analysis from Epoch AI Research finds the Strait of Hormuz conflict's energy shock is manageable for AI infrastructure, but the real threat is the potential drying up of Gulf capital investment, crucial for projects like Stargate UAE.
GPT-4 Held ECI Lead for 18 Months, Epoch AI Data Shows
GPT-4 led the ECI for 18 months, the longest reign. GPT-4o and Claude 3.5 Sonnet broke the streak in September 2024.
AI Data Center Scale Doubles Every 7 Months, Epoch Finds
Epoch AI finds AI data center scale doubles every 7 months, driven by Google, Microsoft, and Amazon investments. This accelerates beyond the earlier 12-month cycle, raising training cost projections to $10 billion by 2028.
ECDLP Explained: Why Elliptic-Curve Crypto Still Holds
Epoch AI explains ECDLP, the math securing Bitcoin and Ethereum. No classical break exists; quantum Shor's remains theoretical, keeping 256-bit curves safe.
AI Finds First Rational Diophantine Septuple, Cracking 80-Year Conjecture
Epoch AI used AI-guided search to find the first rational Diophantine septuple, ending an 80-year conjecture by Erdős and Graham. The result, verified by formal proof, shows AI can solve long-stalled number theory problems.
OpenAI Stargate Abilene: $100B data center cluster breaks ground in Texas
OpenAI's $100B Stargate Abilene data center cluster in Texas targets 5 GW capacity by 2028, the largest single AI compute build. Epoch AI estimates power consumption equivalent to 5 nuclear reactors.
OSWorld 2.0 Launches, Tests AI Agents on 1,500 Desktop Tasks
Epoch AI released OSWorld 2.0 with 1,500 desktop tasks, up from 369 in v1, testing AI agents on adversarial and cross-application workflows.
MirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits
Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.
MirrorCode Rebuilds Programs from Behavior Alone, Beats GPT-4o by 37%
Epoch AI's MirrorCode reconstructs programs from I/O behavior alone, scoring 67.3% on SWE-bench—37% above GPT-4o—without source code or traces.
Nvidia B200 Costs $6,400 to Produce, Gross Margin Hits 82%
Epoch AI estimates Nvidia's B200 GPU costs $5,700–$7,300 to produce, with HBM memory and advanced packaging accounting for two-thirds of the cost. At a $30k–$40k sale price, chip-level gross margins reach ~82%, though rack-scale margins may be lower.
Open-Weight Models Trail Frontier AI by Four Months: EpochAI
EpochAI finds open-weight models trail frontier closed-source models by four months, a small gap reflecting rapid catch-up.
GPT-5.5 Pro Leapfrogs on Epoch Benchmark; Base Model Beats Prior Pro
A tweet from @kimmonismus reveals GPT-5.5 Pro shows significant Epoch benchmark gains, and the non-Pro GPT-5.5 surpasses GPT-5.4 Pro, suggesting major efficiency improvements at OpenAI.
WiT: Waypoint Diffusion Transformers Achieve FID 2.09 on ImageNet 256×256 in 265 Epochs, Matching JiT-L/16 Efficiency
Researchers introduced WiT, a diffusion transformer that uses semantic waypoints from pretrained vision models to resolve trajectory conflicts in pixel-space flow matching. It matches the performance of JiT-L/16 at 600 epochs in just 265 epochs, achieving an FID of 2.09 on ImageNet 256×256.
Frontier AI Labs Used Only 21% of Global Compute in 2025
Frontier labs used only 21% of global AI compute in 2025, per EpochAI, challenging the narrative of compute concentration.
gdb: Benchmarks Saturate Too Fast for Reliable AI Progress Tracking
@gdb notes benchmarks saturate quickly. This undermines AI progress tracking and may force shift to dynamic evaluations.
ByteDance Finds AI Agents Double Learning Speed Every 3 Months
ByteDance's Seed AI team discovered that AI agents double learning speed every three months via real-world interaction, per a Thursday paper. EdgeBench benchmark with 134 tasks ≥12 hours each underpins the finding.
Amazon Emissions Jump 16% as AI Data Center Demand Grows
Amazon emissions rose 16% to 82.3M metric tons in 2025, the largest jump since 2019, driven by AI data center electricity. The company maintains its net-zero pledge but faces skepticism.
Water Joins Power as AI Data Center Site Selection Constraint
Water joins power as AI data center site selection constraint, per Emergence Water and Nimbus. Communities now scrutinize water use like electrical demand, affecting campus viability.
Colossus 2: xAI's Memphis Cluster Hits 300,000 GPUs
xAI's Colossus 2 hits 300,000 GPUs, targeting 1M by year-end. Training Grok-3, the $6B cluster challenges OpenAI and Google.
SemiAnalysis: Pretraining Dead for All but Frontier Labs
@SemiAnalysis_ declares pretraining dead for non-frontier labs, citing 'Pretrainitis' as vanity-driven waste. Prompt engineering offers higher ROI.
Time's First AI A-List: Alibaba, ByteDance, Zhipu AI Make Cut
Time magazine named Alibaba, ByteDance, and Zhipu AI among its first AI-specific top 10 list, alongside six US companies and France's Mistral AI. The recognition highlights China's growing global influence through open-source models and consumer AI apps.
A Practical Guide to Fine-Tuning Open-Source LLMs for AI Agents
This Portuguese-language Medium article is Part 2 of a series on LLM engineering for AI agents. It provides a hands-on guide to fine-tuning an open-source model, building on a foundation of clean data and established baselines from Part 1.