Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

epoch ai

30 articles about epoch ai in AI news

Epoch AI's 5-Variable Polynomial System Generates All Primes

Epoch AI published a five-variable polynomial system whose positive values are exactly the primes, a concrete Matiyasevich construction. The result ties to Hilbert's tenth problem and signals Epoch's pivot into pure math.

85% relevant

Epoch AI Opens FrontierMath's Unsolved Problems to Public Scrutiny After 2 Years

Epoch AI opened FrontierMath's unsolved problems to public scrutiny after two years, aiming to verify AI claims. The benchmark includes 1,000+ original math problems, with transparency seen as a step against benchmark gaming.

100% relevant

Epoch AI: Parallelization limits could delay intelligence explosion

Epoch AI argues parallelization limits could delay an intelligence explosion by years, as scaling beyond 10^28 FLOP faces diminishing returns from hardware constraints.

99% relevant

Epoch AI: Google's Colossus 1 Training Compute Hits 1e26 FLOP

Google's Colossus 1 used 1e26 FLOP at $4.6B, per Epoch AI. It is the largest known training run, signaling a new capital scale.

100% relevant

Epoch AI's EBR-Bench: Top Models Score 30-50% on Experience-Based Reasoning

Epoch AI's EBR-Bench tests experience-based reasoning. Top models score 30-50%, with Google Gemini 3 Pro leading at 48.2%, revealing a gap between pattern matching and true learning.

100% relevant

SciCode: Epoch AI Launches Benchmark Measuring AI Research Ability

Epoch AI launched SciCode benchmark testing LLMs on real research coding tasks. Top models score below 30%, exposing gap between coding benchmarks and scientific ability.

95% relevant

Epoch AI's CursorBench Benchmarks AI Code Editing at Scale

Epoch AI launched CursorBench, a 500-task benchmark for AI code editors. It reveals a 15% accuracy gap vs. humans and 3x latency variance.

95% relevant

MirrorCode: Epoch AI Tests If AI Can Rebuild 25 Unix Tools From Scratch

Epoch AI released MirrorCode, a 25-program benchmark testing AI's ability to reimplement software from scratch without source access, requiring exact stdout/stderr match.

82% relevant

Epoch AI: Hormuz LNG Shock Absorbed by Chip Margins, Gulf Investment is AI Risk

A new analysis from Epoch AI Research finds the Strait of Hormuz conflict's energy shock is manageable for AI infrastructure, but the real threat is the potential drying up of Gulf capital investment, crucial for projects like Stargate UAE.

85% relevant

GPT-4 Held ECI Lead for 18 Months, Epoch AI Data Shows

GPT-4 led the ECI for 18 months, the longest reign. GPT-4o and Claude 3.5 Sonnet broke the streak in September 2024.

93% relevant

AI Data Center Scale Doubles Every 7 Months, Epoch Finds

Epoch AI finds AI data center scale doubles every 7 months, driven by Google, Microsoft, and Amazon investments. This accelerates beyond the earlier 12-month cycle, raising training cost projections to $10 billion by 2028.

95% relevant

ECDLP Explained: Why Elliptic-Curve Crypto Still Holds

Epoch AI explains ECDLP, the math securing Bitcoin and Ethereum. No classical break exists; quantum Shor's remains theoretical, keeping 256-bit curves safe.

88% relevant

AI Finds First Rational Diophantine Septuple, Cracking 80-Year Conjecture

Epoch AI used AI-guided search to find the first rational Diophantine septuple, ending an 80-year conjecture by Erdős and Graham. The result, verified by formal proof, shows AI can solve long-stalled number theory problems.

85% relevant

OpenAI Stargate Abilene: $100B data center cluster breaks ground in Texas

OpenAI's $100B Stargate Abilene data center cluster in Texas targets 5 GW capacity by 2028, the largest single AI compute build. Epoch AI estimates power consumption equivalent to 5 nuclear reactors.

100% relevant

OSWorld 2.0 Launches, Tests AI Agents on 1,500 Desktop Tasks

Epoch AI released OSWorld 2.0 with 1,500 desktop tasks, up from 369 in v1, testing AI agents on adversarial and cross-application workflows.

95% relevant

MirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits

Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.

77% relevant

MirrorCode Rebuilds Programs from Behavior Alone, Beats GPT-4o by 37%

Epoch AI's MirrorCode reconstructs programs from I/O behavior alone, scoring 67.3% on SWE-bench—37% above GPT-4o—without source code or traces.

100% relevant

Nvidia B200 Costs $6,400 to Produce, Gross Margin Hits 82%

Epoch AI estimates Nvidia's B200 GPU costs $5,700–$7,300 to produce, with HBM memory and advanced packaging accounting for two-thirds of the cost. At a $30k–$40k sale price, chip-level gross margins reach ~82%, though rack-scale margins may be lower.

100% relevant

Open-Weight Models Trail Frontier AI by Four Months: EpochAI

EpochAI finds open-weight models trail frontier closed-source models by four months, a small gap reflecting rapid catch-up.

79% relevant

GPT-5.5 Pro Leapfrogs on Epoch Benchmark; Base Model Beats Prior Pro

A tweet from @kimmonismus reveals GPT-5.5 Pro shows significant Epoch benchmark gains, and the non-Pro GPT-5.5 surpasses GPT-5.4 Pro, suggesting major efficiency improvements at OpenAI.

99% relevant

WiT: Waypoint Diffusion Transformers Achieve FID 2.09 on ImageNet 256×256 in 265 Epochs, Matching JiT-L/16 Efficiency

Researchers introduced WiT, a diffusion transformer that uses semantic waypoints from pretrained vision models to resolve trajectory conflicts in pixel-space flow matching. It matches the performance of JiT-L/16 at 600 epochs in just 265 epochs, achieving an FID of 2.09 on ImageNet 256×256.

85% relevant

Frontier AI Labs Used Only 21% of Global Compute in 2025

Frontier labs used only 21% of global AI compute in 2025, per EpochAI, challenging the narrative of compute concentration.

91% relevant

gdb: Benchmarks Saturate Too Fast for Reliable AI Progress Tracking

@gdb notes benchmarks saturate quickly. This undermines AI progress tracking and may force shift to dynamic evaluations.

75% relevant

ByteDance Finds AI Agents Double Learning Speed Every 3 Months

ByteDance's Seed AI team discovered that AI agents double learning speed every three months via real-world interaction, per a Thursday paper. EdgeBench benchmark with 134 tasks ≥12 hours each underpins the finding.

100% relevant

Amazon Emissions Jump 16% as AI Data Center Demand Grows

Amazon emissions rose 16% to 82.3M metric tons in 2025, the largest jump since 2019, driven by AI data center electricity. The company maintains its net-zero pledge but faces skepticism.

82% relevant

Water Joins Power as AI Data Center Site Selection Constraint

Water joins power as AI data center site selection constraint, per Emergence Water and Nimbus. Communities now scrutinize water use like electrical demand, affecting campus viability.

95% relevant

Colossus 2: xAI's Memphis Cluster Hits 300,000 GPUs

xAI's Colossus 2 hits 300,000 GPUs, targeting 1M by year-end. Training Grok-3, the $6B cluster challenges OpenAI and Google.

98% relevant

SemiAnalysis: Pretraining Dead for All but Frontier Labs

@SemiAnalysis_ declares pretraining dead for non-frontier labs, citing 'Pretrainitis' as vanity-driven waste. Prompt engineering offers higher ROI.

85% relevant

Time's First AI A-List: Alibaba, ByteDance, Zhipu AI Make Cut

Time magazine named Alibaba, ByteDance, and Zhipu AI among its first AI-specific top 10 list, alongside six US companies and France's Mistral AI. The recognition highlights China's growing global influence through open-source models and consumer AI apps.

74% relevant

A Practical Guide to Fine-Tuning Open-Source LLMs for AI Agents

This Portuguese-language Medium article is Part 2 of a series on LLM engineering for AI agents. It provides a hands-on guide to fine-tuning an open-source model, building on a foundation of clean data and established baselines from Part 1.

74% relevant