peer review
30 articles about peer review in AI news
Chinese Team Claims Carbon Nanotube CFET Breakthrough; Challenges TSMC at 2nm
Chinese team claims 3x carbon nanotube CFET gain over silicon at 2nm, bypassing EUV. No peer review; skepticism warranted.
Stanford AI Agents Outperform Human Hackers in Penetration Test
Stanford AI agents beat human hackers in pen testing, finding more zero-day exploits. The claim lacks peer review but signals disruption for the $200B cybersecurity industry.
Study: Rhetoric Shifts AI Review Scores Across 4,200 Papers
Study across 4,200 manuscripts shows rhetoric alone shifts AI review scores. Maps 6 dimensions and 42K+ reviews to identify exploitable stylistic levers.
Claude Agentic Framework Uses 20 Specialized Agents to Enforce a 3-Stage
The Claude Agentic Framework enforces a Spec → Build → Review pipeline with 20 specialized agents and PowerShell hooks, preventing Claude Code from coding too early or finishing incomplete.
Microsoft's Study Proves Claude Code Boosts PR Output by 24%
Claude Code users merge 24% more PRs per Microsoft's study. Drive adoption via peer visibility, not mandates. Retention correlates with coding activity, not demographics.
Tencent's HY3 AI Model Has 295B Params, Led by Ex-OpenAI Researcher
Tencent unveiled its HY3 preview model, its most powerful yet with 295 billion parameters. It's already deployed in consumer app Yuanbao and coding assistant CodeBuddy.
Humwork AI Launches A2P Marketplace, Shifts Humans to On-Demand Fallback
Humwork AI has launched a marketplace where AI agents execute work end-to-end, fundamentally shifting the labor model from peer-to-peer (P2P) to agent-to-peer (A2P). This repositions humans from default workers to an on-demand fallback layer, a significant threshold for AI agent economics.
LLMs Show 'Privileged Access' to Own Policies in Introspect-Bench, Explaining Self-Knowledge via Attention Diffusion
Researchers formalize LLM introspection as computation over model parameters, showing frontier models outperform peers at predicting their own behavior. The study provides causal evidence for how introspection emerges via attention diffusion without explicit training.
Intuition First or Reflection Before Judgment? How Evaluation Sequence Polarizes Consumer Ratings
New research reveals that asking for a star rating *before* a written review leads to more extreme, polarized scores. This 'Rating-First' design amplifies gut reactions, significantly impacting perceived product quality and platform credibility.
Claude Code's Architecture Explained
Master Claude Code's 6-layer architecture—harness, context window, subagents, MCP, skills, hooks, CLAUDE.md—to optimize token usage, parallelism, and agent reliability. First sentence: Claude Code's 6 layers (harness, context window, subagents, MCP, skills, hooks, CLAUDE.md) determine your setup's success.
DeepSeek V4 Pro 1.5T Beats Nemotron3 Ultra on Agentic Tasks
DeepSeek V4 Pro 0813 (1.5T) and Flash 0731 beat Nemotron3 Ultra on agentic tasks per SemiAnalysis, with Flash using 4.2x fewer active params. No benchmarks disclosed.
AI Finds First Rational Diophantine Septuple, Cracking 80-Year Conjecture
Epoch AI used AI-guided search to find the first rational Diophantine septuple, ending an 80-year conjecture by Erdős and Graham. The result, verified by formal proof, shows AI can solve long-stalled number theory problems.
Mirage Avatar X Claims Identity Preservation Breakthrough
Mirage's Avatar X, per tester @hasantoxr, is the first AI avatar to excel at both identity preservation and expression believability, setting a new standard.
AI Disproves 87-Year-Old Conjecture, Finds Counterexample Humans Missed
AI disproves 87-year-old math conjecture, finding a counterexample humans missed, per @rohanpaul_ai.
China's 14nm AI Chip Hits 520 TFLOPS Via Architecture, Not Shrink
China's 14nm AI chip claims 520 TFLOPS and 6.4TB/s bandwidth via software-defined and 3D near-memory architecture, bypassing advanced node restrictions.
Claude Fable 5 Solves String Theory Problem Stalled for Six Months
Claude Fable 5 solved a string theory problem stalled for six months. Professor Yuji Tachikawa says the model made a non-trivial observation and used SymPy for verification.
POSTECH 10+ Layer Chip Stack Hits 4× Density of 12-Hi HBM
POSTECH developed 10+ layer chip stacking with 4× HBM density, targeting AI inference memory bottlenecks.
AI data centers could add 1.4°C to global warming by 2060, paper finds
AI data centers could add 1.4°C to global warming by 2060, per a new arXiv preprint, assuming 30% annual compute growth. The paper highlights the need for policy intervention.
NVIDIA Blackwell Cuts DeepSeek V4 Token Costs 5x in One Month
NVIDIA claims Blackwell inference stack cut DeepSeek V4 token costs 5x in one month, per a newly published report shared by @rohanpaul_ai.
We Cut Embedding Storage Costs by ~90% — Replacing S3 with PostgreSQL
A team cut embedding storage costs by ~90% by migrating from S3 to PostgreSQL with pgvector, enabling efficient vector search and on-demand retrieval for RAG and recommender systems, with no performance loss.
Qwen 2.5 7B Expresses Near-Constant Confidence Whether It Is Right or Wrong, Study Finds
A June 2026 arXiv preprint from University of Minnesota researchers tested Qwen 2.5 7B on structured clinical prediction data and found its verbalized confidence scores are essentially uninformative -- clustering between 0.856 and 0.937 no matter how well or badly the model performs. Combining SHAP-
OpenAI Says GPT-5.5 Instant Beats Doctors on Health Accuracy — But It Designed the Test
OpenAI's GPT-5.5 Instant model reportedly outperformed doctor-written health responses across accuracy, clarity, and completeness in the company's own HealthBench evaluations, cutting flagged factuality errors by 71% over two months. The catch: OpenAI built the benchmark, organized the physician pan
Midjourney Plans 60-Second Ultrasound Spa in SF by 2027
Midjourney plans a 2027 SF spa with 60-second ultrasound scans, aiming for 100x faster than MRI.
Tensordyne Claims 10x Efficiency Gain with Napier Architecture
Tensordyne claims 10x efficiency over Nvidia in inference with Napier gen, but lacks data or verification.
NVIDIA Blackwell Sweeps MLPerf Training 6.0, GB300 Hits 1.6x Speedup
NVIDIA Blackwell swept MLPerf Training 6.0 across all seven benchmarks. GB300 NVL72 delivered 1.6x speedup over GB200 NVL72 using NVFP4 and 8,192 GPUs.
AI editor matches pro on 84% of video cuts in blind test
AI editor matched pro on 84% of video cuts in blind test of 4-hour project. Suggests editorial judgment is partially automatable.
Wiwynn Shows First SCADA Server: 2.9PB, No CPU for I/O
Wiwynn showed first Nvidia SCADA server at Computex 2026: 2.9 PB storage, 528M IOPS, GPUs bypass CPU for I/O. Marks shift in AI storage architecture.
MIT Spinoff's Nuclear-Inspired Cooling Targets Data Center Water Use
MIT spinoff Infinite Cooling unveiled a nuclear-inspired cooling system that recycles data center heat and water, targeting 40% water use reduction. The tech faces competition from liquid cooling but offers retrofits for existing towers.
Anthropic's Fable 5 Beta Shows 10x Drug Design Speedup Ahead of IPO
Anthropic's Fable 5 beta achieved 10x speedup in protein design before being pulled from testing, signaling enterprise monetization ahead of IPO.
MiniMax-M3 Scores 55 on AI Index, Open-Source Lead Looms
MiniMax-M3 scored 55 on the Artificial Analysis Intelligence Index, set to become the leading open-source model once weights are released.