Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

ev

30 articles about ev in AI news

vLLM's Force-Merged eval() Bug Shows LLM Host Takeover Risk

CVE-2025-9141 showed vLLM's eval() parser bug enabled LLM host takeover. The lead maintainer force-merged the vulnerable PR despite Gemini's critical warning, revealing systemic security gaps.

93% relevant

The Evaluation Stack: Why Metrics That Predict Production Quality Are

Towards AI's piece argues metrics that seem to improve can degrade production quality. It introduces the evaluation stack, a framework for metrics that predict real-world LLM performance, crucial for AI teams.

75% relevant

Claude Code vs. eve: Ora's Live-Site Benchmark Shows 7% Fewer Steps

Ora's benchmark shows Vercel's eve beats Claude Code with 7% fewer steps and 2x native success on live sites. Claude Code users should evaluate eve for web-integration tasks.

100% relevant

Papers with Code Adds Chat with PDF on Every Paper Page

Hugging Face's Papers with Code adds native PDF chat to every paper page. Niels Rogge announced the feature on X, positioning it to displace standalone PDF chat tools.

80% relevant

HarnessEval-W: New Benchmark Audits World Models via Sub-Agents

HarnessEval-W applies harness paradigm to world model eval, using sub-agents for auditable scoring. Announced via @HuggingPapers; technical details pending.

78% relevant

Tencent's EVIE Hits ViDoRe SOTA with 128D Embeddings

Tencent's EVIE model achieves SOTA on ViDoRe with 128D embeddings, a major efficiency gain. The release lacks technical details, but the compact vector size signals a shift toward cost-efficient retrieval.

85% relevant

MobileMem: On-Device Memory From a Year of Phone Data

MobileMem trains an LLM on a year of mobile data for on-device memory. Paper and code released, but no benchmarks disclosed.

85% relevant

NIQ reports 34% AI-native revenue growth as agentic commerce product nears

NIQ reported 34% AI-native revenue growth as its agentic commerce product nears launch. The move signals agentic AI is maturing in retail data, with NIQ competing against Google and Alipay in this emerging category.

80% relevant

Tara Lipinski Revives Every Bow at Devil Wears Prada 2 Premiere

Tara Lipinski wore a bow-covered look to the Devil Wears Prada 2 premiere in New York, per WWD, reviving the trend on the red carpet.

75% relevant

Study: Rhetoric Shifts AI Review Scores Across 4,200 Papers

Study across 4,200 manuscripts shows rhetoric alone shifts AI review scores. Maps 6 dimensions and 42K+ reviews to identify exploitable stylistic levers.

80% relevant

KAT-Coder-V2.5-Dev: The Open Agentic Coding Model That Could Rival Claude Code

KAT-Coder-V2.5-Dev lets you run agentic coding locally. Pair it with Claude Code for hybrid workflows: use it for private code, keep Claude for complex reasoning.

90% relevant

Musk: xAI to Hit 10GW Compute, $300-500B Revenue by 2027

Musk told SpaceX staff xAI will hit 10GW by 2027, projecting $300-500B revenue. The 7x expansion faces unprecedented supply chain and monetization challenges.

100% relevant

Nvidia to Invest Up to $3B in Stargate Power Developer Lancium

Nvidia to invest up to $3B in Lancium, Stargate's power developer, per The Information. Deal signals Nvidia's move into energy infrastructure.

85% relevant

Shepherd: Stanford's Git-Like Reversible AI Agent Runs

Stanford's Shepherd makes agent runs reversible via Git-like commits and syscall-level permissions. Copy-on-write forks are 5x faster than docker commit with 95% KV cache reuse.

85% relevant

Havaianas Debuts First-Ever Heeled Silhouette at Copenhagen Fashion Week

Havaianas debuted its first-ever heeled silhouette at Copenhagen Fashion Week, marking a significant expansion beyond its classic flip-flop lineup. Pricing, release date, and materials remain undisclosed.

78% relevant

OpenAI's First Device: Doughnut-Shaped Speaker, No Display

OpenAI's first device is a doughnut-shaped, display-less smart speaker, per Bloomberg. It's battery-powered with moving parts for interaction, targeting voice-first AI computing.

87% relevant

AMD 2Q26: Revenue Hits Record $11.54B, DC Segment Doubles

AMD reported record 2Q26 revenue of $11.54B, up 50% y/y with 56% gross margin, as data center segment revenue is set to more than double year-over-year.

82% relevant

Bobby Hundreds Revives '90s Fantasia Tee for Disney Debut

Bobby Hundreds revives a '90s Harajuku-found Fantasia tee for his debut Disney collaboration, per @hypebeast. The release marks his first official Disney partnership, leveraging vintage sourcing.

78% relevant

Model Routing Cut This Dev's Claude API Spend 35%

A developer cut Claude API spend 35% by routing tasks across Opus, Sonnet, and Haiku. Quality on hard tasks rose because Opus stopped handling busywork.

88% relevant

Opus 5 Backlash Signals Anthropic's Souring Developer Mood

Opus 5 launch draws harsh developer criticism per X post, signaling Anthropic's eroding technical brand amid competitive pressure.

80% relevant

SemiAnalysis Tests Qwen3.8-Max-Preview, 2.4T Params

SemiAnalysis tested Qwen3.8-Max-Preview, a 2.4T-param model, per a tweet. No results disclosed, but independent eval is notable.

87% relevant

Satisfy and Levi's Return With Second Climbing Gear Collab

Satisfy and Levi's released their second climbing gear collab, per @highsnobiety. The repeat drop signals the first capsule's success, though details remain sparse.

78% relevant

Ye Previews New YZY Footwear Wave on X Post

Ye previewed new YZY footwear on X, per Hypebeast, with no pricing or dates. Signals possible YZY SZN revival after quiet period.

75% relevant

BEAMS 50th Anniversary: First-Ever Teatora Patterned Pieces

BEAMS partners with Teatora for 50th anniversary, releasing first-ever patterned pieces from the workwear label. Details on pricing and availability remain undisclosed.

85% relevant

Supabase's Evals Benchmark Just Gave Claude Code a Real-World Report Card

Supabase Evals is an open-source benchmark that scores Claude Code, Codex, and OpenCode on real Supabase tasks. Run `supabase eval` on your repo to find agent weaknesses.

100% relevant

Nike and Zellerfeld Reveal 3D-Printed AirWorks Concept by Motoi Hatsuki

Nike and Zellerfeld revealed a 3D-printed AirWorks concept by Motoi Hatsuki. The fully printed sneaker signals Nike's push into on-demand additive manufacturing, though no release plans were disclosed.

75% relevant

Travis Scott's Nike Odyssey Sneakers: Exclusive Cast Gifts Revealed

Travis Scott revealed exclusive Nike sneakers for 'The Odyssey' cast, per WWD. The non-retail drop signals a gifting strategy over commercial sales.

78% relevant

METR's 'Expenditure Horizon': AI Agents Break Even at $3,300

METR's expenditure horizon metric shows AI agents break even at $0–$3,300 on NanoGPT, vs $2,500 per 1% speedup for humans. GPT-5 and Opus-4.1 pro lead, but blind spots remain.

90% relevant

NIQ Report: AI Personalization Boosts Retail Revenue 10-30%—Here’s How

NIQ reports AI in personalized shopping boosts retail revenue 10-30% by transforming product discovery via predictive analytics. This matters as retailers seek competitive edge through customer experience.

98% relevant

AWS Unveils Production Blueprint for Evaluating AI Agents with Strands and

AWS released Strands and AgentCore, a production blueprint for evaluating AI agents. It generates realistic scenarios and tracks metrics like completion rate and cost, addressing the gap between lab benchmarks and real-world performance—critical for retail AI deployments.

88% relevant