ev
30 articles about ev in AI news
vLLM's Force-Merged eval() Bug Shows LLM Host Takeover Risk
CVE-2025-9141 showed vLLM's eval() parser bug enabled LLM host takeover. The lead maintainer force-merged the vulnerable PR despite Gemini's critical warning, revealing systemic security gaps.
The Evaluation Stack: Why Metrics That Predict Production Quality Are
Towards AI's piece argues metrics that seem to improve can degrade production quality. It introduces the evaluation stack, a framework for metrics that predict real-world LLM performance, crucial for AI teams.
Claude Code vs. eve: Ora's Live-Site Benchmark Shows 7% Fewer Steps
Ora's benchmark shows Vercel's eve beats Claude Code with 7% fewer steps and 2x native success on live sites. Claude Code users should evaluate eve for web-integration tasks.
Papers with Code Adds Chat with PDF on Every Paper Page
Hugging Face's Papers with Code adds native PDF chat to every paper page. Niels Rogge announced the feature on X, positioning it to displace standalone PDF chat tools.
HarnessEval-W: New Benchmark Audits World Models via Sub-Agents
HarnessEval-W applies harness paradigm to world model eval, using sub-agents for auditable scoring. Announced via @HuggingPapers; technical details pending.
Tencent's EVIE Hits ViDoRe SOTA with 128D Embeddings
Tencent's EVIE model achieves SOTA on ViDoRe with 128D embeddings, a major efficiency gain. The release lacks technical details, but the compact vector size signals a shift toward cost-efficient retrieval.
MobileMem: On-Device Memory From a Year of Phone Data
MobileMem trains an LLM on a year of mobile data for on-device memory. Paper and code released, but no benchmarks disclosed.
NIQ reports 34% AI-native revenue growth as agentic commerce product nears
NIQ reported 34% AI-native revenue growth as its agentic commerce product nears launch. The move signals agentic AI is maturing in retail data, with NIQ competing against Google and Alipay in this emerging category.
Tara Lipinski Revives Every Bow at Devil Wears Prada 2 Premiere
Tara Lipinski wore a bow-covered look to the Devil Wears Prada 2 premiere in New York, per WWD, reviving the trend on the red carpet.
Study: Rhetoric Shifts AI Review Scores Across 4,200 Papers
Study across 4,200 manuscripts shows rhetoric alone shifts AI review scores. Maps 6 dimensions and 42K+ reviews to identify exploitable stylistic levers.
KAT-Coder-V2.5-Dev: The Open Agentic Coding Model That Could Rival Claude Code
KAT-Coder-V2.5-Dev lets you run agentic coding locally. Pair it with Claude Code for hybrid workflows: use it for private code, keep Claude for complex reasoning.
Musk: xAI to Hit 10GW Compute, $300-500B Revenue by 2027
Musk told SpaceX staff xAI will hit 10GW by 2027, projecting $300-500B revenue. The 7x expansion faces unprecedented supply chain and monetization challenges.
Nvidia to Invest Up to $3B in Stargate Power Developer Lancium
Nvidia to invest up to $3B in Lancium, Stargate's power developer, per The Information. Deal signals Nvidia's move into energy infrastructure.
Shepherd: Stanford's Git-Like Reversible AI Agent Runs
Stanford's Shepherd makes agent runs reversible via Git-like commits and syscall-level permissions. Copy-on-write forks are 5x faster than docker commit with 95% KV cache reuse.
Havaianas Debuts First-Ever Heeled Silhouette at Copenhagen Fashion Week
Havaianas debuted its first-ever heeled silhouette at Copenhagen Fashion Week, marking a significant expansion beyond its classic flip-flop lineup. Pricing, release date, and materials remain undisclosed.
OpenAI's First Device: Doughnut-Shaped Speaker, No Display
OpenAI's first device is a doughnut-shaped, display-less smart speaker, per Bloomberg. It's battery-powered with moving parts for interaction, targeting voice-first AI computing.
AMD 2Q26: Revenue Hits Record $11.54B, DC Segment Doubles
AMD reported record 2Q26 revenue of $11.54B, up 50% y/y with 56% gross margin, as data center segment revenue is set to more than double year-over-year.
Bobby Hundreds Revives '90s Fantasia Tee for Disney Debut
Bobby Hundreds revives a '90s Harajuku-found Fantasia tee for his debut Disney collaboration, per @hypebeast. The release marks his first official Disney partnership, leveraging vintage sourcing.
Model Routing Cut This Dev's Claude API Spend 35%
A developer cut Claude API spend 35% by routing tasks across Opus, Sonnet, and Haiku. Quality on hard tasks rose because Opus stopped handling busywork.
Opus 5 Backlash Signals Anthropic's Souring Developer Mood
Opus 5 launch draws harsh developer criticism per X post, signaling Anthropic's eroding technical brand amid competitive pressure.
SemiAnalysis Tests Qwen3.8-Max-Preview, 2.4T Params
SemiAnalysis tested Qwen3.8-Max-Preview, a 2.4T-param model, per a tweet. No results disclosed, but independent eval is notable.
Satisfy and Levi's Return With Second Climbing Gear Collab
Satisfy and Levi's released their second climbing gear collab, per @highsnobiety. The repeat drop signals the first capsule's success, though details remain sparse.
Ye Previews New YZY Footwear Wave on X Post
Ye previewed new YZY footwear on X, per Hypebeast, with no pricing or dates. Signals possible YZY SZN revival after quiet period.
BEAMS 50th Anniversary: First-Ever Teatora Patterned Pieces
BEAMS partners with Teatora for 50th anniversary, releasing first-ever patterned pieces from the workwear label. Details on pricing and availability remain undisclosed.
Supabase's Evals Benchmark Just Gave Claude Code a Real-World Report Card
Supabase Evals is an open-source benchmark that scores Claude Code, Codex, and OpenCode on real Supabase tasks. Run `supabase eval` on your repo to find agent weaknesses.
Nike and Zellerfeld Reveal 3D-Printed AirWorks Concept by Motoi Hatsuki
Nike and Zellerfeld revealed a 3D-printed AirWorks concept by Motoi Hatsuki. The fully printed sneaker signals Nike's push into on-demand additive manufacturing, though no release plans were disclosed.
Travis Scott's Nike Odyssey Sneakers: Exclusive Cast Gifts Revealed
Travis Scott revealed exclusive Nike sneakers for 'The Odyssey' cast, per WWD. The non-retail drop signals a gifting strategy over commercial sales.
METR's 'Expenditure Horizon': AI Agents Break Even at $3,300
METR's expenditure horizon metric shows AI agents break even at $0–$3,300 on NanoGPT, vs $2,500 per 1% speedup for humans. GPT-5 and Opus-4.1 pro lead, but blind spots remain.
NIQ Report: AI Personalization Boosts Retail Revenue 10-30%—Here’s How
NIQ reports AI in personalized shopping boosts retail revenue 10-30% by transforming product discovery via predictive analytics. This matters as retailers seek competitive edge through customer experience.
AWS Unveils Production Blueprint for Evaluating AI Agents with Strands and
AWS released Strands and AgentCore, a production blueprint for evaluating AI agents. It generates realistic scenarios and tracks metrics like completion rate and cost, addressing the gap between lab benchmarks and real-world performance—critical for retail AI deployments.