rl
30 articles about rl in AI news
ClawGym II Boosts Agent RL Pass@1 by 10-15 Points
ClawGym II claims 10-15 Pass@1 point gains on ClawGym-Bench via black-box RL with Qwen3-30A3B, targeting complex agent harnesses. Details on baselines and compute remain undisclosed.
HarnessEval-W: New Benchmark Audits World Models via Sub-Agents
HarnessEval-W applies harness paradigm to world model eval, using sub-agents for auditable scoring. Announced via @HuggingPapers; technical details pending.
VibeWorlding: 2,616 Assets, 6,828 Queries for 3D Agent Training
VibeWorlding, announced by @HuggingPapers, offers a 3D world-building benchmark with 2,616 assets and 6,828 queries, plus an RL gym. No baselines or paper details disclosed yet.
Tencent's UI-Mate-27B Hits 77.0 OSWorld, Learns From One Demo
Tencent's UI-Mate-27B scores 77.0 OSWorld-Verified, 66.2 WindowsAgentArena, learning GUI tasks from one demo. Released on Hugging Face.
Tencent's WorldClaw Generates Explorable 3D Worlds from Text
Tencent's WorldClaw generates explorable 3D worlds from text via agentic planning. No benchmarks or release date disclosed, but the agentic decomposition approach marks a structural shift.
Timberland, New Era Drop MLB Caps & Boots Aug 15
Timberland and New Era drop a limited MLB collection Aug 15, covering seven teams with caps and boots. Pricing and quantities remain undisclosed.
Ferran Torres on World Cup Glory, Fashion Ambitions, and the Red Hat That Divided Fans
Ferran Torres discussed World Cup-winning goal and fashion ambitions in WWD. Red hat controversy addressed; clothing line planned for 2027.
OrcaMan Open-Sources BoundaryBench for Enterprise RL
OrcaMan open-sourced BoundaryBench, a benchmark for enterprise RL boundary detection. The tool targets reliability gaps in production AI.
InAgent Hits 90.2% on OSWorld, First Agent Past 90%
InAgent scored 90.2% on OSWorld, first above 90%, with 100% on system-level tasks, surpassing OpenAI, Google, and Anthropic records. Harness engineering, not raw model power, drove the result.
RLSVR Turns Open-Ended Tasks Into Spy Game for Verifiable Rewards
RLSVR uses a 'spy' game to generate verifiable rewards for open-ended LLM tasks, removing the judge model. Announced via tweet, no benchmark data disclosed.
Intel EMIB-T Wins Second Cloud ASIC, Early 2028 Production
Intel EMIB-T wins second cloud ASIC project; MediaTek CEO confirms early 2028 production on track despite packaging yield and capacity challenges.
Charli XCX's 2026 Top Shoe Moments: Fashion Week to Tour
WWD highlighted Charli XCX's top 2026 shoe moments across red carpets and tours, reinforcing her fashion influence. Specific brands were not disclosed.
BYD HyWorldVLA Hits 90.59 PDMS on NAVSIM v1
BYD's HyWorldVLA achieved 90.59 PDMS on NAVSIM v1, a new SOTA, using a hybrid pixel-latent world model. It marks BYD's entry into autonomous driving foundation models.
NVIDIA's Molt: 9.2K-Line RL Framework Scales to 1T-Parameter MoE Models
NVIDIA released Molt, a 9.2K-line PyTorch RL framework scaling to 1T-parameter MoE models via vLLM, targeting agentic tasks with fully-async rollout.
AlayaRenderer-Flash Hits 31.54 FPS, Enables Playable Generative Worlds
AlayaRenderer-Flash accelerates real-time scene synthesis from 0.56 to 31.54 FPS, a 56x speedup enabling playable generative worlds.
Under Armour Drops Solid Gold Rings for Spain World Cup Winners
Under Armour collaborates with Hoorsenbuhs and 424 on solid gold rings for Spain's World Cup winners. Luxury sports memorabilia without disclosed pricing.
GigaWorld-Policy-0.5 Hits 85ms on RTX 4090 for Robot Control
GigaWorld-Policy-0.5 runs robot control at 85ms on an RTX 4090, using a Mixture-of-Transformers architecture for real-time local deployment.
90 Hours of Black Myth: Wukong Fuel New World Model Benchmark
A new survey and benchmark rethinks interactive world models as game engines, with a data engine collecting over 90 hours of Black Myth: Wukong gameplay.
HG-RAG Beats Flat Retrieval on Graph Queries Across 800-Node Worlds
HG-RAG uses graph-traversal over knowledge graphs for RAG, beating flat retrieval on hierarchical and multi-hop queries across worlds up to 800 nodes.
Beyoncé Wears $895 Jimmy Choo x Timberland Boots by Designer Caroline Hu
Designer Caroline Hu co-created Jimmy Choo x Timberland boots Beyoncé wore, retailing at $895. Hu discovered the moment on social media, feeling shocked.
BRAID Fuses Text-Image Reasoning Into One RL Objective
BRAID unifies multi-turn text-image reasoning as a Markov decision process, enabling joint RL optimization of both modalities with a single objective.
BAAI Orca World Model Matches π0.5 With No Action Labels
BAAI's Orca world model matches specialized π0.5 on five robotics tasks, trained on 125,000 hours of video without action labels, predicting abstract world states.
Crusoe Launches Serverless Fine-Tuning, Targets AI Lifecycle Beyond GPUs
Crusoe launched serverless fine-tuning and inference, targeting enterprise AI teams. IDC says GPU access is no longer the differentiator; portability is now a procurement requirement.
LLM agents fail nonlinearly as tasks lengthen, 27-paper synthesis finds
27-paper synthesis finds LLM agent failures compound nonlinearly with task length. Six failure clusters identified across 19 benchmarks.
Free RL Textbook 'Math Foundations' Hits 16.2K GitHub Stars
Free RL textbook by Shiyu Zhao hits 16.2K GitHub stars and 2.1M video views, filling a gap in RL education with rigorous math and a unified grid-world example.
Alibaba Open-Sources Qwen-AgentWorld for Generalist Agent Training
Alibaba open-sourced Qwen-AgentWorld and Wan-Streamer v0.1 on Hugging Face, targeting generalist agent training and real-time streaming. The releases include 8 additional papers on agent benchmarks and architectures.
OSWorld 2.0 Launches, Tests AI Agents on 1,500 Desktop Tasks
Epoch AI released OSWorld 2.0 with 1,500 desktop tasks, up from 369 in v1, testing AI agents on adversarial and cross-application workflows.
World Action Models Survey Unifies 100+ Methods Under One Taxonomy
A survey reviews 100+ world action models, unifying world models, video generation, and VLA policies under one taxonomy.
Gemini 3.5 Flash Scores 78.4 on OSWorld, Matching GPT-5.5
Google integrated Computer Use into Gemini 3.5 Flash, scoring 78.4 on OSWorld — matching GPT-5.5 and undercutting on cost.
World Model MCP: Memory Layer That Cut SWE-bench Repeat Mistakes by +10.2 Points
World Model MCP adds a temporal knowledge graph to Claude Code that learns from corrections, prevents repeated mistakes, and re-injects context after compaction — proven with +10.2 pts on SWE-bench.