Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

rl

30 articles about rl in AI news

ClawGym II Boosts Agent RL Pass@1 by 10-15 Points

ClawGym II claims 10-15 Pass@1 point gains on ClawGym-Bench via black-box RL with Qwen3-30A3B, targeting complex agent harnesses. Details on baselines and compute remain undisclosed.

85% relevant

HarnessEval-W: New Benchmark Audits World Models via Sub-Agents

HarnessEval-W applies harness paradigm to world model eval, using sub-agents for auditable scoring. Announced via @HuggingPapers; technical details pending.

78% relevant

VibeWorlding: 2,616 Assets, 6,828 Queries for 3D Agent Training

VibeWorlding, announced by @HuggingPapers, offers a 3D world-building benchmark with 2,616 assets and 6,828 queries, plus an RL gym. No baselines or paper details disclosed yet.

85% relevant

Tencent's UI-Mate-27B Hits 77.0 OSWorld, Learns From One Demo

Tencent's UI-Mate-27B scores 77.0 OSWorld-Verified, 66.2 WindowsAgentArena, learning GUI tasks from one demo. Released on Hugging Face.

94% relevant

Tencent's WorldClaw Generates Explorable 3D Worlds from Text

Tencent's WorldClaw generates explorable 3D worlds from text via agentic planning. No benchmarks or release date disclosed, but the agentic decomposition approach marks a structural shift.

85% relevant

Timberland, New Era Drop MLB Caps & Boots Aug 15

Timberland and New Era drop a limited MLB collection Aug 15, covering seven teams with caps and boots. Pricing and quantities remain undisclosed.

82% relevant

Ferran Torres on World Cup Glory, Fashion Ambitions, and the Red Hat That Divided Fans

Ferran Torres discussed World Cup-winning goal and fashion ambitions in WWD. Red hat controversy addressed; clothing line planned for 2027.

72% relevant

OrcaMan Open-Sources BoundaryBench for Enterprise RL

OrcaMan open-sourced BoundaryBench, a benchmark for enterprise RL boundary detection. The tool targets reliability gaps in production AI.

85% relevant

InAgent Hits 90.2% on OSWorld, First Agent Past 90%

InAgent scored 90.2% on OSWorld, first above 90%, with 100% on system-level tasks, surpassing OpenAI, Google, and Anthropic records. Harness engineering, not raw model power, drove the result.

100% relevant

RLSVR Turns Open-Ended Tasks Into Spy Game for Verifiable Rewards

RLSVR uses a 'spy' game to generate verifiable rewards for open-ended LLM tasks, removing the judge model. Announced via tweet, no benchmark data disclosed.

85% relevant

Intel EMIB-T Wins Second Cloud ASIC, Early 2028 Production

Intel EMIB-T wins second cloud ASIC project; MediaTek CEO confirms early 2028 production on track despite packaging yield and capacity challenges.

85% relevant

Charli XCX's 2026 Top Shoe Moments: Fashion Week to Tour

WWD highlighted Charli XCX's top 2026 shoe moments across red carpets and tours, reinforcing her fashion influence. Specific brands were not disclosed.

72% relevant

BYD HyWorldVLA Hits 90.59 PDMS on NAVSIM v1

BYD's HyWorldVLA achieved 90.59 PDMS on NAVSIM v1, a new SOTA, using a hybrid pixel-latent world model. It marks BYD's entry into autonomous driving foundation models.

100% relevant

NVIDIA's Molt: 9.2K-Line RL Framework Scales to 1T-Parameter MoE Models

NVIDIA released Molt, a 9.2K-line PyTorch RL framework scaling to 1T-parameter MoE models via vLLM, targeting agentic tasks with fully-async rollout.

91% relevant

AlayaRenderer-Flash Hits 31.54 FPS, Enables Playable Generative Worlds

AlayaRenderer-Flash accelerates real-time scene synthesis from 0.56 to 31.54 FPS, a 56x speedup enabling playable generative worlds.

85% relevant

Under Armour Drops Solid Gold Rings for Spain World Cup Winners

Under Armour collaborates with Hoorsenbuhs and 424 on solid gold rings for Spain's World Cup winners. Luxury sports memorabilia without disclosed pricing.

85% relevant

GigaWorld-Policy-0.5 Hits 85ms on RTX 4090 for Robot Control

GigaWorld-Policy-0.5 runs robot control at 85ms on an RTX 4090, using a Mixture-of-Transformers architecture for real-time local deployment.

85% relevant

90 Hours of Black Myth: Wukong Fuel New World Model Benchmark

A new survey and benchmark rethinks interactive world models as game engines, with a data engine collecting over 90 hours of Black Myth: Wukong gameplay.

78% relevant

HG-RAG Beats Flat Retrieval on Graph Queries Across 800-Node Worlds

HG-RAG uses graph-traversal over knowledge graphs for RAG, beating flat retrieval on hierarchical and multi-hop queries across worlds up to 800 nodes.

82% relevant

Beyoncé Wears $895 Jimmy Choo x Timberland Boots by Designer Caroline Hu

Designer Caroline Hu co-created Jimmy Choo x Timberland boots Beyoncé wore, retailing at $895. Hu discovered the moment on social media, feeling shocked.

75% relevant

BRAID Fuses Text-Image Reasoning Into One RL Objective

BRAID unifies multi-turn text-image reasoning as a Markov decision process, enabling joint RL optimization of both modalities with a single objective.

85% relevant

BAAI Orca World Model Matches π0.5 With No Action Labels

BAAI's Orca world model matches specialized π0.5 on five robotics tasks, trained on 125,000 hours of video without action labels, predicting abstract world states.

100% relevant

Crusoe Launches Serverless Fine-Tuning, Targets AI Lifecycle Beyond GPUs

Crusoe launched serverless fine-tuning and inference, targeting enterprise AI teams. IDC says GPU access is no longer the differentiator; portability is now a procurement requirement.

75% relevant

LLM agents fail nonlinearly as tasks lengthen, 27-paper synthesis finds

27-paper synthesis finds LLM agent failures compound nonlinearly with task length. Six failure clusters identified across 19 benchmarks.

90% relevant

Free RL Textbook 'Math Foundations' Hits 16.2K GitHub Stars

Free RL textbook by Shiyu Zhao hits 16.2K GitHub stars and 2.1M video views, filling a gap in RL education with rigorous math and a unified grid-world example.

83% relevant

Alibaba Open-Sources Qwen-AgentWorld for Generalist Agent Training

Alibaba open-sourced Qwen-AgentWorld and Wan-Streamer v0.1 on Hugging Face, targeting generalist agent training and real-time streaming. The releases include 8 additional papers on agent benchmarks and architectures.

82% relevant

OSWorld 2.0 Launches, Tests AI Agents on 1,500 Desktop Tasks

Epoch AI released OSWorld 2.0 with 1,500 desktop tasks, up from 369 in v1, testing AI agents on adversarial and cross-application workflows.

95% relevant

World Action Models Survey Unifies 100+ Methods Under One Taxonomy

A survey reviews 100+ world action models, unifying world models, video generation, and VLA policies under one taxonomy.

87% relevant

Gemini 3.5 Flash Scores 78.4 on OSWorld, Matching GPT-5.5

Google integrated Computer Use into Gemini 3.5 Flash, scoring 78.4 on OSWorld — matching GPT-5.5 and undercutting on cost.

100% relevant

World Model MCP: Memory Layer That Cut SWE-bench Repeat Mistakes by +10.2 Points

World Model MCP adds a temporal knowledge graph to Claude Code that learns from corrections, prevents repeated mistakes, and re-injects context after compaction — proven with +10.2 pts on SWE-bench.

95% relevant