A new survey and benchmark rethinks interactive world models as game engines. The data engine collected over 90 hours of Black Myth: Wukong gameplay.
Key facts
- Data engine collected over 90 hours of Black Myth: Wukong gameplay.
- Benchmark analyzes four key dimensions of world models.
- Frames world models as specialized game engines.
- Source does not disclose specific benchmark dimensions or results.
A new survey and benchmark is reframing how the research community evaluates interactive world models — not as general-purpose simulators, but as game engines. The work, shared on X by @HuggingPapers, analyzes four key dimensions of world model performance, though the specific dimensions are not detailed in the source. The centerpiece is a new data engine that collected over 90 hours of gameplay from Black Myth: Wukong, the action RPG from Game Science that became a global hit in 2024.
The Game Engine Framing

The core shift is treating world models as specialized game engines rather than universal physics simulators. This mirrors how game developers build custom engines for specific titles — trading generality for fidelity. By focusing on a single game with rich visual and interactive complexity (Black Myth: Wukong), the benchmark can measure how well models predict state transitions, render novel views, and maintain temporal coherence in a constrained but challenging domain.
Data Engine and Benchmark Dimensions
The 90+ hours of gameplay collected by the data engine likely captures diverse scenarios: combat, exploration, boss fights, environmental interactions. The benchmark evaluates four key dimensions, though the source does not name them. According to @HuggingPapers, the work is a comprehensive survey and benchmark, suggesting it also catalogs existing approaches to interactive world models and positions them relative to the new game-engine framing.
Implications for World Model Research

Most world model benchmarks (e.g., DMControl, Atari, Habitat) treat environments as generic testbeds. By contrast, this work argues that high-fidelity prediction requires domain-specific priors — exactly what game engines provide. The Black Myth: Wukong data set is visually dense and computationally demanding, which may stress-test model capacity in ways simpler environments cannot. The survey component likely contextualizes this against prior work like Dreamer (Hafner et al. 2020), IRIS (Micheli et al. 2022), and DayDreamer (Wu et al. 2022).
What to watch
Watch for the full paper release with named benchmark dimensions and baseline results. If the benchmark reveals that game-engine-style world models outperform general-purpose simulators by >20% on in-distribution prediction, it could shift research toward domain-specific world models.









