Dyna-2, a world-action model trained on over 1M hours of egocentric video, reportedly reveals new scaling laws. The model jointly predicts future video and actions, reasoning about outcomes before acting.
Key facts
- Training corpus: over 1M hours of egocentric video
- Jointly predicts future video and future actions
- No parameter count or compute disclosed
- No benchmark results published yet
- Claim of new scaling laws unverified
The claim, posted by @omarsar0, describes Dyna-2 as a world-action model that unifies visual prediction with action selection. According to @omarsar0, the training corpus exceeds 1 million hours of egocentric human video, a scale that would dwarf prior embodied datasets like Ego4D's ~3,000 hours of daily-life activity. The joint prediction of future video and future actions positions the model within the world-model lineage that traces back to Ha and Schmidhuber's 2018 World Models paper, but with a critical difference: action prediction is folded into the objective rather than treated as a downstream policy head.
What the scaling-law claim implies
The phrase "new scaling laws" is the most consequential part of the announcement. Language-model scaling laws, formalized by Kaplan et al. 2020 and refined by Hoffmann et al. 2022, describe compute-optimal training regimes for next-token prediction. If Dyna-2's joint video-plus-action objective yields different exponents, it would suggest that embodied prediction scales differently from passive text prediction. The claim of new scaling laws, if substantiated, would extend the power-law framework beyond language tokens into embodied action spaces. However, the source provides no exponent values, no loss curves, and no benchmark comparisons. The company did not disclose the figure for parameter count or training compute.
What remains unverified
No evaluation results accompany the announcement. The tweet does not specify whether Dyna-2 outperforms existing video prediction models like VideoPoet or action-conditioned world models on standard benchmarks. The absence of a technical report, a paper, or a model release makes it impossible to verify the scaling-law claim. Given the pattern of unverified world-model announcements in recent months, the burden of proof rests on the authors to publish loss-vs-compute curves and downstream task results. Watch for the release of the technical report, the model's parameter count, and whether Dyna-2's scaling exponents differ from those of language models.
Key Takeaways
- Dyna-2, trained on 1M+ hours of egocentric video, jointly predicts future video and actions.
- Claims new scaling laws but no benchmarks or technical details released.
What to watch
Watch for the release of the Dyna-2 technical report or arXiv paper, which should include parameter counts, compute budgets, and loss-vs-compute scaling curves. The key update would be whether the reported scaling exponents differ from Kaplan et al. 2020 and Hoffmann et al. 2022. Also track whether any benchmark results on embodied tasks like manipulation or navigation are published, and whether the model is open-sourced.








