Alibaba released RynnBrain 1.1 on Hugging Face, a family of embodied foundation models at 2B, 9B, and 122B-A10B scales. The models target robot manipulation tasks spanning perception, spatial reasoning, and contact-point prediction.
Key facts
- Three model sizes: 2B, 9B, 122B-A10B.
- 122B-A10B is sparse MoE with 10B active parameters.
- Supports 5 capabilities: perception, spatial reasoning, localization, planning, contact-point prediction.
- Released on Hugging Face under Alibaba's account.
- No benchmark results or training dataset disclosed.
Alibaba released RynnBrain 1.1 on Hugging Face, a family of embodied foundation models at 2B, 9B, and 122B-A10B scales. The models support perception, spatial reasoning, localization, planning, and contact-point prediction for robot manipulation According to @HuggingPapers.
The 122B-A10B variant—a sparse mixture-of-experts architecture with 122B total parameters and 10B active per token—extends Alibaba's push into large-scale embodied AI. By contrast, the 2B and 9B models target edge deployment on resource-constrained robots.
RynnBrain 1.1 covers the full pipeline from visual input to motor output: object localization, spatial reasoning about reachability, task planning, and fine-grained contact-point prediction for grasping. The company did not disclose training dataset size, compute budget, or benchmark results.
Unique Take
RynnBrain 1.1's multi-scale release mirrors the trend seen in language models (e.g., Llama 3.1 8B/70B/405B) but applied to embodied AI—a category where most open models remain single-scale. The 122B-A10B sparse MoE is particularly notable: it likely uses a router to activate task-specific experts for perception vs. planning, reducing inference cost while maintaining capacity. This architectural choice suggests Alibaba expects robots to run multiple specialized sub-tasks simultaneously.
The lack of benchmark numbers is conspicuous. Without standardized embodied benchmarks (e.g., RT-2's success rate on manipulation tasks or Habitat's navigation metrics), comparing RynnBrain 1.1 to Google's RT-2 or Meta's Habitat remains impossible. The open release on Hugging Face at least enables third-party evaluation.
What to Watch
Watch for third-party evaluations on standardized embodied benchmarks (e.g., ManiSkill2 or RLBench) within 60 days. If RynnBrain 1.1's 122B model outperforms RT-2 on manipulation tasks, it signals Alibaba's leadership in open embodied AI—and likely triggers a response from Google or Meta.








