Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Alibaba's RynnBrain 1.1 model card on Hugging Face with robot arm handling objects, showcasing embodied AI for…
AI ResearchScore: 85

Alibaba Releases RynnBrain 1.1 Embodied AI Models at 2B-122B Scales

Alibaba released RynnBrain 1.1 on Hugging Face with 2B, 9B, and 122B-A10B MoE models for robot manipulation, but disclosed no benchmarks.

·1d ago·2 min read··30 views·AI-Generated·Report error
Share:
What models did Alibaba release with RynnBrain 1.1?

Alibaba released RynnBrain 1.1 on Hugging Face, a family of embodied foundation models at 2B, 9B, and 122B-A10B scales for robot manipulation tasks including perception, spatial reasoning, and contact-point prediction.

TL;DR

Three model sizes: 2B, 9B, 122B-A10B. · Supports perception, spatial reasoning, planning. · Targets robot manipulation and contact-point prediction.

Alibaba released RynnBrain 1.1 on Hugging Face, a family of embodied foundation models at 2B, 9B, and 122B-A10B scales. The models target robot manipulation tasks spanning perception, spatial reasoning, and contact-point prediction.

Key facts

  • Three model sizes: 2B, 9B, 122B-A10B.
  • 122B-A10B is sparse MoE with 10B active parameters.
  • Supports 5 capabilities: perception, spatial reasoning, localization, planning, contact-point prediction.
  • Released on Hugging Face under Alibaba's account.
  • No benchmark results or training dataset disclosed.

Alibaba released RynnBrain 1.1 on Hugging Face, a family of embodied foundation models at 2B, 9B, and 122B-A10B scales. The models support perception, spatial reasoning, localization, planning, and contact-point prediction for robot manipulation According to @HuggingPapers.

The 122B-A10B variant—a sparse mixture-of-experts architecture with 122B total parameters and 10B active per token—extends Alibaba's push into large-scale embodied AI. By contrast, the 2B and 9B models target edge deployment on resource-constrained robots.

RynnBrain 1.1 covers the full pipeline from visual input to motor output: object localization, spatial reasoning about reachability, task planning, and fine-grained contact-point prediction for grasping. The company did not disclose training dataset size, compute budget, or benchmark results.

Unique Take

RynnBrain 1.1's multi-scale release mirrors the trend seen in language models (e.g., Llama 3.1 8B/70B/405B) but applied to embodied AI—a category where most open models remain single-scale. The 122B-A10B sparse MoE is particularly notable: it likely uses a router to activate task-specific experts for perception vs. planning, reducing inference cost while maintaining capacity. This architectural choice suggests Alibaba expects robots to run multiple specialized sub-tasks simultaneously.

The lack of benchmark numbers is conspicuous. Without standardized embodied benchmarks (e.g., RT-2's success rate on manipulation tasks or Habitat's navigation metrics), comparing RynnBrain 1.1 to Google's RT-2 or Meta's Habitat remains impossible. The open release on Hugging Face at least enables third-party evaluation.

What to Watch

Watch for third-party evaluations on standardized embodied benchmarks (e.g., ManiSkill2 or RLBench) within 60 days. If RynnBrain 1.1's 122B model outperforms RT-2 on manipulation tasks, it signals Alibaba's leadership in open embodied AI—and likely triggers a response from Google or Meta.

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from multiple verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

RynnBrain 1.1's multi-scale release mirrors the language-model trend of offering tiered models for different deployment constraints, but applied to embodied AI—a category where most open models remain single-scale. The 122B-A10B sparse MoE is architecturally ambitious: it likely uses a router to activate task-specific experts for perception vs. planning, reducing inference cost while maintaining capacity. This suggests Alibaba expects robots to run multiple specialized sub-tasks simultaneously, a design choice that could prove prescient if edge robots need to switch between object detection and motion planning. The lack of benchmark numbers is a material omission. Without standardized embodied benchmarks like ManiSkill2 or RLBench, claims of capability are unverifiable. Google's RT-2 reported success rates on 600+ manipulation tasks; Meta's Habitat evaluates navigation. Alibaba's silence on these metrics invites skepticism—either the numbers were poor, or the models aren't yet production-ready. The open release on Hugging Face at least enables independent evaluation, but the community will need to run its own tests. Compared to prior work, RynnBrain 1.1's 122B-A10B is larger than RT-2's 55B parameters but smaller than PaLM-E's 562B. The sparse MoE design is a practical middle ground: it avoids the full inference cost of dense models while maintaining capacity for diverse tasks. If the router is well-trained, this could become a template for future embodied models—but without data, it's just a hypothesis.
This story is part of
Hugging Face Becomes the Neutral Ground Where Google and Anthropic's Agent Protocol War Converges
As Claude Code's MCP dominance threatens Google Cloud, Hugging Face's unique position as partner to both players creates an unexpected convergence zone
Compare side-by-side
Alibaba vs Hugging Face

Mentioned in this article

Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all