🤖 AI Summary
This work addresses the generalization bottlenecks in embodied intelligence concerning perception, spatial reasoning, and action alignment by proposing a unified spatiotemporal–physical grounding framework. It introduces a series of embodied foundation models spanning 2B to 122B parameters, featuring native 3D semantic alignment, contact-point prediction, and a cross-embodiment action space with embodiment-specific masking strategies to enable joint multi-task and multi-robot training. Evaluated on VSI-Bench, MMSI, and RefSpatial-Bench, the models consistently outperform existing open- and closed-source approaches. Real-robot experiments demonstrate that the proposed initialization strategy significantly surpasses Qwen baselines and mainstream general-purpose vision-language-action (VLA) models, with multi-task joint training yielding substantial gains in both task success rates and procedural scores.
📝 Abstract
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.