RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
Summary
RynnBrain 1.1 is a family of embodied foundation models (2B, 9B, 122B-A10B) that improve perception, spatial reasoning, and manipulation, achieving state-of-the-art results on VSI-Bench, MMSI, and RefSpatial-Bench, and outperforming baselines in real-robot experiments.
View Cached Full Text
Cached at: 07/21/26, 06:35 AM
Paper page - RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
Source: https://huggingface.co/papers/2607.17977 Published on Jul 20
#3 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
WepresentRynnBrain1.1,afamilyofembodiedfoundationmodelsspanning2B,9B,and122B-A10Bscales.Trainedwithaunifiedspatio-temporalandphysicallygroundedframework,RynnBrain1.1supportsembodiedperception,spatialreasoning,localization,andplanning.ComparedwithRynnBrain1.0,itfurtherintroducescontact-pointpredictionacrossthemodelfamilyandnative3Dgroundingforthe2Band9Bmodels,yieldingrepresentationsandoutputsthataremoredirectlyalignedwithrobotmanipulation.WealsodevelopRynnBrain-VLAwithaunifiedcross-embodimentactionspaceandembodiment-specificmasking,anddeployitonUnitreeG1,Astribot-S1,andTianji-Wuji.RynnBrain1.1achievesstrongresultsonembodiedcognition,localization,and3Dgrounding,withthe122B-A10Bmodeloutperformingallevaluatedproprietaryandopen-sourcemodelsonVSI-Bench,MMSI,andRefSpatial-Bench.Real-robotexperimentsshowthatRynnBrain-initializedpoliciesoutperformQwen-basedandrepresentativegeneralistVLAs,whilejointmulti-taskandmulti-embodimenttrainingimprovesprocessscoresandsuccessratesoverper-tasktraining.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.17977
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.17977 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.17977 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.17977 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Embodied-R1.5 is a unified embodied foundation model that achieves state-of-the-art performance on 16 out of 24 embodied vision-language benchmarks using multi-task balanced reinforcement learning. It introduces a Planner-Grounder-Corrector closed-loop framework for long-horizon tasks and is open-sourced to facilitate future research.
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
RxBrain is an embodied cognition foundation model that jointly reasons with language and visual imagination to represent embodied plans, using a unified multimodal Mixture-of-Transformers architecture. It achieves promising real-robot performance without large-scale action data.
tencent/Hy-Embodied-RxBrain-1.0 · Hugging Face
Tencent releases Hy-Embodied-RxBrain-1.0, a unified multimodal foundation model for embodied cognition that combines language reasoning with visual imagination for understanding, world state prediction, and subgoal planning.
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
GigaBrain-0.7 is an embodied foundation model that uses a three-system architecture and large-scale heterogeneous pretraining to improve generalization across diverse robot embodiments, achieving enhanced zero-shot capabilities and task success rates.
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
RynnWorld-4D is a generative world model that co-produces future RGB, depth, and optical flow from a single RGB-D image and language instruction using a unified diffusion process, enabling efficient robotic manipulation through inverse dynamics policy learning. It achieves state-of-the-art on real-world bimanual manipulation tasks.