RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Summary
RxBrain is an embodied cognition foundation model that jointly reasons with language and visual imagination to represent embodied plans, using a unified multimodal Mixture-of-Transformers architecture. It achieves promising real-robot performance without large-scale action data.
View Cached Full Text
Cached at: 07/20/26, 09:41 AM
Paper page - RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Source: https://huggingface.co/papers/2607.14187 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Embodiedcognitionrequiresagentstoconnecthigh-leveltaskreasoningwiththephysicalstatestobeachieved.WeintroduceHy-Embodied-RxBrain,anembodiedcognitionfoundationmodelwithjointlanguage-visualreasoningandimagination.Unlikevision-languagemodelsthatemphasizesceneunderstandingandtextualdecisionmaking,orgenerativeworldmodelsthatmainlypredictfuturevisualstates,RxBrainrepresentsembodiedplansinasingleplanningsequencewherelanguageandvisualimaginationplaycomplementaryroles.Languageprovidestheabstractstructureofaplan,includingtaskdecomposition,planningprimitives,constraints,temporalorder,anddecisionlogic,whilevisualimaginationgroundsthisstructurethroughworldstatepredictionandjointsubgoalplanning,associatingeachplanningstepwithintermediateandfinalphysicalstates.RxBrainadoptsaunifiedmultimodalMixture-of-Transformersarchitecturethatsupportslanguage,image,andvideounderstandingandgenerationwithinonemodel.Totrainthiscapability,webuildanautomaticpipelinethatconvertsembodiedvideosintojointtext-visualplanningsupervisionbydecomposingvideosintoplanningstepsandaligningthemwithvisualstatetransitions.WefurtherintroduceRxBrain-Benchtoevaluatewhethermodelscanrepresentembodiedplansthroughjointtextualandvisualcomponentsratherthanseparateunderstandingorgeneration.ExperimentsshowthatRxBrainmaintainsembodiedunderstandingandgenerationabilities,andproducesplanswithcoupledtextualreasoning,worldstateprediction,andjointsubgoalplanning.WealsoextendRxBraintocontinuousrobotactiongeneration,whereitshowspromisingreal-robotperformancewithoutlarge-scaleaction-datapretraining.Theseresultsprovideaninitialsteptowardfoundationmodelsforembodiedcognition.
View arXiv pageView PDFProject pageGitHub90Add to collection
Models citing this paper1
#### tencent/Hy-Embodied-RxBrain-1.0 Any-to-Any• 6B• Updated2 days ago • 174 • 42
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.14187 in a dataset README.md to link it from this page.
Spaces citing this paper2
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
tencent/Hy-Embodied-RxBrain-1.0 · Hugging Face
Tencent releases Hy-Embodied-RxBrain-1.0, a unified multimodal foundation model for embodied cognition that combines language reasoning with visual imagination for understanding, world state prediction, and subgoal planning.
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
RynnBrain 1.1 is a family of embodied foundation models (2B, 9B, 122B-A10B) that improve perception, spatial reasoning, and manipulation, achieving state-of-the-art results on VSI-Bench, MMSI, and RefSpatial-Bench, and outperforming baselines in real-robot experiments.
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Embodied-R1.5 is a unified embodied foundation model that achieves state-of-the-art performance on 16 out of 24 embodied vision-language benchmarks using multi-task balanced reinforcement learning. It introduces a Planner-Grounder-Corrector closed-loop framework for long-horizon tasks and is open-sourced to facilitate future research.
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
BrainJanus is the first unified brain model that integrates brain, vision, and language via a shared Omni space, enabling bidirectional mapping between neural activity and sensory stimuli through tokenized representation and autoregressive next-token prediction.
PhysBrain 1.0 Technical Report
PhysBrain 1.0 is a technical report presenting a method that uses human egocentric video to generate physical commonsense supervision for vision-language-action models, achieving state-of-the-art results on embodied control benchmarks including ERQA, PhysBench, SimplerEnv-WidowX, LIBERO, and RoboCasa.