InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
Summary
InternW0 is a foundational physical world model from Shanghai AI Laboratory that jointly learns visual dynamics and robot control for efficient real-world interactions, trained on heterogeneous data and evaluated on scientific tasks.
View Cached Full Text
Cached at: 09/24/26, 03:39 AM
Paper page - InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
Source: https://huggingface.co/papers/2609.27656 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Physicalintelligencerequiresmorethanpredictinghowtheworldmayevolve:predictionsmustremainactionableastheworldcontinuestochange.WeintroduceInternW0,thefirstinstantiationoftheInternWphysicalworldmodelseriesfromShanghaiAILaboratory,builtaroundomnimodalinterfaces,asynchronousmulti-frequencyprocessing,andlocalphysicalmodelingunderpartialobservationsandexternalinfluences.InternW0jointlylearnsfuturevisualdynamicsandcontinuousrobotcontrolthroughanasymmetricvideo--actionarchitecturewithflowmatching.Ahigh-capacityvideoexpertprovideslonger-horizonpredictivecontext,whilealightweightactionexpertoperatesatafastertimescale.Insteadofregeneratingthefutureforeveryactionupdate,InternW0reuseslayerwiseK/Vandadaptsittonewlyobservedstatesthroughobservation-conditionedcontextrouting.Domain-specificinterfacesandsoftpromptssupportheterogeneousembodiments,whilecontact-awarepost-trainingincorporatesforceandtactilesignalsforcontact-richmanipulation.WetrainInternW0onapproximately7,200hoursofheterogeneousrobotandegocentricdata,includingEgoLab,a275-hourreal-laboratoryegocentricdataset.Evaluationspanssimulationbenchmarksandreal-worldscientifictasks,includinga15-stagemetal--organicframeworksynthesisworkflowand5-stagecontact-andforce-awaredexterousmanipulationforgeneral-purposequantitativepipetting.Theseresultsadvancescalable,asynchronous,andscience-nativephysicalworldmodelsforuniversalandefficientreal-worldinteractions.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.27656
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.27656 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.27656 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.27656 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
In-Context World Modeling for Robotic Control
This paper introduces In-Context World Modeling (ICWM), a framework that enables robot policies to infer system variables from self-generated interactions, allowing adaptation to novel configurations without parameter updates by treating system identification as an in-context adaptation problem. It outperforms standard VLA baselines on novel camera viewpoints in simulation and real-world experiments.
PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation
PAIWorld enhances diffusion-transformer world models with geometric awareness and cross-view attention to improve multi-view 3D consistency for robotic manipulation tasks, achieving state-of-the-art results on benchmarks.
A Tutorial on World Models and Physical AI
This tutorial presents a coherent framework unifying diverse world modeling approaches for physical AI, covering explicit and implicit world models and their role in prediction, reasoning, and planning.
τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation
τ_0-WM is a unified video-action world model for robotic manipulation that integrates policy learning, video prediction, and action evaluation using a shared video diffusion backbone. It shows superior performance on challenging long-horizon and fine-grained tasks.
World in World: Explore the World with World Models
The paper presents World in World, a training-free interface that enables flexible camera and time control in frozen autoregressive video world models by using correspondence-guided queries and evidence-wise attention guidance.