Cosmos 3: Omnimodal World Models for Physical AI
Summary
Cosmos 3 is a family of omnimodal world models from NVIDIA that jointly processes language, image, video, audio, and action sequences using a unified mixture-of-transformers architecture, achieving state-of-the-art performance in understanding and generation tasks for Physical AI.
View Cached Full Text
Cached at: 06/04/26, 03:41 AM
Paper page - Cosmos 3: Omnimodal World Models for Physical AI
Source: https://huggingface.co/papers/2606.02800 Published on Jun 1
#1 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Cosmos 3 is an omnimodal world model that processes and generates multiple data types through a unified mixture-of-transformers architecture, achieving state-of-the-art performance in various understanding and generation tasks.
We introduce Cosmos 3, a family ofomnimodal world modelsdesigned to jointly process and generate language, image, video, audio, and action sequences within a unifiedmixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities forPhysical AI-- effectively subsumingvision-language models,video generators,world simulators, andworld-action modelsinto a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstratingomnimodal world modelsas scalable, general-purpose backbones forembodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment inPhysical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation’s OpenMDW-1.1 https://openmdw.ai/license/1-1/ License at https://github.com/nvidia/cosmos}{github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3 . The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3 .
View arXiv pageView PDFProject pageGitHub8.68kAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.02800 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.02800 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.02800 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@mli0603: This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI - Cosm…
Announcing Cosmos 3, an omnimodal world model for Physical AI that can understand and generate language, images, video, audio, and actions within a unified architecture.
nvidia/Cosmos3-Super
NVIDIA released Cosmos3, a collection of omnimodal world foundation models for Physical AI, capable of generating video, image, audio, and action commands from various inputs, with versions for different tasks like policy learning and image-to-video generation.
nvidia/Cosmos3-Nano
NVIDIA releases Cosmos3-Nano, an omnimodal world model for Physical AI that generates video, image, audio, and action commands from text, image, video, and action inputs, targeting robotics, autonomous driving, and smart space applications.
How Cosmos 3 Helps Physical AI Think Before It Acts
NVIDIA announces Cosmos 3, an open world foundation model that combines vision reasoning, multimodal generation, and action prediction to help robots, autonomous vehicles, and AI agents understand and predict real-world dynamics.
Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action
NVIDIA Cosmos 3 is an open omni-model for physical AI that unifies world generation, reasoning, and action generation into a single model, available on Hugging Face with various resources.