EchoWM: Open and Enterable Omnimodal World Models
Summary
EchoWM introduces an open omnimodal world model that generates synchronized video, sound, music, and speech with continuous 6-DoF navigation for enterable generative media.
View Cached Full Text
Cached at: 08/25/26, 08:34 AM
Paper page - EchoWM: Open and Enterable Omnimodal World Models
Source: https://huggingface.co/papers/2608.23189 Published on Aug 24
#3 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
EchoWM is an omnimodal world model that generates synchronized high-resolution video, sound, music, and speech while following continuous 6-DoF navigation trajectories across first- and third-person views.
We present EchoWM, anomnimodal world modelforenterable generative mediathat responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction aroundcamera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative6-DoF trajectory, withdataset-level calibrationpreserving motion magnitude across heterogeneous data. To jointly learnaudio-visual generationand trajectory control, we construct a complementary data engine and adoptprogressive trainingfollowed byautoregressive post-trainingforlong-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on publicworld-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech overlong-horizon generation.
View arXiv pageView PDFProject pageGitHub1.88kAdd to collection
Get this paper in your agent:
hf papers read 2608\.23189
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.23189 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.23189 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.23189 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Open Omnimodal World Models (GitHub Repo)
JoyAI-Echo is an open-source GitHub repository that provides long-horizon audio-visual generation and omnimodal world models for creating persistent stories and interactive worlds.
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
minWM is a full-stack open-source framework that converts bidirectional video diffusion models into real-time interactive video world models with controllable camera, low-latency rollout, and modular architecture.
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
HY-World 2.0 is a multi-modal world model framework that generates high-fidelity 3D Gaussian Splatting scenes from text, images, and videos through specialized modules for panorama generation, trajectory planning, and scene composition, achieving state-of-the-art performance among open-source approaches.
Wonder: Video World Model Done Better
Wonder is a general-purpose video world model that enables real-time, camera-controllable world exploration from an image or conditional video. It introduces camera conditioning via dense coordinate fields, a sparse attention memory mechanism, and techniques to improve distillation, allowing minute-scale video generation at 16 FPS.
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Ex-Omni-2D is an omni-modal dialogue framework that generates coordinated text, speech, and reference-conditioned video responses via a visual thought plan and a distilled streaming video generator, achieving a practical quality-efficiency trade-off.