video-world-models

Tag

Cards List
#video-world-models

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Hugging Face Daily Papers · 5d ago Cached

SolarWM introduces an open framework and unified training recipe for building interactive video world models with scalable training across diverse data sources, enabling long-horizon real-time rollouts.

0 favorites 0 likes
#video-world-models

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Hugging Face Daily Papers · 2026-08-14 Cached

ForgeWM is a progressive framework that distills bidirectional video generators into efficient few-step interactive world models, supporting low-latency interaction and replay-time refinement with demonstrated improvements on Minecraft and FPS gameplay.

0 favorites 0 likes
#video-world-models

Memory in Video World Models (6 minute read)

TLDR AI · 2026-08-12 Cached

A best-paper research from NVIDIA and collaborators introduces WorldTrace, a training-free framework that keeps compressed memory addressable in autoregressive video world models by assigning fixed slot-rank positions, enabling coherent long rollouts and long-range recall beyond the training horizon.

0 favorites 0 likes
#video-world-models

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Hugging Face Daily Papers · 2026-08-10 Cached

Introduces Latent Dynamics Reasoning (LDR), a video world model that integrates kinematic dynamics in a structured latent space, enabling extrapolation of learned dynamics far beyond training distributions while using far fewer parameters and running much faster than video diffusion baselines.

0 favorites 0 likes
#video-world-models

Addressable Memory for Video World Models

Hugging Face Daily Papers · 2026-08-07 Cached

This paper introduces WorldTrace, a training-free memory framework for long-horizon video world models that keeps compressed cache addressable, plus LoopBench, a benchmark for episodic recall after long detours. It improves temporal consistency by +15.5% and episodic recall by +19.5% on LoopBench.

0 favorites 0 likes
#video-world-models

HelloWorld: Enabling Socially Interactive Characters in Video World Models

Hugging Face Daily Papers · 2026-08-05 Cached

HelloWorld is a video world model that enables socially interactive characters, allowing users to prompt on-screen characters to respond via a single button press. It uses self-distillation and training-free cross-attention masking to naturalize interactions, and introduces HelloWorldBench for evaluation.

0 favorites 0 likes
#video-world-models

MiniWorld: Democratizing the Training of Video World Models from Scratch

Hugging Face Daily Papers · 2026-08-02 Cached

MiniWorld is a reproducible framework for training video world models from scratch using a block-causal Video Diffusion Transformer with Flow Matching, enabling efficient streaming generation and trainable in days on a single 8-GPU server.

0 favorites 0 likes
#video-world-models

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

Hugging Face Daily Papers · 2026-07-30 Cached

ShadowDancer proposes a method for any-action, frame-level control of interactive video world models by learning unified dynamics representations from a video and its shadow, enabling transferable action control without labels or motion estimators. Experiments show improved action transfer and rollout performance over baselines.

0 favorites 0 likes
#video-world-models

AlayaWorld: Long-Horizon and Playable Video World Generation

Hugging Face Daily Papers · 2026-07-07 Cached

AlayaWorld is an open-source framework for building interactive generative worlds that enables real-time user interaction and supports diverse actions. It unifies the complete development pipeline from data preparation to deployment.

0 favorites 0 likes
#video-world-models

MemLearner: Learning to Query Context memory for Video World Models

Hugging Face Daily Papers · 2026-06-30 Cached

MemLearner proposes a learning-based adaptive context query method using query tokens to improve scene consistency and memory in video world models, particularly for long sequences with occlusions and dynamic objects.

0 favorites 0 likes
#video-world-models

Holo-World: Unified Camera, Object and Weather Control for Video World Model

Hugging Face Daily Papers · 2026-06-18 Cached

Holo-World presents a unified controllable video world model that generates videos from a single image with explicit control over camera, object motion, and weather. It introduces a novel dataset and techniques to preserve scene structure while transferring to target weather states.

0 favorites 0 likes
#video-world-models

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Hugging Face Daily Papers · 2026-06-08 Cached

This paper introduces MBench, a benchmark for evaluating the memory capabilities of video world models across entity, environment, and causal consistency over long temporal horizons.

0 favorites 0 likes
#video-world-models

Latent Spatial Memory for Video World Models

Hugging Face Daily Papers · 2026-06-08 Cached

This paper introduces latent spatial memory for video world models, storing 3D scene information directly in diffusion latent space to avoid costly pixel-space reconstruction. The proposed Mirage framework achieves up to 10.57x faster generation and 55x memory reduction while achieving state-of-the-art performance on WorldScore and RealEstate10K.

0 favorites 0 likes
#video-world-models

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

Hugging Face Daily Papers · 2026-05-29 Cached

StressDream enhances video world models by steering diffusion-based imaginations toward high-impact yet plausible outcomes through optimized noise initialization with semantic and plausibility objectives, enabling robust policy evaluation and improvement.

0 favorites 0 likes
#video-world-models

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

Hugging Face Daily Papers · 2026-05-28 Cached

minWM is a full-stack open-source framework that converts bidirectional video diffusion models into real-time interactive video world models with controllable camera, low-latency rollout, and modular architecture.

0 favorites 0 likes
#video-world-models

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

Hugging Face Daily Papers · 2026-05-18 Cached

Incantation presents an interactive video world model that uses natural language as the action interface for fine-grained multi-entity control and cross-entity generalization, achieving high performance and real-time streaming through novel attention and distillation techniques.

0 favorites 0 likes
#video-world-models

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

Hugging Face Daily Papers · 2026-04-20 Cached

MultiWorld is a unified framework for multi-agent multi-view video world modeling that achieves accurate control of multiple agents while maintaining multi-view consistency through a Multi-Agent Condition Module and Global State Encoder.

0 favorites 0 likes
← Back to home

Submit Feedback