Tag
HDR is a unified framework integrating hierarchical latents into causal video generation for multi-step visual reasoning, achieving better reasoning consistency, lower latency, and strong data efficiency compared to baselines.
This video was entirely produced by AI agents within Raft, with only human direction, showcasing agent-driven content creation.
KeyFrame-Compass is a benchmark and evaluation framework for keyframe-conditioned video generation, designed to assess how well models reproduce given keyframes while maintaining video quality across diverse settings.
This paper rethinks interactive world models as game engines by examining four key dimensions—action control, state dynamics, state-observation persistence, and real-time generation—and introduces a scalable data engine for Black Myth: Wukong with over 90 hours of gameplay data to advance state-aware game world modeling.
This paper introduces VideoRAE, a representation autoencoder that leverages frozen video foundation models to create compact, reconstruction-capable, and generation-friendly video latents. It achieves state-of-the-art results on UCF-101 with faster convergence than competing autoencoders.
LingBot-World 2.0 achieves stable 720p 60fps world model simulation for up to an hour, overcoming common failure modes like texture smearing and warped geometry.
A new state-of-the-art agentic pipeline has been introduced for easy music video creation, leveraging AI to streamline the process.
A user expresses admiration for a video generated by an AI model and asks which model was used, implying a significant advancement in AI video generation.
GenCeption repurposes pre-trained video generative models into a single unified feed-forward vision model that achieves state-of-the-art performance across multiple tasks with exceptional data efficiency, marking a shift toward general-purpose visual intelligence.
Singapore-based video-generation startup PixVerse raised $439 million in a Series C extension, pushing its valuation past $2 billion. The company plans to expand its world model offerings and reach global customers.
This tweet explains LingBot-World-Infinity, an open-weight video generation model that uses a training technique to recover from errors, enabling coherent hour-long videos across multiple scenes.
An open model that predicts a robot's actions from a control signal, raising questions about whether it constitutes a world model or just a video generator.
A proof-of-concept In-Context LoRA adapter for LTX-Video 2.3 that re-renders video scenes from new camera angles using a fixed vocabulary prompt, trained on synthetic multi-view data.
Wan-Dancer is a hierarchical framework for generating long-duration, coherent dance videos from music, with model weights and inference code released on Hugging Face.
This paper proposes that large-scale text-to-video generation can serve as a powerful pre-training paradigm for computer vision, introducing GenCeption which achieves state-of-the-art performance across diverse vision tasks with high data efficiency and emergent generalization to unseen domains.
The user tested Grok's video generation capability, showcasing a World Cup quarterfinal preview video, and noted significant improvement, fast generation speed, and impressive visual effects.
The paper proposes a post-training acceleration framework for video diffusion models that integrates dynamic structural sparsification with few-step distillation, achieving significant speedup while maintaining quality.
OPSD-V improves few-step autoregressive video diffusion models by using real long-video data as temporal context during training, providing dense trajectory-level supervision that enhances visual quality and motion dynamics without altering inference mechanisms.
OpenCoF introduces a reasoning video dataset and a fine-tuned video generation model that improves temporal reasoning through diverse supervision and explicit reasoning tokens, showing significant gains on four video reasoning benchmarks.
LingBot-Video is the first open-source large-scale MoE video generation model for embodied intelligence, featuring efficient MoE architecture, massive embodied data training, and multi-reward system for high aesthetics, physical rationality, and task completion.