world-model

Tag

Cards List
#world-model

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

arXiv cs.LG · yesterday Cached

This reproduction study independently reimplements LeWorldModel on the TwoRoom environment, reaching 94% goals rather than the reported 87%, and shows that four undocumented evaluation conventions determine the outcome. It also finds that one-step prediction error does not reliably predict long-horizon planning success and that batch normalization can inflate validation loss.

0 favorites 0 likes
#world-model

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

arXiv cs.AI · 6d ago Cached

DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to track latent state and predict future risks, enabling interventions before unsafe actions execute. It outperforms baselines on benchmarks and online evaluation with 25ms latency.

0 favorites 0 likes
#world-model

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Hugging Face Daily Papers · 2026-08-06 Cached

Presents UA-NWM, an uncertainty-aware latent world model for aerial image-goal navigation that decomposes prediction-goal discrepancy into uncertainty-explainable and unexplainable components, enabling robust trajectory scoring without multiple future samples.

0 favorites 0 likes
#world-model

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

Hugging Face Daily Papers · 2026-08-06 Cached

EnvACE introduces world rehearsal, an agentic reinforcement learning method that replaces external environment interaction by having the policy rehearse environment responses internally, achieving strong performance across multiple benchmarks.

0 favorites 0 likes
#world-model

UniNav: A Unified World-Action Diffusion Model for Visual Navigation

arXiv cs.AI · 2026-08-05 Cached

UniNav is a unified world-action diffusion model for image-goal visual navigation that jointly predicts future visual observations and waypoint trajectories in a single diffusion process, achieving strong benchmark results with efficient inference.

0 favorites 0 likes
#world-model

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Hugging Face Daily Papers · 2026-07-31 Cached

Introduces WCM, a World Critic Model that jointly predicts future latent states and estimates values to improve temporal modeling for Vision-Language-Action reinforcement learning, achieving state-of-the-art results across robotic manipulation benchmarks.

0 favorites 0 likes
#world-model

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

Hugging Face Daily Papers · 2026-07-30 Cached

This paper introduces ODEWorld, a continuous-time latent world model using Physical-Time Flow (PT-Flow) that learns a latent velocity field parameterized by an ordinary differential equation, enabling arbitrary temporal resolution, backward prediction, and improved planning for video generation and robotic control.

0 favorites 0 likes
#world-model

PhiZero: A World Model Built Around Physical Language

Hugging Face Daily Papers · 2026-07-30 Cached

PhiZero is a physical world model that learns a compact discrete representation called 'physical language' from videos and uses it to reason about world state transitions before rendering future videos, improving physical coherence in generation and understanding tasks.

0 favorites 0 likes
#world-model

The World Model and Spatial Intelligence Era: Governing AI Beyond Language

Reddit r/ArtificialInteligence · 2026-07-29 Cached

Stanford HAI discusses the emergence of world models and spatial intelligence in AI, calling for governance frameworks that address capabilities beyond language processing.

0 favorites 0 likes
#world-model

@svpino: The new WorldDiT robotics model is really cool: It's very small (< 1B parameters), yet it can perform prediction and co…

X AI KOLs Timeline · 2026-07-28 Cached

WorldDiT is a new, small (<1B parameters) robotics model that unifies world prediction and control, achieving top performance on the LIBERO benchmark without requiring a VLM.

0 favorites 0 likes
#world-model

@iScienceLuvr: Music-JEPA: Learning a World Model of Sound from Action "we propose to learn a world model of piano sound using JEPA by…

X AI KOLs Following · 2026-07-27 Cached

This paper proposes Music-JEPA, a world model that learns piano sound representations by framing audio as a state and piano roll as an action. It captures action-sound relationships and enables downstream tasks like beat tracking and piano transcription via planning.

0 favorites 0 likes
#world-model

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

Hugging Face Daily Papers · 2026-07-26 Cached

N₀-TWAM is a tactile-native world-action model for contact-rich manipulation, trained at scale on visuo-tactile data from 6 embodiments and 450 tasks. The authors release code and pretrained checkpoints, positioning it as the first tactile world-action model trained at scale.

0 favorites 0 likes
#world-model

@mattshumer_: Opus 5 enables this crazy new live AI tutor concept I'm prototyping for @AlphaSchool. It powers chat + direction under …

X AI KOLs Timeline · 2026-07-24 Cached

Opus 5 powers a live AI tutor prototype for AlphaSchool, combining chat, direction, and a custom tiny world model for real-time rendering and interactivity.

0 favorites 0 likes
#world-model

Flux 3 X Mimic: The Next Generation of Video-Action Models

Hacker News Top · 2026-07-24 Cached

Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly generates audio-visual content and, via collaboration with mimic robotics, enables video-action prediction for robot control, tested at Audi.

0 favorites 0 likes
#world-model

Lightricks/LTX-2.5

Hugging Face Models Trending · 2026-07-23 Cached

Lightricks releases LTX-2.5, an open-weights world model for generating synchronized video and audio from text, image, and video inputs, with features like native multishot generation and a new diffusion video decoder.

0 favorites 0 likes
#world-model

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Hugging Face Daily Papers · 2026-07-23 Cached

WorldWeaver (W²) introduces cross-agent world state registers to multi-agent video diffusion models, enabling shared world state persistence across agents and views, improving logical consistency in two-agent Minecraft video generation.

0 favorites 0 likes
#world-model

AlayaWorld is a full-stack, open-source video world model that supports 720p, 24 FPS streaming video generation with camera control

Reddit r/singularity · 2026-07-22

AlayaWorld is a full-stack, open-source video world model capable of generating 720p, 24 FPS streaming video with camera control, enabling dynamic scene creation.

0 favorites 0 likes
#world-model

Generative World Renderer at the Speed of Play

Hugging Face Daily Papers · 2026-07-21 Cached

This paper introduces AlayaRenderer-Flash, a real-time generative world renderer that accelerates rendering from 0.56 FPS to 31.54 FPS using a few-step autoregressive streaming model and lightweight distilled codecs, enabling interactive play with a physics engine.

0 favorites 0 likes
#world-model

Introducing Cosmos 3 Edge

Hugging Face Blog · 2026-07-20 Cached

NVIDIA released Cosmos 3 Edge, a 4-billion-parameter open world model for edge devices that helps robots and vision AI agents understand surroundings, reason in real time, and generate actions. It achieves best-in-class throughput and accuracy among similar-sized models.

0 favorites 0 likes
#world-model

DSWorld: A Data Science World Model for Efficient Autonomous Agents

arXiv cs.AI · 2026-07-20 Cached

DSWorld introduces a Data Science World Model that predicts environment state transitions to reduce costly trial-and-error in autonomous agents, achieving 14x acceleration in RL training and 3-6x in inference while maintaining competitive performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback