World Pilot: Steering Vision-Language-Action Models with World-Action Priors

Hugging Face Daily Papers 06/10/26, 12:00 AM Papers

vision-language-action world-action-model robotics manipulation zero-shot trajectory-priors dynamic-scene

Summary

World Pilot enhances Vision-Language-Action models by incorporating dynamic scene evolution and trajectory priors from a World-Action Model, achieving state-of-the-art zero-shot performance on manipulation tasks.

Vision-Language-Action (VLA) models inherit semantic grounding from large-scale pretraining and perform competently across in-distribution manipulation tasks. This grounding, however, is built on static image-text pairs, whereas manipulation is a continuous, contact-rich process whose dynamics this pretraining cannot capture. We present World Pilot, a VLA framework that augments the policy with priors from a World-Action Model (WAM), routed into the decision chain through two complementary pathways. Latent Steering conditions the perception layer on a scene-evolution latent, and Action Steering supplies an anticipated trajectory as a motion prior to the action generator. Together the two priors equip the VLA with an anticipated view of the scene and a trajectory-level motion hint alongside its semantic conditioning, and the scene-evolution prior remains effective even when supplied by a video-pretrained world model that has not been action-post-trained. World Pilot attains a state-of-the-art Total success rate of 84.7% on the LIBERO-Plus zero-shot OOD benchmark and the highest success rate on every real-robot setting across four manipulation tasks, with the largest margins under shifts in viewpoint, geometry, deformable state, and pose. Project Website: https://world-pilot.github.io/

Original Article

View Cached Full Text

Cached at: 06/11/26, 01:41 PM

Paper page - World Pilot: Steering Vision-Language-Action Models with World-Action Priors

Source: https://huggingface.co/papers/2606.12403

Abstract

World Pilot enhances Vision-Language-Action models by incorporating dynamic scene evolution and trajectory priors from a World-Action Model, achieving superior performance in zero-shot out-of-distribution manipulation tasks.

Vision-Language-Action (VLA) models inherit semantic grounding from large-scale pretraining and perform competently across in-distributionmanipulation tasks. This grounding, however, is built on static image-text pairs, whereas manipulation is a continuous, contact-rich process whose dynamics this pretraining cannot capture. We present World Pilot, a VLA framework that augments the policy with priors from aWorld-Action Model(WAM), routed into the decision chain through two complementary pathways.Latent Steeringconditions the perception layer on ascene-evolution latent, andAction Steeringsupplies ananticipated trajectoryas amotion priorto the action generator. Together the two priors equip the VLA with an anticipated view of the scene and a trajectory-level motion hint alongside its semantic conditioning, and the scene-evolution prior remains effective even when supplied by a video-pretrained world model that has not been action-post-trained. World Pilot attains a state-of-the-art Total success rate of 84.7% on the LIBERO-Pluszero-shot OOD benchmarkand the highest success rate on every real-robot setting across fourmanipulation tasks, with the largest margins under shifts in viewpoint, geometry, deformable state, and pose. Project Website: https://world-pilot.github.io/

View arXiv page View PDF Project page GitHub9 Add to collection

Get this paper in your agent:

hf papers read 2606\.12403

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### Chedan86/WorldPilot-LIBERO Robotics• Updatedabout 12 hours ago • 1

Datasets citing this paper1

#### Chedan86/WorldPilot-LIBERO-precompute Updatedabout 12 hours ago • 841 • 1

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.12403 in a Space README.md to link it from this page.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors

Paper page - World Pilot: Steering Vision-Language-Action Models with World-Action Priors

Abstract

Models citing this paper1

Datasets citing this paper1

Spaces citing this paper0

Collections including this paper1

Similar Articles

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Learning POMDP World Models from Observations with Language-Model Priors

The DAWN of World-Action Interactive Models

Submit Feedback

Similar Articles

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Learning POMDP World Models from Observations with Language-Model Priors

The DAWN of World-Action Interactive Models