Tag
The paper introduces Reinforced Planning, a method that learns to improve multi-step plans using latent world models, achieving near-perfect success in tasks like visual navigation and robotic manipulation with significantly higher efficiency than hand-designed algorithms.
This paper introduces decision-metric alignment diagnostics and action-conditioned objectives to improve latent world models for model-predictive control. The proposed DA-LeWM method enhances convergence and success rates in planning tasks.
The paper investigates why latent world models fail at long-horizon planning and finds the bottleneck is the planning objective (squared latent distance), not the predictor's accuracy; replacing the objective with a learned cost dramatically improves planning performance.
This paper addresses objective mismatch in model-based RL by proposing offline diagnostics to predict closed-loop performance of latent world models. On LunarLander-v3, the Reward Observability Fraction (ROF) and a Composite score (CROF) enable selecting checkpoints that yield strong MPC and model-based RL policies with far fewer real-environment interactions.