Tag
The paper introduces Reinforced Planning, a method that learns to improve multi-step plans using latent world models, achieving near-perfect success in tasks like visual navigation and robotic manipulation with significantly higher efficiency than hand-designed algorithms.
UniNav is a unified world-action diffusion model for image-goal visual navigation that jointly predicts future visual observations and waypoint trajectories in a single diffusion process, achieving strong benchmark results with efficient inference.
Apple Research introduces Weblica, a framework for creating scalable and reproducible training environments for visual web agents using HTTP caching and LLM-based synthesis.