Tag
This paper identifies 'Preference Coverage Collapse' as a failure mode in hindsight relabeling for multi-objective reinforcement learning and introduces 'her_mix' to mitigate it, improving performance across various settings.
Introduces Learning from Hindsight (LfH), a method that applies hindsight relabeling to RL post-training of vision-language-action models. By relabeling failed robot rollouts with the tasks they actually achieved, LfH achieves 5x improvement in sample efficiency on out-of-distribution manipulation tasks.