hindsight-relabeling

Tag

Cards List
#hindsight-relabeling

On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning

arXiv cs.LG ↗ · yesterday Cached

This paper identifies 'Preference Coverage Collapse' as a failure mode in hindsight relabeling for multi-objective reinforcement learning and introduces 'her_mix' to mitigate it, improving performance across various settings.

0 favorites 0 likes
#hindsight-relabeling

Learning More from Less: Reinforcement Learning from Hindsight

arXiv cs.LG ↗ · 2026-07-13 Cached

Introduces Learning from Hindsight (LfH), a method that applies hindsight relabeling to RL post-training of vision-language-action models. By relabeling failed robot rollouts with the tasks they actually achieved, LfH achieves 5x improvement in sample efficiency on out-of-distribution manipulation tasks.

0 favorites 0 likes
← Back to home

Submit Feedback