Tag
This paper identifies 'Preference Coverage Collapse' as a failure mode in hindsight relabeling for multi-objective reinforcement learning and introduces 'her_mix' to mitigate it, improving performance across various settings.
This paper identifies group preference collapse in personalized multimodal large language models and proposes PrefMoE, a preference-centric framework that separates profile information from preferences to improve personalization and reduce collapse.