Tag
This paper introduces LEMUR, a framework that combines multi-objective reinforcement learning with preference-based learning from multiple human feedback to learn Pareto-optimal policies without predefined reward functions.
AETDICE proposes a unified framework for nonlinear multi-objective reinforcement learning in offline settings, bridging SER and ESR paradigms via density-ratio estimation.