Tag
RLCD is explained as a schema-conditioned Plackett–Luce objective that advances reward modeling from scalar rewards to pairwise preferences to multiway calibrated decisions, simplifying the understanding of Jev.
This paper proposes MoPLEx, an algorithm for learning mixtures of Plackett-Luce models to handle heterogeneous preferences in AI alignment, showing improved clustering and ranking accuracy over baselines.
This paper proposes a distributionally robust listwise preference optimization method for LLM alignment that handles ranking-label uncertainty, with a tractable objective and strong convergence guarantees.