plackett-luce

Tag

Cards List
#plackett-luce

What Is RLCD? The Secret Behind Jev

Hacker News Top ↗ · yesterday Cached

RLCD is explained as a schema-conditioned Plackett–Luce objective that advances reward modeling from scalar rewards to pairwise preferences to multiway calibrated decisions, simplifying the understanding of Jev.

0 favorites 0 likes
#plackett-luce

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper proposes MoPLEx, an algorithm for learning mixtures of Plackett-Luce models to handle heterogeneous preferences in AI alignment, showing improved clustering and ranking accuracy over baselines.

0 favorites 0 likes
#plackett-luce

Distributionally Robust Listwise Preference Optimization

arXiv cs.AI ↗ · 2026-07-03 Cached

This paper proposes a distributionally robust listwise preference optimization method for LLM alignment that handles ranking-label uncertainty, with a tractable objective and strong convergence guarantees.

0 favorites 0 likes
← Back to home

Submit Feedback