Tag
This paper proposes an EM-based algorithm that jointly learns item rewards and worker reliability from pairwise comparisons under a Boltzmann-rational model, using Polya-Gamma latent variables for tractable optimization. Experiments show robustness to spammers and adversarial workers in crowdsourcing and reward learning settings.
This paper addresses the fixed-confidence top-k identification problem from noisy pairwise comparisons, and develops an asymptotically optimal algorithm that minimizes the expected number of comparisons.
This paper critiques the use of pairwise comparisons for learning human preferences, arguing that internal pluralism (multiple conflicting priorities) undermines the standard approach. It proposes a formal model and suggests that allowing indecision can improve learning efficiency.