Tag
This paper introduces Group Alignment-Induced Sycophancy (GAS), a two-sided evaluation framework that measures both the intended gain in opinion alignment and the unintended shift in sycophancy when aligning LLMs to demographic groups, finding that these effects are non-uniform and group-specific.
Introduces Rushes, a large-scale dataset of human engagement preferences in AI-generated branching narratives, revealing that current LLMs like GPT-5 fail to outperform simple baselines in predicting user choices, highlighting the need for personalized alignment.
Introduces PEBS, a per-rater empirical-Bayes shrinkage estimator for calibrating reward models in RLHF, reducing within-user RMSE by over 8.5% on PRISM and over 9.6% on PluriHarms.
Introduces SCPO, a novel reward model training algorithm that incorporates diverse cultural preferences in a balanced manner, achieving up to 7 points improvement and 280% data efficiency over baselines.
This paper argues that standard RLHF's scalarization of human preferences collapses multiple valid interpretations into a single target, mis-measuring alignment in culturally plural societies. Analyzing a Malaysian dataset, they find 79% of prompts have multiple majority-supported responses that single-winner aggregation discards.
This paper argues that aggregating moral evaluations for AI value alignment must account for contextual factors, showing that ignoring context can lead to violations of the weak Pareto principle, analogous to Simpson's paradox.
This paper introduces Internal Coherence Maximization (ICM) to generate persona-specific examples for aligning AI with diverse human values without human supervision, demonstrating that coherent examples generalize better across benchmarks.
This paper introduces a persona-based evaluation framework that uses synthetic cognitive profiles to represent diverse human perspectives for pluralistic alignment in generative AI, addressing the limitations of monolithic benchmarks.
This paper introduces DVMap, a framework for fine-grained pluralistic value alignment in LLMs that uses high-consensus demographic-value mapping instead of coarse national labels, achieving strong generalization across demographics, countries, and values.