pluralistic-alignment

Tag

Cards List
#pluralistic-alignment

Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

arXiv cs.CL · 2026-08-13 Cached

This paper introduces Group Alignment-Induced Sycophancy (GAS), a two-sided evaluation framework that measures both the intended gain in opinion alignment and the unintended shift in sycophancy when aligning LLMs to demographic groups, finding that these effects are non-uniform and group-specific.

0 favorites 0 likes
#pluralistic-alignment

Rushes: A Human Preference Dataset for Pluralistic Alignment

arXiv cs.CL · 2026-07-24 Cached

Introduces Rushes, a large-scale dataset of human engagement preferences in AI-generated branching narratives, revealing that current LLMs like GPT-5 fail to outperform simple baselines in predicting user choices, highlighting the need for personalized alignment.

0 favorites 0 likes
#pluralistic-alignment

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

arXiv cs.LG · 2026-06-29 Cached

Introduces PEBS, a per-rater empirical-Bayes shrinkage estimator for calibrating reward models in RLHF, reducing within-user RMSE by over 8.5% on PRISM and over 9.6% on PluriHarms.

0 favorites 0 likes
#pluralistic-alignment

Steerable Cultural Preference Optimization of Reward Models

arXiv cs.CL · 2026-06-18 Cached

Introduces SCPO, a novel reward model training algorithm that incorporates diverse cultural preferences in a balanced manner, achieving up to 7 points improvement and 280% data efficiency over baselines.

0 favorites 0 likes
#pluralistic-alignment

Hidden Consensus:Preference-Validity Compression in Human Feedback

arXiv cs.CL · 2026-06-10 Cached

This paper argues that standard RLHF's scalarization of human preferences collapses multiple valid interpretations into a single target, mis-measuring alignment in culturally plural societies. Analyzing a Malaysian dataset, they find 79% of prompts have multiple majority-supported responses that single-winner aggregation discards.

0 favorites 0 likes
#pluralistic-alignment

Accounting for Context: Shaping Moral Credences for Value Alignment

arXiv cs.AI · 2026-06-08 Cached

This paper argues that aggregating moral evaluations for AI value alignment must account for contextual factors, showing that ignoring context can lead to violations of the weak Pareto principle, analogous to Simpson's paradox.

0 favorites 0 likes
#pluralistic-alignment

Coherence Maximization Improves Pluralistic Alignment

arXiv cs.CL · 2026-06-03 Cached

This paper introduces Internal Coherence Maximization (ICM) to generate persona-specific examples for aligning AI with diverse human values without human supervision, demonstrating that coherent examples generalize better across benchmarks.

0 favorites 0 likes
#pluralistic-alignment

A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI

arXiv cs.AI · 2026-06-01 Cached

This paper introduces a persona-based evaluation framework that uses synthetic cognitive profiles to represent diverse human perspectives for pluralistic alignment in generative AI, addressing the limitations of monolithic benchmarks.

0 favorites 0 likes
#pluralistic-alignment

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

arXiv cs.AI · 2026-05-15 Cached

This paper introduces DVMap, a framework for fine-grained pluralistic value alignment in LLMs that uses high-consensus demographic-value mapping instead of coarse national labels, achieving strong generalization across demographics, countries, and values.

0 favorites 0 likes
← Back to home

Submit Feedback