pluralistic-alignment

Tag

Cards List
#pluralistic-alignment

Demographic Pluralism: Inference-Time Modeling of Pluralistic Human Preference Distributions

arXiv cs.AI ↗ · yesterday Cached

Amazon Science proposes Demographic Pluralism, an inference-time framework that estimates population-level opinion distributions by generating multiple perspectives within demographically grounded groups, reducing Jensen-Shannon distance by 8.4%-26.4% over Modular Pluralism across four backbones on GlobalOpinionQA and VITAL without requiring opinion-distribution training data.

0 favorites 0 likes
#pluralistic-alignment

Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

arXiv cs.AI ↗ · 2026-09-24 Cached

This paper investigates how to predict objective conflicts and cover trade-offs in steerable pluralistic alignment using Multi-Objective Direct Preference Optimization (MODPO), showing that pre-training measurements can predict alignment for human-annotated data and providing methods for broader trade-off coverage.

0 favorites 0 likes
#pluralistic-alignment

LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts

arXiv cs.CL ↗ · 2026-09-02 Cached

This research paper examines how sociodemographic prompting affects LLMs' judgments in subjective tasks, finding that it often aligns models with majority groups while misrepresenting minority groups, with instruction-tuning identified as a potential cause.

0 favorites 0 likes
#pluralistic-alignment

Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper introduces Group Alignment-Induced Sycophancy (GAS), a two-sided evaluation framework that measures both the intended gain in opinion alignment and the unintended shift in sycophancy when aligning LLMs to demographic groups, finding that these effects are non-uniform and group-specific.

0 favorites 0 likes
#pluralistic-alignment

Rushes: A Human Preference Dataset for Pluralistic Alignment

arXiv cs.CL ↗ · 2026-07-24 Cached

Introduces Rushes, a large-scale dataset of human engagement preferences in AI-generated branching narratives, revealing that current LLMs like GPT-5 fail to outperform simple baselines in predicting user choices, highlighting the need for personalized alignment.

0 favorites 0 likes
#pluralistic-alignment

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

arXiv cs.LG ↗ · 2026-06-29 Cached

Introduces PEBS, a per-rater empirical-Bayes shrinkage estimator for calibrating reward models in RLHF, reducing within-user RMSE by over 8.5% on PRISM and over 9.6% on PluriHarms.

0 favorites 0 likes
#pluralistic-alignment

Steerable Cultural Preference Optimization of Reward Models

arXiv cs.CL ↗ · 2026-06-18 Cached

Introduces SCPO, a novel reward model training algorithm that incorporates diverse cultural preferences in a balanced manner, achieving up to 7 points improvement and 280% data efficiency over baselines.

0 favorites 0 likes
#pluralistic-alignment

Hidden Consensus:Preference-Validity Compression in Human Feedback

arXiv cs.CL ↗ · 2026-06-10 Cached

This paper argues that standard RLHF's scalarization of human preferences collapses multiple valid interpretations into a single target, mis-measuring alignment in culturally plural societies. Analyzing a Malaysian dataset, they find 79% of prompts have multiple majority-supported responses that single-winner aggregation discards.

0 favorites 0 likes
#pluralistic-alignment

Accounting for Context: Shaping Moral Credences for Value Alignment

arXiv cs.AI ↗ · 2026-06-08 Cached

This paper argues that aggregating moral evaluations for AI value alignment must account for contextual factors, showing that ignoring context can lead to violations of the weak Pareto principle, analogous to Simpson's paradox.

0 favorites 0 likes
#pluralistic-alignment

Coherence Maximization Improves Pluralistic Alignment

arXiv cs.CL ↗ · 2026-06-03 Cached

This paper introduces Internal Coherence Maximization (ICM) to generate persona-specific examples for aligning AI with diverse human values without human supervision, demonstrating that coherent examples generalize better across benchmarks.

0 favorites 0 likes
#pluralistic-alignment

A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI

arXiv cs.AI ↗ · 2026-06-01 Cached

This paper introduces a persona-based evaluation framework that uses synthetic cognitive profiles to represent diverse human perspectives for pluralistic alignment in generative AI, addressing the limitations of monolithic benchmarks.

0 favorites 0 likes
#pluralistic-alignment

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

arXiv cs.AI ↗ · 2026-05-15 Cached

This paper introduces DVMap, a framework for fine-grained pluralistic value alignment in LLMs that uses high-consensus demographic-value mapping instead of coarse national labels, achieving strong generalization across demographics, countries, and values.

0 favorites 0 likes
← Back to home

Submit Feedback