Pareto-Guided Teacher Alignment for Fair Personalized Text Generation
Summary
This paper introduces a Pareto-guided teacher alignment method for fair personalized text generation, aiming to balance multiple objectives in language model outputs.
View Cached Full Text
Cached at: 06/10/26, 06:10 AM
# Pareto-Guided Teacher Alignment for Fair Personalized Text Generation Source: [https://arxiv.org/abs/2606.10126](https://arxiv.org/abs/2606.10126) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
PAFO: Pareto Fairness Optimization for Personalized Reward Modeling
This paper proposes PAFO, a Pareto fairness optimization framework to mitigate personalized reward bias in reward models for LLMs, improving accuracy for minority user groups without harming majority groups.
Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
This paper presents the first comparative evaluation of training-free methods for personalizing toxicity sensitivity in language models at inference time, showing that all methods reduce alignment error by 28-47% but reveal a trade-off between alignment, personalization, and language quality.
On the Limits of Steering Vectors for Preference-Aligned Generation
This paper systematically studies the limitations of steering vectors for controlled text generation, finding that their effectiveness varies across traits, degrades on task transfer, and suffers from composition tradeoffs.
When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer
This paper proposes RidgeFT, a lightweight analytic update framework for lifelong machine-generated text attribution that adapts to new text generators without forgetting old ones, achieving strong performance across multiple evaluation settings.
Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation
This paper examines Self-Generated Text Recognition (SGTR) in large language models, revealing that evaluation design choices affect accuracy and that training for SGTR can induce self-preference biases, highlighting key implications for AI safety.