preference-learning

Tag

Cards List
#preference-learning

From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning

arXiv cs.CL ↗ · 2026-05-25 Cached

This paper proposes a unified framework for personalized agentic reinforcement learning that decouples generic task rewards from personalized preference rewards, introducing PARPO and PSGM for preference-aligned policy optimization and skill retrieval.

0 favorites 0 likes
#preference-learning

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

arXiv cs.AI ↗ · 2026-05-20 Cached

This paper formalizes trust calibration for agentic tool use as a preference learning problem, using Gaussian processes and Bayesian optimization to decide when an AI agent's actions should be autonomous or require human approval.

0 favorites 0 likes
#preference-learning

AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering

arXiv cs.CL ↗ · 2026-05-19 Cached

Proposes AMATA, a multi-agent trajectory alignment framework for knowledge-intensive question answering that introduces intra-trajectory preference learning and inter-agent dependency learning to improve factual grounding and interpretability, outperforming baselines on five benchmarks.

0 favorites 0 likes
#preference-learning

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment

arXiv cs.CL ↗ · 2026-05-19 Cached

This paper introduces the Hybrid Reward-Cyclic (HRC) model and Dynamic Self-Play Preference Optimization (DSPPO) to address the cyclic nature of human preferences in LLM alignment, achieving improved performance over Bradley-Terry and General Preference Model baselines.

0 favorites 0 likes
#preference-learning

Learning Transferable Latent User Preferences for Human-Aligned Decision Making

arXiv cs.AI ↗ · 2026-05-14 Cached

This paper introduces CLIPR, a framework that learns transferable latent user preferences from minimal conversational input to improve human-aligned decision making in LLMs.

0 favorites 0 likes
#preference-learning

$\xi$-DPO: Direct Preference Optimization via Ratio Reward Margin

arXiv cs.LG ↗ · 2026-05-13 Cached

This paper introduces xi-DPO, a novel preference optimization method that reformulates the objective to minimize distance to optimal ratio reward margins, addressing hyperparameter tuning challenges in SimPO. Experimental results show that xi-DPO outperforms existing methods on open benchmarks.

0 favorites 0 likes
#preference-learning

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

arXiv cs.CL ↗ · 2026-04-20 Cached

WildFeedback is a novel framework that leverages in-situ user feedback from actual LLM conversations to automatically create preference datasets for aligning language models with human preferences, addressing scalability and bias issues in traditional annotation-based alignment methods.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback