Tag
This paper recasts preference-based reward learning as a human-autonomy team problem, arguing that a teacher who knows the objective can design more efficient training examples than learner-driven query selection. It introduces understanding statements with second-order theory-of-mind to keep the teacher's model of the learner synchronized, showing in simulation that this approach outperforms learner-led selection.
This paper introduces KappaSharp, a method for condition-number shaping in Preferential Bayesian Optimization that corrects the ill-conditioned Hessian from isolated pairwise comparisons, showing up to +10.9% improvement over the standard PairedGP/EUBO baseline on 11 benchmarks including plasma medicine controller tuning.
Crayotter introduces Group-Relative Preference Backpropagation (GRPB), a method for training long-horizon video editing agents using within-task preference rankings rather than global scalar rewards. The resulting 9B model outperforms several proprietary systems on AgenticVBench.
This paper proposes a preference-based learning framework for antibody expression ranking, integrating scarce quantitative data with large-scale weak positive supervision from immunization sequences. The method adapts Direct Preference Optimization to protein language models using a union-masked log-likelihood approximation and IMGT-based alignment, achieving improved ranking performance on a diverse internal dataset.
Introduces 'weights to words', a method that automatically discovers domain-relevant preference dimensions described in natural language from choice data, enabling users to inspect and edit preference model inferences in real time.
The paper proposes the Plan, Learn, Adapt (PLA) framework for personalized on-device itinerary generation, combining feasibility-guaranteed combinatorial planning with human preference learning via a Bradley-Terry reward model. In deployment, it achieved a 91% increase in itinerary completion rates with low latency, outperforming frontier LLMs in feasibility.
This paper introduces DROPJ, a human-centred method for safely training and deploying agent policies by learning a world model from real-world trajectories, then eliciting human preferences with justifications to train a reward model for model predictive control. Experiments show that using human-generated simulated trajectories and justifications improves safety and reduces computational cost.
This paper critiques the use of pairwise comparisons for learning human preferences, arguing that internal pluralism (multiple conflicting priorities) undermines the standard approach. It proposes a formal model and suggests that allowing indecision can improve learning efficiency.
This paper scales the SOLiD lie-detector oversight method to larger LLMs (up to 405B parameters) and evaluates it in realistic preference-learning settings, finding that undetected deception decreases with model scale but that the method is sensitive to distribution shift between training data.
ThoughtFold proposes a framework using introspective preference learning to reduce redundant explorations in Chain-of-Thought reasoning for Large Reasoning Models, achieving ~56% token reduction on DeepSeek-R1-Distill-Qwen-7B without accuracy loss.
This paper introduces two automated metrics, Lexical Alignment Score and Triangulated Preference Shift, to identify lexical overuse in LLMs and attribute it to preference learning stages. The method is tested on six model families using PubMed abstracts, replicating prior findings without manual intervention.
Proposes SelSkill, a dual-granularity preference-learning framework that learns when to invoke skills in agentic tasks, improving task success by 10.9% on ALFWorld and 5.7% on BFCL.
Introduces a curation-free metric (Triangulated Preference Shift) to isolate and quantify lexical biases induced during preference learning in LLMs, without manual curation, across six model families.
This paper formalizes calibration for probabilistic label ranking, introducing a hierarchy of calibration notions and showing that common models are poorly calibrated. It further demonstrates applications to RLHF reward models, where calibration correlates with but is not identical to accuracy.
This paper introduces a method to improve factual consistency in text summarization by aggregating scores from multiple weak metrics via preference learning, achieving consistent factuality gains across various language models.
This paper proposes a unified framework for personalized agentic reinforcement learning that decouples generic task rewards from personalized preference rewards, introducing PARPO and PSGM for preference-aligned policy optimization and skill retrieval.
This paper formalizes trust calibration for agentic tool use as a preference learning problem, using Gaussian processes and Bayesian optimization to decide when an AI agent's actions should be autonomous or require human approval.
Proposes AMATA, a multi-agent trajectory alignment framework for knowledge-intensive question answering that introduces intra-trajectory preference learning and inter-agent dependency learning to improve factual grounding and interpretability, outperforming baselines on five benchmarks.
This paper introduces the Hybrid Reward-Cyclic (HRC) model and Dynamic Self-Play Preference Optimization (DSPPO) to address the cyclic nature of human preferences in LLM alignment, achieving improved performance over Bradley-Terry and General Preference Model baselines.
This paper introduces CLIPR, a framework that learns transferable latent user preferences from minimal conversational input to improve human-aligned decision making in LLMs.