Tag
This paper investigates how to predict objective conflicts and cover trade-offs in steerable pluralistic alignment using Multi-Objective Direct Preference Optimization (MODPO), showing that pre-training measurements can predict alignment for human-annotated data and providing methods for broader trade-off coverage.
The paper proposes DiaWhisper-DPO, an end-to-end model for transcription and role attribution in clinical interviews using failure-mined preference optimization, achieving high accuracy and reducing errors compared to cascaded baselines.
The paper proposes Style-Debiased DPO (SD-DPO), a method for factuality-aware synthetic preference data to improve knowledge elicitation in large language models, addressing issues where standard DPO may incorrectly penalize correct responses due to stylistic differences.
StalePO introduces a token-level preference optimization method for machine translation that uses legacy post-edits to improve model performance, showing significant gains in quality metrics.
The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method for LLMs that uses comparison oracles to avoid likelihood displacement. It includes theoretical guarantees and experimental improvements over existing methods.
This paper proposes a budget-aware online teaching framework for web agents that reduces teacher calls and compute costs while maintaining performance.
Direct Diversity Optimization (DDO) is an offline post-training method that improves successful strategy coverage in LLM agents for sequential decision tasks, outperforming other methods in benchmarks like BabyAI, BabaIsAI, and WebShop.
Research from Stanford and Columbia shows that sycophantic agreement can emerge as an unintended consequence of contrastive preference optimization objectives, with teacher model bias transferring to student models even when preference data appears neutral across the dataset.
The paper introduces FiMI Banking, a sovereign conversational AI model for Indian retail banking, trained using preference optimization and reinforcement learning to enhance safe behavior and tool-use performance within regulatory constraints.
CoMerge is a conflict-driven preference optimization framework for merging multi-task LLMs, using self-supervised strategies to mitigate parameter interference and achieve high performance on benchmarks like MergeBench.
PLC-DPO enhances Direct Preference Optimization by routing noisy preference labels into clean, flipped, or tied cases using policy-reference margins, leading to improved performance across various benchmarks.
This paper proposes MoPLEx, an algorithm for learning mixtures of Plackett-Luce models to handle heterogeneous preferences in AI alignment, showing improved clustering and ranking accuracy over baselines.
This paper proposes Step-KTOder, a framework for code preference optimization that uses function-level execution feedback with unit tests to improve over outcome-only methods like KTO and DPO on benchmarks such as HumanEval+ and MBPP+.
Cross-lingual Ranking Preference Optimization (CRPO) is a novel framework that enhances multilingual LLM alignment by transferring English preference knowledge to target languages through hierarchical ranking optimization, demonstrating improved performance in instruction-following and knowledge utilization across multiple languages.
FAR-DPO is a feasibility-aware and robust direct preference optimization framework that enhances cyclic peptide design for drug discovery by aligning generative models with structural and biophysical constraints, improving success rates on benchmarks like PepGLAD and PepFlow.
The paper identifies manifold drift as a root cause of reward hacking in flow preference optimization and introduces ThermoDPO, a temperature-controlled method that anchors optimization on preferred samples to preserve data manifold integrity and improve performance.
Data-DPO is a target model-oriented data selection method for LLM supervised fine-tuning that learns data preferences through one-step probing and combines them with quality scores and diversity, outperforming baselines on Vision-Flan and LLaVA-CoT datasets.
This paper introduces SFS-DPO, a reinforcement learning two-stage framework for step-level self-verification and self-correction in LLMs, with a teacher-assisted variant SFS-DPO-R. It demonstrates improvements in self-correction effectiveness across multiple LLMs with less training data than prior approaches.
The paper proposes MAP-PO, a multi-agent framework that clusters annotators by labeling behavior and fine-tunes separate LLM agents per cluster using preference optimization, preserving disagreement in sexism detection tasks. Experiments on the EXIST 2024 dataset show that cluster-specific training is necessary and that a shared team-level reward keeps agents calibrated.
Soup is an open-source CLI that simplifies LLM fine-tuning and post-training with a single command, enabling QLoRA-based training on consumer GPUs with as little as 4 GB VRAM via layer streaming. The latest version adds preference losses like DPO, ORPO, SimPO, and KTO without doubling memory requirements.