preference-optimization

Tag

Cards List
#preference-optimization

Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

arXiv cs.AI ↗ · 5d ago Cached

This paper investigates how to predict objective conflicts and cover trade-offs in steerable pluralistic alignment using Multi-Objective Direct Preference Optimization (MODPO), showing that pre-training measurements can predict alignment for human-annotated data and providing methods for broader trade-off coverage.

0 favorites 0 likes
#preference-optimization

DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization

arXiv cs.CL ↗ · 2026-09-16 Cached

The paper proposes DiaWhisper-DPO, an end-to-end model for transcription and role attribution in clinical interviews using failure-mined preference optimization, achieving high accuracy and reducing errors compared to cascaded baselines.

0 favorites 0 likes
#preference-optimization

Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data

arXiv cs.CL ↗ · 2026-09-16 Cached

The paper proposes Style-Debiased DPO (SD-DPO), a method for factuality-aware synthetic preference data to improve knowledge elicitation in large language models, addressing issues where standard DPO may incorrectly penalize correct responses due to stylistic differences.

0 favorites 0 likes
#preference-optimization

StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation

arXiv cs.CL ↗ · 2026-09-16 Cached

StalePO introduces a token-level preference optimization method for machine translation that uses legacy post-edits to improve model performance, showing significant gains in quality metrics.

0 favorites 0 likes
#preference-optimization

A Zeroth-Order Paradigm for LLM Preference Alignment

Hugging Face Daily Papers ↗ · 2026-09-16 Cached

The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method for LLMs that uses comparison oracles to avoid likelihood displacement. It includes theoretical guarantees and experimental improvements over existing methods.

0 favorites 0 likes
#preference-optimization

When and What to Teach: Budget-Aware Online Adaptation for Web Agents

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper proposes a budget-aware online teaching framework for web agents that reduces teacher calls and compute costs while maintaining performance.

0 favorites 0 likes
#preference-optimization

Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training

arXiv cs.CL ↗ · 2026-09-10 Cached

Direct Diversity Optimization (DDO) is an offline post-training method that improves successful strategy coverage in LLM agents for sequential decision tasks, outperforming other methods in benchmarks like BabyAI, BabaIsAI, and WebShop.

0 favorites 0 likes
#preference-optimization

@Memoirs: Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization Camila Blank, Zhuofan Ying, C…

X AI KOLs Following ↗ · 2026-09-04 Cached

Research from Stanford and Columbia shows that sycophantic agreement can emerge as an unintended consequence of contrastive preference optimization objectives, with teacher model bias transferring to student models even when preference data appears neutral across the dataset.

0 favorites 0 likes
#preference-optimization

FiMI Banking: A Sovereign Model for Indian Retail Banking

arXiv cs.AI ↗ · 2026-09-04 Cached

The paper introduces FiMI Banking, a sovereign conversational AI model for Indian retail banking, trained using preference optimization and reinforcement learning to enhance safe behavior and tool-use performance within regulatory constraints.

0 favorites 0 likes
#preference-optimization

CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging

arXiv cs.AI ↗ · 2026-09-03 Cached

CoMerge is a conflict-driven preference optimization framework for merging multi-task LLMs, using self-supervised strategies to mitigate parameter interference and achieve high performance on benchmarks like MergeBench.

0 favorites 0 likes
#preference-optimization

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

PLC-DPO enhances Direct Preference Optimization by routing noisy preference labels into clean, flipped, or tied cases using policy-reference margins, leading to improved performance across various benchmarks.

0 favorites 0 likes
#preference-optimization

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper proposes MoPLEx, an algorithm for learning mixtures of Plackett-Luce models to handle heterogeneous preferences in AI alignment, showing improved clustering and ranking accuracy over baselines.

0 favorites 0 likes
#preference-optimization

Function-Level Execution Feedback for Code Preference Optimization

arXiv cs.AI ↗ · 2026-08-26 Cached

This paper proposes Step-KTOder, a framework for code preference optimization that uses function-level execution feedback with unit tests to improve over outcome-only methods like KTO and DPO on benchmarks such as HumanEval+ and MBPP+.

0 favorites 0 likes
#preference-optimization

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

Hugging Face Daily Papers ↗ · 2026-08-24 Cached

Cross-lingual Ranking Preference Optimization (CRPO) is a novel framework that enhances multilingual LLM alignment by transferring English preference knowledge to target languages through hierarchical ranking optimization, demonstrating improved performance in instruction-following and knowledge utilization across multiple languages.

0 favorites 0 likes
#preference-optimization

FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design

arXiv cs.LG ↗ · 2026-08-21 Cached

FAR-DPO is a feasibility-aware and robust direct preference optimization framework that enhances cyclic peptide design for drug discovery by aligning generative models with structural and biophysical constraints, improving success rates on benchmarks like PepGLAD and PepFlow.

0 favorites 0 likes
#preference-optimization

Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

arXiv cs.AI ↗ · 2026-08-21 Cached

The paper identifies manifold drift as a root cause of reward hacking in flow preference optimization and introduces ThermoDPO, a temperature-controlled method that anchors optimization on preferred samples to preserve data manifold integrity and improve performance.

0 favorites 0 likes
#preference-optimization

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

arXiv cs.LG ↗ · 2026-08-19 Cached

Data-DPO is a target model-oriented data selection method for LLM supervised fine-tuning that learns data preferences through one-step probing and combines them with quality scores and diversity, outperforming baselines on Vision-Flan and LLaVA-CoT datasets.

0 favorites 0 likes
#preference-optimization

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper introduces SFS-DPO, a reinforcement learning two-stage framework for step-level self-verification and self-correction in LLMs, with a teacher-assisted variant SFS-DPO-R. It demonstrates improvements in self-correction effectiveness across multiple LLMs with less training data than prior approaches.

0 favorites 0 likes
#preference-optimization

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

arXiv cs.CL ↗ · 2026-08-06 Cached

The paper proposes MAP-PO, a multi-agent framework that clusters annotators by labeling behavior and fine-tunes separate LLM agents per cluster using preference optimization, preserving disagreement in sexism detection tasks. Experiments on the EXIST 2024 dataset show that cluster-specific training is necessary and that a shared team-level reward keeps agents calibrated.

0 favorites 0 likes
#preference-optimization

Show HN: Fine-tune an 8B model on a 4 GB laptop GPU

Hacker News Top ↗ · 2026-08-04 Cached

Soup is an open-source CLI that simplifies LLM fine-tuning and post-training with a single command, enabling QLoRA-based training on consumer GPUs with as little as 4 GB VRAM via layer streaming. The latest version adds preference losses like DPO, ORPO, SimPO, and KTO without doubling memory requirements.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback