preference-optimization

Tag

Cards List
#preference-optimization

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper studies how preference optimization shapes LLM counselors' behavior in motivational interviewing, finding that penalizing confrontation trades goal persistence for relational attunement rather than teaching the balanced skill of rolling with resistance.

0 favorites 0 likes
#preference-optimization

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

arXiv cs.CL ↗ · 2026-08-03 Cached

Presents TELLER, a dual-path iterative preference optimization approach for table entity linking, with direct-answer and reasoning paths that improve accuracy on TableInstruct and MammoTab V2 benchmarks.

0 favorites 0 likes
#preference-optimization

LARA: Lightweight Adapters in the Residual Stream for Composable Adaptation and Alignment

arXiv cs.LG ↗ · 2026-08-03 Cached

LARA is a method for efficient adaptation that adds low-rank corrections to a frozen model's residual stream instead of modifying weights, matching LoRA's performance while enabling composable behaviors and inference-time steering.

0 favorites 0 likes
#preference-optimization

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

arXiv cs.AI ↗ · 2026-07-29 Cached

This paper introduces DMAPO, a method for preference optimization that uses multi-evaluator consensus to select high-confidence on-policy responses, achieving strong alignment with significantly less data (only 3.45% acceptance rate) and outperforming baselines like SimPO on several benchmarks.

0 favorites 0 likes
#preference-optimization

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

arXiv cs.AI ↗ · 2026-07-29 Cached

This paper demonstrates that the final window of pretraining significantly influences a model's response to post-training alignment, even when SFT performance is identical, suggesting that checkpoint evaluation should include the last training data.

0 favorites 0 likes
#preference-optimization

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

arXiv cs.AI ↗ · 2026-07-24 Cached

Proposes Reliability-Guided Preference Optimization (RGPO) to handle inconsistent human feedback in LLM alignment by estimating annotator reliability and dynamically modulating training based on consensus, achieving superior performance over standard RLHF methods.

0 favorites 0 likes
#preference-optimization

TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue

arXiv cs.LG ↗ · 2026-07-22 Cached

This paper proposes TD-DPO, a token-level difference-aware preference optimization method to mitigate sycophancy in LLMs for clinical autism intervention dialogue, achieving a better trade-off between sycophancy reduction and intervention ability retention.

0 favorites 0 likes
#preference-optimization

Normalized Rewards for Preference Optimization

arXiv cs.LG ↗ · 2026-07-21 Cached

This paper introduces a regularization technique for Direct Alignment Algorithms (DAAs) that maintains normalized response probabilities, mitigating over-optimization and likelihood displacement. The method improves generation quality and benchmark performance, achieving over 20% relative increase on AlpacaEval2 and 9% gains on general benchmarks for Llama-3.1-8B-Instruct.

0 favorites 0 likes
#preference-optimization

RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper proposes RIMS, a three-stage preference optimization framework for small-scale language models in retrieval-augmented generation, using synthetic chain-of-thought data and a differentiable soft aggregation mechanism to improve robustness against noisy evidence. Experiments show consistent gains over baselines on multi-hop QA benchmarks.

0 favorites 0 likes
#preference-optimization

SciForma: Structure-Faithful Generation of Scientific Diagrams

Hugging Face Daily Papers ↗ · 2026-07-20 Cached

Introduces SciForma, a framework for generating scientific methodology diagrams with high structural fidelity, using multi-dimensional conjunctive preference optimization (M-DPO) and a structural inventory to ensure correctness across component, arrow, and text axes. The 9B model surpasses open-source baselines and GPT-Image-1.5 on benchmark evaluations.

0 favorites 0 likes
#preference-optimization

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

arXiv cs.LG ↗ · 2026-07-14 Cached

This paper proposes a bilevel optimization framework for Direct Preference Optimization under noisy preference labels, introducing a metadata-free meta-reweighting method that uses central-difference approximation and LoRA fine-tuning to improve alignment performance.

0 favorites 0 likes
#preference-optimization

D2PO: Optimizing Diffusion Samplers via Dynamic Preference

arXiv cs.LG ↗ · 2026-07-09 Cached

D2PO proposes a dynamic preference optimization framework that aligns diffusion sampling policies with perceptual quality using direct preference optimization, outperforming regression-based methods under low-NFE constraints.

0 favorites 0 likes
#preference-optimization

@helloiamleonie: Working with the @liquidai team on these engineering blogs is just so much fun! Here's what we've been working on: Reas…

X AI KOLs Following ↗ · 2026-07-07 Cached

The post explains doom loops in reasoning models where the model repeats tokens like 'Wait' until the context fills up, and introduces FTPO (Final Token Preference Optimization) as a training-time fix. The associated Antidoom tool reduces doom loop rates significantly (e.g., from 22.9% to 1% on Qwen3.5-4B).

0 favorites 0 likes
#preference-optimization

KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment

arXiv cs.CL ↗ · 2026-07-07 Cached

KARMA proposes a knowledge graph-based approach to generate slot-aligned contrastive candidates and uses Slot-Parallel Alignment (SPA) to apply preference optimization at the entity-slot level, addressing the Resolution Mismatch Problem in LLM reasoning supervision.

0 favorites 0 likes
#preference-optimization

Distributionally Robust Listwise Preference Optimization

arXiv cs.AI ↗ · 2026-07-03 Cached

This paper proposes a distributionally robust listwise preference optimization method for LLM alignment that handles ranking-label uncertainty, with a tractable objective and strong convergence guarantees.

0 favorites 0 likes
#preference-optimization

Multi-Objective Exploration and Preference Optimization via Mutual Information

arXiv cs.CL ↗ · 2026-07-03 Cached

Proposes MI-EPO, an information-theoretic framework for multi-objective alignment of large language models that uses mutual information to enhance exploration and ensure generated responses are distinguishable and aligned with different preference vectors, achieving stable trade-offs across conflicting objectives.

0 favorites 0 likes
#preference-optimization

CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models

arXiv cs.CL ↗ · 2026-07-02 Cached

CAT introduces a framework that leverages model self-certainty signals to autonomously adjust reasoning length based on problem difficulty, reducing overthinking and improving inference efficiency for large reasoning models.

0 favorites 0 likes
#preference-optimization

@dair_ai: New research from Google. LLMs hallucinate with high confidence, miss their own knowledge boundaries, and misreport unc…

X AI KOLs Timeline ↗ · 2026-07-02 Cached

A new research paper introduces RLMF (Reinforcement Learning with Metacognitive Feedback), a two-stage approach that uses the model's own self-judgments to calibrate confidence and express uncertainty faithfully, achieving state-of-the-art calibration across diverse tasks while preserving accuracy and surpassing standard RL by up to 63%.

0 favorites 0 likes
#preference-optimization

Which Pairs to Compare for LLM Post-Training?

arXiv cs.AI ↗ · 2026-06-20 Cached

This paper studies the problem of selecting which completion pairs to label for human preference feedback in LLM post-training. It formulates comparison curation as a sampling-design problem, provides theoretical bounds on DPO's policy optimality gap, and proposes practical sampling designs that improve sample efficiency over common heuristics on synthetic and real benchmarks.

0 favorites 0 likes
#preference-optimization

PolyAlign: Conditional Human-Distribution Alignment

arXiv cs.CL ↗ · 2026-06-12 Cached

PolyAlign is a distribution-aware alignment framework that aligns language models to context-specific human response distributions rather than a single global style, improving naturalness and faithfulness across bilingual settings.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback