Tag
This paper studies how preference optimization shapes LLM counselors' behavior in motivational interviewing, finding that penalizing confrontation trades goal persistence for relational attunement rather than teaching the balanced skill of rolling with resistance.
Presents TELLER, a dual-path iterative preference optimization approach for table entity linking, with direct-answer and reasoning paths that improve accuracy on TableInstruct and MammoTab V2 benchmarks.
LARA is a method for efficient adaptation that adds low-rank corrections to a frozen model's residual stream instead of modifying weights, matching LoRA's performance while enabling composable behaviors and inference-time steering.
This paper introduces DMAPO, a method for preference optimization that uses multi-evaluator consensus to select high-confidence on-policy responses, achieving strong alignment with significantly less data (only 3.45% acceptance rate) and outperforming baselines like SimPO on several benchmarks.
This paper demonstrates that the final window of pretraining significantly influences a model's response to post-training alignment, even when SFT performance is identical, suggesting that checkpoint evaluation should include the last training data.
Proposes Reliability-Guided Preference Optimization (RGPO) to handle inconsistent human feedback in LLM alignment by estimating annotator reliability and dynamically modulating training based on consensus, achieving superior performance over standard RLHF methods.
This paper proposes TD-DPO, a token-level difference-aware preference optimization method to mitigate sycophancy in LLMs for clinical autism intervention dialogue, achieving a better trade-off between sycophancy reduction and intervention ability retention.
This paper introduces a regularization technique for Direct Alignment Algorithms (DAAs) that maintains normalized response probabilities, mitigating over-optimization and likelihood displacement. The method improves generation quality and benchmark performance, achieving over 20% relative increase on AlpacaEval2 and 9% gains on general benchmarks for Llama-3.1-8B-Instruct.
This paper proposes RIMS, a three-stage preference optimization framework for small-scale language models in retrieval-augmented generation, using synthetic chain-of-thought data and a differentiable soft aggregation mechanism to improve robustness against noisy evidence. Experiments show consistent gains over baselines on multi-hop QA benchmarks.
Introduces SciForma, a framework for generating scientific methodology diagrams with high structural fidelity, using multi-dimensional conjunctive preference optimization (M-DPO) and a structural inventory to ensure correctness across component, arrow, and text axes. The 9B model surpasses open-source baselines and GPT-Image-1.5 on benchmark evaluations.
This paper proposes a bilevel optimization framework for Direct Preference Optimization under noisy preference labels, introducing a metadata-free meta-reweighting method that uses central-difference approximation and LoRA fine-tuning to improve alignment performance.
D2PO proposes a dynamic preference optimization framework that aligns diffusion sampling policies with perceptual quality using direct preference optimization, outperforming regression-based methods under low-NFE constraints.
The post explains doom loops in reasoning models where the model repeats tokens like 'Wait' until the context fills up, and introduces FTPO (Final Token Preference Optimization) as a training-time fix. The associated Antidoom tool reduces doom loop rates significantly (e.g., from 22.9% to 1% on Qwen3.5-4B).
KARMA proposes a knowledge graph-based approach to generate slot-aligned contrastive candidates and uses Slot-Parallel Alignment (SPA) to apply preference optimization at the entity-slot level, addressing the Resolution Mismatch Problem in LLM reasoning supervision.
This paper proposes a distributionally robust listwise preference optimization method for LLM alignment that handles ranking-label uncertainty, with a tractable objective and strong convergence guarantees.
Proposes MI-EPO, an information-theoretic framework for multi-objective alignment of large language models that uses mutual information to enhance exploration and ensure generated responses are distinguishable and aligned with different preference vectors, achieving stable trade-offs across conflicting objectives.
CAT introduces a framework that leverages model self-certainty signals to autonomously adjust reasoning length based on problem difficulty, reducing overthinking and improving inference efficiency for large reasoning models.
A new research paper introduces RLMF (Reinforcement Learning with Metacognitive Feedback), a two-stage approach that uses the model's own self-judgments to calibrate confidence and express uncertainty faithfully, achieving state-of-the-art calibration across diverse tasks while preserving accuracy and surpassing standard RL by up to 63%.
This paper studies the problem of selecting which completion pairs to label for human preference feedback in LLM post-training. It formulates comparison curation as a sampling-design problem, provides theoretical bounds on DPO's policy optimality gap, and proposes practical sampling designs that improve sample efficiency over common heuristics on synthetic and real benchmarks.
PolyAlign is a distribution-aware alignment framework that aligns language models to context-specific human response distributions rather than a single global style, improving naturalness and faithfulness across bilingual settings.