Tag
PRO-Step introduces a step-level process reward optimization method for retrieval-augmented generation to improve multi-hop reasoning by evaluating logical validity and evidential grounding at each step.
PLC-DPO enhances Direct Preference Optimization by routing noisy preference labels into clean, flipped, or tied cases using policy-reference margins, leading to improved performance across various benchmarks.
This paper investigates failures in a 2B model for dialogue games and introduces a diagnosis-guided post-training recipe using SFT, DPO, and LoRA to boost performance while maintaining general capabilities.
This paper proposes inference-time mitigation strategies using Chain of Thought prompting and Direct Preference Optimization to reduce adversarial political bias in large language models, demonstrating significant improvements in political neutrality scores.
ProteinDPO is a method that uses LLM preference learning techniques to improve the stability of protein models, developed by researchers at Arc Institute.
Introduces ELMER, an evolutionary language model that searches over natural-language policy descriptions and compiles them into executable programs, using fine-tuned Qwen3-8B with Direct Preference Optimization to control mutation strength and improve search efficiency.
DIRECT is a framework for sequence labeling using large language models that improves domain alignment through Direct Preference Optimization (DPO) after supervised fine-tuning and increases inference efficiency via controlled decoding with template-filling and KV cache reuse.
Proposes AutoThinkSQL, a framework that integrates an auto-thinking mechanism into SFT and DPO for Text-to-SQL, enabling the model to dynamically skip reasoning for simple queries and invoke deep CoT for complex ones, achieving gains on Spider and BIRD benchmarks while reducing output tokens by 24.6% and latency by 17.1%.
This paper proposes diversity-oriented fine-tuning strategies to improve uncertainty-based hallucination detection in LLMs by encouraging varied generations, making hallucinations more detectable via semantic entropy.
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.
DharmaOCR, a model specialized for Brazilian Portuguese OCR, outperforms newer models like Mistral OCR4 through domain-specific fine-tuning and Direct Preference Optimization. The article explains the training pipeline and presents benchmark results.
Proposes MI-EPO, an information-theoretic framework for multi-objective alignment of large language models that uses mutual information to enhance exploration and ensure generated responses are distinguishable and aligned with different preference vectors, achieving stable trade-offs across conflicting objectives.
This paper introduces Emergent Alignment, a self-supervised method that endows LLMs with a conscience step to review their own outputs and uses Direct Preference Optimization to steer away from unethical behavior, enabling online alignment without external judges.
This paper presents an empirical study of Direct Preference Optimization (DPO) for fine-tuning a large language model, showing that DPO simplifies the training pipeline and achieves competitive performance while addressing training instability.
Direct Preference Optimization (DPO) is applied to OCR tasks beyond chatbots, showing significant reduction in text degeneration across multiple model families, with an average reduction of 59.4%.
This paper proposes Staged-Competence, a curriculum learning framework for DPO-based safety alignment that organizes preference data by difficulty, improving robustness and data efficiency while preserving general capabilities.
This paper applies Direct Preference Optimization (DPO) to align Audio LLMs for transcribing English-Mandarin code-switching speech, achieving up to 89.6% MER reduction in-distribution and 20% out-of-distribution. It identifies three failure modes—language omission, translation instead of transcription, and hallucination—and shows that preference-based alignment effectively elicits correct code-switching behavior from multilingual Audio LLMs.
Proposes AttentionPO, a token-weighted direct preference optimization method that uses attention from the LLM itself to estimate token weights, improving alignment performance on AlpacaEval, MT-Bench, and ArenaHard without requiring a separate reward model.
This paper analyzes spurious correlation learning in preference optimization methods like DPO, identifying mechanisms such as mean spurious bias and causal-spurious leakage. It proposes 'tie training' using equal-utility preference pairs as a mitigation strategy to reduce reliance on spurious features without degrading causal learning.
This paper introduces xi-DPO, a novel preference optimization method that reformulates the objective to minimize distance to optimal ratio reward margins, addressing hyperparameter tuning challenges in SimPO. Experimental results show that xi-DPO outperforms existing methods on open benchmarks.