dpo

Tag

Cards List
#dpo

EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy

arXiv cs.CL ↗ · 2026-09-22 Cached

EAVer introduces an end-to-end agentic policy for verifying facts in long-form text by grouping claims, efficiently searching, and reusing evidence, achieving superior performance on benchmarks like VeriFastScore and FaStFact-Bench with fewer searches.

0 favorites 0 likes
#dpo

SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning

arXiv cs.AI ↗ · 2026-09-16 Cached

SKIP is a self-knowledge-guided step-wise preference learning framework that improves reasoning compression in large language models, mitigating performance degradation and reducing overthinking by using DPO to guide efficient reasoning.

0 favorites 0 likes
#dpo

Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data

arXiv cs.CL ↗ · 2026-09-16 Cached

The paper proposes Style-Debiased DPO (SD-DPO), a method for factuality-aware synthetic preference data to improve knowledge elicitation in large language models, addressing issues where standard DPO may incorrectly penalize correct responses due to stylistic differences.

0 favorites 0 likes
#dpo

StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation

arXiv cs.CL ↗ · 2026-09-16 Cached

StalePO introduces a token-level preference optimization method for machine translation that uses legacy post-edits to improve model performance, showing significant gains in quality metrics.

0 favorites 0 likes
#dpo

@Memoirs: Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization Camila Blank, Zhuofan Ying, C…

X AI KOLs Following ↗ · 2026-09-04 Cached

Research from Stanford and Columbia shows that sycophantic agreement can emerge as an unintended consequence of contrastive preference optimization objectives, with teacher model bias transferring to student models even when preference data appears neutral across the dataset.

0 favorites 0 likes
#dpo

MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

arXiv cs.CL ↗ · 2026-08-28 Cached

MemToC is a controlled benchmark for evaluating how large language models resolve conflicts between parametric memory and tool outputs, with findings that fine-tuning methods like SFT and DPO can improve correctness but require joint assessment with tool use and robustness.

0 favorites 0 likes
#dpo

VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models

arXiv cs.CL ↗ · 2026-08-24 Cached

This paper proposes VA-DPO, a method for controllable emotion generation in language models using continuous valence-arousal dimensions, which improves over prompting techniques without degrading model performance.

0 favorites 0 likes
#dpo

Show HN: I trained a 125M model to autocomplete piano on-device

Hacker News Top ↗ · 2026-08-20 Cached

A developer trained a 125M-parameter transformer model to autocomplete piano performances in real-time on-device using MIDI data and DPO post-training, and released the RollTab app for iOS.

0 favorites 0 likes
#dpo

@FinanceYF5: Tonight, skip a TV show and finish this 2-hour 34-minute Stanford course. It covers from Tokenization, BPE to Transformer, pre-training, RLHF, DPO, and token-by-token generation, fully deconstructing how large models like ChatGPT and Claude are built…

X AI KOLs Following ↗ · 2026-08-20 Cached

A recommended Stanford course on AI that details the principles behind building large language models, covering Tokenization, BPE, Transformer, pre-training, RLHF, and DPO.

0 favorites 0 likes
#dpo

MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment

arXiv cs.AI ↗ · 2026-08-18 Cached

MINT introduces min-selection preference distillation to balance multiple objectives in language agent alignment by ranking candidates based on their weakest objective, significantly reducing imbalance.

0 favorites 0 likes
#dpo

Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring

arXiv cs.CL ↗ · 2026-08-14 Cached

The paper introduces RA-DPO, a reliability-aware direct preference optimization method that combines annotator agreement, model confidence, and token-level uncertainty for sexism detection, improving training efficiency and enabling selective prediction.

0 favorites 0 likes
#dpo

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper introduces Preference Tree Optimization (PTO), a framework that generates preference data via look-ahead simulations to iteratively improve goal-oriented dialogue agents, with experiments showing gains in Motivational Interviewing settings.

0 favorites 0 likes
#dpo

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper presents ConnectED, a human-centered AI system for Vietnamese education that uses the VietEduQwen LLM to support curriculum-aligned lesson planning and interactive student learning, reducing teacher prep time from hours to ~30-45 minutes while achieving 87% accuracy on national exam questions.

0 favorites 0 likes
#dpo

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper studies how preference optimization shapes LLM counselors' behavior in motivational interviewing, finding that penalizing confrontation trades goal persistence for relational attunement rather than teaching the balanced skill of rolling with resistance.

0 favorites 0 likes
#dpo

Normalized Rewards for Preference Optimization

arXiv cs.LG ↗ · 2026-07-21 Cached

This paper introduces a regularization technique for Direct Alignment Algorithms (DAAs) that maintains normalized response probabilities, mitigating over-optimization and likelihood displacement. The method improves generation quality and benchmark performance, achieving over 20% relative increase on AlpacaEval2 and 9% gains on general benchmarks for Llama-3.1-8B-Instruct.

0 favorites 0 likes
#dpo

@h100envy: Ex-JPMorgan engineer who wrote the LLM Course explained everything about fine-tuning and merging in 18 minutes - better…

X AI KOLs Timeline ↗ · 2026-07-15 Cached

An ex-JPMorgan engineer behind the LLM Course offers an 18-minute guide on fine-tuning and merging models using LoRA, QLoRA, DPO, KTO, and mergekit, claiming it outperforms costly bootcamps.

0 favorites 0 likes
#dpo

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

arXiv cs.LG ↗ · 2026-07-14 Cached

This paper proposes a bilevel optimization framework for Direct Preference Optimization under noisy preference labels, introducing a metadata-free meta-reweighting method that uses central-difference approximation and LoRA fine-tuning to improve alignment performance.

0 favorites 0 likes
#dpo

D2PO: Optimizing Diffusion Samplers via Dynamic Preference

arXiv cs.LG ↗ · 2026-07-09 Cached

D2PO proposes a dynamic preference optimization framework that aligns diffusion sampling policies with perceptual quality using direct preference optimization, outperforming regression-based methods under low-NFE constraints.

0 favorites 0 likes
#dpo

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs

arXiv cs.CL ↗ · 2026-06-30 Cached

PASTA is a novel framework for knowledge updating in LLMs that combines data augmentation, question-answering generation, and self-learning DPO to integrate factual information from news articles, achieving accuracy improvement from 0.02 to 0.82 while preserving general capabilities.

0 favorites 0 likes
#dpo

Which Pairs to Compare for LLM Post-Training?

arXiv cs.AI ↗ · 2026-06-20 Cached

This paper studies the problem of selecting which completion pairs to label for human preference feedback in LLM post-training. It formulates comparison curation as a sampling-design problem, provides theoretical bounds on DPO's policy optimality gap, and proposes practical sampling designs that improve sample efficiency over common heuristics on synthetic and real benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback