dpo

Tag

Cards List
#dpo

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

arXiv cs.AI · 2026-08-03 Cached

This paper presents ConnectED, a human-centered AI system for Vietnamese education that uses the VietEduQwen LLM to support curriculum-aligned lesson planning and interactive student learning, reducing teacher prep time from hours to ~30-45 minutes while achieving 87% accuracy on national exam questions.

0 favorites 0 likes
#dpo

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

arXiv cs.CL · 2026-08-03 Cached

This paper studies how preference optimization shapes LLM counselors' behavior in motivational interviewing, finding that penalizing confrontation trades goal persistence for relational attunement rather than teaching the balanced skill of rolling with resistance.

0 favorites 0 likes
#dpo

Normalized Rewards for Preference Optimization

arXiv cs.LG · 2026-07-21 Cached

This paper introduces a regularization technique for Direct Alignment Algorithms (DAAs) that maintains normalized response probabilities, mitigating over-optimization and likelihood displacement. The method improves generation quality and benchmark performance, achieving over 20% relative increase on AlpacaEval2 and 9% gains on general benchmarks for Llama-3.1-8B-Instruct.

0 favorites 0 likes
#dpo

@h100envy: Ex-JPMorgan engineer who wrote the LLM Course explained everything about fine-tuning and merging in 18 minutes - better…

X AI KOLs Timeline · 2026-07-15 Cached

An ex-JPMorgan engineer behind the LLM Course offers an 18-minute guide on fine-tuning and merging models using LoRA, QLoRA, DPO, KTO, and mergekit, claiming it outperforms costly bootcamps.

0 favorites 0 likes
#dpo

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

arXiv cs.LG · 2026-07-14 Cached

This paper proposes a bilevel optimization framework for Direct Preference Optimization under noisy preference labels, introducing a metadata-free meta-reweighting method that uses central-difference approximation and LoRA fine-tuning to improve alignment performance.

0 favorites 0 likes
#dpo

D2PO: Optimizing Diffusion Samplers via Dynamic Preference

arXiv cs.LG · 2026-07-09 Cached

D2PO proposes a dynamic preference optimization framework that aligns diffusion sampling policies with perceptual quality using direct preference optimization, outperforming regression-based methods under low-NFE constraints.

0 favorites 0 likes
#dpo

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs

arXiv cs.CL · 2026-06-30 Cached

PASTA is a novel framework for knowledge updating in LLMs that combines data augmentation, question-answering generation, and self-learning DPO to integrate factual information from news articles, achieving accuracy improvement from 0.02 to 0.82 while preserving general capabilities.

0 favorites 0 likes
#dpo

Which Pairs to Compare for LLM Post-Training?

arXiv cs.AI · 2026-06-20 Cached

This paper studies the problem of selecting which completion pairs to label for human preference feedback in LLM post-training. It formulates comparison curation as a sampling-design problem, provides theoretical bounds on DPO's policy optimality gap, and proposes practical sampling designs that improve sample efficiency over common heuristics on synthetic and real benchmarks.

0 favorites 0 likes
#dpo

Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train (11 minute read)

TLDR AI · 2026-06-12 Cached

This research introduces a method using interpretability to predict which behaviors DPO will amplify or suppress from a preference dataset before training, enabling data debugging to prevent undesired effects. The technique achieves R²=0.9 prediction accuracy and is integrated into Goodfire's Silico platform.

0 favorites 0 likes
#dpo

Beyond the Golden Teacher: Enhancing Graph Learning through LLM-GNN Co-teaching

arXiv cs.LG · 2026-06-11 Cached

This paper proposes LLM-GNN Co-Teaching, a bidirectional framework for few-shot graph learning on text-attributed graphs. The LLM and GNN exchange confident pseudo-labels and use round-based preference optimization (RPL-PO) to mutually improve, outperforming prior methods on benchmarks.

0 favorites 0 likes
#dpo

Fine-tuned Qwen2.5-7B to 96% of Claude Haiku on a domain-specific task using ~$3 of API calls and zero human labelers

Reddit r/LocalLLaMA · 2026-06-10

Presented DV-DPO, a method to fine-tune Qwen2.5-7B on domain-specific tasks using only ~$3 in API calls and zero human labelers, achieving 96% composite performance of Claude Haiku via adversarial cross-examination.

0 favorites 0 likes
#dpo

Improving Multimodal Reasoning via Worst Dimension Optimization

arXiv cs.AI · 2026-06-09 Cached

This paper introduces Multimodal Multi-Dimensional Scalarization Process Reward Modeling (MMS-PRM), which enforces the worst dimension's robustness in multimodal reasoning to prevent failures like visual hallucinations from being masked by strong text logic.

0 favorites 0 likes
#dpo

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

arXiv cs.LG · 2026-06-09 Cached

DOG-DPO is a training-free data selection framework that treats preference pairs as structured geometric signals, decomposing multi-dataset preference geometry into anchor and residual subspaces to select diverse subsets for safety alignment. It achieves strong utility-robustness trade-offs using only 11% of preference pairs across six safety benchmarks.

0 favorites 0 likes
#dpo

@TheTuringPost: 15 Policy Optimization and Preference Optimization techniques important in 2026 GRPO DPO REINFORCE++ DAPO (Dynamic sAmp…

X AI KOLs Timeline · 2026-06-07 Cached

A comprehensive guide to 15 policy optimization and preference optimization techniques important in 2026, including GRPO, DPO, REINFORCE++, and many newer variants, mapping the landscape of reasoning RL methods.

0 favorites 0 likes
#dpo

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

arXiv cs.CL · 2026-06-04 Cached

VCIFBench is a new benchmark for evaluating complex instruction following in video understanding, featuring 306 test instructions with content, format, style, and structure constraints, plus a DPO preference dataset. Experiments on 10 MLLMs reveal that joint constraint satisfaction remains challenging, and DPO training on the benchmark data improves instruction-following performance.

0 favorites 0 likes
#dpo

Direct Preference Optimization Beyond Chatbots

Hugging Face Blog · 2026-06-03 Cached

Direct Preference Optimization (DPO) is applied to OCR tasks beyond chatbots, showing significant reduction in text degeneration across multiple model families, with an average reduction of 59.4%.

0 favorites 0 likes
#dpo

@yuwen_lu_: I'm halfway through, damn why did no one ever tell me RL is this fun

X AI KOLs Timeline · 2026-05-30 Cached

Sanbu 散步 released a modern RL tutorial Hands-On Modern RL, covering from CartPole+PPO basics to LLM post-training (RLHF, DPO, GRPO) and Agentic RL, code-first, English version coming soon.

0 favorites 0 likes
#dpo

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

arXiv cs.CL · 2026-05-27 Cached

This paper introduces CroCo, a method for cross-lingual contrastive preference tuning on self-generated responses, showing that a reward model trained on English preferences can effectively rank responses in other languages, improving model performance across 14 languages without language-specific annotations.

0 favorites 0 likes
#dpo

@neural_avb: Next video is on training tiny (<1B) models for preference tuning. Plus how to generate preference datasets with local …

X AI KOLs Timeline · 2026-05-26 Cached

Announces an upcoming video on training tiny models for preference tuning, covering reward models, RLHF, DPO, ORPO with Unsloth and TRL.

0 favorites 0 likes
#dpo

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

arXiv cs.AI · 2026-05-22 Cached

This paper proves that the equivalence between Direct Preference Optimization (DPO) and Reinforcement Learning from Human Feedback (RLHF) is conditional and often violated in practice, revealing failure modes where DPO optimizes relative advantage rather than absolute alignment. The authors introduce Constrained Preference Optimization (CPO) for provable alignment and demonstrate state-of-the-art performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback