Tag
EAVer introduces an end-to-end agentic policy for verifying facts in long-form text by grouping claims, efficiently searching, and reusing evidence, achieving superior performance on benchmarks like VeriFastScore and FaStFact-Bench with fewer searches.
SKIP is a self-knowledge-guided step-wise preference learning framework that improves reasoning compression in large language models, mitigating performance degradation and reducing overthinking by using DPO to guide efficient reasoning.
The paper proposes Style-Debiased DPO (SD-DPO), a method for factuality-aware synthetic preference data to improve knowledge elicitation in large language models, addressing issues where standard DPO may incorrectly penalize correct responses due to stylistic differences.
StalePO introduces a token-level preference optimization method for machine translation that uses legacy post-edits to improve model performance, showing significant gains in quality metrics.
Research from Stanford and Columbia shows that sycophantic agreement can emerge as an unintended consequence of contrastive preference optimization objectives, with teacher model bias transferring to student models even when preference data appears neutral across the dataset.
MemToC is a controlled benchmark for evaluating how large language models resolve conflicts between parametric memory and tool outputs, with findings that fine-tuning methods like SFT and DPO can improve correctness but require joint assessment with tool use and robustness.
This paper proposes VA-DPO, a method for controllable emotion generation in language models using continuous valence-arousal dimensions, which improves over prompting techniques without degrading model performance.
A developer trained a 125M-parameter transformer model to autocomplete piano performances in real-time on-device using MIDI data and DPO post-training, and released the RollTab app for iOS.
A recommended Stanford course on AI that details the principles behind building large language models, covering Tokenization, BPE, Transformer, pre-training, RLHF, and DPO.
MINT introduces min-selection preference distillation to balance multiple objectives in language agent alignment by ranking candidates based on their weakest objective, significantly reducing imbalance.
The paper introduces RA-DPO, a reliability-aware direct preference optimization method that combines annotator agreement, model confidence, and token-level uncertainty for sexism detection, improving training efficiency and enabling selective prediction.
This paper introduces Preference Tree Optimization (PTO), a framework that generates preference data via look-ahead simulations to iteratively improve goal-oriented dialogue agents, with experiments showing gains in Motivational Interviewing settings.
This paper presents ConnectED, a human-centered AI system for Vietnamese education that uses the VietEduQwen LLM to support curriculum-aligned lesson planning and interactive student learning, reducing teacher prep time from hours to ~30-45 minutes while achieving 87% accuracy on national exam questions.
This paper studies how preference optimization shapes LLM counselors' behavior in motivational interviewing, finding that penalizing confrontation trades goal persistence for relational attunement rather than teaching the balanced skill of rolling with resistance.
This paper introduces a regularization technique for Direct Alignment Algorithms (DAAs) that maintains normalized response probabilities, mitigating over-optimization and likelihood displacement. The method improves generation quality and benchmark performance, achieving over 20% relative increase on AlpacaEval2 and 9% gains on general benchmarks for Llama-3.1-8B-Instruct.
An ex-JPMorgan engineer behind the LLM Course offers an 18-minute guide on fine-tuning and merging models using LoRA, QLoRA, DPO, KTO, and mergekit, claiming it outperforms costly bootcamps.
This paper proposes a bilevel optimization framework for Direct Preference Optimization under noisy preference labels, introducing a metadata-free meta-reweighting method that uses central-difference approximation and LoRA fine-tuning to improve alignment performance.
D2PO proposes a dynamic preference optimization framework that aligns diffusion sampling policies with perceptual quality using direct preference optimization, outperforming regression-based methods under low-NFE constraints.
PASTA is a novel framework for knowledge updating in LLMs that combines data augmentation, question-answering generation, and self-learning DPO to integrate factual information from news articles, achieving accuracy improvement from 0.02 to 0.82 while preserving general capabilities.
This paper studies the problem of selecting which completion pairs to label for human preference feedback in LLM post-training. It formulates comparison curation as a sampling-design problem, provides theoretical bounds on DPO's policy optimality gap, and proposes practical sampling designs that improve sample efficiency over common heuristics on synthetic and real benchmarks.