dpo

Tag

Cards List
#dpo

@akshay_pachaar: The Operating System for Al Research Labs. TransformerLab orchestrates GPUs across any cloud and runs any training or e…

X AI KOLs Following · 2026-05-20 Cached

TransformerLab is an open-source platform that orchestrates GPUs across clouds and provides pre-built templates for AI training and evaluation workflows like LoRA, DPO, and MMLU.

0 favorites 0 likes
#dpo

@anyscalecompute: LLM post-training is the new baseline. Picking the wrong method or GPU config is how you waste a 36-hour run. Introduci…

X AI KOLs Following · 2026-05-15 Cached

Anyscale introduces a new Agent Skill for LLM post-training that automatically selects the optimal fine-tuning method (SFT, DPO, GRPO, etc.) and generates ready-to-launch configs, helping avoid wasted GPU runs.

0 favorites 0 likes
#dpo

Enhancing Multilingual Counterfactual Generation through Alignment-as-Preference Optimization

arXiv cs.CL · 2026-05-13 Cached

The paper introduces Macro, a preference alignment framework using DPO to improve the validity and minimality of self-generated counterfactual explanations across multiple languages.

0 favorites 0 likes
#dpo

talkie-lm/talkie-1930-13b-it

Hugging Face Models Trending · 2026-04-20 Cached

Talkie-1930-13b-it is a 13B parameter instruction-tuned language model trained on pre-1931 text and fine-tuned using reinforcement learning with DPO.

0 favorites 0 likes
#dpo

Where does output diversity collapse in post-training?

arXiv cs.CL · 2026-04-20 Cached

This paper investigates where and why output diversity collapses during post-training of language models, analyzing three OLMo 3 lineages (Think, Instruct, RL-Zero) across multiple tasks and metrics. The authors find that diversity collapse is primarily determined by training data composition and embedded in model weights during training, not addressable at inference time alone.

0 favorites 0 likes
#dpo

GroupDPO: Memory efficient Group-wise Direct Preference Optimization

arXiv cs.CL · 2026-04-20 Cached

GroupDPO introduces a memory-efficient algorithm for group-wise direct preference optimization that leverages multiple candidate responses per prompt while reducing peak memory usage through decoupled backpropagation. The method demonstrates consistent improvements over standard DPO across offline and online alignment settings.

0 favorites 0 likes
#dpo

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Papers with Code Trending · 2024-01-05 Cached

DeepSeek LLM is an open-source language model project that develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback