@Orange41324306: TailRL and TailSFT

X AI KOLs Timeline Papers

Summary

Sadhika Malladi proposes TailSFT, a lightweight and principled method to improve coverage and enhance post-RL performance, building on previous research that criticized xent SFT for preparing RL.

TailRL and TailSFT
Original Article
View Cached Full Text

Cached at: 09/13/26, 01:11 PM

TailRL and TailSFT

Sadhika Malladi (@SadhikaMalladi): RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https://t.co/uniFmpbDOn). Now, we propose TailSFT (https://t.co/75yejb8voI), a lightweight + principled way to directly improve coverage and get better post-RL perf.

Similar Articles

Tail-Likelihood Reinforcement Learning

arXiv cs.LG

The paper proposes Tail-Likelihood Reinforcement Learning (TailRL), an optimization method that focuses on the upper tails of reward distributions to improve policy performance in generative tasks, demonstrated across various applications.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

arXiv cs.LG

Harvard researchers challenge the standard LLM training pipeline by showing RL can be effectively applied during pre-training rather than only after SFT, finding that data composition matters more than model scale, and proposing parallel averaging of RL and SFT objectives that outperforms sequential approaches while preserving general capabilities.