Tag
SMOPD is a loss-only stabilization method for multi-turn on-policy self-distillation that uses selective token-entropy masking to improve accuracy in dirty-history settings, demonstrating improvements with Qwen3 models.
Proposes EKSFT, a selective fine-tuning method for large language models that masks tokens with high entropy or high KL divergence from a reference model, preserving pre-trained distribution while injecting task knowledge. Experiments on mathematical reasoning benchmarks show it outperforms standard SFT and improves subsequent RL fine-tuning.