Tag
This paper presents a theoretical and empirical analysis showing that SFT suffers from task conflicts in multi-task LLM training while RL enables stable coexistence, proposing the Parallel-RL paradigm for efficient multi-task training.
This paper investigates why larger models outperform smaller ones, attributing it to reduced gradient interference and better resource allocation, allowing them to learn rare and complex tasks even with infinite data. Experiments on synthetic data and OLMo models verify that larger models avoid overwriting rare-task features due to weaker gradient updates for common tasks.