Tag
This paper demonstrates that the final window of pretraining significantly influences a model's response to post-training alignment, even when SFT performance is identical, suggesting that checkpoint evaluation should include the last training data.