Understanding Reasoning from Pretraining to Post-Training

Hugging Face Daily Papers Papers

Summary

This paper studies how pretraining choices (model size, data) affect returns from RL post-training on reasoning tasks, using chess as a controlled testbed. It finds that post-RL performance is well-predicted by pretraining loss and that RL amplifies correct moves on easy puzzles while surfacing correct moves on hard puzzles, with findings transferring to math domains.

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:41 AM

Paper page - Understanding Reasoning from Pretraining to Post-Training

Source: https://huggingface.co/papers/2607.16097

Abstract

Reinforcementlearning(RL)hasbecomecentraltoimprovinglargelanguagemodels(LLMs)oncomplexreasoningtasks,yetRLpost-trainingislargelystudiedinisolationfromthepretrainingthatprecedesit.Asaresult,twobasicquestionsremainopen:(1)howdopretrainingchoices(modelsize,data)shapethereturnstoRLcompute,and(2)whatdoesRLactuallydotothemodel?ThesequestionsaredifficulttostudyinthestandardLLMsetting:pretrainingcorporaarevastanduncontrolled,makingithardtoattributebehaviorstopretrainingversusRL,andsystematiccomputesweepsacrossbothstagesareprohibitivelyexpensive.Toaddressthesechallenges,weusechessasacontrolledtestbedforstudyingreasoningacrossthefullpretraining-to-post-trainingpipeline.WefollowthestandardLLMtrainingpipelinebypretraininglanguagemodelsfrom5Mto1Bparametersonhumanchessgames,supervisedfine-tuningonsyntheticreasoningtraces,andrunningRLonchesspuzzleswithverifiablerewards.Usingthisframework,wefindthatthepost-RLperformanceatgivenRLcomputeleveliswell-predictedfromthepretrainingloss,andslopeoftheRLrewardcurvesimprovesapproximatelylinearlywiththepretrainingtokens.Beyondscaling,wefindthatRLdoesnotsimplysharpentheSFTpolicy:oneasypuzzlesitamplifiescorrectmovestheSFTpolicyalreadypreferred,whileonhardpuzzlesitsurfacescorrectmovesthatwerenearlyabsentunderSFT.Wefurthertestwhetherourfindingstransferbeyondchessbytraininga1Blanguagemodelonmath-domaintext,wherethesamepredictivepatternemerges:longer-pretrainedcheckpointsreachhigherpost-RLperformanceandimprovefasterunderRL.Insum,weprovideaquantitativeaccountofthepretraining-to-RLinterfaceandacontrolledtestbedforstudyingthescienceofreasoningacrossthefullpretraining-to-post-trainingpipeline.

View arXiv pageView PDFGitHub1Add to collection

Get this paper in your agent:

hf papers read 2607\.16097

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper2

#### pavelslab-nyu/Chess-SFT-Models Text Generation• Updatedabout 7 hours ago #### pavelslab-nyu/Chess-Pretrain-Models Text Generation• Updatedabout 8 hours ago

Datasets citing this paper2

#### pavelslab-nyu/pretrain_v1_54B Updatedabout 8 hours ago • 1.01k • 1 #### pavelslab-nyu/chess_puzzle_benchmark Viewer• Updatedabout 7 hours ago • 2.36k • 387

Spaces citing this paper1

Collections including this paper2

Similar Articles

Understanding Reasoning from Pretraining to Post-Training

arXiv cs.CL

This paper investigates the relationship between pretraining and reinforcement learning (RL) post-training for large language models using chess as a controlled testbed. It establishes a scaling law connecting pretraining loss to post-RL performance and shows that RL amplifies correct moves on easy puzzles while surfacing new correct moves on hard ones.

How Post-Training Shapes Biological Reasoning Models

Hugging Face Daily Papers

This paper investigates how post-training stages such as continued pre-training, supervised fine-tuning, and reinforcement learning affect generalization in biological reasoning models, finding that these stages have distinct impacts on in-domain and out-of-domain performance.

Revisiting Complete Reasoning Traces for Post-Training

Hugging Face Daily Papers

This paper finds that large language models can gain reasoning improvements from truncated reasoning trajectories rather than full ones during post-training, reducing redundancy while benefiting methods like supervised fine-tuning and reinforcement learning.