Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models

arXiv cs.CL Papers

Summary

The paper proposes a mean-field framework to model chain-of-thought reasoning in LLMs as a guided discovery process on a clue graph, deriving an ODE for the fraction of discovered clues and validating it experimentally.

arXiv:2608.05152v1 Announce Type: new Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization. In this study, we introduce a framework that seeks statistical regularities and theoretical interpretations in LLM reasoning without simplifying the model architecture or making analogies to existing physical systems. We formulate LLM reasoning as a guided discovery process on a clue graph, and derive a one-dimensional ordinary differential equation for the fraction of discovered clues using the mean-field approximation. Experimentally, clue tokens are identified using the normalized surprisal of a student LLM on the outputs of a teacher LLM, and statistical regularities are obtained by averaging over many reasoning chains of thought. Our experiments show that the resulting statistical regularities are reproducible within the same dataset and can be fitted by the solving the proposed theoretical equation.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:48 AM

# Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
Source: [https://arxiv.org/html/2608.05152](https://arxiv.org/html/2608.05152)
aihao\.phys@gmail\.com\]

###### Abstract

Large language models \(LLMs\) with chain\-of\-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization\. In this study, we introduce a framework that seeks statistical regularities and theoretical interpretations in LLM reasoning without simplifying the model architecture or making analogies to existing physical systems\. We formulate LLM reasoning as a guided discovery process on a clue graph, and derive a one\-dimensional ordinary differential equation for the fraction of discovered clues using the mean\-field approximation\. Experimentally, clue tokens are identified using the normalized surprisal of a student LLM on the outputs of a teacher LLM, and statistical regularities are obtained by averaging over many reasoning chains of thought\. Our experiments show that the resulting statistical regularities are reproducible within the same dataset and can be fitted by the solving the proposed theoretical equation\.

Introduction–Chain\-of\-thought\-based large language models \(LLMs\) have recently demonstrated strong reasoning capabilities, with performance approaching the human level on selected tasks such as mathematical reasoning\[[39](https://arxiv.org/html/2608.05152#bib.bib1),[19](https://arxiv.org/html/2608.05152#bib.bib2),[37](https://arxiv.org/html/2608.05152#bib.bib3),[13](https://arxiv.org/html/2608.05152#bib.bib4)\], competitive programming\[[13](https://arxiv.org/html/2608.05152#bib.bib4),[10](https://arxiv.org/html/2608.05152#bib.bib5)\], medical question answering\[[28](https://arxiv.org/html/2608.05152#bib.bib6),[29](https://arxiv.org/html/2608.05152#bib.bib7)\], and professional or academic examinations\[[8](https://arxiv.org/html/2608.05152#bib.bib8),[38](https://arxiv.org/html/2608.05152#bib.bib9),[27](https://arxiv.org/html/2608.05152#bib.bib10)\]\. Understanding the internal mechanisms or statistical regularities underlying LLMs\[[22](https://arxiv.org/html/2608.05152#bib.bib11),[36](https://arxiv.org/html/2608.05152#bib.bib12),[17](https://arxiv.org/html/2608.05152#bib.bib13),[25](https://arxiv.org/html/2608.05152#bib.bib14),[4](https://arxiv.org/html/2608.05152#bib.bib15),[42](https://arxiv.org/html/2608.05152#bib.bib16)\]may provide useful guidance for model optimization, efficiency improvement, and cost reduction\. It has therefore become an important direction in theoretical research\.

Existing theoretical studies have explored deep learning systems such as LLMs from physics\-inspired perspectives\. One common approach is to start from first principles\. Unlike most physical systems, the internal architecture of a deep learning model is fully known, allowing researchers to formulate interpretable theories based on the elementary architecture of neural networks\[[3](https://arxiv.org/html/2608.05152#bib.bib17),[24](https://arxiv.org/html/2608.05152#bib.bib18),[5](https://arxiv.org/html/2608.05152#bib.bib19),[23](https://arxiv.org/html/2608.05152#bib.bib20),[35](https://arxiv.org/html/2608.05152#bib.bib21)\]\. Nevertheless, the architectures of modern LLMs are already highly complex, making it difficult for such theories to account for all of their specialized structures\. In practice, omitting architectural components such as residual connections\[[15](https://arxiv.org/html/2608.05152#bib.bib22)\], multi\-head attention\[[34](https://arxiv.org/html/2608.05152#bib.bib23)\], or MoE\[[11](https://arxiv.org/html/2608.05152#bib.bib24)\]can substantially reduce LLM performance, thereby limiting the universality of these theories based on simplified architectures\. Another approach is to compare LLMs with well\-studied physical systems, such as spin glasses\[[21](https://arxiv.org/html/2608.05152#bib.bib25)\], systems near phase transitions\[[2](https://arxiv.org/html/2608.05152#bib.bib26),[9](https://arxiv.org/html/2608.05152#bib.bib27)\], and BKT\-type statistical\-mechanical models\[[32](https://arxiv.org/html/2608.05152#bib.bib28)\]\. However, these physical systems were not built to describe LLMs, and the physical concepts imported into LLM studies are often difficult to define rigorously\. This weakens the rigor of such theoretical approaches\.

Rather than decomposing LLMs from first principles, this work treats them at the level of collective behavior and seek statistical regularities in their reasoning processes\. Instead of comparing LLMs with existing physical systems, we seek to construct a theoretical model tailored to LLMs\. Using a mean\-field approximation, this theory yields an ordinary differential equation that captures the statistical patterns observed in experiments\.

Theoretically, we formulate chain\-of\-thought reasoning in LLMs as a clue discovery process\. We assume that solving a problem requires a set of clues, each of which is either known or unknown\. During reasoning, the LLM progressively turns unknown clues into known ones\. The clues form a directed acyclic clue graph, in which known upstream clues facilitate the discovery of downstream clues\. Within a mean\-field approximation, we derive an ordinary differential equation for the time evolution of the fraction of known clues\. We refer to this equation as the guided discovery equation\. Its solution gives the time\-dependent clue discovery rate\.

We validate the theory experimentally by using a stronger teacher LLM to generate chains of thought and a student LLM to scan them\. The observable for clues is the token\-level normalized surprisal, which measures the capability gap between the two LLMs and therefore implicitly identifies key clues for solving the problem\. By collecting a large number of chains of thought and averaging the resulting clue discovery rate curves, we obtain statistical patterns that demonstrate two central results: 1\) under the same dataset and the same LLMs, the clue discovery rate follows reproducible statistical regularities; and 2\) the averaged clue discovery rate can be partially fitted by the guided discovery equation\. These findings indicate that LLM reasoning can be treated as a physical system with statistical regularities, and that such regularities can, in certain regimes, be captured by a simple equation\.

Chain\-of\-thought reasoning as clue discovery–We formulate chain\-of\-thought reasoning in LLMs as guided clue discovery on a clue graph, as shown in Fig\.[1](https://arxiv.org/html/2608.05152#S0.F1)\(a\)\. Solving a problem is assumed to requires to knowNNclues, whose dependencies form a directed acyclic graph\. The LLM acts as a discovery agent on this graph, progressively converting unknown clues into known ones\. The time variable of this process corresponds to the token position in the reasoning chain\. Each clue has a binary stateXi∈\{0,1\}X\_\{i\}\\in\\\{0,1\\\}, where0denotes unknown and11denotes known\. For each unknown clueii, its state may switch from0to11within a small intervald​tdtwith a certain probability\. The discovery probability depends on how many of its nearest upstream neighbors are already known\. Since the upstream neighbors are the clues that can directly support the inference of clueii, the guided discovery rate of clueiiincreases with the number of such upstream clues that have already been discovered\.

![Refer to caption](https://arxiv.org/html/2608.05152v1/x1.png)
Figure 1:\(a\) Schematic of clue discovery\. A set of clues, shown as circles, forms a directed acyclic graph\. Orange\-filled and empty circles denote known and unknown clues, respectively\. Known clues can guide the discovery of unknown ones, but only those within the attention window are attended to\. The attention window is marked by the blue dashed circle\. \(b\) Discovery rate for a single chain of thought\. Tokens with normalized surprisal above a threshold are identified as clue tokens\. The vertical lines show their normalized surprisal, and the blue solid line shows the resulting discovery rate curve\. \(c\) Average discovery rate curve over many chains of thought, revealing a statistical regularity\.In addition, although an LLM can in principle access the full context when generating each new token, its attention mechanism aggregates contextual information through weighted averages over attention scores\[[33](https://arxiv.org/html/2608.05152#bib.bib29)\]\. As a result, the LLM may not be able to fully, uniformly, and comprehensively use all known clues\. We therefore introduce an attention window into our theory\. At each time, the agent attends only to a finite subset of the known clues, and only those attended clues can contribute to the guided discovery of new clues\. For a typical unknown clueii, we assume that each known nearest upstream clue enters the attention window with probabilityρ\\rho\. The introduction ofρ\\rhoimplicitly assumes that the LLM has the ability to select relevant clues from a large number of known clues\. In contrast, if the LLM agent can only randomly sample from the known clues, the selection process reduces to a hypergeometric distribution\. Specifically, suppose that there areMMknown clues in total, among whichccare nearest upstream neighbors of clueii\. If the agent randomly drawskkclues from theMMknown clues into the attention window, then the guided discovery rate of clueiidepends on the numberrrof selected clues that are also nearest upstream neighbors of clueii\.

Mean\-field theory–For an individual problem, the clue discovery graph may contain substantial randomness, and the discovery trajectory of a single chain of thought may also be highly uncertain\. Nevertheless, when a large number of samples are collected and averaged, the resulting behavior can exhibit certain regularities\. We therefore introduce a mean\-field approximation, a standard approach in statistical physics for reducing many\-body interactions to an effective averaged field\[[7](https://arxiv.org/html/2608.05152#bib.bib30)\]\. First, we assume that every clue has the same number of nearest upstream clues, denoted bydd\. Second, ifMMof theNNclues are known at a certain time, then for a representative unknown clueii, each of its nearest upstream clues is known with probabilityM/NM/N\. Third, as discussed earlier, a known nearest upstream clue ofiiis attended to with probabilityρ\\rho\. As a result, under the mean\-field approximation, the guided discovery process can be formulated through a binomial distribution, i\.e\., the guided discovery rate of clueiiis determined by the numberrrof itsddnearest upstream clues that are both known and selected into the attention window, and each upstream clue satisfies this condition with probabilityρ​m\\rho m\.

In addition to this guided discovery effect, we further account for the accidental discovery of clues, which represents the LLM’s prior or intrinsic knowledge of the corresponding dataset\. Consider that at most one new clue can be discovered at any moment, the clue discovery process can be reduced to an one\-dimensional ordinary differential equation,

d​md​t\\displaystyle\\frac\{dm\}\{dt\}=\(1−m\)​\[ϵ\+β​G​\(m\)\],\\displaystyle=\(1\-m\)\\big\[\\epsilon\+\\beta G\(m\)\\big\],\(1\)G​\(m\)\\displaystyle G\(m\)=∑r=0df​\(r\)​\(dr\)​\(ρ​m\)r​\(1−ρ​m\)d−r,\\displaystyle=\\sum\_\{r=0\}^\{d\}f\(r\)\\binom\{d\}\{r\}\(\\rho m\)^\{r\}\(1\-\\rho m\)^\{d\-r\},wheremmis the dependent variable and represents the proportion of known clues, namelym=M/Nm=M/N\.ϵ\\epsilondenotes the coefficient of accidental discovery, whileβ\\betadenotes the coefficient of guided discovery\. The functionf​\(r\)f\(r\)is the guided discovery kernel\. It describes the dependence of the guided discovery probability of an unknown clueiion the numberrrof upstream clues that are both known and included in the attention window\. It is typically chosen as a monotonically increasing function\. In the subsequent experiments, we takef​\(r\)=\(r/d\)0\.1f\(r\)=\(r/d\)^\{0\.1\}, which grows rapidly whenrris small and then gradually approaches saturation\.

Observable: normalized surprisal–We next introduce the experimental setup for validating the theory, together with the main observable, normalized surprisal\. In the experiment, a strong teacher LLM is asked to repeatedly answer multiple questions from a dataset, thereby producing a large collection of chains of thought\. We then use a weaker student LLM to scan the chains of thought generated by the teacher LLM, and identify the tokens that are difficult for the student LLM to predict\. Since the teacher LLM generally performs much better than the student LLM, such hard\-to\-predict tokens in the teacher LLM’s reasoning process can be viewed as the crucial elements that guide the LLM toward the better answers\. Accordingly, they correspond to the clues in our theory, and we call them clue tokens\. Introducing both a teacher LLM and a student LLM is necessary for obtaining surprisals in our setting, and the framework is also broadly used in tasks such as knowledge distillation\[[16](https://arxiv.org/html/2608.05152#bib.bib31),[18](https://arxiv.org/html/2608.05152#bib.bib32)\]and weak\-to\-strong generalization\[[6](https://arxiv.org/html/2608.05152#bib.bib33)\]\.

The surprisalsts\_\{t\}reflects the degree to which the student LLM is surprised by thett\-th token in the teacher LLM’s chain of thought\[[26](https://arxiv.org/html/2608.05152#bib.bib34)\]\. It is defined as

st=−log⁡pθ​\(xt∣𝒞,x<t\),s\_\{t\}=\-\\log p\_\{\\theta\}\\left\(x\_\{t\}\\mid\\mathcal\{C\},x\_\{<t\}\\right\),\(2\)wherextx\_\{t\}is thett\-th token in the teacher LLM’s chain of thought,x<tx\_\{<t\}denotes all previous tokens before positiontt,𝒞\\mathcal\{C\}denotes the prompt and problem context, andpθp\_\{\\theta\}is the next\-token probability distribution predicted by the student LLM\. However, surprisal itself does not fully reflect the capability gap between the two LLMs\. If a token receives a high surprisal under the student LLM, this does not necessarily mean that the student model is incapable of generating that token\. It may simply be that, due to the sentence structure, the token at that position is inherently uncertain\. For example, there are often many possible choices for the first token of a sentence, so its surprisal is naturally high\. To address this issue, we perform z\-score normalization on surprisal\. Using the student LLM’s forward predictive entropy and predictive varentropy\[[20](https://arxiv.org/html/2608.05152#bib.bib41),[1](https://arxiv.org/html/2608.05152#bib.bib42)\], we obtain the normalized surprisalztz\_\{t\},

zt=st−HtVt,z\_\{t\}=\\frac\{s\_\{t\}\-H\_\{t\}\}\{\\sqrt\{V\_\{t\}\}\},\(3\)where

Ht=−∑v∈𝒱pθ​\(v∣𝒞,x<t\)​log⁡pθ​\(v∣𝒞,x<t\),H\_\{t\}=\-\\sum\_\{v\\in\\mathcal\{V\}\}p\_\{\\theta\}\\left\(v\\mid\\mathcal\{C\},x\_\{<t\}\\right\)\\log p\_\{\\theta\}\\left\(v\\mid\\mathcal\{C\},x\_\{<t\}\\right\),\(4\)and

Vt=∑v∈𝒱pθ​\(v∣𝒞,x<t\)​\[−log⁡pθ​\(v∣𝒞,x<t\)−Ht\]2\.V\_\{t\}=\\sum\_\{v\\in\\mathcal\{V\}\}p\_\{\\theta\}\\left\(v\\mid\\mathcal\{C\},x\_\{<t\}\\right\)\\left\[\-\\log p\_\{\\theta\}\\left\(v\\mid\\mathcal\{C\},x\_\{<t\}\\right\)\-H\_\{t\}\\right\]^\{2\}\.\(5\)Here,HtH\_\{t\}is the forward prediction entropy of the student model, which is also the expectation of surprisal\. It reflects the semantic uncertainty of the student model during prediction\. The predictive varentropyVtV\_\{t\}is the variance of surprisal, and reflects how concentrated the predictive probability distribution is\.

Finally, we obtain the statistical regularities of clues from normalized surprisal\. For each individual chain of thought produced in a specific answer, we record the normalized surprisal at every token position\. A token with a larger normalized surprisal is more likely to serve as a clue token\. For practical simplicity, we introduce a fixed thresholdλ\\lambdato determine whether each token is a clue token\. Although this thresholding procedure can be rough for a single chain of thought, it is reasonable at the statistical level\. In this way, normalized surprisal is transformed into a binary variable,z^t=Θ​\(zt−λ\)\\hat\{z\}\_\{t\}=\\Theta\(z\_\{t\}\-\\lambda\), wherez^t\\hat\{z\}\_\{t\}equals11if the token is a clue token and equals0otherwise\. In the above guided discovery theory, the core observable is the clue discovery rate, defined as the number of clues discovered within a short time interval\. In the experiment, this quantity corresponds to the number of clue tokens in a local neighborhood around a given token\. For smoothness, we use a Gaussian kernel to count the number of clues within such a local neighborhood,

z~t=∑τ=1TKσ​\(t−τ\)​z^τ∑τ=1TKσ​\(t−τ\),Kσ​\(t−τ\)=exp⁡\[−\(t−τ\)22​σ2\],\\tilde\{z\}\_\{t\}=\\frac\{\\sum\_\{\\tau=1\}^\{T\}K\_\{\\sigma\}\(t\-\\tau\)\\hat\{z\}\_\{\\tau\}\}\{\\sum\_\{\\tau=1\}^\{T\}K\_\{\\sigma\}\(t\-\\tau\)\},\\quad K\_\{\\sigma\}\(t\-\\tau\)=\\exp\\left\[\-\\frac\{\(t\-\\tau\)^\{2\}\}\{2\\sigma^\{2\}\}\\right\],\(6\)whereTTis the length of the chain of thought, andσ\\sigmais the bandwidth of the Gaussian smoothing kernel\. At this point, we have obtained the clue discovery rate curve for a single chain of thought, as shown in Fig\.[1](https://arxiv.org/html/2608.05152#S0.F1)\(b\)\. The clue discovery rate curve of a single chain of thought appears random, but the average of a large number of such curves exhibits statistical regularities, as shown in Fig\.[1](https://arxiv.org/html/2608.05152#S0.F1)\(c\)\. Specifically, we normalize the horizontal coordinate of each clue discovery rate curve to the interval from0to11according to the token position, and then average the normalized curves across many chains of thought\. The resulting averaged curve gives the clue discovery rate in the statistical sense, and corresponds to the clue discovery rated​m/d​tdm/dtobtained by solving the clue discovery equation Eq\. \([1](https://arxiv.org/html/2608.05152#S0.E1)\) under the mean\-field approximation\.

Experimental results–In our experiments, Qwen3\-Max serves as the teacher LLM and Qwen3\-8B serves as the student LLM\[[40](https://arxiv.org/html/2608.05152#bib.bib35)\]\. Experiments are performed on four textual reasoning datasets, namely MuSR\[[31](https://arxiv.org/html/2608.05152#bib.bib37)\], CLUTRR\[[30](https://arxiv.org/html/2608.05152#bib.bib38)\], StrategyQA\[[12](https://arxiv.org/html/2608.05152#bib.bib39)\], and FOLIO\[[14](https://arxiv.org/html/2608.05152#bib.bib40)\]\. For each dataset, we choose100100questions and sample1010independent chains of thought for each question, producing a total of1 0001\\ 000chains of thought per dataset\. The experiments have two main objectives\. The first is to verify that, with normalized surprisal as the observable, statistical regularities can be observed\. The second is to adjust the model parameters and show that the mean\-field clue discovery equation Eq\. \([1](https://arxiv.org/html/2608.05152#S0.E1)\) is able to fit the empirical statistical curves\.

![Refer to caption](https://arxiv.org/html/2608.05152v1/x2.png)
Figure 2:Two\-fold experiments demonstrate the existence of the statistical regularities\. The four subplots correspond to the four datasets, namely MuSR, CLUTRR, StrategyQA, and FOLIO\. For each dataset, the generated chains of thought are divided by question into two non\-overlapping subsets\. The clue discovery rates are then computed from normalized surprisal for the two folds, shown in orange and green\. The results from the two folds agree well with each other\.![Refer to caption](https://arxiv.org/html/2608.05152v1/x3.png)
Figure 3:Two\-fold experiments using GLM\-4\.7 as the teacher\. The student LLM is kept to be Qwen3\-8B\. The subplots indicate that the statistical regularities hold for different teacher LLMs, but the resulting discovery rate curves are different between Qwen3\-Max and GLM\-4\.7\.To demonstrate that the surprisal\-based clue discovery ratez~t\\tilde\{z\}\_\{t\}has statistical regularities, we divide the100100problems in each dataset into two non\-overlapping parts, each containing5050problems, denoted as fold\-1 and fold\-2\. We then compute the averagedz~t\\tilde\{z\}\_\{t\}curves for the two folds separately, as shown in Fig\.[2](https://arxiv.org/html/2608.05152#S0.F2)\. As a result, thez~t\\tilde\{z\}\_\{t\}curves from the two folds exhibit strong consistency within each dataset\. This indicates that the clue discovery ratez~t\\tilde\{z\}\_\{t\}indeed contains certain statistical regularities, rather than being a randomly varying curve\.

![Refer to caption](https://arxiv.org/html/2608.05152v1/x4.png)
Figure 4:Fitting the theory to the experiments\. The four subplots correspond to the four datasets\. In each subplot, the blue solid line denotes the experimentally measured average clue discovery rate, and the blue shaded region spans the first to the third quartile\. The red dashed line shows the theoretical fit\. The theory agrees well with the experimental results in the first half of the reasoning process\.![Refer to caption](https://arxiv.org/html/2608.05152v1/x5.png)
Figure 5:Fitting the theory to the experiments\. GLM\-4\.7 serves as the teacher LLM, and Qwen3\-8B is the student LLM\. The theory also agrees well with the experimental results in the first half of the curve\.Next, to validate the theoretical modeling, we tune the hyperparameters of the guided discovery equation so that the solvedd​m/d​tdm/dtfits the experimental clue discovery rate curve of each dataset\. Since a small number of clue tokens may be provided by the question or the prompt, the initial valuem​\(0\)m\(0\)of the equation is also treated as a tunable hyperparameter, and it satisfies0<m​\(0\)≪10<m\(0\)\\ll 1\. In addition, we apply a linear transformation to thed​m/d​tdm/dtobtained by solving the equation in order to fit the experimental result, namelyz~t∼a​\[d​m​\(t\)/d​t\]\+b,\\tilde\{z\}\_\{t\}\\sim a\[dm\(t\)/dt\]\+b,wherem​\(t\)m\(t\)is the solution of the equation\. This is because the theoretical equation gives the discovery rate of the fractionmmof known clues, whereas the experimental observable is the discovery rate of the number of clue tokens\. Therefore, we multiplyd​m/d​tdm/dtby a coefficientaato adjust the scale\. Moreover, because the student LLM is still less capable than the teacher LLM in extracting existing clues from the context, even already discovered clues may still produce high surprisal under the student model with some probability\. We therefore introduce a bias termbbto reduce the influence of this mechanism\. The fitting results on the four datasets are shown in Fig\.[4](https://arxiv.org/html/2608.05152#S0.F4)\. It can be seen that the theory proposed in this study can fit the experimental results well in the first half of the evolution, which demonstrates the validity of the theory within a certain range\.

In addition, we also conduct experiments using GLM\-4\.7\[[41](https://arxiv.org/html/2608.05152#bib.bib36)\]as the teacher LLM, and the corresponding results are reported in the Fig\.[3](https://arxiv.org/html/2608.05152#S0.F3)and Fig\.[5](https://arxiv.org/html/2608.05152#S0.F5)\.

Conclusion and limitations–In summary, this study focuses on a theoretical interpretation of LLM chain\-of\-thought reasoning\. Our theory neither breaks down the key components of LLMs to obtain an interpretable reduced model, nor explains LLMs by comparison with well\-studied physical systems\. Instead, this study offers a new perspective\. It seeks to identify statistical regularities through experimental design, metric transformation, and averaging over a large number of samples\. We then regard these statistical regularities as the solution of an underlying differential equation, construct a theoretical model, derive the differential equation through a mean\-field approximation, and tune its hyperparameters so that the theory agrees with the experiments\.

Specifically, this study regards chain\-of\-thought reasoning in LLMs as the guided discovery of clues\. We introduce the clue graph and the attention window, and derive the clue discovery equation based on a mean\-field approximation\. In the experiments, we identify key clues through the capability gap between teacher and student LLMs, use normalized surprisal to characterize clue tokens, and then obtain the discovery rate curve statistically\. The experimental results demonstrate that the curves are consistent across samples from the same dataset, confirming the existence of statistical regularities\. With appropriate hyperparameters, the theoretical equation can fit the experimental results within a certain regime, indicating that the proposed theory at least partially captures the physical regularities underlying the real system\.

As an early attempt, this study still has several limitations\. First, both the theoretical modeling and the experimental setup involve many hyperparameters\. Although each hyperparameter is chosen in a reasonable way, this reduces the simplicity and generality of the theory\. Second, the experimental results indicate that the statistical regularities of the clue discovery rate do not exhibit consistency across different datasets and models\. This may arise from differences in reasoning paths across datasets and differences in the internal knowledge of the LLMs\. Finally, the experiments depend on two LLMs, a teacher and a student, which may introduce further uncontrolled factors affecting the universality of the regularities\. In future research, we will look for experimental settings and observables that rely only on a single LLM, while being more interpretable and more universal, and we will develop theories to explain them\.

## References

- \[1\]F\. Ahmed, Y\. J\. Ong, and C\. DeLuca\(2026\)LogitScope: a framework for analyzing llm uncertainty through information metrics\.arXiv preprint arXiv:2603\.24929\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p12.9)\.
- \[2\]T\. Aoyama and E\. Wilcox\(2025\-07\)Language models grow less humanlike beyond phase transition\.InACL 2025,Vienna, Austria,pp\. 24938–24958\.External Links:[Link](https://aclanthology.org/2025.acl-long.1214/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1214),ISBN 979\-8\-89176\-251\-0Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[3\]Y\. Bahri, J\. Kadmon, J\. Pennington, S\. S\. Schoenholz, J\. Sohl\-Dickstein, and S\. Ganguli\(2020\)Statistical mechanics of deep learning\.Annual review of condensed matter physics11\(1\),pp\. 501–528\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[4\]M\. Belkin, D\. Hsu, S\. Ma, and S\. Mandal\(2019\)Reconciling modern machine\-learning practice and the classical bias–variance trade\-off\.Proceedings of the National Academy of Sciences116\(32\),pp\. 15849–15854\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[5\]A\. Bhaskar, A\. Wettig, D\. Friedman, and D\. Chen\(2024\)Finding transformer circuits with edge pruning\.InNeurIPS 2024,Vol\.37,pp\. 18506–18534\.External Links:[Document](https://dx.doi.org/10.52202/079017-0587),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/20fdaf67581e6d7157376d1ed584040a-Paper-Conference.pdf)Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[6\]C\. Burns, P\. Izmailov, J\. H\. Kirchner, B\. Baker, L\. Gao, L\. Aschenbrenner, Y\. Chen, A\. Ecoffet, M\. Joglekar, J\. Leike,et al\.\(2023\)Weak\-to\-strong generalization: eliciting strong capabilities with weak supervision\.arXiv preprint arXiv:2312\.09390\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p11.1)\.
- \[7\]P\. M\. Chaikin, T\. C\. Lubensky, and T\. A\. Witten\(1995\)Principles of condensed matter physics\.Vol\.10,Cambridge university press Cambridge\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p9.11)\.
- \[8\]H\. W\. Chung, L\. Hou, S\. Longpre, B\. Zoph, Y\. Tay, W\. Fedus, Y\. Li, X\. Wang, M\. Dehghani, S\. Brahma,et al\.\(2024\)Scaling instruction\-finetuned language models\.Journal of Machine Learning Research25\(70\),pp\. 1–53\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[9\]H\. Cui, F\. Behrens, F\. Krzakala, and L\. Zdeborová\(2024\)A phase transition between positional and semantic learning in a solvable model of dot\-product attention\.Advances in Neural Information Processing Systems37,pp\. 36342–36389\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[10\]A\. El\-Kishky, A\. Wei, A\. Saraiva, B\. Minaiev, D\. Selsam, D\. Dohan, F\. Song, H\. Lightman, I\. Clavera, J\. Pachocki,et al\.\(2025\)Competitive programming with large reasoning models\.arXiv preprint arXiv:2502\.06807\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[11\]W\. Fedus, B\. Zoph, and N\. Shazeer\(2022\)Switch transformers: scaling to trillion parameter models with simple and efficient sparsity\.Journal of Machine Learning Research23\(120\),pp\. 1–39\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[12\]M\. Geva, D\. Khashabi, E\. Segal, T\. Khot, D\. Roth, and J\. Berant\(2021\)Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies\.Transactions of the Association for Computational Linguistics9,pp\. 346–361\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p14.3)\.
- \[13\]D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[14\]S\. Han, H\. Schoelkopf, Y\. Zhao, Z\. Qi, M\. Riddell, W\. Zhou, J\. Coady, D\. Peng, Y\. Qiao, L\. Benson,et al\.\(2024\)Folio: natural language reasoning with first\-order logic\.InEMNLP 2024,pp\. 22017–22031\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p14.3)\.
- \[15\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2016\)Deep residual learning for image recognition\.InCVPR 2016,Vol\.,pp\. 770–778\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2016.90)Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[16\]C\. Hsieh, C\. Li, C\. Yeh, H\. Nakhost, Y\. Fujii, A\. Ratner, R\. Krishna, C\. Lee, and T\. Pfister\(2023\)Distilling step\-by\-step\! outperforming larger language models with less training data and smaller model sizes\.InFindings of ACL 2023,pp\. 8003–8017\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p11.1)\.
- \[17\]J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei\(2020\)Scaling laws for neural language models\.arXiv preprint arXiv:2001\.08361\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[18\]Y\. Kim and A\. M\. Rush\(2016\)Sequence\-level knowledge distillation\.InEMNLP 2016,pp\. 1317–1327\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p11.1)\.
- \[19\]A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo,et al\.\(2022\)Solving quantitative reasoning problems with language models\.Advances in neural information processing systems35,pp\. 3843–3857\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[20\]X\. Li, E\. Callanan, X\. Zhu, M\. Sibue, A\. Papadimitriou, M\. Mahfouz, Z\. Ma, and X\. Liu\(2025\)Entropy\-aware branching for improved mathematical reasoning\.arXiv e\-prints,pp\. arXiv–2503\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p12.9)\.
- \[21\]Y\. Li, R\. Bai, and H\. Huang\(2025\)Spin\-glass model of in\-context learning\.Physical Review E112\(1\),pp\. L013301\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[22\]K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov\(2022\)Locating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[23\]C\. Olsson, N\. Elhage, N\. Nanda, N\. Joseph, N\. DasSarma, T\. Henighan, B\. Mann, A\. Askell, Y\. Bai, A\. Chen,et al\.\(2022\)In\-context learning and induction heads\.arXiv preprint arXiv:2209\.11895\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[24\]D\. A\. Roberts, S\. Yaida, and B\. Hanin\(2022\)The principles of deep learning theory\.Vol\.46,Cambridge University Press Cambridge, MA, USA\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[25\]R\. Schaeffer, B\. Miranda, and S\. Koyejo\(2023\)Are emergent abilities of large language models a mirage?\.Advances in neural information processing systems36,pp\. 55565–55581\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[26\]C\. E\. Shannon\(1948\)A mathematical theory of communication\.The Bell system technical journal27\(3\),pp\. 379–423\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p12.2)\.
- \[27\]P\. Shetty, A\. Upadhayaya, P\. M\. Shah, S\. Jagabathula, S\. Nayak, and A\. J\. Fee\(2025\)Advanced financial reasoning at scale: a comprehensive evaluation of large language models on cfa level iii\.arXiv preprint arXiv:2507\.02954\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[28\]K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl,et al\.\(2023\)Large language models encode clinical knowledge\.Nature620\(7972\),pp\. 172–180\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[29\]K\. Singhal, T\. Tu, J\. Gottweis, R\. Sayres, E\. Wulczyn, M\. Amin, L\. Hou, K\. Clark, S\. R\. Pfohl, H\. Cole\-Lewis,et al\.\(2025\)Toward expert\-level medical question answering with large language models\.Nature medicine31\(3\),pp\. 943–950\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[30\]K\. Sinha, S\. Sodhani, J\. Dong, J\. Pineau, and W\. L\. Hamilton\(2019\)CLUTRR: a diagnostic benchmark for inductive reasoning from text\.InEMNLP\-IJCNLP 2019,pp\. 4506–4515\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p14.3)\.
- \[31\]Z\. Sprague, X\. Ye, K\. Bostrom, S\. Chaudhuri, and G\. Durrett\(2024\)Musr: testing the limits of chain\-of\-thought with multistep soft reasoning\.InICLR,Vol\.2024,pp\. 14670–14728\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p14.3)\.
- \[32\]Y\. Toji, J\. Takahashi, V\. Roychowdhury, and H\. Miyahara\(2026\-01\)Berezinskii\-kosterlitz\-thouless transition in a context\-sensitive random language model\.Phys\. Rev\. E113,pp\. 015305\.External Links:[Document](https://dx.doi.org/10.1103/s7nf-bwzd),[Link](https://link.aps.org/doi/10.1103/s7nf-bwzd)Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[33\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.Advances in neural information processing systems30\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p8.11)\.
- \[34\]E\. Voita, D\. Talbot, F\. Moiseev, R\. Sennrich, and I\. Titov\(2019\-07\)Analyzing multi\-head self\-attention: specialized heads do the heavy lifting, the rest can be pruned\.InACL 2019,Florence, Italy,pp\. 5797–5808\.External Links:[Link](https://aclanthology.org/P19-1580/),[Document](https://dx.doi.org/10.18653/v1/P19-1580)Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[35\]J\. Von Oswald, E\. Niklasson, E\. Randazzo, J\. Sacramento, A\. Mordvintsev, A\. Zhmoginov, and M\. Vladymyrov\(2023\)Transformers learn in\-context by gradient descent\.InICML 2023,pp\. 35151–35174\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p3.1)\.
- \[36\]K\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt\(2022\)Interpretability in the wild: a circuit for indirect object identification in gpt\-2 small\.arXiv preprint arXiv:2211\.00593\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[37\]X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. Zhou\(2022\)Self\-consistency improves chain of thought reasoning in language models\.arXiv preprint arXiv:2203\.11171\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[38\]Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.\(2024\)Mmlu\-pro: a more robust and challenging multi\-task language understanding benchmark\.Advances in Neural Information Processing Systems37,pp\. 95266–95290\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[39\]J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.
- \[40\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p14.3)\.
- \[41\]A\. Zeng, X\. Lv, Q\. Zheng, Z\. Hou, B\. Chen, C\. Xie, C\. Wang, D\. Yin, H\. Zeng, J\. Zhang,et al\.\(2025\)Glm\-4\.5: agentic, reasoning, and coding \(arc\) foundation models\.arXiv preprint arXiv:2508\.06471\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p17.1)\.
- \[42\]B\. Žunkovič and E\. Ilievski\(2024\)Grokking phase transitions in learning local rules with gradient descent\.Journal of Machine Learning Research25\(199\),pp\. 1–52\.Cited by:[Mean\-Field Dynamics of Chain\-of\-Thought Reasoning in Large Language Models](https://arxiv.org/html/2608.05152#p2.1)\.

Similar Articles

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations

arXiv cs.CL

This paper presents a comprehensive empirical evaluation of how large language models handle corruptions in chain-of-thought reasoning steps, testing 13 models across 5 perturbation types (MathError, UnitConversion, Sycophancy, SkippedSteps, ExtraSteps) on mathematical reasoning tasks. The findings reveal heterogeneous vulnerability patterns with implications for deploying LLMs in multi-stage reasoning pipelines.

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts

arXiv cs.CL

This research paper from MediaTek and National Taiwan University challenges the assumption that reasoning chains must be dense and sequential, showing that models can extract answers from sparse, shuffled, and noisy reasoning traces. The findings suggest that answer extraction is robust and order-independent, potentially enabling more efficient, parallelized reasoning generation.

Reasoning emerges from constrained inference manifolds in large language models

arXiv cs.LG

This paper investigates reasoning in LLMs as an intrinsic dynamical process, finding that inference-time representations self-organize into low-dimensional manifolds. It proposes a label-free diagnostic based on internal dynamics to assess reasoning quality, suggesting that effective reasoning is governed by geometric and informational constraints.

Not All LLM Reasoning is Visible in the Chain-of-Thought

arXiv cs.CL

This paper demonstrates that frontier language models can perform 'invisible reasoning' using semantically irrelevant filler tokens, improving accuracy on synthetic reasoning tasks by up to 13 percentage points, which undermines the assumption that chain-of-thought monitoring captures all reasoning.