量子模型是否像LLMs一样扩展?
摘要
本文研究了在来自Rydberg atom arrays的量子测量数据上训练的transformers是否表现出类似于大型语言模型的神经缩放定律,发现缩放行为取决于数据的统计结构,尤其是在临界点附近。
arXiv:2609.20912v1 Announce Type: new
Abstract: In this work, we study the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data gathered from interacting Rydberg atom arrays. The quantum system is known to exhibit a finite-size remnant of a critical point as the laser detuning parameter is varied. We find that near the critical point the transformer loss as a function of training dataset size is well described by a power-law with a loss floor correction. However, away from criticality the quality of the power-law description is substantially reduced. We then compare the statistical structure of both Rydberg measurements and natural-language corpora using an entropy-normalised, finite sample corrected mutual information "two-point" function. We find that near-critical statistics of the two point functions are closest to those observed in natural-language, whilst other qubit configurations far from the critical point have two-point functions that decay more rapidly. This supports the hypothesis that multi-scale dependence contributes to stable neural scaling, and that scaling behaviour should be viewed as a property of the model-data pair.
查看缓存全文
缓存时间: 2026/09/21 09:13
# Do Quantum Models Scale Like LLMs?
Source: [https://arxiv.org/html/2609.20912](https://arxiv.org/html/2609.20912)
David S\. BermanAffiliation:Centre for Theoretical Physics, Queen Mary University of London, Mile End Road, London E1 4NS, United KingdomRoger G\. MelkoAffiliation:Department of Physics and Astronomy, University of Waterloo, Waterloo, Ontario N2L 3G1, CanadaAffiliation:Perimeter Institute for Theoretical Physics, Waterloo, Ontario N2L 2Y5, CanadaAlexander G\. StapletonAffiliation:Centre for Theoretical Physics, Queen Mary University of London, Mile End Road, London E1 4NS, United Kingdom
###### Abstract
In this work, we study the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data gathered from interacting Rydberg atom arrays\. The quantum system is known to exhibit a finite\-size remnant of a critical point as the laser detuning parameter is varied\. We find that near the critical point the transformer loss as a function of training dataset size is well described by a power\-law with a loss floor correction\. However, away from criticality the quality of the power\-law description is substantially reduced\. We then compare the statistical structure of both Rydberg measurements and natural\-language corpora using an entropy\-normalised, finite sample corrected mutual information ‘two\-point’ function\. We find that near\-critical statistics of the two point functions are closest to those observed in natural\-language, whilst other qubit configurations far from the critical point have two\-point functions that decay more rapidly\. This supports the hypothesis that multi\-scale dependence contributes to stable neural scaling, and that scaling behaviour should be viewed as a property of the model–data pair\.
## IIntroduction
In recent years, transformer models\[[22](https://arxiv.org/html/2609.20912#bib.bib4)\]have become pervasive across a wide range of applications\. Of these, arguably the most ubiquitous use of transformers is in models which produce and process natural language\. Such networks often have a vast number of parameters, and are thus almost universally referred to aslarge language models\(LLMs\)\. Undoubtedly, LLMs represent one of the most significant technological advances of the past century; however, as with many modern and competitive machine\-learning models, their training is computationally expensive, time\-consuming, and associated with substantial environmental costs\. Understanding how large models scale is thus of central importance\. By relating the computational resources required for training to the resulting predictive loss, an effective neural scaling law should enable the forecasting resources, allow one to make cheap comparisons between learning regimes, and assess how efficiently a model captures the statistical structure of its training data\.
In this paper, we ask whether transformers trained on quantum measurement data exhibit similar neural scaling laws, and therefore whether their performance could reliably be forecast as increasingly large datasets from quantum devices become available for training\.
In general, common ‘neural scaling laws’ are actually almost exclusivelynatural languageneural scaling laws\. Of these, the two most famous areHoffmann scaling\[[11](https://arxiv.org/html/2609.20912#bib.bib1)\]andKaplan scaling\[[13](https://arxiv.org/html/2609.20912#bib.bib3)\]\. As the number of model parameters and the training tokens \(i\.e\. subwords\) increase, the predictive loss is shown to follow a remarkably regular and predictable trend\. The success of these frameworks motivates a broader question: is scaling behaviour a distinctive property of natural language, or does it reflect a more general feature of learning from structured probability distributions with a transformer\-based architecture?
Whilst some works have endeavoured to derive scaling properties of natural language from their statistics, for example\[[5](https://arxiv.org/html/2609.20912#bib.bib16)\], the generalisation of similar statistical measures to broader contexts remains remarkably under\-studied\. Recent theoretical works studying datasets motivated by physical systems suggest that neural scaling can depend crucially on the statistical structure of the training distribution\[[1](https://arxiv.org/html/2609.20912#bib.bib17),[2](https://arxiv.org/html/2609.20912#bib.bib18),[18](https://arxiv.org/html/2609.20912#bib.bib19)\]\.
In this work, we investigate this question using RydbergGPT\[[9](https://arxiv.org/html/2609.20912#bib.bib2)\], an open\-source transformer model trained on synthetic measurement samples generated by quantum Monte Carlo simulations of a neutral Rydberg atom array\.
Rydberg atom arrays are programmable quantum computing devices, that combine flexible control of qubit interactions with single\-atom resolved preparation and measurement\[[7](https://arxiv.org/html/2609.20912#bib.bib20)\]\.
They are composed of atoms in which one valence electron is excited to a state with a very large principal quantum numbernn\. In such states the electron is only weakly bound and occupies an orbital whose radius scales approximately asn2n^\{2\}, giving rise to exaggerated atomic properties\. In particular, the electric dipole moment and polarisability become very large, meaning that two Rydberg atoms can interact strongly even when separated by relatively large distances; see\[[19](https://arxiv.org/html/2609.20912#bib.bib5),[4](https://arxiv.org/html/2609.20912#bib.bib6)\]for reviews of the platform and its use in quantum simulation\.
TheRydberg modeltreats each atom as an effective two\-level qubit system consisting of an electronic ground state\|g⟩\\ket\{g\}and a Rydberg state\|r⟩\\ket\{r\}\. Following the notation of RydbergGPT\[[9](https://arxiv.org/html/2609.20912#bib.bib2)\]\(and the standard square\-lattice Rydberg\-array Hamiltonians used in studies of density\-wave order and quantum phase transitions\[[20](https://arxiv.org/html/2609.20912#bib.bib7),[12](https://arxiv.org/html/2609.20912#bib.bib8)\]\), for a squareL×LL\\times Larray of atoms at position vectors\{ri\}i=1N\\\{r\_\{i\}\\\}\_\{i=1\}^\{N\}, whereN=L2N=L^\{2\}, the Hamiltonian may be written as
H^=∑i<jC6‖ri−rj‖6n^in^j−δ∑i=1Nn^i−Ω2∑i=1Nσ^ix,\\hat\{H\}=\\sum\_\{i<j\}\\frac\{C\_\{6\}\}\{\\\|r\_\{i\}\-r\_\{j\}\\\|^\{6\}\}\\hat\{n\}\_\{i\}\\hat\{n\}\_\{j\}\-\\delta\\sum\_\{i=1\}^\{N\}\\hat\{n\}\_\{i\}\-\\frac\{\\Omega\}\{2\}\\sum\_\{i=1\}^\{N\}\\hat\{\\sigma\}\_\{i\}^\{x\},\(1\)where
σ^ix=\|g⟩i⟨r\|i\+\|r⟩i⟨g\|i,n^i=12\(σ^i\+1\)=\|r⟩i⟨r\|i,\\hat\{\\sigma\}\_\{i\}^\{x\}=\\ket\{g\}\_\{i\}\\bra\{r\}\_\{i\}\+\\ket\{r\}\_\{i\}\\bra\{g\}\_\{i\},\\quad\\hat\{n\}\_\{i\}=\\frac\{1\}\{2\}\\left\(\\hat\{\\sigma\}\_\{i\}\+1\\right\)=\\ket\{r\}\_\{i\}\\bra\{r\}\_\{i\},σ^i=\|r⟩i⟨r\|i−\|g⟩i⟨g\|i\.\\hat\{\\sigma\}\_\{i\}=\\ket\{r\}\_\{i\}\\bra\{r\}\_\{i\}\-\\ket\{g\}\_\{i\}\\bra\{g\}\_\{i\}\.\(2\)The interaction strength may equivalently be parameterised by the blockade radiusRbR\_\{b\}and lattice spacingaa,
C6=Ω\(Rba\)6,Vij=a6‖ri−rj‖6\.C\_\{6\}=\\Omega\\left\(\\frac\{R\_\{b\}\}\{a\}\\right\)^\{6\},\\qquad V\_\{ij\}=\\frac\{a^\{6\}\}\{\\\|r\_\{i\}\-r\_\{j\}\\\|^\{6\}\}\.\(3\)
The parameterΩ\\Omegasets the coherent drive between\|g⟩\\ket\{g\}and\|r⟩\\ket\{r\}, while the laser detuningδ\\deltabiases the energy cost of creating a Rydberg excitation through the term−δ∑in^i\-\\delta\\sum\_\{i\}\\hat\{n\}\_\{i\}\. The ratioRb/aR\_\{b\}/afixes the effective interaction scale throughC6C\_\{6\}, the matrixVVencodes the lattice geometry via the separations‖ri−rj‖\\\|r\_\{i\}\-r\_\{j\}\\\|\. The final physical parameter is the Rydberg blockade radiusRbR\_\{b\}which penalises simultaneous excitation of nearby atoms\. In addition, in the synthetic RydbergGPT dataset\[[10](https://arxiv.org/html/2609.20912#bib.bib21)\]produced by world\-line QMC, an effective inverse temperatureβΩ\\beta\\Omegais included in order to produce thermal ensembles\.
The competition between laser driving, detuning, interactions and temperature produces structured many\-body measurement distributions\.
### I\.1RydbergGPT
RydbergGPT is an autoregressive transformer model, parameterised by weightsθ\\theta, designed to learn measurement distributions of Rydberg systems\[[9](https://arxiv.org/html/2609.20912#bib.bib2)\]\. Architecturally, it inherits the transformer/autoregressive modelling structure introduced in\[[22](https://arxiv.org/html/2609.20912#bib.bib4)\], but conditions generation on Hamiltonian data rather than natural language context\. In the notation of the previous section, the physically controllable parameters of the model are encoded by
x=\(Ω,δ/Ω,Rb/a,V,βΩ\),x=\\left\(\\Omega,\\,\\delta/\\Omega,\\,R\_\{b\}/a,\\,V,\\,\\beta\\Omega\\right\),\(4\)and the target sequence is a binary occupation\-basis measurement
σ=\{σ1,σ2,…,σN\}\.\\sigma=\\\{\\sigma\_\{1\},\\sigma\_\{2\},\\ldots,\\sigma\_\{N\}\\\}\.\(5\)The model learns the conditional probabilities
pθ\(σ,x\)=∏i=1Npθ\(σi∣σ<i;x\),p\_\{\\theta\}\(\\sigma;x\)=\\prod\_\{i=1\}^\{N\}p\_\{\\theta\}\\\!\\left\(\\sigma\_\{i\}\\mid\\sigma\_\{<i\};x\\right\),\(6\)whereθ\\thetadenotes the trainable neural\-network parameters\.
In this machine\-learning formulation,xxis the condition supplied to the encoder: changingxxinduces a change in the physical properties of the distribution that the decoder learns\. The sequence variablesσi\\sigma\_\{i\}play the role of tokens, whileθ\\thetacontrols the expressive capacity of the transformer used to approximate the Rydberg array measurement distribution\. Each configuration is canonically tokenised such thatσi\\sigma\_\{i\}correspond to bits in\{0,1\}\\\{0,1\\\}, representing\|g⟩\\ket\{g\}and\|r⟩\\ket\{r\}respectively\.
For visual reference, Figure[1](https://arxiv.org/html/2609.20912#S1.F1)shows four single occupation\-basis measurements from the same ensembles\. The samples are selected reproducibly without conditioning on their spatial pattern\. They illustrate the binary configurations presented to the model, rather than an ensemble\-averaged order parameter or correlation function\.
Figure 1:Representative randomly selected occupation\-basis configurations from an open6×66\\times 6Rydberg array at the four labelled detunings\[[10](https://arxiv.org/html/2609.20912#bib.bib21)\]\. Light sites are in the ground state and dark sites represent Rydberg excitations\. The samples were selected with a fixed pseudorandom seed from the same data archive used for the data\-scaling and mutual\-information analyses\. They are individual measurements and are shown for visual orientation only\.As was alluded to earlier, the physical interpretability ofxxmakes RydbergGPT a useful test case for scaling behaviour beyond natural language\. In language, loss scaling is usually interpreted in terms of the number of text tokens, the number of model parameters \(which vicariously encode the compute budget\), and the acceptable loss value\. Conversely, in the Rydberg setting, the amount of data is only one part of the story\. Samples at different detunings possess very different internal structure, ranging from weakly dependent random fields in the very high temperature regime to more constrained spatial patterns nearer criticality, for example\.
One way to consider this is through effective dataset size\. In a maximally degenerate dataset the number of data samples to fully specify the target distribution is simply unity, whereas in the random noise limit, the number of data samples to specify the distribution becomes infinite\. Of course, in large limits the distribution of text is effectively not random \(for example Zipf’s law applies\) in the same way that quantum systems at criticality are not random\. Schematically, one may write the effective size of the datasetDeffD\_\{\\text\{eff\}\}as a scaling of a Zipfian text dataset, namelyDeff=ΓDtextD\_\{\\text\{eff\}\}=\\Gamma D\_\{\\text\{text\}\}\. In the case thatΓ\>1\\Gamma\>1, the training dataset in some sense encodes more information than an equivalently sized corpus for natural language\. Conversely, ifΓ<1\\Gamma<1, then the corpus encodes less information than an equivalently sized natural language corpus\.
A study of the dependence on effective dataset size is crucial since it asks whether the empirical scaling principles developed for language models are properties of text itself111Natural language exhibits incredible structure: Zipf’s law\[[14](https://arxiv.org/html/2609.20912#bib.bib15)\]holds almost universally, for example\[[3](https://arxiv.org/html/2609.20912#bib.bib14)\]\.or of a broader class of structured distributions\. If similar scaling behaviour appears in a controlled physical setting, it becomes possible to relate model performance not only to dataset size and parameter count, but also to interpretable features of the underlying quantum system\. A natural way to make this dependence structure precise is through mutual information\. The information\-theoretic foundation for this is entropy\[[21](https://arxiv.org/html/2609.20912#bib.bib9)\]; mutual\-information functions have long been used to diagnose statistical dependence in natural language\[[15](https://arxiv.org/html/2609.20912#bib.bib10),[8](https://arxiv.org/html/2609.20912#bib.bib11)\]\. The connection between mutual\-information decay, the structure of natural language, and statistical physics of states at criticality was made explicit by Lin and Tegmark\[[17](https://arxiv.org/html/2609.20912#bib.bib12)\]; although more recent work has also used mutual information as a scaling object for long\-context language modelling\[[6](https://arxiv.org/html/2609.20912#bib.bib13)\]\. These references motivate treating raw token or sample count as only one coordinate of the learning problem\.
### I\.2A General Review of Hoffmann Scaling
Whilst there are many known scaling laws which empirically provide a good fit to experimental data, arguably one of the most ubiquitous isHoffmann scaling222An equally prevalent scaling law is known asKaplan scaling, initially presented in\[[13](https://arxiv.org/html/2609.20912#bib.bib3)\]\.\. Hoffmann scaling laws were developed for transformer language models trained under a fixed compute budget\[[11](https://arxiv.org/html/2609.20912#bib.bib1)\], building on the broader empirical scaling\-law programme of\[[13](https://arxiv.org/html/2609.20912#bib.bib3)\]\. The main qualitative lesson is that compute\-optimal training should not only increase the number of parameters; it should also increase the number of training tokens\. A model that is too large for its dataset is under\-trained, while a model that is too small may fail to use available data efficiently\. In the usual language\-modelling notation, the expected loss is treated as a function of parameter countPPand token countDD, for example through an empirical decomposition of the form
ℒ\(P,D\)≈ℒ∞\+AP−α\+BD−γ,\\mathcal\{L\}\(P,D\)\\approx\\mathcal\{L\}\_\{\\infty\}\+AP^\{\-\\alpha\}\+BD^\{\-\\gamma\},\(7\)whereℒ∞\\mathcal\{L\}\_\{\\infty\}is an irreducible loss floor and the two power\-law terms describe limitations from finite model size and finite data\. Chinchilla scaling then asks howPPandDDshould co\-vary when the training compute budget is held fixed\.
Since quantum data does not exhibit a universal structure \(like that in natural language which gives rise to Zipf’s law\), a useful empirical model for the RydbergGPT loss arises from promoting the scaling constants in Equation \([7](https://arxiv.org/html/2609.20912#S1.E7)\) to functions of the physical parameters, i\.e\.
ℒRyd\(P,D,x\)≈ℒ∞\(x\)\+A\(x\)P−α\(x\)\+B\(x\)D\(x\)−γ\(x\)\.\\mathcal\{L\}\_\{\\mathrm\{Ryd\}\}\(P,D;x\)\\approx\\mathcal\{L\}\_\{\\infty\}\(x\)\+A\(x\)P^\{\-\\alpha\(x\)\}\+B\(x\)D\(x\)^\{\-\\gamma\(x\)\}\.\(8\)Following the spirit of\[[13](https://arxiv.org/html/2609.20912#bib.bib3),[11](https://arxiv.org/html/2609.20912#bib.bib1)\], this form is not assumed to be exact, but rather to separate three possible causes of improved validation loss: increased model capacity, increased raw sample count, and changes in the physical complexity of the data distribution\. In the remainder of this work, we hold all parameters inxxfixed except the laser detuningδ\\delta, and use that as the physical control parameter\.
## IIFixed\-Parameter Data Scaling of RydbergGPT
We first consider data scaling at fixed model size\. LetP0P\_\{0\}denote the fixed number of trainable parameters and letDDdenote the number of training configurations, or equivalently the number of lattice\-token sequences used during training\. The dataset size is swept at fixedL=6L=6,Rb/a=1\.15R\_\{b\}/a=1\.15, andβΩ=8\\beta\\Omega=8for the completed detuning set
δ/Ω∈\{−0\.36,0\.93,1\.05,1\.17,1\.52,2\.94\}\.\\delta/\\Omega\\in\\\{\-0\.36,0\.93,1\.05,1\.17,1\.52,2\.94\\\}\.\(9\)The detunings1\.051\.05and1\.171\.17bracket the known quantum critical regime of this model\.
Note, due to the small lattice sizes considered here, we only have access to the finite\-size remnant of the true quantum phase transition\. We nonetheless refer to this of delta as thecritical pointin the below discussion\. As was found in\[[12](https://arxiv.org/html/2609.20912#bib.bib8)\]which employed quantum Monte Carlo simulations to model the near\-critical region,δ/Ω≃1\.1\\delta/\\Omega\\simeq 1\.1may be used as a reference value for the critical point\.
For each detuning, the validation loss is measured as
ℒ^val\(D;δ\)∼−1Mval∑m=1Mval∑t=1TmlogpP0,D\(σt\(m\)∣σ<t\(m\),δ\),\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{val\}\}\(D;\\delta\)\\sim\-\\frac\{1\}\{M\_\{\\mathrm\{val\}\}\}\\sum\_\{m=1\}^\{M\_\{\\mathrm\{val\}\}\}\\sum\_\{t=1\}^\{T\_\{m\}\}\\log p\_\{P\_\{0\},D\}\\\!\\left\(\\sigma\_\{t\}^\{\(m\)\}\\mid\\sigma\_\{<t\}^\{\(m\)\},\\delta\\right\),\(10\)whereσt\(m\)\\sigma\_\{t\}^\{\(m\)\}is thett\-th site token in themm\-th held\-out draw from the Rydberg ensemble, andpP0,Dp\_\{P\_\{0\},D\}is the trained RydbergGPT predictive distribution at fixed parameter countP0P\_\{0\}\. For each data fraction and seed, we take the minimum logged training or validation loss, average over ten seeds, and fit the mean curve\. The data fractions range from10%10\\%to100%100\\%of the available training data\.
At fixed model size, we fit the aggregate curves to the empirical form
ℒ\(f,δ\)=ℒ∞\(δ\)\+BD\(δ\)D\(f\)−γD\(δ\),\\mathcal\{L\}\(f;\\delta\)=\\mathcal\{L\}\_\{\\infty\}\(\\delta\)\+B\_\{D\}\(\\delta\)D\(f\)^\{\-\\gamma\_\{D\}\(\\delta\)\},\(11\)whereffis the training\-data percentage,ℒ∞\\mathcal\{L\}\_\{\\infty\}is the fitted loss floor, andγD\\gamma\_\{D\}is the data\-scaling exponent\.
Figure[2](https://arxiv.org/html/2609.20912#S2.F2)compares the two completed detunings that bracket the critical point\. Both the training and validation curves decrease systematically with training\-set fraction\.
Figure 2:Fixed\-model RydbergGPT data\-scaling curves on the two sides of the critical detuning point\. The system parameters areL=6L=6,Rb/a=1\.15R\_\{b\}/a=1\.15andβΩ=8\\beta\\Omega=8\. Solid and dashed curves denote the seed\-averaged training and validation losses, respectively; shaded regions show one standard deviation across seeds\. Dotted and dash\-dotted curves are fits to Equation \([11](https://arxiv.org/html/2609.20912#S2.E11)\)\. The scientific\-notation multiplier is shown once per vertical axis\.Table 1:Aggregate fits to Equation \([11](https://arxiv.org/html/2609.20912#S2.E11)\)\. Each fit uses ten data fractions\. The validation loss floor is lower atδ/Ω=1\.17\\delta/\\Omega=1\.17than at1\.051\.05, while the validation exponent remains of order unity on both sides of the critical point\.The fitted validation floor decreases by3\.2%3\.2\\%betweenδ/Ω=1\.05\\delta/\\Omega=1\.05and1\.171\.17, from0\.48390\.4839to0\.46830\.4683\. The corresponding exponent increases from0\.6560\.656to0\.7910\.791\. The high fit qualities in Table[1](https://arxiv.org/html/2609.20912#S2.T1)show that the observed aggregate curves are adequately represented by the three\-parameter form within this narrow interval\. They do not, by themselves, identify a universal exponent or a discontinuity\.
Figure 3:Normalised aggregate loss curves and fitted forms for the two detunings in Figure[2](https://arxiv.org/html/2609.20912#S2.F2)\. Points are observed seed means and lines are fits to Equation \([11](https://arxiv.org/html/2609.20912#S2.E11)\)\. Each series is normalised by its largest observed loss for display only\.### II\.1Dynamics of the Fit Parameters
Whilst the previous section considers data close to the critical point, a natural question is: how do the fitted parameters evolve as the detuning is varied away from criticality? To answer this question, this section analyses the induced flow of the data\-scaling exponentγD\\gamma\_\{D\}, the goodness of fit of the power\-law ansatz as quantified byR2R^\{2\}, and the asymptotic loss floorℒ∞\\mathcal\{L\}\_\{\\infty\}across the low\-detuning regime, through the critical point, and into the high\-detuning regime\. The values ofδ\\deltaare chosen such that they do not deviatetoofar fromδ=1\.1\\delta=1\.1and saturate the classical or resonant regimes\.
Figure 4:Validation\-fit parameters for the six retained detunings\. Faint points denote fits to individual random seeds, while circles and error bars indicate the corresponding mean and standard deviation at each detuning\. The dashed vertical line marks the critical point atδ/Ω=1\.10\\delta/\\Omega=1\.10\.Figure[4](https://arxiv.org/html/2609.20912#S2.F4)extends the preceding analysis to six detuning values by displaying the results of validation fits performed independently for each random seed\. The plotted spread directly characterises seed\-to\-seed variability in the fitted parameters and should not be interpreted as a confidence interval for a thermodynamic critical exponent\. The values closest to the critical point atδ/Ω=1\.10\\delta/\\Omega=1\.10are consistent with the aggregate results reported in Table[1](https://arxiv.org/html/2609.20912#S2.T1)\.
The strongest agreement with the power\-law ansatz is observed in the vicinity of the critical point\. Near the dashed line in Figure[4](https://arxiv.org/html/2609.20912#S2.F4), the fittedR2R^\{2\}values are both high and relatively tightly clustered, while the corresponding values ofγD\\gamma\_\{D\}remain of order unity\. This behaviour is consistent with the aggregate fits in Table[1](https://arxiv.org/html/2609.20912#S2.T1)\. Atδ/Ω=1\.05\\delta/\\Omega=1\.05and1\.171\.17, the training fits yieldR2=0\.998R^\{2\}=0\.998and0\.9970\.997, respectively, while the corresponding validation fits yieldR2=0\.989R^\{2\}=0\.989and0\.9830\.983\. Likewise, the fitted curves in Figure[3](https://arxiv.org/html/2609.20912#S2.F3)closely track the normalised seed\-averaged losses at these detunings\.
By contrast, the fits become less stable towards the endpoints of the detuning sweep\. In particular, atδ/Ω=−0\.36\\delta/\\Omega=\-0\.36and2\.942\.94, the seed\-levelR2R^\{2\}values are lower and exhibit a broader spread\. The fitted exponents and asymptotic loss floors also vary more strongly between seeds\. Although the Hoffmann form can still be fitted at an individual detuning far from the critical point, the resulting parameters are substantially less reproducible\.
As can be seen in Figure[2](https://arxiv.org/html/2609.20912#S2.F2), both the training and validation losses decrease smoothly with increasing dataset size atδ/Ω=1\.05\\delta/\\Omega=1\.05and1\.171\.17\. The fitted curves in Figure[3](https://arxiv.org/html/2609.20912#S2.F3)closely follow the corresponding normalised seed means, and all four aggregate fits reported in Table[1](https://arxiv.org/html/2609.20912#S2.T1)achieve high values ofR2R^\{2\}against the power\-law ansatz\.
Taken together, these observations indicate that a power\-law ansatz provides the most accurate empirical description of the sampled data near the critical point\.
A possible explanation for the enhanced stability of this scaling form near the critical point is provided by the behaviour of the correlation length\. At a continuous phase transition in the thermodynamic limit, the correlation lengthξ\\xidiverges, and the system develops fluctuations over arbitrarily large spatial scales\. Although finite lattices and non\-zero temperatures preclude a true divergence, one nevertheless expectsξ\\xito become large relative to its values away from the critical point; the resulting distribution contains correlations over a broad range of length scales, rather than being dominated by either short\-range fluctuations or a single characteristic scale\. As the dataset size is increased, the model gains access to progressively better estimates of correlations at increasingly large separations and of higher\-order configurations involving multiple spatial scales\. Each additional increment of data presumably resolves further structure in the distribution, producing a gradual power\-law improvement in the loss\.
Away from the critical point, where the correlation length is shorter, the relevant statistical structure may be exhausted more rapidly\. In that regime, additional samples predominantly refine already\-learned local statistics, and the simple power\-law ansatz need not remain stable over the sampled range of dataset sizes as they are primarily informing the model about what it has already learnt\.
These correlations provide additional learnable structure as the amount of training data is increased\. The observed scaling form therefore suggests that Hoffmann\-like behaviour may arise more generally from the organised and learnable structure of a probability distribution, rather than being specific to language\. This analogy should nevertheless be interpreted cautiously: it does not imply that the fitted exponents are universal, nor that quantum\-measurement data and natural language data share a common mechanism\.
It is in this restricted sense that RydbergGPT near the critical point exhibits scaling behaviour analogous to that observed in natural language models\. Natural language is organised at several levels that unfold over different lengths of a sequence\. Nearby words are linked by grammar and local meaning, while longer spans of text carry sentence structure, and paragraphs communicate ideas\. These patterns operate over a combination of short, medium, and large distances, since a word can depend not only on the words immediately before it, but also on information introduced much earlier\. The fact that quantum models trained on near\-critical data obey the same correspondence as language is non\-trivial and somewhat remarkable\.
## IIIMutual\-information scaling for Rydberg and natural language data
The scaling analysis in Section[II](https://arxiv.org/html/2609.20912#S2)shows that the fixed\-model power\-law ansatz is most stable near the critical point\. We postulated that this observation could suggest that the scaling of the predictive loss may be connected to the organisation of the data distribution across varying spatial scales\. Near the critical point, dependence persists over a broad range of separations, so additional samples can continue to refine statistical structure that is not confined to nearest\-neighbour configurations\. Natural language is also structured across several sequential scales\. It is therefore useful to ask whether the two data sources possess quantitatively comparable ranges of statistical dependence, rather than relying only on an analogy between their loss curves\.
Taking inspiration from\[[16](https://arxiv.org/html/2609.20912#bib.bib22)\], we use mutual information as a candidate observable for this comparison333A limitation of this approach is that it is sensitive to broken permutation symmetries within the Rydberg snapshot which the transformer is not\.\. It detects arbitrary pairwise statistical dependence without assuming a linear relation, and it is defined for both binary occupations and categorical language symbols\. It is also invariant under invertible re\-labellings of a fixed alphabet\. Its raw magnitude is not, however, directly comparable across the two domains\. The local entropies are different, the Rydberg correlations are distributed over a two\-dimensional lattice, and a language distance expressed in tokens changes with the tokenisation convention\.
Of course, these considerations inform the comparison used here\. The same null\-corrected and entropy\-normalised metric is constructed in both systems; language is measured on a fixed character\-level basis; and the resulting second\-moment length is divided by the largest separation included in the measurement\.
### III\.1Definition of the common estimator
LetXiX\_\{i\}denote the discrete observable at locationii\. For the Rydberg data it is the binary occupation of a lattice site\. For language it is the character\-level symbol at a fixed position in the underlying text\. The latter representation has a fixed alphabet and a fixed microscopic coordinate; it is consequently unchanged by any subsequent word or subword tokenisation\.
For Rydberg sites at lattice coordinates\(xi,yi\)\(x\_\{i\},y\_\{i\}\)and\(xj,yj\)\(x\_\{j\},y\_\{j\}\), their separation in lattice\-spacing units is
rij=\(xi−xj\)2\+\(yi−yj\)2\.r\_\{ij\}=\\sqrt\{\(x\_\{i\}\-x\_\{j\}\)^\{2\}\+\(y\_\{i\}\-y\_\{j\}\)^\{2\}\}\.The displacement vector from siteiito sitejjis the ordered coordinate difference\(Δx,Δy\)=\(xj−xi,yj−yi\)\(\\Delta x,\\Delta y\)=\(x\_\{j\}\-x\_\{i\},y\_\{j\}\-y\_\{i\}\), so the two components give the signed horizontal and vertical offsets, respectively\. The magnitude of the pair\(Δx,Δy\)\(\\Delta x,\\Delta y\)is given byrijr\_\{ij\}\. Let𝒫s\\mathcal\{P\}\_\{s\}denote the multiset of observed value pairs at separationss\. For the Rydberg data,𝒫s\\mathcal\{P\}\_\{s\}contains\(Xi\(q\),Xj\(q\)\)\(X\_\{i\}^\{\(q\)\},X\_\{j\}^\{\(q\)\}\)for every configurationqqand every lattice pair withrij=sr\_\{ij\}=s\. For language, it contains\(Xt,Xt\+s\)\(X\_\{t\},X\_\{t\+s\}\)for every within\-document positionttfor which both symbols are present\. WritingNs=\|𝒫s\|N\_\{s\}=\|\\mathcal\{P\}\_\{s\}\|, the separation\-conditioned empirical joint law and its marginals are
p^s\(a,b\)\\displaystyle\\widehat\{p\}\_\{s\}\(a,b\)=1Ns\#\{\(x,y\)∈𝒫s:x=a,y=b\},\\displaystyle=\\frac\{1\}\{N\_\{s\}\}\\\#\\bigl\\\{\(x,y\)\\in\\mathcal\{P\}\_\{s\}:x=a,\\ y=b\\bigr\\\},p^s,L\(a\)\\displaystyle\\widehat\{p\}\_\{s,L\}\(a\)=∑bp^s\(a,b\),p^s,R\(b\)=∑ap^s\(a,b\)\.\\displaystyle=\\sum\_\{b\}\\widehat\{p\}\_\{s\}\(a,b\),\\qquad\\widehat\{p\}\_\{s,R\}\(b\)=\\sum\_\{a\}\\widehat\{p\}\_\{s\}\(a,b\)\.Thusssenters the estimator by selecting either a Pythagorean radial shell in the Rydberg array or a positional lag in the language sequence\. The mutual information at that separation is
I^\(s\)=∑a,bp^s\(a,b\)logp^s\(a,b\)p^s,L\(a\)p^s,R\(b\)\.\\widehat\{I\}\(s\)=\\sum\_\{a,b\}\\widehat\{p\}\_\{s\}\(a,b\)\\log\\frac\{\\widehat\{p\}\_\{s\}\(a,b\)\}\{\\widehat\{p\}\_\{s,L\}\(a\)\\widehat\{p\}\_\{s,R\}\(b\)\}\.\(12\)Mutual information estimated from a finite sample is generally non\-zero even when the underlying variables are independent\. This sampling bias is estimated by breaking the original pairing between observations\. Here, a pair comprises two values observed at separationss\. For the Rydberg data,\(Xi\(q\),Xj\(q\)\)\(X\_\{i\}^\{\(q\)\},X\_\{j\}^\{\(q\)\}\)contains the occupations of sitesiiandjjin the same configurationqq\. For language,\(Xt,Xt\+s\)\(X\_\{t\},X\_\{t\+s\}\)contains the symbols at positionsttandt\+st\+sin the same document\. Breaking the pairing means retaining every observed value but changing which value at the first location is matched to which value at the second\. For a fixed Rydberg site pair\(i,j\)\(i,j\), the occupations measured at siteiiare left in their original order, whereas those measured at sitejjare shuffled between configurations\. Equivalently,Xi\(q\)X\_\{i\}^\{\(q\)\}is paired withXj\(π\(q\)\)X\_\{j\}^\{\(\\pi\(q\)\)\}, whereπ\\piis a random permutation of the configuration labels; the two values therefore generally originate from different configurations\. For language, the same operation is applied at each lagss: the first symbol in every observed pair is retained, while the second symbols are shuffled between pairs\. Because only the ordering is changed, the occupation counts and the frequency of each language symbol are unchanged\. The original co\-occurrence pattern is nevertheless destroyed, so systematic dependence no longer contributes on average\. Each shuffled data set is called a permutation\-null sample\. Its mutual information estimates the finite\-sample contribution expected in the absence of dependence\. LetBBdenote the number of independent permutation\-null samples generated at each separation, and letI^null\(m\)\(s\)\\widehat\{I\}^\{\(m\)\}\_\{\\rm null\}\(s\)denote the mutual information estimated from samplem∈\{1,…,B\}m\\in\\\{1,\\ldots,B\\\}\. The empirical marginal entropies areH^L\(s\)=−∑ap^s,L\(a\)logp^s,L\(a\)\\widehat\{H\}\_\{L\}\(s\)=\-\\sum\_\{a\}\\widehat\{p\}\_\{s,L\}\(a\)\\log\\widehat\{p\}\_\{s,L\}\(a\)andH^R\(s\)=−∑bp^s,R\(b\)logp^s,R\(b\)\\widehat\{H\}\_\{R\}\(s\)=\-\\sum\_\{b\}\\widehat\{p\}\_\{s,R\}\(b\)\\log\\widehat\{p\}\_\{s,R\}\(b\), with0log00\\log 0defined to be zero\.
Equation \([13](https://arxiv.org/html/2609.20912#S3.E13)\) combines the two adjustments required for comparison across the Rydberg and language data\. First, the mean permutation\-null mutual information is subtracted from the observed value to remove the contribution expected from finite\-sample bias\. Second, the result is divided by the geometric mean of the two marginal entropies\. Raw mutual information is measured in nats and its attainable magnitude depends on the uncertainty of the two variables; the entropy denominator removes this local scale\. The geometric mean is symmetric under exchange of the two locations\. Moreover, for exact distributions,I\(X,Y\)≤min\{H\(X\),H\(Y\)\}≤H\(X\)H\(Y\)I\(X;Y\)\\leq\\min\\\{H\(X\),H\(Y\)\\\}\\leq\\sqrt\{H\(X\)H\(Y\)\}, so the corresponding uncorrected entropy\-normalised mutual information lies between zero and one\. The resulting bias\-corrected, dimensionless profile is
ρ^MI\(s\)=I^\(s\)−B−1∑m=1BI^null\(m\)\(s\)H^L\(s\)H^R\(s\)\.\\widehat\{\\rho\}\_\{\\mathrm\{MI\}\}\(s\)=\\frac\{\\widehat\{I\}\(s\)\-B^\{\-1\}\\sum\_\{m=1\}^\{B\}\\widehat\{I\}^\{\(m\)\}\_\{\\rm null\}\(s\)\}\{\\sqrt\{\\widehat\{H\}\_\{L\}\(s\)\\widehat\{H\}\_\{R\}\(s\)\}\}\.\(13\)For each language corpus, positional stationarity is assumed over the analysed prefix: the probability of observing a given symbol is taken not to depend on its absolute position within a document\. The left and right marginal entropies in Equation \([13](https://arxiv.org/html/2609.20912#S3.E13)\) are therefore estimated by the same corpus\-wide symbol entropy\. This assumption concerns only the single\-symbol distribution; dependence between symbols at different positions remains the quantity being measured\.
For the Rydberg data, the occupation probability, and hence the binary entropy, can differ between sites because the array has open boundaries\. Equation \([13](https://arxiv.org/html/2609.20912#S3.E13)\) is therefore evaluated separately for each site pair\(i,j\)\(i,j\): the null\-subtracted mutual information for that pair is divided byH^iH^j\\sqrt\{\\widehat\{H\}\_\{i\}\\widehat\{H\}\_\{j\}\}before any spatial average is taken, whereH^i\\widehat\{H\}\_\{i\}andH^j\\widehat\{H\}\_\{j\}are the empirical occupation entropies of the two sites\. For a fixed displacement\(Δx,Δy\)\(\\Delta x,\\Delta y\), the resulting pair\-normalised values are averaged over every origin for which both sites lie inside the array\. These displacement averages are then combined for all displacements with the same Pythagorean lengths=Δx2\+Δy2s=\\sqrt\{\\Delta x^\{2\}\+\\Delta y^\{2\}\}\. Averaging in this order gives each displacement equal weight even when different numbers of its translations fit within the open boundary\.
The population value of mutual information is non\-negative, but the null\-subtracted finite\-sample estimate in Equation \([13](https://arxiv.org/html/2609.20912#S3.E13)\) can fluctuate below zero\. No monotonicity constraint is imposed: mutual information is not required to decrease at every successive separation, and shell\-dependent or lag\-dependent structure should remain in the estimator\. Negative estimates are also not set to zero, since pointwise clipping would retain positive fluctuations while discarding comparable negative fluctuations and would therefore bias the range upwards\.
The signed estimates are instead used directly in the cumulative zeroth and second moments\. For a cutoffRR, the second\-moment length is
ξ^MI,22\(R\)=∑0<sk≤Rgd\(sk\)sk2ρ^MI\(sk\)2d∑0<sk≤Rgd\(sk\)ρ^MI\(sk\)\.\\widehat\{\\xi\}\_\{\\mathrm\{MI\},2\}^\{2\}\(R\)=\\frac\{\\displaystyle\\sum\_\{0<s\_\{k\}\\leq R\}g\_\{d\}\(s\_\{k\}\)s\_\{k\}^\{2\}\\widehat\{\\rho\}\_\{\\mathrm\{MI\}\}\(s\_\{k\}\)\}\{\\displaystyle 2d\\sum\_\{0<s\_\{k\}\\leq R\}g\_\{d\}\(s\_\{k\}\)\\widehat\{\\rho\}\_\{\\mathrm\{MI\}\}\(s\_\{k\}\)\}\.\(14\)HereRRis the largest separation included in the sum,d=2d=2for the Rydberg array andd=1d=1for language, whilegd\(s\)g\_\{d\}\(s\)is the number of distinct displacement vectors, or signed coordinate offsets, whose magnitude isss\. In one dimension the two directions supply a common factor which cancels\. The factor2d2dis the standard dimensional normalisation of a second\-moment length\. Appreciable dependence at large separations increasesξ^MI,2\(R\)\\widehat\{\\xi\}\_\{\\mathrm\{MI\},2\}\(R\)through the factorsk2s\_\{k\}^\{2\}\. The length is defined when the cumulative numerator and denominator are both positive; this condition holds throughout the cutoff ranges reported below\.
### III\.2Data and estimation
The Rydberg calculation uses the open6×66\\times 6occupation arrays atRb/a=1\.15R\_\{b\}/a=1\.15,βΩ=8\\beta\\Omega=8and the six detunings employed in the scaling analysis\. Each profile is estimated from10510^\{5\}independent configurations, such thatRmax=50R\_\{\\max\}=\\sqrt\{50\}in lattice\-spacing units \(for a 6x6 array one has a maximum of 5 spaces in each direction, giving Pythagorean distance50\\sqrt\{50\}\)\.
For language, the two common open natural language datasets AG News and Yahoo Answers are used\. Each dataset is represented by a prefix containing 200k word\-level symbols after a fixed text normalisation\. Pairs are pooled only within documents, so no dependence is introduced across an artificial document boundary\. Symbol lags1≤s≤601\\leq s\\leq 60are retained\.
Figure 5:Entropy\-normalised excess mutual\-information profiles for the Rydberg occupations \(left\) and natural language symbols \(right\)\. Markers show the signed permutation\-corrected estimates in Equation \([13](https://arxiv.org/html/2609.20912#S3.E13)\), and lines join successive separations\. The symmetric\-logarithmic vertical scale is linear for\|ρ^MI\|≤10−4\|\\widehat\{\\rho\}\_\{\\mathrm\{MI\}\}\|\\leq 10^\{\-4\}, so small negative finite\-sample estimates remain visible\. Separations are divided by the largest available cutoff within each system\. The left panel contains the six open6×66\\times 6Rydberg detunings\. The right panel contains the AG News and Yahoo answers profiles in the fixed character\-level representation; the dashed line is their pointwise arithmetic mean at each separation\.Figure[5](https://arxiv.org/html/2609.20912#S3.F5)shows that the detunings with the highest Hoffmann\-fitR2R^\{2\}values in Section[II](https://arxiv.org/html/2609.20912#S2)are also those whose Rydberg MI two\-point functions are closest in shape to their natural language counterparts\. The language profiles fall rapidly at short separation and then cross over to a shallower positive tail\. The Rydberg curves at these high\-R2R^\{2\}detunings display a similar hierarchy: strong local dependence is accompanied by progressively weaker dependence over larger lattice distances\. This differs from the remote configurations atδ/Ω=−0\.36\\delta/\\Omega=\-0\.36and2\.942\.94, for which the corrected MI rapidly approaches the null level\. For an autoregressive learner, a positive tail means that increasingly distant observations can still carry predictive information\. Since this signal becomes weaker with separation, more samples are required to estimate it reliably; the model can therefore continue to learn structure at larger scales after the dominant local statistics have been resolved\. By contrast, once the profile has collapsed to the null level, increasing the context range supplies little additional pairwise information\. The numerical proximity of the corresponding dimensionless ranges should not be over\-interpreted; what is important is the scaling behaviour\. The coexistence of strong local information and a weaker extended tail during the critical point therefore provides supporting evidence for the hypothesis posed in Section[II](https://arxiv.org/html/2609.20912#S2): the enhanced stability of Hoffmann\-type scaling near the critical region is associated with learnable statistical dependence distributed across a hierarchy of scales\.
## IVConclusions
With the advent of modern neural networks, neural scaling laws have become a crucial means of relating attainable loss to the number of available training tokens and model parameters\. Their success in language modelling nevertheless leaves open a structural question: which properties of a training distribution permit a stable power\-law learning curve? Do models trained on quantum data exhibit the same scaling behaviour as those trained on natural language? To address this, we tested fixed\-model Hoffmann scaling across a detuning\-controlled finite\-size Rydberg critical point while holding the architecture, objective, and training procedure fixed\. By comparing the goodness\-of\-fit of a power\-law Hoffmann ansatz across several detunings, we assess whether quantum data exhibits language\-like scaling across different physical regimes\.
The central result is that, within an envelope surrounding the critical point, the loss is accurately described by a power\-law ansatz consisting of an asymptotic loss floor and a power\-law dependence on dataset size\. The fitted curves attain high values ofR2R^\{2\}and exhibit limited seed\-to\-seed variation\. At detunings far from the critical region, both fit quality and parameter stability are substantially reduced\. The evidence therefore supports the Hoffmann form as an empirical description of near\-critical data, but not as a detuning\-independent scaling law for quantum data in general\. More broadly, this suggests that stable power\-law scaling is not determined by model architecture alone, but depends on the statistical regime of the training distribution\.
Mutual\-information analysis provides independent support for this interpretation\. Rydberg configurations near the critical point show the closest agreement in functional form with natural language: both display strong short\-range dependence followed by a decaying positive tail over a substantial fraction of the normalised observation range\. By contrast, profiles at detunings distant from the critical point decay rapidly\. Although this does not imply equality of correlation lengths or a universality\-like equivalence between the two distributions, it indicates that the near\-critical Rydberg data has the most language\-like statistical organisation\.
Taken together, the simultaneous occurrence of stable power\-law fits and dependence distributed across several scales identifies multi\-scale statistical structure as a plausible explanation for the remarkable validity of the scaling ansatz near the critical point\. One interpretation is that such distributions continue to provide statistically useful structure as dataset size increases, allowing progressively smaller but systematic improvements in loss\. Around the critical region, the correlation length is very large, meaning many samples are required to fully learn the distribution suggesting that scaling behaviour should be viewed as a property of the model–data pair, rather than of the model alone\.
Several limitations constrain the scope of this conclusion\. The mutual\-information estimator depends on the construction of the finite\-sample reference, and thus breaks the permutation andℤ2\\mathbb\{Z\}\_\{2\}symmetry of the array\. Quantum fluctuations, open boundaries, and the restricted lattice size also limit the physical interpretation, however finite\-size scaling in the underlying Monte Carlo data are well\-understood\[[10](https://arxiv.org/html/2609.20912#bib.bib21)\]\.
Nevertheless, given these results, we hypothesise that neural scaling laws may emerge most clearly for datasets whose statistical dependencies are distributed over multiple scales\. This suggests that for quantum simulators, like Rydberg atom arrays, that are expected to continue producing larger and richer qubit measurement datasets in the future, robust neural scaling laws are most likely to be observed when the system is tuned near a quantum critical point\. More generally, our work has provided a possible link between the structure of a training distribution and the form of its learning curve\.
## Data Availability Statement
## Acknowledgements
We thank Yi\-Hong Teoh and Schuyler Moss for crucial discussions on the RydbergGPT model\. DSB and AGS are grateful to Pierre Andurand for his donation supporting this research\. DSB is partially supported by the Science and Technology Facilities Council \(STFC\) Consolidated Grant ST/X00063X/1 “Amplitudes, Strings & Duality”\. RGM would like to acknowledge the support of the Natural Sciences and Engineering Research Council of Canada \(NSERC\)\. This research was also supported in part by grant NSF PHY\-2309135 to the Kavli Institute for Theoretical Physics \(KITP\)\. Research at the Perimeter Institute is supported in part by the Government of Canada through the Department of Innovation, Science and Economic Development Canada and by the Province of Ontario through the Ministry of Economic Development, Job Creation and Trade\. YJK acknowlegdes the support from the National Science and Technology Council of Taiwan \(NSTC\) through grants 113\-2112\-M\-002\-033\-MY3 and 115\-2124\-M\-001\-015\. We acknowledge the National Center for High\-Performance Computing in Taiwan for the computational resources used in this research\.
## Appendix AMutual\-information dependence in Rydberg samples
For occupations at sitesiiandjj, we estimate the plug\-in pairwise mutual information
I^ij=∑a,b∈\{0,1\}p^ij\(a,b\)lnp^ij\(a,b\)p^i\(a\)p^j\(b\),\\widehat\{I\}\_\{ij\}=\\sum\_\{a,b\\in\\\{0,1\\\}\}\\widehat\{p\}\_\{ij\}\(a,b\)\\ln\\frac\{\\widehat\{p\}\_\{ij\}\(a,b\)\}\{\\widehat\{p\}\_\{i\}\(a\)\\widehat\{p\}\_\{j\}\(b\)\},\(15\)where zero\-probability terms are omitted\. For sites at lattice coordinates\(xi,yi\)\(x\_\{i\},y\_\{i\}\)and\(xj,yj\)\(x\_\{j\},y\_\{j\}\), we define the radial separation in lattice\-spacing units by the Pythagorean distance
rij=\(xi−xj\)2\+\(yi−yj\)2\.r\_\{ij\}=\\sqrt\{\(x\_\{i\}\-x\_\{j\}\)^\{2\}\+\(y\_\{i\}\-y\_\{j\}\)^\{2\}\}\.\(16\)We average the pairwise values over all open\-boundary pairs at a common separationrij=rr\_\{ij\}=rto obtainI^\(r\)\\widehat\{I\}\(r\)for the same six detunings presented in Section[II](https://arxiv.org/html/2609.20912#S2), using 10,000 configurations sourced from Quantum Monte Carlo data\.
At criticality, the mutual\-information scaling ansatz has the algebraic form below, withη\\etathe mutual\-information scaling exponent\. As a finite\-array diagnostic, we fit the 19 radial shells in the interval1≤r≤501\\leq r\\leq\\sqrt\{50\}to
I\(r\)=Ar−η\.I\(r\)=Ar^\{\-\\eta\}\.\(17\)ThuslogI\(r\)=logA−ηlogr\\log I\(r\)=\\log A\-\\eta\\log ris fitted by ordinary least squares over the positive radial\-shell estimates\. In the near\-critical reference region, the high fit qualities in Figure[6](https://arxiv.org/html/2609.20912#A1.F6)show that this ansatz describes the measured radial dependence well\. The fittedη\\etais nevertheless a descriptive finite\-array parameter; it is not identified with a universal critical exponent\.
Figure 6:Fit quality of the mutual\-information power\-law ansatz over detuning\. Points give theR2R^\{2\}of the log–log fit in Equation \([17](https://arxiv.org/html/2609.20912#A1.E17)\); the dashed line marks the critical point\.The fitted exponent decreases fromη=2\.24\\eta=2\.24atδ/Ω=0\.93\\delta/\\Omega=0\.93\(R2=0\.978R^\{2\}=0\.978\) toη=1\.56\\eta=1\.56at1\.051\.05\(R2=0\.970R^\{2\}=0\.970\) andη=1\.08\\eta=1\.08at1\.171\.17\(R2=0\.942R^\{2\}=0\.942\)\. Within the adopted ansatz, this corresponds to a progressively slower radial decay on approaching and passing the critical point\. The valuesR2=0\.978R^\{2\}=0\.978,0\.9700\.970and0\.9420\.942atδ/Ω=0\.93\\delta/\\Omega=0\.93,1\.051\.05and1\.171\.17, respectively, provide direct support for the use of the mutual information ansatz as a proxy for correlation length\. At the more remote settings, the fit is less stable:η=1\.57\\eta=1\.57atδ/Ω=−0\.36\\delta/\\Omega=\-0\.36\(R2=0\.465R^\{2\}=0\.465\),η=0\.163\\eta=0\.163at1\.521\.52\(R2=0\.499R^\{2\}=0\.499\), andη=1\.71\\eta=1\.71at2\.942\.94\(R2=0\.335R^\{2\}=0\.335\)\. The power\-law parameters therefore provide a compact description of the measured radial dependence, but do not establish a universal exponent or a sharply defined phase transition\. This limitation is physically expected for a finite, thermally mixed quantum system\.
## References
- \[1\]Y\. Bahri, E\. Dyer, J\. Kaplan, J\. Lee, and U\. Sharma\(2024\)Explaining neural scaling laws\.Proceedings of the National Academy of Sciences121\(27\),pp\. e2311878121\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2311878121)Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p4.1)\.
- \[2\]M\. Barkeshli, A\. Alfarano, and A\. Gromov\(2026\)On the origin of neural scaling laws: from random graphs to natural language\.arXiv preprint arXiv:2601\.10684\.Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p4.1)\.
- \[3\]D\. S\. Berman and A\. G\. Stapleton\(2026\)A path to natural language through tokenisation and transformers\.arXiv preprint\.External Links:2601\.03368,[Link](https://arxiv.org/abs/2601.03368)Cited by:[footnote 1](https://arxiv.org/html/2609.20912#footnote1)\.
- \[4\]A\. Browaeys and T\. Lahaye\(2020\)Many\-body physics with individually controlled Rydberg atoms\.Nature Physics16,pp\. 132–142\.External Links:[Document](https://dx.doi.org/10.1038/s41567-019-0733-z)Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p7.1)\.
- \[5\]F\. Cagnetta, A\. Raventós, S\. Ganguli, and M\. Wyart\(2026\)Deriving neural scaling laws from the statistics of natural language\.arXiv preprint\.External Links:2602\.07488,[Link](https://arxiv.org/abs/2602.07488)Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p4.1)\.
- \[6\]Z\. Chen, O\. Mayné i Comas, Z\. Jin, D\. Luo, and M\. Soljačić\(2025\)L2ML^\{2\}M: mutual information scaling law for long\-context language modeling\.arXiv preprint\.External Links:2503\.04725Cited by:[§I\.1](https://arxiv.org/html/2609.20912#S1.SS1.p6.1)\.
- \[7\]S\. Ebadi, T\. T\. Wang, H\. Levine, A\. Keesling, G\. Semeghini, A\. Omran, D\. Bluvstein, R\. Samajdar, H\. Pichler, W\. W\. Ho, S\. Choi, S\. Sachdev, M\. Greiner, V\. Vuletić, and M\. D\. Lukin\(2021\)Quantum phases of matter on a 256\-atom programmable quantum simulator\.Nature595\(7866\),pp\. 227–232\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-03582-4),[Link](https://doi.org/10.1038/s41586-021-03582-4),ISSN 1476\-4687Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p6.1)\.
- \[8\]W\. Ebeling and T\. Pöschel\(1994\)Entropy and long\-range correlations in literary english\.Europhysics Letters26\(4\),pp\. 241–246\.External Links:[Document](https://dx.doi.org/10.1209/0295-5075/26/4/001)Cited by:[§I\.1](https://arxiv.org/html/2609.20912#S1.SS1.p6.1)\.
- \[9\]D\. Fitzek, Y\. Hong Teoh, H\. P\. Cyrus Fung, G\. A\. Dagnew, E\. Merali, M\. Schuyler Moss, B\. MacLellan, and R\. G\. Melko\(2025\)RydbergGPT\.Machine Learning: Science and Technology6\(4\),pp\. 045057\.External Links:[Document](https://dx.doi.org/10.1088/2632-2153/ae1d0b),[Link](https://doi.org/10.1088/2632-2153/ae1d0b)Cited by:[§I\.1](https://arxiv.org/html/2609.20912#S1.SS1.p1.1),[§I](https://arxiv.org/html/2609.20912#S1.p5.1),[§I](https://arxiv.org/html/2609.20912#S1.p8.1),[Data Availability Statement](https://arxiv.org/html/2609.20912#Sx1.p1.1)\.
- \[10\]D\. Fitzek, Y\. H\. Teoh, H\. P\. Fung, G\. A\. Dagnew, E\. Merali, M\. S\. Moss, B\. MacLellan, and R\. G\. Melko\(2024\)PennyLane Datasets for RydbergGPT\.Note:[https://pennylane\.ai/datasets/rydberggpt](https://pennylane.ai/datasets/rydberggpt)Cited by:[Figure 1](https://arxiv.org/html/2609.20912#S1.F1),[Figure 1](https://arxiv.org/html/2609.20912#S1.F1.4),[§I](https://arxiv.org/html/2609.20912#S1.p9.1),[§IV](https://arxiv.org/html/2609.20912#S4.p5.1)\.
- \[11\]J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, J\. W\. Rae, O\. Vinyals, and L\. Sifre\(2022\)Training compute\-optimal large language models\.External Links:2203\.15556,[Link](https://arxiv.org/abs/2203.15556)Cited by:[§I\.2](https://arxiv.org/html/2609.20912#S1.SS2.p1.1),[§I\.2](https://arxiv.org/html/2609.20912#S1.SS2.p2.2),[§I](https://arxiv.org/html/2609.20912#S1.p3.1)\.
- \[12\]M\. Kalinowski, R\. Samajdar, R\. G\. Melko, M\. D\. Lukin, S\. Sachdev, and S\. Choi\(2022\)Bulk and boundary quantum phase transitions in a square Rydberg atom array\.Physical Review B105,pp\. 174417\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevB.105.174417)Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p8.1),[§II](https://arxiv.org/html/2609.20912#S2.p2.1)\.
- \[13\]J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei\(2020\)Scaling laws for neural language models\.arXiv preprint arXiv:2001\.08361\.External Links:2001\.08361Cited by:[§I\.2](https://arxiv.org/html/2609.20912#S1.SS2.p1.1),[§I\.2](https://arxiv.org/html/2609.20912#S1.SS2.p2.2),[§I](https://arxiv.org/html/2609.20912#S1.p3.1),[footnote 2](https://arxiv.org/html/2609.20912#footnote2)\.
- \[14\]G\. Kingsley Zipf\(1932\)Selected studies of the principle of relative frequency in language\.Harvard University Press\.External Links:ISBN 9780674432048,[Link](http://dx.doi.org/10.4159/harvard.9780674434929),[Document](https://dx.doi.org/10.4159/harvard.9780674434929)Cited by:[footnote 1](https://arxiv.org/html/2609.20912#footnote1)\.
- \[15\]W\. Li\(1989\)Mutual information functions of natural language texts\.External Links:[Link](https://api.semanticscholar.org/CorpusID:15270663)Cited by:[§I\.1](https://arxiv.org/html/2609.20912#S1.SS1.p6.1)\.
- \[16\]W\. Li\(1990\)Mutual information functions versus correlation functions\.Journal of Statistical Physics60,pp\. 823–837\.External Links:[Link](https://api.semanticscholar.org/CorpusID:10721009)Cited by:[§III](https://arxiv.org/html/2609.20912#S3.p2.1)\.
- \[17\]H\. W\. Lin and M\. Tegmark\(2017\)Critical behavior in physics and probabilistic formal languages\.Entropy19\(7\),pp\. 299\.External Links:[Document](https://dx.doi.org/10.3390/e19070299)Cited by:[§I\.1](https://arxiv.org/html/2609.20912#S1.SS1.p6.1)\.
- \[18\]G\. Peraza Coppola, M\. Helias, and Z\. Ringel\(2025\)Renormalization group for deep neural networks: universality of learning and scaling laws\.arXiv preprint arXiv:2510\.25553\.Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p4.1)\.
- \[19\]M\. Saffman, T\. G\. Walker, and K\. Mølmer\(2010\)Quantum information with Rydberg atoms\.Reviews of Modern Physics82,pp\. 2313–2363\.External Links:[Document](https://dx.doi.org/10.1103/RevModPhys.82.2313)Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p7.1)\.
- \[20\]R\. Samajdar, W\. W\. Ho, H\. Pichler, M\. D\. Lukin, and S\. Sachdev\(2020\)Complex density wave orders and quantum phase transitions in a model of square\-lattice Rydberg atom arrays\.Physical Review Letters124,pp\. 103601\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevLett.124.103601)Cited by:[§I](https://arxiv.org/html/2609.20912#S1.p8.1)\.
- \[21\]C\. E\. Shannon\(1948\)A mathematical theory of communication\.The Bell System Technical Journal27\(3\),pp\. 379–423\.External Links:[Document](https://dx.doi.org/10.1002/j.1538-7305.1948.tb01338.x)Cited by:[§I\.1](https://arxiv.org/html/2609.20912#S1.SS1.p6.1)\.
- \[22\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin\(2023\)Attention is all you need\.External Links:1706\.03762,[Link](https://arxiv.org/abs/1706.03762)Cited by:[§I\.1](https://arxiv.org/html/2609.20912#S1.SS1.p1.1),[§I](https://arxiv.org/html/2609.20912#S1.p1.1)\.相似文章
神经语言模型的缩放规律
基础性实证研究,展示了语言模型性能与模型规模、数据集大小和计算预算之间的幂律缩放关系,对最优训练资源分配和样本效率有重要启示。
将量子算子与大语言模型对齐
本文介绍了一种将幺正算子映射到大语言模型潜在空间的方法,实现了量子电路合成以及语言条件化的门约束指定,并在Clifford+T电路合成上取得了与现有方法相竞争的结果。
训练前预测:粒子物理基础模型的缩放定律
本文证明,在小规模Transformer模型上拟合的缩放定律能够准确预测在粒子物理喷注数据上训练的更大模型的损失,从而在大规模训练之前将计算预算转化为预期的物理性能。他们发布了五个预训练模型以及完整的训练方案。
论大型语言模型缩放指数的微小性
本文讨论了大型语言模型的小缩放指数,认为它们在能源资源方面指示了一种不可持续的状态。还探讨了'pedestal effect',并类比流体湍流以评论数据的平滑性。
Muon优化器的谱缩放定律
本文首次系统研究了大语言模型训练过程中Muon优化器动量矩阵奇异值谱的行为规律,发现了在不同模型规模(77M至2.8B参数)下清晰的幂律缩放关系。研究结果为从业者提供了有理论依据、感知层级的Newton–Schulz迭代配置指南,在前沿规模下无需额外计算即可保持正交归一化质量。