The Silent Freeze: Predicting When Low-Precision Training Stops Learning

arXiv cs.LG Papers

Summary

This paper identifies a deterministic, predictable freeze in low-precision training when gradient updates fall below half the unit in the last place of the weight, and demonstrates that the condition can be forecast from a high-precision trajectory alone, without low-precision data.

arXiv:2607.09800v1 Announce Type: new Abstract: Training in reduced floating-point precision can silently halt learning: when a gradient-descent weight update falls below half the unit in the last place (ULP) of the weight, it rounds away and that coordinate freezes while its gradient is still nonzero. The freeze is deterministic, governed by a per-coordinate half-ULP condition, and predictable from a high-precision trajectory and the target mantissa length alone, without low-precision data. In a small GPT trained under the standard AdamW-plus-cosine recipe with bf16-equivalent stored weights, training proceeds normally and then permanently freezes just past mid-run, within four steps of the a-priori prediction. In a $124$-million-parameter GPT-2 transformer whose weights are constrained to the $8$-bit floating-point grid after every optimizer step, with no master weights, the dense weights freeze at initialization in both fp8 formats -- predicted \emph{a priori} from an fp32 reference -- and validation loss plateaus while full precision keeps improving. Stochastic rounding removes the persistent freeze, and the same reference predicts that too. The condition transfers across frozen-feature regression, a mantissa-truncation emulator spanning $128\times$ in precision, small networks, and a CNN on MNIST: a computable axis of low-precision training, not diffuse noise.
Original Article
View Cached Full Text

Cached at: 07/14/26, 04:14 AM

# Predicting When Low-Precision Training Stops Learning
Source: [https://arxiv.org/html/2607.09800](https://arxiv.org/html/2607.09800)
###### Abstract

Training in reduced floating\-point precision can silently halt learning: when a gradient\-descent weight update falls below half the unit in the last place \(ULP\) of the weight, it rounds away and that coordinate freezes while its gradient is still nonzero\. The freeze is deterministic, governed by a per\-coordinate half\-ULP condition, and predictable from a high\-precision trajectory and the target mantissa length alone, without low\-precision data\. In a small GPT trained under the standard AdamW\-plus\-cosine recipe with bf16\-equivalent stored weights, training proceeds normally and then permanently freezes just past mid\-run, within four steps of the a\-priori prediction\. In a124124\-million\-parameter GPT\-2 transformer whose weights are constrained to the88\-bit floating\-point grid after every optimizer step, with no master weights, the dense weights freeze at initialization in both fp8 formats—predicted*a priori*from an fp32 reference—and validation loss plateaus while full precision keeps improving\. Stochastic rounding removes the persistent freeze, and the same reference predicts that too\. The condition transfers across frozen\-feature regression, a mantissa\-truncation emulator spanning128×128\\timesin precision, small networks, and a CNN on MNIST: a computable axis of low\-precision training, not diffuse noise\.

Reduced\-precision arithmetic is now pervasive in training, and the bit width keeps falling:bfloat16is routine and88\-bit floating point has reached frontier transformer training\. As the bits have fallen, independent groups across scientific machine learning, low\-precision optimizer design, and distributed post\-training have collided with the same unnamed failure: training runs on, the loss is finite, the gradient is nonzero—and weights silently stop moving \(Table[1](https://arxiv.org/html/2607.09800#S0.T1)\)\. Each group engineered a workaround; none could say in advance when the failure would strike\. Low\-precision failure modes are usually treated as diffuse “noise\.” This one is not noise at all but a deterministic, per\-coordinate, and*predictable*event: the gradient\-underflow freeze\. It is, concretely, the failure mode that fp32 master weights exist to prevent and that pure low\-mantissa weight updates expose\. In gradient descent a weight is updated asw←w−η​∇ℓw\\leftarrow w\-\\eta\\nabla\\ell; in finite precision the result is rounded back onto the floating\-point grid\. When the updateη​\|∇ℓ\|\\eta\|\\nabla\\ell\|is smaller than half the unit in the last place \(ULP\) ofww, the rounded result iswwitself—the coordinate stops moving while its true gradient is still nonzero, and once enough coordinates cross this threshold the weight vector stops advancing even though the loss gradient has not vanished\. As∥w∥\\lVert w\\rVertgrows during training its ULP grows with it, so a learning run that began well below the resolution limit can cross into the frozen regime partway through and halt\.

Table 1:Six independent reports of the same arithmetic hazard: weight updates too small to survive the destination grid\.Each is a manifestation of half\-ULP update swamping—in stored\-weight training, mixed\-precision pipelines, or cast\-and\-synchronize post\-training—though the surrounding mechanisms differ \(Related work\)\. None predicts when the per\-coordinate condition\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)engages, and only one\[[9](https://arxiv.org/html/2607.09800#bib.bib9)\]states the condition, as a gate for a fix rather than a forecast\.Our central claim is that this freeze is*governed and predictable*\. The governing event is per\-coordinate: coordinateiifreezes when\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\), and the fraction of coordinates this flags is the frozen fraction\. When the weights are homogeneous in scale this per\-coordinate rule aggregates into a single dimensionless number,

ρ≡η​∥g∥ε​∥w∥,\\rho\\equiv\\frac\{\\eta\\,\\lVert g\\rVert\}\{\\varepsilon\\,\\lVert w\\rVert\},\(1\)the ratio of the GD step size to the mantissa resolutionε​∥w∥\\varepsilon\\lVert w\\rVert\(withε=2−\(m\+1\)\\varepsilon=2^\{\-\(m\+1\)\}formmmantissa bits\), whose onset is the thresholdρ⋆=𝒪​\(1\)\\rho^\{\\star\}=\\mathcal\{O\}\(1\); where the weights span many scales the per\-coordinate condition remains exact while this pooled scalar loses its meaning\. More strongly, the freeze*time*τ⋆\\tau^\{\\star\}follows a priori from a single high\-precision trajectory and the target mantissa length alone—no low\-precision run is needed to predict when low precision will fail\. We establish the mechanism in a controlled frozen\-feature setting, then show it transfers unchanged to a mantissa\-truncation emulator across128×128\\timesin precision and to a trainable\-feature network under plain full\-batch GD, that it bites mid\-training in a small GPT under the standard AdamW\-plus\-cosine recipe at a step named in advance, and finally that it governs a124124\-million\-parameter GPT\-2 transformer whose weights are constrained to the88\-bit floating\-point grid at every step under its standard nanoGPT\-style optimizer and schedule—where the dense weights freeze at initialization, predicted*a priori*from a cheap fp32 reference\. The freeze is thus a property of the arithmetic, not of any one model or scale\.

Because it is a property of the floating\-point update itself, the freeze matters wherever gradient descent runs in reduced precision; six independent reports have encountered the underlying half\-ULP swamping in one form or another \(Table[1](https://arxiv.org/html/2607.09800#S0.T1)and Related work\)\. The broader aim is to move precision selection from folklore to forecast\. A recent mechanistic account of the bf16 FlashAttention loss explosion\[[10](https://arxiv.org/html/2607.09800#bib.bib10)\]closed by calling for the same treatment of other precision\-failure channels; this paper answers that call for the weight\-update channel—a control parameter, a derived threshold, and an a\-priori forecast of when a given format under a given schedule stops learning, available before any low\-precision hardware is committed\. The first of our test systems, a frozen\-feature linear regression, we draw from a study of predictability inference\[[1](https://arxiv.org/html/2607.09800#bib.bib1),[2](https://arxiv.org/html/2607.09800#bib.bib2)\]\.

## Related work

Documented encounters with the update\-channel freeze\.Weight\-update underflow has been hit independently across several communities—in all but one case\[[9](https://arxiv.org/html/2607.09800#bib.bib9)\]without being named, and in no case made predictive \(Table[1](https://arxiv.org/html/2607.09800#S0.T1)\)\. In scientific machine learning, L\-BFGS optimization of physics\-informed networks halts prematurely when fp32 updates become numerically invisible\[[4](https://arxiv.org/html/2607.09800#bib.bib4)\], and fp16 weight updates underflow to zero in mixed\-precision training\[[6](https://arxiv.org/html/2607.09800#bib.bib6)\]\. In low\-precision optimizer design, pure\-bf16 Adam diverges or stalls without an fp32 master copy\[[5](https://arxiv.org/html/2607.09800#bib.bib5)\], and a production PyTorch operator carries1616extra mantissa bits precisely so that small updates survive the bf16 grid\[[7](https://arxiv.org/html/2607.09800#bib.bib7)\]\. In distributed reinforcement\-learning post\-training, roughly99%99\\%of per\-step bf16 weight updates are invisible after the cast at typical learning rates—there the higher\-precision learner keeps training, and the invisibility of the cast updates is exploited for communication compression\[[12](https://arxiv.org/html/2607.09800#bib.bib12)\]\. And in the controlled comparison of Ref\.\[[9](https://arxiv.org/html/2607.09800#bib.bib9)\], naive bf16 fine\-tuning of a T5 transformer collapses outright\. Each report answers the symptom with a workaround, and one of them—Yu\[[9](https://arxiv.org/html/2607.09800#bib.bib9)\], taken up next—states the per\-coordinate half\-ULP condition itself as the gate for its fix\. What no report supplies, including that one, is a derived population threshold for the onset or an a\-priori prediction of*when*the condition will engage\.

Closest prior art\.Yu\[[9](https://arxiv.org/html/2607.09800#bib.bib9)\]independently identifies the same arithmetic dead zone—“swamping,” in the classical vocabulary of floating\-point summation\[[18](https://arxiv.org/html/2607.09800#bib.bib18)\]—stating the half\-ULP condition\|δ\|<12​u​\(θ\)\|\\delta\|<\\tfrac\{1\}\{2\}\\,u\(\\theta\)and gating updates on the per\-coordinate ratio\|δ\|/u​\(θ\)\|\\delta\|/u\(\\theta\), which in our notation is\|Δ​wi\|/ULPm​\(wi\)\|\\Delta w\_\{i\}\|/\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)\. This raw ULP ratio—writeρYu≡\|Δ​wi\|/ULPm​\(wi\)\\rho\_\{\\mathrm\{Yu\}\}\\equiv\|\\Delta w\_\{i\}\|/\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)—is*not*our pooledρ\\rho: it relates to our per\-coordinateρi≡\|Δ​wi\|/\(ε​\|wi\|\)\\rho\_\{i\}\\equiv\|\\Delta w\_\{i\}\|/\(\\varepsilon\|w\_\{i\}\|\)byρYu=2ϕi−1​ρi\\rho\_\{\\mathrm\{Yu\}\}=2^\{\\phi\_\{i\}\-1\}\\rho\_\{i\}, withϕi=frac​\(log2⁡\|wi\|\)\\phi\_\{i\}=\\mathrm\{frac\}\(\\log\_\{2\}\|w\_\{i\}\|\)locatingwiw\_\{i\}in its binade, so Yu’s dead\-zone gateρYu<12\\rho\_\{\\mathrm\{Yu\}\}<\\tfrac\{1\}\{2\}is exactly our per\-coordinate conditionρi<2−ϕi\\rho\_\{i\}<2^\{\-\\phi\_\{i\}\}\. The pooled scalarρ=η​∥g∥/\(ε​∥w∥\)\\rho=\\eta\\lVert g\\rVert/\(\\varepsilon\\lVert w\\rVert\)and its derived thresholdρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}are separate homogeneous\-population objects \(Methods\), not a relabeling of his ratio\. That work’s contribution is a*fix*: conditionally amplifying swamped updates to±1\\pm 1ULP so that fully low\-precision training proceeds without master weights\. Ours is orthogonal and complementary: we derive the population thresholdρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}from the arithmetic, show that the freeze*time*follows a priori from a single high\-precision trajectory plus the target mantissa length—before any low\-precision run is launched—and validate that prediction from controlled regression through a124124\-million\-parameter fp8 transformer\. Prediction and fix compose naturally: a forecast of which tensors freeze, and when, is what would let any such mitigation be deployed selectively rather than blanket\.

Sibling failure channels, not instances\.Two adjacent low\-precision failure modes share the finite\-mantissa root cause but are distinct mechanisms, and we do not claim them as instances of the weight freeze\. Qiu and Yao\[[10](https://arxiv.org/html/2607.09800#bib.bib10)\]trace the long\-standing bf16 FlashAttention loss explosion to biased rounding in the forward\-pass attention product, compounding along shared low\-rank directions \(see also\[[26](https://arxiv.org/html/2607.09800#bib.bib26)\]\); that is a forward\-pass accumulation bias with the opposite signature \(sudden explosion\), whereas the freeze studied here is a rounding event in the weight update itself, whose signature is silent stagnation\. Topollai and Choromanska\[[11](https://arxiv.org/html/2607.09800#bib.bib11)\]model the staleness of quantized*optimizer states*, a third channel in which the EMA memory, not the weight, stops updating\. Qiu and Yao close by calling for the same mechanistic treatment of other precision\-failure channels, formats, and scales; the present work supplies exactly that for the weight\-update channel, through fp8 at transformer scale\.

Mitigations and the missing forecast\.The standard defenses bracket the freeze without predicting it\. An fp32 master copy of the weights\[[13](https://arxiv.org/html/2607.09800#bib.bib13)\], the default that made bf16 training routine\[[14](https://arxiv.org/html/2607.09800#bib.bib14)\]and that eight\-bit training has retained from its first demonstrations\[[15](https://arxiv.org/html/2607.09800#bib.bib15)\]through frontier practice\[[16](https://arxiv.org/html/2607.09800#bib.bib16)\], keeps the accumulating update where its ULP is fine enough to register\. Stochastic rounding removes the deterministic dead zone entirely, with a mature convergence theory\[[8](https://arxiv.org/html/2607.09800#bib.bib8),[19](https://arxiv.org/html/2607.09800#bib.bib19),[20](https://arxiv.org/html/2607.09800#bib.bib20),[21](https://arxiv.org/html/2607.09800#bib.bib21)\]and a growing practice at the edge\[[22](https://arxiv.org/html/2607.09800#bib.bib22)\]; unbiased quantization schemes for four\-bit formats push the same idea further\[[23](https://arxiv.org/html/2607.09800#bib.bib23),[24](https://arxiv.org/html/2607.09800#bib.bib24)\], and unit scaling renormalizes tensors into the representable sweet spot\[[25](https://arxiv.org/html/2607.09800#bib.bib25)\]\. Scaling studies quantify the aggregate cost of precision\[[17](https://arxiv.org/html/2607.09800#bib.bib17)\]\. What none of these supplies is the run\-specific forecast—*which*coordinates of*which*tensors will freeze, at*which*step, in a given format under a given schedule—obtainable before any low\-precision hardware is committed\. That forecast is this paper’s contribution\.

## Results

### The half\-ULP freeze condition

The freeze is a property of the arithmetic, not of any one model\. In a format withmmmantissa bits the representable spacing atwiw\_\{i\}isULPm​\(wi\)=2⌊log2⁡\|wi\|⌋−m\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)=2^\{\\lfloor\\log\_\{2\}\|w\_\{i\}\|\\rfloor\-m\}, and round\-to\-nearest returnswiw\_\{i\}unchanged whenever\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\); the fraction of coordinates this flags is the frozen fraction\. Written in units ofρ\\rho\[Eq\. \([1](https://arxiv.org/html/2607.09800#S0.E1)\), withε=2−\(m\+1\)\\varepsilon=2^\{\-\(m\+1\)\}\], the condition becomesρi<2−ϕi\\rho\_\{i\}<2^\{\-\\phi\_\{i\}\}, whereϕi=frac​\(log2⁡\|wi\|\)∈\[0,1\)\\phi\_\{i\}=\\mathrm\{frac\}\(\\log\_\{2\}\|w\_\{i\}\|\)\\in\[0,1\)locateswiw\_\{i\}within its binade, so the half\-ULP boundary is a band\[12,1\]\[\\tfrac\{1\}\{2\},1\]exactly one bit wide\. When theϕi\\phi\_\{i\}are equidistributed—which holds to high accuracy in our runs—the median of this band is2−1/22^\{\-1/2\}, fixingρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}with no fitted parameter \(full derivation in Methods\); the measured thresholds \(0\.710\.71emulator,0\.720\.72–0\.740\.74networks\) sit on this value\. Because the band is one bit wide independently of the coordinate count, the freeze onset has an intrinsic, size\-independent width—a crossover rather than a sharpening transition, which we verify directly below\.

As training proceeds∥w∥\\lVert w\\rVertgrows and∥g∥\\lVert g\\rVertdecays, so in these runsρ\\rhofalls toward and through an𝒪​\(1\)\\mathcal\{O\}\(1\)threshold, at which the iterate freezes\. We measure the freeze directly as the fraction of weight coordinates left bitwise unchanged by a step,f=⟨𝟙​\[fl​\(w−η​g\)=w\]⟩f=\\langle\\mathbb\{1\}\[\\,\\mathrm\{fl\}\(w\-\\eta g\)=w\\,\]\\rangle, and define the freeze timeτ⋆\\tau^\{\\star\}as the step at whichfffirst reaches12\\tfrac\{1\}\{2\}\. In frozen\-feature GD on realbfloat16hardware, the freeze is logged at sparse checkpoints:f=0f=0at step100100andf=1\.0f=1\.0by step300300, after which the iterate stays frozen for the final90%90\\%of training \(Fig\.[1](https://arxiv.org/html/2607.09800#Sx2.F1)a\)\. The hardware checkpoints bracket rather than pin the crossing; readingρ\\rhoat the first fully\-frozen point givesρ⋆≈0\.72\\rho^\{\\star\}\\approx 0\.72, consistent with the densely\-logged emulator \(ρ⋆=0\.71\\rho^\{\\star\}=0\.71, below\)\.

### The freeze time is predictable a priori

The decisive test is whether we can predictτ⋆\\tau^\{\\star\}*without ever running the low\-precision model*\. We compute a single high\-precision \(fp64\) trajectory\{wt,gt\}\\\{w\_\{t\},g\_\{t\}\\\}, which carries nomm\-bit rounding, and from it alone form the predicted frozen fraction at each step,f^t​\(m\)=⟨𝟙​\[η​\|gt,i\|<12​ULPm​\(wt,i\)\]⟩\\hat\{f\}\_\{t\}\(m\)=\\langle\\mathbb\{1\}\[\\,\\eta\|g\_\{t,i\}\|<\\tfrac\{1\}\{2\}\\,\\mathrm\{ULP\}\_\{m\}\(w\_\{t,i\}\)\\,\]\\rangle, using only the target mantissa lengthmm\. The predicted freeze time is wheref^t\\hat\{f\}\_\{t\}crosses12\\tfrac\{1\}\{2\}\. This predictor contains zero low\-precision data; if the freeze were a complicated dynamical artifact its*shape*could easily be wrong\. It is not: across precisions and learning rates the predicted and measuredτ⋆\\tau^\{\\star\}agree across a384×384\\timesrange of freeze times \(1111to42224222steps\)\. On the trainable\-network population \(n=55n=55clean cells\) the median predicted/measured ratio is1\.001\.00and∼95%\\sim\\\!95\\%\(94\.5%94\.5\\%\) of cells fall within15%15\\%; the mantissa\-emulator points corroborate the same identity across the128×128\\timesprecision span \(Fig\.[1](https://arxiv.org/html/2607.09800#Sx2.F1)b\)\. Thesen=55n=55cells are not statistically independent—they share feature seeds and high\-precision trajectories across mantissa lengths and rates—so we read the spread as a coverage range, not as5555independent draws\. The prediction tracks individual trajectories, not just averages: where one run freezes two orders of magnitude earlier than its neighbors, the a\-priori predictor follows it there\.

### Transfer across precision, rate, and architecture

The freeze is a property of the arithmetic, so it should not depend on the model\. We test this on four increasingly different systems\. \(i\) Frozen\-feature linear GD on realbfloat16, above\. \(ii\) A mantissa\-truncation emulator that performs all arithmetic in fp64 \(wide exponent, no denormal underflow at our magnitudes\) but rounds every weight tommbits after each step, sweepingm∈\{5,6,7,8,10,12\}m\\in\\\{5,6,7,8,10,12\\\}—a128×128\\timesrange inε\\varepsilon\. Validated against realbfloat16atm=7m=7, the emulator reproduces the freeze curve to RMSE0\.0020\.002, so it is a faithful variable\-resolution stand\-in\. Across the full grid its freeze threshold isρ⋆=0\.71\\rho^\{\\star\}=0\.71\(coefficient of variation0\.060\.06over128×128\\timesin precision and4×4\\timesin rate, at a single feature seed; the across\-seed spread is quantified for the trainable network below\)\. \(iii\) A two\-layertanh\\tanhnetwork with*both*layers trainable, trained by plain full\-batch GD on a teacher–student regression—a non\-convex problem with no frozen features and no closed\-form target\. The freeze survives intact: the emulator again matches realbfloat16\(RMSE0\.0070\.007\), the threshold sits atρ⋆≈0\.74\\rho^\{\\star\}\\approx 0\.74\(clean across\-seed median; within\-cell across\-seed coefficient of variation, median0\.210\.21\), and the a\-priori prediction holds at the accuracy quoted above \(Fig\.[1](https://arxiv.org/html/2607.09800#Sx2.F1)c\)\. The same𝒪​\(1\)\\mathcal\{O\}\(1\)threshold and the same a\-priori predictor thus carry over from a convex linear readout to the tested non\-convex trainable network—one teacher–studenttanh\\tanhregression—establishing transfer to a genuinely different architecture rather than universality over all models\. \(iv\) A convolutional network \(∼105\\sim\\\!10^\{5\}parameters: two convolution–pool stages and two fully connected layers\) trained by GD with the cross\-entropy loss on real MNIST digits—a recognizable architecture on real data, not a synthetic target\. We probe it two ways\. In the mantissa\-truncation emulator, sweepingm∈\{5,6,7,8,10\}m\\in\\\{5,6,7,8,10\\\}over three seeds \(1515precision–seed cells\), the a\-priori predictor is exact: from the fp64 trajectory and2−\(m\+1\)2^\{\-\(m\+1\)\}alone the predicted freeze step equals the measured one in every cell \(median ratio1\.0001\.000, maximum error0%0\\%\)\. On genuinebfloat16hardware \(them=7m=7point, three seeds\) the emulator reproduces the freeze curve to RMSE0\.020\.02and the same freeze appears on schedule\. Here the pooledρ\\rhois no longer a single meaningful number—the layers sit at very different scales, and the tiny bias vectors, which never freeze, inflate anL2L^\{2\}aggregate, so the pooled value sprawls fromρ≈2\\rho\\approx 2toρ≈170\\rho\\approx 170across the1515emulator cells, far from theO​\(1\)O\(1\)threshold and uninformative about the freeze\. The per\-coordinate rule has no such problem\. The exact half\-ULP test\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)recovers the frozen set by construction; the nontrivial finding is that even its*scalarized*form—flagging coordinateiiby the single constant thresholdρi<ρ⋆=1/2\\rho\_\{i\}<\\rho^\{\\star\}=1/\\sqrt\{2\}, with no per\-coordinate binade correction—still tracks the measured frozen fraction to within0\.00110\.0011across all1515cells \(median agreement0\.00040\.0004\), even as that fraction itself ranges from0\.510\.51to0\.770\.77from cell to cell\. The pooled scalar fails by two orders of magnitude where the per\-coordinate diagnostic holds, confirming that the freeze lives at the coordinate level, not in any aggregate\. The functional cost is consistent with the freeze\-*time*\-not\-*loss*asymmetry discussed below: all seeds freeze, but whether the freeze caught the network before or after it had learned makes the final test accuracy range, on realbfloat16hardware across seeds, from chance \(0\.110\.11\) to near the fp64 value \(0\.930\.93–0\.950\.95\), even though the freeze step itself was predicted exactly\.

### Transfer across optimizer, loss, and architecture

The freeze condition\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)is stated per coordinate for an*arbitrary*updateΔ​w\\Delta w, so neither the freeze nor its a\-priori predictor should depend on howΔ​w\\Delta wis produced\. We test this by perturbing one axis at a time away from the baseline \(two\-layertanh\\tanh, squared error, full\-batch GD\): the optimizer \(GD→\{\\to\}Adam, whose adaptive step removes∥g∥\\lVert g\\rVertfrom the update entirely\), the loss \(squared error→\{\\to\}cross\-entropy, under which∥w∥\\lVert w\\rVertgrows without bound so the freeze bar12​ε​\|wi\|\\tfrac\{1\}\{2\}\\varepsilon\|w\_\{i\}\|*rises*over training rather than falls\), the depth \(two→\{\\to\}three weight layers\), and the architecture \(the convolutional network on MNIST\)\. In every case the predictor holds: the median predicted/measured freeze\-time ratio stays between0\.9990\.999and1\.031\.03across all six systems \(Table[2](https://arxiv.org/html/2607.09800#Sx2.T2)\), with9090–96%96\\%of cells within15%15\\%and the CNN exact in every cell\. The protocols and per\-cell tables are in Methods and the Supplementary Information\. The freeze is thus a property of the rounded update, not of any optimizer, loss, or architecture\.

Table 2:The a\-priori freeze\-time predictor transfers across optimizer, loss, depth, and architecture\.Each row perturbs one axis away from the two\-layer squared\-error full\-batch\-GD baseline\. The predictor uses only a single fp64 trajectory and the target mantissa length—no low\-precision data; the ratio is predicted/measured freeze time over the non\-degenerate cells of a mantissa sweep \(m∈\{5,6,7,8,10,12\}m\\in\\\{5,6,7,8,10,12\\\}\), pooled over seeds\. The per\-coordinate half\-ULP condition is identical in every row; only the rule producingΔ​w\\Delta wchanges\.Cells are precision–seed combinations on a mantissa sweep \(emulator rowsm∈\{5,6,7,8,10,12\}m\\in\\\{5,6,7,8,10,12\\\}over three learning rates; the CNN row usesm∈\{5,6,7,8,10\}m\\in\\\{5,6,7,8,10\\\}\);nncounts the*non\-degenerate*cells, i\.e\. those that freeze after step11\. Cells that freeze at step11\(low\-mm, low\-η\\etacorners where the first update already underflows\) carry no dynamical information and are excluded; see Methods\. The network rows pool four seeds over them×ηm\\times\\etagrid; the mantissa emulator is a single feature seed; the CNN pools three seeds over five precisions \(1515precision–seed cells\)\. The CNN frozen\-fraction prediction is exact in every precision–seed cell \(maximum error0%0\\%\); “—” not separately tabulated\.

### A crossover, not a critical transition

A statistical\-physics transition sharpens without bound as the system grows; a crossover keeps a fixed intrinsic width\. We test which the freeze is by finite\-size scaling\. The intensive width of the onset is the spread of the per\-coordinate log\-ratioui=log2⁡\|Δ​wi\|−log2⁡\|wi\|u\_\{i\}=\\log\_\{2\}\|\\Delta w\_\{i\}\|\-\\log\_\{2\}\|w\_\{i\}\|at the half\-freeze epoch \(this is the distribution whose threshold crossing*is*the frozen fraction; the seed\-averagedffcurve narrows only through trivialN−1/2N^\{\-1/2\}sampling and is not used\)\. Sweeping the input\-layer width overN=128N=128to40964096coordinates \(32×32\\times\) at three precisions and six seeds, the width does not shrink: a fitw∼N−aw\\sim N^\{\-a\}givesa=−0\.05a=\-0\.05to−0\.08\-0\.08acrossmm\(if anything a slight broadening\), consistent with the one\-bit floor derived above\. The freeze is therefore a fixed\-width crossover with a computable location, not a critical transition—which is the honest reading of the per\-coordinate mechanism\.

![Refer to caption](https://arxiv.org/html/2607.09800v1/x1.png)Figure 1:Theρ\\rho\-controlled precision freeze \(protocol in Methods\)\. \(a\) On realbfloat16hardware the fraction of weight coordinates left bitwise unchanged per step reaches1\.01\.0by step∼300\\sim\\\!300\(sparse checkpoints\) and stays there; the iterate freezes once the update falls below half a ULP, bracketing the threshold atρ⋆≈0\.72\\rho^\{\\star\}\\approx 0\.72\. \(b\) The freeze time is predictable*a priori*: predicted vs\. measuredτ⋆\\tau^\{\\star\}from a single fp64 trajectory plus the target mantissa length, over a384×384\\timesrange \(1111–42224222steps\); the identity statistic \(median ratio1\.001\.00,∼95%\\sim\\\!95\\%within15%15\\%\) is for the trainable\-network population, with mantissa\-emulator points corroborating across precision\. \(c\) Transfer: the freeze thresholdρ⋆\\rho^\{\\star\}stays𝒪​\(1\)\\mathcal\{O\}\(1\)across frozen\-feature linear GD, the mantissa\-truncation emulator \(128×128\\timesin precision\), and a trainable two\-layer network under plain full\-batch GD\.
### Mid\-training freeze in a real GPT under the standard recipe

Before moving to fp8 scale we ask whether the freeze survives contact with a realistic training recipe—mini\-batch AdamW with warmup, cosine decay, weight decay, and gradient clipping—or whether schedule and optimizer dynamics suppress it\. We train a four\-layer,0\.80\.8\-million\-parameter character\-level GPT \(nanoGPT defaults\) with master weights*off*: the stored weights are constrained tommmantissa bits by the validated emulator after each AdamW step, with all compute and Adam moments in fp32, anchored against a run whose weights are stored in realtorch\.bfloat16\(protocol in Methods\)\. The frozen\-fraction trajectory has the mechanism’s signature shape \(Fig\.[2](https://arxiv.org/html/2607.09800#Sx2.F2)\): the cold initialization freezes, warmup*melts*the freeze \(minimumf≈0\.21f\\approx 0\.21\), and, as cosine decay shrinks the update while the weights grow, the network re\-freezes—permanently—in mid\-training\. At the bf16\-equivalentm=7m=7the persistent freeze \(first step after whichf≥12f\\geq\\tfrac\{1\}\{2\}for the rest of training\) lands at steps16441644and16661666of30003000in the two seeds; the a\-priori predictions from the fp32 reference alone are16401640and16691669—within four steps, with no low\-precision data\. The emulated trajectory matches the real\-bf16 anchor to RMSE<0\.01<0\.01in both seeds, and the higher\-precision control atm=10m=10neither freezes nor is predicted to \(final frozen fraction0\.350\.35measured,0\.350\.35predicted\)\. This is the silent version of the failure: the loss is finite, gradients flow, the recipe is untouched—and a persistent majority of the weights—more than half on every subsequent step—stops moving a little past mid\-run, at a step named in advance by a run that never saw low precision\. \(Atm=5m=5the network freezes at initialization, the degenerate regime we exclude throughout; at bf16*with*master weights the freeze never fires, which is precisely the workaround this failure mode necessitates\.\)

![Refer to caption](https://arxiv.org/html/2607.09800v1/x2.png)Figure 2:The freeze bites mid\-training in a real GPT under the standard recipe, at a step named in advance\.Frozen fractionffof the 2D \(decay\-group\) weights per AdamW step for a four\-layer character\-level GPT \(nanoGPT defaults: AdamW, warmup100100, cosine decay, weight decay, gradient clipping\) with master weights off\. Realbfloat16stored weights \(solid\) and the validatedm=7m=7emulator \(dashed\) agree to RMSE<0\.01<0\.01; warmup melts the cold\-start freeze, and cosine decay re\-freezes the network permanently at step16441644\(persistentf≥12f\\geq\\tfrac\{1\}\{2\}, dotted line\)\. The vertical line marks the a\-priori prediction, step16401640, computed from the fp32 reference trajectory and the target mantissa length alone, before any low\-precision run\. Them=10m=10control neither freezes nor is predicted to\.
### The freeze at transformer scale in fp8

The controlled systems above isolate the mechanism; the question for practice is whether it survives at the scale and in the format where low\-precision training is actually being pushed\. We test the freeze in a124124\-million\-parameter GPT\-2 transformer \(1212layers,1212heads, width768768, context10241024\) with weights constrained to the88\-bit floating\-point grid under its standard nanoGPT\-style recipe \(AdamW, cosine schedule with warmup, weight decay0\.10\.1, gradient clipping\), in a deliberately short diagnostic run\. The stored weights are re\-quantized onto the fp8 grid after every step, with*no*fp32 master copy—the regime the freeze condition addresses; the optimizer moments and the forward/backward compute stay in fp32, so the stored\-weight rounding is isolated from any matmul or accumulation effect \(Methods\)\. We report both OCP fp8 formats, E5M2 \(primary,99\.5%99\.5\\%of weights normal at initialization\) and E4M3 \(99\.0%99\.0\\%normal under a fixed power\-of\-two scale\); a cheap fp32 reference, which reproduces the fp64 freeze condition at this precision, supplies the a\-priori prediction, so no fp64 run is needed at scale\.

At GPT\-2’s standard learning rate \(6×10−46\\times 10^\{\-4\}\), training on OpenWebText, the dense attention and feed\-forward weights—the compute\-bearing matmul parameters—freeze at*initialization*: the per\-step bitwise\-frozen fraction of dense weights is already a majority at the first step, and over600600steps the most active single step still leaves8484–93%93\\%of dense weights bitwise unchanged \(\>99%\>\\\!99\\%by the end\), in both formats and both seeds\. A per\-step minority keeps moving in early training—enough to pull the loss well below chance before it plateaus \(below\)—but without fp32 master weights the dense weights are frozen from the outset rather than training freely\. The a\-priori predictor, built from the fp32 reference and the target fp8 grid alone, tracks the measured dense frozen\-fraction trajectory to RMSE0\.0100\.010–0\.0110\.011across formats and seeds; the freeze is corpus\-insensitive, a byte\-pair\-encoded Shakespeare corpus giving similar numbers \(RMSE0\.0130\.013–0\.0140\.014; Methods\)\. The freeze\-at\-initialization is not a knife\-edge in the learning rate: at the sampled1×1\\timesand20×20\\timesrates the dense weights stay frozen at initialization, and at the sampled80×80\\timesrate the freeze instead engages early in training, after a brief initial melt, where the predictor again follows the trajectory\.

We report the dense matmul weights on their own\. The∼38\\sim\\\!38\-million\-parameter token embedding \(about31%31\\%of the decayed weights at this vocabulary\) receives a gradient at only the few token rows present in each minibatch, so its untouched rows freeze for a sparsity reason distinct from update underflow; pooling it in would inflate the frozen fraction for the wrong reason\. Broken out separately it freezes earliest, confirming the distinction, but the headline rests on the dense weights, where the freeze*is*the half\-ULP update underflow\. As in the small systems the half\-freeze*step*is fragile in fp8, so we lead with the frozen\-fraction trajectory and the post\-warmup frozen fraction, not the crossing step \(Methods\)\.

The freeze has a measurable downstream cost\. Over a longer30003000\-step OpenWebText run at the standard rate, the held\-out validation loss of the frozen fp8 model descends through the early steps—while a per\-step minority of weights still moves—and then plateaus, at6\.706\.70nats \(E5M2\) and6\.226\.22nats \(E4M3\), whereas the fp32 model under the*identical*recipe and schedule keeps improving, reaching5\.045\.04and still descending at step30003000\. The held\-out gap therefore widens at every logged evaluation rather than closing, from\+0\.9\+0\.9to\+1\.65\+1\.65nats \(E5M2\) and\+0\.4\+0\.4to\+1\.17\+1\.17nats \(E4M3\) by step30003000, with a parallel∼8\\sim\\\!8\-point next\-token\-accuracy gap, in both seeds \(Fig\.[3](https://arxiv.org/html/2607.09800#Sx2.F3)a; Methods\)\. Because both runs share the same cosine schedule, the schedule alone—ordinary cosine decay absent low\-precision rounding—cannot account for the difference; the plateau is instead associated with, and mechanistically consistent with, the measured half\-ULP freeze of the stored weights\. We are careful not to overstate this: the frozen model is not pinned at chance—enough coordinates move in early training to carry it well below the chance lossln⁡50304≈10\.83\\ln 50304\\approx 10\.83—and the gap is measured over30003000steps, not proven to grow without bound\. It is the freeze\-*time*\-not\-*loss*asymmetry made concrete at scale: the freeze onset is predicted a priori, the loss it ultimately costs is not\.

Finally, the freeze is defeated by the standard fix, and that too is predictable\. Replacing round\-to\-nearest with*stochastic*rounding—which rounds a sub\-ULP update up with probability equal to its fraction of a ULP\[[8](https://arxiv.org/html/2607.09800#bib.bib8)\]—gives every sub\-ULP update a nonzero probability of moving its coordinate, removing the deterministic persistence that keeps a coordinate frozen while its successive updates remain sub\-ULP\. The never\-moved fraction of dense weights—coordinates whose stored value is bitwise unchanged across the entire run—collapses to≈2×10−3\\approx 2\\times 10^\{\-3\}\(E5M2\) and≈9×10−5\\approx 9\\times 10^\{\-5\}\(E4M3\) under stochastic rounding, from a round\-to\-nearest baseline in which a majority of dense coordinates never move \(for the primary E5M2 format about31%31\\%move at least once, mostly in early training;49%49\\%for E4M3; Methods\)\. Under stochastic rounding essentially every dense coordinate moves at least once\. The same cheap fp32 reference predicts this collapse a priori, to RMSE<0\.004<0\.004across formats, seeds, and both corpora \(Fig\.[3](https://arxiv.org/html/2607.09800#Sx2.F3)b; Methods\)—so the a\-priori predictor describes not only when the freeze occurs but also when the mitigation removes it\.

![Refer to caption](https://arxiv.org/html/2607.09800v1/x3.png)Figure 3:The freeze at transformer scale \(GPT\-2\-124M, fp8, no master weights\)\.\(a\) Downstream cost on OpenWebText \(default rate6×10−46\\times 10^\{\-4\},30003000steps\): the held\-out validation loss of the fp32 reference keeps descending while the frozen E5M2 and E4M3 models plateau early, so the gap widens to\+1\.65\+1\.65\(E5M2\) and\+1\.17\+1\.17\(E4M3\) nats and has not converged\. Both arms share the same cosine schedule, so the schedule alone cannot explain the difference; the plateau is instead mechanistically consistent with the measured stored\-weight freeze\. The frozen models sit well below the chance lossln⁡50304\\ln 50304\(dotted\), i\.e\. not pinned at chance\. Faint curves are the second seed\. \(b\) Stochastic rounding defeats the freeze and the cheap fp32 reference predicts it: the never\-moved fraction of dense weights—bitwise unchanged across the run—collapses to≈2×10−3\\approx 2\\times 10^\{\-3\}\(E5M2\) and≈9×10−5\\approx 9\\times 10^\{\-5\}\(E4M3\) under stochastic rounding, and the a\-priori predictor \(hatched\) matches the measured collapse to RMSE<0\.004<0\.004\. Bars show the two fp8 formats, averaged over both seeds \(OWT; the Shakespeare\-corpus values are tabulated in the Supplement\)\. The never\-moved fraction is the cumulative persistence observable \(bitwise unchanged across*all*steps\), distinct from the per\-step frozen fraction reported in the fp8 freeze trajectories \(Methods and SI\)\.

## Discussion

The gradient\-underflow freeze is a deterministic, per\-coordinate failure mode of low\-precision GD, and our central result is that it is*predictable*: a single fp64 trajectory and the target mantissa length fix the freeze time a priori, across a384×384\\timesrange and from controlled systems to a transformer\-scale fp8 test\. This reframes reduced\-precision “noise” as a governed crossover with a computable location and a thresholdρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}fixed by the arithmetic\.

The freeze also gives a first\-principles account of a standard mixed\-precision practice\. An fp32 master copy of the weights—updated in fp32 and cast down only for the forward and backward passes—is exactly what defeats the half\-ULP freeze: the update accumulates in the fp32 master copy, whose ULP is fine enough that successive sub\-grid increments are not each discarded but build up until the cast\-down value changes\. It is distinct from the underflow that loss scaling targets, which rescues gradient values that would flatten to zero in the narrow exponent range of fp16; bf16 shares fp32’s exponent range yet still freezes, because this is a mantissa\-resolution effect in the weight update, not an exponent\-range effect in the gradient\. The freeze is, in this sense, the failure mode that master weights are for, and naming it separates the two classic low\-precision underflows and the mitigation each requires\. Stochastic rounding is the complementary fix: rather than carrying the update in a finer format, it gives every sub\-ULP update a nonzero probability of landing, removing the freeze’s deterministic, absorbing character—and our experiments confirm it removes the persistent freeze in fp8, predictably\. When neither is used, the freeze exacts a downstream cost: the frozen fp8 transformer’s loss plateaus early while a full\-precision run under the same recipe keeps improving, the gap widening over training\. Consistent with the time\-not\-loss ceiling, the freeze onset is anticipated from a high\-precision run while the loss it costs is not\.

We are careful about what is and is not established\. This is*one*mechanism—the half\-ULP weight freeze—shown to transfer across the architectures we tested, not two independent phenomena that happen to agree, and not yet a survey of all models\. Its reach is bounded in an instructive way: the predictor fixes the freeze*time*, not the freeze\-induced loss\. The final low\-precision loss is the high\-precision loss at the freeze step plus a cumulative rounding penalty that grows with the learning rate and has no high\-precision proxy, so predicting*when*precision bites needs only a high\-precision run, whereas predicting*how much*it costs requires the low\-precision run itself\.

Two limitations bound the claim\. First, the freeze is established on a single GPU*vendor*; because it is a property of IEEEbfloat16arithmetic rather than of any device, it reproduces across hardware—on two NVIDIA architectures \(RTX 4090 and RTX 5060\) the logged bf16 freeze curve and weight norm are identical at every recorded checkpoint \(frozen fraction and weight norm agree to RMSE0\), and the mantissa\-truncation emulator \(validated to RMSE0\.0020\.002against real hardware\) is a portable stand\-in across precisions\. Non\-NVIDIA bf16 and hardware\-fused update paths remain to be checked\. Second, the controlled demonstrations are full\-batch GD on small synthetic networks, a small MNIST subset, and small\-initialization regimes chosen to place the freeze mid\-training; the character\-level GPT shows the same mid\-training freeze under the standard mini\-batch AdamW recipe at0\.80\.8M parameters, and the fp8 GPT\-2 result extends the mechanism to a124124\-million\-parameter transformer under its standard mini\-batch recipe, but there the freeze occurs at initialization and is logged over the first hundreds of steps on a standard byte\-pair\-encoded corpus, not across a full pretraining run\. The arithmetic mechanism is general, but the empirical reach is these tested regimes, not a survey of production\-scale training\.

That the mechanism is real and consequential is corroborated by independent practice\. At least six recent efforts run into weight\-update underflow and respond to the symptom \(Table[1](https://arxiv.org/html/2607.09800#S0.T1)\): switching to fp64, realigning the optimizer with the bf16 grid, carrying extra mantissa bits so small updates survive, amplifying swamped updates to a full ULP, or exploiting the invisibility for communication compression—but none predicts when the freeze bites, and the one report that names the arithmetic condition\[[9](https://arxiv.org/html/2607.09800#bib.bib9)\]deploys it as a gate for a fix, not as a forecast\. The per\-coordinate freeze condition and its a\-priori predictor supply exactly that missing diagnosis—which precision will stall, and at which step—from a single high\-precision run, the gap these workarounds leave open\.

Gradient underflow thus joins the small set of low\-precision failure modes that can be anticipated rather than merely observed: a single high\-precision run fixes when training will freeze, before any reduced\-precision hardware is touched\.

## Methods

### Experimental systems

The freeze is demonstrated across four learning systems\. \(i\)*Frozen\-feature linear GD\.*A frozen random feature mapϕ​\(𝐱\)=tanh⁡\(𝐱​W1\)\\phi\(\\mathbf\{x\}\)=\\tanh\(\\mathbf\{x\}W\_\{1\}\)withW1∼𝒩​\(0,σw2\)W\_\{1\}\\sim\\mathcal\{N\}\(0,\\sigma\_\{w\}^\{2\}\),σw=0\.1\\sigma\_\{w\}=0\.1, and no bias projects short temporal contexts of a dynamical signal top=8p=8dimensions; a linear readout is fit by GD \(η=5×10−3\\eta=5\\times 10^\{\-3\},τmax=3000\\tau\_\{\\max\}=3000steps\) minimizing1−cos⁡\(y−y^\)1\-\\cos\(y\-\\hat\{y\}\)for a circular target \(Kuramoto\-oscillator data\) or squared error for a real target \(Lorenz\-63 data\)\. Only the readout is trained; the fixed features make the high\-precision trajectory a clean reference for the a\-priori test\. This system is drawn from a predictability\-inference study\[[1](https://arxiv.org/html/2607.09800#bib.bib1),[2](https://arxiv.org/html/2607.09800#bib.bib2)\]; the full data\-generation protocol is in the Supplementary Information\[[27](https://arxiv.org/html/2607.09800#bib.bib27)\]\. \(ii\)*Mantissa emulator*, below\. \(iii\)*Trainable two\-layertanh\\tanhnetwork*under plain full\-batch GD on a teacher–student regression—a non\-convex problem with no frozen features and no closed\-form target\. \(iv\)*Convolutional network on MNIST*: two3×33\\times 3convolution layers \(1→16→321\\\!\\to\\\!16\\\!\\to\\\!32channels\), each with2×22\\times 2max\-pooling, then fully connected layers1568→64→101568\\\!\\to\\\!64\\\!\\to\\\!10\(∼1\.06×105\\sim\\\!1\.06\\times 10^\{5\}parameters\), cross\-entropy loss, full\-batch GD, with small initialization so the weights grow into the freeze rather than starting on it\.

The freeze condition uses the round\-to\-nearest half\-ULPε=2−\(m\+1\)\\varepsilon=2^\{\-\(m\+1\)\}formmmantissa bits, equal to2−82^\{\-8\}forbfloat16; the high\-precision reference trajectory is computed in fp64 \(ε≈2\.2×10−16\\varepsilon\\approx 2\.2\\times 10^\{\-16\}\)\. The reduced\-precision runs that establish the freeze fall into two classes that share one weight representation\. On genuinebfloat16hardware the weights are stored inbfloat16and the update is applied directly in that format, so the rounding is the hardware’s\. The mantissa\-truncation emulator and the trainable\-network bridge instead compute the raw update in fp64 \(wide exponent, no denormal underflow at our magnitudes\) and then round the updated weight tommmantissa bits after each step; this isolates the half\-ULP weight rounding from any exponent or accumulation effect, and atm=7m=7it reproduces the realbfloat16freeze curve to RMSE0\.0020\.002\. Both are the regime of low\-precision stored weights without an fp32 master copy, not mixed\-precision training with an fp32 master copy and fp32 accumulation, which would not exhibit the same weight\-level freeze\. \(Some supplemental controls deliberately use fp32 master weights; these are labeled as such and are not part of the freeze evidence\.\)

### Mid\-training GPT freeze \(bf16\-equivalent\)

The mid\-training demonstration uses the nanoGPT reference implementation: a four\-layer,44\-head,d=128d=128character\-level GPT \(block size128128, vocabulary6565, no biases,≈0\.8\\approx 0\.8M parameters,811,136811\{,\}136of them in the 2D decay\-group matrices\) trained on the Shakespeare character corpus for30003000steps with the standard recipe intact: AdamW\(β1,β2\)=\(0\.9,0\.95\)\(\\beta\_\{1\},\\beta\_\{2\}\)=\(0\.9,0\.95\), peak learning rate6×10−46\\times 10^\{\-4\},100100\-step linear warmup, cosine decay to6×10−56\\times 10^\{\-5\}, weight decay0\.10\.1, gradient clipping at1\.01\.0, batch size3232, two seeds\. Standard “bf16” training in this codebase is autocast—fp32 master weights with bf16 compute—under which the freeze never fires; to expose the stored\-weight freeze the master copy is turned off, and the stored weight itself is constrained tommmantissa bits with the validated emulator \(rounded after each AdamW step; all compute and Adam moments in fp32, which yields freeze predictions bitwise identical to an fp64 reference\[[27](https://arxiv.org/html/2607.09800#bib.bib27)\]\)\. The frozen fraction pools the 2D decay\-group weights—a coordinate counts as frozen on a step if its stored value is bitwise unchanged by the full AdamW step—andτ⋆\\tau^\{\\star\}is the*persistent*freeze step: the first step after whichf≥12f\\geq\\tfrac\{1\}\{2\}for the remainder of training, robust to the cold\-init/warmup transient\. The a\-priori arm runs the same recipe in fp32 with no rounding and flags, at each step and for each targetmm, the coordinates with\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\,\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\); the anchor arm stores the weights in realtorch\.bfloat16with master weights off\. The measured freeze steps carry a few steps of run\-to\-run nondeterminism from the fused\-AdamW kernels; the a\-priori predictions are deterministic\.

### fp8 transformer at scale

The scale test is a GPT\-2 transformer \(nlayer=12n\_\{\\text\{layer\}\}=12,nhead=12n\_\{\\text\{head\}\}=12,nembd=768n\_\{\\text\{embd\}\}=768, context10241024, vocabulary5030450304, no bias;∼124\\sim\\\!124M parameters\) trained on OpenWebText with the conventional recipe: AdamW \(β=0\.9,0\.95\\beta=0\.9,0\.95\), cosine schedule with100100warmup steps, weight decay0\.10\.1, gradient clipping1\.01\.0\. The architecture and recipe are unchanged from a reference GPT\-2 implementation; only the stored\-weight precision of the decayed weights is altered\. The weights are constrained to the88\-bit floating\-point grid by re\-quantization after every optimizer step,w←Qfp8​\(w\)w\\leftarrow Q\_\{\\text\{fp8\}\}\(w\), with no fp32 master copy; the AdamW moments and the forward/backward compute are carried in fp32\. This isolates the stored\-weight half\-ULP freeze, the paper’s mechanism, from fp8 matmul or accumulation effects, which are a separate systems axis we do not enable here—and because fp8 GEMM is not used, the run requires no fp8\-matmul hardware\. We test both OCP fp8 formats: E5M2 \(22mantissa bits,ε=2−3\\varepsilon=2^\{\-3\}\) at unit scale, the clean primary, and E4M3 \(33mantissa bits,ε=2−4\\varepsilon=2^\{\-4\}\) under a fixed power\-of\-two scales=128s=128\(an exact exponent shift\) that lifts the weights into the format’s normal band; at initialization99\.5%99\.5\\%\(E5M2\) and99\.0%99\.0\\%\(E4M3\) of weights are normal, so representation underflow does not contaminate the update\-freeze measurement\. The quantizerQfp8Q\_\{\\text\{fp8\}\}is an analytic round\-to\-nearest onto the fp8 grid; it agrees bitwise with the hardwaretorch\.float8cast over10710^\{7\}weights, so it is an exact stand\-in for the hardware rounding\. The stored weights are held as fp32 tensors constrained exactly to the representable fp8 grid \(rather than in atorch\.float8container\), so the optimizer reads and writes the same fp8 grid values the hardware would store while the fp8 GEMM path stays disabled; the bitwise validation guarantees the two are identical value sets\.

The reference is fp32, not fp64\. The freeze condition is a sign test on which side of a half\-ULP fp8 boundary the updateη​gi\\eta g\_\{i\}falls; computing\(wi,gi\)\(w\_\{i\},g\_\{i\}\)in fp32 versus fp64 perturbs each by a relative∼2−24\\sim\\\!2^\{\-24\}, so the decision changes only for the vanishing fraction of coordinates whose update lands within∼2−24\\sim\\\!2^\{\-24\}of a boundary that is itself spaced by the much coarser fp8 ULP \(2−22^\{\-2\}to2−32^\{\-3\}of a binade\)\. Empirically the substitution is exact: in a controlled sweep of reference mantissa length on the small systems, an fp32 reference predicts the freeze step bitwise\-identically to an fp64 reference \(zero step difference across all2525clean cells, and the same for a1616\-bit reference; SI\), so the freeze\-step prediction is independent of reference precision\. The cheaper fp32 trajectory therefore supplies the same a\-priori prediction, and no fp64 run is needed at scale\. We run33learning rates \(6×10−46\\times 10^\{\-4\},1\.2×10−21\.2\\times 10^\{\-2\},4\.8×10−24\.8\\times 10^\{\-2\}\)×\\times22seeds×\\times600600steps\. For each cell the fp32 reference yields the predicted frozen\-fraction trajectory and each fp8 format the measured one; we report the trajectory RMSE and the post\-warmup frozen fraction, both robust, and not the half\-freeze step, which is fragile in fp8 \(the frozen fraction jitters across12\\tfrac\{1\}\{2\}\)\. The freeze is corpus\-insensitive—being an initialization\-versus\-update\-magnitude effect of the first steps—and a byte\-pair\-encoded Shakespeare corpus reproduces the same dense freeze\-at\-initialization and trajectory RMSE; we report OpenWebText as the primary corpus\. Weights are grouped by function—attention, feed\-forward, token embedding, positional embedding—and the headline*dense*fraction pools only the attention and feed\-forward matmul weights\. The token embedding is reported separately: at vocabulary5030450304it is∼38\\sim\\\!38M parameters \(∼31%\\sim\\\!31\\%of the decayed weights\) and receives gradients only at the token rows present in each minibatch, so its untouched rows freeze through gradient sparsity rather than update underflow, a distinct effect that must not inflate the dense headline\.

*Downstream loss\.*To tie the freeze to a number a reader cares about, we run a longer single\-rate experiment at the standard rate \(6×10−46\\times 10^\{\-4\},22seeds,30003000steps on OpenWebText\) and evaluate held\-out quality on a fixed validation set shared identically by the fp32 reference and every fp8 model: the cross\-entropy validation loss and the next\-token accuracy, against the chance lossln⁡50304≈10\.83\\ln 50304\\approx 10\.83\. The frozen fp8 models and the fp32 reference use the same data order, schedule, and seed, so the only difference is the stored\-weight precision\. We report the loss trajectories and the fp8\-minus\-fp32 gap; because the freeze fixes the*time*but not the cumulative rounding penalty, this downstream gap is measured, not predicted a priori\.

*Stochastic rounding\.*As a mitigation test we replace the round\-to\-nearest store with unbiased stochastic rounding: an update landing a fractionrrof a ULP above a grid point rounds up with probabilityrrand down otherwise, so its expectation is exact and sub\-ULP updates are no longer discarded\[[8](https://arxiv.org/html/2607.09800#bib.bib8)\]\. The persistence observable is then the*never\-moved*fraction—coordinates whose stored value is bitwise unchanged across all steps\. Under round\-to\-nearest this is a majority of dense coordinates \(for E5M2 about31%31\\%move at least once, mostly early, so it is well below the high per\-step frozen fraction and is not the same observable;49%49\\%for E4M3\); under stochastic rounding it should collapse toward zero\. The fp32 reference predicts this collapse a priori asp^never​\(t\)=⟨∏s≤t\(1−rs,i\)⟩\\hat\{p\}\_\{\\text\{never\}\}\(t\)=\\langle\\prod\_\{s\\leq t\}\(1\-r\_\{s,i\}\)\\ranglefrom the reference update fractionsrs,ir\_\{s,i\}alone; we report the measured never\-moved fraction, this prediction, and their RMSE, per format and seed on both corpora\.

### The freeze condition and theρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}threshold

Write the per\-coordinate update asΔ​wi=η​gi\\Delta w\_\{i\}=\\eta g\_\{i\}withg=∇wℓg=\\nabla\_\{w\}\\ell\. The exact per\-coordinate freeze condition is\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\,\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)\. WithULPm​\(wi\)=2⌊log2⁡\|wi\|⌋​2−m\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)=2^\{\\lfloor\\log\_\{2\}\|w\_\{i\}\|\\rfloor\}2^\{\-m\}andε=2−\(m\+1\)\\varepsilon=2^\{\-\(m\+1\)\}, dividing through byε​\|wi\|\\varepsilon\|w\_\{i\}\|gives the per\-coordinate ratio in units ofρ\\rho,

ρi≡\|Δ​wi\|ε​\|wi\|<2⌊log2⁡\|wi\|⌋−log2⁡\|wi\|=2−ϕi,ϕi=frac​\(log2⁡\|wi\|\)∈\[0,1\)\.\\rho\_\{i\}\\equiv\\frac\{\|\\Delta w\_\{i\}\|\}\{\\varepsilon\|w\_\{i\}\|\}<2^\{\\lfloor\\log\_\{2\}\|w\_\{i\}\|\\rfloor\-\\log\_\{2\}\|w\_\{i\}\|\}=2^\{\-\\phi\_\{i\}\},\\quad\\phi\_\{i\}=\\mathrm\{frac\}\(\\log\_\{2\}\|w\_\{i\}\|\)\\in\[0,1\)\.\(2\)A coordinate freezes precisely whenρi<2−ϕi\\rho\_\{i\}<2^\{\-\\phi\_\{i\}\}\. The right\-hand side ranges over the band\[12,1\]\[\\tfrac\{1\}\{2\},1\]asϕi\\phi\_\{i\}sweeps a binade, a band exactly one bit wide regardless of coordinate count, learning rate, or precision\. When theϕi\\phi\_\{i\}are equidistributed in\[0,1\)\[0,1\)—the generic case for weights spread across a binade—the thresholdT=2−ϕT=2^\{\-\\phi\}hasmedian​\(T\)=geomean​\(T\)=2−1/2≈0\.7071\\mathrm\{median\}\(T\)=\\mathrm\{geomean\}\(T\)=2^\{\-1/2\}\\approx 0\.7071and𝔼​\[T\]=1/\(2​ln⁡2\)≈0\.7213\\mathbb\{E\}\[T\]=1/\(2\\ln 2\)\\approx 0\.7213, so the half\-freeze threshold isρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}with no fitted parameter\. We validate the uniform\-ϕ\\phiassumption directly: feeding the measured per\-coordinateρi\\rho\_\{i\}through the analytic frozen fractionfpred=⟨clip​\(−log2⁡ρi,0,1\)⟩f\_\{\\mathrm\{pred\}\}=\\langle\\mathrm\{clip\}\(\-\\log\_\{2\}\\rho\_\{i\},0,1\)\\ranglereproduces the measured bitwise frozen fraction at the half\-freeze epoch across all widths, precisions, and seeds of the finite\-size study with mean absolute error\|Δ\|<0\.007\|\\Delta\|<0\.007\(worst case0\.090\.09, in a single small\-NNlow\-mantissa cell where coordinate count is smallest\)\. The measured thresholds—0\.7070\.707\(emulator\),0\.720\.72\(frozen\-feature hardware\),≈0\.74\\approx 0\.74\(trainable network\)—bracket this prediction\.

The aggregateρ=η​∥g∥/\(ε​∥w∥\)\\rho=\\eta\\lVert g\\rVert/\(\\varepsilon\\lVert w\\rVert\)\[Eq\. \([1](https://arxiv.org/html/2607.09800#S0.E1)\)\] is the pooled form of this per\-coordinate condition\. It is not the per\-coordinate ratio but their weight\-magnitude\-weighted root mean square:ρ=⟨ρi2⟩w\\rho=\\sqrt\{\\langle\\rho\_\{i\}^\{2\}\\rangle\_\{w\}\}, where⟨⋅⟩w\\langle\\cdot\\rangle\_\{w\}averages with weightwi2/∥w∥2w\_\{i\}^\{2\}/\\lVert w\\rVert^\{2\}\. The single scalar therefore stands in for the typical coordinate, and the half\-freeze thresholdρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}transfers from theρi\\rho\_\{i\}toρ\\rho, only under two assumptions: \(i\) the per\-coordinate ratiosρi\\rho\_\{i\}are concentrated—i\.e\. the update\-to\-weight ratio\|Δ​wi\|/\|wi\|\|\\Delta w\_\{i\}\|/\|w\_\{i\}\|is approximately coordinate\-independent, so the weighted RMS coincides with the medianρi\\rho\_\{i\}—and \(ii\) the binade phasesϕi\\phi\_\{i\}are equidistributed across the coordinates that carry the weight norm\. Both hold when the weights are homogeneous in scale and fail when they are not: where weights span many scales \(the convolutional network\) the per\-coordinate condition remains exact while the pooled scalar loses its meaning\. Across the mantissa\-emulator grid \(66values ofm×3m\\times 3learning rates, single feature seed\) the threshold isρ⋆=0\.707\\rho^\{\\star\}=0\.707with coefficient of variation0\.0550\.055\(range0\.670\.67–0\.820\.82over128×128\\timesinε\\varepsilonand4×4\\timesinη\\eta\)\. For the trainable two\-layer network the clean across\-seed median isρ⋆≈0\.74\\rho^\{\\star\}\\approx 0\.74; the within\-cell across\-seed coefficient of variation has median0\.210\.21, maximum0\.350\.35over the1212non\-degenerate cells—looser because the network’sρ\\rhopoolsW1W\_\{1\},W2W\_\{2\}, and biases at different scales\.

### A\-priori prediction of the freeze time

The non\-circular core of the result is that the freeze time follows from a single high\-precision trajectory plus the target mantissa length, with zero low\-precision data in the predictor\. We run one fp64 trajectory\{wt,gt\}\\\{w\_\{t\},g\_\{t\}\\\}and form the predicted frozen fractionf^t​\(m\)=⟨𝟙​\[η​\|gt,i\|<12​ULPm​\(wt,i\)\]⟩\\hat\{f\}\_\{t\}\(m\)=\\langle\\mathbb\{1\}\[\\eta\|g\_\{t,i\}\|<\\tfrac\{1\}\{2\}\\,\\mathrm\{ULP\}\_\{m\}\(w\_\{t,i\}\)\]\\rangle, taking the predicted freeze timeτ^⋆\\hat\{\\tau\}^\{\\star\}as the first crossing of12\\tfrac\{1\}\{2\}; the measuredτ⋆\\tau^\{\\star\}comes from an independent low\-precision \(emulated or real\) run\. We exclude degenerate cells that freeze at step11\(low\-mm, low\-η\\etacorners where the first update already underflows\), for whichτ⋆\\tau^\{\\star\}carries no dynamical information\. In the trainable\-network bridge, pooling four seeds over all non\-degenerate cells \(n=55n=55pairs\) gives a median predicted/measured ratio of1\.0001\.000\(mean1\.0191\.019\) with94\.5%94\.5\\%within15%15\\%, spanning measured freeze times from1111to42224222steps \(384×384\\times\)\. Thesen=55n=55cells share feature seeds and high\-precision trajectories across mantissa lengths and rates, so the spread is a coverage range, not5555independent draws\. The predictor is genuinely per\-trajectory: in one cell a single seed froze at step6161while its siblings froze near18001800, and the predictor tracked it to6161\. The mantissa\-emulator sweep corroborates across precision \(n=18n=18cells, median ratio1\.031\.03\)\.

### Mantissa\-truncation emulator and across\-rate invariance

To sweep precision continuously, all arithmetic is carried in fp64 \(a wide exponent, so no denormal underflow at our weight magnitudes\), but after every GD step each weight is rounded tommmantissa bits,w←roundm​\(w\)w\\leftarrow\\mathrm\{round\}\_\{m\}\(w\), withε=2−\(m\+1\)\\varepsilon=2^\{\-\(m\+1\)\}\. Settingm=7m=7reproduces thebfloat16mantissa; the sweep spansm∈\{5,6,7,8,10,12\}m\\in\\\{5,6,7,8,10,12\\\}\(128×128\\timesinε\\varepsilon\)\. The emulator isolates the mantissa\-resolution effect from the exponent range and the specific hardware rounding path\. Against realtorch\.bfloat16atm=7m=7the emulated freeze curvef​\(t\)f\(t\)matches to RMSE1\.8×10−31\.8\\times 10^\{\-3\}in the frozen\-feature linear system and6\.7×10−36\.7\\times 10^\{\-3\}in the trainable two\-layer network, and0\.0220\.022in the convolutional network—faithful in all three regimes\.

Becauseε\\varepsilonentersρ\\rhoby definition, the across\-precision collapse ofρ⋆\\rho^\{\\star\}is in part definitional and we do not claim it as a discovery\. The content lies in three places that are not built in: \(a\) at*fixed*ε\\varepsilonthe threshold is invariant across more than an order of magnitude in learning rate—sweepingη\\etaover16×16\\times\(0\.001250\.00125–0\.020\.02\) in the frozen\-featurebfloat16system leavesρ⋆\\rho^\{\\star\}flat at0\.690\.69–0\.770\.77\(mean0\.720\.72, CV0\.040\.04\); \(b\) the a\-priori predictor gets the freeze\-time*shape*right over384×384\\times, which a definitional identity does not guarantee; and \(c\) the threshold and predictor transfer across architectures\.

### Finite\-size scaling

A genuine statistical\-physics transition sharpens without bound as the number of degrees of freedomNNgrows; a crossover retains a fixed intrinsic width\. The seed\-averaged frozen fractionf​\(t\)f\(t\)narrows withNN, but only through trivialN−1/2N^\{\-1/2\}sampling; it is not the intrinsic width\. The intensive observable is the spread of the per\-coordinate log\-ratioui=log2⁡\|Δ​wi\|−log2⁡\|wi\|u\_\{i\}=\\log\_\{2\}\|\\Delta w\_\{i\}\|\-\\log\_\{2\}\|w\_\{i\}\|\(withΔ​wi=η​gi\\Delta w\_\{i\}=\\eta g\_\{i\}the fp64 update\), whose threshold crossing atc​\(m\)=−\(m\+1\)c\(m\)=\-\(m\+1\)*is*the frozen fraction by construction\. We measurew​\(N\)=stdi​\(ui\)w\(N\)=\\mathrm\{std\}\_\{i\}\(u\_\{i\}\)at the half\-freeze epoch on the input\-layer \(W1W\_\{1\}\) weights alone \(8​H8Hcoordinates sharing a single,HH\-independent initialization scale\)\. SweepingH∈\{16,32,64,128,256,512\}H\\in\\\{16,32,64,128,256,512\\\}, i\.e\.N=8​HN=8Hfrom128128to40964096coordinates \(32×32\\times\), atm∈\{6,7,8\}m\\in\\\{6,7,8\\\}and six seeds \(50005000epochs,η=0\.1\\eta=0\.1\), the intrinsic width does not shrink: a fitw∼N−aw\\sim N^\{\-a\}givesa=−0\.08a=\-0\.08\(m=6m\{=\}6\),−0\.05\-0\.05\(m=7m\{=\}7\),−0\.06\-0\.06\(m=8m\{=\}8\)—all negative, so the width is flat\-to\-broadening, consistent with the one\-bit threshold band whose spread is dominated by binade\-position jitter independent ofNN\. The freeze is therefore a fixed\-width crossover with a computable location, not a critical transition\. The half\-freezeρ\\rhostays in0\.660\.66–0\.950\.95across the full32×32\\timesrange, reconfirming the𝒪​\(1\)\\mathcal\{O\}\(1\)threshold well beyond the main\-text system sizes\.

### Transfer across optimizer, loss, and depth

The freeze condition\|Δ​wi\|<12​ULPm​\(wi\)\|\\Delta w\_\{i\}\|<\\tfrac\{1\}\{2\}\\mathrm\{ULP\}\_\{m\}\(w\_\{i\}\)is stated per coordinate for an arbitrary updateΔ​w\\Delta w, so neither the freeze nor its predictor should depend on howΔ​w\\Delta wis produced\. We perturb one axis at a time from the baseline \(two\-layertanh\\tanh, squared error, full\-batch GD\), keeping Adam moments in fp64 so only the weight is rounded\.*Optimizer\.*Adam emitsΔ​w=η​m^/\(v^\+δ\)\\Delta w=\\eta\\,\\hat\{m\}/\(\\sqrt\{\\hat\{v\}\}\+\\delta\), which removes∥g∥\\lVert g\\rVertfromρ\\rho; the per\-coordinate predictor applies unchanged\. Overn=60n=60non\-degenerate cells the predicted/measured freeze\-time ratio has median1\.0001\.000\(mean0\.9940\.994\),90%90\\%within15%15\\%\.*Loss\.*Under cross\-entropy on a teacher–student classification with10%10\\%label flips,∥w∥\\lVert w\\rVertgrows without bound, so the freeze bar12​ε​\|wi\|\\tfrac\{1\}\{2\}\\varepsilon\|w\_\{i\}\|*rises*over training—the opposite of the squared\-error case—yet the freeze and predictor survive: at depth two \(n=64n=64\) median ratio1\.0161\.016\(92%92\\%within15%15\\%\); at depth three \(n=25n=25\) median0\.9990\.999\(96%96\\%within15%15\\%\)\. The aggregateρ⋆\\rho^\{\\star\}stays𝒪​\(1\)\\mathcal\{O\}\(1\)throughout but its tightness depends on coordinate homogeneity: tight for the shallow cross\-entropy network \(medianρ⋆=0\.76\\rho^\{\\star\}=0\.76, CV0\.080\.08\), loose for Adam \(pooled median0\.960\.96, CV0\.890\.89\) because the pooled update norm is dominated by a few still\-moving high\-scale coordinates while most have frozen\. In every case the per\-coordinate predictor remains accurate \(median ratio within2%2\\%of unity\) where the pooled scalar is loose, which is why we treat the per\-coordinate condition, not the aggregate, as load\-bearing\.

*Convolutional network\.*On real MNIST \(three seeds,m∈\{5,6,7,8,10\}m\\in\\\{5,6,7,8,10\\\},300300epochs; fp64 reference0\.930\.93test accuracy\) the a\-priori predictor is exact: the predicted freeze step equals the measured one in all1515precision–seed cells \(median and mean ratio1\.0001\.000, maximum error0%0\\%\), and the emulatedm=7m=7freeze curve matches realbfloat16to RMSE0\.0220\.022\. Because the kernels, dense layers, and bias vectors live at very different scales, theL2L^\{2\}\-pooledρ\\rhois a poor summary: the never\-freezing biases dominate the aggregate, so at the half\-freeze epoch the pooled value sprawls from≈2\\approx 2to≈170\\approx 170across the1515cells, far from the threshold and uninformative about the freeze\. The per\-coordinate rule has no such problem\. Even its*scalarized*form—flagging coordinateiiby the single constantρi<ρ⋆=1/2\\rho\_\{i\}<\\rho^\{\\star\}=1/\\sqrt\{2\}, with no per\-coordinate binade correction—tracks the measured frozen fraction to within0\.00110\.0011across all1515cells \(median0\.00040\.0004\), even as that fraction itself ranges from0\.510\.51to0\.770\.77from cell to cell\. The functional outcome illustrates the freeze\-time/freeze\-loss split: all three seeds freeze on the predicted step, but the real\-bfloat16final accuracy ranges across seeds from chance \(0\.110\.11, freeze caught before learning\) to near the fp64 value \(0\.930\.93–0\.950\.95, freeze caught after\), so the freeze*time*is predicted exactly while the freeze*loss*is not\.

### Hardware and reproducibility

fp64 runs use an NVIDIA Titan V \(Volta, CC 7\.0\);bfloat16runs an RTX 5060 \(Blackwell, CC 12\.0\), both under PyTorch 2\.x with CUDA 13 and deterministic kernels\. The hardwarebfloat16freeze reproduction \(Fig\.[1](https://arxiv.org/html/2607.09800#Sx2.F1)a\) re\-runs the identical RTX 5060bfloat16GD kernel on a single candidate with per\-step logging of the weight norm, update norm, and bitwise\-frozen fraction\. Re\-running the same kernel on a second NVIDIA architecture \(RTX 4090, Ada, CC 8\.9\) gives logged freeze curves that are identical at every recorded checkpoint: the frozen fraction and weight norm agree to RMSE0\(both cross to fully frozen between steps100100and300300and report∥w∥=9\.612×10−4\\lVert w\\rVert=9\.612\\times 10^\{\-4\}thereafter\)\. We log the freeze curve and weight norm at each checkpoint rather than full weight tensors, so this establishes identity of the recorded quantities, not a tensor\-level bitwise hash; the freeze is nonetheless a deterministic property of IEEEbfloat16arithmetic, not an artifact of a particular device or kernel schedule\.

The load\-bearing quantities are accumulated per seed\. In an early version of the network bridge only the test\-error degradation was seed\-averaged, andρ⋆\\rho^\{\\star\}and the predictor ratios were inadvertently read from a single seed; multi\-seeding with per\-seed accumulators overturned two single\-seed artifacts \(the cleanρ⋆\\rho^\{\\star\}mean moved from0\.920\.92to the across\-seed value, and one alarming73%73\\%predictor miss was revealed as single\-seed noise, the other three seeds agreeing to within5%5\\%\)\. We therefore report across\-seed statistics throughout and flag that the load\-bearing quantities, not merely the headline outcome, must be multi\-seeded\. Further protocol details, theρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}derivation, the finite\-size\-scaling tables, and the per\-cell predictor tables \(including the optimizer, loss, depth, and convolutional\-network transfer experiments\) are in the Supplementary Information\[[27](https://arxiv.org/html/2607.09800#bib.bib27)\]\.

## Data availability

The data that support the findings of this study—the high\-precision and reduced\-precision GD trajectories, the mantissa\-emulator and trainable\-network freeze grids, the a\-priori predictor outputs, the cross\-architecturebfloat16logs, the mid\-training GPT freeze trajectories, the fp8 GPT\-2 frozen\-fraction trajectories and per\-group records, and the CNN/MNIST freeze and per\-coordinate records—are archived in a Zenodo repository \(DOI 10\.5281/zenodo\.20726265\)\. A machine\-readable manifest maps each displayed statistic and figure panel to its source data file and generating script\. Dataset locations are resolved through the manifest and a configurable data root; absolute paths appearing inside archived run metadata are provenance records, not load paths, and no analysis or figure script depends on them\. The MNIST digits are the standard public dataset\[[3](https://arxiv.org/html/2607.09800#bib.bib3)\]; a download\-and\-verify script reproduces the exact raw files against published SHA\-256 checksums\.

## Code availability

The code that generates the trajectories, the mantissa\-truncation emulator, the a\-priori freeze\-time predictor, and the finite\-size\-scaling and cross\-architecture analyses is archived with the data under the Zenodo DOI above\. The release includes a pinned environment specification, the MNIST fetch\-and\-checksum utility, the per\-claim provenance manifest, and a one\-command path that regenerates the figures from the stored data\.

## References

- \[1\]G\. Boffetta, M\. Cencini, M\. Falcioni, and A\. Vulpiani, “Predictability: a way to characterize complexity,”*Phys\. Rep\.*356, 367–474 \(2002\)\.
- \[2\]E\. Bollt, “On explaining the surprising success of reservoir computing forecaster of chaos? The universal machine learning dynamical system with contrast to VAR and DMD,”*Chaos*31, 013108 \(2021\)\.
- \[3\]Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner, “Gradient\-based learning applied to document recognition,”*Proc\. IEEE*86, 2278–2324 \(1998\), doi: 10\.1109/5\.726791\.
- \[4\]C\. Xu, D\. Liu, A\. Nassereldine, and J\. Xiong, “FP64 is all you need: rethinking failure modes in physics\-informed neural networks,” arXiv:2505\.10949 \(2025\)\.
- \[5\]X\. Liang, S\. Loeschcke, M\. Toftrup, and A\. Anandkumar, “M\+Adam: low\-precision training via mantissa–exponent optimization,”*International Conference on Machine Learning \(ICML 2026, Poster\)*;[https://icml\.cc/virtual/2026/poster/63374](https://icml.cc/virtual/2026/poster/63374)\.
- \[6\]J\. Hayford, J\. Goldman\-Wetzler, E\. Wang, and L\. Lu, “Speeding up and reducing memory usage for scientific machine learning via mixed precision,”*Comput\. Methods Appl\. Mech\. Eng\.*428, 117093 \(2024\); arXiv:2401\.16645\.
- \[7\]imoneoi, “bf16\-fused\-adam: abfloat16fused Adam operator for PyTorch,” software,[https://github\.com/imoneoi/bf16\_fused\_adam](https://github.com/imoneoi/bf16_fused_adam)\(2024\)\.
- \[8\]S\. Gupta, A\. Agrawal, K\. Gopalakrishnan, and P\. Narayanan, “Deep learning with limited numerical precision,” in*Proceedings of the 32nd International Conference on Machine Learning \(ICML\)*, 1737–1746 \(2015\);[https://arxiv\.org/abs/1502\.02551](https://arxiv.org/abs/1502.02551)\.
- \[9\]Z\. Yu, “Escaping the swamping zone: fully low\-precision training via relative ULP noise injection,” in*Proceedings of the 2026 International Conference on Artificial Intelligence and Control \(CAIC ’26\)*, ACM, 568–572 \(2026\), doi: 10\.1145/3807246\.3807335\.
- \[10\]H\. Qiu and Q\. Yao, “Why low\-precision transformer training fails: an analysis on flash attention,”*International Conference on Learning Representations \(ICLR 2026\)*; arXiv:2510\.04212\.
- \[11\]K\. Topollai and A\. Choromanska, “Understanding quantization of optimizer states in LLM pre\-training: dynamics of state staleness and effectiveness of state resets,” arXiv:2603\.16731 \(2026\)\.
- \[12\]E\. Miahi and E\. Belilovsky, “Understanding and exploiting weight update sparsity for communication\-efficient distributed RL,” arXiv:2602\.03839 \(2026\)\.
- \[13\]P\. Micikevicius, S\. Narang, J\. Alben, G\. Diamos, E\. Elsen, D\. Garcia, B\. Ginsburg, M\. Houston, O\. Kuchaiev, G\. Venkatesh, and H\. Wu, “Mixed precision training,” in*International Conference on Learning Representations \(ICLR\)*\(2018\); arXiv:1710\.03740\.
- \[14\]D\. Kalamkar*et al\.*, “A study of BFLOAT16 for deep learning training,” arXiv:1905\.12322 \(2019\)\.
- \[15\]N\. Wang, J\. Choi, D\. Brand, C\.\-Y\. Chen, and K\. Gopalakrishnan, “Training deep neural networks with 8\-bit floating point numbers,” in*Advances in Neural Information Processing Systems 31 \(NeurIPS\)*\(2018\)\.
- \[16\]DeepSeek\-AI, “DeepSeek\-V3 technical report,” arXiv:2412\.19437 \(2024\)\.
- \[17\]T\. Kumar, Z\. Ankner, B\. F\. Spector, B\. Bordelon, N\. Muennighoff, M\. Paul, C\. Pehlevan, C\. Ré, and A\. Raghunathan, “Scaling laws for precision,” arXiv:2411\.04330 \(2024\)\.
- \[18\]N\. J\. Higham,*Accuracy and Stability of Numerical Algorithms*, 2nd ed\. \(SIAM, Philadelphia, 2002\)\.
- \[19\]M\. P\. Connolly, N\. J\. Higham, and T\. Mary, “Stochastic rounding and its probabilistic backward error analysis,”*SIAM J\. Sci\. Comput\.*43, A566–A585 \(2021\)\.
- \[20\]L\. Xia, S\. Massei, and M\. E\. Hochstenbach, “On the convergence of the gradient descent method with stochastic fixed\-point rounding errors under the Polyak–Łojasiewicz inequality,”*Comput\. Optim\. Appl\.*90, 753–799 \(2025\); arXiv:2301\.09511\.
- \[21\]M\. Kwun, D\. Morwani, C\. H\. Su, S\. Gil, N\. Anand, and S\. Kakade, “LOTION: smoothing the optimization landscape for quantized training,” arXiv:2510\.08757 \(2025\)\.
- \[22\]T\. Liu, M\. Andronic, D\. Gündüz, and G\. A\. Constantinides, “Training with fewer bits: unlocking edge LLMs training with stochastic rounding,” in*Findings of the Association for Computational Linguistics: EMNLP 2025*, 14531–14546 \(2025\), doi: 10\.18653/v1/2025\.findings\-emnlp\.784; arXiv:2511\.00874\.
- \[23\]A\. Tseng, T\. Yu, and Y\. Park, “Training LLMs with MXFP4,” arXiv:2502\.20586 \(2025\)\.
- \[24\]A\. Panferov, E\. Schultheis, S\. Tabesh, and D\. Alistarh, “Quartet II: accurate LLM pre\-training in NVFP4 by improved unbiased gradient estimation,” arXiv:2601\.22813 \(2026\)\.
- \[25\]C\. Blake, D\. Orr, and C\. Luschi, “Unit scaling: out\-of\-the\-box low\-precision training,” in*Proceedings of the 40th International Conference on Machine Learning \(ICML\)*, 2548–2576 \(2023\)\.
- \[26\]A\. Golden, S\. Hsia, F\. Sun, B\. Acun, B\. Hosmer, Y\. Lee, Z\. DeVito, J\. Johnson, G\.\-Y\. Wei, D\. Brooks, and C\.\-J\. Wu, “Is flash attention stable?,” arXiv:2405\.02803 \(2024\)\.
- \[27\]See Supplementary Information for the full experimental protocol and data\-generation details, the mantissa\-truncation emulator and its validation, theρ⋆=1/2\\rho^\{\\star\}=1/\\sqrt\{2\}derivation, the finite\-size\-scaling tables, the a\-priori predictor protocol and per\-cell tables, the mid\-training GPT per\-seed table, and the transfer experiments \(Adam, cross\-entropy, depth, and the convolutional network\)\.

Similar Articles