Nothing Changed but the Model: CellFill -- Bounded In-Cell Learning for Bit-Identical, Revocable Updates to Quantized LLMs

arXiv cs.LG Papers

Summary

Proposes in-cell learning for bit-identical, revocable updates to quantized LLMs, introducing the CellFill algorithm to update models within quantization cells without altering the original artifact.

arXiv:2608.20873v1 Announce Type: new Abstract: Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint, and with it every evaluation and cache that referred to those exact bits. We instead learn inside the dequantization gap: with the integer codes and scales of a 4-bit release frozen, new knowledge is written only into the per-weight residual that lives strictly inside each quantization decision cell. Re-quantization then returns the released artifact bit-for-bit, a machine-checkable guarantee; updates are exactly revocable by dropping the residual; and drift is bounded. We give six propositions and three training paths, including CellFill, a bounded reparameterization that makes invariance structural rather than enforced. Exact invariance turns out to be nearly free: across three paired seeds the constrained dense path matches an unconstrained reference whose weights provably escape the artifact (58.9 vs 59.3 percent fact recall; paired difference -0.5 points, 95% CI [-5.0,+4.0]), and is better on held-out cross-domain perplexity. Against the natural null hypothesis -- serving the same update as an unmerged adapter -- projecting into the cells reduces cross-domain forgetting in every run that converged, and a diverged control shows the boundary: projection is a trust region, not a repair. What no method escapes is the cost of knowledge itself, and the apparent free lunch of in-domain perplexity improving past the anchor is an artifact of rehearsal sharing a corpus with the metric. Methods differ threefold at matched rehearsal in knowledge bought per point of cross-domain perplexity, a ranking that is not the recall ranking. The method transfers to a 27B hybrid linear-attention model (2.4e10 constrained weights, verified bit-identical), where matched recall costs about half as much cross-domain perplexity as at 1.7B.
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:34 AM

# Nothing Changed but the Model:CellFill — Bounded In-Cell Learning for Bit-Identical,Revocable Updates to Quantized LLMs
Source: [https://arxiv.org/html/2608.20873](https://arxiv.org/html/2608.20873)
Zifeng LiuAffiliation:Institute for Frontier Interdisciplinary Research in Health Sciences and Technology, Sun Yat\-sen UniversityAffiliation:Guangdong Engineering Research Center of Medical Artificial Intelligence Multimodal SystemZhiyong DuAffiliation:Sun Yat\-sen University Institute of Artificial Intelligence School of Business, Sun Yat\-sen UniversityYiming MaoAffiliation:School of Computer Science, China University of Geosciences \(Wuhan\)\[0\.35em\] Zhenhe WangWenqi ShiZhengkun JingAffiliation:School of Public Health, Sun Yat\-sen University Hospital of Stomatology, Sun Yat\-sen University\[0\.3em\] Correspondence:liuzf@mail\.sysu\.edu\.cn\[0\.8em\]Big Data and Artificial Intelligence Center, The Third Affiliated Hospital of Sun Yat\-sen University

Draft — August 21, 2026

###### Abstract

Deployed language models face a structural tension: every update—full fine\-tuning, adapter merging\[[25](https://arxiv.org/html/2608.20873#bib.bib25),[26](https://arxiv.org/html/2608.20873#bib.bib26)\], or model editing\[[16](https://arxiv.org/html/2608.20873#bib.bib16),[28](https://arxiv.org/html/2608.20873#bib.bib28)\]—produces a*new*model whose prior evaluations, certifications, documentation\[[31](https://arxiv.org/html/2608.20873#bib.bib31)\]and caches are invalidated\. We propose*in\-cell learning*: with the integer codes and scales of a 4\-bit quantized release frozen, new knowledge is written only into the per\-weight residual that lives strictly inside each weight’s quantization decision cell—the interval the release discards\. Re\-quantization then reproduces the released artifact bit\-for\-bit, a guarantee that is machine\-checkable and that we check on every merge; updates stay exactly revocable, since dropping the residual restores the released model; and every weight stays inside a box whose radii the quantization grid fixes before training begins\. Drift within that box we measure rather than certify, and report as a calibrated function of the radius\. We give six elementary but load\-bearing propositions—bitwise invariance, projection optimality, a nested refinement code, a capacity bound via the data\-processing inequality, a second\-order forgetting budget, and a geometric law of plasticity decay—and three algorithms that realize the paradigm: post\-hoc clip\-merge of a LoRA update, projected dense fine\-tuning, andCellFill, a bounded reparameterizationW=W^\+M⊙tanh⁡\(s​B​A⊤\)W=\\hat\{W\}\+M\\odot\\tanh\(s\\,BA^\{\\\!\\top\}\)whose learnable object is a position inside the cell, so invariance holds by construction rather than by enforcement\.

On Qwen3\-1\.7B, injecting 1,000 provably\-unseen synthetic facts under exact invariance recovers 24–61% of them depending on the path\. Exact invariance turns out to be nearly free: across three paired seeds the constrained dense path matches an unconstrained reference whose weights provably escape the artifact \(±7\.0%58\.9\\\!\\pm\\\!7\.0\\%vs\.±8\.7%59\.3\\\!\\pm\\\!8\.7\\%; paired difference−0\.5\-0\.5points, 95% CI\[−5\.0,\+4\.0\]\[\-5\.0,\+4\.0\]\), and on held\-out cross\-domain perplexity it is better in both of the two seeds where that metric was recorded\. Against the natural null hypothesis—serving the same update as an unmerged adapter—projecting into the cells reduces cross\-domain forgetting in all nine runs that recorded both stages and converged, by0\.970\.97to271\.3271\.3perplexity points, while a deliberately diverged control shows the effect’s boundary: projection is a trust region, not a repair\. What no method escapes is the cost of knowledge itself: cross\-domain perplexity degrades for all of them, and we show that the apparent “free lunch” of in\-domain perplexity improving past the anchor is an artifact of rehearsal sharing a corpus with the metric\. Methods differ threefold in efficiency at matched rehearsal, and sixfold across the table—knowledge bought per point of cross\-domain perplexity given up—and that ranking is not the recall ranking\. The method transfers to a 27B hybrid linear\-attention model \(2\.4×10102\.4\\times 10^\{10\}constrained weights, verified bit\-identical\), where matched recall costs about half as much cross\-domain perplexity as at 1\.7B\. Efficiency measured per point of perplexity appears to scale asN0\.42N^\{0\.42\}, but we show that71%71\\%of that exponent is the base models’ own perplexity scaling; the scale\-invariant statement is that the*relative*cross\-domain cost is flat at0\.400\.40nats across a12×12\\timesparameter range and two families\. Absorption scales with exponent≈0\.8\\approx\\\!0\.8from10310^\{3\}to10410^\{4\}facts at about 1% cell\-space utilization\. A partition ablation locates where new knowledge is cheapest to write, and it is not where the editing literature locates existing knowledge\. Finally, sequential updates reveal a law we did not expect and can derive: a bounded update consumes a constant*fraction*of the room left in each cell, so plasticity decays geometrically \(β=0\.825±0\.002\\beta=0\.825\\pm 0\.002over four tasks; a direct measurement of the fill refutes our first explanation of why, and we report both\) and the knowledge absorbable without ever changing the artifact is finite even though no constraint is ever violated\. Code, the fifty\-five archived result files behind every table, and the table\- and figure\-generating scripts:[https://github\.com/sumsliu/cellfill](https://github.com/sumsliu/cellfill)\.

## 1Introduction

A 4\-bit checkpoint is released once and depended upon many times\. Between release and retirement it accumulates a benchmark report, a red\-team review, integration tests, downstream caches, and increasingly an external sign\-off, and every one of those artifacts refers to a specific set of integer codes and group scales\. When the operator later needs the model to know something it did not know at release—a corrected specification, a new API, a changed policy—every available mechanism \(full fine\-tuning, adapter merging, closed\-form editing\) returns a*different*checkpoint, and the whole apparatus above must be rebuilt against it\. The binding cost of continual learning in deployment is not GPU time; it is loss of certification\.

We therefore ask what can be learned while changing the released artifact not at all\. Quantization error is normally treated as damage to be minimized; we instead treat the interval between a weight’s dequantized anchor and the boundaries of its rounding cell—the*dequantization gap*—as addressable storage\. With the codes and scales of the release frozen, the update is written only into a full\-precision residual confined to the interior of each cell\. Re\-quantization under those frozen scales then returns the released codes bit for bit at any point in the model’s life \(Prop\.[1](https://arxiv.org/html/2608.20873#Thmproposition1)\); the guarantee is checked by integer comparison rather than argued, and it must be stated against a frozen\(a,s\)\(a,s\)rather than against re\-running a quantizer, since absmax scales are data\-dependent and a single in\-cell move can silently reassign an untouched neighbor \(§[A](https://arxiv.org/html/2608.20873#A1)\)\. Dropping the residual restores the release exactly, and the excursion is bounded per weight by radii the grid supplies for free \(Prop\.[5](https://arxiv.org/html/2608.20873#Thmproposition5)\)\.

What this buys is narrower than it may first appear, and we state it precisely\. The served model isW^\+δ\\hat\{W\}\+\\delta, and it is a different function fromW^\\hat\{W\}: its cross\-domain perplexity moves with every unit of knowledge absorbed, so we do*not*claim that the release’s validation transfers to the updated model\. What is inherited is the*reference*\. The certified configuration remains recoverable bit\-for\-bit by truncation at any time, the update is exactly revocable, and its drift is bounded by a budget whose radii are known before training begins\. An update becomes an auditable increment against a fixed baseline rather than an opaque replacement of it\.

Empirically, on Qwen3\-1\.7B with 1,000 synthetic facts whose novelty is certain by construction, three constraint\-realizing paths recover±1\.7%24\.0\\\!\\pm\\\!1\.7\\%\(clip\-merge,n=3n\{=\}3\) to±10\.6%61\.4\\\!\\pm\\\!10\.6\\%\(projected dense,n=2n\{=\}2\) of the facts under exact invariance \(196 matrices,1\.409×1091\.409\\times 10^\{9\}constrained weights, zero violations\)\. At the one rehearsal fraction where a matched unconstrained reference exists, the constrained dense path recovers±7\.0%58\.9\\\!\\pm\\\!7\.0\\%against the reference’s±8\.7%59\.3\\\!\\pm\\\!8\.7\\%, which buys that gap by moving7\.67\.6M weights out of their cells—and the gap does not survive seed noise, so we report the reference as a reference and not as an upper bound\. Cross\-domain perplexity \(LAMBADA, untouched by rehearsal\) rises monotonically with absorption across the five paths run at matched rehearsal,28\.6728\.67at the anchor to31\.731\.7,33\.133\.1,37\.437\.4,40\.740\.7and52\.252\.2\(Table[1](https://arxiv.org/html/2608.20873#S5.T1)\): the cost of knowledge is real and measurable, and the in\-domain “free lunch” reported by an earlier version of this work is an artifact of rehearsal sharing a corpus with the metric \(§[5\.5](https://arxiv.org/html/2608.20873#S5.SS5)\)\. Three findings then cut against expectation\. First, the constraint is not only a tax: against the natural null hypothesis of serving the same adapter unmerged—which preserves the artifact trivially, by never touching it—projecting the update into the cells*reduces*cross\-domain forgetting in all nine runs that recorded both stages and converged, by0\.970\.97to271\.3271\.3perplexity points, with the effect largest exactly where drift is worst—and we report the boundary of that effect too: on an update whose own optimization has already blown up, projection amplifies the damage instead of repairing it \(§[5\.6](https://arxiv.org/html/2608.20873#S5.SS6)\)\. Second, the mechanism survives scale and architecture: on a 27B hybrid model whose blocks are gated linear attention rather than softmax attention, all2\.435×10102\.435\\times 10^\{10\}constrained weights across 496 matrices verify bit\-identical after merging, and matched recall \(25\.4%25\.4\\%at 27B, single seed, against±1\.7%24\.0\\\!\\pm\\\!1\.7\\%for the 1\.7B clip\-merge arm\) costs\+1\.29\+1\.29LAMBADA points over its own anchor where the 1\.7B arm pays\+3\.07\+3\.07over its own, that figure being a two\-seed mean as in Table[1](https://arxiv.org/html/2608.20873#S5.T1)\(§[5\.11](https://arxiv.org/html/2608.20873#S5.SS11)\)\. Third, knowledge is most cheaply written into the gate and up projections at*all*depths \(130130absorbed bits per million trainable parameters, and911911bits per point of cross\-domain perplexity given up\), rather than intodown\_proj\(7272\); the early and middle MLP partitions absorbed nothing distinguishable from the guessing floor at all, so we report them as non\-absorbing rather than as a low yield—contra the localization prescriptions of ROME and MEMIT\[[16](https://arxiv.org/html/2608.20873#bib.bib16),[17](https://arxiv.org/html/2608.20873#bib.bib17)\], a discrepancy we read as a difference between*editing*an existing association and*writing*a new one under a norm constraint \(§[5\.10](https://arxiv.org/html/2608.20873#S5.SS10)\)\. Absorption grows with exponent≈0\.8\\approx\\\!0\.8from10310^\{3\}to10410^\{4\}presented facts at roughly1%1\\%cell\-space utilization: in this range the regime is optimization\-limited, not capacity\-limited \(§[5](https://arxiv.org/html/2608.20873#S5)\)\.

Contributions\.\(1\) A reframing of the quantized release from compression artifact to*fixed reference frame*for continued learning, with a formal update contract—bitwise artifact invariance, exact revocability, and a computable drift budget—and the elementary propositions that establish it \(§[2](https://arxiv.org/html/2608.20873#S2)\)\. \(2\) A law of plasticity in bounded learning: a bounded update spends a constant fraction of the remaining room, so capacity decays geometrically and lifetime capacity without re\-quantization is finite \(Proposition[6](https://arxiv.org/html/2608.20873#Thmproposition6)\)—a derivation that predicts a constant we then measure to 1\.8% \(§[5\.12](https://arxiv.org/html/2608.20873#S5.SS12)\)\. \(3\) Three training paths that realize the constraint, including CellFill, a bounded reparameterization whose invariance holds by construction at every step so that the trained and shipped models coincide \(§[3](https://arxiv.org/html/2608.20873#S3)\)\. \(4\) A measurement methodology for knowledge capacity in bits—synthetic corpora with certain novelty and exactly countable information content, a declared guessing floor, and cross\-domain forgetting metrics disjoint from the rehearsal corpus \(§[4](https://arxiv.org/html/2608.20873#S4)\)—and the map of*where*in a network new knowledge is cheapest to write that it makes possible \(§[5\.10](https://arxiv.org/html/2608.20873#S5.SS10)\)\. \(5\) An empirical study across methods, replay recipes, radii, corpus sizes, scales, and architectures that locates the learning–forgetting frontier of invariant updates, and reports two negative results we consider load\-bearing: a fixed rehearsal buffer is memorized and is worse than no rehearsal at all, and the diagonal\-Fisher budget of Prop\.[5](https://arxiv.org/html/2608.20873#Thmproposition5)is a scaling law, not a numerical certificate \(§[5](https://arxiv.org/html/2608.20873#S5), §[5\.9](https://arxiv.org/html/2608.20873#S5.SS9)\)\. \(6\) A nested refinement code turning updates into prefix\-compatible\(4\+k\)\(4\{\+\}k\)\-bit checkpoints whose 4\-bit truncation is always the original release \(§[6](https://arxiv.org/html/2608.20873#S6)\)\.

## 2Setting and Guarantees

###### Definition 1\(Frozen artifact\)\.

A blockwise quantizer with frozen group scaless∈ℝ\>0N/gs\\in\\mathbb\{R\}^\{N/g\}\_\{\>0\}and level tableLLmaps codesa∈\{0,…,15\}Na\\in\\\{0,\\dots,15\\\}^\{N\}to anchorsw^i=sg⁡\(i\)​L​\[ai\]\\hat\{w\}\_\{i\}=s\_\{g\(i\)\}\\,L\[a\_\{i\}\]\. The artifact is𝒜=\(a,s\)\\mathcal\{A\}=\(a,s\)\. The*cell*CiC\_\{i\}of weightiiis the round\-to\-nearest decision region ofaia\_\{i\}under frozenss; the invariant set is𝒞=∏iCi\\mathcal\{C\}=\\prod\_\{i\}C\_\{i\}\.

###### Proposition 1\(Bitwise invariance\)\.

Ifw′∈int⁡𝒞w^\{\\prime\}\\in\\operatorname\{int\}\\mathcal\{C\}, re\-quantization under frozen scales returns exactlyaa\. Consequently any composition of updates whose images lie in𝒞\\mathcal\{C\}preserves the artifact bit\-for\-bit\.

###### Proof\.

Immediate from the definition of the decision regions\. The content of the proposition is the precise statement of what must be frozen: recomputing absmax scales after training silently reassigns entire blocks \(we exhibit a two\-weight counterexample in the appendix\), so invariance must be defined against frozen\(a,s\)\(a,s\), never against re\-running a quantizer\. ∎

###### Proposition 2\(Projection optimality\)\.

Coordinatewise clipping to𝒞\\mathcal\{C\}is the Euclidean projectionP𝒞P\_\{\\mathcal\{C\}\}; clip\-merge returns the invariant model nearest to any proposed update, and projected SGD on𝒞\\mathcal\{C\}is proximal descent on the box indicator\.

###### Proposition 3\(Nested refinement code\)\.

For bin positionxi∈\[0,1\)x\_\{i\}\\in\[0,1\)letri\(k\)=⌊2k​xi⌋r^\{\(k\)\}\_\{i\}=\\lfloor 2^\{k\}x\_\{i\}\\rfloorwith sub\-cell\-center reconstruction\. Then \(i\)r\(k\)=r\(k\+j\)≫jr^\{\(k\)\}=r^\{\(k\+j\)\}\\\!\\gg\\\!j; \(ii\) the reconstruction error is at mostwidthi/2k\+1\\mathrm\{width\}\_\{i\}/2^\{k\+1\}; \(iii\) reconstructions are strictly interior, so every truncation depth satisfies Proposition[1](https://arxiv.org/html/2608.20873#Thmproposition1)\.

###### Proposition 4\(Capacity identity\)\.

An update messagem=\(mask,r\(k\)\)m=\(\\text\{mask\},r^\{\(k\)\}\)has length\|m\|≤k​Nt\+log2⁡\(NNt\)\|m\|\\leq kN\_\{t\}\+\\log\_\{2\}\\binom\{N\}\{N\_\{t\}\}bits, and for any fact setFFand updated modelW′=f⁡\(W^,m\)W^\{\\prime\}=f\(\\hat\{W\},m\), the data\-processing inequality\[[30](https://arxiv.org/html/2608.20873#bib.bib30)\]givesI⁡\(F,W′\)≤H⁡\(m\)≤\|m\|I\(F;W^\{\\prime\}\)\\leq H\(m\)\\leq\|m\|\. Absorbed knowledge is bounded by shipped bits as a theorem, not an estimate\.

###### Proposition 5\(Forgetting budget\)\.

For\|δi\|≤ρ​Δi/2\|\\delta\_\{i\}\|\\leq\\rho\\,\\Delta\_\{i\}/2,𝔼xKL\(pW^∥pW^\+δ\)=12δ⊤Fδ\+O\(∥δ∥3\),\\mathbb\{E\}\_\{x\}\\,\\mathrm\{KL\}\\\!\\left\(p\_\{\\hat\{W\}\}\\,\\\|\\,p\_\{\\hat\{W\}\+\\delta\}\\right\)=\\tfrac\{1\}\{2\}\\,\\delta^\{\\top\}F\\,\\delta\+O\(\\\|\\delta\\\|^\{3\}\),so to second order the drift of any update confined to that box is controlled by the Fisher form on radii the quantization grid supplies for free, and scales asρ2\\rho^\{2\}\. Which updates the hypothesis covers is worth stating exactly, because the three paths differ\. CellFill moves inside the symmetric box\|δi\|≤roomi=min⁡\(w^i−ℓi,ui−w^i\)\|\\delta\_\{i\}\|\\leq\\mathrm\{room\}\_\{i\}=\\min\(\\hat\{w\}\_\{i\}\-\\ell\_\{i\},\\,u\_\{i\}\-\\hat\{w\}\_\{i\}\), and sinceroomi≤Δi/2\\mathrm\{room\}\_\{i\}\\leq\\Delta\_\{i\}/2\(Appendix[B](https://arxiv.org/html/2608.20873#A2)\) the proposition applies to it directly\. The projected paths do not: clip\-merge and projected dense clamp to the full cell\[ℓi,ui\]\[\\ell\_\{i\},u\_\{i\}\], so a weight may travel up tomax⁡\(w^i−ℓi,ui−w^i\)\\max\(\\hat\{w\}\_\{i\}\-\\ell\_\{i\},\\,u\_\{i\}\-\\hat\{w\}\_\{i\}\), which exceedsΔi/2\\Delta\_\{i\}/2wherever the cell is asymmetric about its anchor—which, for the non\-uniform NF4 level table, is fourteen of the sixteen codes \(the two extremes are symmetric only because capping mirrors their inner half\-gap\)\. For those paths the proposition bounds the symmetric part of the excursion and not the whole of it\.

We stress what this does and does not give\. The expression is exact to second order, but evaluating it requires the fullFF, and replacingFFby its diagonal gives two different surrogates that must not be confused\. The supremum over the loose box isBF​\(ρ\)=ρ28​∑iFi​i​Δi2B\_\{F\}\(\\rho\)=\\tfrac\{\\rho^\{2\}\}\{8\}\\sum\_\{i\}F\_\{ii\}\\Delta\_\{i\}^\{2\}; the*mean*over a uniform fill of the true cells isB¯F​\(ρ\)=ρ26​∑iFi​i​roomi2\\bar\{B\}\_\{F\}\(\\rho\)=\\tfrac\{\\rho^\{2\}\}\{6\}\\sum\_\{i\}F\_\{ii\}\\,\\mathrm\{room\}\_\{i\}^\{2\}\(Appendix[B](https://arxiv.org/html/2608.20873#A2)\), and it isB¯F\\bar\{B\}\_\{F\}that our geometry sweep evaluates\. Neither is a usable numerical bound\. Measured against random in\-cell perturbations of a 1\.7B model,B¯F\\bar\{B\}\_\{F\}underestimates the true drift by15\.9×15\.9\\timesatρ=0\.125\\rho\{=\}0\.125rising to383×383\\timesatρ=1\\rho\{=\}1\(§[5\.9](https://arxiv.org/html/2608.20873#S5.SS9)\); sinceroomi≤Δi/2\\mathrm\{room\}\_\{i\}\\leq\\Delta\_\{i\}/2we haveBF≥3​B¯FB\_\{F\}\\geq 3\\bar\{B\}\_\{F\}, so the supremum form underestimates by at most128×128\\timesatρ=1\\rho\{=\}1—two orders of magnitude either way\. The reason is*not*that off\-diagonal Fisher mass dominates: for the zero\-mean coordinatewise perturbations we apply, the off\-diagonal entries cancel*exactly*in the mean ofδ⊤​F​δ\\delta^\{\\\!\\top\}F\\deltaand enter only its variance \(Appendix[B](https://arxiv.org/html/2608.20873#A2)\)—an identity, not a measurement\. Whether one realization sits near that mean is the empirical question, and two perturbation seeds bound it only crudely: their measuredΔ\\DeltaNLL spreads1414–39%39\\%about the mean forρ≥0\.25\\rho\\geq 0\.25, two orders of magnitude below the196196–383×383\\timesgap, so fluctuation cannot account for it\. Atρ=0\.125\\rho\{=\}0\.125the two seeds do not agree even in sign and we draw nothing from the16×16\\timesfigure there\. Two explanations remain, and our data do not separate them: the empirical Fisher we estimate from squared gradients is a poor proxy for curvature at a converged point, and the second\-order truncation itself may fail at this magnitude—‖δ‖\\\|\\delta\\\|is macroscopic even though every coordinate moves less than a cell half\-width\. The measured exponent is close to but not equal to the predicted22, and it does not even move consistently in one direction:2\.432\.43at 1\.7B and1\.71\.7–1\.91\.9at 4B \(§[5\.9](https://arxiv.org/html/2608.20873#S5.SS9)\), which is itself an argument against reading much into either\. We therefore treatρ\\rhoas a calibrated dial rather than an a\-priori certificate, and report the measured curve\.

###### Proposition 6\(Geometric plasticity decay\)\.

Writea=w^−ℓa=\\hat\{w\}\-\\ellandb=u−w^b=u\-\\hat\{w\}for a weight’s distances to the two walls of its cell andr=min⁡\(a,b\)r=\\min\(a,b\)for the inradius about its current anchor, and let a bounded update move it tow′=w^\+r​tw^\{\\prime\}=\\hat\{w\}\+r\\,twith\|t\|<1\|t\|<1, as CellFill does witht=tanh⁡\(s​B​A⊤\)t=\\tanh\(s\\,BA^\{\\top\}\)\. Foldingw′w^\{\\prime\}into the anchor for the next task leaves inradius

r′=min⁡\(a\+r​t,b−r​t\)≥r⁡\(1−\|t\|\),r^\{\\prime\}\\;=\\;\\min\\\!\\big\(a\+rt,\\;b\-rt\\big\)\\;\\geq\\;r\\,\(1\-\|t\|\),with equality whenever the update moves toward the nearer wall, and always when the cell is symmetric about its anchor \(a=ba=b\)\. Each update therefore consumes at most a constant*fraction*of the room rather than a constant amount, and the geometric law below is the conservative edge of that inequality: an update that moves toward the*far*wall recentres the anchor and can leave more room than it started with\. Which case dominates in practice is an empirical question, and §[5\.12](https://arxiv.org/html/2608.20873#S5.SS12)answers it—the measured decay tracks the geometric law closely\. Writingr¯k\\bar\{r\}\_\{k\}for the mean inradius afterkktasks and assuming\|t\|\|t\|is uncorrelated withrracross weights,

r¯k=r¯0​βk,β=1−𝔼​\|t\|\.\\bar\{r\}\_\{k\}=\\bar\{r\}\_\{0\}\\,\\beta^\{k\},\\qquad\\beta=1\-\\mathbb\{E\}\|t\|\.If moreover a task’s absorption is proportional to the room available to it,ak=c​r¯k−1a\_\{k\}=c\\,\\bar\{r\}\_\{k\-1\}, thenak=a1​βk−1a\_\{k\}=a\_\{1\}\\beta^\{\\,k\-1\}and the knowledge absorbable without ever changing the artifact is finite:

∑k≥1ak=a11−β<∞\.\\sum\_\{k\\geq 1\}a\_\{k\}=\\frac\{a\_\{1\}\}\{1\-\\beta\}<\\infty\.

Two consequences are worth separating\. First, the bound is never violated and no single update fails—the sequence simply becomes unable to learn, which is a failure mode that a constraint\-violation check cannot detect\. Second, per\-task throughputak→0a\_\{k\}\\to 0, so a policy that never re\-quantizes has asymptotically zero learning rate; consolidating everymmtasks restoresr¯\\bar\{r\}and yields a positive steady statea1​\(1−βm\)/\(m⁡\(1−β\)\)a\_\{1\}\(1\-\\beta^\{m\}\)/\\big\(m\(1\-\\beta\)\\big\), decreasing inmm\. Consolidation is therefore not a fallback for when something goes wrong but a structural requirement of learning in a bounded space, andmmis the dial trading learning throughput against how often the released artifact changes\.

## 3Methods

A \(clip\-merge\)\.Train LoRA on the quantized base; materializeΔ\\Deltalayerwise and setW′=P𝒞​\(base\+Δ\)W^\{\\prime\}=P\_\{\\mathcal\{C\}\}\(\\text\{base\}\+\\Delta\)\. Cheapest; optimal in weights, not in loss \(Prop\.[2](https://arxiv.org/html/2608.20873#Thmproposition2)\)\.A\+ \(healing\)\.Continue fromW′W^\{\\prime\}with projected dense fine\-tuning: the optimizer sees the walls and re\-routes clipped\-away knowledge into remaining room\.B \(projected dense\)\.Constraint\-aware from step zero; the reference solution of the constrained problem\.CellFill\.Two names at two levels\. A weight’s constraint set is its*cell*; any procedure that confines learning to the cells is*in\-cell learning*, and all three paths in this section are instances of it\.*CellFill*is the algorithm that realizes the paradigm by construction rather than by enforcement, and it is the one we recommend\. It is not a LoRA variant, and the distinction is not cosmetic:BBandAAparameterize a*position*in\(−1,1\)\(\-1,1\)per weight rather than a weight delta, and both the elementwisetanh\\tanhand the elementwiseMMdestroy low\-rank structure, so a rank\-rrparameterization produces a*full\-rank*update\. In a512×512512\\times 512probe atr=16r\{=\}16the LoRA productB​A⊤BA^\{\\top\}has rank1616whileM⊙tanh⁡\(s​B​A⊤\)M\\odot\\tanh\(sBA^\{\\top\}\)has rank512512and stable rank110110\. This is the most likely reason CellFill at rank 64 matches projected dense recall with20×20\\timesfewer trainable parameters \(§[5\.2](https://arxiv.org/html/2608.20873#S5.SS2)\)\.W=W^\+M⊙tanh⁡\(s​B​A⊤\)W=\\hat\{W\}\+M\\odot\\tanh\(s\\,BA^\{\\top\}\)withMMthe in\-bin half\-width: the learnable object*is*the bin position; invariance is structural, and training\-time and shipped models coincide\. The clip rate of path A is a computable diagnostic of the divergence betweenP𝒞​\(arg⁡min⁡f\)P\_\{\\mathcal\{C\}\}\(\\arg\\min f\)andarg⁡min𝒞⁡f\\arg\\min\_\{\\mathcal\{C\}\}f: below∼\\sim1% the paths agree; above∼\\sim5%, escalate to A\+/B\.

## 4Experimental setup

Synthetic biography facts \(novelty certain by construction;21\.721\.7bits per fact of countable attribute entropy; three cloze probes per fact\), fresh per\-epoch rehearsal from wikitext\-train \(a fixed replay buffer repeated across epochs is itself memorized and*raises*test perplexity from 24\.6 to 172—a negative result we document\), forgetting measured by wikitext\-test and cross\-domain LAMBADA perplexity\. Unless stated otherwise the base model is Qwen3\-1\.7B\-Base; the scale ladder \(0\.6B–27B, including a hybrid linear\-attention model\) and the second model family are reported in §[5\.11](https://arxiv.org/html/2608.20873#S5.SS11)\. Quantization: NF4, blocksize 64, double quantization, frozen scales; invariance is asserted per layer, in the integer domain, on every merge\.

Each fact carries21\.721\.7bits of attribute entropy, but the three cloze probes cover only city, occupation and employer:11\.911\.9bits\. All capacity numbers use the probed11\.911\.9bits\. Two probes continue the training phrasing and one paraphrases it, so we report recall per probe kind as well as pooled\. A model that has learned the answer vocabulary but no name–attribute bindings scores6\.5%6\.5\\%under greedy decoding; that is the floor against which small recall numbers should be read\.

## 5Results

Figure 1:The invariant learning–forgetting frontier \(Qwen3\-1\.7B, 1,000 synthetic facts, 24 epochs; log perplexity axes, error bars±\\pm1 std over three seeds\)\. Every diamond, circle and square re\-quantizes bit\-identically to the released 4\-bit artifact; the cross does not\.The two panels disagree, and that disagreement is the point\.Rehearsal is drawn from WikiText\-train, so in \(a\) every method sits at or below the anchor and in\-cell learning looks free; LAMBADA in \(b\) never saw the rehearsal corpus, and there every method is far above it\. A reader who measures only \(a\) concludes that nothing was given up\. The two gray squares in \(a\) ablate rehearsal: omitting it doubles perplexity, and re\-using a fixed buffer across epochs is far worse still \(63\.863\.8\), because the buffer itself gets memorized\. Note also that ordering by recall is not ordering by cost: CellFill at rank 64 absorbs nearly what projected dense does while adding12\.012\.0points of cross\-domain perplexity over the anchor against dense’s42\.342\.3\.### 5\.1The invariant frontier \(Qwen3\-1\.7B, 1k facts, 24 epochs\)

methodinvariancereh\.recallWikiText\-2LAMBADAbits/ptseeds4\-bit anchor\(reference\)——11\.7128\.67——original fp32\(reference\)——10\.6626\.13——A: clip\-merge✓0\.1±1\.724\.0\\\!\\pm\\\!1\.7%±0\.0210\.29\\\!\\pm\\\!0\.02±0\.3731\.74\\\!\\pm\\\!0\.379293/2CellFillr=16r\{=\}16✓ \(structural\)0\.136\.9%10\.5433\.0610001A\+: clip\-merge→\\toheal✓0\.1±2\.140\.2\\\!\\pm\\\!2\.1%±0\.2711\.69\\\!\\pm\\\!0\.27±1\.0937\.40\\\!\\pm\\\!1\.095473/2CellFillr=64r\{=\}64✓ \(structural\)0\.156\.7%11\.6440\.665631B: projected dense✓0\.1±10\.661\.4\\\!\\pm\\\!10\.6%±0\.0513\.53\\\!\\pm\\\!0\.05±4\.8252\.15\\\!\\pm\\\!4\.823112B: projected dense✓0\.3±7\.058\.9\\\!\\pm\\\!7\.0%±1\.9413\.39\\\!\\pm\\\!1\.94±5\.4370\.98\\\!\\pm\\\!5\.431663/2unconstrained dense×\\times0\.3±8\.759\.3\\\!\\pm\\\!8\.7%±0\.1312\.27\\\!\\pm\\\!0\.13±5\.5774\.83\\\!\\pm\\\!5\.571533/2Table 1:The invariant learning–forgetting frontier \(Qwen3\-1\.7B, 1,000 facts, 24 epochs\)\. Every ✓ row re\-quantizes bit\-identically to the released artifact—asserted per layer, in the integer domain, over all1\.409×1091\.409\\times 10^\{9\}constrained weights, on every merge\. Cells are mean±\\pmstd; the last column gives the number of seeds \(recall/LAMBADA, where they differ because the cross\-domain harness was added after the first seed\)\.The two perplexity columns disagree by construction: rehearsal is drawn from WikiText\-train, so WikiText\-test flatters every method, while LAMBADA is untouched by rehearsal and prices the true cost\. Rows A and A\+ are the same runs before and after healing\. The*reh\.*column gives the fresh\-rehearsal fraction: the bottom two rows were run at0\.30\.3because the paired test of §[5\.3](https://arxiv.org/html/2608.20873#S5.SS3)needs the constrained and unconstrained arms at the same fraction as each other; the matched0\.10\.1dense row directly above them is the one to compare against A, A\+ and CellFill\.*bits/pt*is the efficiency measure of §[5\.2](https://arxiv.org/html/2608.20873#S5.SS2): probed bits of knowledge absorbed per point of LAMBADA perplexity given up\. This table is generated from the archived result files byscripts/make\_tables\.py\.
### 5\.2Cost is not proportional to absorption: methods differ in efficiency

It would be convenient if cross\-domain damage simply tracked how much a method absorbs\. It does not, and the gap between methods is large\. Ranking the invariant paths by*bits absorbed per point of cross\-domain perplexity*\(Table[1](https://arxiv.org/html/2608.20873#S5.T1),*bits/pt*\) gives three tiers, not a five\-way ordering\. Computed per seed rather than as a ratio of means, and at a common rehearsal fraction of0\.10\.1: clip\-merge±183918\\\!\\pm\\\!183and CellFillr=16r\{=\}1610001000\(n=1n\{=\}1\); A\+±88569\\\!\\pm\\\!88and CellFillr=64r\{=\}64563563\(n=1n\{=\}1\); projected dense±10312\\\!\\pm\\\!10\. The tiers separate by3\.2×3\.2\\timesand that separation is far outside the spreads\. The ordering*within*a tier is not resolved and we do not claim it—indeed A\+ and CellFillr=64r\{=\}64change places depending on whether the ratio is taken per seed or between means, which is precisely why we now take it per seed\. What the tiers show is that the exchange rate is*not*the recall ordering: projected dense has the highest recall and the worst exchange rate, CellFillr=16r\{=\}16nearly the reverse\.

This is the quantity a deployment actually optimizes, and it separates two regimes that raw recall conflates\. Bounded low\-rank paths stay near the anchor and buy knowledge cheaply but cannot buy much of it; full\-rank paths buy a lot at a steeply rising price\. The efficient operating points are in the middle tier: CellFill at rank 64 absorbs1\.5×1\.5\\timeswhat CellFillr=16r\{=\}16does at roughly56%56\\%of its efficiency, while projected dense at the same rehearsal fraction absorbs8%8\\%more than CellFillr=64r\{=\}64for2\.0×2\.0\\timesthe cross\-domain damage \(23\.523\.5points against12\.012\.0\) and without the structural guarantee\. That last comparison is a tier crossing and survives the spreads; the choice between A\+ and CellFillr=64r\{=\}64inside the middle tier does not, and should be made on the guarantee rather than on these numbers\. Unless recall is the only thing that matters, the bounded reparameterization dominates the projected paths\.

### 5\.3What does exact invariance cost?

The obvious worry is that forbidding the0\.54%0\.54\\%of weights that an unconstrained run moves out of its cell from going where the optimizer wants must cost accuracy\. We measure it directly: identical dense training, identical seeds, the only difference being whether the projection is applied after each step\. Across three paired seeds the constrained arm recovers±7\.0%58\.9\\\!\\pm\\\!7\.0\\%of facts and the unconstrained arm±8\.7%59\.3\\\!\\pm\\\!8\.7\\%; the paired difference is−0\.47\-0\.47recall points with a 95% interval of\[−5\.0,\+4\.0\]\[\-5\.0,\+4\.0\]\(t=−0\.45t=\-0\.45,df=2\\mathrm\{df\}=2\)\. We therefore cannot exclude a cost of five recall points, nor a gain of four; the point estimate is indistinguishable from zero, while the unconstrained arm moves7\.57\.5–7\.87\.8million weights \(0\.54%0\.54\\%\) out of their cells per run and destroys the artifact\.

On the strict axis the constrained arm is ahead: LAMBADA70\.9870\.98versus74\.8374\.83\(seeds 1–2 only; seed 0 predates the cross\-domain harness, so this pair isn=2n\{=\}2where the recall test above isn=3n\{=\}3\)\. Constraining the update is not merely affordable here; it is weakly better, for the reason developed in §[5\.6](https://arxiv.org/html/2608.20873#S5.SS6)\. We flag that an earlier draft of this work reported the cost as “1\.8 recall points” from single\-seed runs\. That number was an artifact ofn=1n=1against a between\-seed standard deviation of77–99points, and we report the paired analysis instead\.

#### What it costs to run, against the thing it replaces\.

The unconstrained arm above is the closest control we have to full\-parameter fine\-tuning of the released model: every one of the1\.409×1091\.409\\times 10^\{9\}constrained weights is trainable and free to move\. Priced against it, CellFill at rank 64 trains69,730,30469\{,\}730\{,\}304parameters—20×20\\timesfewer—in12\.012\.0minutes against20\.120\.1, with the Adam state falling from10\.510\.5GB to0\.520\.52GB, and it recovers56\.7%56\.7\\%of the facts against±8\.7%59\.3\\\!\\pm\\\!8\.7\\%, a gap inside the control’s own between\-seed spread\. It gives up less than half as much cross\-domain ability in the process \(LAMBADA40\.6640\.66against74\.8374\.83\)\. And the two arms differ in kind at the end: the control has moved7\.57\.5–7\.87\.8million weights out of their cells and produced a new checkpoint that replaces the release, while the constrained run ships an increment that re\-quantizes to the release bit\-for\-bit and can be dropped to recover it exactly\. The comparison that matters for deployment is therefore not only cheaper—it returns a different kind of object\.

### 5\.4Attribution: forgetting is training drift, not clipping

With facts\-only training \(no rehearsal\), the adapter model degrades wikitext PPL from 11\.71 to 24\.60 while the clipped merge scores 22\.59: clipping*reduces*drift \(it pulls weights toward anchors\) and the entire degradation is attributable to the training distribution, not the constraint\. Kernel/dtype numerics are excluded by an anchor\-rebuild control \(bnb\-4bit 11\.727 vs\. fp32 rebuild 11\.710\)\.

### 5\.5Rehearsal: fresh or fatal, and what it does not buy

A fixed replay buffer repeated across 24 epochs is itself memorized and sharpens the model onto its own rehearsal set: test perplexity explodes from24\.624\.6to172172for the unmerged adapter, and from22\.622\.6to63\.863\.8after clip\-merge\. Redrawing rehearsal snippets every epoch removes in\-domain forgetting entirely: the merged model reaches WikiText PPL 10\.78, below its own 4\-bit anchor \(11\.71\), while absorbing 17\.4% of the facts; a 10% replay fraction dominates 30% on both axes \(recall→25\.1%17\.4\\\!\\to\\\!25\.1\\%, PPL→10\.3110\.78\\\!\\to\\\!10\.31\)\.

This in\-domain gain is a rehearsal effect, not a property of the dequantization gap, and it does not survive a domain shift\.Two controls establish this and we state them prominently because an earlier draft of this work drew the opposite conclusion\. First, the*unmerged*adapter —which is not constrained to the cells at all—scores essentially the same in\-domain perplexity as the box\-constrained merge \(10\.299 vs\. 10\.271 on seed 1; the gap stays under0\.030\.03across the three 10%\-rehearsal seeds and never exceeds0\.290\.29in any fresh\-rehearsal run\), so the improvement cannot be attributed to in\-cell storage\. Second, rehearsal is drawn from WikiText\-train while the perplexity metric is WikiText\-test: on held\-out LAMBADA every method*degrades*relative to the anchor \(28\.67\), monotonically in how much it absorbs among the arms run at matched rehearsal \(Table[1](https://arxiv.org/html/2608.20873#S5.T1)\)\. The learning–forgetting trade\-off is real; the free lunch was an artifact of measuring rehearsal on its own domain\.

### 5\.6The constraint is not only a tax: it reduces cross\-domain forgetting

The natural null hypothesis for this entire paper is “why not simply keep the adapter unmerged?” — servingW^\\hat\{W\}plus an unconstrained LoRA costs nothing and preserves the artifact trivially, by never touching it\. Comparing the two is a single\-intervention experiment: identical training, identical adapter, the only difference being whether the update is projected into the cells\. On held\-out LAMBADA the projected model is better in*every*run whose optimization converged—all nine, spanning four model sizes, two families and both attention mechanisms:

runmodelanchorunmergedprojectedΔ\\Deltaexp26\_lora\_lr1e\-3Qwen3\-1\.7B28\.67311\.0839\.81−271\.27\-271\.27exp17\_aplus10k\_xdomQwen3\-1\.7B28\.67108\.6336\.84−71\.79\-71\.79exp8\_e96\_qwen3\-1\.7bQwen3\-1\.7B28\.6746\.3333\.71−12\.62\-12\.62exp12b\_27b\_e24Qwen3\.8\-27B20\.0330\.9221\.25−9\.66\-9\.66exp20\_aplus\_0p6Qwen3\-0\.6B38\.9947\.3842\.74−4\.64\-4\.64exp25\_mistral\_a\_lr5e\-5Mistral\-7B\-v0\.314\.0017\.6615\.58−2\.09\-2\.09exp1c\_s1Qwen3\-1\.7B28\.6733\.9032\.00−1\.90\-1\.90exp1c\_s2Qwen3\-1\.7B28\.6732\.7531\.47−1\.28\-1\.28exp12\_27b\_e8Qwen3\.8\-27B20\.0322\.2921\.32−0\.97\-0\.97*unmerged optimization already diverged \(\>100×\>100\\timesanchor\):*exp23\_mistral\_aMistral\-7B\-v0\.314\.007512\.283165\.85−4346\.43\-4346\.43exp26\_lora\_lr3e\-3Qwen3\-1\.7B28\.676420\.07115458\.72\+109038\.65\+109038\.65
The box acts as a per\-weight trust region whose radii the quantization grid supplies for free, and clipping pulls the update back toward the anchor along exactly the coordinates that moved farthest\. The effect is largest where drift is largest:−271\.3\-271\.3perplexity points on a rank\-16 LoRA at lr10−310^\{\-3\}, whose unmerged adapter reaches LAMBADA311\.1311\.1against39\.839\.8after projection, and−71\.8\-71\.8on the10410^\{4\}\-fact A\+ run—so it is self\-strengthening exactly when it is most needed\. The last two rows bound the claim\. Both are runs whose*unmerged*optimization had already exceeded100×100\\timesits anchor, so neither tests whether projection regularizes training; they test what projection does to a diverged update, and the answer is not stable\. On Mistral it removes more than half the damage \(7512→31667512\\to 3166\); on plain LoRA at lr3×10−33\\times 10^\{\-3\}it makes matters far worse \(6420→115,4596420\\to 115\{,\}459\), because clipping70\.2%70\.2\\%of the weights of an incoherent update leaves an update that is incoherent*and*truncated, retaining18\.2%18\.2\\%of its norm\. Projection is a trust region, not a repair: it constrains a converging optimization and cannot rescue one that has already failed\. This is the robustness argument for the constraint that survives cross\-domain scrutiny—and unlike the in\-domain perplexity gains of §[5\.5](https://arxiv.org/html/2608.20873#S5.SS5), it is attributable to the constraint itself, since projection is the only variable that differs\.

### 5\.7Post\-hoc projection vs\. constraint\-aware training

The clip rate diagnoses the gap betweenP𝒞​\(arg⁡min⁡f\)P\_\{\\mathcal\{C\}\}\(\\arg\\min f\)andarg⁡min𝒞⁡f\\arg\\min\_\{\\mathcal\{C\}\}f\(Prop\.[2](https://arxiv.org/html/2608.20873#Thmproposition2)\)\. At 8 epochs the LoRA delta clips at0\.14%0\.14\\%and projection is free; at 24 epochs it clips at2\.42\.4–4\.6%4\.6\\%, costing 3\.7 recall points without rehearsal and1\.91\.9–2\.42\.4with it; at10410^\{4\}facts it clips at30\.4%30\.4\\%and path A collapses to 4\.4% recall—while four epochs of projected healing recover 19\.1%,*above*the unconstrained rank\-16 adapter itself \(14\.7%\): dense in\-bin degrees of freedom compensate the low\-rank bottleneck\. Halvingρ\\rhocosts only13%13\\%of healed recall \(31\.5→27\.3%31\.5\\to 27\.3\\%,n=1n\{=\}1\), which is what makesρ\\rhoa usable dial; how the damage itself scales is measured directly in §[5\.9](https://arxiv.org/html/2608.20873#S5.SS9), where halvingρ\\rhoshrinks measuredΔ\\DeltaNLL by a factor0\.190\.19—an exponent of2\.42\.4, not22\. Two further ablations locate the binding constraint\. Quadrupling exposures \(96 epochs, rank 16\) moves healed recall only→40\.5%37\.9\\\!\\to\\\!40\.5\\%while the clip rate quadruples: more optimization time is not what is missing\. Rank is\. At matched rehearsal, raising CellFill’s rank from 16 to 64 lifts recall from36\.9%36\.9\\%to56\.7%56\.7\\%while cross\-domain perplexity rises from33\.0633\.06to40\.6640\.66—so rank buys knowledge, and buys it at a worsening exchange rate \(10001000down to563563bits per point\)\. Capacity in this regime is rank\-limited rather than time\-limited, and the rank knob is a position on the efficiency curve rather than a free improvement\.

We report this ablation twice because the first version of it was wrong in an instructive way\. The rank\-16 and rank\-64 runs were originally compared across different rehearsal fractions \(0\.30\.3vs\.0\.10\.1\), which inflated the apparent rank effect from\+19\.8\+19\.8to\+31\.1\+31\.1recall points—and the paper’s own §[5\.5](https://arxiv.org/html/2608.20873#S5.SS5)supplies the evidence that rehearsal fraction matters that much\. Rerunning rank 16 at matched rehearsal also moved its cross\-domain perplexity from45\.8145\.81to33\.0633\.06, i\.e\. the confound was distorting both axes at once\. Comparisons in Table[1](https://arxiv.org/html/2608.20873#S5.T1)now carry their rehearsal fraction explicitly\.

### 5\.8Capacity

Figure 2:Absorbed vs\. presented facts \(log–log\), two points per path\. Both invariant paths scale with exponent≈0\.8\\approx\\\!0\.8at0\.740\.74–1\.03%1\.03\\%cell\-space utilization\. Two points cannot resolve a saturation knee; the claim is only that absorption has not flattened by10410^\{4\}facts\.From10310^\{3\}to10410^\{4\}presented facts, absorbed facts grow→1913315\\\!\\to\\\!1913for A\+ \(α≈0\.78\\alpha\\\!\\approx\\\!0\.78\) and→4130612\\\!\\to\\\!4130for projected dense \(α≈0\.83\\alpha\\\!\\approx\\\!0\.83\), with bin\-space saturation of0\.740\.74–1\.03%1\.03\\%: the regime is optimization\-limited \(training loss1\.431\.43at10410^\{4\}facts, 24 exposures\), not capacity\-limited—five orders of magnitude below thek​NtkN\_\{t\}shipped\-bit ceiling of Prop\.[4](https://arxiv.org/html/2608.20873#Thmproposition4)and far below the∼\\sim2\-bit/parameter full\-precision ceiling of\[[18](https://arxiv.org/html/2608.20873#bib.bib18)\]\. At10410^\{4\}facts the projected dense path absorbs4,1304\{,\}130facts \(4130×11\.9≈494130\\times 11\.9\\approx 49kbit of probed attribute entropy\) under exact invariance, on a single consumer GPU in 3\.3 hours\. Two points cannot resolve a knee, so we claim only that none is visible in this range\.

The cross\-domain measurement at10410^\{4\}facts, which we added after the audit described in §[5\.5](https://arxiv.org/html/2608.20873#S5.SS5), is the most severe result in this paper\. Projected dense training absorbs41\.641\.6kbit while WikiText perplexity stays*below*the 4\-bit anchor \(→11\.6911\.73\\\!\\to\\\!11\.69\)—and LAMBADA rises from28\.7028\.70to129\.74129\.74, a factor of4\.54\.5\. A practitioner watching in\-domain perplexity would conclude nothing had gone wrong\. The recipe matters enormously here: rank\-64 healing with10%10\\%rehearsal absorbs73\.473\.4kbit at the same scale—76%76\\%more knowledge—for64\.4364\.43rather than129\.74129\.74\(20542054bits per point against412412\)\. At10410^\{4\}facts, method choice is worth a factor of five in efficiency, far more than it is worth at10310^\{3\}\.

We also observed a reproducibility issue worth recording: two runs with identical configuration and seed at10410^\{4\}facts differed by 6\.4 recall points \(41\.3%41\.3\\%vs\.34\.9%34\.9\\%\), far beyond probe\-sampling error\. Dense training at this scale is not reproducible to better than a few points on our stack, so the capacity exponents above should be read with that uncertainty\. A10510^\{5\}\-fact point and a bits\-per\-parameter comparison across the scale ladder are the natural next measurements; both are running at the time of writing and neither is reported here\.

### 5\.9Geometry of the safe region

Figure 3:\(a\) Walking from the 4\-bit anchor to the original fp32 weights and past them\. Loss falls monotonically along the whole segment while essentially no weight leaves its cell; beyond the original weights, loss turns up and codes break catastrophically \(19\.6% escaped att=1\.25t\{=\}1\.25\)\. \(b\) Cost of random in\-cell perturbation at radiusρ\\rho, two seeds\. The diagonal\-Fisher budget tracks the wrong magnitude by up to383×383\\times; the measured exponent is 2\.43 against a predicted 2\.How much room is there, and what does using it cost? Two measurements answer this directly \(Fig\.[3](https://arxiv.org/html/2608.20873#S5.F3)\)\.

First, the segment from the anchorW^\\hat\{W\}to the original weightsWorigW\_\{\\mathrm\{orig\}\}lies inside the cells—by the RTN property, and confirmed by code assignment:0%0\\%of weights escape up tot=0\.5t\{=\}0\.5and0\.61%0\.61\\%att=1t\{=\}1\(those are original weights sitting inside the 1% safety margin\)\. Along it, perplexity falls monotonically \(11\.34→10\.3011\.34\\to 10\.30in\-domain,28\.54→26\.0228\.54\\to 26\.02on LAMBADA\): the dequantization gap is not a loss barrier but a downhill corridor\. Extrapolating pastt=1t\{=\}1reverses both: loss rises and escapes jump to19\.6%19\.6\\%then32\.7%32\.7\\%\. The safe region and the useful region coincide, which is the geometric reason this method works at all\.

Second, filling the cells with*random*content costs11\.34→12\.911\.34\\to 12\.9in\-domain and28\.5→34\.128\.5\\to 34\.1cross\-domain atρ=1\\rho\{=\}1, but only11\.34→11\.3911\.34\\to 11\.39atρ=0\.25\\rho\{=\}0\.25: three quarters of the radius is nearly free and the last quarter carries most of the damage, which is what makesρ\\rhoa usable knob for trading capacity against stability \(§[5](https://arxiv.org/html/2608.20873#S5)\)\.

Repeating the whole measurement on Qwen3\-4B replicates the structure and sharpens one conclusion while weakening another\. The segment behaves identically—monotone descent from anchor to original weights \(8\.99→8\.468\.99\\to 8\.46\),0\.81%0\.81\\%of weights outside their cells att=1t\{=\}1, then19\.7%19\.7\\%and32\.7%32\.7\\%att=1\.25t\{=\}1\.25and1\.51\.5, within a tenth of a point of the 1\.7B escape fractions, so the geometry of the safe region appears not to depend much on scale\. What does depend on scale is the price of using it: filling the cells atρ=1\\rho\{=\}1costs the 4B model\+6\.8%\+6\.8\\%in\-domain and\+6\.1%\+6\.1\\%cross\-domain against\+13\.6%\+13\.6\\%and\+19\.4%\+19\.4\\%at 1\.7B—roughly half —which is the same direction as the 27B absorption results \(§[5\.11](https://arxiv.org/html/2608.20873#S5.SS11)\)\. Against that, the exponent we measure is1\.71\.7–1\.91\.9at 4B where it was2\.432\.43at 1\.7B: the quadratic prediction of Proposition[5](https://arxiv.org/html/2608.20873#Thmproposition5)brackets both, but the deviation does not replicate even in sign, and we draw no conclusion from it\. The diagonal Fisher underestimates by278×278\\timesat 4B against383×383\\timesat 1\.7B, so that failure is not a small\-model artifact either\.

### 5\.10Where in the network should knowledge be written?

Not all weights are equally good storage\. We partition the constrained matrices and let CellFill \(r=64r\{=\}64\) write into one partition at a time, freezing the rest—structural invariance means no clipping confound enters the comparison\. Normalization matters: dividing absorbed bits by*constrained*weights asks which substrate holds the most, while dividing by*trainable*parameters asks where a fixed low\-rank budget is best spent\. The two disagree, so we give both, plus the price in cross\-domain perplexity\.

Each partition run quantizes only its own matrices and leaves the rest of the model at original precision, so the seven runs do*not*share a baseline; every row is therefore priced against the anchor archived with it\. \(An earlier version of this analysis priced all seven against the fully\-quantized model’s anchor of 28\.67, which made the two smallest partitions appear to cost nothing\. They do not\.\)

Three findings, two of which cut against the standard picture\.

Module type barely matters per unit of budget\.gate\+up \(130\), full MLP \(118\) and attention \(117\) fall within 11% of each other\. Measured per*constrained*weight, attention appears74%74\\%better than MLP— but that gap is entirely an artifact of attention’s smaller matrices under a per\-matrix rank budget, and it disappears under the normalization an implementer actually pays\. We report this explicitly because the wrong normalization here yields a confident and false headline\.

Depth dominates module type, but only one end of it is measurable\.Late MLP layers absorb113113bits per million trainable parameters against4949and4646for early and middle\. We do not report that as a2\.3–2\.5×2\.3\\text\{\-\-\}2\.5\\timesratio, because the two denominators are not measurements: early and middle MLP recall5\.8%5\.8\\%and6\.0%6\.0\\%, both at or below the6\.5%6\.5\\%marginal\-guess floor of §[9](https://arxiv.org/html/2608.20873#S9)\. What the partition sweep establishes is one\-sided — late MLP absorbs, early and middle did not absorb anything distinguishable from guessing under this budget — and a ratio computed against a floor would put a number on the difference that the data cannot support\. This runs against the localization results of ROME and MEMIT\[[16](https://arxiv.org/html/2608.20873#bib.bib16),[17](https://arxiv.org/html/2608.20873#bib.bib17)\], which place factual associations in*middle*MLP layers, and against the prescription to write throughdown\_proj\(worst non\-degenerate partition here, 72\)\. We read the discrepancy as a genuine difference of task rather than a contradiction: that literature locates where an existing association can be*found and edited*, whereas we measure where a*new*association is most cheaply*written*under a norm constraint\. The middle layers are not idle—they show the highest cell saturation of any partition \(16\.6%\), i\.e\. the optimizer pushes hardest there and gets the least back\.

The best target is the same under both prices\.Ranking by damage rather than by budget—bits absorbed per point of LAMBADA given up—keeps gate\+up on top \(911911\) and keepsdown\_projat the bottom \(210210\), with attention \(705705\) ahead of full MLP \(594594\)\. Early and middle layers are not free, as an incorrect shared\-anchor calculation initially suggested; they degrade their own anchors by1\.51\.5points, which is little only because they absorb little\. A practitioner optimizing absorbed knowledge per unit of forgetting should write to the gate and up projections across all depths— not todown\_proj, and not to the middle of the network\.

### 5\.11Scale, architecture, and family

The mechanism touches only quantized linear layers, so it should be indifferent to what surrounds them\. We test that on a12×12\\timesparameter ladder within one family, on a second family with a different tokenizer and initialization, and on a hybrid model whose blocks are gated linear attention rather than softmax attention\.

The first five rows are CellFillr=64r\{=\}64under one recipe; the last is clip\-merge at 24 epochs, because our CellFill implementation does not fit on one 80 GB A100 at 27B\. The reason is an implementation choice rather than anything in the method: each wrapped layer stores its per\-weight half\-widthMMdensely, which over2\.435×10102\.435\\times 10^\{10\}constrained weights is48\.748\.7GB in bf16 on top of the12\.212\.2GB of 4\-bit weights\. Gradient checkpointing of the fill \(which we added, and which suffices at 8B\) does not close a6161GB gap, and restricting the fill togate\_projandup\_projdoes not either, since those are the largest matrices in the model\.MMis a deterministic function of the frozen codes and scales, so it can be recomputed per block instead of stored; we have not implemented that, and it is the single change that would put 27B CellFill within reach\. Raw recall is*not*monotone in scale—0\.6B absorbs more than 1\.7B—and we do not have an explanation we can defend; a smaller model has less to protect, but we have not tested that\. Efficiency measured as knowledge per point of cross\-domain perplexity rises steeply and cleanly with scale,η∝N0\.42\\eta\\propto N^\{0\.42\}across the five matched runs—but most of that slope is not ours\. WritingΔ=log⁡\(ppl1/ppl0\)\\Delta=\\log\(\\mathrm\{ppl\}\_\{1\}/\\mathrm\{ppl\}\_\{0\}\)for the*relative*cross\-domain damage and differentiatinglog⁡η=log⁡K−log⁡ppl0−log⁡\(eΔ−1\)\\log\\eta=\\log K\-\\log\\mathrm\{ppl\}\_\{0\}\-\\log\(e^\{\\Delta\}\-1\)inlog⁡N\\log Ngives

aη=aK−appl0−a\(eΔ−1\),a\_\{\\eta\}\\;=\\;a\_\{K\}\\;\-\\;a\_\{\\mathrm\{ppl\}\_\{0\}\}\\;\-\\;a\_\{\(e^\{\\Delta\}\-1\)\},and on our ladderaK=\+0\.09a\_\{K\}=\+0\.09andaΔ=−0\.03a\_\{\\Delta\}=\-0\.03are both near zero whileappl0=−0\.30a\_\{\\mathrm\{ppl\}\_\{0\}\}=\-0\.30: the identity closes \(0\.09\+0\.30\+0\.03=0\.420\.09\+0\.30\+0\.03=0\.42\) and71%71\\%of the exponent is the base models’ own perplexity scaling, inherited by any metric with perplexity points in its denominator\. A point of perplexity simply means more at a model whose perplexity is lower\.

The scale\-invariant statement is the one worth making\. Knowledge per*nat*of relative damage,K/ΔK/\\Delta, varies by1\.5×1\.5\\timesacross the ladder \(17,973 to 27,472, cv17%17\\%\) whereη\\etavaries by3\.2×3\.2\\times, and the relative damage itself is flat atΔ=0\.40\\Delta=0\.40\(cv15%15\\%\) across a12×12\\timesrange in parameters and two model families\. What in\-cell learning holds roughly fixed with scale is not the absolute price of knowledge but the*fraction*of the anchor’s cross\-domain ability it costs\. The 27B point sits far above theη\\etatrend, but it is a different path, a different attention mechanism and a single seed, so we quote it as an observation rather than part of any fit\.

We resist the reading this invites\. It is tempting to say the price of in\-cell learning falls as models grow—the models actually shipped as 4\-bit artifacts are the large ones—but our own scale\-invariant metric does not show that:Δ\\Deltais flat at0\.400\.40\(cv15%15\\%\) andK/ΔK/\\Deltaspans only1\.5×1\.5\\timeswith cv17%17\\%across a12×12\\timesparameter range\. The falling price is carried by the 27B point, which the paragraph above excludes from the fit for three separate reasons\. A point cannot be excluded from a fit and then be allowed to carry the conclusion\. What the ladder supports is that the relative cost does*not*rise with scale, which is the property a deployment needs; whether it falls is unresolved here\. The 27B run also shows the difference is not only quantity: at 1\.7B recall is dominated by the probe that continues the training sentence almost verbatim \(city65\.9%65\.9\\%against company17\.7%17\.7\\%at10410^\{4\}facts\), while at 27B the paraphrased probe reaches53\.5%53\.5\\%and the gap between probe kinds narrows sharply \(city74\.8%74\.8\\%, occupation88\.0%88\.0\\%\)\.

#### A divergence, and what it says about the bound\.

Our first Mistral\-7B run did not work at all: transplanting the 1\.7B hyperparameters \(r=16r\{=\}16, lr2×10−42\\times 10^\{\-4\}\) drove plain LoRA to perplexity23112311*before any merging*, with a39\.6%39\.6\\%clip rate and only40%40\\%of the update norm surviving projection—all signatures of a diverged optimization, faithfully reproduced by clip\-merge\. We report it because the archived file is public and because the contrast is informative: CellFill on the same model at*five times*that learning rate \(10−310^\{\-3\}\) trained cleanly to the71\.6%71\.6\\%in the table\. This is what the parameterization predicts\. Since the fill isM⊙tanh⁡\(⋅\)M\\odot\\tanh\(\\cdot\)with\|tanh\|<1\|\\tanh\|<1, the update cannot leave the cell however large the gradients become, so the bound is a stability property and not only an invariance property\. Lowering the learning rate to5×10−55\\times 10^\{\-5\}makes plain LoRA work on Mistral, and work well—34\.0%34\.0\\%recall at25572557bits per point, the second most efficient run in this paper—so the divergence was a transplanted hyperparameter and nothing more\. What survives is the margin: the bounded parameterization was stable at a learning rate20×20\\timeslarger than the one plain LoRA needed, which is the practical form the guarantee takes\.

#### The controlled version of that observation\.

The anecdote above confounds two things, so we ran the matched sweep: one model \(Qwen3\-1\.7B\), one seed, rank1616, rehearsal0\.10\.1,2424epochs, and only the parameterization and the learning rate varying\.

At10−310^\{\-3\}plain LoRA has already damaged the model before anything is merged—WikiText34\.3034\.30against an anchor of11\.7111\.71, LAMBADA311\.1311\.1against28\.6728\.67—and clip\-merge, which must clip63\.3%63\.3\\%of the weights and keeps37\.5%37\.5\\%of the update norm, delivers6\.3%6\.3\\%recall\. Tripling the learning rate destroys it outright\. CellFill is stable at both settings and*better*at the larger one: recall34\.0→54\.2%34\.0\\to 54\.2\\%, for WikiText10\.51→11\.8210\.51\\to 11\.82and LAMBADA32\.85→39\.5532\.85\\to 39\.55\. The mechanism is in the last column\. The excess gradient never leaves the cells; it is absorbed by saturation, which rises from0\.1%0\.1\\%to47\.5%47\.5\\%of coordinates pressed against\|tanh\|→1\|\\tanh\|\\to 1\. Divergence is a cliff and saturation is a ceiling, and the parameterization converts the first into the second\. This is the paper’s most direct evidence on catastrophic forgetting, in the term’s original sense of an abrupt collapse of prior capability rather than a graded cost: at3×10−33\\times 10^\{\-3\}the LoRA arm loses everything it had—WikiText2\.1×1042\.1\\times 10^\{4\}, LAMBADA1\.2×1051\.2\\times 10^\{5\}, recall00—while the bounded arm at the identical learning rate, seed and rank keeps WikiText within0\.110\.11of the anchor and returns its best recall of the sweep\. The collapse is not avoided by tuning; it is unavailable, because no gradient can move a weight past a wall the grid fixed before training began\. What remains is a graded cost, and §[5\.3](https://arxiv.org/html/2608.20873#S5.SS3)and §[5\.6](https://arxiv.org/html/2608.20873#S5.SS6)measure it\. Notably, the higher learning rate at rank1616buys most of what quadrupling the rank buys \(54\.2%54\.2\\%against56\.7%56\.7\\%forr=64r\{=\}64at10−310^\{\-3\}, at592592against563563bits/pt\), so rank and step size are partly substitutable inside the cells\.

Two caveats\. This is one seed per cell, so we report the direction, which is large, and not the size\. And the CellFill10−310^\{\-3\}cell is a repeat of the configuration in Table[1](https://arxiv.org/html/2608.20873#S5.T1), which recorded36\.9%36\.9\\%: the same code, the same seed and the same hyperparameters on different GPUs \(RTX 4090 against A100\) differ by2\.92\.9recall points, since we do not force deterministic kernels\. That gap is a lower bound on the noise in everyn=1n\{=\}1comparison in this paper, and we have not subtracted it anywhere\.

### 5\.12The lifecycle: four sequential updates, one artifact

Everything so far injects knowledge once\. The claim in our title is about a*sequence*, so we run one: four disjoint 500\-fact tasks arriving in order, each written into the same residual with CellFill \(r=64r\{=\}64\), with the artifact verified after every task\. Between tasks the accumulated in\-cell position becomes the next task’s anchor and the remaining room shrinks to the distance from there to the cell walls, so the budget is literally consumed as the sequence proceeds\.

Four things happen at once, and only the first is good news\. The artifact survives the whole sequence bit\-for\-bit—zero violations after every one of the four tasks, over1\.4×1091\.4\\times 10^\{9\}weights each time—so the guarantee is not a single\-shot property\. Cross\-domain ability degrades monotonically and we report it here rather than only the recall columns: LAMBADA runs39\.42→43\.5739\.42\\to 43\.57across the sequence, which against the28\.6728\.67anchor is a37%37\\%cost for the first task rising to52%52\\%after the fourth, while in\-domain perplexity moves only11\.21→11\.6211\.21\\to 11\.62—the same divergence between the two metrics that §[5\.5](https://arxiv.org/html/2608.20873#S5.SS5)attributes to rehearsal sharing a corpus with WikiText\. We have no unconstrained control sequence to price that52%52\\%against, so it is a measurement of this method’s cost over four tasks and not a comparison\. Old tasks are forgotten: T0 retains31%31\\%of its recall after three subsequent updates\. And, distinctively for this setting,*plasticity itself decays*: each task learns less than the one before \(50\.9→42\.7→35\.8→29\.7%50\.9\\to 42\.7\\to 35\.8\\to 29\.7\\%\) as the writable room falls by44%44\\%\. The second effect is ordinary catastrophic forgetting; the third is specific to learning in a bounded space\.

#### The decay is geometric, and it is predicted rather than fitted\.

Proposition[6](https://arxiv.org/html/2608.20873#Thmproposition6)says a bounded update consumes a constant*fraction*of the remaining room, because the displacement is scaled by the room itself\. The measured ratios arer¯k/r¯k−1=0\.8271,0\.8254,0\.8238\\bar\{r\}\_\{k\}/\\bar\{r\}\_\{k\-1\}=0\.8271,\\,0\.8254,\\,0\.8238: geometric to three decimal places,β=0\.825±0\.002\\beta=0\.825\\pm 0\.002\. New\-task absorption follows the same law,0\.8377,0\.8391,0\.83050\.8377,\\,0\.8391,\\,0\.8305, and is proportional to the room available before the task \(ak/r¯k−1=142,144,145a\_\{k\}/\\bar\{r\}\_\{k\-1\}=142,\\,144,\\,145fork≥2k\\geq 2\)\.

An earlier version of this paper offered a second, sharper claim here, and measuring it refuted it\. The argument was that sinceβ=1−𝔼​\|t\|\\beta=1\-\\mathbb\{E\}\|t\|, the*implied*𝔼​\|t\|=1−r¯k/r¯k−1\\mathbb\{E\}\|t\|=1\-\\bar\{r\}\_\{k\}/\\bar\{r\}\_\{k\-1\}must be constant across tasks; it is \(0\.1729,0\.1746,0\.17620\.1729,\\,0\.1746,\\,0\.1762, a spread of1\.8%1\.8\\%\), and we read that as independent confirmation\. It is not independent: those three numbers are one minus the three ratios above, so their constancy andβ\\beta’s are the same statement\. Testing the claim requires𝔼​\|tanh\|\\mathbb\{E\}\|\\tanh\|measured from the fill actually applied at each fold, which we now log\. It is0\.4429,0\.4617,0\.5027,0\.51800\.4429,\\,0\.4617,\\,0\.5027,\\,0\.5180—neither equal to the implied value \(they differ by2\.8×2\.8\\times\) nor constant \(it rises17%17\\%as the room falls\)\. The identityβ=1−𝔼​\|t\|\\beta=1\-\\mathbb\{E\}\|t\|is therefore false as an identity, exactly as Proposition[6](https://arxiv.org/html/2608.20873#Thmproposition6)now states: it provesr′≥r⁡\(1−\|t\|\)r^\{\\prime\}\\geq r\(1\-\|t\|\), and the data confirm the inequality is strict and loose,0\.820\.82against the0\.510\.51the naive law predicts\. The slack is the cell asymmetry giving room back—a fill that moves toward the far wall recentres its anchor—and the implied geometry factorβ/\(1−𝔼​\|tanh\|\)\\beta/\(1\-\\mathbb\{E\}\|\\tanh\|\)is1\.52,1\.65,1\.721\.52,\\,1\.65,\\,1\.72\.

What survives is the empirical law and not our explanation of it: the room decays geometrically withβ\\betastable to1\.7%1\.7\\%over four tasks, and that is what makes lifetime capacity finite\. Whyβ\\betais so stable while both of its putative factors drift is open\.

Two numbers follow\. Without consolidation the lifetime capacity isa1/\(1−β\)=292%a\_\{1\}/\(1\-\\beta\)=292\\%of a first\-task equivalent—finite, though no constraint is ever violated on the way—and90%90\\%of it is spent by task 13\. This is what we mean by the budget being operational: it can be estimated after two tasks and it predicts when the model will stop being able to learn\.

#### Consolidation renews the budget, at an explicit price\.

Repeating the sequence with a re\-quantization after T1—a deliberate major version, not silent drift—restores the room \(2\.49→2\.98×10−32\.49\\to 2\.98\\times 10^\{\-3\},99%99\\%of its original value\) and with it the plasticity: the fourth task recovers33\.4%33\.4\\%instead of29\.7%29\.7\\%, and for the first time in the sequence a later task outperforms its predecessor \(32\.8→33\.4%32\.8\\to 33\.4\\%\)\. The price is stated exactly:19\.0%19\.0\\%of the 4\-bit codes change at that step, which is precisely why it must be a version bump rather than an update, and the oldest task’s retention drops further \(15\.7→12\.2%15\.7\\to 12\.2\\%\)\.

Consolidation is not free in the short run either, and the cumulative absorption shows why\. Re\-quantizing injects fresh quantization error, so the task immediately after it does*worse*than it would have without \(32\.8%32\.8\\%against35\.8%35\.8\\%\), and only the following task pulls ahead \(33\.4%33\.4\\%against29\.7%29\.7\\%\)\. Cumulatively the two policies cross between the third and fourth task \(129\.4%129\.4\\%vs\.126\.4%126\.4\\%, then159\.1%159\.1\\%vs\.159\.8%159\.8\\%\): consolidation takes about two tasks to repay its own cost\. A policy consolidating everymmtasks therefore has to be chosen against the sequence length, and Proposition[6](https://arxiv.org/html/2608.20873#Thmproposition6)gives the throughput to compare—46\.5%46\.5\\%per task atm=2m\{=\}2against39\.1%39\.1\\%atm=4m\{=\}4, versus zero in the limit for a policy that never consolidates\.

Minor versions keep the artifact and spend the budget; a major version buys the budget back by giving up the artifact\. We report this from a single sequence on one model, so the constantβ\\betashould be read as a demonstration that the law is measurable, not as a value that transfers\.

## 6Shipping the update: nested checkpoints

An update that cannot be distributed is not an update\. Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)says the in\-cell position can be encoded withkkbits per weight such that truncating the stream returns the released artifact exactly; here we measure what a recipient gets at each depth\. We take the healed weights of the10410^\{4\}\-fact dense run \(§[5\.8](https://arxiv.org/html/2608.20873#S5.SS8)\), re\-encode every weight’s cell position atk∈\{1,2,3,4\}k\\in\\\{1,2,3,4\\\}, and evaluate the reconstruction\.

Every depth re\-quantizes bit\-identically—checked in the integer domain over all1\.409×1091\.409\\times 10^\{9\}constrained weights at everykk, zero violations—so the nesting property of Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)holds in practice and a truncated stream is a valid, older, still\-certified checkpoint\.

The quality curve is not what a compression result usually looks like\. Two bits per weight recover99\.4%99\.4\\%of the knowledge that a full fp16 residual carries \(34\.7%34\.7\\%vs\.34\.9%34\.9\\%\) at one eighth the payload, and*three bits exceed it*\(35\.5%35\.5\\%\)\. Coarsening the residual is not merely lossy storage: it quantizes each weight back toward its anchor, which is the same shrinkage that makes projection help in §[5\.6](https://arxiv.org/html/2608.20873#S5.SS6)\. The effect is visible on the strict axis, where every depth beats the fp16 residual— LAMBADA87\.087\.0atk=1k\{=\}1against129\.2129\.2for fp16—and in the efficiency column, wherek∈\{1,2\}k\\in\\\{1,2\\\}reach523523bits per point against fp16’s413413\. The refinement depth is therefore a third operating knob alongside rank and radius, not just a transport format:k=1k\{=\}1ships16×16\\timesless and forgets less, at the price of a quarter of the knowledge\.

We report this as a single\-configuration study on one archived run\. The comparison we cannot yet make is to BitDelta\[[12](https://arxiv.org/html/2608.20873#bib.bib12)\], which compresses an*unconstrained*fine\-tuning delta to one bit; the constrained and unconstrained deltas are different objects and a fair comparison needs both pipelines on the same task\.

## 7Related work

Frozen quantized bases with residuals\.QLoRA\[[1](https://arxiv.org/html/2608.20873#bib.bib1)\]freezes an NF4 base and trains BF16 low\-rank adapters; LoftQ\[[2](https://arxiv.org/html/2608.20873#bib.bib2)\]initializes adapters from the SVD of the quantization residualW−W^W\-\\hat\{W\}—the very object our constraint set is built around\. Neither constrains the merged weights to the quantization cells, so merging changes the artifact\. QA\-LoRA\[[3](https://arxiv.org/html/2608.20873#bib.bib3)\]and low\-rank QAT\[[4](https://arxiv.org/html/2608.20873#bib.bib4)\]take the complementary road: they alter the int4 weights to absorb the adapter\. We occupy the unexplored corner: dense residual learning under a bitwise*no\-change*guarantee for the artifact\.

Sub\-cell structure carries signal\.AdaRound\[[5](https://arxiv.org/html/2608.20873#bib.bib5)\]and BRECQ\[[6](https://arxiv.org/html/2608.20873#bib.bib6)\]showed that choosing rounding directions within quantization cells measurably changes loss—evidence that the dequantization gap has usable information capacity, which we exploit for continual learning rather than calibration\.

Nested and progressive precision\.Matryoshka representation learning\[[29](https://arxiv.org/html/2608.20873#bib.bib29)\]and Matryoshka quantization\[[7](https://arxiv.org/html/2608.20873#bib.bib7)\], any\-precision LLMs\[[8](https://arxiv.org/html/2608.20873#bib.bib8)\], NestQuant\[[9](https://arxiv.org/html/2608.20873#bib.bib9)\], MatGPTQ\[[10](https://arxiv.org/html/2608.20873#bib.bib10)\], and recurrent residual quantization\[[11](https://arxiv.org/html/2608.20873#bib.bib11)\]build bit\-sliced models whose truncations serve multiple precisions\. Our refinement code \(Prop\.[3](https://arxiv.org/html/2608.20873#Thmproposition3)\) shares the nesting mechanics but deploys it along the*version/time*axis: low bits are writable by learning, high bits are immutable across releases\. BitDelta\[[12](https://arxiv.org/html/2608.20873#bib.bib12)\]compresses fine\-tuning deltas to 1 bit for multi\-tenant serving \(cf\. S\-LoRA\[[13](https://arxiv.org/html/2608.20873#bib.bib13)\]\); our deltas are additionally box\-bounded, giving the invariance and rollback properties\.

Continual learning and editing\.EWC\[[14](https://arxiv.org/html/2608.20873#bib.bib14)\]softly anchors weights by Fisher information; rehearsal\[[15](https://arxiv.org/html/2608.20873#bib.bib15)\]replays old data\. Our box constraint is a hard, per\-weight trust region whose radii come free from the quantization grid, and our results show it must be combined with fresh rehearsal \(direction shaping\) since it bounds only magnitude\. Knowledge editing \(ROME\[[16](https://arxiv.org/html/2608.20873#bib.bib16)\], MEMIT\[[17](https://arxiv.org/html/2608.20873#bib.bib17)\], MEND\[[28](https://arxiv.org/html/2608.20873#bib.bib28)\]\) writes individual facts via closed\-form updates without invariance guarantees\. GRACE\[[27](https://arxiv.org/html/2608.20873#bib.bib27)\]is the closest existing analogue to our revocability claim—it stores edits in a discrete codebook that can be removed—but it does so in an auxiliary memory consulted at inference, whereas we leave the released artifact bit\-identical and add no inference\-time structure; knowledge capacity scaling laws\[[18](https://arxiv.org/html/2608.20873#bib.bib18)\]\(∼\\sim2 bits/parameter; int4 degrades capacity\) supply the measurement methodology and ceiling reference for our bit accounting\. Weight interpolation studies \(linear mode connectivity\[[19](https://arxiv.org/html/2608.20873#bib.bib19)\], model soups\[[20](https://arxiv.org/html/2608.20873#bib.bib20)\], WiSE\-FT\[[21](https://arxiv.org/html/2608.20873#bib.bib21)\]\) established that base–finetune segments are often low\-loss; the quantization segment\[W^,W\]\[\\hat\{W\},W\]is a within\-cell special case that we measure directly\.

Quantization stacks\.We build on bitsandbytes/LLM\.int8\[[22](https://arxiv.org/html/2608.20873#bib.bib22)\]NF4\. GPTQ\[[23](https://arxiv.org/html/2608.20873#bib.bib23)\]and AWQ\[[24](https://arxiv.org/html/2608.20873#bib.bib24)\]produce uniform asymmetric grids, for which the cell arithmetic of §[2](https://arxiv.org/html/2608.20873#S2)applies unchanged \(the level table becomes uniform\); we have not run those stacks and do not claim results on them\.

## 8Discussion and outlook

What the contract changes about shipping a model\.The practical content of Propositions[1](https://arxiv.org/html/2608.20873#Thmproposition1)–[5](https://arxiv.org/html/2608.20873#Thmproposition5)is that an update stops being a replacement and becomes an increment stated against a fixed baseline\. Four consequences follow, none of which requires trusting the party that produced the update\. First, verification is mechanical and performable by a third party: given the released\(a,s\)\(a,s\), the recipient assigns codes to the served weights in the integer domain and compares—a hash over the code tensor suffices—and this is exactly the check we run on every merge in this paper, over1\.409×1091\.409\\times 10^\{9\}constrained weights at 1\.7B and2\.435×10102\.435\\times 10^\{10\}at 27B \(§[5\.11](https://arxiv.org/html/2608.20873#S5.SS11)\)\. Two conditions make the check meaningful and both are easy to get wrong: codes must be assigned under the*shipped*scales rather than by re\-running the quantizer \(§[A](https://arxiv.org/html/2608.20873#A1)\), and the residual’s storage dtype must respect the cell margin\. Second, rollback is exact rather than approximate: discardingδ\\deltareturnsW^\\hat\{W\}bit\-for\-bit, so revoking an update—a fact that must be deleted, a bad training batch—is a truncation, not a retrain\. Third, the excursion is bounded by a radius chosen before training and measured after it\. Fourth, because the integer codes and scales are identical across every residual derived from them, one cached base can back many concurrent residuals, in the manner of multi\-tenant adapter serving\[[13](https://arxiv.org/html/2608.20873#bib.bib13),[12](https://arxiv.org/html/2608.20873#bib.bib12)\], with the addition that each tenant’s*merged*weights still certify against the shared base bits\. We should be plain about the cost side of that last point: our residuals are dense, and akk\-bit dense refinement overNNweights costsk​NkNbits, which atk=4k\{=\}4is the size of the 4\-bit artifact itself\. Making the increment small is precisely what the sparsity mask of Proposition[4](https://arxiv.org/html/2608.20873#Thmproposition4)and the nested code of Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)are for, and §[6](https://arxiv.org/html/2608.20873#S6)measures what that costs: two bits per weight recover99\.4%99\.4\\%of the fp16 residual’s recall at one eighth the payload\.

A version lifecycle\.The contract suggests an obvious release discipline\. Minor versions write refinement bits into the cells: the shipped 4\-bit file is unchanged across the whole sequence, each increment is verifiable and revocable, and capacity is drawn down from the writable radius\. A major version*consolidates*: the current weights are re\-quantized into a new frozen artifact, the residual is zeroed, and the writable radius is restored\. The artifact changes here—by design, at a declared moment, with a new hash—which is the opposite of the silent drift that ordinary fine\-tuning produces\.experiments/exp\_seq\.pyimplements both modes, including the fold that lets the residual accumulate across tasks and the re\-anchoring that follows consolidation\. §[5\.12](https://arxiv.org/html/2608.20873#S5.SS12)runs both: capacity is renewable—consolidation restores99%99\\%of the writable room—but at a stated price,19\.0%19\.0\\%of the codes change, and cumulative forgetting is measurable, with the first task retaining31%31\\%of its recall after three later updates\. How the geometric constantβ\\betaand the consolidation break\-even point vary with model, task and optimizer remains open\.

Where the guarantee is load\-bearing\.It is worth being narrow\. The contract pays for itself where*re\-certification*, not compute, is the bottleneck: a fleet whose evaluation harness, caches, or acceptance tests are keyed to a specific checkpoint; an on\-device deployment that must be able to revert exactly and cheaply; a multi\-tenant server that wants a shared, verifiable base under many customer\-specific updates; a provider that must be able to demonstrate which bits a given update did and did not change\. It does*not*pay where retraining and re\-evaluating from scratch is affordable\. And a bit\-identical artifact establishes something narrower than it may sound: it establishes that the reference configuration is unchanged and recoverable, that the update is a bounded, revocable increment, and that the increment can be audited\. It does*not*establish that the served modelW^\+δ\\hat\{W\}\+\\deltainherits any property that was certified ofW^\\hat\{W\}—our own measurements say it does not, since cross\-domain perplexity rises with absorption, up to28\.67→71\.028\.67\\to 71\.0for the most damaging path in Table[1](https://arxiv.org/html/2608.20873#S5.T1)\. We make no claim about regulatory sufficiency; what the mechanism supplies is a stable referent for a re\-evaluation, not a substitute for one\.

Relationship to retrieval\.In\-weight injection and retrieval augment different failure modes—retrieval keeps facts editable and attributable at query time and pays for them in context length and latency at every call; in\-weight facts are paid for once and are available to inference without a retrieval step, but are diffuse and, absent something like this contract, hard to revoke\. They are complementary rather than competing, and we want to be explicit that*we have not run the comparison*: no RAG baseline appears anywhere in this paper, and none of our numbers speak to when one should be preferred over the other\.

Open problems, ranked\.\(1\)*Skills, not facts\.*Everything demonstrated here is factual recall on synthetic biographies, chosen because novelty is certain and information content is countable\. Whether procedural or stylistic capability can be written into the cells is untested, and we regard it as the single largest open question about the method’s reach: the per\-weight box bounds magnitude, and skills may require coordinated displacement that the box shapes badly\. \(2\)*Gauge freedom as a way to enlarge the writable radius\.*The cells are defined per coordinate by the frozen grid, so a weight’s writable half\-width depends on where it happens to sit relative to its walls; the transformer parameterization has exact function\-preserving symmetries \(head permutations carried through the output projection; positive rescalings absorbed by a following normalization\) that move weights without changing the function\. We flag as*conjecture*that gauge\-fixing along such an orbit before freezing could increase total writable volume, and note one thing that follows immediately: a uniform rescaling buys nothing, since blockwise absmax rescales with it and the relative half\-widths are unchanged\. Any gain must come from per\-channel or per\-head freedom that redistributes magnitude*across*quantization blocks\. Note also that this is a release\-time choice—it changes the artifact—not a post\-hoc trick\. \(3\)*Interference\-limited capacity\.*Absorption grows with exponent≈0\.8\\approx\\\!0\.8from10310^\{3\}to10410^\{4\}facts at under1%1\\%cell\-space utilization and roughly10−610^\{\-6\}of the shipped\-bit ceiling of Proposition[4](https://arxiv.org/html/2608.20873#Thmproposition4)\(§[5](https://arxiv.org/html/2608.20873#S5)\), i\.e\. we are nowhere near any bound we can compute\. The regime where facts begin to interfere with one another rather than with the base model—plausibly10510^\{5\}–10610^\{6\}facts—is where the interesting capacity law lives, and we have not reached it; two\-point extrapolations from our data are order\-of\-magnitude conjecture and nothing more\. \(4\)*Uniform\-grid anchors\.*The cell arithmetic of §[2](https://arxiv.org/html/2608.20873#S2)applies verbatim to the uniform asymmetric grids of GPTQ\[[23](https://arxiv.org/html/2608.20873#bib.bib23)\]and AWQ\[[24](https://arxiv.org/html/2608.20873#bib.bib24)\]—only the level table changes—but we have implemented and measured only NF4, and until those stacks are run the claim of generality across quantizers is an argument, not a result\.

#### A speculation, labelled as one\.

It is tempting to read this work as a step toward models that keep learning on their own after release, and we want to state precisely how far the evidence reaches in that direction and where it stops\. What we demonstrate is that a released artifact can absorb a sequence of operator\-issued increments over its deployed lifetime while the bits everyone else depends on never change, and that the resulting lifecycle—spend the budget across minor versions, declare a major version when it runs out—has measurable exchange rates rather than metaphorical ones\. Extrapolating from that to autonomous self\-improvement requires crossing three boundaries this paper does not cross\. The increments here are*facts*, not skills or reasoning, and nothing we measure says the latter fit in a cell\. Every update is*operator\-driven*: a curated corpus, a chosen rehearsal fraction, a chosen radius; the model initiates nothing\. And the budget is*finite and decaying*—plasticity falls by42%42\\%over four tasks—so continuation depends on a human\-declared consolidation that, by design, changes the artifact\. A model that accumulates knowledge indefinitely without supervision is not what we built, and the failure modes we measured are the reasons it does not follow\. What we do claim is narrower and, we think, more useful: the update loop that such a system would need— auditable, revocable, bounded, and cheap enough to run nightly—turns out to exist, and to cost less than the field assumed\.

## 9Limitations

We state the boundaries of the evidence in more detail than is customary, because an earlier draft of this work contained three headline claims that the project’s own archived results contradicted, and the corrections \(§[5\.5](https://arxiv.org/html/2608.20873#S5.SS5), §[5\.3](https://arxiv.org/html/2608.20873#S5.SS3)\) were only possible because the raw files were kept\. Every caveat below is checkable againstresults/\*\.json\.

#### One model family, one quantizer, one architecture per claim\.

Every number in §[5](https://arxiv.org/html/2608.20873#S5)comes from a single base model, Qwen3\-1\.7B\-Base, quantized with a single stack \(bitsandbytes NF4\[[22](https://arxiv.org/html/2608.20873#bib.bib22)\], blocksize 64, double quantization\), except §[5\.11](https://arxiv.org/html/2608.20873#S5.SS11)and the Qwen3\-4B replication of the geometry sweep in §[5\.9](https://arxiv.org/html/2608.20873#S5.SS9)\. The 27B linear\-attention runs—two of them, at 8 and 24 epochs—are the only evidence that the mechanism is architecture\-agnostic\. A second model family and tokenizer \(Mistral\-7B\-v0\.3\) has been run, atn=1n\{=\}1per configuration; no second quantization grid has been\. The claim in §[8](https://arxiv.org/html/2608.20873#S8)that GPTQ\[[23](https://arxiv.org/html/2608.20873#bib.bib23)\]and AWQ\[[24](https://arxiv.org/html/2608.20873#bib.bib24)\]uniform grids admit the same cell arithmetic is an argument about the level table, not a measurement\. Results should be read as a case study on one artifact, not as a property of 4\-bit models in general\.

#### Seed counts, and what the error bars can and cannot exclude\.

Five arms have three seeds, and none has more\. Recall and in\-domain perplexity are averaged over seeds\{0,1,2\}\\\{0,1,2\\\}for path A, path A\+, path B at rehearsal0\.30\.3, the unconstrained control, and CellFillr=16r\{=\}16at rehearsal0\.30\.3\(±5\.2%25\.6\\\!\\pm\\\!5\.2\\%\); cross\-domain LAMBADA is averaged over two seeds for all of them, because that harness was added after the seed\-0 runs\. Two points are not a confidence interval, and we do not treat the LAMBADA spreads — or path B at rehearsal0\.10\.1, which has two seeds — as one\. Everything else in the paper isn=1n\{=\}1: both CellFill rows of Table[1](https://arxiv.org/html/2608.20873#S5.T1)as configured there \(rank 16 at rehearsal0\.10\.1, and rank 64, the strongest constrained operating point\), all seven partitions of §[5\.10](https://arxiv.org/html/2608.20873#S5.SS10), both capacity points at10410^\{4\}facts, the 96\-epoch exposure ablation, theρ=0\.5\\rho\{=\}0\.5radius ablation, both rehearsal ablations, and the entire 27B result\. The between\-seed standard deviations we do measure are large relative to the differences being discussed —±5\.2\\pm 5\.2recall points for CellFillr=16r\{=\}16,±2\.1\\pm 2\.1for A\+,±7\.0\\pm 7\.0for B,±8\.7\\pm 8\.7for the unconstrained control, and±10\.6\\pm 10\.6across the two matched\-rehearsal B seeds — so single\-seed comparisons of a few points are not interpretable, and we avoid making them\. The paired test in §[5\.3](https://arxiv.org/html/2608.20873#S5.SS3)hasdf=2\\mathrm\{df\}\{=\}2and correspondingly almost no power: its interval\[−5\.0,\+4\.0\]\[\-5\.0,\+4\.0\]is consistent with exact invariance costing five recall points and with it gaining four\. We claim only that no cost is detectable at this sample size, not that none exists\. Finally, the seed also selects the fact corpus \(synth\_facts\.generateis seeded\), so a reported standard deviation mixes optimization noise with corpus resampling; the compensating benefit is that arms sharing a seed see an identical fact set and rehearsal stream, which is what licenses the paired analysis\.

#### Arms in Table[1](https://arxiv.org/html/2608.20873#S5.T1)are not matched on rehearsal\.

Five of the seven rows now share rehearsal0\.10\.1; the unconstrained control and its paired path\-B partner remain at0\.30\.3, because the paired test of §[5\.3](https://arxiv.org/html/2608.20873#S5.SS3)requires them to share seeds and rehearsal with each other, not with the rest of the table\. Those last two rows therefore cannot be compared row\-to\-row with the first five, and the166166against311311bits/pt gap between the two*projected*dense arms is mostly the rehearsal difference: our ablation prices30%30\\%against10%10\\%at7\.77\.7recall points and0\.470\.47perplexity points on path A \(17\.4%17\.4\\%/10\.7810\.78versus25\.1%25\.1\\%/10\.3110\.31\)\. Within the matched block the individual differences are readable; across the boundary only the ordering is\.

#### In\-domain perplexity is contaminated by rehearsal, everywhere\.

Rehearsal snippets are drawn from the WikiText train split \(auto\-switching to WikiText\-103 train for pools above 4,000 paragraphs\) while the in\-domain metric is the WikiText\-2 test split\. The splits are disjoint, the domain is not\. This affects*every*WikiText number in this paper without exception: the frontier column of Table[1](https://arxiv.org/html/2608.20873#S5.T1), the10\.7810\.78and10\.3110\.31of §[5\.5](https://arxiv.org/html/2608.20873#S5.SS5), the11\.6111\.61/11\.7611\.76of the capacity experiments, and the 27B figure of7\.257\.25against its7\.447\.44anchor — the 27B result reproduces the same below\-anchor pattern and should be read the same way, as a rehearsal effect rather than as repair of quantization damage\. The controls in §[5\.5](https://arxiv.org/html/2608.20873#S5.SS5)\(an unmerged, entirely unconstrained adapter scoring10\.29910\.299against the box\-constrained merge’s10\.27110\.271\) establish that in\-cell storage is not what produces the gain\. We keep the in\-domain column only for continuity with the ablation history and because it is the axis on which the fixed\-buffer failure mode \(24\.6→17224\.6\\to 172\) is visible; no claim in the paper should rest on it\.

#### The headline capacity number and its cross\-domain price come from different runs\.

The4,1304\{,\}130\-fact headline comes from the run that recorded no LAMBADA\. The cross\-domain price quoted in §[5\.8](https://arxiv.org/html/2608.20873#S5.SS8)\(28\.70→129\.7428\.70\\to 129\.74\) comes from a re\-run at identical configuration and seed that absorbed3,4943\{,\}494facts—the6\.46\.4\-point reproducibility gap reported there—so the largest absorption and its measured price must not be quoted as one operating point\. At10310^\{3\}facts the same method degrades LAMBADA by roughly150%150\\%over its anchor\.

#### Capacity exponents are ratios of two points, not fits\.

α≈0\.78\\alpha\\approx 0\.78\(A\+\) andα≈0\.83\\alpha\\approx 0\.83\(B\) each come from exactly two measurements,10310^\{3\}and10410^\{4\}presented facts, one seed each\. A power law through two points has zero residual degrees of freedom: the exponent is an arithmetic ratio, curvature is invisible by construction, and no saturation knee could be detected even if one existed inside the range\. Cell occupancy at10410^\{4\}facts is0\.74%0\.74\\%\(B\) and1\.03%1\.03\\%\(A\+\), so the regime is far from full, but “no knee is visible between10310^\{3\}and10410^\{4\}” is the entire claim\. The order\-of\-magnitude extrapolations in our internal analysis \(a10510^\{5\}point, a∼\\sim1 Mbit interference ceiling from linear extrapolation of occupancy\) are not claims of this paper and are not reported in it\.

#### Recall is exact\-match verbatim recall, and only one probe of three is a paraphrase\.

Recall is exact prefix match under greedy decoding with at most eight new tokens\. Each fact carries three cloze probes;cityandjobcontinue the training sentence’s own phrasing nearly verbatim, and onlycompanyrephrases it \(“is employed at” against the trained “works as a … at”\)\. Where the per\-kind breakdown was recorded, the paraphrase probe is far weaker than the verbatim one: the unconstrained dense control decomposes as88\.7%/74\.2%/33\.9%88\.7\\%/74\.2\\%/33\.9\\%\(city/job/company\) at a pooled65\.6%65\.6\\%on seed 2, and85\.8%/41\.0%/21\.3%85\.8\\%/41\.0\\%/21\.3\\%at a pooled49\.4%49\.4\\%on seed 1\. Pooled recall therefore overstates transferable knowledge by roughly a factor of two, and the paraphrase axis is the one that degrades first\. The breakdown was instrumented late but is now archived for twelve runs, including both seeds of therehearsal\-0\.10\.1projected\-dense row of Table[1](https://arxiv.org/html/2608.20873#S5.T1), which decomposes as86\.9%/72\.0%/47\.7%86\.9\\%/72\.0\\%/47\.7\\%at a pooled68\.9%68\.9\\%and93\.6%/49\.3%/18\.8%93\.6\\%/49\.3\\%/18\.8\\%at a pooled53\.9%53\.9\\%\. The company probe is the weakest of the three in every run we have, and the spread between those two seeds on it \(47\.747\.7against18\.818\.8\) is larger than the spread on either verbatim probe\. Separately, the attribute vocabularies are small and closed \(20 cities, 16 occupations, 12 employers\), so the probes measure binding within a known answer set rather than open\-vocabulary retrieval, and exact match discards correct answers that are worded differently\. We do not claim recall is a calibrated measure of knowledge in either direction\.

#### The6\.5%6\.5\\%marginal\-guess floor\.

A model that has acquired the answer vocabulary but no name–attribute bindings, and that guesses uniformly within each attribute set, scores1/201/20,1/161/16and1/121/12, i\.e\.6\.5%6\.5\\%pooled\. Any recall at or below that figure carries no evidence of learning: this includes the under\-stepped CellFill variant at4\.0%4\.0\\%, path A’s collapse to4\.4%4\.4\\%at10410^\{4\}facts under a30\.4%30\.4\\%clip rate, and—most consequentially, because a headline was once computed against them—the early and middle MLP partitions of §[5\.10](https://arxiv.org/html/2608.20873#S5.SS10)at5\.8%5\.8\\%and6\.0%6\.0\\%\. The floor is a construction, not a measured baseline — greedy decoding is not uniform sampling, and the untrained base model in fact scores00–0\.07%0\.07\\%— so it should be read as the level below which a number cannot be distinguished from vocabulary acquisition, not as an observed chance rate\. Note also that21\.721\.7bits per fact is the entropy of the full attribute tuple including birth year and month, which no probe ever tests; all capacity accounting uses the probed11\.911\.9bits\.

#### The diagonal\-Fisher budget is not a usable numerical certificate, and its exponent is fit\-window dependent\.

Proposition[5](https://arxiv.org/html/2608.20873#Thmproposition5)is exact to second order, but both practical surrogates substitute the diagonal of an empirical Fisher estimated from≈\\approx49k tokens of WikiText\-train\. Against measured in\-cell perturbations of a 1\.7B model the uniform\-fill meanB¯F​\(ρ\)\\bar\{B\}\_\{F\}\(\\rho\)—the one the geometry sweep evaluates—underestimates drift by15\.9×15\.9\\timesatρ=0\.125\\rho\{=\}0\.125and383×383\\timesatρ=1\\rho\{=\}1, and the supremum formBF​\(ρ\)B\_\{F\}\(\\rho\), which is at least3×3\\timeslarger, by at most128×128\\times\. It must not be used to certify anything; we useρ\\rhoas a calibrated dial and report the measured curve\. Even the empirical scaling is softer than §[5\.9](https://arxiv.org/html/2608.20873#S5.SS9)may suggest: the log–log slope of measuredΔ\\DeltaNLL is3\.403\.40over all five radii,2\.492\.49overρ≥0\.25\\rho\\geq 0\.25and2\.392\.39overρ≥0\.5\\rho\\geq 0\.5on the in\-domain metric, and2\.172\.17–2\.202\.20on LAMBADA\. The quoted2\.432\.43is one member of that family\. The smallest radius carries no signal at all: its measuredΔ\\DeltaNLL is8\.4×10−58\.4\\times 10^\{\-5\}nats while the two perturbation seeds differ from each other by8\.1×10−48\.1\\times 10^\{\-4\}nats, an order of magnitude more, so including it is what drives the steepest fit\. Two perturbation seeds per radius at 1\.7B and only one at 4B, one interpolation trace per model, two models\. The independent radius ablation that reports a factor0\.22≈ρ20\.22\\approx\\rho^\{2\}isn=1n\{=\}1and compares perplexity differences rather than theΔ\\DeltaNLL of the geometry sweep\.

#### The localization experiment’s partitions do not share an anchor\.

In §[5\.10](https://arxiv.org/html/2608.20873#S5.SS10), the target filter selects both which matrices carry the residual and which matrices are replaced by their dequantized anchors in the evaluated model, so the non\-target matrices remain at original fp32 precision\. Each partition therefore has its*own*anchor — LAMBADA26\.3326\.33to27\.9227\.92across the seven runs — rather than the fully quantized28\.6728\.67anchor of Table[1](https://arxiv.org/html/2608.20873#S5.T1), and absolute values in that table are not comparable to the rest of the paper\. Priced against each partition’s matched anchor the efficiency ranking becomes gate\+up911911, attention705705, full MLP594594, early MLP456456, middle MLP455455, late MLP243243anddown\_proj210210bits per LAMBADA point\. The recommendation to write into the gate and up projections survives that correction; the claim that early and middle layers are “nearly free” does not — against their own anchors they cost\+1\.52\+1\.52and\+1\.58\+1\.58points\. Each partition is a single seed, and the early/middle difference \(5\.8%5\.8\\%versus6\.0%6\.0\\%\) is far inside the seed noise measured elsewhere\.

#### The 1\.7B/27B comparison is not a matched experiment\.

The 27B result reported in §[5\.11](https://arxiv.org/html/2608.20873#S5.SS11)matches the 1\.7B table’s 24 epochs \(its 8\-epoch companion is reported alongside it\), but differs in every other respect: no healing stage, a2%2\\%cell margin rather than1%1\\%, an fp16 rather than fp32 residual map, and bf16 rather than fp32 evaluation; its anchor is the bitsandbytes 4\-bit model because the fp32 anchor\-rebuild control was skipped for memory reasons, whereas the 1\.7B table uses fp32\-rebuilt anchors\. The anchor\-rebuild numerics control \(bnb11\.72711\.727against fp3211\.71011\.710\) exists only at 1\.7B\. “Larger models pay less at matched recall” is thus a two\-point observation across two different recipes, which is why §[5\.11](https://arxiv.org/html/2608.20873#S5.SS11)declines to call it a scaling law\.

#### The constraint is nearly non\-binding in the regime we measure\.

Over 24 epochs the unconstrained control moves only7\.537\.53–7\.767\.76million of1\.409×1091\.409\\times 10^\{9\}weights \(0\.530\.53–0\.55%0\.55\\%\) out of their cells\. That the cost of invariance is undetectable here is partly a statement about this regime rather than about the constraint: where the constraint does bind hard — path A at10410^\{4\}facts, clipping30\.4%30\.4\\%of the delta — recall collapses from the adapter’s14\.7%14\.7\\%to4\.4%4\.4\\%, below the guess floor, and is recovered only by switching to constraint\-aware training\. Nothing here establishes that invariance stays cheap at larger update budgets, longer schedules, or higher clip rates\.

#### Facts are demonstrated; skills are not\.

Every training signal in this paper is a closed\-vocabulary attribute binding on synthetic biographies\. There is no evaluation of reasoning, instruction following, code, mathematics, or any downstream capability benchmark, and none is claimed: we do not show that a skill, a behavior, or a capability can be injected into the dequantization gap\. The forgetting axis is likewise perplexity only — two corpora, one in\-domain and one out\-of\-domain — and perplexity is a proxy that is known to move differently from task accuracy\. LAMBADA is scored here as perplexity over concatenated passages, not as last\-word accuracy, so our figures are not comparable to published LAMBADA accuracies\. There is also no locality or specificity evaluation in the sense of the editing literature\[[16](https://arxiv.org/html/2608.20873#bib.bib16),[17](https://arxiv.org/html/2608.20873#bib.bib17)\]: we do not measure whether injecting one fact perturbs an unrelated one\. Comparison to the∼\\sim2\-bit/parameter full\-precision ceiling of\[[18](https://arxiv.org/html/2608.20873#bib.bib18)\]is a reference point for bit accounting, not a matched replication\.

#### “Continual” is, so far, one four\-task sequence\.

With the exception of §[5\.12](https://arxiv.org/html/2608.20873#S5.SS12), every result in this paper injects a single batch of knowledge once\. That exception is one sequence: four disjoint 500\-fact tasks, one model, one seed, one rank, one consolidation variant, and no unconstrained control sequence against which to price the forgetting\. The decay constantβ=0\.825\\beta=0\.825is measured from three ratios inside that single sequence, so it demonstrates that the law is measurable rather than fixing a value that transfers, and nothing here tests a sequence longer than four updates or a model other than Qwen3\-1\.7B\.

#### What the guarantee does not say\.

Proposition[1](https://arxiv.org/html/2608.20873#Thmproposition1)protects the artifact, not the update\. Bitwise invariance implies nothing about whether the served modelW^\+δ\\hat\{W\}\+\\deltais safe, accurate, or aligned: a harmful residual is exactly as harmful as a harmful fine\-tune, merely revocable and bounded in per\-weight magnitude\. The box bounds each coordinate, not the direction of the update and not the resulting function — the random\-perturbation sweep shows LAMBADA moving28\.5→34\.128\.5\\to 34\.1with every code intact\. Certification does not transfer: the served model is a different function from the certified one, and §[5](https://arxiv.org/html/2608.20873#S5)measures its drift precisely so that this is not confused\. Three further boundaries of the contract deserve stating explicitly\. \(i\) Consolidation*does*change the artifact, by design: renewing capacity requires re\-quantizing into a new frozen\(a,s\)\(a,s\), which is an explicit major version and which invalidates exactly the certification the minor\-version regime preserves\. Capacity within one grid is a buffer, not unbounded memory\. \(ii\) Invariance is defined against the shipped\(a,s\)\(a,s\)and holds only for consumers who verify under those scales; re\-running a quantizer can reassign untouched weights, as the two\-weight counterexample of Appendix[A](https://arxiv.org/html/2608.20873#A1)shows\. This is a requirement on the deployment convention, not a property of the residual\. \(iii\) The guarantee is bounded by storage precision and margin: bounds are shrunk by1%1\\%of cell width per side \(2%2\\%at 27B\), forfeiting that fraction of writable range for tie safety, and fp32/fp16 storage is safe at1%1\\%while bf16 requires a margin of at least5%5\\%\. All archived invariance checks are on fp32 \(or fp16 at 27B\) residual maps; a bf16 end\-to\-end serialization has not been verified\.

#### Verification is our own integer arithmetic, not a vendor round\-trip\.

“Bit\-identical” means that codes re\-derived by our NF4 code\-assignment routine under the frozen scales match the shipped codes exactly, over every constrained weight\. That is the correct definition \(Appendix[A](https://arxiv.org/html/2608.20873#A1)\) and it is asserted rather than sampled, but it is checked by our reimplementation\. An end\-to\-end byte\-level reload through stock bitsandbytes has not been run\. Relatedly, the served object in all experiments is a dense fp32/fp16 weight matrix, not a 4\-bit one: nothing here demonstrates that an in\-cell update can be*served*at 4\-bit cost\. The refinement code of Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)is designed for exactly that and its nesting and error bounds are machine\-verified\. §[6](https://arxiv.org/html/2608.20873#S6)measures the quality–size curve, but on a single configuration in one archived run, evaluated by reconstructing dense weights rather than by a 4\-bit reload, and with no BitDelta baseline\.

#### Evaluation harness truncation\.

For wall\-clock reasons, in\-domain perplexity is computed over at most40×1024=40,96040\\times 1024=40\{,\}960tokens of the WikiText\-2 test split and cross\-domain perplexity over the first 400 LAMBADA test passages, rather than over the full corpora that published perplexities normally use\. Training runs under bf16 autocast while final 1\.7B evaluation rebuilds the model in fp32, so training\-time and evaluation\-time numerics differ\. Dataset loads pin name, configuration and split but not a revision hash\.

## 10Reproducibility

#### Stack and commands\.

The core package is pure PyTorch \(torch≥\\,\\geq\\,2\.1\) and runs on CPU; the experiment layer additionally requiresbitsandbytes≥\\,\\geq\\,0\.43,transformers≥\\,\\geq\\,4\.44,peft≥\\,\\geq\\,0\.12,datasetsandaccelerateon CUDA\.server/setup\.shprovisions a Python 3\.12 environment withuv, is idempotent, and prints the resolved bitsandbytes / transformers / peft versions together with the GPU name, driver version and VRAM before running the test suite; those lines are captured in each run’s log\. The 1\.7B ladder ran on a single 48 GB consumer GPU and the 27B model on an A100\. Each experiment is a single command —exp0\_clip\_rate\.pyfor paths A/A\+/B and the unconstrained control,exp5\_qil\.pyfor CellFill and the localization partitions,exp\_geom\.pyfor the geometry sweep,exp\_seq\.pyfor the sequential lifecycle — and the exact invocations, including the ones behind every archived file, are the queue scripts inexperiments/\. Wall clock is recorded per run:1212–2020minutes for the 1\.7B frontier arms,3030minutes at 96 epochs,4646minutes for 27B clip\-merge at 8 epochs,6\.76\.7minutes for the geometry sweep\. Each run writes one JSON whoseconfigblock is the complete serialized argument namespace, so any archived number can be traced to the command that produced it; the tables in this paper are generated from those files byscripts/make\_tables\.pyrather than typed, and each cell carries the number of seeds it was averaged over\.

#### What is deterministic\.

Seeding is explicit and covers the data as well as the optimizer\. A seed fixestorch\.manual\_seed, the synthetic fact corpus \(the same seed yields the identical set of names, cities, occupations, employers and birth dates, with name uniqueness enforced and machine\-checked\), the rehearsal draw from the WikiText train split, the probe subsample when one is used, and the shuffle order of the healing phase\. Two arms run at the same seed therefore see an identical fact set and an identical rehearsal stream, which is what makes the paired comparison of §[5\.3](https://arxiv.org/html/2608.20873#S5.SS3)a genuine single\-intervention experiment\. The invariance layer is exactly reproducible because it is integer arithmetic: codes are re\-derived under the frozen\(codes,absmax,blocksize\)\(\\text\{codes\},\\text\{absmax\},\\text\{blocksize\}\)extracted once from eachLinear4bitand compared as integers\. No floating\-point comparison of dequantized values occurs anywhere in the verification path, precisely because such a comparison would not be reproducible across devices\. The theory layer reproduces on CPU in about a second: 25 unit tests covering NF4 code assignment, anchors as interior fixed points, in\-cell invariance atρ∈\{0\.25,0\.5,0\.9\}\\rho\\in\\\{0\.25,0\.5,0\.9\\\}, detection of boundary crossings, the frozen\-scale counterexample of Appendix[A](https://arxiv.org/html/2608.20873#A1), bf16 storage at a wide margin, projection of oversized deltas, sequential folding, consolidation renewing room, the Matryoshka nesting and error bounds of Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)atk∈\{1,2,4,8\}k\\in\\\{1,2,4,8\\\}, and the fact/probe invariants including probe\-kind balance and the declared chance levels\.

#### What is not\.

GPU results are not bit\-reproducible\. Training runs undertorch\.autocastin bf16, deterministic algorithm enforcement and cuBLAS workspace pinning are not enabled, and kernel reduction order varies; the same seed on the same GPU can differ slightly and different GPUs will differ more\. Recall and perplexity should be expected to reproduce to about the between\-seed spread reported in Table[1](https://arxiv.org/html/2608.20873#S5.T1), not to the digit\. Dataset loads pin the dataset name, configuration and split but not a revision hash\.

#### Frozen\-artifact convention\.

The artifact is the pair\(a,s\)\(a,s\)— integer codes and per\-group absmax scales at blocksize 64 with double quantization — extracted once from the loaded 4\-bit model and never recomputed thereafter\. All per\-weight bounds derive from that frozen pair, with the outermost cells capped so every code addresses a finite interval, and with each interval shrunk by a margin of1%1\\%of its width per side \(2%2\\%for the 27B run\) so stored values remain strictly interior to their decision regions\. Residual maps are archived in fp32 \(fp16 for 27B\)\. The released code implements code assignment under frozen scales directly, because quantization libraries do not expose that operation and recomputing absmax silently re\-bins whole blocks \(Appendix[A](https://arxiv.org/html/2608.20873#A1)\)\.

#### Invariance is asserted, not spot\-checked\.

On every merge, the per\-layer check runs over*all*constrained weights of*all*quantized matrices — 196 matrices and1\.409×1091\.409\\times 10^\{9\}weights at 1\.7B, 496 matrices and2\.435×10102\.435\\times 10^\{10\}weights at 27B — and raises on the first layer with any mismatched code, so a violated run produces no result file at all\. The projected\-healing path re\-runs the same exhaustive check over every trainable weight after training and raises if any weight moved, and CellFill’s materialization performs the check when it converts the bounded fill into fp32 weights\. No sampling, thresholding or tolerance is involved\. Consequently every archived constrained run recordsinvariance\_violations=0\\,=0, and the deliberately unconstrained control records the exact escape count instead —7,534,3467\{,\}534\{,\}346,7,619,1427\{,\}619\{,\}142and7,759,1957\{,\}759\{,\}195weights for seeds00,11and22— which is what demonstrates that the constraint is non\-vacuous rather than merely satisfied\.

#### Data\.

In\-domain perplexity uses the WikiText\-2 raw test split; rehearsal is drawn from the corresponding train split, switching automatically to WikiText\-103 raw train once the pool exceeds 4,000 paragraphs so that snippets stay fresh across epochs; cross\-domain perplexity uses the LAMBADA \(OpenAI\) English test split\. The fact corpus is generated in\-process and requires no download\. Code, unit tests and the complete set of archived result files is at[https://github\.com/sumsliu/cellfill](https://github.com/sumsliu/cellfill)\.

## Appendix AWhy invariance must be defined against frozen scales

Proposition[1](https://arxiv.org/html/2608.20873#Thmproposition1)is stated for a*frozen*artifact\(a,s\)\(a,s\)\. It is tempting to define invariance operationally instead—“re\-run the quantizer and check the codes”—but that definition is false, because group scales are data\-dependent\. The following two\-weight counterexample is executed as a unit test in the released code\.

Take one block underabsmax\\mathrm\{absmax\}scaling with NF4 levels, where the decision boundary between the top two codes sits at\(L14\+L15\)/2=0\.8615\(L\_\{14\}\+L\_\{15\}\)/2=0\.8615in normalized units\. Letw1=1\.00w\_\{1\}=1\.00\(the block maximum, sos=1\.00s=1\.00\) andw2=0\.80w\_\{2\}=0\.80; thenw1w\_\{1\}takes code 15 andw2w\_\{2\}, at0\.80<0\.86150\.80<0\.8615, takes code 14\.

Now movew1w\_\{1\}to0\.870\.87, an update that stays inside its own cell and that Proposition[1](https://arxiv.org/html/2608.20873#Thmproposition1)therefore permits\. Under the frozen scale both codes are unchanged, as the proposition promises\. But re\-running the quantizer setss′=0\.87s^\{\\prime\}=0\.87, andw2/s′=0\.80/0\.87=0\.920\>0\.8615w\_\{2\}/s^\{\\prime\}=0\.80/0\.87=0\.920\>0\.8615: the untouched weightw2w\_\{2\}silently migrates from code 14 to code 15\. The artifact changed without any weight leaving its cell\.

Hence the artifact must ship its scales, and verification must assign codes under those scales rather than recompute them\. Every invariance check in this paper does so, in the integer domain; floating\-point comparison of dequantized values is not used, since it is not reproducible across devices\.

## Appendix BProof sketches

This appendix proves Propositions[2](https://arxiv.org/html/2608.20873#Thmproposition2)–[5](https://arxiv.org/html/2608.20873#Thmproposition5)against the objects the code actually manipulates\. Proposition[1](https://arxiv.org/html/2608.20873#Thmproposition1)and the frozen\-scale subtlety are treated separately in Appendix[A](https://arxiv.org/html/2608.20873#A1)\.

#### Notation, fixed to the implementation\.

LetL⁡\[0\]<⋯<L⁡\[15\]L\[0\]<\\dots<L\[15\]be the NF4 level table andM⁡\[j\]=\(L⁡\[j\]\+L⁡\[j\+1\]\)/2M\[j\]=\(L\[j\]\{\+\}L\[j\{\+\}1\]\)/2,j=0,…,14j=0,\\dots,14, the round\-to\-nearest \(RTN\) decision boundaries in normalized units\. Under a*frozen*block scales\>0s\>0, code assignment isai=\#⁡\{j:M⁡\[j\]<wi/s\}a\_\{i\}=\\\#\\\{j:M\[j\]<w\_\{i\}/s\\\}; equivalently the decision region of codejjiss⋅\(M⁡\[j−1\],M⁡\[j\]\]s\\cdot\(M\[j\{\-\}1\],M\[j\]\]withM⁡\[−1\]=−∞M\[\-1\]=\-\\infty,M⁡\[15\]=\+∞M\[15\]=\+\\infty\(ties fall to the lower code, matchingtorch\.bucketizewithright=False\)\. The two outer regions are half\-infinite, so the implementation*caps*them by mirroring the inner half\-gap about the level,lo⁡\[0\]=L⁡\[0\]−\(M⁡\[0\]−L⁡\[0\]\)\\mathrm\{lo\}\[0\]=L\[0\]\-\(M\[0\]\-L\[0\]\)andhi⁡\[15\]=L⁡\[15\]\+\(L⁡\[15\]−M⁡\[14\]\)\\mathrm\{hi\}\[15\]=L\[15\]\+\(L\[15\]\-M\[14\]\), which is required for the refinement code \(a code must address a finite interval\)\. Finally a relative marginη\\eta\(η=0\.01\\eta=0\.01throughout\) shrinks each interval byη\\etaof its width per side\. WriteWiW\_\{i\}for the capped cell width of weightii, and

Ci=\[ℓi,ui\],ℓi=loi\+η​Wi,ui=hii−η​Wi,Δi=ui−ℓi=\(1−2​η\)​Wi,C\_\{i\}=\[\\ell\_\{i\},u\_\{i\}\],\\qquad\\ell\_\{i\}=\\mathrm\{lo\}\_\{i\}\+\\eta W\_\{i\},\\quad u\_\{i\}=\\mathrm\{hi\}\_\{i\}\-\\eta W\_\{i\},\\qquad\\Delta\_\{i\}=u\_\{i\}\-\\ell\_\{i\}=\(1\-2\\eta\)W\_\{i\},𝒞=∏i=1NCi\\mathcal\{C\}=\\prod\_\{i=1\}^\{N\}C\_\{i\}, andw^i=sg⁡\(i\)​L​\[ai\]\\hat\{w\}\_\{i\}=s\_\{g\(i\)\}L\[a\_\{i\}\]for the anchor\. Because the anchor may sit off\-center in a non\-uniform level table, the*symmetric*radius actually available at weightiiis

roomi=min⁡\(w^i−ℓi,ui−w^i\)≤Δi/2,\\mathrm\{room\}\_\{i\}=\\min\(\\hat\{w\}\_\{i\}\-\\ell\_\{i\},\\;u\_\{i\}\-\\hat\{w\}\_\{i\}\)\\;\\leq\\;\\Delta\_\{i\}/2,with equality only in the two capped outer cells\. This is the quantity the code carries \(bin\_bounds, andMMin CellFill\); statements below that useΔi/2\\Delta\_\{i\}/2are the loose form and are marked as such\.

#### Fact A\.1 \(the margined cell is compactly interior\)\.

For everyii,CiC\_\{i\}is a nonempty compact interval contained in the*open*decision region of codeaia\_\{i\}, at distance at leastη​Wi\\eta W\_\{i\}from either RTN boundary\. Indeedℓi\>loi≥s​M​\[ai−1\]\\ell\_\{i\}\>\\mathrm\{lo\}\_\{i\}\\geq s\\,M\[a\_\{i\}\-1\]andui<hii≤s​M​\[ai\]u\_\{i\}<\\mathrm\{hi\}\_\{i\}\\leq s\\,M\[a\_\{i\}\]wheneverη\>0\\eta\>0andWi\>0W\_\{i\}\>0\. Two consequences are used repeatedly: \(a\)𝒞\\mathcal\{C\}is a nonempty compact convex box, so Euclidean projection onto it exists and is unique; \(b\) no point of𝒞\\mathcal\{C\}lies on a tie, so the tie\-breaking convention never carries information, and a storage rounding of magnitude<η​Wi<\\eta W\_\{i\}cannot change a code—which is why the safe\-dtype question of Appendix[A](https://arxiv.org/html/2608.20873#A1)is a statement aboutη\\etaalone\.

### B\.1Proposition[2](https://arxiv.org/html/2608.20873#Thmproposition2)\(projection optimality\)

#### \(i\) Clipping is the Euclidean projection\.

Lety∈ℝNy\\in\\mathbb\{R\}^\{N\}be any proposed weight vector \(in path A,y=W^\+dy=\\hat\{W\}\+dwithddthe materialized LoRA update; we write the proposed update asdd, reservingΔi\\Delta\_\{i\}for the cell width\)\. Since‖y−z‖22=∑i\(yi−zi\)2\\\|y\-z\\\|\_\{2\}^\{2\}=\\sum\_\{i\}\(y\_\{i\}\-z\_\{i\}\)^\{2\}separates over coordinates and𝒞\\mathcal\{C\}is a product set, minimizing over𝒞\\mathcal\{C\}decouples intoNNone\-dimensional problemsminzi∈\[ℓi,ui\]⁡\(yi−zi\)2\\min\_\{z\_\{i\}\\in\[\\ell\_\{i\},u\_\{i\}\]\}\(y\_\{i\}\-z\_\{i\}\)^\{2\}\. Each is a strictly convex function on a nonempty compact interval, so it has the unique minimizerzi=median⁡\(ℓi,yi,ui\)=clamp⁡\(yi,ℓi,ui\)z\_\{i\}=\\operatorname\{median\}\(\\ell\_\{i\},y\_\{i\},u\_\{i\}\)=\\operatorname\{clamp\}\(y\_\{i\},\\ell\_\{i\},u\_\{i\}\): ifyi∈\[ℓi,ui\]y\_\{i\}\\in\[\\ell\_\{i\},u\_\{i\}\]the objective is00; otherwise the objective is strictly monotone on the interval and is minimized at the nearer endpoint\. HenceP𝒞​\(y\)i=clamp⁡\(yi,ℓi,ui\)P\_\{\\mathcal\{C\}\}\(y\)\_\{i\}=\\operatorname\{clamp\}\(y\_\{i\},\\ell\_\{i\},u\_\{i\}\), which is exactlyclip\_merge\. By Fact A\.1\(a\) this projection is single\-valued and11\-Lipschitz \(firmly nonexpansive\), as for any projection onto a closed convex set\.

#### \(ii\) It is the proximal operator of the box indicator\.

Letι𝒞\\iota\_\{\\mathcal\{C\}\}be the indicator of𝒞\\mathcal\{C\}\(00on𝒞\\mathcal\{C\},\+∞\+\\inftyoff it\); it is proper, closed and convex because𝒞\\mathcal\{C\}is nonempty, closed and convex\. For anyγ\>0\\gamma\>0,

proxγ​ι𝒞⁡\(y\)=arg⁡minz​\{ι𝒞​\(z\)\+12​γ​‖z−y‖22\}=arg⁡minz∈𝒞​‖z−y‖22=P𝒞​\(y\),\\operatorname\{prox\}\_\{\\gamma\\iota\_\{\\mathcal\{C\}\}\}\(y\)=\\arg\\min\_\{z\}\\Big\\\{\\iota\_\{\\mathcal\{C\}\}\(z\)\+\\tfrac\{1\}\{2\\gamma\}\\\|z\-y\\\|\_\{2\}^\{2\}\\Big\\\}=\\arg\\min\_\{z\\in\\mathcal\{C\}\}\\\|z\-y\\\|\_\{2\}^\{2\}=P\_\{\\mathcal\{C\}\}\(y\),independently ofγ\\gamma\. Therefore the projected\-gradient iteration used in paths A\+/B,wt\+1=P𝒞\(wt−γ∇f\(wt\)\)=proxγ​ι𝒞\(wt−γ∇f\(wt\)\)w^\{t\+1\}=P\_\{\\mathcal\{C\}\}\\\!\\left\(w^\{t\}\-\\gamma\\nabla f\(w^\{t\}\)\\right\)=\\operatorname\{prox\}\_\{\\gamma\\iota\_\{\\mathcal\{C\}\}\}\\\!\\left\(w^\{t\}\-\\gamma\\nabla f\(w^\{t\}\)\\right\), is literally forward–backward splitting onf\+ι𝒞f\+\\iota\_\{\\mathcal\{C\}\}\. This is the precise sense in which “the optimizer sees the walls” \(§[3](https://arxiv.org/html/2608.20873#S3)\)\.

#### \(iii\) Coordinatewise shrinkage\.

Writeκ=P𝒞​\(W^\+d\)−W^\\kappa=P\_\{\\mathcal\{C\}\}\(\\hat\{W\}\+d\)\-\\hat\{W\}for the retained update ande=\(W^\+d\)−P𝒞​\(W^\+d\)e=\(\\hat\{W\}\+d\)\-P\_\{\\mathcal\{C\}\}\(\\hat\{W\}\+d\)for the clipped\-away part, sod=κ\+ed=\\kappa\+e\. Sincew^i∈Ci\\hat\{w\}\_\{i\}\\in C\_\{i\}, fordi\>0d\_\{i\}\>0we haveκi=min⁡\(w^i\+di,ui\)−w^i∈\[0,di\]\\kappa\_\{i\}=\\min\(\\hat\{w\}\_\{i\}\+d\_\{i\},u\_\{i\}\)\-\\hat\{w\}\_\{i\}\\in\[0,d\_\{i\}\], and symmetrically fordi<0d\_\{i\}<0; hence

signκi=signei=signdi,\|κi\|≤\|di\|for everyi\.\\operatorname\{sign\}\\kappa\_\{i\}=\\operatorname\{sign\}e\_\{i\}=\\operatorname\{sign\}d\_\{i\},\\qquad\|\\kappa\_\{i\}\|\\leq\|d\_\{i\}\|\\quad\\text\{for every \}i\.So projection shrinks the update coordinatewise toward the anchor, and only along coordinates that tried to leave their cell\. Two computable corollaries:‖κ‖≤‖d‖\\\|\\kappa\\\|\\leq\\\|d\\\|\(also immediate from nonexpansiveness, sinceP𝒞​\(W^\)=W^P\_\{\\mathcal\{C\}\}\(\\hat\{W\}\)=\\hat\{W\}\), and, because⟨κ,e⟩≥0\\langle\\kappa,e\\rangle\\geq 0,

‖e‖2≤‖d‖2−‖κ‖2=‖d‖2​\(1−ϕ2\),ϕ:=‖κ‖/‖d‖,\\\|e\\\|^\{2\}\\;\\leq\\;\\\|d\\\|^\{2\}\-\\\|\\kappa\\\|^\{2\}\\;=\\;\\\|d\\\|^\{2\}\\left\(1\-\\phi^\{2\}\\right\),\\qquad\\phi:=\\\|\\kappa\\\|/\\\|d\\\|,whereϕ\\phiis thenorm\_kept\_fraclogged at every merge\. This is the formal content of the “per\-weight trust region” reading of §[5\.6](https://arxiv.org/html/2608.20873#S5.SS6): it shows the projected update is never larger than the proposal and is shrunk exactly where the proposal was most extreme\. It does*not*predict the sign of the resulting perplexity change; that the projection helps cross\-domain perplexity is an empirical finding \(§[5\.6](https://arxiv.org/html/2608.20873#S5.SS6)\), not a corollary of the geometry\.

#### \(iv\) What this does*not*give\.

Optimality here is in*weight space*only\. Post\-hoc projection solvesminw∈𝒞⁡‖w−y‖\\min\_\{w\\in\\mathcal\{C\}\}\\\|w\-y\\\|exactly; it does not solveminw∈𝒞⁡f⁡\(w\)\\min\_\{w\\in\\mathcal\{C\}\}f\(w\), and the two can differ already for convexff\. TakeN=2N=2,𝒞=\[−1,1\]2\\mathcal\{C\}=\[\-1,1\]^\{2\},f⁡\(w\)=\(w1\+w2−4\)2f\(w\)=\(w\_\{1\}\+w\_\{2\}\-4\)^\{2\}\. The unconstrained minimizers form the linew1\+w2=4w\_\{1\}\+w\_\{2\}=4; the constrained minimum is44, attained at\(1,1\)\(1,1\)\. Projecting the particular unconstrained minimizer\(4,0\)\(4,0\)gives\(1,0\)\(1,0\)withf=9\>4f=9\>4\. Note also that projecting the unconstrained minimizer\(2,2\)\(2,2\)gives\(1,1\)\(1,1\), which*is*optimal: the map “P𝒞​\(arg⁡min⁡f\)P\_\{\\mathcal\{C\}\}\(\\arg\\min f\)” is not even well defined as a function offf, since it depends on which unconstrained solution the optimizer reached\. This is the structural reason paths A and A\+ can differ at all, and it is not a nonconvexity artifact\.

For nonconvexff\(the actual case\) the honest statement about projected descent is: if∇f\\nabla fisLL\-Lipschitz andγ∈\(0,1/L\]\\gamma\\in\(0,1/L\], projected gradient descent is a descent method whose gradient mappingGγ\(w\)=γ−1\(w−P𝒞\(w−γ∇f\(w\)\)\)G\_\{\\gamma\}\(w\)=\\gamma^\{\-1\}\\big\(w\-P\_\{\\mathcal\{C\}\}\(w\-\\gamma\\nabla f\(w\)\)\\big\)obeysmint<T⁡‖Gγ​\(wt\)‖2≤2​\(f⁡\(w0\)−inf𝒞f\)/\(γ​T\)\\min\_\{t<T\}\\\|G\_\{\\gamma\}\(w^\{t\}\)\\\|^\{2\}\\leq 2\\big\(f\(w^\{0\}\)\-\\inf\_\{\\mathcal\{C\}\}f\\big\)/\(\\gamma T\), so accumulation points are*stationary*for the constrained problem, i\.e\. satisfy−∇f​\(w\)∈N𝒞​\(w\)\-\\nabla f\(w\)\\in N\_\{\\mathcal\{C\}\}\(w\)\. Nothing here implies convergence to a global constrained optimum, and we do not claim it\. Two further gaps between that theorem and our runs should be stated plainly: the gradients are stochastic, and the update is projected AdamW \(clamp after every optimizer step\), which is a projection in the Euclidean metric applied to a step taken in a preconditioned metric—so it is not the proximal step of the classical analysis, and none of the above rates apply to it verbatim\. Paths A\+/B are therefore justified by measurement \(§[5](https://arxiv.org/html/2608.20873#S5)\), not by a convergence guarantee\.

#### \(v\) The clip rate as a diagnostic, made quantitative\.

If no coordinate is clipped theny∈𝒞y\\in\\mathcal\{C\}andP𝒞​\(y\)=yP\_\{\\mathcal\{C\}\}\(y\)=y; in particular, ifyyis a global unconstrained minimizer that happens to lie in𝒞\\mathcal\{C\}, it is also a global constrained minimizer, and post\-hoc projection is exactly free\. More generally, with∇f\\nabla fLL\-Lipschitz andeethe clipped\-away vector of \(iii\),

f⁡\(P𝒞​\(y\)\)−f⁡\(y\)≤−⟨∇f​\(y\),e⟩\+L2​‖e‖2≤‖∇f​\(y\)‖​‖e‖\+L2​‖e‖2,f\\big\(P\_\{\\mathcal\{C\}\}\(y\)\\big\)\-f\(y\)\\;\\leq\\;\-\\langle\\nabla f\(y\),e\\rangle\+\\tfrac\{L\}\{2\}\\\|e\\\|^\{2\}\\;\\leq\\;\\\|\\nabla f\(y\)\\\|\\,\\\|e\\\|\+\\tfrac\{L\}\{2\}\\\|e\\\|^\{2\},so the loss cost of projection is controlled by the clipped\-away mass, which is measured at every merge \(clipped\_frac,max\_excess\_halfwidths, and‖e‖≤‖d‖​1−ϕ2\\\|e\\\|\\leq\\\|d\\\|\\sqrt\{1\-\\phi^\{2\}\}\)\. SinceLLis not measured, this is a scaling argument that explains the observed pattern—projection free at0\.14%0\.14\\%clip, costly at33–5%5\\%, catastrophic at30\.4%30\.4\\%\(§[5](https://arxiv.org/html/2608.20873#S5)\)—and not a numerical certificate\.

### B\.2Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)\(nested refinement code\)

Fix a weight with margined cell\[ℓ,u\]\[\\ell,u\]of widthΔ=u−ℓ\>0\\Delta=u\-\\ell\>0, and letx=clamp⁡\(\(w−ℓ\)/Δ,0,1\)∈\[0,1\]x=\\operatorname\{clamp\}\\\!\\big\(\(w\-\\ell\)/\\Delta,\\,0,\\,1\\big\)\\in\[0,1\]be its clamped in\-cell position\. The codec is

r\(k\)=min⁡\(⌊2k​x⌋,2k−1\)∈\{0,…,2k−1\},w^\(k\)=ℓ\+\(r\(k\)\+12\)​Δ/2k,r^\{\(k\)\}=\\min\\\!\\big\(\\lfloor 2^\{k\}x\\rfloor,\\;2^\{k\}\-1\\big\)\\in\\\{0,\\dots,2^\{k\}\-1\\\},\\qquad\\hat\{w\}^\{\(k\)\}=\\ell\+\\big\(r^\{\(k\)\}\+\\tfrac\{1\}\{2\}\\big\)\\Delta/2^\{k\},for1≤k≤81\\leq k\\leq 8\(pack\_residual/unpack\_residual\)\.

#### \(i\) Nesting\.

We use the floor\-of\-floor identity: for realy≥0y\\geq 0and integerm≥1m\\geq 1,⌊⌊y⌋/m⌋=⌊y/m⌋\\lfloor\\lfloor y\\rfloor/m\\rfloor=\\lfloor y/m\\rfloor\. Proof: putn=⌊y⌋n=\\lfloor y\\rfloorandq=⌊n/m⌋q=\\lfloor n/m\\rfloor, soq​m≤n≤\(q\+1\)​m−1qm\\leq n\\leq\(q\+1\)m\-1\. Theny≥n≥q​my\\geq n\\geq qmgives⌊y/m⌋≥q\\lfloor y/m\\rfloor\\geq q, andy<n\+1≤\(q\+1\)​my<n\+1\\leq\(q\+1\)mgives⌊y/m⌋≤q\\lfloor y/m\\rfloor\\leq q\. Applying it withy=2k\+j​xy=2^\{k\+j\}xandm=2jm=2^\{j\}, and using that≫j\\gg jis exactly⌊⋅/2j⌋\\lfloor\\cdot/2^\{j\}\\rflooron nonnegative integers,

⌊2k\+j​x⌋≫j=⌊⌊2k\+j​x⌋/2j⌋=⌊2k​x⌋\.\\lfloor 2^\{k\+j\}x\\rfloor\\gg j=\\big\\lfloor\\lfloor 2^\{k\+j\}x\\rfloor/2^\{j\}\\big\\rfloor=\\lfloor 2^\{k\}x\\rfloor\.This settlesx∈\[0,1\)x\\in\[0,1\)\. The boundary casex=1x=1—which occurs exactly whenw≥uw\\geq u, i\.e\. when the encoder saturates—must be checked separately because there the floor overflows and themin\\minclamp fires: thenr\(k\+j\)=2k\+j−1r^\{\(k\+j\)\}=2^\{k\+j\}\-1, whose binary expansion is all ones, sor\(k\+j\)≫j=2k−1=r\(k\)r^\{\(k\+j\)\}\\gg j=2^\{k\}\-1=r^\{\(k\)\}, the clamp having fired at depthkkas well\. Hencer\(k\)=r\(k\+j\)≫jr^\{\(k\)\}=r^\{\(k\+j\)\}\\gg jholds on all of\[0,1\]\[0,1\]: the saturating clamp is compatible with truncation precisely because it saturates to an all\-ones code\. Akk\-bit stream is therefore a prefix of every deeper one, and dropping trailing bits is a valid decode at the shallower depth—down tok=0k=0, which returns the anchor itself and the released artifact unchanged\.

#### \(ii\) Error bound\.

Forx∈\[0,1\)x\\in\[0,1\)the clamp is inactive andr=⌊2k​x⌋r=\\lfloor 2^\{k\}x\\rfloorsatisfiesr≤2k​x<r\+1r\\leq 2^\{k\}x<r\+1, i\.e\.xxlies in the sub\-cell\[r​2−k,\(r\+1\)​2−k\)\[r2^\{\-k\},\(r\{\+\}1\)2^\{\-k\}\)of length2−k2^\{\-k\}, whose center is the reconstructed positionx^=\(r\+12\)​2−k\\hat\{x\}=\(r\+\\tfrac\{1\}\{2\}\)2^\{\-k\}\. Hence\|x−x^\|≤2−\(k\+1\)\|x\-\\hat\{x\}\|\\leq 2^\{\-\(k\+1\)\}and

\|w−w^\(k\)\|=Δ​\|x−x^\|≤Δ/2k\+1=\(1−2​η\)​W/2k\+1≤W/2k\+1,\|w\-\\hat\{w\}^\{\(k\)\}\|=\\Delta\\,\|x\-\\hat\{x\}\|\\;\\leq\\;\\Delta/2^\{k\+1\}\\;=\\;\(1\-2\\eta\)\\,W/2^\{k\+1\}\\;\\leq\\;W/2^\{k\+1\},which is the claimed bound, stated in the raw cell widthWW; the margin only tightens it\. The bound is attained \(atx=r​2−kx=r2^\{\-k\}\), so it cannot be improved\. If the encoder saturated \(w∉\[ℓ,u\]w\\notin\[\\ell,u\]\), the same computation bounds the distance to the*clamped*position, so the total error isΔ/2k\+1\+dist⁡\(w,\[ℓ,u\]\)\\Delta/2^\{k\+1\}\+\\operatorname\{dist\}\(w,\[\\ell,u\]\); in the training paths this second term is zero by construction, since every shipped weight is already the output of a projection or of a boundedtanh\\tanhfill\.

#### \(iii\) Strict interiority, at every depth\.

Since0≤r≤2k−10\\leq r\\leq 2^\{k\}\-1, the reconstructed position satisfiesx^∈\[2−\(k\+1\),1−2−\(k\+1\)\]⊂\(0,1\)\\hat\{x\}\\in\[2^\{\-\(k\+1\)\},\\,1\-2^\{\-\(k\+1\)\}\]\\subset\(0,1\), so

ℓ\+Δ/2k\+1≤w^\(k\)≤u−Δ/2k\+1\.\\ell\+\\Delta/2^\{k\+1\}\\;\\leq\\;\\hat\{w\}^\{\(k\)\}\\;\\leq\\;u\-\\Delta/2^\{k\+1\}\.Combining with Fact A\.1, the reconstruction clears either RTN decision boundary by at leastη​W\+\(1−2​η\)​W/2k\+1\>η​W\\eta W\+\(1\-2\\eta\)W/2^\{k\+1\}\>\\eta W, a clearance that is bounded below*uniformly inkk*\. By Proposition[1](https://arxiv.org/html/2608.20873#Thmproposition1), re\-quantization under the frozen scales returns the original codeaia\_\{i\}, for every truncation depth and for every weight\. Since the 4\-bit codes are never written by the codec, the 4\-bit prefix of the shipped file is bit\-identical to the release by construction rather than by verification\.

#### Machine\-verified content\.

tests/test\_codec\.pychecks, on randomly quantized blocks: nesting \(r\(4\)=r\(8\)≫4r^\{\(4\)\}=r^\{\(8\)\}\\\!\\gg\\\!4,r\(2\)=r\(8\)≫6r^\{\(2\)\}=r^\{\(8\)\}\\\!\\gg\\\!6,r\(2\)=r\(4\)≫2r^\{\(2\)\}=r^\{\(4\)\}\\\!\\gg\\\!2\); the error bound fork∈\{1,2,4,8\}k\\in\\\{1,2,4,8\\\}; strict interiority fork∈\{1,4,8\}k\\in\\\{1,4,8\\\}; monotone decrease of mean error inkk; and integer\-domain invariance under frozen scales after every round trip\.

### B\.3Proposition[4](https://arxiv.org/html/2608.20873#Thmproposition4)\(capacity bound\)

#### The channel\.

LetFFbe the random fact corpus, drawn by the synthetic generator of §[4](https://arxiv.org/html/2608.20873#S4)by sampling each entity’s attributes independently and uniformly from finite lists\. LetRRbe the publisher’s independent randomness \(initialization, data order, rehearsal draw\)\. A training procedure produces a messagem=T⁡\(F,R\)m=T\(F,R\), and the subscriber, holding the public artifactW^\\hat\{W\}, formsW′=f⁡\(W^,m\)W^\{\\prime\}=f\(\\hat\{W\},m\)\. BecauseW^\\hat\{W\}is fixed beforeFFexists andW′W^\{\\prime\}depends onFFonly throughmm, the variables form a Markov chainF→m→W′F\\to m\\to W^\{\\prime\}\.

#### Message length\.

The message is \(mask,r\(k\)r^\{\(k\)\}\): a setS⊆\{1,…,N\}S\\subseteq\\\{1,\\dots,N\\\}ofNt=\|S\|N\_\{t\}=\|S\|written coordinates andkkrefinement bits each\. Two regimes must be distinguished, and this is the first place a reviewer should push\.*\(a\) Declared mask\.*IfSSis fixed and public—announced in the release contract, as in every experiment here: paths A/A\+/B write all quantized linear weights, and the localization study of §[5\.10](https://arxiv.org/html/2608.20873#S5.SS10)uses masks named in advance by module and depth—then the receiver already knowsSSand the index term is*not*paid:\|m\|=k​Nt\|m\|=kN\_\{t\}\.*\(b\) Data\-dependent mask\.*IfSSis chosen as a function of the run \(hence ofFF\), it must be transmitted; enumerative coding of anNtN\_\{t\}\-subset costs⌈log2⁡\(NNt\)⌉\\lceil\\log\_\{2\}\\binom\{N\}\{N\_\{t\}\}\\rceilbits, plusO⁡\(log⁡N\)O\(\\log N\)to conveyNtN\_\{t\}if it is not fixed\. Then\|m\|≤k​Nt\+log2⁡\(NNt\)\+O⁡\(log⁡N\)\|m\|\\leq kN\_\{t\}\+\\log\_\{2\}\\binom\{N\}\{N\_\{t\}\}\+O\(\\log N\), which is the form stated in Proposition[4](https://arxiv.org/html/2608.20873#Thmproposition4)\. The distinction is not cosmetic: forNt≪NN\_\{t\}\\ll N,log2⁡\(NNt\)≈Nt​log2⁡\(e​N/Nt\)\\log\_\{2\}\\binom\{N\}\{N\_\{t\}\}\\approx N\_\{t\}\\log\_\{2\}\(eN/N\_\{t\}\)can exceedk​NtkN\_\{t\}\. The proposition states the worst case; our runs live in case \(a\)\.

#### The bound\.

Assumemmis encoded in a prefix\-free or fixed\-length binary code of length\|m\|\|m\|\. ThenH⁡\(m\)≤\|m\|H\(m\)\\leq\|m\|, and the data\-processing inequality alongF→m→W′F\\to m\\to W^\{\\prime\}gives

I⁡\(F,W′\)≤I⁡\(F,m\)≤H⁡\(m\)≤\|m\|\.I\(F;W^\{\\prime\}\)\\;\\leq\\;I\(F;m\)\\;\\leq\\;H\(m\)\\;\\leq\\;\|m\|\.WithN=1\.409×109N=1\.409\\times 10^\{9\}constrained weights andk=4k=4, case \(a\) gives a ceiling of5\.6×1095\.6\\times 10^\{9\}bits\.

#### What is bounded, and what is not\.

This is the second place a reviewer should push, and the honest answer has two parts\.

*The DPI bounds mutual information, not recall\.*To convert the proposition into a statement about the measured quantity we need the converse direction, Fano’s inequality\. The cloze probe supplies the subject name and asks for an attribute, so the decoder isA^=g⁡\(W′,names\)\\hat\{A\}=g\(W^\{\\prime\},\\text\{names\}\)and the relevant quantity is the conditional mutual informationI⁡\(A;W′∣names\)I\(A;W^\{\\prime\}\\mid\\text\{names\}\), whereAAcollects the probed attributes; names and attributes are independent by construction, and the same chain conditioned on names givesI⁡\(A;W′∣names\)≤\|m\|I\(A;W^\{\\prime\}\\mid\\text\{names\}\)\\leq\|m\|\. IfAAis uniform on a product of finite attribute sets—true by construction,H⁡\(A\)=11\.9H\(A\)=11\.9bits per fact for the three probed attributes—and the probe errs with probabilityPeP\_\{e\}, Fano givesH⁡\(A∣W′,names\)≤h⁡\(Pe\)\+Pe​H​\(A\)H\(A\\mid W^\{\\prime\},\\text\{names\}\)\\leq h\(P\_\{e\}\)\+P\_\{e\}H\(A\), withhhthe binary entropy andh≤1h\\leq 1, and hence

\(1−Pe\)​H​\(A\)−1≤I⁡\(A;W′∣names\)≤\|m\|\.\(1\-P\_\{e\}\)\\,H\(A\)\-1\\;\\leq\\;I\(A;W^\{\\prime\}\\mid\\text\{names\}\)\\;\\leq\\;\|m\|\.The left\-hand side is exactly the accounting used throughout the paper,recall×Nfacts×11\.9\\text\{recall\}\\times N\_\{\\text\{facts\}\}\\times 11\.9bits: the “absorbed bits” column is*not*a Fano lower bound, and an earlier version of this appendix said it was\. Applied per fact, Fano givesI≥\(1−Pe\)​H​\(A\)−h⁡\(Pe\)I\\geq\(1\-P\_\{e\}\)H\(A\)\-h\(P\_\{e\}\)for each, so the bound on the whole corpus isrecall×N×11\.9−N​h​\(Pe\)\\text\{recall\}\\times N\\times 11\.9\-N\\,h\(P\_\{e\}\)bits, which sits below the column by13%13\\%at the highest\-recall arm and28%28\\%at the lowest — a recall\-dependent gap, not a constant factor, so it does not cancel in the efficiency ratios either\. What the column is, is a recall\-weighted count of probed attribute bits: a consistent measure of absorption across arms, and the one all our comparisons use\. The information\-theoretic statement that does hold is the upper bound,I⁡\(A;W′∣names\)≤\|m\|I\(A;W^\{\\prime\}\\mid\\text\{names\}\)\\leq\|m\|, which is Proposition[4](https://arxiv.org/html/2608.20873#Thmproposition4)and is what bounds capacity\. Three caveats travel with it\. Recall is measured by exact\-match greedy decoding, so information the model holds but cannot emit is not counted, and the estimate is conservative in that direction\. The probes are not independent of one another and the6\.5%6\.5\\%vocabulary floor \(§[4](https://arxiv.org/html/2608.20873#S4)\) must be subtracted before reading small recalls as information\. And Fano requires the decoder to use onlyW′W^\{\\prime\}and the public subject list, which is why the probe never conditions on the corpus\.

*Novelty is a hypothesis of the theorem, not a conclusion\.*IfW^\\hat\{W\}already depended onFF,I⁡\(F,W′\)I\(F;W^\{\\prime\}\)could be large with\|m\|=0\|m\|=0and the identity would say nothing about learning\. The synthetic corpus is generated after the release and the base model’s recall is at the vocabulary floor \(§[4](https://arxiv.org/html/2608.20873#S4)\), so the required independence holds by construction; on real post\-release facts it would have to be argued rather than assumed\.

*The ceiling is not an estimate\.*Nothing above asserts thatk​NtkN\_\{t\}bits are*achievable*; the bound is one\-sided\. Empirically the gap is five orders of magnitude—4,1304\{,\}130facts absorbed at10410^\{4\}presented,≈49\\approx 49kbit of probed attribute entropy, against a5\.65\.6Gbit ceiling \(§[5](https://arxiv.org/html/2608.20873#S5)\)—which is the quantitative form of the claim that the regime is optimization\-limited rather than capacity\-limited\. Finally, the bound scales with what is*shipped*: a full\-precision residual would give\|m\|=32​Nt\|m\|=32N\_\{t\}, eight times weaker atk=4k=4\. Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)is what makes Proposition[4](https://arxiv.org/html/2608.20873#Thmproposition4)bite\.

### B\.4Proposition[5](https://arxiv.org/html/2608.20873#Thmproposition5)\(forgetting budget\)

#### Second\-order expansion\.

Fix an input distribution𝒟\\mathcal\{D\}and writeφx\(δ\)=KL\(pW^\(⋅∣x\)∥pW^\+δ\(⋅∣x\)\)\\varphi\_\{x\}\(\\delta\)=\\mathrm\{KL\}\\big\(p\_\{\\hat\{W\}\}\(\\cdot\\mid x\)\\,\\\|\\,p\_\{\\hat\{W\}\+\\delta\}\(\\cdot\\mid x\)\\big\)\. The model is a softmax head over a network smooth in its weights \(SiLU activations\), sopW​\(y∣x\)\>0p\_\{W\}\(y\\mid x\)\>0andW↦log⁡pW​\(y∣x\)W\\mapsto\\log p\_\{W\}\(y\\mid x\)isC3C^\{3\}on a neighborhood of the compact box𝒞\\mathcal\{C\}\. Thenφx≥0\\varphi\_\{x\}\\geq 0withφx​\(0\)=0\\varphi\_\{x\}\(0\)=0, soδ=0\\delta=0is a global minimum and∇φx​\(0\)=0\\nabla\\varphi\_\{x\}\(0\)=0; explicitly,∇δφx\(0\)=−𝔼y∼pW^\[∇WlogpW\(y∣x\)\]=−∇W∑ypW\(y∣x\)=−∇W1=0\\nabla\_\{\\delta\}\\varphi\_\{x\}\(0\)=\-\\mathbb\{E\}\_\{y\\sim p\_\{\\hat\{W\}\}\}\\big\[\\nabla\_\{W\}\\log p\_\{W\}\(y\\mid x\)\\big\]=\-\\nabla\_\{W\}\\sum\_\{y\}p\_\{W\}\(y\\mid x\)=\-\\nabla\_\{W\}1=0, the zero\-mean\-score identity\. Differentiating once more and using the same identity,∇2φx\(0\)=𝔼y∼pW^\[∇logp∇logp⊤\]=Fx\\nabla^\{2\}\\varphi\_\{x\}\(0\)=\\mathbb\{E\}\_\{y\\sim p\_\{\\hat\{W\}\}\}\\big\[\\nabla\\log p\\,\\nabla\\log p^\{\\\!\\top\}\\big\]=F\_\{x\}, the Fisher information atW^\\hat\{W\}\. Taylor’s theorem with remainder, uniformly on the compact set𝒞\\mathcal\{C\}, gives

𝔼x∼𝒟KL\(pW^∥pW^\+δ\)=12δ⊤Fδ\+O\(∥δ∥3\),F=𝔼x∼𝒟Fx,\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\}\\,\\mathrm\{KL\}\\big\(p\_\{\\hat\{W\}\}\\\|p\_\{\\hat\{W\}\+\\delta\}\\big\)=\\tfrac\{1\}\{2\}\\,\\delta^\{\\\!\\top\}F\\delta\+O\(\\\|\\delta\\\|^\{3\}\),\\qquad F=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\}F\_\{x\},which is Proposition[5](https://arxiv.org/html/2608.20873#Thmproposition5)\. The remainder constant is uniform over the box but is*not measured*; everything below turns on that\.

#### Whereρ2\\rho^\{2\}comes from\.

Two differentρ2\\rho^\{2\}statements are in play and they should not be conflated\.

*Homogeneity \(exact, no approximation\)\.*The constraint set at radiusρ\\rhoisρ​ℬ\\rho\\mathcal\{B\}withℬ\\mathcal\{B\}the fixed box of half\-widthsroomi\\mathrm\{room\}\_\{i\}, andδ↦12​δ⊤​F​δ\\delta\\mapsto\\tfrac\{1\}\{2\}\\delta^\{\\\!\\top\}F\\deltais homogeneous of degree two, so for*any*PSDFF, diagonal or not,supδ∈ρ​ℬ12​δ⊤​F​δ=ρ2​supδ∈ℬ12​δ⊤​F​δ\\sup\_\{\\delta\\in\\rho\\mathcal\{B\}\}\\tfrac\{1\}\{2\}\\delta^\{\\\!\\top\}F\\delta=\\rho^\{2\}\\sup\_\{\\delta\\in\\mathcal\{B\}\}\\tfrac\{1\}\{2\}\\delta^\{\\\!\\top\}F\\delta\. Theρ2\\rho^\{2\}law is thus a property of the geometry, and survives even when the diagonal surrogate does not\.

*Worst case over the box, diagonal surrogate\.*Dropping off\-diagonal terms and using the loose radiusΔi/2≥roomi\\Delta\_\{i\}/2\\geq\\mathrm\{room\}\_\{i\},sup\|δi\|≤ρ​Δi/212​∑iFi​i​δi2=ρ28​∑iFi​i​Δi2=:BF​\(ρ\)\\sup\_\{\|\\delta\_\{i\}\|\\leq\\rho\\Delta\_\{i\}/2\}\\tfrac\{1\}\{2\}\\sum\_\{i\}F\_\{ii\}\\delta\_\{i\}^\{2\}=\\tfrac\{\\rho^\{2\}\}\{8\}\\sum\_\{i\}F\_\{ii\}\\Delta\_\{i\}^\{2\}=:B\_\{F\}\(\\rho\), the expression quoted in §[2](https://arxiv.org/html/2608.20873#S2)\.

*Uniform in\-cell fill \(what is measured\)\.*The geometry experiment setsδi=ui​ρ​roomi\\delta\_\{i\}=u\_\{i\}\\,\\rho\\,\\mathrm\{room\}\_\{i\}withui∼𝒰⁡\(−1,1\)u\_\{i\}\\sim\\mathcal\{U\}\(\-1,1\)i\.i\.d\., which is the natural model of “filling the cells with random content”\. Then𝔼⁡\[δi\]=0\\mathbb\{E\}\[\\delta\_\{i\}\]=0,𝔼⁡\[δi2\]=ρ2​roomi2/3\\mathbb\{E\}\[\\delta\_\{i\}^\{2\}\]=\\rho^\{2\}\\mathrm\{room\}\_\{i\}^\{2\}/3and𝔼⁡\[δi​δj\]=0\\mathbb\{E\}\[\\delta\_\{i\}\\delta\_\{j\}\]=0fori≠ji\\neq j, so

𝔼⁡\[12​δ⊤​F​δ\]=12​∑i,jFi​j​𝔼​\[δi​δj\]=ρ26​∑iFi​i​roomi2,\\mathbb\{E\}\\Big\[\\tfrac\{1\}\{2\}\\delta^\{\\\!\\top\}F\\delta\\Big\]=\\tfrac\{1\}\{2\}\\sum\_\{i,j\}F\_\{ij\}\\,\\mathbb\{E\}\[\\delta\_\{i\}\\delta\_\{j\}\]=\\frac\{\\rho^\{2\}\}\{6\}\\sum\_\{i\}F\_\{ii\}\\,\\mathrm\{room\}\_\{i\}^\{2\},which is the predictor implemented inexperiments/exp\_geom\.py\. \(The same symmetry kills the first\-order term of the*evaluation*loss, which does not vanish atW^\\hat\{W\}:𝔼​⟨∇L,δ⟩=0\\mathbb\{E\}\\langle\\nabla L,\\delta\\rangle=0\. So the second\-order prediction applies to the measuredΔ\\DeltaNLL as well as to the KL\.\)

#### Why the diagonal surrogate fails empirically, and what that rules out\.

Measured againstρ\\rho\-scaled uniform in\-cell fills of Qwen3\-1\.7B \(1\.409×1091\.409\\times 10^\{9\}perturbed coordinates, two perturbation seeds,∑iFi​i​roomi2=2\.03×10−3\\sum\_\{i\}F\_\{ii\}\\mathrm\{room\}\_\{i\}^\{2\}=2\.03\\times 10^\{\-3\}\), the predictor underestimates the measuredΔ\\DeltaNLL by16×16\\timesatρ=0\.125\\rho=0\.125,196×196\\timesat0\.250\.25,291×291\\timesat0\.50\.5,340×340\\timesat0\.750\.75and383×383\\timesatρ=1\\rho=1; fitting the measured curve overρ∈\[0\.25,0\.75\]\\rho\\in\[0\.25,0\.75\]against its value atρ=1\\rho=1gives exponent2\.432\.43, not22\(§[5\.9](https://arxiv.org/html/2608.20873#S5.SS9),results/exp\_geom\_1p7b\.json\)\. The surrogate is therefore not a usable numerical bound, as §[2](https://arxiv.org/html/2608.20873#S2)states\.

The displayed computation, however, rules out one tempting explanation\. Under independent zero\-mean coordinatewise perturbations the off\-diagonal entries ofFFcontribute nothing to the*mean*ofδ⊤​F​δ\\delta^\{\\\!\\top\}F\\delta; they enter only its variance\. Whether a single realization sits near that mean is empirical, and two seeds bound it only crudely: the per\-seedΔ\\DeltaNLL, recovered fromresults/exp\_geom\_1p7b\.jsonaslog⁡pplρ−log⁡pplt=0\\log\\mathrm\{ppl\}\_\{\\rho\}\-\\log\\mathrm\{ppl\}\_\{t=0\}, ranges over39%39\\%,16%16\\%,14%14\\%and15%15\\%of its mean atρ=0\.25\\rho=0\.25,0\.50\.5,0\.750\.75and11, and atρ=0\.125\\rho=0\.125the two seeds do not agree even in sign\. \(The raw perplexities agree to within2%2\\%at every radius, but that is not the quantity at issue: the signal is their small difference from the anchor\.\) A1414–39%39\\%fluctuation is still two orders of magnitude short of the196196–383×383\\timesgap, so high dimension alone does not make off\-diagonal Fisher mass inflate the second\-order prediction\. Two candidate explanations remain, and we can separate neither with the present data\. \(a\)*The truncation itself fails at this magnitude*: the perturbation is full\-dimensional, so‖δ‖\\\|\\delta\\\|is macroscopic even though each coordinate moves less than a cell half\-width, and theO⁡\(‖δ‖3\)O\(\\\|\\delta\\\|^\{3\}\)remainder need not be small; the measured exponent2\.43\>22\.43\>2is direct evidence of this\. \(b\)*The estimator ofFFis biased low*: the implementation squares the gradient of a loss averaged over a10241024\-token chunk, and for approximately independent zero\-mean per\-token gradients𝔼⁡\[\(1n​∑tgt\)2\]≈1n​𝔼​\[g2\]\\mathbb\{E\}\[\(\\tfrac\{1\}\{n\}\\sum\_\{t\}g\_\{t\}\)^\{2\}\]\\approx\\tfrac\{1\}\{n\}\\mathbb\{E\}\[g^\{2\}\], an underestimate of the per\-token Fisher diagonal by a factor of ordern=1024n=1024— the right order of magnitude for the observed gap\. A further, unquantified gap is that the measured quantity is the change in evaluation NLL, whose Hessian equalsFFonly at a stationary point, andW^\\hat\{W\}is not one: perplexity falls monotonically along the whole segment fromW^\\hat\{W\}toWorigW\_\{\\mathrm\{orig\}\}\(§[5\.9](https://arxiv.org/html/2608.20873#S5.SS9)\)\.

#### Status\.

What Proposition[5](https://arxiv.org/html/2608.20873#Thmproposition5)delivers is \(i\) an exact second\-order identity with the Fisher form, \(ii\) an exactρ2\\rho^\{2\}homogeneity statement for that form on the box, and \(iii\) the observation that the box radii are supplied by the quantization grid at no cost\. What it does not deliver is a computable a\-priori certificate: the diagonal surrogate is off by1616–383×383\\timesand the measured exponent is2\.432\.43\. We therefore treatρ\\rhoas a calibrated dial and report the measured curve \(Fig\.[3](https://arxiv.org/html/2608.20873#S5.F3)b\) rather than a bound\.

## Appendix CHyperparameters

Every value that appears in any archivedconfigblock\.

## Appendix DRun manifest

Archived filenames useqil, the working name under which CellFill was developed \(exp5\_qil\.py,exp10\_qil\_r64\_\*,exp16\_qil\_r16\_\*\); the method and the files are otherwise identical\. Every number in this paper comes from one of the files below, which are released with the code; tables and figures are generated from them byscripts/make\_tables\.pyandscripts/fig\_\*\.pyrather than transcribed\. Wall\-clock times are for a single RTX 4090 \(48 GB\) or one A100 \(80 GB\) for the 8B and larger runs\.

## Appendix EStorage precision of the residual

The refinement code of Proposition[3](https://arxiv.org/html/2608.20873#Thmproposition3)places reconstructions at sub\-cell centers, strictly interior to the cell\. Checkpoint dtype nonetheless matters: at LLM weight scales, bf16’s 8\-bit mantissa can round a stored value across a decision boundary\. With the 1% cell margin used at 1\.7B \(2% at 27B\), fp32 and fp16 storage are safe and bf16 is not; a 5% margin makes bf16 safe at the cost of 8% of the writable range\. This is verified by unit test over randomly perturbed blocks\.

## References

- \[1\]T\. Dettmers, A\. Pagnoni, A\. Holtzman, L\. Zettlemoyer\. QLoRA: Efficient finetuning of quantized LLMs\.*NeurIPS*, 2023\. arXiv:2305\.14314\.
- \[2\]Y\. Li, Y\. Yu, C\. Liang, P\. He, N\. Karampatziakis, W\. Chen, T\. Zhao\. LoftQ: LoRA\-fine\-tuning\-aware quantization for large language models\.*ICLR*, 2024\. arXiv:2310\.08659\.
- \[3\]Y\. Xu, L\. Xie, X\. Gu, X\. Chen, H\. Chang, H\. Zhang, Z\. Chen, X\. Zhang, Q\. Tian\. QA\-LoRA: Quantization\-aware low\-rank adaptation of large language models\.*ICLR*, 2024\. arXiv:2309\.14717\.
- \[4\]Y\. Bondarenko, R\. Del Chiaro, M\. Nagel\. Low\-rank quantization\-aware training for LLMs\. 2024\. arXiv:2406\.06385\.
- \[5\]M\. Nagel, R\. A\. Amjad, M\. van Baalen, C\. Louizos, T\. Blankevoort\. Up or down? Adaptive rounding for post\-training quantization\.*ICML*, 2020\. arXiv:2004\.10568\.
- \[6\]Y\. Li, R\. Gong, X\. Tan, Y\. Yang, P\. Hu, Q\. Zhang, F\. Yu, W\. Wang, S\. Gu\. BRECQ: Pushing the limit of post\-training quantization by block reconstruction\.*ICLR*, 2021\. arXiv:2102\.05426\.
- \[7\]P\. Nair, P\. Datta, J\. Dean, P\. Jain, A\. Kusupati\. Matryoshka quantization\. 2025\. arXiv:2502\.06786\.
- \[8\]Y\. Park, J\. Hyun, S\. Cho, B\. Sim, J\. W\. Lee\. Any\-precision LLM: Low\-cost deployment of multiple, different\-sized LLMs\.*ICML*, 2024\. arXiv:2402\.10517\.
- \[9\]S\. Savkin, E\. Porat, O\. Ordentlich, Y\. Polyanskiy\. NestQuant: Nested lattice quantization for matrix products and LLMs\. 2025\. arXiv:2502\.09720\.
- \[10\]M\. Kleinegger, E\. Crnčević, D\. Alistarh\. MatGPTQ: Accurate and efficient post\-training Matryoshka quantization\. 2026\. arXiv:2602\.03537\.
- \[11\]Y\. Luo, B\. Dong, W\. Cheng, H\. Shen\. Recurrent residual quantization: A progressive multi\-precision representation for LLMs\. 2026\. arXiv:2608\.04048\.
- \[12\]J\. Liu, G\. Xiao, K\. Li, J\. D\. Lee, S\. Han\. BitDelta: Your fine\-tune may only be worth one bit\.*NeurIPS*, 2024\. arXiv:2402\.10193\.
- \[13\]Y\. Sheng, S\. Cao, D\. Li, C\. Hooper, N\. Lee, S\. Yang, C\. Chou, B\. Zhu, L\. Zheng, K\. Keutzer, J\. Gonzalez, I\. Stoica\. S\-LoRA: Serving thousands of concurrent LoRA adapters\.*MLSys*, 2024\. arXiv:2311\.03285\.
- \[14\]J\. Kirkpatrick, R\. Pascanu, N\. Rabinowitz, et al\. Overcoming catastrophic forgetting in neural networks\.*PNAS*114\(13\), 2017\.
- \[15\]A\. Robins\. Catastrophic forgetting, rehearsal and pseudorehearsal\.*Connection Science*7\(2\), 1995\.
- \[16\]K\. Meng, D\. Bau, A\. Andonian, Y\. Belinkov\. Locating and editing factual associations in GPT\.*NeurIPS*, 2022\. arXiv:2202\.05262\.
- \[17\]K\. Meng, A\. S\. Sharma, A\. Andonian, Y\. Belinkov, D\. Bau\. Mass\-editing memory in a transformer\.*ICLR*, 2023\. arXiv:2210\.07229\.
- \[18\]Z\. Allen\-Zhu, Y\. Li\. Physics of language models: Part 3\.3, knowledge capacity scaling laws\. 2024\. arXiv:2404\.05405\.
- \[19\]J\. Frankle, G\. K\. Dziugaite, D\. M\. Roy, M\. Carbin\. Linear mode connectivity and the lottery ticket hypothesis\.*ICML*, 2020\. arXiv:1912\.05671\.
- \[20\]M\. Wortsman, G\. Ilharco, S\. Gadre, et al\. Model soups: Averaging weights of multiple fine\-tuned models improves accuracy without increasing inference time\.*ICML*, 2022\. arXiv:2203\.05482\.
- \[21\]M\. Wortsman, G\. Ilharco, J\. W\. Kim, et al\. Robust fine\-tuning of zero\-shot models\.*CVPR*, 2022\. arXiv:2109\.01903\.
- \[22\]T\. Dettmers, M\. Lewis, Y\. Belkada, L\. Zettlemoyer\. LLM\.int8\(\): 8\-bit matrix multiplication for transformers at scale\.*NeurIPS*, 2022\. arXiv:2208\.07339\.
- \[23\]E\. Frantar, S\. Ashkboos, T\. Hoefler, D\. Alistarh\. GPTQ: Accurate post\-training quantization for generative pre\-trained transformers\.*ICLR*, 2023\. arXiv:2210\.17323\.
- \[24\]J\. Lin, J\. Tang, H\. Tang, S\. Yang, W\.\-M\. Chen, W\.\-C\. Wang, G\. Xiao, X\. Dang, C\. Gan, S\. Han\. AWQ: Activation\-aware weight quantization for LLM compression and acceleration\.*MLSys*, 2024\. arXiv:2306\.00978\.
- \[25\]G\. Ilharco, M\. T\. Ribeiro, M\. Wortsman, S\. Gururangan, L\. Schmidt, H\. Hajishirzi, A\. Farhadi\. Editing models with task arithmetic\.*ICLR*, 2023\. arXiv:2212\.04089\.
- \[26\]P\. Yadav, D\. Tam, L\. Choshen, C\. Raffel, M\. Bansal\. TIES\-Merging: Resolving interference when merging models\.*NeurIPS*, 2023\. arXiv:2306\.01708\.
- \[27\]T\. Hartvigsen, S\. Sankaranarayanan, H\. Palangi, Y\. Kim, M\. Ghassemi\. Aging with GRACE: Lifelong model editing with discrete key\-value adaptors\.*NeurIPS*, 2023\. arXiv:2211\.11031\.
- \[28\]E\. Mitchell, C\. Lin, A\. Bosselut, C\. Finn, C\. D\. Manning\. Fast model editing at scale\.*ICLR*, 2022\. arXiv:2110\.11309\.
- \[29\]A\. Kusupati, G\. Bhatt, A\. Rege, et al\. Matryoshka representation learning\.*NeurIPS*, 2022\. arXiv:2205\.13147\.
- \[30\]T\. M\. Cover, J\. A\. Thomas\.*Elements of Information Theory*, 2nd ed\. Wiley, 2006\.
- \[31\]M\. Mitchell, S\. Wu, A\. Zaldivar, et al\. Model cards for model reporting\.*FAT\**, 2019\. arXiv:1810\.03993\.

Similar Articles

Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference

arXiv cs.AI

This paper presents the first systematic study of zero-knowledge-friendly quantization for LLMs, formalizing the concept and evaluating nine models (including Qwen2.5-14B and Qwen3-30B-A3B) across weight, activation, and nonlinear lookup table precision. It finds that activation precision and RMSNorm inverse-square-root lookups dominate ZK proving costs and that conventional low-bit quantization heuristics do not translate to proportional proving savings.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

arXiv cs.CL

Mix-Quant proposes a phase-aware quantization framework for agentic LLMs, using NVFP4 quantization for the prefilling stage to accelerate computation while preserving BF16 precision for decoding to maintain accuracy. The method achieves up to 3x speedup in prefilling with minimal performance degradation on agentic benchmarks.

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

Hugging Face Daily Papers

This paper presents a framework for quantizing vision-language models to 2.7 bits per parameter, enabling efficient mobile deployment by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB while preserving performance on visual QA tasks.