Exact ReLU realization of affine one-dimensional refinement iterates via residual memory and offset frames

arXiv cs.LG Papers

Summary

This paper proves that every finite affine iterate of vector-valued affine refinement operators admits an exact fixed-width ReLU realization with depth O(n) for M>=3, using a residual memory controller and offset frames. The result extends to arbitrary compactly supported continuous piecewise linear forcing terms.

arXiv:2607.20586v1 Announce Type: new Abstract: We study vector-valued affine refinement operators of the form [ (W\gamma)(t)=\sum_{j\in\mathbb{Z}} A_j\gamma(Mt-j)+B(t), ] with finitely supported matrix mask and compactly supported continuous piecewise linear input and forcing data. Building on the homogeneous realization theorem for (B\equiv 0), we prove that, for (M\ge 3), every finite affine iterate (W^n\gamma) admits an exact fixed-width ReLU realization whose depth is (O(n)). The main new ingredient is a residual memory controller. It replaces the noninvertible residual dynamics by an injective skew-product and permits exact backward replay of the residual states required by a Horner-type evaluation of the affine forcing sum. Offset frames align the forcing atoms away from residual seams, allowing complementary loop readouts to recover their values exactly. The remaining branch-selection ambiguity occurs only where the accumulated affine state has already vanished. For (M\ge 3), the result applies to arbitrary compactly supported continuous piecewise linear forcing terms. For (M=2), the same construction applies to ordinary-frame seam-separated forcing. We also prove a stage-dependent extension for forcing terms in a fixed finite-dimensional continuous piecewise linear span and record the resulting linear-depth upgrade for open-curve, finite-state, and Hilbert- and Morton-type recursive constructions.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:12 AM

# Exact ReLU realization of affine one-dimensional refinement iterates via residual memory and offset frames
Source: [https://arxiv.org/html/2607.20586](https://arxiv.org/html/2607.20586)
Tsogtgerel GantumurMcGill UniversityNational University of MongoliaInstitute of Mathematics and Digital Technology, Mongolian Academy of Sciences

\(July 20, 2026\)

###### Abstract

We study vector\-valued affine refinement operators of the form

\(W​γ\)​\(t\)=∑j∈ℤAj​γ​\(M​t−j\)\+B​\(t\),\(W\\gamma\)\(t\)=\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}\\gamma\(Mt\-j\)\+B\(t\),with finitely supported matrix mask and compactly supported continuous piecewise linear input and forcing data\. Building on the homogeneous realization theorem forB≡0B\\equiv 0, we prove that, forM≥3M\\geq 3, every finite affine iterateWn​γW^\{n\}\\gammaadmits an exact fixed\-width ReLU realization whose depth isO​\(n\)O\(n\)\.

The main new ingredient is a residual memory controller\. It replaces the noninvertible residual dynamics by an injective skew\-product and permits exact backward replay of the residual states required by a Horner\-type evaluation of the affine forcing sum\. Offset frames align the forcing atoms away from residual seams, allowing complementary loop readouts to recover their values exactly\. The remaining branch\-selection ambiguity occurs only where the accumulated affine state has already vanished\.

ForM≥3M\\geq 3, the result applies to arbitrary compactly supported continuous piecewise linear forcing terms\. ForM=2M=2, the same construction applies to ordinary\-frame seam\-separated forcing\. We also prove a stage\-dependent extension for forcing terms in a fixed finite\-dimensional continuous piecewise linear span and record the resulting linear\-depth upgrade for open\-curve, finite\-state, and Hilbert\- and Morton\-type recursive constructions\.

2020 Mathematics Subject Classification\.Primary 68T07; Secondary 41A30, 65D17\.

Keywords\.ReLU neural networks, affine refinement operators, vector\-valued refinement, exact realization, continuous piecewise linear functions, recursive curves\.

## 1Introduction

### 1\.1Background and motivation

A central question in neural\-network approximation theory\[[1](https://arxiv.org/html/2607.20586#bib.bib1),[2](https://arxiv.org/html/2607.20586#bib.bib2)\]is to explain why deep networks can represent highly oscillatory, recursive, or self\-similar functions with relatively small width and depth\. A basic result in this direction is the scalar binary theorem for refinement operators\[[3](https://arxiv.org/html/2607.20586#bib.bib3)\]: if the input function is compactly supported and continuous piecewise linear \(CPwL\), then its finite refinement iterates admit exact fixed\-width ReLU realizations whose depth grows only linearly with the number of refinement steps\.

The scalar binary construction of\[[3](https://arxiv.org/html/2607.20586#bib.bib3)\]naturally suggests extensions to vector\-valued,MM\-ary refinement\. A homogeneous theory in this setting was developed in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]using an exact loop controller for the residual dynamics\. That work proves fixed\-width, depth\-O​\(n\)O\(n\)realizations for homogeneous refinement iterates and also treats affine refinement, but with depthO​\(n2\)O\(n^\{2\}\)\. The corresponding linear\-depth realization of genuinely affine forcing sums therefore remained open for theMM\-ary vector\-valued systems arising in recursive curve generation\.

The main motivation is geometric\. Finite approximants to recursive constructions, including Hilbert\-type curves and Morton\-type traversals, can be generated by vector\-valued refinement rules with affine connector terms\. This parametrized\-generation viewpoint complements\[[5](https://arxiv.org/html/2607.20586#bib.bib5)\], where neural networks are used to approximate indicators or classifiers of fractal sets\. In the present paper, we take the homogeneous theorem of\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]as a black box and address the missing affine problem: the exact fixed\-width realization of the forcing sum with depthO​\(n\)O\(n\)\.

The affine refinement operator considered here has the form

\(W​γ\)​\(t\)=∑j∈ℤAj​γ​\(M​t−j\)\+B​\(t\),\(W\\gamma\)\(t\)=\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}\\gamma\(Mt\-j\)\+B\(t\),\(1\)whereM≥2M\\geq 2, the matricesAj∈ℝp×pA\_\{j\}\\in\\mathbb\{R\}^\{p\\times p\}are fixed and vanish for all but finitely manyjj, andB:ℝ→ℝpB:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}is compactly supported and CPwL\. Writing the homogeneous part as

\(Vγ\)\(t\):=∑j∈ℤAjγ\(Mt−j\),\(V\\gamma\)\(t\):=\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}\\gamma\(Mt\-j\),one obtains the affine iterate identity

Wn​γ=Vn​γ\+∑r=0n−1Vr​B\.W^\{n\}\\gamma=V^\{n\}\\gamma\+\\sum\_\{r=0\}^\{n\-1\}V^\{r\}B\.\(2\)The first term is covered by the homogeneous theorem of\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\. The new question is whether the affine forcing sum can also be realized exactly with fixed width and depthO​\(n\)O\(n\)\.

This question is subtler than the homogeneous one\. In the homogeneous case, the cascade can be organized as a forward adjoint iteration along the residual orbit, while the scalar terminal factor makes certain selector ambiguities harmless\. For affine forcing, the terms enter at many different stages, and the efficient organization is instead a backward Horner\-type evaluation of the affine sum\. This creates a chronology problem: after the residual controller has advanced to depthnn, the network must recover the earlier residual states in the reverse order required by the affine recursion\.

The main new device of this paper is therefore a residual memory controller\. It augments the loop state by a fixed\-dimensional memory coordinate so that the forward controller is injective on the relevant state space and has a CPwL inverse on its image\. After running forward to depthnn, the network can consequently replay the residual states backward exactly and feed them into the affine recursion in the correct order\.

The second new ingredient is the use of offset frames\. Informally, an offset frame is a shifted unit\-cell decomposition of the line\. By combining the ordinary frame with a suitable admissible offset frame, one can express a general CPwL forcing term as a finite sum of elementary hats whose supports avoid the relevant cell seams in at least one frame\. This supplies seam\-safe forcing readouts for the backward affine recursion\. Such a nontrivial admissible offset exists for everyM≥3M\\geq 3; the binary case therefore leads to a more restricted forcing class\.

### 1\.2Main results

The homogeneous finite\-iterate theorem from\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]is recalled in §[2\.5](https://arxiv.org/html/2607.20586#S2.SS5)\. Our main result is its affine linear\-depth counterpart, which may be summarized as follows\.

###### Theorem 1\.1\(Informal statement of the main result\)\.

LetM≥3M\\geq 3, and letWWbe the refinement operator in \([1](https://arxiv.org/html/2607.20586#S1.E1)\)\. Assume that the associated homogeneous operator preserves a compact support window containing the supports of the CPwL inputγ\\gammaand forcing termBB\. Then, for everyn≥1n\\geq 1, the affine iterateWn​γW^\{n\}\\gammahas an exact ReLU realization whose width is bounded independently ofnnand whose depth isO​\(n\)O\(n\)\.

The proof realizes the affine forcing sum in \([2](https://arxiv.org/html/2607.20586#S1.E2)\) by a backward Horner recursion\. The memory controller supplies the residual states in reverse chronological order, while the offset\-frame decomposition supplies seam\-safe evaluations of the forcing term\. The remaining selector ambiguity occurs only where the accumulated affine state has already vanished\. ForM=2M=2, the same conclusion holds for the restricted class of ordinary\-frame seam\-separated forcing terms defined later\.

We also prove a fixed\-span stage\-dependent extension\. Suppose thatWr​γ=V​γ\+BrW\_\{r\}\\gamma=V\\gamma\+B\_\{r\}, where the homogeneous refinement rule is fixed andBr=∑α=0Nλr,α​B\(α\)B\_\{r\}=\\sum\_\{\\alpha=0\}^\{N\}\\lambda\_\{r,\\alpha\}B^\{\(\\alpha\)\}belongs to a fixed finite\-dimensional CPwL span\. The same memory\-Horner construction then gives exact fixed\-width, depth\-O​\(n\)O\(n\)realizations ofWn−1​⋯​W1​W0​γW\_\{n\-1\}\\cdots W\_\{1\}W\_\{0\}\\gamma\. The weight bounds depend on the fixed refinement data and onmaxr<n,α⁡\|λr,α\|\\max\_\{r<n,\\alpha\}\|\\lambda\_\{r,\\alpha\}\|\.

Finally, the anchored\-profile, finite\-state, and copy\-and\-connector reductions developed in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]inherit the improved linear\-depth bound whenever their forcing sequences lie in a fixed finite\-dimensional CPwL span\. This includes the Hilbert\- and Morton\-type recursive constructions discussed there\.

### 1\.3Structure of the paper

Section[2](https://arxiv.org/html/2607.20586#S2)fixes the notation, recalls the vectorized refinement formalism and the homogeneous theorem from\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\], and introduces offset frames and special basis curves\. Section[3](https://arxiv.org/html/2607.20586#S3)proves the core affine forcing result in one admissible frame\. It constructs the residual memory controller, exact backward replay, seam\-safe forcing readouts, and the memory\-Horner recursion\. Section[4](https://arxiv.org/html/2607.20586#S4)combines the ordinary and offset frames to treat arbitrary compactly supported CPwL forcing forM≥3M\\geq 3, proves the main affine theorem, and records the restricted binary result\. Section[5](https://arxiv.org/html/2607.20586#S5)gives the fixed\-span stage\-dependent extension\. Section[6](https://arxiv.org/html/2607.20586#S6)records the resulting linear\-depth upgrade for the geometric reductions of\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\. The final section summarizes the scope of the construction and the remaining binary limitation\.

## 2Preliminaries and notation

This section fixes the notation used in the affine construction\. We recall theMM\-ary digit maps, offset vectorizations, block transition matrices, special basis curves, and the affine iterate formula\. We also record the homogeneous linear\-depth theorem of\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\], which will be used as a black box for the termVn​γV^\{n\}\\gamma\.

### 2\.1Network classes and CPwL functions

For integersW,L,d,N≥1W,L,d,N\\geq 1, we writeΥW,L​\(ReLU;d,N\)\\Upsilon\_\{W,L\}\(\\operatorname\{ReLU\};d,N\)for the class of outputs of fully connected ReLU networks with widthWW, depthLL, input dimensiondd, and output dimensionNN\. The final realization theorems have input dimensiond=1d=1, although several intermediate controller maps act on fixed higher\-dimensional state spaces\.

We use CPwL as shorthand for continuous piecewise linear\. A mapf:ℝd→ℝNf:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{N\}is CPwL if every compact subset ofℝd\\mathbb\{R\}^\{d\}admits a finite polyhedral subdivision on each cell of whichffis affine\.

### 2\.2MM\-ary digit maps, residuals, and offset frames

Fix an integerM≥2M\\geq 2and an offsets∈\[0,1\)s\\in\[0,1\)\. The corresponding*ss\-offset frame*uses the local coordinatex∈\[0,1\)x\\in\[0,1\)on the physical interval\[s,s\+1\)\[s,s\+1\), and more generally on its translates\[s\+k−1,s\+k\)\[s\+k\-1,s\+k\)\. The cases=0s=0is called the*ordinaryMM\-ary frame*\. We use half\-open intervals for the digit dynamics; equalities between continuous functions are extended tox=1x=1by continuity\.

Forx∈\[0,1\)x\\in\[0,1\), define the digit and residual maps by

Q​\(x\):=⌊M​x\+\(M−1\)​s⌋,R​\(x\):=M​x\+\(M−1\)​s−Q​\(x\)\.Q\(x\):=\\lfloor Mx\+\(M\-1\)s\\rfloor,\\qquad R\(x\):=Mx\+\(M\-1\)s\-Q\(x\)\.ThenR​\(x\)∈\[0,1\)R\(x\)\\in\[0,1\)\. The same formula extendsRRto a11\-periodic map onℝ\\mathbb\{R\}, and hence to a map on the circleℝ/ℤ\\mathbb\{R\}/\\mathbb\{Z\}\. This circle dynamics will provide the geometric model for the loop and memory controllers\.

We callss*admissible*if\(M−1\)​s∈ℤ\(M\-1\)s\\in\\mathbb\{Z\}, and then also call the corresponding frame admissible\. Ifc:=\(M−1\)​s∈ℤc:=\(M\-1\)s\\in\\mathbb\{Z\}, then

Q​\(x\)=⌊M​x⌋\+c,R​\(x\)=M​x−⌊M​x⌋,Q\(x\)=\\lfloor Mx\\rfloor\+c,\\qquad R\(x\)=Mx\-\\lfloor Mx\\rfloor,and the digit set isD:=Q​\(\[0,1\)\)=\{c,c\+1,…,c\+M−1\}D:=Q\(\[0,1\)\)=\\\{c,c\+1,\\ldots,c\+M\-1\\\}\. Thus admissibility leaves the ordinaryMM\-ary residual dynamics unchanged and merely shifts the digit labels bycc\. If\(M−1\)​s∉ℤ\(M\-1\)s\\notin\\mathbb\{Z\}, the digit set hasM\+1M\+1elements\.

A nontrivial admissible offset exists for everyM≥3M\\geq 3, for examples=1/\(M−1\)s=1/\(M\-1\)\. ForM=2M=2, no such offset exists; accordingly, the binary result below applies only to a restricted class of forcing terms compatible with the ordinary frame\.

SetR0​\(x\):=xR^\{0\}\(x\):=x\. Forj≥1j\\geq 1, define the successive residuals and digits by

Rj\(x\):=R\(Rj−1\(x\)\),qj\(x\):=Q\(Rj−1\(x\)\)\.R^\{j\}\(x\):=R\\bigl\(R^\{j\-1\}\(x\)\\bigr\),\\qquad q\_\{j\}\(x\):=Q\\bigl\(R^\{j\-1\}\(x\)\\bigr\)\.They satisfy

M​Rj−1​\(x\)\+\(M−1\)​s=qj​\(x\)\+Rj​\(x\),j≥1\.MR^\{j\-1\}\(x\)\+\(M\-1\)s=q\_\{j\}\(x\)\+R^\{j\}\(x\),\\qquad j\\geq 1\.

### 2\.3The vectorized cascade formalism

Fix matricesAj∈ℝp×pA\_\{j\}\\in\\mathbb\{R\}^\{p\\times p\},j∈ℤj\\in\\mathbb\{Z\}, withAj=0A\_\{j\}=0for all but finitely manyjj, and define the homogeneous refinement operator by

\(V​γ\)​\(t\)=∑j∈ℤAj​γ​\(M​t−j\)\.\(V\\gamma\)\(t\)=\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}\\gamma\(Mt\-j\)\.\(3\)We work with a fixed support window\[0,L\]\[0,L\], whereL≥1L\\geq 1is an integer, and assume that

supp⁡f⊂\[0,L\]⟹supp⁡\(V​f\)⊂\[0,L\]\.\\operatorname\{supp\}f\\subset\[0,L\]\\quad\\Longrightarrow\\quad\\operatorname\{supp\}\(Vf\)\\subset\[0,L\]\.The inputγ\\gammaand all forcing terms considered below are supported in this window\.

#### Offset vectorization\.

Fix an offsets∈\[0,1\)s\\in\[0,1\)\. The translated unit intervals\[s\+k−1,s\+k\]\[s\+k\-1,s\+k\]that meet the support window\[0,L\]\[0,L\]in a set of positive length are indexed by

Js=\{\{1,…,L\},s=0,\{0,1,…,L\},0<s<1\.J\_\{s\}=\\begin\{cases\}\\\{1,\\ldots,L\\\},&s=0,\\\\\[2\.84526pt\] \\\{0,1,\\ldots,L\\\},&0<s<1\.\\end\{cases\}We writeLs:=\|Js\|L\_\{s\}:=\|J\_\{s\}\|, so thatLs=LL\_\{s\}=Lin the ordinary frame andLs=L\+1L\_\{s\}=L\+1in a nontrivial offset frame\.

For anyf:ℝ→ℝpf:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}supported in\[0,L\]\[0,L\], define itskk\-th block in thess\-offset frame byfk\(s\)​\(x\):=f​\(x\+s\+k−1\)f\_\{k\}^\{\(s\)\}\(x\):=f\(x\+s\+k\-1\), fork∈Jsk\\in J\_\{s\}andx∈\[0,1\]x\\in\[0,1\]\. Its*ss\-offset vectorization*is

Vecs⁡\(f\)​\(x\):=\(fk\(s\)​\(x\)\)k∈Js𝖳∈ℝp​Ls,x∈\[0,1\],\\operatorname\{Vec\}\_\{s\}\(f\)\(x\):=\\bigl\(f\_\{k\}^\{\(s\)\}\(x\)\\bigr\)\_\{k\\in J\_\{s\}\}^\{\\mathsf\{T\}\}\\in\\mathbb\{R\}^\{pL\_\{s\}\},\\qquad x\\in\[0,1\],where the blocks are ordered by increasingkk\. When the offset frame is fixed, we abbreviateVecs⁡\(f\)\\operatorname\{Vec\}\_\{s\}\(f\)toVec⁡\(f\)\\operatorname\{Vec\}\(f\)\.

In particular, setG:=Vecs⁡\(γ\)G:=\\operatorname\{Vec\}\_\{s\}\(\\gamma\)and, forn≥0n\\geq 0,

Gn:=Vecs⁡\(Vn​γ\)\.G^\{n\}:=\\operatorname\{Vec\}\_\{s\}\(V^\{n\}\\gamma\)\.ThusG0=GG^\{0\}=G, andGnG^\{n\}records all translated unit\-interval pieces ofVn​γV^\{n\}\\gammain a common local coordinate\.

#### Block transition matrices\.

For each digitq∈D:=Q​\(\[0,1\)\)q\\in D:=Q\(\[0,1\)\), defineTq∈ℝp​Ls×p​LsT\_\{q\}\\in\\mathbb\{R\}^\{pL\_\{s\}\\times pL\_\{s\}\}by prescribing itsp×pp\\times pblocks as

\(Tq\)k​ℓ:=Aq\+M​\(k−1\)−\(ℓ−1\),k,ℓ∈Js\.\(T\_\{q\}\)\_\{k\\ell\}:=A\_\{\\,q\+M\(k\-1\)\-\(\\ell\-1\)\},\\qquad k,\\ell\\in J\_\{s\}\.
To explain this definition, fixx∈\[0,1\)x\\in\[0,1\),k∈Jsk\\in J\_\{s\}, and sett:=x\+s\+k−1t:=x\+s\+k\-1\. Ifq=Q​\(x\)q=Q\(x\), then the definition of the residual gives

M​t−j=R​\(x\)\+s\+\(q\+M​\(k−1\)−j\)\.Mt\-j=R\(x\)\+s\+\\bigl\(q\+M\(k\-1\)\-j\\bigr\)\.This argument belongs to theℓ\\ell\-th translated block precisely whenℓ−1=q\+M​\(k−1\)−j\\ell\-1=q\+M\(k\-1\)\-j, or equivalently whenj=q\+M​\(k−1\)−\(ℓ−1\)j=q\+M\(k\-1\)\-\(\\ell\-1\)\. The corresponding coefficient is therefore\(Tq\)k​ℓ\(T\_\{q\}\)\_\{k\\ell\}\. Integer shifts not represented byJsJ\_\{s\}contribute zero by the support convention\. Thus the matrixTqT\_\{q\}records the passage from the translated blocks of a function to those of its refinement on the digit branchqq\.

#### Cascade identities\.

The block matrices encode one refinement step as follows\.

###### Proposition 2\.2\(One\-step cascade identity\)\.

For everyf:ℝ→ℝpf:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}supported in\[0,L\]\[0,L\]and everyx∈\[0,1\)x\\in\[0,1\), one has

Vecs⁡\(V​f\)​\(x\)=Tq1​\(x\)​Vecs⁡\(f\)​\(R​\(x\)\)\.\\operatorname\{Vec\}\_\{s\}\(Vf\)\(x\)=T\_\{q\_\{1\}\(x\)\}\\operatorname\{Vec\}\_\{s\}\(f\)\(R\(x\)\)\.

###### Proof\.

Fixx∈\[0,1\)x\\in\[0,1\), setq:=q1​\(x\)=Q​\(x\)q:=q\_\{1\}\(x\)=Q\(x\), and letk∈Jsk\\in J\_\{s\}\. Thekk\-th block of the left\-hand side is

\[Vecs⁡\(V​f\)​\(x\)\]k=∑j∈ℤAj​f​\(M​x\+M​s\+M​\(k−1\)−j\)\.\\bigl\[\\operatorname\{Vec\}\_\{s\}\(Vf\)\(x\)\\bigr\]\_\{k\}=\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}f\\bigl\(Mx\+Ms\+M\(k\-1\)\-j\\bigr\)\.SinceM​x\+\(M−1\)​s=R​\(x\)\+qMx\+\(M\-1\)s=R\(x\)\+q, the argument offfequalsR​\(x\)\+s\+q\+M​\(k−1\)−jR\(x\)\+s\+q\+M\(k\-1\)\-j\. Settingℓ:=q\+M​\(k−1\)−j\+1\\ell:=q\+M\(k\-1\)\-j\+1, we obtain

\[Vecs⁡\(V​f\)​\(x\)\]k=∑ℓ∈JsAq\+M​\(k−1\)−\(ℓ−1\)​\[Vecs⁡\(f\)​\(R​\(x\)\)\]ℓ\.\\bigl\[\\operatorname\{Vec\}\_\{s\}\(Vf\)\(x\)\\bigr\]\_\{k\}=\\sum\_\{\\ell\\in J\_\{s\}\}A\_\{\\,q\+M\(k\-1\)\-\(\\ell\-1\)\}\\bigl\[\\operatorname\{Vec\}\_\{s\}\(f\)\(R\(x\)\)\\bigr\]\_\{\\ell\}\.Terms withℓ∉Js\\ell\\notin J\_\{s\}vanish by the support convention\. By the definition ofTqT\_\{q\}, the last expression is thekk\-th block ofTq​Vecs⁡\(f\)​\(R​\(x\)\)T\_\{q\}\\operatorname\{Vec\}\_\{s\}\(f\)\(R\(x\)\)\. ∎

Takingf=γf=\\gammagivesG1​\(x\)=Tq1​\(x\)​G​\(R​\(x\)\)G^\{1\}\(x\)=T\_\{q\_\{1\}\(x\)\}G\(R\(x\)\)\. Repeated application yields the following full cascade identity\.

###### Corollary 2\.3\(Iterated cascade identity\)\.

For everyn≥1n\\geq 1andx∈\[0,1\)x\\in\[0,1\), one has

Gn​\(x\)=Tq1​\(x\)​Tq2​\(x\)​⋯​Tqn​\(x\)​G​\(Rn​\(x\)\)\.G^\{n\}\(x\)=T\_\{q\_\{1\}\(x\)\}T\_\{q\_\{2\}\(x\)\}\\cdots T\_\{q\_\{n\}\(x\)\}G\(R^\{n\}\(x\)\)\.

###### Proof\.

Applying the one\-step identity toVn−1​γV^\{n\-1\}\\gamma, we haveGn​\(x\)=Tq1​\(x\)​Gn−1​\(R​\(x\)\)G^\{n\}\(x\)=T\_\{q\_\{1\}\(x\)\}G^\{n\-1\}\(R\(x\)\)\. The result follows by induction fromqj​\(R​\(x\)\)=qj\+1​\(x\)q\_\{j\}\(R\(x\)\)=q\_\{j\+1\}\(x\)andRj​\(R​\(x\)\)=Rj\+1​\(x\)R^\{j\}\(R\(x\)\)=R^\{j\+1\}\(x\)\. ∎

### 2\.4Special hats and basis curves

Let0<ϱ<120<\\varrho<\\frac\{1\}\{2\}be a support margin\. In a one\-frame construction it may be chosen arbitrarily, while in the two\-frame construction it will be chosen after the nontrivial offsetssso that2​ϱ<min⁡\{s,1−s\}2\\varrho<\\min\\\{s,1\-s\\\}\.

###### Definition 2\.4\(Special hat\)\.

A scalar functionh:ℝ→ℝh:\\mathbb\{R\}\\to\\mathbb\{R\}is called a*special hat*if it is nonnegative and CPwL and satisfies

supp⁡h⊂\[ϱ,1−ϱ\]\.\\operatorname\{supp\}h\\subset\[\\varrho,1\-\\varrho\]\.

Leteμ∈ℝpe\_\{\\mu\}\\in\\mathbb\{R\}^\{p\},μ=1,…,p\\mu=1,\\ldots,p, denote the standard coordinate vectors\. In the fixedss\-offset frame, a*special basis curve*is a function of the form

γ​\(t\)=h​\(t−δ\)​eμ,δ∈ℤ\+s,\\gamma\(t\)=h\(t\-\\delta\)e\_\{\\mu\},\\qquad\\delta\\in\\mathbb\{Z\}\+s,wherehhis a special hat\. Writeδ=j0\+s\\delta=j\_\{0\}\+swithj0∈ℤj\_\{0\}\\in\\mathbb\{Z\}\. Fork∈Jsk\\in J\_\{s\}, itskk\-th block is

γk\(s\)​\(x\)=h​\(x\+k−1−j0\)​eμ,x∈\[0,1\]\.\\gamma\_\{k\}^\{\(s\)\}\(x\)=h\(x\+k\-1\-j\_\{0\}\)e\_\{\\mu\},\\qquad x\\in\[0,1\]\.The support condition onhhimplies that all blocks vanish except the one indexed byk=j0\+1k=j\_\{0\}\+1\. Hence, wheneverj0\+1∈Jsj\_\{0\}\+1\\in J\_\{s\}, the offset vectorization is

Vecs⁡\(γ\)​\(x\)=h​\(x\)​uj0\+1,μ,\\operatorname\{Vec\}\_\{s\}\(\\gamma\)\(x\)=h\(x\)u\_\{j\_\{0\}\+1,\\mu\},whereuj0\+1,μ∈ℝp​Lsu\_\{j\_\{0\}\+1,\\mu\}\\in\\mathbb\{R\}^\{pL\_\{s\}\}is the coordinate vector corresponding to componentμ\\muof blockj0\+1j\_\{0\}\+1\.

Thus the offsetssdetermines the frame in which the scalar hat is read, while the integerj0=δ−sj\_\{0\}=\\delta\-sdetermines the active vectorization block\. This separates the scalar residual readout from the matrix\-valued cascade update\.

### 2\.5Affine decomposition and the imported homogeneous theorem

LetB:ℝ→ℝpB:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}be compactly supported and CPwL, withsupp⁡B⊂\[0,L\]\\operatorname\{supp\}B\\subset\[0,L\], and define the affine refinement operator by

W​γ:=V​γ\+B\.W\\gamma:=V\\gamma\+B\.\(4\)A direct induction gives, for everyn≥1n\\geq 1,

Wn​γ=Vn​γ\+Sn,Sn:=∑r=0n−1Vr​B\.W^\{n\}\\gamma=V^\{n\}\\gamma\+S\_\{n\},\\qquad S\_\{n\}:=\\sum\_\{r=0\}^\{n\-1\}V^\{r\}B\.\(5\)Thus the problem separates into the homogeneous termVn​γV^\{n\}\\gammaand the forcing sumSnS\_\{n\}\. The homogeneous contribution is supplied by the following theorem from\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\.

###### Theorem 2\.5\(HomogeneousMM\-ary vector\-valued realization,\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\)\.

LetM≥2M\\geq 2, and letVVbe the homogeneous refinement operator from \([3](https://arxiv.org/html/2607.20586#S2.E3)\), with finitely supported matrix mask and invariant support window\[0,L\]\[0,L\]\. Ifγ:ℝ→ℝp\\gamma:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}is CPwL and supported in\[0,L\]\[0,L\], then there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnn, such that

Vn​γ∈ΥC0,C1​n​\(ReLU;1,p\),n≥1\.V^\{n\}\\gamma\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,p\),\\qquad n\\geq 1\.The realizing networks may moreover be chosen with weights and biases growing at most exponentially innn\.

In view of \([5](https://arxiv.org/html/2607.20586#S2.E5)\), it therefore remains to construct an exact fixed\-width, depth\-O​\(n\)O\(n\)realization of the forcing sumSnS\_\{n\}\.

## 3Affine forcing in one admissible frame

In this section we construct a fixed\-width, linear\-depth realization of the affine forcing sum for forcing terms decomposed into special basis curves aligned with one admissible frame\. The next section combines two such framewise constructions to treat arbitrary compactly supported CPwL forcing whenM≥3M\\geq 3\.

Throughout the section, fix an admissible offsets∈\[0,1\)s\\in\[0,1\)and setc:=\(M−1\)​s∈ℤc:=\(M\-1\)s\\in\\mathbb\{Z\}\. By the discussion in §[2\.2](https://arxiv.org/html/2607.20586#S2.SS2), admissibility givesQ​\(x\)=⌊M​x⌋\+cQ\(x\)=\\lfloor Mx\\rfloor\+candR​\(x\)=M​x−⌊M​x⌋R\(x\)=Mx\-\\lfloor Mx\\rfloorforx∈\[0,1\)x\\in\[0,1\)\. We therefore define

dj​\(x\):=qj​\(x\)−c∈\{0,…,M−1\},𝖳r:=Tc\+r,r=0,…,M−1\.d\_\{j\}\(x\):=q\_\{j\}\(x\)\-c\\in\\\{0,\\ldots,M\-1\\\},\\qquad\\mathsf\{T\}\_\{r\}:=T\_\{c\+r\},\\quad r=0,\\ldots,M\-1\.ThenTqj​\(x\)=𝖳dj​\(x\)T\_\{q\_\{j\}\(x\)\}=\\mathsf\{T\}\_\{d\_\{j\}\(x\)\}for everyj≥1j\\geq 1\.

### 3\.1The vectorized forcing sum and Horner recursion

LetB:ℝ→ℝpB:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}be CPwL and supported in\[0,L\]\[0,L\], and writeb:=Vecs⁡\(B\)b:=\\operatorname\{Vec\}\_\{s\}\(B\)\. We assume thatbbadmits a finite frame\-aligned special\-hat expansion

b​\(x\)=∑ν=1NBaν​hν​\(x\)​vν,supp⁡hν⊂\[ϱ,1−ϱ\],b\(x\)=\\sum\_\{\\nu=1\}^\{N\_\{B\}\}a\_\{\\nu\}h\_\{\\nu\}\(x\)v\_\{\\nu\},\\qquad\\operatorname\{supp\}h\_\{\\nu\}\\subset\[\\varrho,1\-\\varrho\],\(6\)whereaν∈ℝa\_\{\\nu\}\\in\\mathbb\{R\},vν∈ℝp​Lsv\_\{\\nu\}\\in\\mathbb\{R\}^\{pL\_\{s\}\}, and eachhνh\_\{\\nu\}is a special hat\. In particular, ifB​\(t\)=h​\(t−δ\)​eμB\(t\)=h\(t\-\\delta\)e\_\{\\mu\}is a special basis curve withδ∈ℤ\+s\\delta\\in\\mathbb\{Z\}\+s, then the expansion has one term andvνv\_\{\\nu\}is the corresponding coordinate vector of the active block\.

Forn≥1n\\geq 1, setSn:=∑r=0n−1Vr​BS\_\{n\}:=\\sum\_\{r=0\}^\{n\-1\}V^\{r\}BandGn:=Vecs⁡\(Sn\)G\_\{n\}:=\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\. Fork≥0k\\geq 0, the cascade identity gives

Vecs⁡\(Vk​B\)​\(x\)=𝖳d1​\(x\)​⋯​𝖳dk​\(x\)​b​\(Rk​x\),x∈\[0,1\),\\operatorname\{Vec\}\_\{s\}\(V^\{k\}B\)\(x\)=\\mathsf\{T\}\_\{d\_\{1\}\(x\)\}\\cdots\\mathsf\{T\}\_\{d\_\{k\}\(x\)\}b\(R^\{k\}x\),\\qquad x\\in\[0,1\),where the empty product is the identity\. Summing overkk, we obtain

Gn​\(x\)=∑k=0n−1𝖳d1​\(x\)​⋯​𝖳dk​\(x\)​b​\(Rk​x\)\.G\_\{n\}\(x\)=\\sum\_\{k=0\}^\{n\-1\}\\mathsf\{T\}\_\{d\_\{1\}\(x\)\}\\cdots\\mathsf\{T\}\_\{d\_\{k\}\(x\)\}b\(R^\{k\}x\)\.
This sum is evaluated efficiently by a backward Horner recursion\. For fixednn, defineUn−1​\(x\):=b​\(Rn−1​x\)U\_\{n\-1\}\(x\):=b\(R^\{n\-1\}x\), and then set

Uj​\(x\):=b​\(Rj​x\)\+𝖳dj\+1​\(x\)​Uj\+1​\(x\),j=n−2,…,0\.U\_\{j\}\(x\):=b\(R^\{j\}x\)\+\\mathsf\{T\}\_\{d\_\{j\+1\}\(x\)\}U\_\{j\+1\}\(x\),\\qquad j=n\-2,\\ldots,0\.\(7\)
###### Lemma 3\.1\(Horner form of the forcing sum\)\.

For everyn≥1n\\geq 1andx∈\[0,1\)x\\in\[0,1\), one has

U0​\(x\)=Gn​\(x\)=Vecs⁡\(Sn\)​\(x\)\.U\_\{0\}\(x\)=G\_\{n\}\(x\)=\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\(x\)\.

###### Proof\.

Expanding the recursion gives

U0​\(x\)=b​\(x\)\+∑k=1n−1𝖳d1​\(x\)​⋯​𝖳dk​\(x\)​b​\(Rk​x\),U\_\{0\}\(x\)=b\(x\)\+\\sum\_\{k=1\}^\{n\-1\}\\mathsf\{T\}\_\{d\_\{1\}\(x\)\}\\cdots\\mathsf\{T\}\_\{d\_\{k\}\(x\)\}b\(R^\{k\}x\),which is the preceding expression forGn​\(x\)G\_\{n\}\(x\)\. ∎

The following elementary consequence will make the later selector ambiguity harmless\.

###### Lemma 3\.2\(Vanishing accumulated state\)\.

Fixj∈\{0,…,n−2\}j\\in\\\{0,\\ldots,n\-2\\\}\. Ifb​\(Ri​x\)=0b\(R^\{i\}x\)=0for everyi=j\+1,…,n−1i=j\+1,\\ldots,n\-1, thenUj\+1​\(x\)=0U\_\{j\+1\}\(x\)=0\.

###### Proof\.

Starting fromUn−1​\(x\)=b​\(Rn−1​x\)=0U\_\{n\-1\}\(x\)=b\(R^\{n\-1\}x\)=0, the conclusion follows by backward induction from \([7](https://arxiv.org/html/2607.20586#S3.E7)\)\. ∎

### 3\.2The residual memory controller

The affine Horner recursion requires the residual states in reverse chronological order\. Since the residual map isMM\-to\-one on the circle,Rj​xR^\{j\}xcannot be recovered continuously fromRj\+1​xR^\{j\+1\}xalone\. We therefore augment the residual loop by a memory coordinate and replace its noninjective dynamics by an injective skew\-product\.

Let𝕋:=ℝ/ℤ\\mathbb\{T\}:=\\mathbb\{R\}/\\mathbb\{Z\}, and choose a simple polygonal embeddingE:𝕋→Γ⊂\[−1,1\]2E:\\mathbb\{T\}\\to\\Gamma\\subset\[\-1,1\]^\{2\}\. We takeEEto be CPwL with respect to a finite subdivision of the circle\. Define the inverse\-branch separation constant by

ΔE:=min1≤a≤M−1t∈𝕋⁡‖E​\(t\+a/M\)−E​\(t\)‖∞\.\\Delta\_\{E\}:=\\min\_\{\\begin\{subarray\}\{c\}1\\leq a\\leq M\-1\\\\ t\\in\\mathbb\{T\}\\end\{subarray\}\}\\\|E\(t\+a/M\)\-E\(t\)\\\|\_\{\\infty\}\.\(8\)For eacha=1,…,M−1a=1,\\ldots,M\-1, the pointsttandt\+a/Mt\+a/Mare distinct on𝕋\\mathbb\{T\}\. SinceEEis an embedding and the minimum is taken over a compact set, one hasΔE\>0\\Delta\_\{E\}\>0\.

Following\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\], let the degree\-MMcircle mapt↦M​t\(mod1\)t\\mapsto Mt\\pmod\{1\}induceFΓ:Γ→ΓF\_\{\\Gamma\}:\\Gamma\\to\\Gammathrough

FΓ​\(E​\(t\)\):=E​\(M​t\)\.F\_\{\\Gamma\}\(E\(t\)\):=E\(Mt\)\.After subdividingΓ\\Gammaat the images underEEof the breakpoints ofEEand their preimages under the degree\-MMcircle map,FΓF\_\{\\Gamma\}is affine on each resulting edge\. Extending this finite polyhedral complex to a triangulation of a sufficiently large polygon gives a global CPwL extensionF:ℝ2→ℝ2F:\\mathbb\{R\}^\{2\}\\to\\mathbb\{R\}^\{2\}\.

SetC:=\[−1,1\]2C:=\[\-1,1\]^\{2\}andX:=Γ×C⊂ℝ4X:=\\Gamma\\times C\\subset\\mathbb\{R\}^\{4\}\. Chooseα,β\>0\\alpha,\\beta\>0so that

α\+β≤1,2​α<β​ΔE\.\\alpha\+\\beta\\leq 1,\\qquad 2\\alpha<\\beta\\Delta\_\{E\}\.\(9\)Such a choice is possible by fixing anyβ∈\(0,1\)\\beta\\in\(0,1\)and then choosing

0<α<min⁡\{1−β,β​ΔE/2\}\.0<\\alpha<\\min\\\{1\-\\beta,\\beta\\Delta\_\{E\}/2\\\}\.Define the memory update by

𝒬​\(r,y\):=\(F​\(r\),β​r\+α​y\),\(r,y\)∈ℝ2×ℝ2\.\\mathcal\{Q\}\(r,y\):=\\bigl\(F\(r\),\\,\\beta r\+\\alpha y\\bigr\),\\qquad\(r,y\)\\in\\mathbb\{R\}^\{2\}\\times\\mathbb\{R\}^\{2\}\.\(10\)Forx∈\[0,1\]x\\in\[0,1\], define the valid memory states by

Z0​\(x\):=\(E​\(x\),0\),Zj​\(x\):=𝒬j​\(Z0​\(x\)\)\.Z\_\{0\}\(x\):=\(E\(x\),0\),\\qquad Z\_\{j\}\(x\):=\\mathcal\{Q\}^\{j\}\(Z\_\{0\}\(x\)\)\.In particular, the first coordinate ofZj​\(x\)Z\_\{j\}\(x\)isE​\(Mj​x\)E\(M^\{j\}x\), with the argument understood modulo11\.

###### Lemma 3\.3\(Injective residual memory map\)\.

The memory update has the following properties\.

1. \(i\)𝒬​\(X\)⊂X\\mathcal\{Q\}\(X\)\\subset X\.
2. \(ii\)The restriction𝒬\|X\\mathcal\{Q\}\|\_\{X\}is injective\.
3. \(iii\)The restriction𝒬\|X\\mathcal\{Q\}\|\_\{X\}is piecewise affine on a finite polyhedral complex\.
4. \(iv\)Its inverse on𝒬​\(X\)\\mathcal\{Q\}\(X\)is piecewise affine and admits a global CPwL extensionP:ℝ4→ℝ4P:\\mathbb\{R\}^\{4\}\\to\\mathbb\{R\}^\{4\}\. In particular, P​\(𝒬​\(Z\)\)=Z,Z∈X\.P\(\\mathcal\{Q\}\(Z\)\)=Z,\\qquad Z\\in X\.

###### Proof\.

Let\(r,y\)∈X\(r,y\)\\in X\. Sincer∈Γr\\in\\Gamma, one hasF​\(r\)=FΓ​\(r\)∈ΓF\(r\)=F\_\{\\Gamma\}\(r\)\\in\\Gamma\. Moreover,

‖β​r\+α​y‖∞≤β​‖r‖∞\+α​‖y‖∞≤α\+β≤1,\\\|\\beta r\+\\alpha y\\\|\_\{\\infty\}\\leq\\beta\\\|r\\\|\_\{\\infty\}\+\\alpha\\\|y\\\|\_\{\\infty\}\\leq\\alpha\+\\beta\\leq 1,and hence𝒬​\(r,y\)∈X\\mathcal\{Q\}\(r,y\)\\in X\. This proves \(i\)\.

To prove injectivity, suppose that𝒬​\(E​\(t\),y\)=𝒬​\(E​\(u\),y~\)\\mathcal\{Q\}\(E\(t\),y\)=\\mathcal\{Q\}\(E\(u\),\\widetilde\{y\}\)\. Equality of the first coordinates givesE​\(M​t\)=E​\(M​u\)E\(Mt\)=E\(Mu\)\. SinceEEis injective on𝕋\\mathbb\{T\}, it follows thatu≡t\+a/M\(mod1\)u\\equiv t\+a/M\\pmod\{1\}for somea∈\{0,…,M−1\}a\\in\\\{0,\\ldots,M\-1\\\}\.

Ifa=0a=0, thenE​\(u\)=E​\(t\)E\(u\)=E\(t\), and equality of the second coordinates givesy=y~y=\\widetilde\{y\}\. Ifa≠0a\\neq 0, the definition ofΔE\\Delta\_\{E\}and equality of the memory coordinates imply

β​ΔE≤β​‖E​\(t\)−E​\(u\)‖∞=α​‖y~−y‖∞≤2​α,\\beta\\Delta\_\{E\}\\leq\\beta\\\|E\(t\)\-E\(u\)\\\|\_\{\\infty\}=\\alpha\\\|\\widetilde\{y\}\-y\\\|\_\{\\infty\}\\leq 2\\alpha,contradicting \([9](https://arxiv.org/html/2607.20586#S3.E9)\)\. Thusa=0a=0, proving \(ii\)\.

For \(iii\), subdivideΓ\\Gammaso thatFΓF\_\{\\Gamma\}is affine on every edge, triangulateCC, and triangulate the resulting product cells inX=Γ×CX=\\Gamma\\times C\. On every simplex of this finite complex, bothr↦F​\(r\)r\\mapsto F\(r\)and\(r,y\)↦β​r\+α​y\(r,y\)\\mapsto\\beta r\+\\alpha yare affine\.

Finally, refine this complex if necessary so that𝒬\\mathcal\{Q\}is affine on every simplex\. Because𝒬\|X\\mathcal\{Q\}\|\_\{X\}is globally injective, the images of two simplices intersect only in the image of their intersection\. The image simplices therefore form a finite polyhedral complex on𝒬​\(X\)\\mathcal\{Q\}\(X\), and the inverse is affine on each image simplex\. These affine pieces agree on their common faces, so the inverseP:𝒬​\(X\)→XP:\\mathcal\{Q\}\(X\)\\to Xis piecewise affine\.

Choose a sufficiently large boxK⊂ℝ4K\\subset\\mathbb\{R\}^\{4\}containing𝒬​\(X\)\\mathcal\{Q\}\(X\)in its interior, and extend the image complex to a finite triangulation ofKK\. Assign the prescribed values on the vertices of𝒬​\(X\)\\mathcal\{Q\}\(X\), assign zero values on the boundary vertices ofKK, and choose arbitrary values at the remaining interior vertices\. Affine extension over the simplices, followed by the zero extension outsideKK, gives a global CPwL mapP:ℝ4→ℝ4P:\\mathbb\{R\}^\{4\}\\to\\mathbb\{R\}^\{4\}\. ∎

The following is immediate\.

###### Corollary 3\.4\(Exact backward replay\)\.

For every0≤j≤n0\\leq j\\leq n, one has

Pj​\(Zn​\(x\)\)=Zn−j​\(x\)\.P^\{j\}\(Z\_\{n\}\(x\)\)=Z\_\{n\-j\}\(x\)\.

### 3\.3Lifted forcing readouts

The loop state records the residual point on𝕋\\mathbb\{T\}, but no single continuous scalar readout onΓ\\Gammacan recover its representative in\[0,1\]\[0,1\]across the seam\. Following the loop\-coordinate construction of\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\], we therefore use two complementary seam\-compatible readouts\.

Fix0<ε<ϱ0<\\varepsilon<\\varrho, and definer−,r\+:\[0,1\]→\[0,1\]r^\{\-\},r^\{\+\}:\[0,1\]\\to\[0,1\]by

r−​\(t\)=\{t,0≤t≤1−ε,1−εε​\(1−t\),1−ε≤t≤1,r^\{\-\}\(t\)=\\begin\{cases\}t,&0\\leq t\\leq 1\-\\varepsilon,\\\\\[2\.84526pt\] \\dfrac\{1\-\\varepsilon\}\{\\varepsilon\}\(1\-t\),&1\-\\varepsilon\\leq t\\leq 1,\\end\{cases\}and

r\+​\(t\)=\{1−1−εε​t,0≤t≤ε,t,ε≤t≤1\.r^\{\+\}\(t\)=\\begin\{cases\}1\-\\dfrac\{1\-\\varepsilon\}\{\\varepsilon\}t,&0\\leq t\\leq\\varepsilon,\\\\\[2\.84526pt\] t,&\\varepsilon\\leq t\\leq 1\.\\end\{cases\}Sincer−​\(0\)=r−​\(1\)=0r^\{\-\}\(0\)=r^\{\-\}\(1\)=0andr\+​\(0\)=r\+​\(1\)=1r^\{\+\}\(0\)=r^\{\+\}\(1\)=1, these functions define CPwL readoutsρ±:Γ→\[0,1\]\\rho^\{\\pm\}:\\Gamma\\to\[0,1\]throughρ±​\(E​\(t\)\):=r±​\(t\)\\rho^\{\\pm\}\(E\(t\)\):=r^\{\\pm\}\(t\)\. We fix global CPwL extensionsρ±:ℝ2→ℝ\\rho^\{\\pm\}:\\mathbb\{R\}^\{2\}\\to\\mathbb\{R\}\.

For a special hathh, define the lifted scalar readout

Hh​\(r,y\):=min⁡\{h​\(ρ−​\(r\)\),h​\(ρ\+​\(r\)\)\},\(r,y\)∈ℝ2×ℝ2\.H\_\{h\}\(r,y\):=\\min\\bigl\\\{h\(\\rho^\{\-\}\(r\)\),h\(\\rho^\{\+\}\(r\)\)\\bigr\\\},\\qquad\(r,y\)\\in\\mathbb\{R\}^\{2\}\\times\\mathbb\{R\}^\{2\}\.
###### Lemma 3\.5\(Lifted special\-hat readout\)\.

The mapHh:ℝ4→ℝH\_\{h\}:\\mathbb\{R\}^\{4\}\\to\\mathbb\{R\}is CPwL and satisfies

Hh​\(E​\(t\),y\)=h​\(t\),t∈\[0,1\],y∈ℝ2\.H\_\{h\}\(E\(t\),y\)=h\(t\),\\qquad t\\in\[0,1\],\\quad y\\in\\mathbb\{R\}^\{2\}\.Consequently, for everyx∈\[0,1\)x\\in\[0,1\)andj≥0j\\geq 0, one has

Hh​\(Zj​\(x\)\)=h​\(Rj​x\)\.H\_\{h\}\(Z\_\{j\}\(x\)\)=h\(R^\{j\}x\)\.

###### Proof\.

The mapHhH\_\{h\}is CPwL because compositions and pointwise minima of scalar CPwL functions are CPwL\. Ift∈\[ε,1−ε\]t\\in\[\\varepsilon,1\-\\varepsilon\], thenr−​\(t\)=r\+​\(t\)=tr^\{\-\}\(t\)=r^\{\+\}\(t\)=t, and the identity is immediate\.

Ift∈\[0,ε\]t\\in\[0,\\varepsilon\], thenr−​\(t\)=tr^\{\-\}\(t\)=tandh​\(t\)=0h\(t\)=0, sinceε<ϱ\\varepsilon<\\varrhoandsupp⁡h⊂\[ϱ,1−ϱ\]\\operatorname\{supp\}h\\subset\[\\varrho,1\-\\varrho\]\. The nonnegativity ofhhtherefore gives

min⁡\{h​\(r−​\(t\)\),h​\(r\+​\(t\)\)\}=0=h​\(t\)\.\\min\\bigl\\\{h\(r^\{\-\}\(t\)\),h\(r^\{\+\}\(t\)\)\\bigr\\\}=0=h\(t\)\.The same argument applies on\[1−ε,1\]\[1\-\\varepsilon,1\], usingr\+​\(t\)=tr^\{\+\}\(t\)=t\. This proves the first identity\.

The first coordinate ofZj​\(x\)Z\_\{j\}\(x\)isE​\(Mj​x\)=E​\(Rj​x\)E\(M^\{j\}x\)=E\(R^\{j\}x\)\. Substitution into the first identity therefore yields the valid\-state formula\. ∎

Using the frame\-aligned expansion \([6](https://arxiv.org/html/2607.20586#S3.E6)\), define

ℬ​\(Z\):=∑ν=1NBaν​Hhν​\(Z\)​vν,Z∈ℝ4\.\\mathcal\{B\}\(Z\):=\\sum\_\{\\nu=1\}^\{N\_\{B\}\}a\_\{\\nu\}H\_\{h\_\{\\nu\}\}\(Z\)v\_\{\\nu\},\\qquad Z\\in\\mathbb\{R\}^\{4\}\.
###### Corollary 3\.6\(Exact lifted forcing readout\)\.

The mapℬ:ℝ4→ℝp​Ls\\mathcal\{B\}:\\mathbb\{R\}^\{4\}\\to\\mathbb\{R\}^\{pL\_\{s\}\}is CPwL and, for everyx∈\[0,1\)x\\in\[0,1\)andj≥0j\\geq 0, satisfies

ℬ​\(Zj​\(x\)\)=b​\(Rj​x\)\.\\mathcal\{B\}\(Z\_\{j\}\(x\)\)=b\(R^\{j\}x\)\.

###### Proof\.

The claim follows by applying[Section3\.3](https://arxiv.org/html/2607.20586#S3.SS3)termwise in \([6](https://arxiv.org/html/2607.20586#S3.E6)\)\. ∎

### 3\.4Selectors and the selected matrix action

The digit indicators are discontinuous at theMM\-ary breakpoints and therefore cannot be represented exactly by continuous loop readouts\. We replace them by CPwL selectors that are exact outside small transition intervals\. On those intervals, exactness of the matrix action will follow from the vanishing of the accumulated Horner state\.

Fix0<δ¯<10<\\bar\{\\delta\}<1\. For a prescribed final depthn≥1n\\geq 1, setδn:=δ¯​ϱ​M−\(n\+1\)\\delta\_\{n\}:=\\bar\{\\delta\}\\,\\varrho M^\{\-\(n\+1\)\}and define the transition set

Jn:=⋃k=0M−1\[kM,kM\+δn\]\.J\_\{n\}:=\\bigcup\_\{k=0\}^\{M\-1\}\\left\[\\frac\{k\}\{M\},\\frac\{k\}\{M\}\+\\delta\_\{n\}\\right\]\.Sinceδn<1/M\\delta\_\{n\}<1/M, these intervals are pairwise disjoint and contained in\[0,1\)\[0,1\)\.

Letϑ\(n\)=\(ϑ0\(n\),…,ϑM−1\(n\)\)\\vartheta^\{\(n\)\}=\(\\vartheta^\{\(n\)\}\_\{0\},\\ldots,\\vartheta^\{\(n\)\}\_\{M\-1\}\)be the continuous piecewise affine map from\[0,1\]\[0,1\]into the standard simplex defined as follows\. On each transition interval\[k/M,k/M\+δn\]\[k/M,k/M\+\\delta\_\{n\}\], it interpolates linearly from the coordinate vector corresponding tok−1\(modM\)k\-1\\pmod\{M\}to the coordinate vector corresponding tokk\. Between successive transition intervals, it is equal to the coordinate vector corresponding to the current digit\. Thus we have

∑r=0M−1ϑr\(n\)​\(t\)=1,0≤ϑr\(n\)​\(t\)≤1,\\sum\_\{r=0\}^\{M\-1\}\\vartheta^\{\(n\)\}\_\{r\}\(t\)=1,\\qquad 0\\leq\\vartheta^\{\(n\)\}\_\{r\}\(t\)\\leq 1,and, whenevert∈\[0,1\)∖Jnt\\in\[0,1\)\\setminus J\_\{n\}, one has

ϑr\(n\)​\(t\)=\{1,r=⌊M​t⌋,0,r≠⌊M​t⌋\.\\vartheta^\{\(n\)\}\_\{r\}\(t\)=\\begin\{cases\}1,&r=\\lfloor Mt\\rfloor,\\\\ 0,&r\\neq\\lfloor Mt\\rfloor\.\\end\{cases\}At the endpoints, this construction givesϑ\(n\)​\(0\)=ϑ\(n\)​\(1\)=eM−1\\vartheta^\{\(n\)\}\(0\)=\\vartheta^\{\(n\)\}\(1\)=e\_\{M\-1\}\. Hence the selectors descend continuously to the loop through

χr\(n\)​\(E​\(t\)\):=ϑr\(n\)​\(t\)\.\\chi^\{\(n\)\}\_\{r\}\(E\(t\)\):=\\vartheta^\{\(n\)\}\_\{r\}\(t\)\.We fix global CPwL extensions toℝ2\\mathbb\{R\}^\{2\}, and then lift them to the memory space by settingχr\(n\)​\(r,y\):=χr\(n\)​\(r\)\\chi^\{\(n\)\}\_\{r\}\(r,y\):=\\chi^\{\(n\)\}\_\{r\}\(r\)\.

We use the following elementary exact gate, which is a simpler variant of the product gadget introduced in\[[3](https://arxiv.org/html/2607.20586#bib.bib3), Lemma 9\]and used in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\. Fora\>0a\>0,λ∈ℝ\\lambda\\in\\mathbb\{R\}, andY∈ℝp​LsY\\in\\mathbb\{R\}^\{pL\_\{s\}\}, define

Πa​\(λ,Y\):=ReLU⁡\(Y−a​\(1−λ\)​𝟏\)−ReLU⁡\(−Y−a​\(1−λ\)​𝟏\),\\Pi\_\{a\}\(\\lambda,Y\):=\\operatorname\{ReLU\}\\bigl\(Y\-a\(1\-\\lambda\)\\mathbf\{1\}\\bigr\)\-\\operatorname\{ReLU\}\\bigl\(\-Y\-a\(1\-\\lambda\)\\mathbf\{1\}\\bigr\),where𝟏∈ℝp​Ls\\mathbf\{1\}\\in\\mathbb\{R\}^\{pL\_\{s\}\}denotes the all\-ones vector, and all ReLU operations are applied componentwise\. Forλ∈\[0,1\]\\lambda\\in\[0,1\]andY∈\[−a,a\]p​LsY\\in\[\-a,a\]^\{pL\_\{s\}\}, one has

Πa​\(1,Y\)=Y,Πa​\(0,Y\)=0,Πa​\(λ,0\)=0\.\\Pi\_\{a\}\(1,Y\)=Y,\\qquad\\Pi\_\{a\}\(0,Y\)=0,\\qquad\\Pi\_\{a\}\(\\lambda,0\)=0\.
To choose the scale, set

B∗:=supt∈\[0,1\]‖b​\(t\)‖∞,τ:=max0≤r≤M−1⁡‖𝖳r‖∞→∞\.B\_\{\*\}:=\\sup\_\{t\\in\[0,1\]\}\\\|b\(t\)\\\|\_\{\\infty\},\\qquad\\tau:=\\max\_\{0\\leq r\\leq M\-1\}\\\|\\mathsf\{T\}\_\{r\}\\\|\_\{\\infty\\to\\infty\}\.The Horner recursion gives

‖Uj​\(x\)‖∞≤B∗​∑ℓ=0n−1−jτℓ\.\\\|U\_\{j\}\(x\)\\\|\_\{\\infty\}\\leq B\_\{\*\}\\sum\_\{\\ell=0\}^\{n\-1\-j\}\\tau^\{\\ell\}\.We may therefore take

an:=1\+τ​B∗​∑ℓ=0n−1τℓ,a\_\{n\}:=1\+\\tau B\_\{\*\}\\sum\_\{\\ell=0\}^\{n\-1\}\\tau^\{\\ell\},so that‖𝖳r​Uj​\(x\)‖∞≤an\\\|\\mathsf\{T\}\_\{r\}U\_\{j\}\(x\)\\\|\_\{\\infty\}\\leq a\_\{n\}for everyr,jr,j, andxx\. In particular, there are constantsC,Λ\>0C,\\Lambda\>0, depending only on the fixed data, such thatan≤C​Λna\_\{n\}\\leq C\\Lambda^\{n\}\.

Define the selected matrix action by

𝒯n​\(Z,U\):=∑r=0M−1Πan​\(χr\(n\)​\(Z\),𝖳r​U\)\.\\mathcal\{T\}\_\{n\}\(Z,U\):=\\sum\_\{r=0\}^\{M\-1\}\\Pi\_\{a\_\{n\}\}\\bigl\(\\chi^\{\(n\)\}\_\{r\}\(Z\),\\mathsf\{T\}\_\{r\}U\\bigr\)\.
###### Lemma 3\.7\(Exact selected matrix action on Horner states\)\.

For everyj=0,…,n−2j=0,\\ldots,n\-2andx∈\[0,1\)x\\in\[0,1\), one has

𝒯n​\(Zj​\(x\),Uj\+1​\(x\)\)=𝖳dj\+1​\(x\)​Uj\+1​\(x\)\.\\mathcal\{T\}\_\{n\}\\bigl\(Z\_\{j\}\(x\),U\_\{j\+1\}\(x\)\\bigr\)=\\mathsf\{T\}\_\{d\_\{j\+1\}\(x\)\}U\_\{j\+1\}\(x\)\.

###### Proof\.

Sett:=Rj​xt:=R^\{j\}x\. Ift∉Jnt\\notin J\_\{n\}, then the selectors are exact anddj\+1​\(x\)=⌊M​t⌋d\_\{j\+1\}\(x\)=\\lfloor Mt\\rfloor\. Exactly one selector equals11, while all the others vanish, so the gating identities give

𝒯n​\(Zj​\(x\),Uj\+1​\(x\)\)=𝖳dj\+1​\(x\)​Uj\+1​\(x\)\.\\mathcal\{T\}\_\{n\}\\bigl\(Z\_\{j\}\(x\),U\_\{j\+1\}\(x\)\\bigr\)=\\mathsf\{T\}\_\{d\_\{j\+1\}\(x\)\}U\_\{j\+1\}\(x\)\.
Suppose now thatt∈Jnt\\in J\_\{n\}\. Thent=k/M\+δt=k/M\+\\deltafor somek∈\{0,…,M−1\}k\\in\\\{0,\\ldots,M\-1\\\}and0≤δ≤δn0\\leq\\delta\\leq\\delta\_\{n\}\. Sinceδn<1/M\\delta\_\{n\}<1/M, the first residual isR​\(t\)=M​δR\(t\)=M\\delta\. Moreover, for everyi=j\+1,…,n−1i=j\+1,\\ldots,n\-1, no wrap\-around occurs and

Ri​x=Mi−j​δ≤Mn−1​δn=δ¯​ϱ​M−2<ϱ\.R^\{i\}x=M^\{\\,i\-j\}\\delta\\leq M^\{n\-1\}\\delta\_\{n\}=\\bar\{\\delta\}\\,\\varrho M^\{\-2\}<\\varrho\.Every special hat in \([6](https://arxiv.org/html/2607.20586#S3.E6)\) vanishes on\[0,ϱ\)\[0,\\varrho\), and henceb​\(Ri​x\)=0b\(R^\{i\}x\)=0fori=j\+1,…,n−1i=j\+1,\\ldots,n\-1\. By[Section3\.1](https://arxiv.org/html/2607.20586#S3.SS1), it follows thatUj\+1​\(x\)=0U\_\{j\+1\}\(x\)=0\.

Consequently,𝖳r​Uj\+1​\(x\)=0\\mathsf\{T\}\_\{r\}U\_\{j\+1\}\(x\)=0for every branchrr\. The identityΠa​\(λ,0\)=0\\Pi\_\{a\}\(\\lambda,0\)=0then gives

𝒯n​\(Zj​\(x\),Uj\+1​\(x\)\)=0=𝖳dj\+1​\(x\)​Uj\+1​\(x\),\\mathcal\{T\}\_\{n\}\\bigl\(Z\_\{j\}\(x\),U\_\{j\+1\}\(x\)\\bigr\)=0=\\mathsf\{T\}\_\{d\_\{j\+1\}\(x\)\}U\_\{j\+1\}\(x\),which completes the proof\. ∎

### 3\.5The backward memory\-Horner network

We now combine exact backward replay, the lifted forcing readout, and the selected matrix action\. First compute the terminal memory stateZn​\(x\)=𝒬n​\(Z0​\(x\)\)Z\_\{n\}\(x\)=\\mathcal\{Q\}^\{n\}\(Z\_\{0\}\(x\)\), and initialize

Z^n−1​\(x\):=P​\(Zn​\(x\)\),U^n−1​\(x\):=ℬ​\(Z^n−1​\(x\)\)\.\\widehat\{Z\}\_\{n\-1\}\(x\):=P\(Z\_\{n\}\(x\)\),\\qquad\\widehat\{U\}\_\{n\-1\}\(x\):=\\mathcal\{B\}\(\\widehat\{Z\}\_\{n\-1\}\(x\)\)\.Forj=n−2,…,0j=n\-2,\\ldots,0, define the backward recursion by

Z^j​\(x\):=P​\(Z^j\+1​\(x\)\),U^j​\(x\):=ℬ​\(Z^j​\(x\)\)\+𝒯n​\(Z^j​\(x\),U^j\+1​\(x\)\)\.\\widehat\{Z\}\_\{j\}\(x\):=P\(\\widehat\{Z\}\_\{j\+1\}\(x\)\),\\qquad\\widehat\{U\}\_\{j\}\(x\):=\\mathcal\{B\}\(\\widehat\{Z\}\_\{j\}\(x\)\)\+\\mathcal\{T\}\_\{n\}\(\\widehat\{Z\}\_\{j\}\(x\),\\widehat\{U\}\_\{j\+1\}\(x\)\)\.
###### Lemma 3\.8\(Exactness of the backward memory\-Horner recursion\)\.

For everyj=0,…,n−1j=0,\\ldots,n\-1andx∈\[0,1\)x\\in\[0,1\), one has

Z^j​\(x\)=Zj​\(x\),U^j​\(x\)=Uj​\(x\)\.\\widehat\{Z\}\_\{j\}\(x\)=Z\_\{j\}\(x\),\\qquad\\widehat\{U\}\_\{j\}\(x\)=U\_\{j\}\(x\)\.In particular,U^0​\(x\)=Vecs⁡\(Sn\)​\(x\)\\widehat\{U\}\_\{0\}\(x\)=\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\(x\)\.

###### Proof\.

By exact backward replay, we haveZ^n−1​\(x\)=P​\(Zn​\(x\)\)=Zn−1​\(x\)\\widehat\{Z\}\_\{n\-1\}\(x\)=P\(Z\_\{n\}\(x\)\)=Z\_\{n\-1\}\(x\)\. The lifted forcing readout then gives

U^n−1​\(x\)=ℬ​\(Zn−1​\(x\)\)=b​\(Rn−1​x\)=Un−1​\(x\)\.\\widehat\{U\}\_\{n\-1\}\(x\)=\\mathcal\{B\}\(Z\_\{n\-1\}\(x\)\)=b\(R^\{n\-1\}x\)=U\_\{n\-1\}\(x\)\.
Suppose that the two identities hold at levelj\+1j\+1\. Exact backward replay givesZ^j​\(x\)=P​\(Zj\+1​\(x\)\)=Zj​\(x\)\\widehat\{Z\}\_\{j\}\(x\)=P\(Z\_\{j\+1\}\(x\)\)=Z\_\{j\}\(x\)\. Using[Sections3\.3](https://arxiv.org/html/2607.20586#S3.SS3)and[3\.4](https://arxiv.org/html/2607.20586#S3.SS4), we therefore obtain

U^j​\(x\)\\displaystyle\\widehat\{U\}\_\{j\}\(x\)=ℬ​\(Zj​\(x\)\)\+𝒯n​\(Zj​\(x\),Uj\+1​\(x\)\)\\displaystyle=\\mathcal\{B\}\(Z\_\{j\}\(x\)\)\+\\mathcal\{T\}\_\{n\}\(Z\_\{j\}\(x\),U\_\{j\+1\}\(x\)\)=b​\(Rj​x\)\+𝖳dj\+1​\(x\)​Uj\+1​\(x\)=Uj​\(x\)\.\\displaystyle=b\(R^\{j\}x\)\+\\mathsf\{T\}\_\{d\_\{j\+1\}\(x\)\}U\_\{j\+1\}\(x\)=U\_\{j\}\(x\)\.Backward induction proves the claim, and the final identity follows from[Section3\.1](https://arxiv.org/html/2607.20586#S3.SS1)\. ∎

###### Theorem 3\.9\(Vectorized forcing realization in one admissible frame\)\.

Fix an admissibless\-offset frame, and letB:ℝ→ℝpB:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}be CPwL and supported in\[0,L\]\[0,L\]\. Suppose thatb=Vecs⁡\(B\)b=\\operatorname\{Vec\}\_\{s\}\(B\)has a frame\-aligned expansion of the form \([6](https://arxiv.org/html/2607.20586#S3.E6)\)\. SetSn:=∑r=0n−1Vr​BS\_\{n\}:=\\sum\_\{r=0\}^\{n\-1\}V^\{r\}B\. Then there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnn, and networksΦn∈ΥC0,C1​n​\(ReLU;1,p​Ls\)\\Phi\_\{n\}\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,pL\_\{s\}\)such that

Φn​\(x\)=Vecs⁡\(Sn\)​\(x\),x∈\[0,1\]\.\\Phi\_\{n\}\(x\)=\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\(x\),\\qquad x\\in\[0,1\]\.The realizing networks may moreover be chosen with weights and biases growing at most exponentially innn\.

###### Proof\.

Fix a global CPwL extension of the mapx↦\(E​\(x\),0\)x\\mapsto\(E\(x\),0\)from\[0,1\]\[0,1\]toℝ\\mathbb\{R\}\. The input is first mapped toZ0​\(x\)=\(E​\(x\),0\)Z\_\{0\}\(x\)=\(E\(x\),0\), after whichnncopies of the fixed memory update𝒬\\mathcal\{Q\}produceZn​\(x\)Z\_\{n\}\(x\)\. The backward computation usesnnstage blocks formed fromPP,ℬ\\mathcal\{B\}, and𝒯n\\mathcal\{T\}\_\{n\}\. All state dimensions are fixed: the memory state lies inℝ4\\mathbb\{R\}^\{4\}, while the accumulated affine state lies inℝp​Ls\\mathbb\{R\}^\{pL\_\{s\}\}\.

The maps𝒬\\mathcal\{Q\},PP, andℬ\\mathcal\{B\}are fixed CPwL maps on fixed\-dimensional spaces and hence have exact ReLU realizations of fixed width and depth\. The selected\-action map𝒯n\\mathcal\{T\}\_\{n\}also has an architecture independent ofnn: the number of selectors and their affine pieces is fixed, and the product gadget has fixed size\. Itsnn\-dependence enters only through the selector slopes, which are proportional toδn−1\\delta\_\{n\}^\{\-1\}, and through the scaleana\_\{n\}\. Sinceδn−1=O​\(Mn\)\\delta\_\{n\}^\{\-1\}=O\(M^\{n\}\)andan≤C​Λna\_\{n\}\\leq C\\Lambda^\{n\}, these parameters grow at most exponentially innn\.

By[Section3\.5](https://arxiv.org/html/2607.20586#S3.SS5), the outputΦn\\Phi\_\{n\}of the backward recursion satisfiesΦn​\(x\)=Vecs⁡\(Sn\)​\(x\)\\Phi\_\{n\}\(x\)=\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\(x\)forx∈\[0,1\)x\\in\[0,1\)\. Both the network output andVecs⁡\(Sn\)\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)are continuous on\[0,1\]\[0,1\], so the equality also holds atx=1x=1\. The forward and backward phases useO​\(n\)O\(n\)fixed\-size blocks in total\. Consequently, the width is bounded independently ofnn, the depth isO​\(n\)O\(n\), and the parameter bound is at most exponential\. ∎

### 3\.6From vectorized blocks to the global forcing curve

LetSn,k:\[0,1\]→ℝpS\_\{n,k\}:\[0,1\]\\to\\mathbb\{R\}^\{p\},k∈Jsk\\in J\_\{s\}, denote the blocks ofVecs⁡\(Sn\)\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\. Then, on the translated interval associated withkk, one has

Sn​\(t\)=Sn,k​\(t−s−k\+1\),t∈\[s\+k−1,s\+k\]\.S\_\{n\}\(t\)=S\_\{n,k\}\(t\-s\-k\+1\),\\qquad t\\in\[s\+k\-1,s\+k\]\.BecauseSnS\_\{n\}is continuous, the neighboring block formulas agree at their common endpoints\. These finitely many block realizations can be assembled into a global realization by the standard CPwL gluing construction used in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\.

###### Corollary 3\.10\(Fixed\-frame forcing realization onℝ\\mathbb\{R\}\)\.

Under the assumptions of[Theorem3\.9](https://arxiv.org/html/2607.20586#S3.Thmtheorem9), there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnn, such that

Sn=∑r=0n−1Vr​B∈ΥC0,C1​n​\(ReLU;1,p\),n≥1\.S\_\{n\}=\\sum\_\{r=0\}^\{n\-1\}V^\{r\}B\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,p\),\\qquad n\\geq 1\.These networks may again be chosen with weights and biases growing at most exponentially innn\.

###### Proof\.

By[Theorem3\.9](https://arxiv.org/html/2607.20586#S3.Thmtheorem9), all blocksSn,kS\_\{n,k\}are realized simultaneously by a fixed\-width, depth\-O​\(n\)O\(n\)network\. For eachk∈Jsk\\in J\_\{s\}, precomposition with the affine mapt↦t−s−k\+1t\\mapsto t\-s\-k\+1produces the corresponding realization on\[s\+k−1,s\+k\]\[s\+k\-1,s\+k\]\.

The block\-to\-global gluing construction of\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]combines these finitely many translated outputs and sets the result to zero outside the support window\. Since the numberLsL\_\{s\}of blocks is fixed, this requires only a fixed enlargement of the width and anO​\(1\)O\(1\)increase in depth\. The resulting network therefore has fixed width and depthO​\(n\)O\(n\), while its parameter bound remains exponential after adjusting the constants\. ∎

## 4General forcing and the affine theorem

Section[3](https://arxiv.org/html/2607.20586#S3)treated forcing terms aligned with a single admissible frame\. ForM≥3M\\geq 3, we now combine the ordinary frame with a nontrivial admissible offset frame and decompose arbitrary compactly supported CPwL forcing into atoms aligned with one of the two\. The one\-frame theorem then applies atom by atom, and combining the resulting forcing realization with[Theorem2\.5](https://arxiv.org/html/2607.20586#S2.Thmtheorem5)yields the main affine theorem\.

Throughout this section, fixs∈\(0,1\)s\\in\(0,1\)such that\(M−1\)​s∈ℤ\(M\-1\)s\\in\\mathbb\{Z\}, and chooseϱ\>0\\varrho\>0so that2​ϱ<min⁡\{s,1−s\}2\\varrho<\\min\\\{s,1\-s\\\}\. Thus both the ordinary frame and thess\-offset frame are admissible\.

### 4\.1Two\-frame special\-hat decomposition

The following nodal decomposition expresses arbitrary compactly supported CPwL data as a finite sum of special hats whose translations lie inΛs:=ℤ∪\(ℤ\+s\)\\Lambda\_\{s\}:=\\mathbb\{Z\}\\cup\(\\mathbb\{Z\}\+s\), and hence are aligned with one of the two admissible frames\.

###### Lemma 4\.1\(Two\-frame special\-hat decomposition\)\.

LetF:ℝ→ℝpF:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}be compactly supported and CPwL\. ThenFFadmits a representation

F​\(t\)=∑ν=1N∗aν​hν​\(t−δν\)​eμν,F\(t\)=\\sum\_\{\\nu=1\}^\{N\_\{\*\}\}a\_\{\\nu\}h\_\{\\nu\}\(t\-\\delta\_\{\\nu\}\)e\_\{\\mu\_\{\\nu\}\},whereaν∈ℝa\_\{\\nu\}\\in\\mathbb\{R\},δν∈Λs\\delta\_\{\\nu\}\\in\\Lambda\_\{s\},μν∈\{1,…,p\}\\mu\_\{\\nu\}\\in\\\{1,\\ldots,p\\\}, and eachhνh\_\{\\nu\}is a special hat with at most three breakpoints\.

More precisely, suppose thatsupp⁡F⊂\[a,b\]\\operatorname\{supp\}F\\subset\[a,b\], witha<ba<b, and that the union of the breakpoint sets of its scalar components contains at mostmmpoints in\(a,b\)\(a,b\)\. Settingη:=min⁡\{s,1−s\}−2​ϱ\>0\\eta:=\\min\\\{s,1\-s\\\}\-2\\varrho\>0, one may choose the decomposition so thatN∗≤p​\(⌊2​\(b−a\)/η⌋\+m\+2\)N\_\{\*\}\\leq p\\bigl\(\\lfloor 2\(b\-a\)/\\eta\\rfloor\+m\+2\\bigr\)\.

###### Proof\.

Choose a common partition of\[a,b\]\[a,b\]containing all breakpoints of the components ofFF\. If the lengths of its original intervals aredid\_\{i\}, subdivide theii\-th interval into⌊2​di/η⌋\+1\\lfloor 2d\_\{i\}/\\eta\\rfloor\+1equal pieces\. Every resulting mesh interval then has length strictly less thanη/2\\eta/2, and the total number of mesh nodes is at most

⌊2​\(b−a\)η⌋\+m\+2\.\\left\\lfloor\\frac\{2\(b\-a\)\}\{\\eta\}\\right\\rfloor\+m\+2\.On this refined grid,FFhas the vector\-valued nodal expansionF​\(t\)=∑iψi​\(t\)​F​\(ti\)F\(t\)=\\sum\_\{i\}\\psi\_\{i\}\(t\)F\(t\_\{i\}\), where eachψi\\psi\_\{i\}is a nonnegative nodal hat\. Splitting the vectorsF​\(ti\)F\(t\_\{i\}\)into their coordinate components gives at mostppscalar\-coordinate terms per mesh node\.

Fix one of the nodal hatsψ\\psi, and writesupp⁡ψ=\[aψ,bψ\]\\operatorname\{supp\}\\psi=\[a\_\{\\psi\},b\_\{\\psi\}\]\. Since an interior nodal hat spans at most two adjacent mesh intervals, its support has lengthbψ−aψ<ηb\_\{\\psi\}\-a\_\{\\psi\}<\\eta\. The translationsδ\\deltafor whichh​\(t\):=ψ​\(t\+δ\)h\(t\):=\\psi\(t\+\\delta\)satisfiessupp⁡h⊂\[ϱ,1−ϱ\]\\operatorname\{supp\}h\\subset\[\\varrho,1\-\\varrho\]form the interval

Iψ:=\[bψ\+ϱ−1,aψ−ϱ\]\.I\_\{\\psi\}:=\[\\,b\_\{\\psi\}\+\\varrho\-1,\\ a\_\{\\psi\}\-\\varrho\\,\]\.Its length satisfies

\|Iψ\|=1−\(bψ−aψ\)−2​ϱ\>max⁡\{s,1−s\}\.\|I\_\{\\psi\}\|=1\-\(b\_\{\\psi\}\-a\_\{\\psi\}\)\-2\\varrho\>\\max\\\{s,1\-s\\\}\.The successive gaps inΛs=ℤ∪\(ℤ\+s\)\\Lambda\_\{s\}=\\mathbb\{Z\}\\cup\(\\mathbb\{Z\}\+s\)have lengthsssand1−s1\-s\. HenceIψI\_\{\\psi\}contains someδ∈Λs\\delta\\in\\Lambda\_\{s\}\. For this choice,hhis a special hat andψ​\(t\)=h​\(t−δ\)\\psi\(t\)=h\(t\-\\delta\)\. Sinceψ\\psiis a nodal hat,hhhas at most three breakpoints\. Applying this construction to every nodal term and coordinate yields the asserted decomposition and the stated bound onN∗N\_\{\*\}\. ∎

### 4\.2General compactly supported forcing

We now treat arbitrary compactly supported CPwL forcing terms\. The two\-frame decomposition aligns each forcing atom with either the ordinary frame or the fixed admissibless\-offset frame\.

###### Theorem 4\.2\(General compactly supported forcing\)\.

LetB:ℝ→ℝpB:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}be CPwL and supported in\[0,L\]\[0,L\], and suppose that the union of the breakpoint sets of its scalar components contains at mostmmpoints in\(0,L\)\(0,L\)\. For the admissible offsetssfixed above, there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnn, such that

Sn​\[B\]:=∑r=0n−1Vr​B∈ΥC0,C1​n​\(ReLU;1,p\),n≥1\.S\_\{n\}\[B\]:=\\sum\_\{r=0\}^\{n\-1\}V^\{r\}B\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,p\),\\qquad n\\geq 1\.The architectural constants depend only onM,ϱ,p,L,m,sM,\\varrho,p,L,m,sand the fixed mask\(Aj\)j\(A\_\{j\}\)\_\{j\}\. The realizing networks may moreover be chosen with weights and biases bounded byCB​ΛnC\_\{B\}\\Lambda^\{n\}, whereΛ\>0\\Lambda\>0depends only on the fixed refinement and frame data, whileCB\>0C\_\{B\}\>0may also depend on the CPwL coefficients ofBB\.

###### Proof\.

By[Section4\.1](https://arxiv.org/html/2607.20586#S4.SS1), the forcing term has a representation

B​\(t\)=∑ν=1N∗aν​hν​\(t−δν\)​eμν,B\(t\)=\\sum\_\{\\nu=1\}^\{N\_\{\*\}\}a\_\{\\nu\}h\_\{\\nu\}\(t\-\\delta\_\{\\nu\}\)e\_\{\\mu\_\{\\nu\}\},whereaν∈ℝa\_\{\\nu\}\\in\\mathbb\{R\},δν∈ℤ∪\(ℤ\+s\)\\delta\_\{\\nu\}\\in\\mathbb\{Z\}\\cup\(\\mathbb\{Z\}\+s\),μν∈\{1,…,p\}\\mu\_\{\\nu\}\\in\\\{1,\\ldots,p\\\}, and eachhνh\_\{\\nu\}is a special hat with at most three breakpoints\. The numberN∗N\_\{\*\}is bounded in terms ofp,L,m,sp,L,m,s, andϱ\\varrho, independently ofnn\.

The forcing sum is linear inBB, and hence we have

Sn\[B\]=∑ν=1N∗aνSn\[hν\(⋅−δν\)eμν\]\.S\_\{n\}\[B\]=\\sum\_\{\\nu=1\}^\{N\_\{\*\}\}a\_\{\\nu\}S\_\{n\}\\\!\\left\[h\_\{\\nu\}\(\\,\\cdot\-\\delta\_\{\\nu\}\)e\_\{\\mu\_\{\\nu\}\}\\right\]\.Ifδν∈ℤ\\delta\_\{\\nu\}\\in\\mathbb\{Z\}, the corresponding atom is aligned with the ordinary frame; ifδν∈ℤ\+s\\delta\_\{\\nu\}\\in\\mathbb\{Z\}\+s, it is aligned with thess\-offset frame\. Both frames are admissible, so[Section3\.6](https://arxiv.org/html/2607.20586#S3.SS6)applies to every summand\.

The number of summands is bounded independently ofnn\. Their realizing networks can therefore be run in parallel and combined by a final affine layer\. This enlarges the width by only a fixed factor and preserves depthO​\(n\)O\(n\)\. The coefficientsaνa\_\{\\nu\}and the CPwL data of the special hats affect only the network parameters, so the same finite combination gives a bound of the formCB​ΛnC\_\{B\}\\Lambda^\{n\}\. ∎

### 4\.3The affine main theorem

We now combine the general forcing theorem with the homogeneous realization theorem imported from\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\.

###### Theorem 4\.3\(Affine realization theorem\)\.

LetM≥3M\\geq 3, and consider the affine refinement operator

\(W​γ\)​\(t\)=∑j∈ℤAj​γ​\(M​t−j\)\+B​\(t\),\(W\\gamma\)\(t\)=\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}\\gamma\(Mt\-j\)\+B\(t\),where the matrix mask\(Aj\)j∈ℤ⊂ℝp×p\(A\_\{j\}\)\_\{j\\in\\mathbb\{Z\}\}\\subset\\mathbb\{R\}^\{p\\times p\}is finitely supported and the associated homogeneous operatorVVpreserves the support window\[0,L\]\[0,L\]\. Suppose thatγ,B:ℝ→ℝp\\gamma,B:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}are CPwL and supported in\[0,L\]\[0,L\]\. Then there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnn, such that

Wn​γ∈ΥC0,C1​n​\(ReLU;1,p\),n≥1\.W^\{n\}\\gamma\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,p\),\\qquad n\\geq 1\.The constants depend only on the fixed refinement data and the CPwL complexity ofγ\\gammaandBB\. The realizing networks may moreover be chosen with weights and biases bounded byC​ΛnC\\Lambda^\{n\}for suitable constantsC,Λ\>0C,\\Lambda\>0\.

###### Proof\.

With the admissible offset and support margin fixed above, the affine iterate identity gives

Wn​γ=Vn​γ\+Sn​\[B\]\.W^\{n\}\\gamma=V^\{n\}\\gamma\+S\_\{n\}\[B\]\.The homogeneous theorem[Theorem2\.5](https://arxiv.org/html/2607.20586#S2.Thmtheorem5)gives a fixed\-width, depth\-O​\(n\)O\(n\)realization ofVn​γV^\{n\}\\gamma, while[Theorem4\.2](https://arxiv.org/html/2607.20586#S4.Thmtheorem2)gives one forSn​\[B\]S\_\{n\}\[B\]\. Running the two networks in parallel and adding their outputs in a final affine layer preserves fixed width up to a constant factor and preserves depthO​\(n\)O\(n\)\. The exponential parameter bound is preserved after enlarging the constants\. ∎

###### Definition 4\.4\(Ordinary\-frame seam\-separated forcing\)\.

A forcing termB:ℝ→ℝpB:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}supported in\[0,L\]\[0,L\]is called*ordinary\-frame seam\-separated*if there existϱ∈\(0,12\)\\varrho\\in\(0,\\frac\{1\}\{2\}\), scalarsaνa\_\{\\nu\}, vectorsvν∈ℝp​Lv\_\{\\nu\}\\in\\mathbb\{R\}^\{pL\}, and nonnegative CPwL functionshνh\_\{\\nu\}such that

Vec0⁡\(B\)​\(x\)=∑ν=1NBaν​hν​\(x\)​vν,supp⁡hν⊂\[ϱ,1−ϱ\]\.\\operatorname\{Vec\}\_\{0\}\(B\)\(x\)=\\sum\_\{\\nu=1\}^\{N\_\{B\}\}a\_\{\\nu\}h\_\{\\nu\}\(x\)v\_\{\\nu\},\\qquad\\operatorname\{supp\}h\_\{\\nu\}\\subset\[\\varrho,1\-\\varrho\]\.

###### Corollary 4\.5\(Restricted binary affine realization\)\.

LetM=2M=2, let the matrix mask be finitely supported, and assume that the associated homogeneous operatorVVpreserves\[0,L\]\[0,L\]\. Letγ,B:ℝ→ℝp\\gamma,B:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}be CPwL and supported in\[0,L\]\[0,L\]\. IfBBis ordinary\-frame seam\-separated, then there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnn, such that

Wn​γ∈ΥC0,C1​n​\(ReLU;1,p\),n≥1\.W^\{n\}\\gamma\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,p\),\\qquad n\\geq 1\.The realizing networks may be chosen with weights and biases growing at most exponentially innn\.

###### Proof\.

Choose a marginϱ\\varrhoand a representation as in[Section4\.3](https://arxiv.org/html/2607.20586#S4.SS3)\. Apply[Section3\.6](https://arxiv.org/html/2607.20586#S3.SS6)in the ordinary binary frame using this margin, and apply[Theorem2\.5](https://arxiv.org/html/2607.20586#S2.Thmtheorem5)to the homogeneous term\. Combining the two realizations as in the proof of[Theorem4\.3](https://arxiv.org/html/2607.20586#S4.Thmtheorem3)gives the result\. ∎

The memory controller supplies the residual states in the reverse order required by the affine recursion\. The offset\-frame decomposition supplies forcing readouts away from the seams, while the remaining selector ambiguity occurs only where the accumulated Horner state has already vanished\. ForM≥3M\\geq 3, two admissible frames suffice for arbitrary compactly supported CPwL forcing\. ForM=2M=2, the present offset\-frame argument yields only the restricted class above, although some important seam\-touching systems admit separate linear\-depth constructions\.

## 5Stage\-dependent forcing in a fixed finite\-dimensional span

The stage\-dependent refinement framework and its iterate formula were considered in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\. Here we record the improvement supplied by the residual memory controller: when the forcing terms belong to a fixed finite\-dimensional CPwL span, their accumulated contribution admits an exact fixed\-width realization of depthO​\(n\)O\(n\), rather than the quadratic\-depth realization obtained by evaluating the individual summands separately\.

Throughout the section, fix a nontrivial admissible offsets∈\(0,1\)s\\in\(0,1\)and chooseϱ\>0\\varrho\>0so that2​ϱ<min⁡\{s,1−s\}2\\varrho<\\min\\\{s,1\-s\\\}\. Thus both the ordinary frame and thess\-offset frame are admissible\.

#### The stage\-dependent Horner recursion\.

Consider the nonstationary affine recursion

γr\+1=V​γr\+Br,r≥0,\\gamma\_\{r\+1\}=V\\gamma\_\{r\}\+B\_\{r\},\\qquad r\\geq 0,where everyBr:ℝ→ℝpB\_\{r\}:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}is CPwL and supported in\[0,L\]\[0,L\]\. Itsnn\-th iterate satisfies

γn=Vn​γ0\+Sn,Sn:=∑r=0n−1Vn−1−r​Br\.\\gamma\_\{n\}=V^\{n\}\\gamma\_\{0\}\+S\_\{n\},\\qquad S\_\{n\}:=\\sum\_\{r=0\}^\{n\-1\}V^\{\\,n\-1\-r\}B\_\{r\}\.This is the standard stage\-dependent iterate formula recalled from\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\. We concentrate on the forcing contributionSnS\_\{n\}\.

First work in one admissible frame, and writebr:=Vecs⁡\(Br\)b\_\{r\}:=\\operatorname\{Vec\}\_\{s\}\(B\_\{r\}\)\. The cascade identity gives

Vecs⁡\(Sn\)​\(x\)=∑k=0n−1𝖳d1​\(x\)​⋯​𝖳dk​\(x\)​bn−1−k​\(Rk​x\),x∈\[0,1\),\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\(x\)=\\sum\_\{k=0\}^\{n\-1\}\\mathsf\{T\}\_\{d\_\{1\}\(x\)\}\\cdots\\mathsf\{T\}\_\{d\_\{k\}\(x\)\}\\,b\_\{n\-1\-k\}\(R^\{k\}x\),\\qquad x\\in\[0,1\),where the empty matrix product is the identity\. This sum has the backward Horner form obtained by settingUn−1​\(x\):=b0​\(Rn−1​x\)U\_\{n\-1\}\(x\):=b\_\{0\}\(R^\{n\-1\}x\)and, forj=n−2,…,0j=n\-2,\\ldots,0, defining

Uj​\(x\):=bn−1−j​\(Rj​x\)\+𝖳dj\+1​\(x\)​Uj\+1​\(x\)\.U\_\{j\}\(x\):=b\_\{n\-1\-j\}\(R^\{j\}x\)\+\\mathsf\{T\}\_\{d\_\{j\+1\}\(x\)\}U\_\{j\+1\}\(x\)\.Expanding the recursion shows thatU0​\(x\)=Vecs⁡\(Sn\)​\(x\)U\_\{0\}\(x\)=\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\(x\)\.

The same backward\-induction argument as in[Section3\.1](https://arxiv.org/html/2607.20586#S3.SS1)shows that ifbn−1−i​\(Ri​x\)=0b\_\{n\-1\-i\}\(R^\{i\}x\)=0for everyi=j\+1,…,n−1i=j\+1,\\ldots,n\-1, thenUj\+1​\(x\)=0U\_\{j\+1\}\(x\)=0\.

### 5\.1Fixed\-span forcing in one admissible frame

We first assume that the vectorized forcing terms belong to a fixed special\-hat span in one admissible frame\. Thus, for fixed special hatshνh\_\{\\nu\}, vectorsvν∈ℝp​Lsv\_\{\\nu\}\\in\\mathbb\{R\}^\{pL\_\{s\}\}, and scalarsaνa\_\{\\nu\}, we suppose that

br​\(x\)=∑ν=1NBλr,ν​aν​hν​\(x\)​vν,r≥0,b\_\{r\}\(x\)=\\sum\_\{\\nu=1\}^\{N\_\{B\}\}\\lambda\_\{r,\\nu\}a\_\{\\nu\}h\_\{\\nu\}\(x\)v\_\{\\nu\},\\qquad r\\geq 0,\(11\)wheresupp⁡hν⊂\[ϱ,1−ϱ\]\\operatorname\{supp\}h\_\{\\nu\}\\subset\[\\varrho,1\-\\varrho\]\. For a final depthnn, set

Kn:=max0≤r<n1≤ν≤NB⁡\|λr,ν\|\.K\_\{n\}:=\\max\_\{\\begin\{subarray\}\{c\}0\\leq r<n\\\\ 1\\leq\\nu\\leq N\_\{B\}\\end\{subarray\}\}\|\\lambda\_\{r,\\nu\}\|\.
For each stagerr, define the lifted forcing readout by

ℬr​\(Z\):=∑ν=1NBλr,ν​aν​Hhν​\(Z\)​vν\.\\mathcal\{B\}\_\{r\}\(Z\):=\\sum\_\{\\nu=1\}^\{N\_\{B\}\}\\lambda\_\{r,\\nu\}a\_\{\\nu\}H\_\{h\_\{\\nu\}\}\(Z\)v\_\{\\nu\}\.By[Section3\.3](https://arxiv.org/html/2607.20586#S3.SS3), it satisfies

ℬr​\(Zj​\(x\)\)=br​\(Rj​x\)\\mathcal\{B\}\_\{r\}\(Z\_\{j\}\(x\)\)=b\_\{r\}\(R^\{j\}x\)for every valid memory state\.

###### Theorem 5\.1\(One\-frame fixed\-span stage\-dependent forcing\)\.

Assume that the vectorized forcing profiles satisfy \([11](https://arxiv.org/html/2607.20586#S5.E11)\) in one admissible frame\. Then there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnnand of the coefficientsλr,ν\\lambda\_\{r,\\nu\}, such that

Sn:=∑r=0n−1Vn−1−r​Br∈ΥC0,C1​n​\(ReLU;1,p\)\.S\_\{n\}:=\\sum\_\{r=0\}^\{n\-1\}V^\{\\,n\-1\-r\}B\_\{r\}\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,p\)\.The realizing networks may moreover be chosen with weights and biases bounded byC2​\(1\+Kn\)​ΛnC\_\{2\}\(1\+K\_\{n\}\)\\Lambda^\{n\}, whereC2,Λ\>0C\_\{2\},\\Lambda\>0depend only on the fixed frame, the matrix mask, and the fixed special\-hat span\.

###### Proof\.

Set

B∗:=∑ν=1NB\|aν\|​‖hν‖L∞​\(0,1\)​‖vν‖∞B\_\{\*\}:=\\sum\_\{\\nu=1\}^\{N\_\{B\}\}\|a\_\{\\nu\}\|\\,\\\|h\_\{\\nu\}\\\|\_\{L^\{\\infty\}\(0,1\)\}\\\|v\_\{\\nu\}\\\|\_\{\\infty\}and

τ:=max0≤r≤M−1⁡‖𝖳r‖∞→∞\.\\tau:=\\max\_\{0\\leq r\\leq M\-1\}\\\|\\mathsf\{T\}\_\{r\}\\\|\_\{\\infty\\to\\infty\}\.Then‖br‖L∞≤B∗​Kn\\\|b\_\{r\}\\\|\_\{L^\{\\infty\}\}\\leq B\_\{\*\}K\_\{n\}, and the stage\-dependent Horner recursion gives

‖Uj​\(x\)‖∞≤B∗​Kn​∑ℓ=0n−1−jτℓ\.\\\|U\_\{j\}\(x\)\\\|\_\{\\infty\}\\leq B\_\{\*\}K\_\{n\}\\sum\_\{\\ell=0\}^\{n\-1\-j\}\\tau^\{\\ell\}\.Consequently, the product\-gadget scale may be chosen so that

an≤C​\(1\+Kn\)​Λna\_\{n\}\\leq C\(1\+K\_\{n\}\)\\Lambda^\{n\}for constants depending only on the fixed data\.

Use the same selectors and selected\-action map as in §[3\.4](https://arxiv.org/html/2607.20586#S3.SS4), with this enlarged scale\. Outside the selector transition set, the correct branch is selected exactly\. On the transition set, all later residuals lie in\[0,ϱ\)\[0,\\varrho\)\. Since every fixed hathνh\_\{\\nu\}vanishes there, one hasbn−1−i​\(Ri​x\)=0b\_\{n\-1\-i\}\(R^\{i\}x\)=0for every later stageii, independently of the coefficientsλr,ν\\lambda\_\{r,\\nu\}\. The vanishing\-state observation above therefore givesUj\+1​\(x\)=0U\_\{j\+1\}\(x\)=0, and the identityΠa​\(λ,0\)=0\\Pi\_\{a\}\(\\lambda,0\)=0makes the selected matrix action exact\.

ComputeZn​\(x\)=𝒬n​\(Z0​\(x\)\)Z\_\{n\}\(x\)=\\mathcal\{Q\}^\{n\}\(Z\_\{0\}\(x\)\), initialize

Z^n−1​\(x\):=P​\(Zn​\(x\)\),U^n−1​\(x\):=ℬ0​\(Z^n−1​\(x\)\),\\widehat\{Z\}\_\{n\-1\}\(x\):=P\(Z\_\{n\}\(x\)\),\\qquad\\widehat\{U\}\_\{n\-1\}\(x\):=\\mathcal\{B\}\_\{0\}\(\\widehat\{Z\}\_\{n\-1\}\(x\)\),and, forj=n−2,…,0j=n\-2,\\ldots,0, set

Z^j​\(x\):=P​\(Z^j\+1​\(x\)\),U^j​\(x\):=ℬn−1−j​\(Z^j​\(x\)\)\+𝒯n​\(Z^j​\(x\),U^j\+1​\(x\)\)\.\\widehat\{Z\}\_\{j\}\(x\):=P\(\\widehat\{Z\}\_\{j\+1\}\(x\)\),\\qquad\\widehat\{U\}\_\{j\}\(x\):=\\mathcal\{B\}\_\{n\-1\-j\}\(\\widehat\{Z\}\_\{j\}\(x\)\)\+\\mathcal\{T\}\_\{n\}\(\\widehat\{Z\}\_\{j\}\(x\),\\widehat\{U\}\_\{j\+1\}\(x\)\)\.Exact backward replay and the preceding selected\-action argument give, by backward induction,Z^j​\(x\)=Zj​\(x\)\\widehat\{Z\}\_\{j\}\(x\)=Z\_\{j\}\(x\)andU^j​\(x\)=Uj​\(x\)\\widehat\{U\}\_\{j\}\(x\)=U\_\{j\}\(x\)\. HenceU^0​\(x\)=Vecs⁡\(Sn\)​\(x\)\\widehat\{U\}\_\{0\}\(x\)=\\operatorname\{Vec\}\_\{s\}\(S\_\{n\}\)\(x\)forx∈\[0,1\)x\\in\[0,1\), and equality atx=1x=1follows by continuity\.

The forward and backward computations useO​\(n\)O\(n\)fixed\-dimensional blocks\. The readout architecture is fixed because the numberNBN\_\{B\}of templates is independent ofnn; only its hardwired coefficients vary with the stage\. Thus the vectorized forcing sum has fixed width, depthO​\(n\)O\(n\), and parameter boundC2​\(1\+Kn\)​ΛnC\_\{2\}\(1\+K\_\{n\}\)\\Lambda^\{n\}\. The standard finite gluing over the translated unit intervals then yields the asserted realization ofSnS\_\{n\}onℝ\\mathbb\{R\}\. ∎

### 5\.2General fixed\-span stage\-dependent affine iterates

LetB\(0\),…,B\(N\):ℝ→ℝpB^\{\(0\)\},\\ldots,B^\{\(N\)\}:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}be fixed CPwL templates supported in\[0,L\]\[0,L\], and define

Br:=∑α=0Nλr,α​B\(α\),r≥0\.B\_\{r\}:=\\sum\_\{\\alpha=0\}^\{N\}\\lambda\_\{r,\\alpha\}B^\{\(\\alpha\)\},\\qquad r\\geq 0\.For a final depthnn, set

Kn:=max0≤r<n0≤α≤N⁡\|λr,α\|\.K\_\{n\}:=\\max\_\{\\begin\{subarray\}\{c\}0\\leq r<n\\\\ 0\\leq\\alpha\\leq N\\end\{subarray\}\}\|\\lambda\_\{r,\\alpha\}\|\.
###### Theorem 5\.2\(Fixed\-span stage\-dependent affine realization\)\.

LetM≥3M\\geq 3, let the mask be finitely supported, and assume that the associated homogeneous operatorVVpreserves\[0,L\]\[0,L\]\. For the forcing terms above, define

Sn:=∑r=0n−1Vn−1−r​Br\.S\_\{n\}:=\\sum\_\{r=0\}^\{n\-1\}V^\{\\,n\-1\-r\}B\_\{r\}\.Then there exist constantsC0,C1\>0C\_\{0\},C\_\{1\}\>0, independent ofnnand of the coefficientsλr,α\\lambda\_\{r,\\alpha\}, such that

Sn∈ΥC0,C1​n​\(ReLU;1,p\),n≥1\.S\_\{n\}\\in\\Upsilon\_\{C\_\{0\},C\_\{1\}n\}\(\\operatorname\{ReLU\};1,p\),\\qquad n\\geq 1\.The realizing networks may moreover be chosen with weights and biases bounded byC2​\(1\+Kn\)​ΛnC\_\{2\}\(1\+K\_\{n\}\)\\Lambda^\{n\}, whereC2,Λ\>0C\_\{2\},\\Lambda\>0depend only on the fixed templates, the matrix mask, the admissible offset, and the support margin\.

Ifγ:ℝ→ℝp\\gamma:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}is CPwL and supported in\[0,L\]\[0,L\], andWr​γ:=V​γ\+BrW\_\{r\}\\gamma:=V\\gamma\+B\_\{r\}, then there exist constantsC0′,C1′,C2′,Λ′\>0C\_\{0\}^\{\\prime\},C\_\{1\}^\{\\prime\},C\_\{2\}^\{\\prime\},\\Lambda^\{\\prime\}\>0, independent ofnnand of the coefficientsλr,α\\lambda\_\{r,\\alpha\}, such that

Wn−1​⋯​W0​γ∈ΥC0′,C1′​n​\(ReLU;1,p\),n≥1,W\_\{n\-1\}\\cdots W\_\{0\}\\gamma\\in\\Upsilon\_\{C\_\{0\}^\{\\prime\},C\_\{1\}^\{\\prime\}n\}\(\\operatorname\{ReLU\};1,p\),\\qquad n\\geq 1,and the realizing networks have weights and biases bounded byC2′​\(1\+Kn\)​\(Λ′\)nC\_\{2\}^\{\\prime\}\(1\+K\_\{n\}\)\(\\Lambda^\{\\prime\}\)^\{n\}\.

###### Proof\.

Apply[Section4\.1](https://arxiv.org/html/2607.20586#S4.SS1)once to each fixed template\. Thus, for everyα\\alpha, one has

B\(α\)​\(t\)=∑ν=1Nαaα,ν​hα,ν​\(t−δα,ν\)​eμα,ν,B^\{\(\\alpha\)\}\(t\)=\\sum\_\{\\nu=1\}^\{N\_\{\\alpha\}\}a\_\{\\alpha,\\nu\}h\_\{\\alpha,\\nu\}\(t\-\\delta\_\{\\alpha,\\nu\}\)e\_\{\\mu\_\{\\alpha,\\nu\}\},whereδα,ν∈ℤ∪\(ℤ\+s\)\\delta\_\{\\alpha,\\nu\}\\in\\mathbb\{Z\}\\cup\(\\mathbb\{Z\}\+s\), and eachhα,νh\_\{\\alpha,\\nu\}is a special hat\. The total number of atoms is fixed independently ofnn\.

Substituting these decompositions intoSnS\_\{n\}, we obtain a finite sum of terms of the form

∑r=0n−1λr,αaα,νVn−1−r\(hα,ν\(⋅−δα,ν\)eμα,ν\)\.\\sum\_\{r=0\}^\{n\-1\}\\lambda\_\{r,\\alpha\}a\_\{\\alpha,\\nu\}V^\{\\,n\-1\-r\}\\bigl\(h\_\{\\alpha,\\nu\}\(\\,\\cdot\-\\delta\_\{\\alpha,\\nu\}\)e\_\{\\mu\_\{\\alpha,\\nu\}\}\\bigr\)\.Each atom is aligned with either the ordinary frame or the fixed admissibless\-offset frame\. For a fixed atom, its stage\-dependent coefficients areλr,α​aα,ν\\lambda\_\{r,\\alpha\}a\_\{\\alpha,\\nu\}, whose absolute values are bounded by a fixed multiple ofKnK\_\{n\}\. Hence[Theorem5\.1](https://arxiv.org/html/2607.20586#S5.Thmtheorem1)applies in the corresponding frame\.

Since only finitely many atoms occur, their networks can be run in parallel and summed by a final affine layer\. This enlarges the width by only a fixed factor, preserves depthO​\(n\)O\(n\), and gives the parameter boundC2​\(1\+Kn\)​ΛnC\_\{2\}\(1\+K\_\{n\}\)\\Lambda^\{n\}\. This proves the assertion forSnS\_\{n\}\.

For the affine iterates, the standard stage\-dependent identity recalled from\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]gives

Wn−1​⋯​W0​γ=Vn​γ\+Sn\.W\_\{n\-1\}\\cdots W\_\{0\}\\gamma=V^\{n\}\\gamma\+S\_\{n\}\.The homogeneous term is realized by[Theorem2\.5](https://arxiv.org/html/2607.20586#S2.Thmtheorem5)\. Running the homogeneous and forcing networks in parallel and adding their outputs proves the final assertion, including the stated parameter bound after enlarging the constants\. ∎

## 6Reductions and geometric consequences

The recursive geometric constructions developed in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]include open curves, finitely many curve states, and stage\-dependent connector data\. Their reduction to vector\-valued refinement is algebraic and does not depend on the particular realization mechanism\. We record only the parts needed to apply the affine and fixed\-span theorems proved above\.

### 6\.1Anchored defects and finite\-state systems

An*anchored profile*is a CPwL mapΓ:ℝ→ℝp\\Gamma:\\mathbb\{R\}\\to\\mathbb\{R\}^\{p\}that is constant outside a bounded interval\. Subtracting such a profile turns an open curve with fixed tails into a compactly supported defect\.

More generally, consider a stage\-dependent recursion

γr\+1=V​γr\+Br\\gamma\_\{r\+1\}=V\\gamma\_\{r\}\+B\_\{r\}and writeγr=Γr\+ηr\\gamma\_\{r\}=\\Gamma\_\{r\}\+\\eta\_\{r\}, whereΓr\\Gamma\_\{r\}is a prescribed anchored profile\. Then the defect satisfies

ηr\+1=V​ηr\+Er,Er:=V​Γr\+Br−Γr\+1\.\\eta\_\{r\+1\}=V\\eta\_\{r\}\+E\_\{r\},\\qquad E\_\{r\}:=V\\Gamma\_\{r\}\+B\_\{r\}\-\\Gamma\_\{r\+1\}\.\(12\)Indeed, this follows immediately by substitutingγr=Γr\+ηr\\gamma\_\{r\}=\\Gamma\_\{r\}\+\\eta\_\{r\}into the recursion\.

The compact support ofErE\_\{r\}is determined entirely by the tails\. To make this explicit, setA∗:=∑j∈ℤAjA\_\{\*\}:=\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}\. IfΓr\\Gamma\_\{r\},Γr\+1\\Gamma\_\{r\+1\}, andBrB\_\{r\}have constant left and right tails denoted byΓr,±\\Gamma\_\{r,\\pm\},Γr\+1,±\\Gamma\_\{r\+1,\\pm\}, andBr,±B\_\{r,\\pm\}, respectively, thenErE\_\{r\}is compactly supported whenever

A∗​Γr,±\+Br,±=Γr\+1,±\.A\_\{\*\}\\Gamma\_\{r,\\pm\}\+B\_\{r,\\pm\}=\\Gamma\_\{r\+1,\\pm\}\.For a stationary anchorΓr=Γ\\Gamma\_\{r\}=\\Gammaand stationary forcingBr=BB\_\{r\}=B, this reduces to the usual compatibility condition forE=W​Γ−ΓE=W\\Gamma\-\\Gamma\.

###### Corollary 6\.1\(Anchored\-profile reduction\)\.

Assume thatVVpreserves a support window\[0,L\]\[0,L\], thatη0\\eta\_\{0\}is CPwL and supported in\[0,L\]\[0,L\], and that the defect forcing in \([12](https://arxiv.org/html/2607.20586#S6.E12)\) has the fixed\-span form

Er=∑α=0Nλr,α​E\(α\)E\_\{r\}=\\sum\_\{\\alpha=0\}^\{N\}\\lambda\_\{r,\\alpha\}E^\{\(\\alpha\)\}for fixed CPwL templatesE\(α\)E^\{\(\\alpha\)\}supported in\[0,L\]\[0,L\]\. IfM≥3M\\geq 3, then everyηn\\eta\_\{n\}admits an exact fixed\-width ReLU realization of depthO​\(n\)O\(n\)\. The same conclusion holds forM=2M=2when the templatesE\(α\)E^\{\(\\alpha\)\}are ordinary\-frame seam\-separated\.

If, in addition, the profilesΓn\\Gamma\_\{n\}belong to a fixed finite\-dimensional CPwL span, then the same conclusion holds forγn=Γn\+ηn\\gamma\_\{n\}=\\Gamma\_\{n\}\+\\eta\_\{n\}\.

###### Proof\.

The defect recursion is a stage\-dependent affine refinement system, so[Theorem5\.2](https://arxiv.org/html/2607.20586#S5.Thmtheorem2)applies toηn\\eta\_\{n\}\. A fixed finite\-dimensional CPwL realization ofΓn\\Gamma\_\{n\}can then be added in parallel without changing the asymptotic width or depth\. ∎

Finite\-state recursive systems require no additional realization argument\. Suppose thatγ=\(γ1,…,γR\)\\gamma=\(\\gamma\_\{1\},\\ldots,\\gamma\_\{R\}\)satisfies

\(𝔚​γ\)a​\(t\)=∑b=1R∑j∈ℤAja​b​γb​\(M​t−j\)\+Ba​\(t\),a=1,…,R\.\(\\mathfrak\{W\}\\gamma\)\_\{a\}\(t\)=\\sum\_\{b=1\}^\{R\}\\sum\_\{j\\in\\mathbb\{Z\}\}A\_\{j\}^\{ab\}\\gamma\_\{b\}\(Mt\-j\)\+B\_\{a\}\(t\),\\qquad a=1,\\ldots,R\.Stacking the state components into the map

G​\(t\):=\(γ1​\(t\),…,γR​\(t\)\)∈ℝp​RG\(t\):=\\bigl\(\\gamma\_\{1\}\(t\),\\ldots,\\gamma\_\{R\}\(t\)\\bigr\)\\in\\mathbb\{R\}^\{pR\}converts this system into an ordinary vector\-valued affine refinement operator whosejj\-th mask matrix has\(a,b\)\(a,b\)\-blockAja​bA\_\{j\}^\{ab\}\. Consequently, the main affine theorem and its stage\-dependent extension apply directly to the stacked system\. Any individual state is recovered by a fixed coordinate projection\.

### 6\.2Recursive curve generators

Copy\-and\-connector constructions provide a common geometric source of the defect recursion \([12](https://arxiv.org/html/2607.20586#S6.E12)\)\. In such a construction, the scaled and transformed copies of the preceding curve determine the fixed homogeneous operatorVV, while connectors and changes of anchor contribute toErE\_\{r\}\. Detailed versions of this reduction, including the corresponding finite\-state formulations, are given in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]\.

The new realization theorem applies once the resulting defect forcing lies in a fixed finite\-dimensional CPwL span\. Notice that this condition concernsErE\_\{r\}itself\. Since

Er=V​Γr\+Br−Γr\+1,E\_\{r\}=V\\Gamma\_\{r\}\+B\_\{r\}\-\\Gamma\_\{r\+1\},its coefficients may depend on the parameters describing both stagesrrandr\+1r\+1; they need not coincide with the coefficients used to representΓr\\Gamma\_\{r\}or the connector data separately\.

###### Corollary 6\.2\(Linear\-depth upgrade for geometric recursions\)\.

Assume thatVVpreserves a support window\[0,L\]\[0,L\], and consider a recursive curve construction whose anchored and, if necessary, stacked defect satisfies

ηr\+1=V​ηr\+∑α=0Nλr,α​E\(α\),\\eta\_\{r\+1\}=V\\eta\_\{r\}\+\\sum\_\{\\alpha=0\}^\{N\}\\lambda\_\{r,\\alpha\}E^\{\(\\alpha\)\},whereη0\\eta\_\{0\}and the fixed CPwL templatesE\(α\)E^\{\(\\alpha\)\}are supported in\[0,L\]\[0,L\]\. IfM≥3M\\geq 3, then thenn\-th defect admits an exact fixed\-width ReLU realization of depthO​\(n\)O\(n\)\. The same conclusion holds forM=2M=2if every templateE\(α\)E^\{\(\\alpha\)\}is ordinary\-frame seam\-separated\.

If

Kn:=max0≤r<n0≤α≤N⁡\|λr,α\|,K\_\{n\}:=\\max\_\{\\begin\{subarray\}\{c\}0\\leq r<n\\\\ 0\\leq\\alpha\\leq N\\end\{subarray\}\}\|\\lambda\_\{r,\\alpha\}\|,the weights and biases may be bounded byC​\(1\+Kn\)​ΛnC\(1\+K\_\{n\}\)\\Lambda^\{n\}, with constants depending only on the fixed refinement and template data\.

###### Proof\.

This is an immediate application of[Theorem5\.2](https://arxiv.org/html/2607.20586#S5.Thmtheorem2), followed, when needed, by the fixed anchored\-profile and coordinate\-projection readouts described above\. ∎

The Hilbert\- and Morton\-type constructions described in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]fit this criterion after their anchored defects are formed: the stage dependence is carried by finitely many fixed anchor and connector profiles, with coefficients governed by the corresponding endpoint scales\. The corollary therefore upgrades the previously available realization of these stage\-dependent affine recursions to fixed width and depthO​\(n\)O\(n\)\. We do not repeat their explicit copy matrices, translations, and connector formulas here\.

## 7Conclusions

We have proved exact fixed\-width, depth\-O​\(n\)O\(n\)ReLU realization results for affine one\-dimensional refinement iterates with vector\-valued CPwL data\. The homogeneous contribution is supplied by the loop\-controller theorem of\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]; the new contribution is a linear\-depth realization of the affine forcing sum\. For everyM≥3M\\geq 3, arbitrary compactly supported CPwL forcing terms are covered by combining the ordinary frame with a nontrivial admissible offset frame\. The resulting networks have weights and biases growing at most exponentially innn, according to the coarse estimates used here\.

The central new mechanism is the residual memory controller\. It replaces the noninvertible residual dynamics by an injective skew\-product and thereby permits exact backward replay of the residual states required by the Horner recursion\. Complementary loop readouts then recover the forcing values exactly on those states\. The remaining branch\-selector ambiguity is handled by a finite\-horizon trapping argument: whenever a residual enters a selector transition strip, all later forcing samples vanish, and hence the accumulated affine state being multiplied is zero\. Thus the memory controller solves the chronology problem, while the vanishing\-state mechanism resolves the remaining selector ambiguity\.

The construction also applies to stage\-dependent forcing terms lying in a fixed finite\-dimensional CPwL span\. In particular, the anchored\-profile, finite\-state, and copy\-and\-connector reductions developed in\[[4](https://arxiv.org/html/2607.20586#bib.bib4)\]inherit the improved linear\-depth bound whenever their defect forcing belongs to such a span\. This includes the Hilbert\- and Morton\-type recursive constructions discussed there\. Open curves are reduced to compactly supported defects by subtracting anchored profiles, while finite\-state systems are reduced to ordinary vector\-valued refinement by stacking their components\.

The present results concern exact realization of finite iterates rather than convergence to limiting fractal objects\. They are also one\-dimensional in the parameter and CPwL in regularity\. ForM=2M=2, no nontrivial admissible offset exists, and the current argument therefore covers only ordinary\-frame seam\-separated forcing terms; arbitrary binary CPwL forcing requires a more general seam\-exact mechanism\. Extending the memory construction to higher\-dimensional residual dynamics and sharpening the present parameter bounds remain natural directions for further work\.

## Acknowledgments

This work was supported by the Natural Sciences and Engineering Research Council of Canada through its Discovery Grants program\.

## References

- \[1\]I\. Daubechies, R\. DeVore, S\. Foucart, B\. Hanin, and G\. Petrova,*Nonlinear Approximation and \(Deep\) ReLU Networks*, Constr\. Approx\.55\(2022\), no\. 1, 127–172\.[DOI: 10\.1007/s00365\-021\-09548\-z](https://doi.org/10.1007/s00365-021-09548-z)
- \[2\]R\. DeVore, B\. Hanin, and G\. Petrova,*Neural Network Approximation*, Acta Numer\.30\(2021\), 327–444\.[DOI: 10\.1017/S0962492921000052](https://doi.org/10.1017/S0962492921000052)
- \[3\]I\. Daubechies, R\. DeVore, N\. Dym, S\. Faigenbaum\-Golovin, S\. Z\. Kovalsky, K\.\-C\. Lin, J\. Park, G\. Petrova, and B\. Sober,*Neural Network Approximation of Refinable Functions*, IEEE Trans\. Inform\. Theory69\(2023\), no\. 1, 482–495\.[DOI: 10\.1109/TIT\.2022\.3199601](https://doi.org/10.1109/TIT.2022.3199601)
- \[4\]B\. Bolorkhuu and T\. Gantumur,*Exact Loop Controllers for ReLU Realization of Homogeneous Curve Refinements*, arXiv:2605\.01655, 2026\.
- \[5\]N\. Dym, B\. Sober, and I\. Daubechies,*Expression of Fractals Through Neural Network Functions*, IEEE J\. Sel\. Areas Inf\. Theory1\(2020\), no\. 1, 57–66\.[DOI: 10\.1109/JSAIT\.2020\.2991422](https://doi.org/10.1109/JSAIT.2020.2991422)
- \[6\]J\. He, L\. Li, J\. Xu, and C\. Zheng,*ReLU Deep Neural Networks and Linear Finite Elements*, J\. Comput\. Math\.38\(2020\), no\. 3, 502–527\.[DOI: 10\.4208/jcm\.1901\-m2018\-0160](https://doi.org/10.4208/jcm.1901-m2018-0160)

Similar Articles

Shallower ReLU Network Representations via Exact Linear Algebra

arXiv cs.LG

This paper improves theoretical bounds on the depth of ReLU networks needed to represent the maximum function, showing exact two-hidden-layer representations for up to 10 inputs and improved depth for larger n via exact linear algebra techniques.

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv cs.AI

A new verifier-free breadth-depth refinement framework improves LLM reasoning at test time by sampling multiple rollouts, iteratively refining each via self-critique, and aggregating with majority voting. It consistently outperforms greedy decoding, majority voting, and verifier-based selection across several math benchmarks and open-weight models.

Learning to Refine Hidden States for Reliable LLM Reasoning

arXiv cs.LG

Proposes ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations in LLMs before decoding, improving reasoning reliability and efficiency compared to chain-of-thought methods.