An Integrable Token Mixing Layer from the Generalized Yang Baxter Equation

arXiv cs.LG Papers

Summary

The paper introduces YB-Mixer, a token-mixing layer derived from the generalized Yang-Baxter equation, which is exactly norm-preserving, depth-stable, and allows order-free and variable-budget inference. It achieves competitive performance on long-range memory tasks with fewer parameters compared to attention and state-space baselines.

arXiv:2606.15085v1 Announce Type: new Abstract: The YB Mixer is a sequence token mixing layer derived from free fermion and generalized Yang Baxter structures. It applies a core principle from integrable systems where a local algebraic constraint guarantees global computational stability. By using the Ising exchange algebra the mixer creates a free fermionic structure that acts as an exactly norm preserving orthogonal map. This algebra also produces commuting transfer matrices which allow inference to be order free and adaptable to any variable budget. To ensure the model can generalize to longer sequence lengths it uses a spectral circulant generator. This generator maintains the crucial orthogonal and commuting properties of the system. The result is a highly stable and mathematically grounded architecture for sequence processing.
Original Article
View Cached Full Text

Cached at: 06/16/26, 11:37 AM

# 1. Introduction
Source: [https://arxiv.org/html/2606.15085](https://arxiv.org/html/2606.15085)
An Integrable Token\-Mixing Layer from the Generalized Yang–Baxter Equation

Snigdha Chandan Khilar1

1Independent Researcher \|snkhilar@gmail\.com

††footnotetext:Correspondence:snkhilar@gmail\.com###### Abstract

We introduceYB\-Mixer, a sequence token\-mixing layer derived from the free\-fermion / generalized Yang–Baxter structure recently used to construct*hidden*transverse\-field Ising models\. The design rests on a single transferable principle from integrable systems: a*local*algebraic constraint on adjacent operations can certify*global*computational guarantees, independently of the representation\. Concretely, the*Ising exchange algebra*\(an extraspecial22\-group relation\) certifies \(i\) a free\-fermionic structure that makes the mixer an exactly norm\-preserving orthogonal map, and \(ii\) commuting transfer matrices that make inference*order\-free*and*variable\-budget*\(“anytime”\)\. We provide a complete, reproducible empirical pipeline of seven experiments\. We verify the generalized Yang–Baxter equation \(gYBE\) numerically to machine precision; prove and verify that the YBE constraint reduces to a well\-conditioned algebraic surrogate, making integrable gates efficiently learnable; build a brick\-wall YB\-Mixer layer that is exactly norm\-preserving and depth\-stable \(Jacobian condition number=1=1at all depths\); verify commuting transfer matrices and the resulting schedule\-invariant inference; train an integrable\-flow model end\-to\-end that matches a self\-attention baseline on a long\-range transport task at∼3\.3×\\sim\\\!3\.3\\timesfewer parameters; and demonstrate exact order\-free, variable\-budget inference that attention lacks\. We compare against orthogonal RNN, diagonal state\-space, attention, and nonlinear\-mixer baselines, finding YB\-Mixer matches or exceeds the structured baselines on long\-range memory at fewer parameters while honestly lagging a nonlinear mixer on content\-dependent recall\. Finally we show that the length\-generalization failure of a*local*generator is fixed by a*spectral*\(non\-local, circulant\) generator that remains orthogonal and commuting: trained atL=16L\{=\}16it generalizes toL=64L\{=\}64with roughly flat accuracy\. Scaled to∼2\.5\\sim\\\!2\.5M parameters across five downstream tasks against*properly\-tuned*baselines \(an S4D\-Lin/HiPPO SSM, LRU, Transformer, and FNet\), the orthogonal spectral mixer is competitive with the strongest baseline—best or tied on three of five tasks at the fewest parameters, reaching84\.8%84\.8\\%on sequential\-CIFAR \(LRA\-Image\) versus5151–72%72\\%for the identically\-scaffolded Transformer, LRU, and FNet—and is one of only two mixers that solve long\-range token retrieval \(Induction Heads\) exactly\. Code:[https://github\.com/nssprogrammer/yb\-mixer](https://github.com/nssprogrammer/yb-mixer)\.

Modern sequence models are built from*token mixers*that exchange information across positions: self\-attention\[[12](https://arxiv.org/html/2606.15085#bib.bib12)\], MLP\-style mixers\[[13](https://arxiv.org/html/2606.15085#bib.bib13)\], and structured state\-space models \(SSMs\)\[[22](https://arxiv.org/html/2606.15085#bib.bib22),[15](https://arxiv.org/html/2606.15085#bib.bib15),[30](https://arxiv.org/html/2606.15085#bib.bib30)\]\. Two recurrent practical concerns are*stability*\(gradients should neither vanish nor explode with depth\) and*inference flexibility*\(the ability to spend variable compute at test time\)\. Orthogonal and unitary recurrent layers\[[14](https://arxiv.org/html/2606.15085#bib.bib14)\]address stability by construction\. Here we ask whether the much richer toolkit of*quantum integrability*—which is, at heart, a theory of when many operations*commute*—can be exported to design token mixers with provable structure\.

Our starting point is a recent construction of*hidden*transverse\-field Ising models \(TFIMs\) from the generalized Yang–Baxter equation\[[1](https://arxiv.org/html/2606.15085#bib.bib1)\]\. That work shows that seemingly interacting multi\-site spin chains are secretly free\-fermionic, integrable, and governed by the*Ising exchange algebra*of their Hamiltonian densities\. The key lesson, abstracted away from physics, is a design pattern:

> *A purely local algebraic relation between neighbouring operations can certify a global, representation\-independent computational property—here, exact diagonalizability \(orthogonality\) and commuting families of operators \(order\-freedom\)\.*

We turn this pattern into a concrete neural layer,YB\-Mixer, and validate every link of the chain numerically\. Our contributions are:

1. 1\.A verified integrable primitive\(§[4\.1](https://arxiv.org/html/2606.15085#S4.SS1)\)\. We construct the gYBERR\-matrix from extraspecial\-22\-group generators and verify the\(d,6,3\)\(d,6,3\)\-gYBE to machine precision\.
2. 2\.A learnability reduction\(§[4\.2](https://arxiv.org/html/2606.15085#S4.SS2)\)\. We prove that for the Baxterized ansatzR​\(λ\)=𝟙\+tan⁡\(λ\)​MR\(\\lambda\)=\\mathbb\{1\}\+\\tan\(\\lambda\)Mthe braided YBE residual vanishes*iff*M2=𝟙M^\{2\}=\\mathbb\{1\}and neighbouring embeddings anticommute, and show that direct residual minimization is ill\-conditioned whereas the equivalent algebraic surrogate reliably yields integrable gates\.
3. 3\.A norm\-preserving, depth\-stable mixer\(§[4\.3](https://arxiv.org/html/2606.15085#S4.SS3)\)\. The free\-fermion gate acts as an orthogonal map on token features; a brick\-wall of such gates has Jacobian condition number exactly11at all depths\.
4. 4\.Commuting transfer matrices and \(scoped\) anytime inference\(§[4\.4](https://arxiv.org/html/2606.15085#S4.SS4), §[4\.6](https://arxiv.org/html/2606.15085#S4.SS6)\)\. We verify\[τ​\(λ\),τ​\(μ\)\]≈0\[\\tau\(\\lambda\),\\tau\(\\mu\)\]\\approx 0and show that an integrable\-*flow*model supports*exact*order\-free, variable\-budget inference\. We are explicit that this is a property of the one\-parameter group, holds only for the integrable flow \(not arbitrary nets containing a YB\-Mixer layer\), and is one realization—with an exactness and order\-freedom guarantee—of the broader adaptive\-computation idea\[[26](https://arxiv.org/html/2606.15085#bib.bib26),[27](https://arxiv.org/html/2606.15085#bib.bib27),[40](https://arxiv.org/html/2606.15085#bib.bib40),[41](https://arxiv.org/html/2606.15085#bib.bib41)\]\.
5. 5\.A rigorous, honest empirical study\(§[4\.5](https://arxiv.org/html/2606.15085#S4.SS5), §[4\.7](https://arxiv.org/html/2606.15085#S4.SS7)\)\. YB\-Mixer matches a self\-attention baseline on a long\-range transport task at far fewer parameters \(multi\-seed\), with a principled initialization recipe; we document a length\-generalization limitation rooted in free\-fermion dispersion\.
6. 6\.Baselines and a length\-generalization fix\(§[4\.8](https://arxiv.org/html/2606.15085#S4.SS8), §[4\.7](https://arxiv.org/html/2606.15085#S4.SS7)\)\. Against orthogonal RNN, diagonal SSM, attention, and a nonlinear mixer, YB\-Mixer matches or exceeds the structured baselines on long\-range memory at fewer parameters, and trails only the nonlinear mixer on content recall\. We further resolve the dispersion\-driven length\-generalization failure with a*spectral*generator that stays orthogonal and commuting and generalizes to4×4\\timesthe training length\.
7. 7\.Scaled benchmarks\(§[4\.9](https://arxiv.org/html/2606.15085#S4.SS9)\)\. At∼\\sim2\.5M parameters against properly\-tuned baselines \(S4D\-Lin/HiPPO, LRU, Transformer, FNet\), the orthogonal spectral mixer is best or tied\-best on three of five tasks at the fewest parameters—tying the tuned SSM on sequential\-CIFAR \(LRA\-Image,84\.8%84\.8\\%vs5151–72%72\\%for Transformer/LRU/FNet\) and solving Induction\-Heads retrieval exactly—while honestly trailing the tuned SSM on IMDB and ListOps\.

All experiments are small, fully reproducible, and provided as standalone scripts in the released code\. We are candid that this is a controlled\-task study: it establishes the architecture and verifies its properties, and does not claim benchmark\-scale accuracy \(§[6](https://arxiv.org/html/2606.15085#S6)\)\.

## 2\. Background theory

We collect the integrable\-systems machinery we use\. Readers familiar with the transverse\-field Ising model, Jordan–Wigner fermionization, and the quantum inverse scattering method may skim to §[2\.3](https://arxiv.org/html/2606.15085#S2.SS3)\.

### 2\.1\. The transverse\-field Ising model and the Ising exchange algebra

The one\-dimensional spin\-12\\tfrac\{1\}\{2\}TFIM onNNsites is

HTFIM=−g​∑jZj−∑jXj​Xj\+1,H\_\{\\mathrm\{TFIM\}\}\\;=\\;\-\\,g\\sum\_\{j\}Z\_\{j\}\\;\-\\;\\sum\_\{j\}X\_\{j\}X\_\{j\+1\},\(1\)whereXj,Yj,ZjX\_\{j\},Y\_\{j\},Z\_\{j\}are Pauli operators acting on sitejj\. Define the local*Hamiltonian densities*

hjz=Zj,hjx​x=Xj​Xj\+1\.h^\{z\}\_\{j\}\\;=\\;Z\_\{j\},\\qquad h^\{xx\}\_\{j\}\\;=\\;X\_\{j\}X\_\{j\+1\}\.\(2\)A direct computation shows they obey the*Ising exchange algebra*:

\[hjz,hkz\]=\[hjx​x,hkx​x\]=0,\[hjz,hkx​x\]=0\(j≠k,k\+1\),\\displaystyle\[h^\{z\}\_\{j\},h^\{z\}\_\{k\}\]=\[h^\{xx\}\_\{j\},h^\{xx\}\_\{k\}\]=0,\\qquad\[h^\{z\}\_\{j\},h^\{xx\}\_\{k\}\]=0\\ \\ \(j\\neq k,\\,k\{\+\}1\),\(3\)\{hjz,hjx​x\}=\{hj\+1z,hkx​x\}=0,\(hjz\)2=\(hjx​x\)2=𝟙\.\\displaystyle\\\{h^\{z\}\_\{j\},h^\{xx\}\_\{j\}\\\}=\\\{h^\{z\}\_\{j\+1\},h^\{xx\}\_\{k\}\\\}=0,\\qquad\(h^\{z\}\_\{j\}\)^\{2\}=\(h^\{xx\}\_\{j\}\)^\{2\}=\\mathbb\{1\}\.The crucial fact\[[9](https://arxiv.org/html/2606.15085#bib.bib9),[1](https://arxiv.org/html/2606.15085#bib.bib1)\]is that \([3](https://arxiv.org/html/2606.15085#S2.E3)\)*alone*—independent of the matrix realization—forces a free\-fermionic spectrum\. A*local*relation between neighbours thus certifies a*global*structural property\. This representation\-independence is what we will exploit\.

### 2\.2\. Jordan–Wigner fermionization

The Jordan–Wigner \(JW\) transformation maps spins to Majorana fermions,

γ2​j−1=\(∏k<jZk\)​Xj,γ2​j=\(∏k<jZk\)​Yj,\{γi,γj\}=2​δi​j​1\.\\gamma\_\{2j\-1\}=\\Big\(\\prod\_\{k<j\}Z\_\{k\}\\Big\)X\_\{j\},\\qquad\\gamma\_\{2j\}=\\Big\(\\prod\_\{k<j\}Z\_\{k\}\\Big\)Y\_\{j\},\\qquad\\\{\\gamma\_\{i\},\\gamma\_\{j\}\\\}=2\\,\\delta\_\{ij\}\\,\\mathbb\{1\}\.\(4\)Under \([4](https://arxiv.org/html/2606.15085#S2.E4)\) the TFIM \([1](https://arxiv.org/html/2606.15085#S2.E1)\) becomes a quadratic \(free\) Majorana HamiltonianH=i​g​∑jγ2​j−1​γ2​j\+i​∑jγ2​j​γ2​j\+1H=i\\,g\\sum\_\{j\}\\gamma\_\{2j\-1\}\\gamma\_\{2j\}\+i\\sum\_\{j\}\\gamma\_\{2j\}\\gamma\_\{2j\+1\}, diagonalizable by a Bogoliubov \(orthogonal\) rotation\. An operator is*free\-fermionic*precisely when it is*quadratic*in theγ\\gamma’s; its action is then completely determined by an antisymmetric*single\-particle*matrix, anO​\(dim\)O\(\\dim\)object rather than the exponential many\-body operator\. This single\-particle reduction is the bridge to a classical neural layer \(§[3](https://arxiv.org/html/2606.15085#S3)\)\.

### 2\.3\. The generalized Yang–Baxter equation

The ordinary Yang–Baxter equation\[[2](https://arxiv.org/html/2606.15085#bib.bib2),[3](https://arxiv.org/html/2606.15085#bib.bib3)\]is the consistency condition for factorized scattering of two\-body interactions\. Its*generalized*\(d,ℓ,m\)\(d,\\ell,m\)form\[[7](https://arxiv.org/html/2606.15085#bib.bib7)\]allowsRR\-matrices supported onℓ\\elladjacent sites, shifted bymm:

R1​⋯​ℓ​\(λ\)​R\(1\+m\)​⋯​\(ℓ\+m\)​\(λ\+μ\)​R1​⋯​ℓ​\(μ\)=R\(1\+m\)​⋯​\(ℓ\+m\)​\(μ\)​R1​⋯​ℓ​\(λ\+μ\)​R\(1\+m\)​⋯​\(ℓ\+m\)​\(λ\),R\_\{1\\cdots\\ell\}\(\\lambda\)\\,R\_\{\(1\+m\)\\cdots\(\\ell\+m\)\}\(\\lambda\{\+\}\\mu\)\\,R\_\{1\\cdots\\ell\}\(\\mu\)=R\_\{\(1\+m\)\\cdots\(\\ell\+m\)\}\(\\mu\)\\,R\_\{1\\cdots\\ell\}\(\\lambda\{\+\}\\mu\)\\,R\_\{\(1\+m\)\\cdots\(\\ell\+m\)\}\(\\lambda\),\(5\)an operator equation on⨂j=1ℓ\+mℋd\\bigotimes\_\{j=1\}^\{\\ell\+m\}\\mathcal\{H\}\_\{d\}, written here in*braided*\(additive\) form\. The construction of\[[1](https://arxiv.org/html/2606.15085#bib.bib1)\]uses multi\-site operatorsMjM\_\{j\}built from generators of*extraspecial22\-groups*, satisfying

Mj2=𝟙,\{Mj,Mj\+1\}=0,\[Mj,Mk\]=0\(\|j−k\|≥2\),M\_\{j\}^\{2\}=\\mathbb\{1\},\\qquad\\\{M\_\{j\},M\_\{j\+1\}\\\}=0,\\qquad\[M\_\{j\},M\_\{k\}\]=0\\ \\ \(\|j\-k\|\\geq 2\),\(6\)together with the*Baxterized*RR\-matrix

R​\(λ\)=1\+tan⁡\(λ\)​M\.R\(\\lambda\)\\;=\\;\\mathbb\{1\}\+\\tan\(\\lambda\)\\,M\.\(7\)The spectral parameter enters through the tangent becauseM2=𝟙M^\{2\}=\\mathbb\{1\}: substituting \([7](https://arxiv.org/html/2606.15085#S2.E7)\) into \([5](https://arxiv.org/html/2606.15085#S2.E5)\) and demanding nontrivial solutions yields the functional equation

a​\(λ1\+λ3\)=a​\(λ1\)\+a​\(λ3\)1−κ​a​\(λ1\)​a​\(λ3\),M2=κ​1,a\(\\lambda\_\{1\}\{\+\}\\lambda\_\{3\}\)=\\frac\{a\(\\lambda\_\{1\}\)\+a\(\\lambda\_\{3\}\)\}\{1\-\\kappa\\,a\(\\lambda\_\{1\}\)a\(\\lambda\_\{3\}\)\},\\qquad M^\{2\}=\\kappa\\,\\mathbb\{1\},\(8\)solved bya​\(λ\)=tan⁡\(λ\)/κa\(\\lambda\)=\\tan\(\\lambda\)/\\sqrt\{\\kappa\}\(hereκ=1\\kappa=1, the tangent addition law\)\. We give the elementary reduction underlying this in Lemma[1](https://arxiv.org/html/2606.15085#Thmlemma1)below, which is also what makes integrable gates learnable\.

### 2\.4\. Quantum inverse scattering and commuting transfer matrices

Given anRR\-matrix solving the \(non\-braided\) YBE, the quantum inverse scattering method \(QISM\)\[[8](https://arxiv.org/html/2606.15085#bib.bib8)\]builds a one\-parameter family of mutually commuting operators\. With an auxiliary spaceaa, the*monodromy*and*transfer*matrices are

Ta​\(λ\)=Ra,N​\(λ\)​⋯​Ra,1​\(λ\),τ​\(λ\)=tra⁡Ta​\(λ\),T\_\{a\}\(\\lambda\)=R\_\{a,N\}\(\\lambda\)\\cdots R\_\{a,1\}\(\\lambda\),\\qquad\\tau\(\\lambda\)=\\operatorname\{tr\}\_\{a\}\\,T\_\{a\}\(\\lambda\),\(9\)and the YBE/R​T​TRTTrelation implies

\[τ​\(λ\),τ​\(μ\)\]=0∀λ,μ\.\[\\tau\(\\lambda\),\\tau\(\\mu\)\]=0\\qquad\\forall\\,\\lambda,\\mu\.\(10\)Equation \([10](https://arxiv.org/html/2606.15085#S2.E10)\) is the algebraic heart of integrability: an entire family of “forward passes” indexed by the spectral parameter mutually commute\. In §[3](https://arxiv.org/html/2606.15085#S3)we read this as*order\-freedom*of inference\.

### 2\.5\. The boost operator \(briefly\)

The conserved chargesIr\+1I\_\{r\+1\}generated byτ​\(λ\)\\tau\(\\lambda\)can be obtained from a single*boost operator*B=∑jj​MjB=\\sum\_\{j\}j\\,M\_\{j\}via the recursionIr\+1=1r​\[B,Ir\]I\_\{r\+1\}=\\tfrac\{1\}\{r\}\[B,I\_\{r\}\]\[[10](https://arxiv.org/html/2606.15085#bib.bib10),[1](https://arxiv.org/html/2606.15085#bib.bib1)\]\. Each charge is a range\-rrbilinear with a string of conserved central elements between its endpoints\. We do not use the boost tower in our experiments but note it as a route to multi\-scale, mutually compatible features \(§[7](https://arxiv.org/html/2606.15085#S7)\)\.

## 3\. From integrable algebra to a neural layer

![Refer to caption](https://arxiv.org/html/2606.15085v1/ybe_1.png)Figure 1:The YB\-Mixer architecture\.\(a\) The integrable\-flow model \(Eq\.[13](https://arxiv.org/html/2606.15085#S3.E13)\): the input sequence is embedded, mixed by a single orthogonal flowU​\(s\)=exp⁡\(s​K\)U\(s\)=\\exp\(sK\)generated by a learned antisymmetric generatorKK, and read out by a small nonlinear head applied once at position0\. \(b\) The brick\-wall YB\-Mixer layer, the discrete realization of the flow: two\-token integrable gates act on the even bonds\(1,2\),\(3,4\),…\(1,2\),\(3,4\),\\dotsand then the odd bonds\(2,3\),\(4,5\),…\(2,3\),\(4,5\),\\dots; stackingΘ​\(L\)\\Theta\(L\)such layers produces a light cone that couples the entire sequence\. \(c\) The free\-fermion gate, the integrable primitive \(Eq\.[12](https://arxiv.org/html/2606.15085#S3.E12)\): a per\-channel2×22\{\\times\}2rotation by angleθ\\theta\. Because the gate is quadratic in Majorana operators, its single\-particle action is an orthogonal matrix, so the mixer is exactly norm\-preserving and depth\-stable, with Jacobian condition number11at all depths \(Table[3](https://arxiv.org/html/2606.15085#S4.T3)\)\.The global structure of YB\-Mixer is shown in Figure[1](https://arxiv.org/html/2606.15085#S3.F1)\. The design reads the integrable operator algebra as a wiring diagram: local \(anti\)commutation of adjacent gates certifies a global free\-fermionic—hence orthogonal—mixing rule, and the Baxterized gate supplies a continuous mixing\-strength dialλ\\lambda\.

### 3\.1\. Design principle

YB\-Mixer is obtained by reading the operator algebra as a wiring diagram with guarantees \(Table[1](https://arxiv.org/html/2606.15085#S3.T1)\):

Table 1:From integrable structure to neural\-layer design: each algebraic property is read as a guarantee on the mixing layer\.
### 3\.2\. The free\-fermion reduction makes the gate orthogonal

A free\-fermion gate is*quadratic*in Majoranas and therefore acts on the*single\-particle*space as an orthogonal matrix\. For example, the two\-qubit gateM=X⊗YM=X\\otimes Yequals, under \([4](https://arxiv.org/html/2606.15085#S2.E4)\), the Majorana bilinear

X⊗Y=−i​γ2​γ4,X\\otimes Y=\-\\,i\\,\\gamma\_\{2\}\\gamma\_\{4\},\(11\)whose single\-particle action is a rotation in the\(γ2,γ4\)\(\\gamma\_\{2\},\\gamma\_\{4\}\)plane\. Consequently a brick\-wall of such gates is an*orthogonal token mixer*: norm\-preserving by construction, with Jacobian singular values identically11\.

### 3\.3\. The brick\-wall YB\-Mixer layer

LetX∈ℝB×L×CX\\in\\mathbb\{R\}^\{B\\times L\\times C\}be a batch ofLL\-token sequences withCCfeature channels\. A YB\-Mixer layer applies a two\-token integrable gate on the even bonds\(1,2\),\(3,4\),…\(1,2\),\(3,4\),\\dotsand then the odd bonds\(2,3\),\(4,5\),…\(2,3\),\(4,5\),\\dots\. In the simplest free\-fermion instantiation the gate is a per\-channel rotation by an angleθ\\theta,

\(xi′xi\+1′\)=\(cos⁡θ−sin⁡θsin⁡θcos⁡θ\)​\(xixi\+1\),\\begin\{pmatrix\}x^\{\\prime\}\_\{i\}\\\\ x^\{\\prime\}\_\{i\+1\}\\end\{pmatrix\}=\\begin\{pmatrix\}\\cos\\theta&\-\\sin\\theta\\\\ \\sin\\theta&\\cos\\theta\\end\{pmatrix\}\\begin\{pmatrix\}x\_\{i\}\\\\ x\_\{i\+1\}\\end\{pmatrix\},\(12\)which is exactly orthogonal and integrable\. StackingΘ​\(L\)\\Theta\(L\)such layers yields a light cone covering the whole sequence\. Crucially, the following reduction makes*learning*an integrable gate well\-posed\.

###### Lemma 1\(YBE reduction\)\.

LetMA,MBM\_\{A\},M\_\{B\}be Hermitian withMA2=MB2=𝟙M\_\{A\}^\{2\}=M\_\{B\}^\{2\}=\\mathbb\{1\}and\{MA,MB\}=0\\\{M\_\{A\},M\_\{B\}\\\}=0, and setRA​\(λ\)=𝟙\+tan⁡\(λ\)​MAR\_\{A\}\(\\lambda\)=\\mathbb\{1\}\+\\tan\(\\lambda\)M\_\{A\},RB​\(λ\)=𝟙\+tan⁡\(λ\)​MBR\_\{B\}\(\\lambda\)=\\mathbb\{1\}\+\\tan\(\\lambda\)M\_\{B\}\. Then the braided YBE

RA​\(λ\)​RB​\(λ\+μ\)​RA​\(μ\)=RB​\(μ\)​RA​\(λ\+μ\)​RB​\(λ\)R\_\{A\}\(\\lambda\)R\_\{B\}\(\\lambda\{\+\}\\mu\)R\_\{A\}\(\\mu\)=R\_\{B\}\(\\mu\)R\_\{A\}\(\\lambda\{\+\}\\mu\)R\_\{B\}\(\\lambda\)holds for allλ,μ\\lambda,\\muif and only iftan⁡\(λ\+μ\)=tan⁡λ\+tan⁡μ1−tan⁡λ​tan⁡μ\\tan\(\\lambda\{\+\}\\mu\)=\\dfrac\{\\tan\\lambda\+\\tan\\mu\}\{1\-\\tan\\lambda\\tan\\mu\}\(the tangent addition law\)\.

###### Proof\.

Writex=tan⁡λx=\\tan\\lambda,z=tan⁡μz=\\tan\\mu,y=tan⁡\(λ\+μ\)y=\\tan\(\\lambda\{\+\}\\mu\)\. UsingMA2=MB2=𝟙M\_\{A\}^\{2\}=M\_\{B\}^\{2\}=\\mathbb\{1\}andMB​MA=−MA​MBM\_\{B\}M\_\{A\}=\-M\_\{A\}M\_\{B\}, expand both sides:

LHS=\(1\+x​z\)\+\(x\+z\)​MA\+y​\(1−x​z\)​MB\+y​\(x−z\)​MA​MB,\\text\{LHS\}=\(1\{\+\}xz\)\+\(x\{\+\}z\)M\_\{A\}\+y\(1\{\-\}xz\)M\_\{B\}\+y\(x\{\-\}z\)\\,M\_\{A\}M\_\{B\},RHS=\(1\+x​z\)\+y​\(1−x​z\)​MA\+\(x\+z\)​MB\+y​\(x−z\)​MA​MB\.\\text\{RHS\}=\(1\{\+\}xz\)\+y\(1\{\-\}xz\)M\_\{A\}\+\(x\{\+\}z\)M\_\{B\}\+y\(x\{\-\}z\)\\,M\_\{A\}M\_\{B\}\.Since𝟙,MA,MB,MA​MB\\mathbb\{1\},M\_\{A\},M\_\{B\},M\_\{A\}M\_\{B\}are linearly independent, equality holds iffx\+z=y​\(1−x​z\)x\+z=y\(1\-xz\), i\.e\.y=\(x\+z\)/\(1−x​z\)y=\(x\{\+\}z\)/\(1\{\-\}xz\)\. ∎

### 3\.4\. The integrable\-flow model and anytime inference

![Refer to caption](https://arxiv.org/html/2606.15085v1/ybe_3.png)Figure 2:Anytime inference from the one\-parameter group structure \(§[3\.4](https://arxiv.org/html/2606.15085#S3.SS4), §[4\.6](https://arxiv.org/html/2606.15085#S4.SS6)\)\.Top: becauseU​\(s\)​U​\(s′\)=U​\(s\+s′\)U\(s\)U\(s^\{\\prime\}\)=U\(s\{\+\}s^\{\\prime\}\), the rapidity budgetssis additive and the model may be read out at any partial budget; test accuracy rises smoothly and saturates \(e\.g\.s=0\.25,0\.5,0\.75,1\.0→0\.51,0\.71,0\.94,1\.00s=0\.25,0\.5,0\.75,1\.0\\to 0\.51,0\.71,0\.94,1\.00; the measured curve for one run is Fig\.[3](https://arxiv.org/html/2606.15085#S4.F3)\), giving a consistent coarse\-to\-fine answer at any stopping point\. Bottom: order\-freedom—splitting a fixed total budgets=1s=1into the same increments applied in different orders yields the identical final stateU​\(1\)U\(1\)\(output spread∼10−16\{\\sim\}10^\{\-16\}across orderings\)\. This is the architectural reading of the commuting transfer matrices\[τ​\(λ\),τ​\(μ\)\]=0\[\\tau\(\\lambda\),\\tau\(\\mu\)\]=0: the increments commute, so inference is order\-free, cacheable, and parallelizable\. The property is exact for the integrable flow as a single one\-parameter group with one readout head; interleaving nonlinearities between mixing layers breaks the global group structure\.Figure[2](https://arxiv.org/html/2606.15085#S3.F2)previews the inference\-time payoff of integrability: unlike a fixed\-compute network, the integrable flow supports variable\-budget \(anytime\) inference and produces a stopping\-point–independent, order\-independent result, which we verify end\-to\-end on a trained model in §[4\.6](https://arxiv.org/html/2606.15085#S4.SS6)\.

Replacing the discrete brick\-wall by its continuous\-time limit gives a particularly clean object\. LetKKbe a learned antisymmetric single\-particle generator and define the*integrable flow*

U​\(s\)=exp⁡\(s​K\),K⊤=−K⇒U​\(s\)∈O​\(L\)\.U\(s\)=\\exp\(sK\),\\qquad K^\{\\top\}=\-K\\ \\Rightarrow\\ U\(s\)\\in O\(L\)\.\(13\)Because\{U​\(s\)\}s\\\{U\(s\)\\\}\_\{s\}is a one\-parameter group,

U​\(s\)​U​\(s′\)=U​\(s\+s′\)=U​\(s′\)​U​\(s\),U\(s\)\\,U\(s^\{\\prime\}\)=U\(s\+s^\{\\prime\}\)=U\(s^\{\\prime\}\)\\,U\(s\),\(14\)the entire sequence\-mixing is*additive*and*order\-independent*: a total “rapidity”ssmay be split into arbitrary increments and applied in any order, cached, or parallelized, all giving the identical result\. This is the architectural manifestation of the commuting family \([10](https://arxiv.org/html/2606.15085#S2.E10)\)\. A variable rapidity budgetssthen provides a consistent coarse\-to\-fine \(*anytime*\) inference mode\. The full model is

f​\(x\)=Head​\(\[U​\(s\)​Emb​\(x\)\]pos​0\),f\(x\)=\\mathrm\{Head\}\\Big\(\\big\[\\,U\(s\)\\,\\mathrm\{Emb\}\(x\)\\,\\big\]\_\{\\text\{pos \}0\}\\Big\),\(15\)with a small nonlinear head applied*once*at readout \(which does not disturb the group structure of the mixing\)\. We emphasize the scope: the anytime property is a property of the integrable*flow*; interleaving nonlinearities*between*mixing layers breaks the global group structure \(§[6](https://arxiv.org/html/2606.15085#S6)\)\.

## 4\. Experiments

All experiments are deterministic and reproducible from the released scripts\. Tables report numbers from reference runs; magnitudes \(not last digits\) are the content\.

### 4\.1\. The integrable primitive is real

We build Majorana operators on66qubits \(6464\-dimensional Hilbert space\), the multi\-siteMM\-operators of\[[1](https://arxiv.org/html/2606.15085#bib.bib1)\], and the BaxterizedR​\(λ\)R\(\\lambda\), then check the algebra \([6](https://arxiv.org/html/2606.15085#S2.E6)\) and the\(d,6,3\)\(d,6,3\)\-gYBE \([5](https://arxiv.org/html/2606.15085#S2.E5)\) \(Table[2](https://arxiv.org/html/2606.15085#S4.T2)\)\.

Table 2:Integrable primitive: algebra and gYBE residuals on66qubits\. All structural identities hold to machine precision, while a deliberately wrong addition law \(control\) gives anO​\(1\)O\(1\)residual\.The gYBE residual sits at machine precision while a deliberately wrong addition law gives anO​\(1\)O\(1\)residual, confirming the test is non\-vacuous: theRR\-matrix is genuinely integrable\.

### 4\.2\. Integrable gates are learnable

We implement a differentiable braided\-YBE residual on a three\-site space and validate it against the free\-fermion anchorM=X⊗YM=X\\otimes Y\(residual∼10−16\\sim\\\!10^\{\-16\}; random Hermitian gates giveO​\(1\)O\(1\)\)\. Direct minimization of the cubic residual over a free Hermitian gate is ill\-conditioned and frequently stalls\. Minimizing the algebraic surrogate‖M12​M23\+M23​M12‖2\\\|M\_\{12\}M\_\{23\}\+M\_\{23\}M\_\{12\}\\\|^\{2\}over the involution parameterizationM=U​D​U†M=UDU^\{\\dagger\}reliably drives the surrogate to≤10−6\\leq 10^\{\-6\}\(best restarts reach∼10−14\\sim\\\!10^\{\-14\}\), and crucially the*full*YBE residual at the solution vanishes \(∼10−7\\sim\\\!10^\{\-7\}, best∼10−14\\sim\\\!10^\{\-14\}\)\. Thus stochastic gradient descent, given only the YBE constraint,*rediscovers*the extraspecial\-22\-group algebra \([6](https://arxiv.org/html/2606.15085#S2.E6)\)\.

### 4\.3\. Norm preservation and depth stability

We compare three gates in a brick\-wall mixer \(§[3](https://arxiv.org/html/2606.15085#S3)\) across depth: \(A\) integrable free\-fermion rotations, \(B\) a random orthogonal gate \(control\), \(C\) a random generic gate \(control\)\. Table[3](https://arxiv.org/html/2606.15085#S4.T3)reports the output/input norm ratio and the Jacobian condition number\.

Table 3:Depth stability\. Integrable \(A\) and random\-orthogonal \(B\) gates keep Jacobian condition number=1=1and norm ratio=1=1at all depths; the generic gate \(C\) explodes\. The \(B\) control is deliberate: depth stability is the*orthogonality*benefit, which integrable gates provide automatically\. Integrability’s extra payoff is §[4\.4](https://arxiv.org/html/2606.15085#S4.SS4)\.
### 4\.4\. Commuting transfer matrices

We build the QISM transfer matrix \([9](https://arxiv.org/html/2606.15085#S2.E9)\) from the free\-fermion gate and measuremax⁡‖\[τ​\(λ\),τ​\(μ\)\]‖\\max\\\|\[\\tau\(\\lambda\),\\tau\(\\mu\)\]\\\|as the gate is perturbed off the algebra \([6](https://arxiv.org/html/2606.15085#S2.E6)\) by an amountε\\varepsilon\(Table[4](https://arxiv.org/html/2606.15085#S4.T4)\)\.

Table 4:Commuting transfer matrices\.max⁡‖\[τ​\(λ\),τ​\(μ\)\]‖\\max\\\|\[\\tau\(\\lambda\),\\tau\(\\mu\)\]\\\|is at machine precision at the integrable point \(ε=0\\varepsilon\{=\}0\) and grows monotonically as the gate is perturbed off theMM\-algebra\.The transfer matrices commute to machine precision exactly at the integrable point and the commutator grows monotonically away from it\. Commuting transfer matrices are therefore an*integrability*\-specific property, not shared by generic orthogonal mixers \(cf\. Table[3](https://arxiv.org/html/2606.15085#S4.T3), control B\)\.

### 4\.5\. Trainability and the initialization recipe

We train a YB\-Mixer on a long\-range*transport*task: inputs are random bitsx∈\{0,1\}Lx\\in\\\{0,1\\\}^\{L\}, the label isxL−1x\_\{L\-1\}, and the classifier reads only*position0*\. The task is solvable only if the mixer transports information across the whole sequence\. Table[5](https://arxiv.org/html/2606.15085#S4.T5)reports test accuracy \(L=16L=16\)\.

Table 5:Trainability\. With near\-π/4\\pi/4initialization YB\-Mixer solves transport and matches the unconstrained mixer; small\-angle initialization fails because the transmitted amplitude∼sin\(θ\)L−1\\sim\\\!\\sin\(\\theta\)^\{L\-1\}vanishes, trapping the optimizer\. The trained mixing layers remain orthogonal to∼10−6\\sim\\\!10^\{\-6\}, i\.e\. the integrable structure survives training\.The initialization finding is a concrete deployable recipe: integrable mixers must be initialized near the “swap” regime to avoid a vanishing\-transmission trap, analogous to gain calibration in orthogonal RNNs\[[14](https://arxiv.org/html/2606.15085#bib.bib14)\]\.

### 4\.6\. End\-to\-end anytime inference

We train the integrable\-flow model \([13](https://arxiv.org/html/2606.15085#S3.E13)\) \(MLP\-free mixing\) on the same task\. It learns perfectly \(test acc\.1\.0001\.000\)\. We then exercise the group structure \([14](https://arxiv.org/html/2606.15085#S3.E14)\):

- •Anytime budget\.Test accuracy as a function of applied rapidityssrises smoothly and saturates at the trained value \(Fig\.[3](https://arxiv.org/html/2606.15085#S4.F3)\):s=0\.25→0\.51s=0\.25\\\!\\to\\\!0\.51,0\.5→0\.690\.5\\\!\\to\\\!0\.69,0\.75→0\.940\.75\\\!\\to\\\!0\.94,1\.0→1\.001\.0\\\!\\to\\\!1\.00\(then a mild overshoot to0\.990\.99at1\.251\.25\)\. One may stop at any budget for a consistent coarse\-to\-fine answer\.
- •Order\-freedom\.Splittings=1s=1into five random increments and applying them in six random orders yields a*bit\-identical*output \(max relative spread∼10−16\\sim\\\!10^\{\-16\}\), whereas increments built from*different*generators \(non\-integrable\) diverge completely \(spread∼1\.2\\sim\\\!1\.2\)\.

This is the architectural payoff of \([10](https://arxiv.org/html/2606.15085#S2.E10)\): order\-free, cacheable, parallelizable, variable\-budget inference, which a standard fixed\-compute network does not provide\.

![Refer to caption](https://arxiv.org/html/2606.15085v1/x1.png)Figure 3:Measured anytime refinement curve\(integrable flow, transport task\)\. Test accuracy as a function of the applied rapidity budgetssrises monotonically from chance ats=0s\{=\}0and saturates at the trained value bys=1s\{=\}1, so any partial budget yields a consistent coarse\-to\-fine answer\. Values from the releasedanytimenotebook:s=0,0\.25,0\.5,0\.75,1\.0,1\.25→0\.50,0\.51,0\.69,0\.94,1\.00,0\.99s=0,0\.25,0\.5,0\.75,1\.0,1\.25\\;\\to\\;0\.50,0\.51,0\.69,0\.94,1\.00,0\.99\. Because each budget is one exact application ofU​\(s\)U\(s\), this is a refinement axis, not a compute\-saving one\.
### 4\.7\. Multi\-seed competitiveness and length generalization

Finally we compare against a self\-attention baseline over three seeds \(Table[6](https://arxiv.org/html/2606.15085#S4.T6)\)\.

Table 6:Competitiveness\. YB\-Mixer matches self\-attention on transport at∼3\.3×\\sim\\\!3\.3\\timesfewer parameters\. The order\-freedom spread over the three trained models is7×10−16±10−167\\times 10^\{\-16\}\\pm 10^\{\-16\}, i\.e\. the anytime property is robust across seeds\.#### Length generalization and the dispersion tension\.

A*translation\-invariant local*flow \(banded antisymmetric Toeplitz generator\), trained atL=16L\{=\}16and applied at largerLLwith budgets∝Ls\\propto L, does*not*transport across longer chains: accuracy is0\.650\.65atL=16L\{=\}16and falls to∼0\.45\\sim\\\!0\.45\(chance\) atL=24,32L\{=\}24,32\. The cause is structural and worth stating precisely: an orthogonal flow generated by an antisymmetric matrix is necessarily*reciprocal*; a*local*reciprocal generator has a curved dispersion relation, so a wavepacket spreads ballistically and precise transport degrades with length\. Clean \(non\-dispersive\) transport requires a*linear*dispersion relation, which a local generator cannot realize—but a*non\-local*one can\.

![Refer to caption](https://arxiv.org/html/2606.15085v1/ybe_2.png)Figure 4:Fourier\-diagonal mixing with the spectral generator \(§[4\.7](https://arxiv.org/html/2606.15085#S4.SS7)\)\.Replacing the local generator with a circulant one diagonalizesKKin the Fourier basis\. The DFT maps the time\-domain sequence to frequency modes; each modemmis rotated independently by a phaseei​φ​\(f\)e^\{i\\varphi\(f\)\}with normalized frequencyf=m/Lf=m/L, whereφ​\(f\)\\varphi\(f\)is parameterized by a small sine/cosine basis; the inverse DFT returns to the time domain\. The resulting flow is still exactly orthogonal \(norm\-preserving\), and because all circulant generators commute it forms an even cleaner commuting family, so the anytime property of §[3\.4](https://arxiv.org/html/2606.15085#S3.SS4)is preserved\. Sinceφ\\varphidepends only onf=m/Lf=m/L, the same generator instantiates at any length; a near\-linearφ​\(f\)\\varphi\(f\)corresponds to a near\-rigid shift, removing the dispersion that causes a local generator to fail to length\-generalize \(Table[7](https://arxiv.org/html/2606.15085#S4.T7)\)\.Figure[4](https://arxiv.org/html/2606.15085#S4.F4)illustrates why the spectral generator length\-generalizes\. An orthogonal flow generated by an antisymmetric matrix is necessarily reciprocal, and a local reciprocal generator has a curved dispersion relation, so a localized signal spreads ballistically and transport degrades with length\. A global circulant generator instead admits a near\-linear dispersion at the same length\-independent parameterization, giving non\-dispersive transport while retaining norm\-preservation and the commuting family\.

#### Resolution: a spectral generator\.

We therefore replace the local generator by a*circulant*\(spectral\) one, parameterizing the per\-mode rotation phaseφ​\(f\)\\varphi\(f\)as a function of the*normalized*frequencyf=m/Lf=m/L\(a small sine/cosine basis\)\. The resulting flow is diagonal in the Fourier basis: it is still exactly orthogonal \(norm\-preserving\), and since all circulant generators commute it forms an even cleaner commuting family, so the anytime property of §[3\.4](https://arxiv.org/html/2606.15085#S3.SS4)is preserved\. Becauseφ\\varphidepends only onf=m/Lf=m/L, the same generator instantiates at any length\. Trained atL=16L\{=\}16it generalizes with roughly flat accuracy \(Table[7](https://arxiv.org/html/2606.15085#S4.T7)\); a localized signal no longer disperses because a near\-linearφ​\(f\)\\varphi\(f\)corresponds to a near\-rigid shift\.

Table 7:Length generalization \(trainL=16L\{=\}16, test longer\)\. The local generator collapses to chance; the spectral \(non\-local but orthogonal and commuting\) generator stays roughly flat out to4×4\\timesthe training length\. This resolves the dispersion tension: non\-locality buys non\-dispersive transport while retaining norm\-preservation and the commuting family\.

### 4\.8\. Critical baselines: orthogonal RNN, SSM, attention, nonlinear mixer

We compare the integrable flow against the most relevant structured competitors on two tasks \(Table[8](https://arxiv.org/html/2606.15085#S4.T8)\): \(A\) long\-range*memory*\(label=x0=x\_\{0\}, read at the last position—the canonical task for orthogonal RNNs\) and \(B\) content\-dependent*associative recall*\(output the value following a query key—the canonical task where content\-based routing matters\)\. Two seeds, matched scale\.

Table 8:Baselines\. On long\-range*memory*, the integrable flow matches attention and the nonlinear mixer and*exceeds*the most direct structured competitors \(orthogonal RNN, diagonal SSM\) at fewer parameters\. On content\-dependent*recall*, all linear/orthogonal models \(FlowYB, orthogonal RNN, attention here\) cluster together and trail the*nonlinear*mixer—an honest, quantified expressivity gap \(the cost of the orthogonality constraint\)\.The picture is deliberately even\-handed: integrability/orthogonality is an asset for stable long\-range memory and a liability for content\-dependent routing\. YB\-Mixer’s distinguishing feature among these structured models is not raw accuracy but the exact order\-free, variable\-budget inference of §[4\.6](https://arxiv.org/html/2606.15085#S4.SS6)\.

### 4\.9\. Scaled benchmarks \(∼\\sim2\.5M parameters, fair baselines\)

To move beyond the controlled synthetic setting we scale the spectral, orthogonal YB\-Mixer \(§[4\.7](https://arxiv.org/html/2606.15085#S4.SS7)\) to∼\\sim2\.5M parameters on five downstream tasks, using a*single shared block scaffold*in which only the token mixer changes \(identical embedding, MLP, pooling, optimizer, and schedule;dim=256\\dim\{=\}256, depth88,5050epochs\)\. The four baselines are*properly tuned*representatives of their families: a diagonal SSM with theS4D\-Lin / HiPPOinitialization\[[29](https://arxiv.org/html/2606.15085#bib.bib29),[28](https://arxiv.org/html/2606.15085#bib.bib28)\]\(the fair SSM, not a minimal stand\-in\), theLRU\[[30](https://arxiv.org/html/2606.15085#bib.bib30)\]linear recurrent unit, aTransformer\[[12](https://arxiv.org/html/2606.15085#bib.bib12)\]block, andFNet\[[35](https://arxiv.org/html/2606.15085#bib.bib35)\]’s fixed22D\-FFT mixing \(the parameter\-free spectral cousin of YB\)\. Tasks span the recognized regimes: permuted\-MNIST \(L=784L\{=\}784\); the LRA\-Image \(sequential\-CIFAR\-10\) and LRA\-Text \(byte\-IMDB\) tasks\[[39](https://arxiv.org/html/2606.15085#bib.bib39)\]\(L=1024L\{=\}1024\); LRAListOps\(L=1024L\{=\}1024, hierarchical reasoning\); and theInduction Headsretrieval task \(L=256L\{=\}256\)\. Results in Table[9](https://arxiv.org/html/2606.15085#S4.T9)\.

Table 9:Matched\-scale downstream accuracy \(validation; same scaffold,dim=256\\dim\{=\}256, depth88,5050epochs, single seed\)\. Best per task in bold\. YB\-Mixer is best or tied\-best on three of five tasks \(permuted\-MNIST, seq\-CIFAR, Induction\) at the*fewest*parameters after LRU/FNet, while the properly\-initialized S4D\-Lin—the strongest baseline—wins the two linguistic/hierarchical tasks \(IMDB, ListOps\), with YB a close second on both\. On seq\-CIFAR, YB \(0\.8480\.848\) and S4D\-Lin \(0\.8480\.848\) are tied and roughly double attention \(0\.5130\.513\) and FNet \(0\.5280\.528\)\. On Induction Heads, YB and LRU achieve perfect retrieval while S4D\-Lin, Transformer, and FNet remain at chance \(≈1/15\\approx 1/15\)\. Parameter counts vary by a few percent across tasks with input vocabulary and length; the permuted\-MNIST and Induction configurations are slightly smaller\.Two findings stand out\. First, on the two longest*perceptual*sequences \(seq\-CIFAR, permuted\-MNIST\) the orthogonal spectral mixer is at the top of the table and attention collapses \(0\.5130\.513on seq\-CIFAR\), consistent with the known difficulty of global attention on very long, low\-level inputs; YB attains this with∼1\.8×\\sim\\\!1\.8\\timesfewer parameters than the Transformer\. Second, theInduction Headsresult is qualitative, not marginal: YB and LRU solve token\-level retrieval*perfectly*while the convolutional SSM \(S4D\-Lin\), the fixed\-FFT mixer \(FNet\), and attention sit at chance\. This task was run*without positional embeddings*to permit length\-extrapolation evaluation \(§[4\.7](https://arxiv.org/html/2606.15085#S4.SS7)\), which specifically disadvantages attention; the salient point is that YB’s structured mixing routes a specific token to the readout with no positional encoding at all, whereas a fixed spectral map \(FNet\) and a bidirectional diagonal convolution \(S4D\-Lin\) cannot\.

We remain careful about scope\. \(i\) The SSM baseline is now the properly HiPPO\-initialized S4D\-Lin—a fair, strong competitor \(it wins IMDB and ListOps\)—so this is a matched comparison against tuned baselines, not against minimal stand\-ins; absolute numbers are still*within our controlled harness*rather than against maximally\-tuned published systems \(tuned S4 reaches∼\\sim88% on LRA\-Image with task\-specific engineering\)\. \(ii\) Results are single\-seed; multi\-seed confirmation and the remaining LRA tasks \(Pathfinder, Retrieval\) are future work\. With those caveats, the picture is consistent and honest: YB\-Mixer is*competitive with a well\-tuned SSM*across five tasks—winning more of them, at fewer parameters—and is one of only two mixers that solve long\-range token retrieval, which we attribute to its global spectral receptive field combined with exact norm\-preservation\.

## 5\. Related work

#### Integrability and machine learning\.

Most prior intersections*use*neural networks to*discover or solve*RR\-matrices and integrable systems, rather than using integrability as an architectural primitive\[[16](https://arxiv.org/html/2606.15085#bib.bib16)\]\. Brick\-wall circuits of Yang–Baxter gates appear in quantum simulation\[[11](https://arxiv.org/html/2606.15085#bib.bib11)\]but not as classical learning layers\. YB\-Mixer instead uses the YBE/commuting\-transfer\-matrix structure*as*the mixer, which is, to our knowledge, new\.

#### Stable and structured mixers\.

Gradient stability via norm preservation is well studied: unitary and orthogonal RNNs\[[14](https://arxiv.org/html/2606.15085#bib.bib14),[17](https://arxiv.org/html/2606.15085#bib.bib17),[18](https://arxiv.org/html/2606.15085#bib.bib18),[19](https://arxiv.org/html/2606.15085#bib.bib19),[34](https://arxiv.org/html/2606.15085#bib.bib34),[32](https://arxiv.org/html/2606.15085#bib.bib32)\]and orthogonal initialization/dynamics\[[20](https://arxiv.org/html/2606.15085#bib.bib20),[21](https://arxiv.org/html/2606.15085#bib.bib21)\]\. YB\-Mixer’s depth stability \(Table[3](https://arxiv.org/html/2606.15085#S4.T3)\) is the*same*orthogonality benefit, here supplied automatically by the free\-fermion algebra rather than imposed; experimentally it matches or beats orthogonal RNNs and a diagonal SSM on long\-range memory \(Table[8](https://arxiv.org/html/2606.15085#S4.T8)\)\. The continuous flow \([13](https://arxiv.org/html/2606.15085#S3.E13)\) is an orthogonal linear state\-space model closely related to structured SSMs\[[22](https://arxiv.org/html/2606.15085#bib.bib22),[29](https://arxiv.org/html/2606.15085#bib.bib29),[23](https://arxiv.org/html/2606.15085#bib.bib23),[28](https://arxiv.org/html/2606.15085#bib.bib28),[15](https://arxiv.org/html/2606.15085#bib.bib15),[30](https://arxiv.org/html/2606.15085#bib.bib30)\]and the broader family of efficient recurrent and linear\-attention sequence mixers\[[36](https://arxiv.org/html/2606.15085#bib.bib36),[33](https://arxiv.org/html/2606.15085#bib.bib33),[38](https://arxiv.org/html/2606.15085#bib.bib38),[37](https://arxiv.org/html/2606.15085#bib.bib37),[31](https://arxiv.org/html/2606.15085#bib.bib31),[35](https://arxiv.org/html/2606.15085#bib.bib35)\]; indeed our*spectral*generator \(§[4\.7](https://arxiv.org/html/2606.15085#S4.SS7)\) is a Fourier\-diagonal \(global\-convolution\) SSM\. What integrability contributes on top is the commuting family and exact order\-free inference\. Physics\-structured architectures such as Hamiltonian neural networks\[[24](https://arxiv.org/html/2606.15085#bib.bib24),[25](https://arxiv.org/html/2606.15085#bib.bib25)\]similarly bake conservation laws into the model; YB\-Mixer bakes in integrability\.

#### Token mixers\.

Relative to attention\[[12](https://arxiv.org/html/2606.15085#bib.bib12)\], MLP\-style mixers\[[13](https://arxiv.org/html/2606.15085#bib.bib13)\], and Fourier mixers\[[35](https://arxiv.org/html/2606.15085#bib.bib35)\], YB\-Mixer is a constrained \(orthogonal, integrable\) mixer: less expressive for content\-based routing, but exactly stable and uniquely order\-free\.

#### Relation to the source construction and our delta\.

The physics we build on—hidden TFIMs, the extraspecial\-22\-groupRR\-matrices, and the boost\-operator charge tower—is due to\[[1](https://arxiv.org/html/2606.15085#bib.bib1)\], building on free\-fermions\-in\-disguise\[[6](https://arxiv.org/html/2606.15085#bib.bib6)\]and the generalized YBE of\[[7](https://arxiv.org/html/2606.15085#bib.bib7)\]\. Our contribution over\[[1](https://arxiv.org/html/2606.15085#bib.bib1)\]is entirely on the machine\-learning side and is fourfold: \(i\) the YBE\-to\-learnability reduction \(Lemma[1](https://arxiv.org/html/2606.15085#Thmlemma1)\) and the well\-conditioned surrogate that makes integrable gates*trainable*; \(ii\) the single\-particle realization of the gate as an*orthogonal token mixer*usable in a classical network; \(iii\) the reading of commuting transfer matrices as*order\-free inference*, demonstrated end\-to\-end on a trained model; and \(iv\) the spectral generator that makes the flow*length\-generalize*\. None of these appears in\[[1](https://arxiv.org/html/2606.15085#bib.bib1)\]\. We use no result from\[[1](https://arxiv.org/html/2606.15085#bib.bib1)\]beyond the verified algebra of §[2](https://arxiv.org/html/2606.15085#S2)\.

## 6\. Limitations

\(1\)Scale and tuning\.We report results at∼\\sim2\.5M parameters on five real downstream tasks \(§[4\.9](https://arxiv.org/html/2606.15085#S4.SS9)\) against properly\-initialized baselines \(the SSM is the HiPPO\-initialized S4D\-Lin, a fair competitor that wins IMDB and ListOps\), but not yet at10710^\{7\}\+ parameters, on the full LRA suite \(Pathfinder, Retrieval\), or with multi\-seed error bars; absolute numbers are still within a controlled harness rather than against maximally\-tuned published systems\. Larger\-scale, multi\-seed evaluation is the natural next step\. \(2\)Scope of anytime\.The exact order\-free/variable\-budget property holds for the integrable*flow*\(mixing as a one\-parameter group\) with a single readout head; interleaving nonlinearities*between*mixing layers breaks the global group structure\. It is therefore a property of a specific architecture class, not of any network containing a YB\-Mixer layer, and—while exact—is one instance of the broader adaptive\-computation idea\[[26](https://arxiv.org/html/2606.15085#bib.bib26),[27](https://arxiv.org/html/2606.15085#bib.bib27),[40](https://arxiv.org/html/2606.15085#bib.bib40),[41](https://arxiv.org/html/2606.15085#bib.bib41)\]\. \(3\)Expressivity\.As an orthogonal mixer, YB\-Mixer cannot perform content\-based routing; Table[8](https://arxiv.org/html/2606.15085#S4.T8)quantifies the gap to a nonlinear mixer on associative recall\. \(4\)Length generalization\.A*local*generator disperses and fails to length\-generalize; our*spectral*generator \(§[4\.7](https://arxiv.org/html/2606.15085#S4.SS7)\) resolves this on the transport task out to4×4\\times, but at a small accuracy cost and only verified on the synthetic setting\.

## 7\. Conclusion

We introduced YB\-Mixer, a token\-mixing layer derived from the generalized Yang–Baxter / free\-fermion structure of hidden Ising models\. The unifying idea is that a*local*algebraic constraint certifies*global*computational guarantees: the Ising exchange algebra makes the mixer exactly orthogonal \(norm\-preserving, depth\-stable\), and commuting transfer matrices make inference order\-free and variable\-budget\. We verified each link of this chain numerically and showed the resulting layer is trainable and, at∼\\sim2\.5M parameters against properly\-tuned baselines, competitive with the strongest of them—best or tied on three of five downstream tasks at the fewest parameters \(notably84\.8%84\.8\\%on LRA\-Image, tying the HiPPO\-initialized S4D\-Lin and far ahead of Transformer/LRU/FNet, and exact retrieval on Induction Heads\), while honestly documenting where a tuned SSM wins \(IMDB, ListOps\)\. Natural next steps include the boost\-operator charge tower as a parameter\-efficient multi\-scale feature generator, directed/non\-dispersive integrable generators for length generalization, and the\(d,2​k,k\)\(d,2k,k\)\-gYBE family for richer multi\-site mixers\.

## References

- \[1\]A\. Sinha, S\. Maity, P\. Padmanabhan, V\. Korepin,*Hidden Ising models from the generalized Yang–Baxter equation*, arXiv:2605\.30007 \(2026\)\.
- \[2\]C\. N\. Yang,*Some exact results for the many\-body problem in one dimension with repulsive delta\-function interaction*, Phys\. Rev\. Lett\.19, 1312 \(1967\)\.
- \[3\]R\. J\. Baxter,*Exactly Solved Models in Statistical Mechanics*, Academic Press \(1982\)\.
- \[4\]P\. Jordan, E\. Wigner,*Über das Paulische Äquivalenzverbot*, Z\. Phys\.47, 631 \(1928\)\.
- \[5\]P\. Pfeuty,*The one\-dimensional Ising model with a transverse field*, Ann\. Phys\.57, 79 \(1970\)\.
- \[6\]P\. Fendley,*Free fermions in disguise*, J\. Phys\. A52, 335002 \(2019\)\.
- \[7\]E\. C\. Rowell, Y\. Zhang, Y\.\-S\. Wu, M\.\-L\. Ge,*Extraspecial two\-groups, generalized Yang–Baxter equations and braiding quantum gates*, Quantum Inf\. Comput\.10, 685 \(2010\)\.
- \[8\]V\. E\. Korepin, N\. M\. Bogoliubov, A\. G\. Izergin,*Quantum Inverse Scattering Method and Correlation Functions*, Cambridge University Press \(1993\)\.
- \[9\]K\. Minami,*Solvable Hamiltonians and fermionization transformations obtained from operators satisfying specific commutation relations*, J\. Phys\. Soc\. Jpn\.85, 024003 \(2016\)\.
- \[10\]K\. Sogo, M\. Wadati,*Boost operator and its application to quantum Gelfand–Levitan equation*, Prog\. Theor\. Phys\.69, 431 \(1983\)\.
- \[11\]A\. Sinha, T\. Justin, P\. Padmanabhan, V\. Korepin,*The Yang–Baxter integrability of the critical Ising chain*, J\. Stat\. Mech\.2025, 103102 \(2025\)\.
- \[12\]A\. Vaswani et al\.,*Attention is all you need*, NeurIPS \(2017\)\.
- \[13\]I\. Tolstikhin et al\.,*MLP\-Mixer: An all\-MLP architecture for vision*, NeurIPS \(2021\)\.
- \[14\]M\. Arjovsky, A\. Shah, Y\. Bengio,*Unitary evolution recurrent neural networks*, ICML \(2016\)\.
- \[15\]A\. Gu, T\. Dao,*Mamba: Linear\-time sequence modeling with selective state spaces*, arXiv:2312\.00752 \(2023\)\.
- \[16\]S\. Lal, S\. Majumder, E\. Sobko,*The R\-mAtrIx Net*, arXiv:2304\.07247 \(2023\)\.
- \[17\]S\. Wisdom, T\. Powers, J\. Hershey, J\. Le Roux, L\. Atlas,*Full\-capacity unitary recurrent neural networks*, NeurIPS \(2016\)\.
- \[18\]Z\. Mhammedi, A\. Hellicar, A\. Rahman, J\. Bailey,*Efficient orthogonal parametrisation of recurrent neural networks using Householder reflections*, ICML \(2017\)\.
- \[19\]E\. Vorontsov, C\. Trabelsi, S\. Kadoury, C\. Pal,*On orthogonality and learning recurrent networks with long term dependencies*, ICML \(2017\)\.
- \[20\]A\. M\. Saxe, J\. L\. McClelland, S\. Ganguli,*Exact solutions to the nonlinear dynamics of learning in deep linear neural networks*, ICLR \(2014\)\.
- \[21\]J\. Pennington, S\. Schoenholz, S\. Ganguli,*Resurrecting the sigmoid in deep learning through dynamical isometry*, NeurIPS \(2017\)\.
- \[22\]A\. Gu, K\. Goel, C\. Ré,*Efficiently modeling long sequences with structured state spaces \(S4\)*, ICLR \(2022\)\.
- \[23\]J\. T\. H\. Smith, A\. Warrington, S\. W\. Linderman,*Simplified state space layers for sequence modeling \(S5\)*, ICLR \(2023\)\.
- \[24\]S\. Greydanus, M\. Dzamba, J\. Yosinski,*Hamiltonian neural networks*, NeurIPS \(2019\)\.
- \[25\]Z\. Chen, J\. Zhang, M\. Arjovsky, L\. Bottou,*Symplectic recurrent neural networks*, ICLR \(2020\)\.
- \[26\]S\. Teerapittayanon, B\. McDanel, H\. T\. Kung,*BranchyNet: Fast inference via early exiting from deep neural networks*, ICPR \(2016\)\.
- \[27\]T\. Schuster et al\.,*Confident adaptive language modeling*, NeurIPS \(2022\)\.
- \[28\]A\. Gu, T\. Dao, S\. Ermon, A\. Rudra, C\. Ré,*HiPPO: Recurrent memory with optimal polynomial projections*, NeurIPS \(2020\)\.
- \[29\]A\. Gu, A\. Gupta, K\. Goel, C\. Ré,*On the parameterization and initialization of diagonal state space models \(S4D\)*, NeurIPS \(2022\)\.
- \[30\]A\. Orvieto, S\. L\. Smith, A\. Gu, A\. Fernando, C\. Gulcehre, R\. Pascanu, S\. De,*Resurrecting recurrent neural networks for long sequences \(LRU\)*, ICML \(2023\)\.
- \[31\]S\. De, S\. L\. Smith, A\. Fernando, A\. Botev, et al\.,*Griffin: Mixing gated linear recurrences with local attention for efficient language models*, arXiv:2402\.19427 \(2024\)\.
- \[32\]K\. E\. Helfrich, D\. Willmott, Q\. Ye,*Orthogonal recurrent neural networks with scaled Cayley transform \(scoRNN\)*, ICML \(2018\)\.
- \[33\]B\. Peng et al\.,*RWKV: Reinventing RNNs for the transformer era*, Findings of EMNLP \(2023\)\.
- \[34\]L\. Jing et al\.,*Tunable efficient unitary neural networks \(EUNN\) and their application to RNNs*, ICML \(2017\)\.
- \[35\]J\. Lee\-Thorp, J\. Ainslie, I\. Eckstein, S\. Ontañón,*FNet: Mixing tokens with Fourier transforms*, NAACL \(2022\)\.
- \[36\]A\. Katharopoulos, A\. Vyas, N\. Pappas, F\. Fleuret,*Transformers are RNNs: Fast autoregressive transformers with linear attention*, ICML \(2020\)\.
- \[37\]M\. Poli et al\.,*Hyena hierarchy: Towards larger convolutional language models*, ICML \(2023\)\.
- \[38\]Y\. Sun, L\. Dong, S\. Huang, S\. Ma, Y\. Xia, J\. Xue, J\. Wang, F\. Wei,*Retentive network: A successor to transformer for large language models*, arXiv:2307\.08621 \(2023\)\.
- \[39\]Y\. Tay et al\.,*Long Range Arena: A benchmark for efficient transformers*, ICLR \(2021\)\.
- \[40\]M\. Dehghani, S\. Gouws, O\. Vinyals, J\. Uszkoreit, Ł\. Kaiser,*Universal transformers*, ICLR \(2019\)\.
- \[41\]A\. Kusupati et al\.,*Matryoshka representation learning*, NeurIPS \(2022\)\.
- \[42\]T\. Dao, D\. Y\. Fu, S\. Ermon, A\. Rudra, C\. Ré,*FlashAttention: Fast and memory\-efficient exact attention with IO\-awareness*, NeurIPS \(2022\)\.

Similar Articles

Mixing Times of Glauber Dynamics on Masked Language Models

arXiv cs.LG

This paper analyzes the global distributional behavior induced by iterative masked-token resampling in masked language models using Glauber dynamics. It introduces a rectangle test for incompatibility, establishes mixing time bounds, and empirically demonstrates phase transitions and metastable semantic basins.

Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time

arXiv cs.CL

This paper introduces OP-Mix, a data mixing algorithm that uses low-rank adapters trained on the current model to cheaply simulate candidate data mixtures, enabling efficient and unified data mixing across pretraining, continual midtraining, and continual instruction tuning. OP-Mix consistently finds near-optimal mixtures while using a fraction of the compute of baselines, improving pretraining perplexity by 6.3% and reducing compute by 66-95% in continual learning settings.