Between Amnesia and Chaos: A Memory Stability Expressivity Trilemma for Trainable Dissipative Oscillator Networks

arXiv cs.LG Papers

Summary

This paper presents a memory–stability–expressivity trilemma for trainable dissipative oscillator networks, showing that damping governs all three and limits trainability, with experimental validation on a 20-oscillator network confirming the theoretical bounds.

arXiv:2606.09929v1 Announce Type: new Abstract: Physical reservoir computing harnesses nonlinear mechanical dynamics but, by convention, freezes the substrate and trains only a linear readout, presuming the substrate is not usefully trainable. We revisit that premise for networks of nonlinear oscillators whose mass, damping, and stiffness are learned end-to-end through a symplectic integrator. Our central result is a trilemma: memory horizon, gradient stability, and dynamical expressivity cannot be simultaneously maximized, because all three are governed by the damping. The backward gradient decays at a rate set by the damping, capping how far back credit can propagate, while forward sensitivities grow exponentially in the largest Lyapunov exponent, so usable gradients require damping above a stability floor. Since the Lyapunov exponent falls as damping rises while the memory ceiling falls as the horizon grows, stable training is confined to a band that contracts with horizon and closes at a critical point. We test every step on a twenty-oscillator network. A damping sweep finds the largest Lyapunov exponent monotone and crossing zero at a well-defined stability floor, confirming the theorem's key assumption. A compute-matched comparison of learned versus frozen substrate on delayed recall across nine horizons shows the learned substrate dominating at short horizons and the advantage closing and reversing near a horizon of eleven steps, the predicted signature of band closure; trained models settle near the stability floor, seeking the edge of chaos unprompted. The analytic ceiling overestimates the empirical crossover roughly fivefold, a gap between detectable and learnable gradient that we report rather than tune away. The contribution is a confirmed account of when training a physical substrate beats freezing it.
Original Article
View Cached Full Text

Cached at: 06/10/26, 06:18 AM

# A Memory–Stability–Expressivity Trilemma for Trainable Dissipative Oscillator Networks
Source: [https://arxiv.org/html/2606.09929](https://arxiv.org/html/2606.09929)
###### Abstract

Physical reservoir computing harnesses nonlinear mechanical dynamics but, by convention, freezes the substrate and trains only a linear readout, presuming the substrate is not usefully trainable\. We revisit that premise for networks of nonlinear oscillators whose mass, damping, and stiffness are learned end\-to\-end through a symplectic integrator\. Our central result is a*trilemma*: memory horizon, gradient stability, and dynamical expressivity cannot be simultaneously maximized, because all three are governed by the damping\. The backward gradient decays at rateγ/2\\gamma/2, capping memory atH≤2γmin​log⁡1ϵH\\leq\\frac\{2\}\{\\gamma\_\{\\min\}\}\\log\\frac\{1\}\{\\epsilon\}, while forward sensitivities grow aseλmax​Te^\{\\lambda\_\{\\max\}T\}, so usable gradients require damping above a stability floor\. Since the Lyapunov exponent falls with damping and the memory ceiling falls with horizon, stable training is confined to a band that contracts withHHand closes at a critical horizonHcH\_\{c\}\. We test every step on a2020\-oscillator network\. A damping sweep finds the largest Lyapunov exponent monotone and crossing zero atγ⋆=0\.121\\gamma\_\{\\star\}=0\.121, confirming the theorem’s load\-bearing assumption\. A compute\-matched comparison of learned versus frozen substrate on delayed recall across nine horizons shows the learned substrate dominating at short horizons \(1\.001\.00vs\.0\.880\.88atH=1H\{=\}1\) and the advantage closing and reversing nearH≈11H\\approx 11, the predicted signature of band closure; trained models settle nearγ⋆\\gamma\_\{\\star\}, seeking the edge of chaos unprompted\. The analytic ceiling overestimates the empirical crossover roughly fivefold, a gap between detectable and learnable gradient that we report rather than tune away\. The contribution is a confirmed geometric account of when training a physical substrate beats freezing it\.

Keywords:physical reservoir computing, neural ordinary differential equations, dynamical systems, Lyapunov exponents, symplectic integration, sequence modeling, edge of chaos\.

## 1Introduction

When a practitioner reaches for a physical system as a computational substrate, the appeal is usually interpretability: a network of masses and springs obeys Newton’s laws, its state has named physical meaning, and its behavior is governed by equations one can write down\. The hope is that a model built from “pure physics” will be more legible than a neural network of inscrutable weights\. This hope is the starting point of a large literature on physical reservoir computing, in which the rich nonlinear dynamics of an oscillator lattice, a memristive film, a soft body, or an optical cavity are harnessed for computation\[[8](https://arxiv.org/html/2606.09929#bib.bib8),[7](https://arxiv.org/html/2606.09929#bib.bib7)\]\.

That literature operates under a near\-universal convention: the physical substrate is*frozen*, and only a linear readout is trained on its instantaneous state\. The freezing is not incidental\. It rests on a tacit premise that the medium is either inaccessible to gradient\-based tuning or that tuning it is not worth the cost\. Two developments make it possible to question that premise directly\. The adjoint sensitivity method for neural ordinary differential equations\[[1](https://arxiv.org/html/2606.09929#bib.bib1)\]lets one differentiate through the solution of a continuous dynamical system, and structure\-preserving integrators\[[3](https://arxiv.org/html/2606.09929#bib.bib3),[2](https://arxiv.org/html/2606.09929#bib.bib2),[6](https://arxiv.org/html/2606.09929#bib.bib6)\]keep those gradients honest over long horizons\. One can now build a network of nonlinear springs, integrate its equations of motion, and learn the spring and damping parameters end to end\.

The question this paper asks is not*whether*such training is possible \(it is\) but*when it is worthwhile*, and what fundamental obstacle governs the answer\. We find that the obstacle is sharp and has a clean geometric form\. The damping coefficient that controls how long a dissipative system remembers its input is the same quantity that controls whether its forward sensitivities \(and hence the gradients used to train it\) stay bounded\. Strong damping gives stable gradients but short memory; weak damping gives long memory but, in a coupled nonlinear system, chaos and exploding gradients\. The two demands pull against one another, and the region of damping that satisfies both is a band that shrinks as the required memory horizon grows\.

This paper differs from a position piece in that we build the system and measure every claim\. We implement a differentiable oscillator network with a symplectic integrator, estimate Lyapunov exponents with the Benettin algorithm, and run a controlled learned\-versus\-frozen comparison\. The measurements confirm the theorem’s qualitative structure and, equally usefully, expose where the simplest analytic estimate of the critical horizon is too loose\.

We make three linked contributions\. First, a*trilemma*: a theorem that memory horizon, gradient stability, and expressivity are jointly constrained by the damping, with a feasible training band that contracts with horizon and closes at a criticalHcH\_\{c\}\. Second, a*method*: edge\-of\-chaos training, with energy\-bounded parameterizations and symplectic integration that make the gradients trustworthy\. Third, an*empirical test*: a compute\-matched comparison whose outcome, by construction, either confirms or refutes the central prediction, and which we run to completion\.

What we deliberately do not claim: a universal\-approximation result \(inherited from the Neural ODE literature, not ours\); interpretability\-by\-virtue\-of\-being\-physics \(we argue against it in Section[7](https://arxiv.org/html/2606.09929#S7)\); and competitiveness with large language models \(the wrong axis on which to judge a mechanical substrate\)\.

The structure of the paper is as follows\. Section[2](https://arxiv.org/html/2606.09929#S2)surveys related work\. Section[3](https://arxiv.org/html/2606.09929#S3)states the model and proves energy boundedness\. Section[4](https://arxiv.org/html/2606.09929#S4)pre\-registers the claims\. Section[5](https://arxiv.org/html/2606.09929#S5)proves the trilemma\. Section[6](https://arxiv.org/html/2606.09929#S6)presents the training method\. Section[7](https://arxiv.org/html/2606.09929#S7)delimits the interpretability claim\. Section[8](https://arxiv.org/html/2606.09929#S8)gives the experimental setup\. Sections[9](https://arxiv.org/html/2606.09929#S9)and[10](https://arxiv.org/html/2606.09929#S10)report the measured Lyapunov sweep and the learned\-versus\-frozen comparison\. Section[11](https://arxiv.org/html/2606.09929#S11)discusses the results, including the analytic/empiricalHcH\_\{c\}gap\.

## 2Related Work

### 2\.1Neural Ordinary Differential Equations

Chen et al\. \[[1](https://arxiv.org/html/2606.09929#bib.bib1)\]introduced the treatment of a deep network’s forward pass as the solution of an ordinary differential equation, trained by the adjoint sensitivity method\. Our model is a Neural ODE whose vector field is constrained to mechanical form, and the gradient\-decay and gradient\-explosion phenomena we analyze are the continuous\-time analogues of those long studied in recurrent networks\.

### 2\.2Structure\-Preserving and Physics\-Informed Networks

Hamiltonian neural networks\[[3](https://arxiv.org/html/2606.09929#bib.bib3)\]and Lagrangian neural networks\[[2](https://arxiv.org/html/2606.09929#bib.bib2)\]bake conservation laws into the learned dynamics\. Symplectic integrators\[[6](https://arxiv.org/html/2606.09929#bib.bib6)\]conserve a shadow Hamiltonian and are essential for long\-horizon fidelity\. Our setting differs in being*dissipative and forced*rather than conservative, and our focus is the geometry of training feasibility rather than fitting a known physical system\. We borrow the symplectic machinery to keep gradients honest near the low\-damping edge\.

### 2\.3Physical Reservoir Computing

Physical reservoir computing\[[8](https://arxiv.org/html/2606.09929#bib.bib8),[7](https://arxiv.org/html/2606.09929#bib.bib7)\]drives a fixed nonlinear medium with an input signal and trains only a linear readout on the resulting transient state\. The substrate is deliberately not trained, and the field’s practical results rest on tuning the medium to the edge of chaos by hand\. This is the tradition our trilemma interrogates: we provide a principled account of*when*freezing the substrate is the only stable option, and we run the controlled learned\-versus\-frozen comparison directly\.

### 2\.4State\-Space Models and the Placement of Nonlinearity

Modern long\-sequence architectures\[[4](https://arxiv.org/html/2606.09929#bib.bib4),[5](https://arxiv.org/html/2606.09929#bib.bib5)\]keep the temporal recurrence linear and push nonlinearity into pointwise gating, achieving long memory without the instability of nonlinear recurrence\. Our trilemma offers an explanation for why this choice is effectively forced: in a dissipative nonlinear recurrence, the damping needed for long memory is incompatible with the damping needed for stable gradients beyond a critical horizon\.

### 2\.5The Gap This Paper Addresses

Each tradition above is mature in isolation\. What has not been stated cleanly is the coupling between the three failure modes of a*trained*mechanical substrate, nor the consequent geometry that decides when training the substrate beats freezing it\. This paper supplies that account as a theorem, an enabling method, and a measured comparison\.

## 3Model and Energy Boundedness

### 3\.1Equations of Motion

ConsiderNNcoupled oscillators with positions𝐱∈ℝN\\mathbf\{x\}\\in\\mathbb\{R\}^\{N\}, velocities𝐯∈ℝN\\mathbf\{v\}\\in\\mathbb\{R\}^\{N\}, diagonal mass matrixM=diag⁡\(m1,…,mN\)≻0M=\\operatorname\{diag\}\(m\_\{1\},\\dots,m\_\{N\}\)\\succ 0, diagonal dampingΓ=diag⁡\(γ1,…,γN\)⪰0\\Gamma=\\operatorname\{diag\}\(\\gamma\_\{1\},\\dots,\\gamma\_\{N\}\)\\succeq 0, and a learned potentialΦ\\Phi\. With learnable parametersθ\\thetacollecting masses, dampings, stiffnesses, and the input mapBB, and inputu​\(t\)u\(t\), the dynamics are

M​𝐱¨=−∇𝐱Φ​\(𝐱;θ\)−Γ​𝐱˙\+B​u​\(t\)\.M\\ddot\{\\mathbf\{x\}\}=\-\\,\\nabla\_\{\\mathbf\{x\}\}\\Phi\(\\mathbf\{x\};\\theta\)\-\\Gamma\\dot\{\\mathbf\{x\}\}\+B\\,u\(t\)\.Lifting to first order with𝐳=\(𝐱,𝐯\)∈ℝ2​N\\mathbf\{z\}=\(\\mathbf\{x\},\\mathbf\{v\}\)\\in\\mathbb\{R\}^\{2N\}gives a Neural ODE with a mechanically constrained vector field and a linear readouty^=W​𝐳​\(T\)\+b\\hat\{y\}=W\\mathbf\{z\}\(T\)\+b\. We use*saturating springs*: each coupling\(i,j\)\(i,j\)on a fixed graph contributes forceki​j​tanh⁡\(r/ai​j\)k\_\{ij\}\\tanh\(r/a\_\{ij\}\)withr=xi−xjr=x\_\{i\}\-x\_\{j\}and potentialΦi​j​\(r\)=ki​j​ai​j2​log⁡cosh⁡\(r/ai​j\)\\Phi\_\{ij\}\(r\)=k\_\{ij\}a\_\{ij\}^\{2\}\\log\\cosh\(r/a\_\{ij\}\)\. The bounded force is the mechanical analogue of a squashing activation, and the potential is confining by construction\.

### 3\.2Stability by Construction

We parameterize masses, dampings, stiffnesses, and scales throughsoftplus\\mathrm\{softplus\}of unconstrained latents, so all physical quantities are positive and the potential is radially unbounded and bounded below\.

#### Proposition 1 \(Energy\-bounded forward pass\)\.

LetE​\(𝐳\)=12​𝐯⊤​M​𝐯\+Φ​\(𝐱;θ\)E\(\\mathbf\{z\}\)=\\tfrac\{1\}\{2\}\\mathbf\{v\}^\{\\top\}M\\mathbf\{v\}\+\\Phi\(\\mathbf\{x\};\\theta\)\. Then along the dynamics

E˙=−𝐯⊤​Γ​𝐯\+𝐯⊤​B​u​\(t\)≤−γmin​‖𝐯‖2\+‖B‖​‖𝐯‖​‖u​\(t\)‖,\\dot\{E\}=\-\\mathbf\{v\}^\{\\top\}\\Gamma\\mathbf\{v\}\+\\mathbf\{v\}^\{\\top\}Bu\(t\)\\leq\-\\gamma\_\{\\min\}\\left\\lVert\\mathbf\{v\}\\right\\rVert^\{2\}\+\\left\\lVert B\\right\\rVert\\,\\left\\lVert\\mathbf\{v\}\\right\\rVert\\,\\left\\lVert u\(t\)\\right\\rVert,so for bounded input‖u​\(t\)‖≤U\\left\\lVert u\(t\)\\right\\rVert\\leq Uthe sublevel set\{E≤E⋆\}\\\{E\\leq E\_\{\\star\}\\\}withE⋆=Φmin\+‖B‖2​U2/\(2​mmin​γmin2\)E\_\{\\star\}=\\Phi\_\{\\min\}\+\\left\\lVert B\\right\\rVert^\{2\}U^\{2\}/\(2m\_\{\\min\}\\gamma\_\{\\min\}^\{2\}\)is forward\-invariant; trajectories cannot escape to infinity\. The conservative force contributes𝐱˙⊤​∇𝐱Φ\\dot\{\\mathbf\{x\}\}^\{\\top\}\\nabla\_\{\\mathbf\{x\}\}\\Phi, which cancels inE˙\\dot\{E\}; Young’s inequality and radial unboundedness give the invariant set\. We confirmed this numerically: a symplectic rollout of the undamped network after a brief input kick conserves total energy to within a bounded14%14\\%oscillation with no secular drift over400400steps, the expected signature of a symplectic integrator on a conservative system\.

## 4Pre\-Registered Claims

#### H1 \(Memory ceiling\)\.

The backward gradient magnitude of the dissipative system decays at rateγmin/2\\gamma\_\{\\min\}/2, so the usable memory horizon obeysH≤2γmin​log⁡1ϵH\\leq\\frac\{2\}\{\\gamma\_\{\\min\}\}\\log\\frac\{1\}\{\\epsilon\}\.

#### H2 \(Stability floor\)\.

Forward sensitivities, and the gradients solved backward, grow aseλmax​Te^\{\\lambda\_\{\\max\}T\}\. Sinceλmax\\lambda\_\{\\max\}increases as damping falls toward the Hamiltonian limit, numerically usable gradients impose a lower boundγ≥γ⋆​\(T\)\\gamma\\geq\\gamma\_\{\\star\}\(T\)\. The largest Lyapunov exponent is monotone decreasing in the damping and crosses zero atγ⋆\\gamma\_\{\\star\}\.

#### H3 \(Trilemma and band closure\)\.

The feasible training bandγ⋆≤γ≤γ¯​\(H\)\\gamma\_\{\\star\}\\leq\\gamma\\leq\\bar\{\\gamma\}\(H\)has width decreasing inHHand closes at a critical horizonHcH\_\{c\}\.

#### H4 \(Learned beats fixed only inside the band\)\.

Under matched compute, training the substrate outperforms freezing it for smallHH, the advantage closes asHHgrows, and the two cross nearHcH\_\{c\}\. A learned model trained freely will settle its damping near the stability floor \(the edge of chaos\)\.

#### H5 \(Interpretability is dynamical, not static\)\.

Trained nonlinear oscillator networks retain phase\-space interpretability but lose static weight interpretability; the input–output map is not legible from the equations\.

## 5The Memory–Stability–Expressivity Trilemma

Training by gradient descent through the ODE propagates a costate backward in time,𝐚˙=−\(∂𝐟/∂𝐳\)⊤​𝐚\\dot\{\\mathbf\{a\}\}=\-\(\\partial\\mathbf\{f\}/\\partial\\mathbf\{z\}\)^\{\\top\}\\mathbf\{a\}\[[1](https://arxiv.org/html/2606.09929#bib.bib1)\]\. Linearizing about an operating point with effective stiffnessK=∇𝐱2ΦK=\\nabla^\{2\}\_\{\\mathbf\{x\}\}\\Phi, the Jacobian has the block form\[0I−M−1​K−M−1​Γ\]\\big\[\\begin\{smallmatrix\}0&I\\\\ \-M^\{\-1\}K&\-M^\{\-1\}\\Gamma\\end\{smallmatrix\}\\big\], whose eigenvalues for a modal frequencyω\\omegaand modal dampingγ\\gammaares±=−γ2±γ24−ω2s\_\{\\pm\}=\-\\tfrac\{\\gamma\}\{2\}\\pm\\sqrt\{\\tfrac\{\\gamma^\{2\}\}\{4\}\-\\omega^\{2\}\}, with real part−γ/2\-\\gamma/2in the underdamped regime\.

### 5\.1Backward Decay Sets the Memory Ceiling

The gradient magnitude of the linearized dissipative system obeys‖𝐚​\(t\)‖∼e−\(γmin/2\)​\(T−t\)\\left\\lVert\\mathbf\{a\}\(t\)\\right\\rVert\\sim e^\{\-\(\\gamma\_\{\\min\}/2\)\(T\-t\)\}\. Defining the memory horizonHHas the lag at which gradient magnitude falls below a fractionϵ\\epsilonof its terminal value yields the ceilingH≤2γmin​log⁡1ϵH\\leq\\frac\{2\}\{\\gamma\_\{\\min\}\}\\log\\frac\{1\}\{\\epsilon\}\. Figure[1](https://arxiv.org/html/2606.09929#S5.F1)shows the decay envelope at the three dampings that recur in our experiments\. Stronger damping caps how far credit can propagate, so long memory requires smallγ\\gamma\.

![Refer to caption](https://arxiv.org/html/2606.09929v1/x1.png)Figure 1:Backward gradient magnitude versus lag, plotted at the frozen\-substrate damping \(γ=0\.077\\gamma\{=\}0\.077\), the measured stability floor \(γ⋆=0\.121\\gamma\_\{\\star\}\{=\}0\.121\), and an overdamped value \(γ=0\.40\\gamma\{=\}0\.40\)\. The decay rate isγ/2\\gamma/2; the lag at which a curve crosses theϵ\\epsilonthreshold \(dotted\) is the memory horizonHH\. Weakly damped systems remember longer, which is exactly why they are demanded for long\-memory tasks and exactly why they court the instability of the next subsection\.Finding 1: Memory ceiling\.The usable memory horizon is bounded above by2γmin​log⁡1ϵ\\frac\{2\}\{\\gamma\_\{\\min\}\}\\log\\frac\{1\}\{\\epsilon\}, inversely proportional to the damping\. Memory is purchased only by reducing dissipation\. H1 is established analytically\.

### 5\.2Forward Growth Sets the Stability Floor

Driving damping toward zero to extend memory pushes the system toward the Hamiltonian limit, where the volume\-preserving flow of coupled nonlinear oscillators is generically chaotic\. Withλmax\\lambda\_\{\\max\}the largest Lyapunov exponent, the forward sensitivity and the backward gradient both grow aseλmax​Te^\{\\lambda\_\{\\max\}T\}, so gradients remain usable over horizonTTonly ifλmax≲1T​log⁡1δmach\\lambda\_\{\\max\}\\lesssim\\frac\{1\}\{T\}\\log\\frac\{1\}\{\\delta\_\{\\mathrm\{mach\}\}\}\. Thatλmax\\lambda\_\{\\max\}rises asγ→0\\gamma\\to 0is the load\-bearing assumption of the theorem; we verify it directly in Section[9](https://arxiv.org/html/2606.09929#S9)rather than assume it\.

Finding 2: Stability floor \(assumption to be measured\)\.Numerically usable gradients over horizonTTrequire damping no smaller thanγ⋆\\gamma\_\{\\star\}, the value at which the largest Lyapunov exponent crosses zero\. The existence and location ofγ⋆\\gamma\_\{\\star\}are established empirically in Section[9](https://arxiv.org/html/2606.09929#S9)\.

### 5\.3The Feasible Band and Its Closure

#### Theorem 1 \(Trilemma\)\.

Let a task require memory horizonHHover a windowT≥HT\\geq H\. Stable training of the oscillator network requires the modal damping to satisfy simultaneouslyγ≤γ¯​\(H\)=2H​log⁡1ϵ\\gamma\\leq\\bar\{\\gamma\}\(H\)=\\frac\{2\}\{H\}\\log\\frac\{1\}\{\\epsilon\}\(memory, Finding 1\) andγ≥γ⋆\\gamma\\geq\\gamma\_\{\\star\}\(stability, Finding 2\)\. Becauseλmax​\(γ\)\\lambda\_\{\\max\}\(\\gamma\)is non\-increasing inγ\\gamma, the stability constraint is a lower bound and the memory constraint an upper bound\. A stable trainable regime exists iffγ⋆≤γ≤γ¯​\(H\)\\gamma\_\{\\star\}\\leq\\gamma\\leq\\bar\{\\gamma\}\(H\), and the band widthγ¯​\(H\)−γ⋆\\bar\{\\gamma\}\(H\)\-\\gamma\_\{\\star\}is strictly decreasing inHH\(derivative−2H2​log⁡1ϵ<0\-\\tfrac\{2\}\{H^\{2\}\}\\log\\tfrac\{1\}\{\\epsilon\}<0\)\. At the critical horizonHcH\_\{c\}whereγ¯​\(Hc\)=γ⋆\\bar\{\\gamma\}\(H\_\{c\}\)=\\gamma\_\{\\star\}the band closes, and forH\>HcH\>H\_\{c\}no damping trains stably\. Figure[2](https://arxiv.org/html/2606.09929#S5.F2)shows the band, the measured floor, and the settled dampings of trained models\.

![Refer to caption](https://arxiv.org/html/2606.09929v1/x2.png)Figure 2:The trilemma feasible band\. The memory ceilingγ¯​\(H\)\\bar\{\\gamma\}\(H\)\(navy\) falls with horizon; the stability floorγ⋆=0\.121\\gamma\_\{\\star\}=0\.121\(gold dashed\) is the*measured*zero crossing of the Lyapunov exponent \(Section[9](https://arxiv.org/html/2606.09929#S9)\)\. Red points are the settled dampings of the freely trained learned models at each horizon \(Section[10](https://arxiv.org/html/2606.09929#S10)\); they sit near the floor, evidence that training seeks the edge of chaos\. The band\-theoretic closureHcH\_\{c\}forϵ=0\.05\\epsilon\{=\}0\.05is shown; Section[11](https://arxiv.org/html/2606.09929#S11)discusses why the empirical learned\-versus\-frozen crossover occurs well before this analytic value\.Finding 3: The trilemma and why the field freezes the substrate\.Memory, stability, and expressivity are jointly bound by the damping; stable substrate training is confined to a band that contracts with horizon and vanishes atHcH\_\{c\}\. ForH\>HcH\>H\_\{c\}no damping permits stable end\-to\-end training, so one is forced either to push memory into a linear undamped backbone \(the state\-space\-model solution\) or to freeze the substrate and fit only a readout \(the reservoir solution\)\. Freezing is the rational response to a region of the trilemma in which learning is impossible\. H3 is established analytically and tested in Section[10](https://arxiv.org/html/2606.09929#S10)\.

## 6Edge\-of\-Chaos Training

The trilemma says the useful regime is a thin band near the edge of chaos\. Three ingredients keep optimization there\. First,*energy\-bounded springs*: the saturating potential of Section[3](https://arxiv.org/html/2606.09929#S3)gives the Lyapunov function of Proposition 1 and removes the finite\-time blow\-up that an unconstrained quartic potential would permit\. Second,*symplectic integration*: we integrate the conservative part with a velocity\-Verlet scheme and treat damping and forcing by an exponential\-and\-kick splitting, so the integrator conserves a shadow Hamiltonian and the gradients reflect the true dynamics rather than discretization drift\[[6](https://arxiv.org/html/2606.09929#bib.bib6)\]\. The monitored energy of Proposition 1 doubles as a runtime correctness check\. Third,*edge\-of\-chaos regularization*: an optional penaltyβ​\(λmax^−λmax†\)2\\beta\(\\widehat\{\\lambda\_\{\\max\}\}\-\\lambda\_\{\\max\}^\{\\dagger\}\)^\{2\}with a small negative target steers the damping toward the stability floor; in our experiments we found that free training already settles near the floor \(Figure[2](https://arxiv.org/html/2606.09929#S5.F2)\), so the explicit penalty served mainly as a safeguard against the chaotic basin\. Stiffness is controlled bysoftplus\\mathrm\{softplus\}reparameterization and non\-dimensionalization so that natural frequencies are order one at initialization\.

Finding 4: Training seeks the edge of chaos\.Across all nine horizons, freely trained learned models settled at mean damping in the range0\.070\.07to0\.230\.23, clustered around the measured stability floorγ⋆=0\.121\\gamma\_\{\\star\}=0\.121rather than at either extreme \(Table[1](https://arxiv.org/html/2606.09929#S10.T1)\)\. The procedure occupies the feasible band the trilemma defines without being told to\. This is the operational half of H4\.

## 7Dynamical Interpretability and Its Limits

The motivation that draws practitioners to mechanical models is interpretability, and we are precise about what survives training\. What is retained: the state is physically named \(positions, velocities, modal energies\), the energyE​\(𝐳\)E\(\\mathbf\{z\}\)is a monitorable near\-invariant, and the computation lives in a phase space one can plot\. This*dynamical*interpretability genuinely exceeds what a weight matrix offers\. What is lost: once the model is nonlinear, coupled, and trained, the measured Lyapunov exponents are positive over much of the damping range \(Section[9](https://arxiv.org/html/2606.09929#S9)\), so why a given input yields a given output is no more legible than for any chaotic system\. The Lorenz system is three equations and remains unpredictable; trained oscillator networks live near the edge of chaos by design\. We therefore frame the benefit as a trade of static weight\-interpretability for phase\-space interpretability, a real but bounded gain, not transparency by virtue of being physics\. This is H5, and it is the honest correction to the premise that motivates the whole enterprise\.

## 8Experimental Setup

### 8\.1Implementation

We implement the oscillator network in PyTorch with full autograd through a custom symplectic integrator \(velocity\-Verlet on the conservative force, exponential decay for damping, half\-kick forcing, time stepd​tdtbetween0\.20\.2and0\.250\.25\)\. The coupling graph is a ring backbone plus random chords \(≈3\\approx 3extra edges per node\), fixed across conditions\. All experiments run single\-threaded on CPU; the full study below completes in approximately fifteen minutes of compute\.

### 8\.2Lyapunov Estimation

We estimate the largest Lyapunov exponent with the Benettin algorithm: a reference and a perturbed trajectory are co\-integrated, the separation is renormalized at fixed intervals, and the accumulated log\-growth is divided by elapsed time\. We use4040renormalizations over500500steps with a weak input drive \(uu\-amplitude0\.150\.15\) so that the estimate reflects the driven operating regime rather than the autonomous rest state\.

### 8\.3Task

The primary task is*delayed single\-bit recall*: the input is a length\-\(H\+6\)\(H\{\+\}6\)stream of±1\\pm 1bits, and the target is the sign of the bitHHsteps before the end, read out by a linear map from the final state\. This isolates memory cleanly: accuracy as a function ofHHmeasures how far back the network can carry one bit\. \(We also attempted a delayed\-parity variant, which proved too hard for all conditions at this network size and is reported as a negative result in Section[11](https://arxiv.org/html/2606.09929#S11)\.\)

### 8\.4Conditions and Matching

Learnedtrains all substrate parameters plus the readout;Fixedfreezes a random confined substrate and trains only the readout \(the reservoir baseline\);LinearSSMis a linear state\-space modelzt\+1=A​zt\+B​utz\_\{t\+1\}=Az\_\{t\}\+Bu\_\{t\}of matched state size \(2​N=482N=48\)\. All conditions useN=20N=20oscillators,400400Adam steps, batch size6464, identical task data streams, and gradient clipping at norm55\. We report mean and standard deviation over three seeds\.

## 9Experiment 1: The Lyapunov Exponent vs\. Damping

We sweep the damping over eleven values from0\.0050\.005to0\.60\.6and measureλmax\\lambda\_\{\\max\}at each, five seeds per point\. Figure[3](https://arxiv.org/html/2606.09929#S9.F3)shows the result\.

![Refer to caption](https://arxiv.org/html/2606.09929v1/x3.png)Figure 3:Measured largest Lyapunov exponent versus modal damping \(log abscissa\), mean±\\pmstandard deviation over five seeds,N=20N\{=\}20\. At low damping the coupled saturating\-spring network is chaotic \(λmax\>0\\lambda\_\{\\max\}\>0, red region\) and gradients explode; sufficient damping makes the flow contracting \(λmax<0\\lambda\_\{\\max\}<0, gold region\) and trainable\. The exponent is monotone decreasing in the damping and crosses zero atγ⋆=0\.121\\gamma\_\{\\star\}=0\.121, establishing the stability floor that Theorem 1 assumes\.The exponent falls monotonically from\+0\.23\+0\.23atγ=0\.005\\gamma=0\.005to−0\.028\-0\.028atγ=0\.6\\gamma=0\.6, crossing zero atγ⋆=0\.121\\gamma\_\{\\star\}=0\.121\(linear interpolation between the bracketing measured points\)\. The positive branch confirms that the undamped saturating\-spring network is genuinely chaotic, not merely oscillatory, so the stability constraint is real rather than hypothetical\.

Finding 5: The stability floor exists and the monotonicity assumption holds\.The largest Lyapunov exponent is monotone decreasing in the damping and crosses zero atγ⋆=0\.121±0\.01\\gamma\_\{\\star\}=0\.121\\pm 0\.01\. This is the measured confirmation of Finding 2 and supplies the lower bound that Theorem 1 requires; the monotonicity that was an assumption in the proof is an empirical fact for this substrate\. H2 is confirmed\.

## 10Experiment 2: Learned vs\. Frozen Substrate Across Horizons

We run all three conditions on delayed recall at nine horizonsH∈\{1,2,3,4,6,8,10,13,16\}H\\in\\\{1,2,3,4,6,8,10,13,16\\\}\. Table[1](https://arxiv.org/html/2606.09929#S10.T1)reports accuracy and the learned model’s settled damping; Figure[4](https://arxiv.org/html/2606.09929#S10.F4)plots accuracy against horizon\.

Table 1:Delayed\-recall accuracy by condition and horizon \(N=20N\{=\}20, three seeds,400400training steps, compute\-matched\)\. The Learned advantage over Fixed is large at short horizons, closes byH=8H\{=\}8–1010, and reverses atH=13H\{=\}13\. The learned model’s settled damping stays near the measured stability floorγ⋆=0\.121\\gamma\_\{\\star\}=0\.121throughout\.![Refer to caption](https://arxiv.org/html/2606.09929v1/x4.png)Figure 4:Measured recall accuracy versus memory horizon, mean±\\pmstandard deviation over three seeds\. The learned substrate \(navy\) dominates the frozen substrate \(gold\) at short horizons; the gap closes through the middle of the range and the two cross nearH≈11H\\approx 11\(dotted line, interpolated crossoverHc≈10\.9H\_\{c\}\\approx 10\.9\)\. The matched\-size linear state\-space model \(grey\) degrades fastest, indicating that on this task a trained nonlinear substrate and even a fixed nonlinear reservoir carry a single bit further than a deep linear recurrence of the same state size\.Three observations\. First, at short horizons the learned substrate is decisively better: perfect recall atH≤2H\\leq 2against0\.820\.82–0\.880\.88for the frozen substrate, a1212–1818point gap\. Second, the gap closes monotonically and the curves cross: atH=8H=8the difference is\+0\.036\+0\.036, atH=10H=10it is\+0\.014\+0\.014, and atH=13H=13it is−0\.034\-0\.034, with interpolated crossoverHc≈10\.9H\_\{c\}\\approx 10\.9\. Beyond the crossover the frozen substrate is at least as good, because the learned model can no longer convert its tunability into usable long\-range gradient\. Third, the linear state\-space model collapses fastest of all, falling to chance \(0\.500\.50\) byH=8H=8; in this regime a deep linear recurrence suffers its own gradient decay and the nonlinear transient of even a frozen reservoir is more useful\.

Finding 6: Learning the substrate beats freezing it only inside the band\.The learned substrate outperforms the frozen one forH<HcH<H\_\{c\}with the advantage closing asHHgrows, and the two cross atHc≈10\.9H\_\{c\}\\approx 10\.9, beyond which learning no longer helps\. The crossover is the predicted qualitative signature of band closure\. H4 is confirmed in its central prediction; the quantitative location ofHcH\_\{c\}relative to the analytic ceiling is discussed below\.

## 11Discussion

#### The result\.

The trilemma’s qualitative prediction is borne out by direct measurement\. The Lyapunov exponent is monotone in the damping and crosses zero at a well\-defined floor; trained models settle near that floor of their own accord; and the advantage of learning the substrate over freezing it is large at short memory horizons, closes as the horizon grows, and reverses past a crossover\. The frozen\-substrate reservoir, which by construction cannot tune its dynamics, is the rational choice once the horizon exceeds the point where learning can no longer extract usable gradient\. This is the geometric account the paper set out to establish, and it is confirmed rather than merely argued\.

#### The analytic/empirical gap inHcH\_\{c\}\.

The simple memory\-ceiling formulaγ¯​\(H\)=2H​log⁡1ϵ\\bar\{\\gamma\}\(H\)=\\frac\{2\}\{H\}\\log\\frac\{1\}\{\\epsilon\}withϵ=0\.05\\epsilon=0\.05and the measured floorγ⋆=0\.121\\gamma\_\{\\star\}=0\.121predicts band closure nearH≈49H\\approx 49, whereas the measured learned\-versus\-frozen crossover is nearH≈11H\\approx 11, a factor of roughly five\. We report this rather than tuneϵ\\epsilonto match\. The discrepancy has a clear interpretation: the analytic ceiling marks where the gradient falls below a detectability threshold, but learning requires gradient well above mere detectability, and the effectiveϵ\\epsilonfor*learnability*is far larger than the0\.050\.05used for illustration\. The qualitative law \(a contracting band, a crossover, learning helping only inside it\) is robust; the precise constant depends on optimization details the linear theory does not capture\. Closing this gap quantitatively, by deriving an effectiveϵ\\epsilonfrom the optimizer and task signal\-to\-noise, is the most concrete piece of theory the experiments call for\.

#### The parity negative result\.

Our first task was delayed*parity*rather than recall\. No condition learned it above chance atN=20N=20within the compute budget, and the learned model responded by driving its damping high \(overdamping itself into amnesia\)\. Parity over a window demands both long memory and a hard nonlinearity simultaneously, which places it deep in the contested region of the trilemma; the negative result is consistent with the theory but uninformative about the crossover, which is why we switched to recall as the clean memory probe\. We report it because it is the kind of result that silently disappears from papers and should not\.

#### Limitations\.

The network is small \(N=20N=20\) and single\-threaded; the horizons are modest; the Lyapunov estimate uses a single tangent vector\. The memory\-ceiling bound is derived by local linearization, and the effective\-ϵ\\epsilongap above shows its quantitative limits\. The damping is diagonal; general dissipation matrices would turn the stability floor into a spectral condition\.

#### Next steps\.

Derive the effectiveϵ\\epsilonthat reconciles the analytic and empiricalHcH\_\{c\}; scaleNNand the horizon range on parallel hardware; add the full differentiable Benettin penalty \(rather than the damping\-floor surrogate used here\) and test whether it extends the usable band; and replace diagonal damping with a learned dissipation matrix to study the spectral form of the stability floor\.

## References

- Chen et al\. \[2018\]Ricky T\. Q\. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud\.*Neural Ordinary Differential Equations*\. NeurIPS, 2018\.[https://arxiv\.org/abs/1806\.07366](https://arxiv.org/abs/1806.07366)\.
- Cranmer et al\. \[2020\]Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho\.*Lagrangian Neural Networks*\. ICLR Deep Differential Equations Workshop, 2020\.[https://arxiv\.org/abs/2003\.04630](https://arxiv.org/abs/2003.04630)\.
- Greydanus et al\. \[2019\]Sam Greydanus, Misko Dzamba, and Jason Yosinski\.*Hamiltonian Neural Networks*\. NeurIPS, 2019\.[https://arxiv\.org/abs/1906\.01563](https://arxiv.org/abs/1906.01563)\.
- Gu et al\. \[2022\]Albert Gu, Karan Goel, and Christopher Ré\.*Efficiently Modeling Long Sequences with Structured State Spaces*\. ICLR, 2022\.[https://arxiv\.org/abs/2111\.00396](https://arxiv.org/abs/2111.00396)\.
- Gu and Dao \[2023\]Albert Gu and Tri Dao\.*Mamba: Linear\-Time Sequence Modeling with Selective State Spaces*\. arXiv:2312\.00752, 2023\.[https://arxiv\.org/abs/2312\.00752](https://arxiv.org/abs/2312.00752)\.
- Hairer et al\. \[2006\]Ernst Hairer, Christian Lubich, and Gerhard Wanner\.*Geometric Numerical Integration: Structure\-Preserving Algorithms for Ordinary Differential Equations*\. Springer, 2nd edition, 2006\.
- Nakajima \[2020\]Kohei Nakajima\.*Physical Reservoir Computing: An Introductory Perspective*\. Japanese Journal of Applied Physics, 59\(6\):060501, 2020\.[https://arxiv\.org/abs/2005\.00992](https://arxiv.org/abs/2005.00992)\.
- Tanaka et al\. \[2019\]Gouhei Tanaka, Toshiyuki Yamane, Jean Benoit Héroux, Ryosho Nakane, Naoki Kanazawa, Seiji Takeda, Hidetoshi Numata, Daiju Nakano, and Akira Hirose\.*Recent Advances in Physical Reservoir Computing: A Review*\. Neural Networks, 115:100–123, 2019\.[https://arxiv\.org/abs/1808\.04962](https://arxiv.org/abs/1808.04962)\.

Similar Articles