Spectral Distillation: From Nonlinear Dynamics to Linear State-Space Models

arXiv cs.LG Papers

Summary

This paper proposes a provable two-stage pipeline that distills nonlinear dynamical systems into compact linear state-space models using convex optimization, with theoretical guarantees and experiments on LDS benchmarks and MuJoCo.

arXiv:2608.05416v1 Announce Type: new Abstract: Can nonlinear dynamical systems be learned through a compact linear state-space representation, without directly solving a non-convex system-identification problem? We give a provable pipeline for doing so. Starting from observations of an unknown nonlinear dynamical system, we first learn an implicit spectral predictor using Observation Spectral Filtering (OSF), a convex method that competes with the best linear observer for the system. We then apply spectral-to-LDS distillation to convert this predictor into an explicit recurrent linear dynamical system. Our main theorem shows that the average prediction error of the distilled LDS decomposes into an exponentially-small distillation term and the OSF learning term governed by the Luenberger complexity of the best observer. The guarantee is dimension-free: it depends on observer complexity rather than on the latent dimension needed to represent the nonlinear system. To our knowledge, this yields the first end-to-end provable method for extracting a best-in-hindsight LDS representation of nonlinear dynamics through convex learning followed by provable distillation. Experiments on linear LDS benchmarks and MuJoCo behavior cloning show that the train-then-distill pipeline produces compact LDS predictors that match or outperform directly trained baselines.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:50 AM

# Spectral Distillation: From Nonlinear Dynamics to Linear State-Space Models
Source: [https://arxiv.org/html/2608.05416](https://arxiv.org/html/2608.05416)
Liane Galanti Princeton University &Devan Shah11footnotemark:1 Princeton University &Shlomo Fortgang Princeton University &Elad Hazan Princeton UniversityEqual contribution; author order chosen alphabetically\. Corresponding emails:galanti@princeton\.edu,devan\.shah@princeton\.edu\.

###### Abstract

Can nonlinear dynamical systems be learned through a compact linear state\-space representation, without directly solving a non\-convex system\-identification problem? We give a provable pipeline for doing so\. Starting from observations of an unknown nonlinear dynamical system, we first learn an implicit spectral predictor using Observation Spectral Filtering \(OSF\), a convex method that competes with the best linear observer for the system\. We then apply spectral\-to\-LDS distillation to convert this predictor into an explicit recurrent linear dynamical system\. Our main theorem shows that the average prediction error of the distilled LDS decomposes into an exponentially\-small distillation term and the OSF learning term governed by the Luenberger complexity of the best observer\. The guarantee is dimension\-free: it depends on observer complexity rather than on the latent dimension needed to represent the nonlinear system\. To our knowledge, this yields the first end\-to\-end provable method for extracting a best\-in\-hindsight LDS representation of nonlinear dynamics through convex learning followed by provable distillation\. Experiments on linear LDS benchmarks and MuJoCo behavior cloning show that the train\-then\-distill pipeline produces compact LDS predictors that match or outperform directly trained baselines\.

## 1Introduction

Modern sequence modeling and control hinge on the ability to represent and learn dynamical systems\. Linear dynamical systems \(LDS\) occupy a central role in this landscape: they admit compact state\-space representations, efficient simulation, and a rich spectral theory that enables both analysis and algorithm design\. However, most real\-world systems are inherently nonlinear, and learning either a nonlinear state\-space model or even its best linear approximation remains a fundamentally non\-convex and fragile problem\.

This paper asks a basic question:*can we extract a simple linear dynamical system that faithfully represents a nonlinear one, without ever solving a non\-convex system identification problem?*

Our answer is affirmative\. We present a provable nonlinear\-to\-LDS distillation pipeline that converts observations from an unknown nonlinear system into a compact linear dynamical system, using only convex optimization followed by a structured distillation step\.

The key idea is to separate learning from representation\. Rather than attempting to fit a state\-space model directly, we first learn an implicit predictor using Observation Spectral Filtering \(OSF\), a convex method that operates in a fixed temporal feature space\. OSF does not attempt to recover latent states; instead, it competes with the best linear observer that could explain the data, achieving regret bounds that depend only on an intrinsic measure of observability complexity\.

We then convert this predictor into an explicit state\-space model\. Leveraging recent results on spectral representations of dynamical systems, we show how a learned spectral predictor can be distilled into a finite\-dimensional LDS whose impulse responses approximate those of the predictor up to exponentially small error\. The result is a standard recurrent model with memory and per\-step cost depending only on the observation and output dimension, suitable for long\-horizon simulation and control\.

### 1\.1Our Contributions

This paper makes the following contributions\.

- •A nonlinear\-to\-LDS distillation pipeline\.We give a two\-stage procedure that first learns a nonlinear dynamical system through the convex OSF feature class, and then distills the resulting spectral predictor into an explicit recurrent LDS\.
- •A dimension\-free prediction guarantee\.We prove that withk=Θ​\(log2⁡T\)k=\\Theta\(\\log^\{2\}T\)spectral filters andh≥kh\\geq kSpectraLDS poles, the distilled LDS satisfies 1T​∑t=1T‖yt−y^tLDS‖2≤O~​\(Q⋆2​log⁡\(Q⋆\)​log7⁡TT\+λmax​\(F\)2​exp⁡\(−2​klog⁡T\)\)\.\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}\\\|^\{2\}\\leq\\widetilde\{O\}\\\!\\left\(\\frac\{Q\_\{\\star\}^\{2\}\\log\(Q\_\{\\star\}\)\\log^\{7\}T\}\{\\sqrt\{T\}\}\+\\lambda\_\{\\max\}\(F\)^\{2\}\\exp\\\!\\left\(\-\\frac\{2k\}\{\\log T\}\\right\)\\right\)\.Thus the error decomposes into an exponentially small distillation term and the OSF learning term governed by the Luenberger complexityQ⋆Q\_\{\\star\}\.
- •An explicit compact representation\.Unlike OSF, which is an implicit convolutional predictor, the output of our method is a standard LDS withO​\(h⋅\(di​n\+do​u​t\)\)O\(h\\cdot\(d\_\{in\}\+d\_\{out\}\)\)memory andO​\(h⋅\(di​n\+do​u​t\)\)O\(h\\cdot\(d\_\{in\}\+d\_\{out\}\)\)per\-step simulation cost\. The guarantee depends on observer complexity rather than on the potentially enormous dimension of a nonlinear lift or discretization\.
- •Empirical validation of train\-then\-distill\.We test the pipeline on linear LDS benchmarks and MuJoCo behavior cloning\. Across these settings, the distilled LDS preserves the behavior of the spectral predictor and often outperforms directly trained LDS baselines\.

Together, these results give a principled route from nonlinear observations to compact linear state\-space models: learn in a convex spectral space, then distill into an interpretable recurrent LDS\.

### 1\.2Related Work

#### Linear dynamical system identification\.

Classical system identificationLjung \([1999](https://arxiv.org/html/2608.05416#bib.bib2)\)fits state\-space parameters\(A,B,C\)\(A,B,C\)directly, typically through non\-convex optimization\. Recent work has obtained non\-asymptotic guarantees for LDS estimation under Gaussian noiseHardtet al\.\([2018](https://arxiv.org/html/2608.05416#bib.bib5)\); Simchowitzet al\.\([2018](https://arxiv.org/html/2608.05416#bib.bib6)\); Oymak and Ozay \([2019](https://arxiv.org/html/2608.05416#bib.bib7)\); Sarkar and Rakhlin \([2019](https://arxiv.org/html/2608.05416#bib.bib8)\); Tsiamis and Pappas \([2019](https://arxiv.org/html/2608.05416#bib.bib9)\), with sample complexity scaling polynomially in the state dimension and condition number of the system\. These results crucially assume the data are generated by a linear system; extending them to nonlinear settings has remained open\.

#### Spectral filtering\.

Spectral filteringHazanet al\.\([2017](https://arxiv.org/html/2608.05416#bib.bib1),[2018](https://arxiv.org/html/2608.05416#bib.bib39)\)sidesteps non\-convex identification by learning in a fixed convolutional feature space derived from the Hankel matrix\. The OSF extensionDogariuet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib23)\)combines this with a Luenberger\-observer analysis to obtain regret guarantees for nonlinear systems through improper learning\. SpectraLDSShahet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib24)\)converts spectral predictors into recurrent LDS form\. Our work composes these ingredients into an end\-to\-end provable pipeline with a unified guarantee, and we describe these techniques more in the next section\.

#### Koopman operator methods\.

Koopman theoryMezić \([2005](https://arxiv.org/html/2608.05416#bib.bib10)\)represents nonlinear dynamics as a linear operator on an infinite\-dimensional observable space, and finite\-dimensional approximations underlie methods such as EDMDWilliamset al\.\([2015](https://arxiv.org/html/2608.05416#bib.bib11)\)and deep Koopman embeddingsLuschet al\.\([2018](https://arxiv.org/html/2608.05416#bib.bib12)\); Kosticet al\.\([2022](https://arxiv.org/html/2608.05416#bib.bib13)\)\. These approaches typically require choosing a dictionary of observables and offer mostly asymptotic or empirical guarantees\. Our pipeline can be viewed in a similar spirit, but the spectral\-filter dictionary is universal and the guarantees are non\-asymptotic\.

#### Modern state\-space sequence models\.

Structured state\-space models such as S4Guet al\.\([2022](https://arxiv.org/html/2608.05416#bib.bib14)\), S5Smithet al\.\([2023](https://arxiv.org/html/2608.05416#bib.bib15)\), the LRUOrvietoet al\.\([2023](https://arxiv.org/html/2608.05416#bib.bib16)\), and MambaGu and Dao \([2023](https://arxiv.org/html/2608.05416#bib.bib17)\)have shown that linear recurrences with carefully chosen parameterizations can match or exceed Transformer performance on long\-sequence tasks\. These works focus on architectural and empirical advances rather than provable system\-identification guarantees\. Our distilled LDS produces a compact recurrence in the same family, but with a provable approximation guarantee for nonlinear data\-generating processes\.

## 2Background

We recall the two ingredients used by our pipeline\. Observation Spectral Filtering \(OSF\)Dogariuet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib23)\)learns a convex spectral predictorHazanet al\.\([2017](https://arxiv.org/html/2608.05416#bib.bib1)\)from past inputs and observations\. Spectral\-to\-LDS distillationShahet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib24)\)then realizes the learned spectral predictor as an explicit recurrent state\-space model\.

### 2\.1Linear systems as convolutional predictors

A controlled linear dynamical system maps inputsut∈ℝduu\_\{t\}\\in\\mathbb\{R\}^\{d\_\{u\}\}to outputsyt∈ℝdyy\_\{t\}\\in\\mathbb\{R\}^\{d\_\{y\}\}through a hidden statext∈ℝdxx\_\{t\}\\in\\mathbb\{R\}^\{d\_\{x\}\}:

xt\+1=A​xt\+B​ut,yt=C​xt\+D​ut\.x\_\{t\+1\}=Ax\_\{t\}\+Bu\_\{t\},\\qquad y\_\{t\}=Cx\_\{t\}\+Du\_\{t\}\.\(2\.1\)The direct termD​utDu\_\{t\}can be replaced by an expanded hidden state, and so we often considerD=0D=0\. Takingx0=0x\_\{0\}=0and suppressing the direct term, the output expands as

yt=∑ℓ=0t−1C​Aℓ​B​ut−ℓ\.y\_\{t\}=\\sum\_\{\\ell=0\}^\{t\-1\}CA^\{\\ell\}B\\,u\_\{t\-\\ell\}\.\(2\.2\)Thus an LDS is equivalently described by its impulse response\{C​Aℓ​B\}ℓ≥0\\\{CA^\{\\ell\}B\\\}\_\{\\ell\\geq 0\}\.

This representation exposes the main difficulty in directly fitting LDSs\. Long\-memory systems require accurately learning high powers ofAA, and the map fromAAto the sequence\{Aℓ\}ℓ≥0\\\{A^\{\\ell\}\\\}\_\{\\ell\\geq 0\}is non\-convex and poorly conditioned whenρ​\(A\)\\rho\(A\)is close to one\. Spectral filteringHazanet al\.\([2017](https://arxiv.org/html/2608.05416#bib.bib1)\)avoids this difficulty by learning in a fixed convolutional feature space instead of identifying the latent transition matrix\.

### 2\.2Spectral filtering

Forα∈\[0,1\]\\alpha\\in\[0,1\], define the length\-LLgeometric response

μL​\(α\)=\(1,α,α2,…,αL−1\)\.\\mu\_\{L\}\(\\alpha\)=\(1,\\alpha,\\alpha^\{2\},\\ldots,\\alpha^\{L\-1\}\)\.Consider the Hankel spectral basis obtained from

ZL=∫01μL​\(α\)​μL​\(α\)⊤​𝑑α\.Z\_\{L\}=\\int\_\{0\}^\{1\}\\mu\_\{L\}\(\\alpha\)\\mu\_\{L\}\(\\alpha\)^\{\\top\}\\,d\\alpha\.\(2\.3\)Letϕ1\+,…,ϕk\+\\phi^\{\+\}\_\{1\},\\ldots,\\phi^\{\+\}\_\{k\}be the topkkeigenvectors ofZLZ\_\{L\}, with corresponding eigenvaluesσ1,…,σk\\sigma\_\{1\},\\ldots,\\sigma\_\{k\}\. The Spectral Filtering algorithmHazanet al\.\([2017](https://arxiv.org/html/2608.05416#bib.bib1)\)is rooted in the observation that the span of these filters uniformly approximates length\-LLgeometric responses, with approximation error decaying exponentially ink/log⁡Lk/\\log L\.

For a signal streamsts\_\{t\}, define the spectral features

Ut,j\(s\)=⟨ϕj\+,st−1:t−L⟩=∑ℓ=1Lϕj\+​\(ℓ\)​st−ℓ,j=1,…,k\.U\_\{t,j\}^\{\(s\)\}=\\left\\langle\\phi\_\{j\}^\{\+\},s\_\{t\-1:t\-L\}\\right\\rangle=\\sum\_\{\\ell=1\}^\{L\}\\phi^\{\+\}\_\{j\}\(\\ell\)\\,s\_\{t\-\\ell\},\\qquad j=1,\\ldots,k\.\(2\.4\)A spectral predictor over this stream has the form

y^tSF=∑r=1mRr​st−r\+∑j=1kMj​Ut,j\(s\)\.\\hat\{y\}\_\{t\}^\{\\mathrm\{SF\}\}=\\sum\_\{r=1\}^\{m\}R\_\{r\}s\_\{t\-r\}\+\\sum\_\{j=1\}^\{k\}M\_\{j\}U\_\{t,j\}^\{\(s\)\}\.\(2\.5\)The temporal filtersϕj\+\\phi^\{\+\}\_\{j\}are fixed in advance, while the readout matricesRrR\_\{r\}andMjM\_\{j\}are learned\. Squared loss is therefore convex in the trainable parameters, even though the resulting predictor can represent long\-memory impulse responses for an LDS with PSD matrixAA\. To extend this result to a symmetric LDSHazanet al\.\([2018](https://arxiv.org/html/2608.05416#bib.bib39)\), where the transition matrixAAis symmetric, we can expand our filter set to include the negative spectral filtersϕi−=\[\(ϕi\+\)t⋅\(−1\)t\]t=0L−1\\phi\_\{i\}^\{\-\}=\[\(\\phi^\{\+\}\_\{i\}\)\_\{t\}\\cdot\(\-1\)^\{t\}\]\_\{t=0\}^\{L\-1\}\. Thus, when we refer tokkspectral filtersϕi\\phi\_\{i\}, we refer tok/2k/2positive filtersϕi\+\\phi^\{\+\}\_\{i\}andk/2k/2negative filtersϕi−\\phi\_\{i\}^\{\-\}\. For ease of explanation, we redefine the spectral features and predictor to use the combined filter setϕ1,…​ϕk\\phi\_\{1\},\\dots\\phi\_\{k\}\.

### 2\.3Observation Spectral Filtering

OSFDogariuet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib23)\)extends spectral filtering to the input\-output prediction setting\. At timett, the learner predictsyty\_\{t\}from both the control history and the observation history\. The form of the predictor is motivated by the Luenberger observer for a controlled linear system:

x^t\+1=A​x^t\+B​ut\+L​\(yt−C​x^t\);y^t=C​x^t\\hat\{x\}\_\{t\+1\}=A\\hat\{x\}\_\{t\}\+Bu\_\{t\}\+L\(y\_\{t\}\-C\\hat\{x\}\_\{t\}\);\\quad\\hat\{y\}\_\{t\}=C\\hat\{x\}\_\{t\}The observer error evolves according to the closed\-loop matrix

A~=A−L​C\.\\widetilde\{A\}=A\-LC\.Expanding the observer prediction expressesyty\_\{t\}as a combination of recent inputs, recent observations, and long convolutions over both streams, with convolution kernels generated by powers ofA~\\widetilde\{A\}\.

The difficulty faced by the best observer is summarized by the Luenberger complexity

Q​\(A~\)=11−ρ​\(A~\)​max⁡\{1,log⁡κdiag​\(A~\)\},Q⋆=infLQ​\(A−L​C\)\.Q\(\\widetilde\{A\}\)=\\frac\{1\}\{1\-\\rho\(\\widetilde\{A\}\)\}\\max\\\{1,\\log\\kappa\_\{\\mathrm\{diag\}\}\(\\widetilde\{A\}\)\\\},\\qquad Q\_\{\\star\}=\\inf\_\{L\}Q\(A\-LC\)\.\(2\.6\)This quantity depends on the spectral gap and diagonalization conditioning of the best closed\-loop observer, rather than directly on the hidden\-state dimension\.

Withkkspectral filters and finite memory ordernn, OSF uses the predictor

y^tOSF=\\displaystyle\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}=∑r=1nJr​ut−r\+∑j=1kσj1/4​Mj​Ut,j\(u\)\+∑r=1nPr​yt−r\+∑j=1kσj1/4​Nj​Ut,j\(y\),\\displaystyle\\sum\_\{r=1\}^\{n\}J\_\{r\}u\_\{t\-r\}\+\\sum\_\{j=1\}^\{k\}\\sigma\_\{j\}^\{1/4\}M\_\{j\}U\_\{t,j\}^\{\(u\)\}\+\\sum\_\{r=1\}^\{n\}P\_\{r\}y\_\{t\-r\}\+\\sum\_\{j=1\}^\{k\}\\sigma\_\{j\}^\{1/4\}N\_\{j\}U\_\{t,j\}^\{\(y\)\},\(2\.7\)where

Ut,j\(u\)=⟨ϕj,ut−1:t−L⟩,Ut,j\(y\)=⟨ϕj,yt−1:t−L⟩\.U\_\{t,j\}^\{\(u\)\}=\\left\\langle\\phi\_\{j\},u\_\{t\-1:t\-L\}\\right\\rangle,\\qquad U\_\{t,j\}^\{\(y\)\}=\\left\\langle\\phi\_\{j\},y\_\{t\-1:t\-L\}\\right\\rangle\.\(2\.8\)HereJr,Mj∈ℝdy×duJ\_\{r\},M\_\{j\}\\in\\mathbb\{R\}^\{d\_\{y\}\\times d\_\{u\}\}andPr,Nj∈ℝdy×dyP\_\{r\},N\_\{j\}\\in\\mathbb\{R\}^\{d\_\{y\}\\times d\_\{y\}\}\. The parameters\(J,M,P,N\)\(J,M,P,N\)are learned over a bounded convex constraint set\. Since the features are fixed, the prediction is linear in all trainable parameters, and squared loss remains convex\. For nonlinear controlled dynamics,

xt\+1=f​\(xt,ut\),yt=g​\(xt\),x\_\{t\+1\}=f\(x\_\{t\},u\_\{t\}\),\\qquad y\_\{t\}=g\(x\_\{t\}\),OSF is used as an improper learner: it does not recover the latent state or fit a nonlinear model, but learns a predictor directly from filtered histories of inputs and observations, competing with the best finite\-dimensional linear observer over the horizon\. Under the boundedness and Lipschitz assumptions ofDogariuet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib23)\), choosingk=Θ​\(log2⁡T\)k=\\Theta\(\\log^\{2\}T\)yields an average prediction bound of the form

1T​∑t=1T‖yt−y^tOSF‖2≤O~​\(Q⋆2​log⁡\(Q⋆\)​log7⁡TT\),\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\\\|^\{2\}\\leq\\widetilde\{O\}\\\!\\left\(\\frac\{Q\_\{\\star\}^\{2\}\\log\(Q\_\{\\star\}\)\\log^\{7\}T\}\{\\sqrt\{T\}\}\\right\),\(2\.9\)after converting the bounded regret guarantee to squared loss\. The important point for our purposes is that the rate is controlled by observer complexityQ⋆Q\_\{\\star\}, not by the dimension of the finite\-dimensional linear observer used in the comparison argument\.

### 2\.4Spectral\-to\-LDS distillation

OSF learns a predictor written in terms of length\-LLconvolutions\. Spectral\-to\-LDS distillation converts those convolutional features into recurrent coordinates\.

Forαi∈\[−1,1\]\\alpha\_\{i\}\\in\[\-1,1\]and a signal streamsts\_\{t\}, define

zt\+1\(s,i\)=αi​zt\(s,i\)\+st\.z\_\{t\+1\}^\{\(s,i\)\}=\\alpha\_\{i\}z\_\{t\}^\{\(s,i\)\}\+s\_\{t\}\.\(2\.10\)The implicit convolution of this recurrence is

\(1,αi,αi2,…,αiL−1\)\.\(1,\\alpha\_\{i\},\\alpha\_\{i\}^\{2\},\\ldots,\\alpha\_\{i\}^\{L\-1\}\)\.Runninghhsuch recurrences in parallel gives a dictionary of geometric filters\. SpectraLDS chooses modesα1,…,αh\\alpha\_\{1\},\\ldots,\\alpha\_\{h\}and a mixing matrixF∈ℝk×hF\\in\\mathbb\{R\}^\{k\\times h\}such that

Φ1:k≈F​μL​\(α1:h\),\\Phi\_\{1:k\}\\approx F\\,\\mu\_\{L\}\(\\alpha\_\{1:h\}\),\(2\.11\)whereΦ1:k=\[ϕ1,…,ϕk\]\\Phi\_\{1:k\}=\[\\phi\_\{1\},\\dots,\\phi\_\{k\}\]stacks the spectral filters andμL​\(α1:h\)=\[μL​\(α1\),…,μL​\(αh\)\]\\mu\_\{L\}\(\\alpha\_\{1:h\}\)=\[\\mu\_\{L\}\(\\alpha\_\{1\}\),\\dots,\\mu\_\{L\}\(\\alpha\_\{h\}\)\]stacks the geometric responses\.

Therefore each spectral feature can be replaced by a linear combination of recurrent states:

⟨ϕj,st−1:t−L⟩≈∑i=1hF​\(j,i\)​zt\(s,i\)\.\\left\\langle\\phi\_\{j\},s\_\{t\-1:t\-L\}\\right\\rangle\\approx\\sum\_\{i=1\}^\{h\}F\(j,i\)\\,z\_\{t\}^\{\(s,i\)\}\.\(2\.12\)Applying this replacement to bothst=uts\_\{t\}=u\_\{t\}andst=yts\_\{t\}=y\_\{t\}in \([2\.7](https://arxiv.org/html/2608.05416#S2.E7)\) turns the learned OSF predictor into a recurrent LDS\. The finite\-memory terms∑r=1mJr​ut−r\\sum\_\{r=1\}^\{m\}J\_\{r\}u\_\{t\-r\}and∑r=1mPr​yt−r\\sum\_\{r=1\}^\{m\}P\_\{r\}y\_\{t\-r\}are implemented by adding recurrent connections\.

The two stages are complementary: OSF learns the predictor in a convex spectral feature space, and distillation realizes its temporal filters as recurrent coordinates\. The resulting LDS’s prediction error decomposes into the OSF learning error and the cost of replacing spectral\-filter convolutions by recurrent geometric responses\.

## 3Nonlinear\-to\-LDS Distillation

[Algorithm1](https://arxiv.org/html/2608.05416#alg1)states the full pipeline for the controlled input\-output setting\. We fix horizonTTand filter lengthL=TL=T\.

Algorithm 1Nonlinear\-to\-LDS Distillation0:Inputs

u1:Tu\_\{1:T\}, observations

y1:Ty\_\{1:T\}, filters

kk, AR order

nn, LDS modes

hh, filter length

L=TL=T
1:Spectral learning\.Run OSFDogariuet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib23)\)with

kkspectral filters

ϕ1,…,ϕk\\phi\_\{1\},\\dots,\\phi\_\{k\}and AR order

nnon the controlled trajectory

\(u1:T,y1:T\)\(u\_\{1:T\},y\_\{1:T\}\)to learn the readout parameters

\(Jr,Mj,Pr,Nj\)\(J\_\{r\},M\_\{j\},P\_\{r\},N\_\{j\}\)and obtain

y^tOSF\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}as in \([2\.7](https://arxiv.org/html/2608.05416#S2.E7)\)\.

2:Distillation\.Apply SpectraLDSShahet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib24)\)to

Φ1:k=\[ϕ1​⋯​ϕk\]⊤\\Phi\_\{1:k\}=\[\\phi\_\{1\}\\cdots\\phi\_\{k\}\]^\{\\top\}to obtain modes

α1,…,αh∈\[−1,1\]\\alpha\_\{1\},\\dots,\\alpha\_\{h\}\\in\[\-1,1\]and

F∈ℝk×hF\\in\\mathbb\{R\}^\{k\\times h\}such that

Φ1:k≈F​μL​\(α1:h\)\.\\Phi\_\{1:k\}\\approx F\\,\\mu\_\{L\}\(\\alpha\_\{1:h\}\)\.
3:Recurrent realization\.Initialize

z1\(u,i\)=0∈ℝduz\_\{1\}^\{\(u,i\)\}=0\\in\\mathbb\{R\}^\{d\_\{u\}\}and

z1\(y,i\)=0∈ℝdyz\_\{1\}^\{\(y,i\)\}=0\\in\\mathbb\{R\}^\{d\_\{y\}\}for

i=1,…,hi=1,\\dots,h\. At prediction time

tt, the states

zt\(u,i\)z\_\{t\}^\{\(u,i\)\}and

zt\(y,i\)z\_\{t\}^\{\(y,i\)\}contain only the histories up to time

t−1t\-1\. Form

U~t,j\(u\)=∑i=1hF​\(j,i\)​zt\(u,i\),U~t,j\(y\)=∑i=1hF​\(j,i\)​zt\(y,i\)\.\\widetilde\{U\}\_\{t,j\}^\{\(u\)\}=\\sum\_\{i=1\}^\{h\}F\(j,i\)z\_\{t\}^\{\(u,i\)\},\\qquad\\widetilde\{U\}\_\{t,j\}^\{\(y\)\}=\\sum\_\{i=1\}^\{h\}F\(j,i\)z\_\{t\}^\{\(y,i\)\}\.
4:Output\.Predict

y^tLDS=∑r=1nJr​ut−r\+∑j=1kσj1/4​Mj​U~t,j\(u\)\+∑r=1nPr​yt−r\+∑j=1kσj1/4​Nj​U~t,j\(y\)\.\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}=\\sum\_\{r=1\}^\{n\}J\_\{r\}u\_\{t\-r\}\+\\sum\_\{j=1\}^\{k\}\\sigma\_\{j\}^\{1/4\}M\_\{j\}\\widetilde\{U\}\_\{t,j\}^\{\(u\)\}\+\\sum\_\{r=1\}^\{n\}P\_\{r\}y\_\{t\-r\}\+\\sum\_\{j=1\}^\{k\}\\sigma\_\{j\}^\{1/4\}N\_\{j\}\\widetilde\{U\}\_\{t,j\}^\{\(y\)\}\.
5:State update\.After time

ttis processed, update

zt\+1\(u,i\)=αi​zt\(u,i\)\+ut,zt\+1\(y,i\)=αi​zt\(y,i\)\+yt\.z\_\{t\+1\}^\{\(u,i\)\}=\\alpha\_\{i\}z\_\{t\}^\{\(u,i\)\}\+u\_\{t\},\\qquad z\_\{t\+1\}^\{\(y,i\)\}=\\alpha\_\{i\}z\_\{t\}^\{\(y,i\)\}\+y\_\{t\}\.The finite\-memory AR terms are implemented by shift registers over past inputs and observations\. The resulting predictor is an LDS with

O​\(\(h\+m\)​\(du\+dy\)\)O\(\(h\+m\)\(d\_\{u\}\+d\_\{y\}\)\)scalar state variables\.

###### Theorem 3\.1\(Nonlinear\-to\-LDS distillation\)

Let\(f,g\)\(f,g\)be a nonlinear dynamical system s\.t\.:

1. 1\.For alltt, the inputs, states, and outputs are bounded:‖ut‖,‖xt‖,‖yt‖≤R\\\|u\_\{t\}\\\|,\\\|x\_\{t\}\\\|,\\\|y\_\{t\}\\\|\\leq R,
2. 2\.The dynamicsffand observationsggare11\-Lipschitz,
3. 3\.The system has an observable discretized linear lifting \(see Appendix[B](https://arxiv.org/html/2608.05416#A2)\),

which are the assumptions ofDogariuet al\.\([2025](https://arxiv.org/html/2608.05416#bib.bib23)\)\. Run[Algorithm1](https://arxiv.org/html/2608.05416#alg1)with filter lengthL=TL=T,k=Θ​\(log2⁡T\)k=\\Theta\(\\log^\{2\}T\)filters, andhhSpectraLDS modes withk≤h≤4​kk\\leq h\\leq 4k\. Then, with probability at least1−δ1\-\\delta,

1T​∑t=1T‖yt−y^tLDS‖2≤O~​\(Q⋆2​log⁡\(Q⋆\)​log7⁡TT\+λmax​\(F\)2​exp⁡\(−2​klog⁡T\)\)\.\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}\\\|^\{2\}\\leq\\widetilde\{O\}\\\!\\left\(\\frac\{Q\_\{\\star\}^\{2\}\\log\(Q\_\{\\star\}\)\\log^\{7\}T\}\{\\sqrt\{T\}\}\+\\lambda\_\{\\max\}\(F\)^\{2\}\\exp\\\!\\left\(\-\\frac\{2k\}\{\\log T\}\\right\)\\right\)\.whereO~​\(⋅\)\\widetilde\{O\}\(\\cdot\)suppresses logarithmic factors inTTas well as fixed problem\-dependent constants includingRR,DD,mm, and dimension factors\. The distilled predictor usesO​\(\(di​n\+do​u​t\)​\(h\+m\)\)O\(\(d\_\{in\}\+d\_\{out\}\)\(h\+m\)\)memory and per\-step computation\.

The bound has two terms\. The OSF learning termO~​\(Q⋆2​log⁡Q⋆​log7⁡T/T\)\\widetilde\{O\}\(Q\_\{\\star\}^\{2\}\\log Q\_\{\\star\}\\log^\{7\}T/\\sqrt\{T\}\)depends on the Luenberger complexity of the best linear observer\. The distillation termλmax​\(F\)2​exp⁡\(−2​klog⁡T\)\\lambda\_\{\\max\}\(F\)^\{2\}\\exp\\\!\\left\(\-\\frac\{2k\}\{\\log T\}\\right\)is the cost of replacing the learned convolutional predictor by a recurrent one\. SpectraLDS observes empirically thatλmax​\(F\)\\lambda\_\{\\textrm\{max\}\}\(F\)is bounded and decreasing inhh\(Shahet al\.,[2025](https://arxiv.org/html/2608.05416#bib.bib24), Appendix A\.2\)\. The proof of[Theorem3\.1](https://arxiv.org/html/2608.05416#S3.Thmtheorem1)is given in[AppendixA](https://arxiv.org/html/2608.05416#A1)\.

## 4Experiments: linear systems and behavior cloning

The central theoretical contribution of this work is the first provable method for extracting a best\-in\-hindsight LDS representation of a nonlinear system via convex optimization\. We showcase this contribution in two regimes\. First, we study synthetic LDS recovery, where the target class is itself a linear dynamical system\. Second, we study MuJoCoTodorovet al\.\([2012](https://arxiv.org/html/2608.05416#bib.bib42)\)behavior cloning, where the data are generated by nonlinear closed\-loop expert policies and performance is measured by rollout reward\. In these experiments, we consider the STU and OSF without the recurrent connectionsJrJ\_\{r\}andPrP\_\{r\}\(i\.e\.,n=0n=0\)\.

### 4\.1Learning linear dynamical systems

We first consider prediction from a ground\-truth LDS\. Each task samples a system

xt\+1=A​xt\+B​ut,yt=C​xtx\_\{t\+1\}=Ax\_\{t\}\+Bu\_\{t\},\\qquad y\_\{t\}=Cx\_\{t\}and trains a model to predictyty\_\{t\}from a length\-mmhistory\. We evaluate on two LDS families:*symmetric*LDSs, whose transition matrices have real eigenvalues, and*asymmetric*LDSs, whose non\-normal transitions typically have complex eigenvalues and oscillatory responses\. We compare the STU→\\toSpectraLDS pipeline against directly training a Diagonal LDS, training General LDS, and against training the more\-general Linear FIR baselines\. For the STU, we usek=48k=48spectral filters\. Each architecture has a\+y\+yvariant, which is observational and receives past outputs, and a−y\-yvariant, which receives only inputs\. Thus, the STU\+y\+yvariant is equivalent to OSF\. For STU, all reported losses are computed after SpectraLDS conversion, so the curves evaluate the Diagonal LDS rather than the intermediate convolutional predictor\.

We sweep learning rates over\{10−1,10−2,10−3,10−4,10−5\}\\\{10^\{\-1\},10^\{\-2\},10^\{\-3\},10^\{\-4\},10^\{\-5\}\\\}, select the value minimizing the geometric mean of final test MSE over the symmetric and asymmetric tasks, and rerun the selected configuration over five seeds\. Full data\-generation, initialization, learning\-rate sweep results, and optimization details are given in Appendix[C](https://arxiv.org/html/2608.05416#A3)\.

#### Long\-memory systems\.

For the most challenging setting, we choosedin=dout=16d\_\{\\mathrm\{in\}\}=d\_\{\\mathrm\{out\}\}=16,dstate=64d\_\{\\mathrm\{state\}\}=64,m=512m=512, and spectral radiusρ​\(A\)=0\.999\\rho\(A\)=0\.999\. In this regime the impulse response decays slowly, so accurate prediction requires retaining information far into the past\. We compare parameter\-matched STU, Diagonal LDS, and General LDS models at roughly24,60024\{,\}600parameters, together with a Linear FIR baseline, which requires more than ten times as many parameters\. We also include STU trained with Online Newton Step \(ONS\), which we can employ as the prediction is linear in the STU trainable parameters; in contrast, the directly trained LDS baselines remain non\-convex in their recurrent parameterizations\. We test the\+y\+yvariant of each architecture\.

[Table1](https://arxiv.org/html/2608.05416#S4.T1)and[Figure1](https://arxiv.org/html/2608.05416#S4.F1)show that STU→\\toSpectraLDS outperforms each baseline under Adam, while STU\+\+ONS improves the error by another three orders of magnitude\. The General LDS is more expressive than the Diagonal LDS, but it still does not close the gap to the spectral pipeline, suggesting that the main advantage is optimization rather than final model class\. The wall\-clock times in[Table1](https://arxiv.org/html/2608.05416#S4.T1)also show that the spectral model is substantially faster to train than the full General LDS, whose recurrent rollout dominates computation\. Linear FIR requiresm⋅\(di​n\+do​u​t\)⋅do​u​tm\\cdot\(d\_\{in\}\+d\_\{out\}\)\\cdot d\_\{out\}parameters, leading to overparameterization that challenges optimization\. In Appendix[E](https://arxiv.org/html/2608.05416#A5), we show that, in this setting, the STU and its distilled SpectraLDS representation have at most1\.5%1\.5\\%relative error\.

Table 1:Linear LDS prediction in the long\-memory regime \(din=dout=16d\_\{\\mathrm\{in\}\}\{=\}d\_\{\\mathrm\{out\}\}\{=\}16,dstate=64d\_\{\\mathrm\{state\}\}\{=\}64,m=512m\{=\}512,ρ​\(A\)=0\.999\\rho\(A\)\{=\}0\.999\)\. Final test MSE over five seeds, parameter count, and wall\-clock time for one 100\-epoch run on an NVIDIA H100\. STU rows report the SpectraLDS\-distilled recurrent predictor\.![Refer to caption](https://arxiv.org/html/2608.05416v1/images/radius999_d16.png)Figure 1:Test MSE versus epoch on the long\-memory LDS task\. Curves are geometric means over five seeds with geometric\-standard\-deviation bands, and theyy\-axis is log\-scaled\. STU→\\toSpectraLDS is competitive under Adam and dominates when optimized with ONS, while Linear FIR fails despite having more than ten times as many parameters\.
#### Shorter\-memory and larger\-scale systems\.

We also test spectral radiusρ​\(A\)=0\.95\\rho\(A\)=0\.95, where transients decay much faster than in the long\-memory experiment\. We evaluate two scales:

\(din,dout,dstate,m\)=\(4,4,8,256\)and\(64,64,64,512\)\.\(d\_\{\\mathrm\{in\}\},d\_\{\\mathrm\{out\}\},d\_\{\\mathrm\{state\}\},m\)=\(4,4,8,256\)\\quad\\text\{and\}\\quad\(64,64,64,512\)\.At each scale we evaluate the\+y\+yand−y\-yvariants of STU→\\toSpectraLDS, parameter\-matched Diagonal LDS, Diagonal LDS withh=2​dstateh=2d\_\{\\mathrm\{state\}\}, General LDS withh=2​dstateh=2d\_\{\\mathrm\{state\}\}, and Linear FIR\.

For the large\-scale experiment, we report the General LDS ath=2​dstateh=2d\_\{\\mathrm\{state\}\}rather than using a fully parameter\-matched General LDS\. As shown in Table[1](https://arxiv.org/html/2608.05416#S4.T1), General LDS is substantially more expensive than the other methods, making rollouts of a parameter\-matched General LDS model prohibitively expensive\. We therefore useh=2​dstateh=2d\_\{\\mathrm\{state\}\}as the feasible General LDS baseline\. For the Diagonal LDS, Figure[2](https://arxiv.org/html/2608.05416#S4.F2)shows that the parameter\-matched andh=2​dstateh=2d\_\{\\mathrm\{state\}\}variants have similar final test MSE\. Full parameter counts are reported in Appendix[C\.4](https://arxiv.org/html/2608.05416#A3.SS4)\.

[Figure2](https://arxiv.org/html/2608.05416#S4.F2)shows that the qualitative pattern persists\. At both scales, STU→\\toSpectraLDS remains a strong predictor on both symmetric and asymmetric dynamics, although surpassed by Linear FIR on the smaller setting and General LDS at the larger setting\. The diagonal LDS models, SpectraLDS and Diagonal LDS, without the\+y\+yobservation struggle on the asymmetric systems, where real diagonal recurrences cannot represent the complex modes of generic non\-symmetric transitions\. Linear FIR is competitive only in the smallest symmetric setting without the\+y\+yobservation, where its parameter count is not yet a bottleneck\.

![Refer to caption](https://arxiv.org/html/2608.05416v1/images/d8_d64_2x2.png)Figure 2:Test MSE versus epoch on linear LDS prediction at spectral radius0\.950\.95\. The top row shows the small scale \(din=dout=4d\_\{\\mathrm\{in\}\}=d\_\{\\mathrm\{out\}\}=4,dstate=8d\_\{\\mathrm\{state\}\}=8,m=256m=256\), and the bottom row shows the large scale \(din=dout=64d\_\{\\mathrm\{in\}\}=d\_\{\\mathrm\{out\}\}=64,dstate=64d\_\{\\mathrm\{state\}\}=64,m=512m=512\)\. Solid curves are\+y\+yvariants and dashed curves are−y\-yvariants\. Curves are geometric means over five seeds with geometric\-standard\-deviation bands, and theyy\-axis is log\-scaled\.

### 4\.2MuJoCo behavior cloning

We next test the same comparison on nonlinear control data\. For each MuJoCo\-v5taskTodorovet al\.\([2012](https://arxiv.org/html/2608.05416#bib.bib42)\), we roll out a deterministic Soft Actor\-Critic \(SAC\)Haarnojaet al\.\([2018](https://arxiv.org/html/2608.05416#bib.bib43)\)expert trained by the Minari teamYouniset al\.\([2024](https://arxiv.org/html/2608.05416#bib.bib44)\)forT=2,000,000T=2\{,\}000\{,\}000steps and train a student to predict the expert action from a length\-mmwindow of past observations and previous actions\. The trained student is then deployed in the environment and evaluated by episode reward\. Note that training uses teacher\-forced next\-action prediction on expert trajectories, while evaluation is closed\-loop reward in the environment\. MuJoCo results therefore serve as empirical evidence that the train\-then\-distill recipe survives this distribution shift, not as a direct test of our prediction\-error bound, which concerns open\-loop prediction\.

All students share the same encoder, decoder, optimizer, batch size, and training horizon\. The only component that changes is the sequence\-mixing block:

1. 1\.STU→\\toSpectraLDS: trained as a tensor\-dot STU approximation withk=16k=16spectral filters and deployed after conversion to a Diagonal LDS\.
2. 2\.Diagonal LDS, hidden\-dimension matched: a directly trained Diagonal LDS with hidden dimensionh=dh=d\.
3. 3\.Diagonal LDS, parameter\-matched: a directly trained Diagonal LDS whose hidden dimension is chosen to match the STU sequence\-block parameter count\.

Every reward reported for STU is therefore the reward of the distilled recurrent LDS, not the pre\-distillation convolutional model\. All models are trained for100100epochs with Adam at learning rate3×10−43\\\!\\times\\\!10^\{\-4\}and batch size10241024, with five seeds per environment\-model pair\.

The students are intentionally compact: all three student architectures use fewer parameters than the SAC actor deployed by the teacher on every environment\. Within the student models, STU and the parameter\-matched Diagonal LDS have essentially identical parameter counts by construction, while the hidden\-dimension\-matched Diagonal LDS is roughly twenty to twenty\-five percent larger\. See Appendices[D](https://arxiv.org/html/2608.05416#A4)and[D\.5](https://arxiv.org/html/2608.05416#A4.SS5)for full details\.

[Table2](https://arxiv.org/html/2608.05416#S4.T2)and[Figure3](https://arxiv.org/html/2608.05416#S4.F3)show that STU→\\toSpectraLDS achieves the highest mean reward on four of the six environments: Walker2d, Hopper, HalfCheetah, and Humanoid\. On Ant and Pusher, it remains close to the best directly trained Diagonal LDS baseline\. The largest separation occurs on Walker2d, where the spectral pipeline exceeds the parameter\-matched Diagonal LDS by roughly1,7001\{,\}700reward\. These results support the same conclusion as the LDS recovery experiments: training in the spectral parameterization and then distilling to an LDS is often more reliable than directly training the recurrent LDS, even when the final deployed model class is the same\.

![Refer to caption](https://arxiv.org/html/2608.05416v1/images/diag_lds_vs_stu_with_y_6envs_bigfont.png)Figure 3:Evaluation reward during behavior cloning on six MuJoCo tasks\. Curves show mean and standard deviation over five seeds\. STU denotes the deployed SpectraLDS\-distilled recurrent model\.Table 2:Best\-checkpoint MuJoCo behavior\-cloning reward, re\-evaluated over1,0001\{,\}000episodes \(mean±\\pmstd, five seeds\)\. STU denotes the deployed SpectraLDS\-distilled recurrent model\.

## 5Conclusion

We showed that a convex spectral predictor can be distilled into a recurrent LDS whose prediction error decomposes into a distillation term that decays exponentially in the LDS state dimension and an OSF learning term controlled by the observer condition numberQ⋆Q\_\{\\star\}, independent of the latent nonlinear dimension\. Empirically, the distilled LDS matches or outperforms directly trained LDS baselines at equal parameter budgets on both linear recovery and MuJoCo behavior cloning, with the advantage appearing to come from the spectral parameterization’s convexity rather than from a richer model class\.

Several directions remain open\. On the theory side, we expect that a refined choice of sampling distribution over the SpectraLDS modesαi\\alpha\_\{i\}would yield sharper guarantees onλmax​\(F\)\\lambda\_\{\\max\}\(F\)in the overparameterized regimeh≳kh\\gtrsim k, where the current bound is dominated by the pseudoinverse norm\. Additionally, the bound in Theorem[3\.1](https://arxiv.org/html/2608.05416#S3.Thmtheorem1)controls average one\-step prediction error under teacher forcing\. It does not cover closed\-loop rollout error, where errors compound over the horizon — the regime relevant to control and long\-horizon simulation\. Closing this gap is the most natural extension\. Empirically, larger\-scale experiments and integration with modern sequence models may establish nonlinear\-to\-LDS distillation as a building block for long\-horizon forecasting and control\.

## References

- \[1\]\(2025\)Universal learning of nonlinear dynamics\.arXiv preprint arXiv:2508\.11990\.External Links:2508\.11990Cited by:[Appendix A](https://arxiv.org/html/2608.05416#A1.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2608.05416#A2.p1.1),[Appendix B](https://arxiv.org/html/2608.05416#A2.p3.4),[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px2.p1.1),[§2\.3](https://arxiv.org/html/2608.05416#S2.SS3.p1.2),[§2\.3](https://arxiv.org/html/2608.05416#S2.SS3.p3.6),[§2](https://arxiv.org/html/2608.05416#S2.p1.1),[Theorem 3\.1](https://arxiv.org/html/2608.05416#S3.Thmtheorem1.p1.6.5),[1](https://arxiv.org/html/2608.05416#alg1.l1)\.
- \[2\]A\. Gu and T\. Dao\(2023\)Mamba: linear\-time sequence modeling with selective state spaces\.arXiv preprint arXiv:2312\.00752\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px4.p1.1)\.
- \[3\]A\. Gu, K\. Goel, and C\. Ré\(2022\)Efficiently modeling long sequences with structured state spaces\.InInternational Conference on Learning Representations,Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px4.p1.1)\.
- \[4\]T\. Haarnoja, A\. Zhou, P\. Abbeel, and S\. Levine\(2018\)Soft actor\-critic: off\-policy maximum entropy deep reinforcement learning with a stochastic actor\.External Links:1801\.01290,[Link](https://arxiv.org/abs/1801.01290)Cited by:[§D\.2](https://arxiv.org/html/2608.05416#A4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.05416#S4.SS2.p1.2)\.
- \[5\]M\. Hardt, T\. Ma, and B\. Recht\(2018\)Gradient descent learns linear dynamical systems\.InJournal of Machine Learning Research,Vol\.19,pp\. 1–44\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px1.p1.1)\.
- \[6\]E\. Hazan, A\. Agarwal, and S\. Kale\(2007/12/01\)Logarithmic regret algorithms for online convex optimization\.Machine Learning69\(2\),pp\. 169–192\.External Links:[Document](https://dx.doi.org/10.1007/s10994-007-5016-8),ISBN 1573\-0565,[Link](https://doi.org/10.1007/s10994-007-5016-8)Cited by:[§C\.3](https://arxiv.org/html/2608.05416#A3.SS3.SSS0.Px2.p1.6)\.
- \[7\]E\. Hazan, H\. Lee, K\. Singh, C\. Zhang, and Y\. Zhang\(2018\)Spectral filtering for general linear dynamical systems\.InAdvances in Neural Information Processing Systems,Vol\.31,pp\. 4634–4643\.External Links:1802\.03981Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2608.05416#S2.SS2.p2.14)\.
- \[8\]E\. Hazan, K\. Singh, and C\. Zhang\(2017\)Learning linear dynamical systems via spectral filtering\.InAdvances in Neural Information Processing Systems,Vol\.30\.External Links:[Link](https://proceedings.neurips.cc/paper/2017/hash/165a59f7cf3b5c4396ba65953d679f17-Abstract.html),1711\.00946Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px2.p1.1),[§2\.1](https://arxiv.org/html/2608.05416#S2.SS1.p2.4),[§2\.2](https://arxiv.org/html/2608.05416#S2.SS2.p1.8),[§2](https://arxiv.org/html/2608.05416#S2.p1.1)\.
- \[9\]V\. Kostic, P\. Novelli, A\. Maurer, C\. Ciliberto, L\. Rosasco, and M\. Pontil\(2022\)Learning dynamical systems via koopman operator regression in reproducing kernel hilbert spaces\.InAdvances in Neural Information Processing Systems,Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px3.p1.1)\.
- \[10\]Y\. I\. Liu, W\. Nguyen, Y\. Devre, E\. Dogariu, A\. Majumdar, and E\. Hazan\(2024\)Flash STU: fast spectral transform units\.arXiv preprint arXiv:2409\.10489\.External Links:2409\.10489Cited by:[§D\.4](https://arxiv.org/html/2608.05416#A4.SS4.SSS0.Px1.p1.6)\.
- \[11\]L\. Ljung\(1999\)System identification: theory for the user\.2nd edition,Prentice Hall\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px1.p1.1)\.
- \[12\]B\. Lusch, J\. N\. Kutz, and S\. L\. Brunton\(2018\)Deep learning for universal linear embeddings of nonlinear dynamics\.Nature Communications9\(1\),pp\. 4950\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px3.p1.1)\.
- \[13\]I\. Mezić\(2005\)Spectral properties of dynamical systems, model reduction and decompositions\.Nonlinear Dynamics41\(1\-3\),pp\. 309–325\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px3.p1.1)\.
- \[14\]A\. Orvieto, S\. L\. Smith, A\. Gu, A\. Fernando, C\. Gulcehre, R\. Pascanu, and S\. De\(2023\)Resurrecting recurrent neural networks for long sequences\.InInternational Conference on Machine Learning,Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px4.p1.1)\.
- \[15\]S\. Oymak and N\. Ozay\(2019\)Non\-asymptotic identification of lti systems from a single trajectory\.American Control Conference,pp\. 5655–5661\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px1.p1.1)\.
- \[16\]T\. Sarkar and A\. Rakhlin\(2019\)Near optimal finite time identification of arbitrary linear dynamical systems\.InInternational Conference on Machine Learning,pp\. 5610–5618\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px1.p1.1)\.
- \[17\]D\. Shah, S\. Fortgang, S\. Druchyna, and E\. Hazan\(2025\)SpectraLDS: provable distillation for linear dynamical systems\.arXiv preprint arXiv:2505\.17868\.External Links:2505\.17868Cited by:[Appendix A](https://arxiv.org/html/2608.05416#A1.SS0.SSS0.Px2.p1.10),[Appendix A](https://arxiv.org/html/2608.05416#A1.SS0.SSS0.Px2.p1.12),[Appendix A](https://arxiv.org/html/2608.05416#A1.SS0.SSS0.Px2.p1.4),[§D\.4](https://arxiv.org/html/2608.05416#A4.SS4.SSS0.Px1.p1.6),[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.05416#S2.p1.1),[§3](https://arxiv.org/html/2608.05416#S3.p2.4),[2](https://arxiv.org/html/2608.05416#alg1.l2)\.
- \[18\]M\. Simchowitz, H\. Mania, S\. Tu, M\. I\. Jordan, and B\. Recht\(2018\)Learning without mixing: towards a sharp analysis of linear system identification\.InConference on Learning Theory,pp\. 439–473\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px1.p1.1)\.
- \[19\]J\. T\. H\. Smith, A\. Warrington, and S\. Linderman\(2023\)Simplified state space layers for sequence modeling\.InInternational Conference on Learning Representations,Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px4.p1.1)\.
- \[20\]E\. Todorov, T\. Erez, and Y\. Tassa\(2012\)MuJoCo: a physics engine for model\-based control\.In2012 IEEE/RSJ International Conference on Intelligent Robots and Systems,pp\. 5026–5033\.External Links:[Document](https://dx.doi.org/10.1109/IROS.2012.6386109)Cited by:[Table 6](https://arxiv.org/html/2608.05416#A4.T6),[§4\.2](https://arxiv.org/html/2608.05416#S4.SS2.p1.2),[§4](https://arxiv.org/html/2608.05416#S4.p1.3)\.
- \[21\]A\. Tsiamis and G\. J\. Pappas\(2019\)Finite sample analysis of stochastic system identification\.InIEEE Conference on Decision and Control,pp\. 3648–3654\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px1.p1.1)\.
- \[22\]M\. O\. Williams, I\. G\. Kevrekidis, and C\. W\. Rowley\(2015\)A data\-driven approximation of the koopman operator: extending dynamic mode decomposition\.Journal of Nonlinear Science25\(6\),pp\. 1307–1346\.Cited by:[§1\.2](https://arxiv.org/html/2608.05416#S1.SS2.SSS0.Px3.p1.1)\.
- \[23\]MinariExternal Links:[Document](https://dx.doi.org/10.5281/zenodo.13767625),[Link](https://doi.org/10.5281/zenodo.13767625)Cited by:[§D\.2](https://arxiv.org/html/2608.05416#A4.SS2.p1.1),[Table 6](https://arxiv.org/html/2608.05416#A4.T6),[§4\.2](https://arxiv.org/html/2608.05416#S4.SS2.p1.2)\.

## Appendix AProof of Theorem[3\.1](https://arxiv.org/html/2608.05416#S3.Thmtheorem1)

The proof has two ingredients\. First, the OSF guarantee gives a mean\-squared prediction bound for the learned spectral predictor\. Second, replacing the OSF filters by their SpectraLDS approximation perturbs the predictions by an amount controlled by the filter approximation error and the norm of the learned readout\. Combining these two bounds gives the result\.

Throughout,O~​\(⋅\)\\widetilde\{O\}\(\\cdot\)suppresses logarithmic factors inTTand fixed problem\-dependent constants\. We state the proof for scalar outputs for readability; the vector\-output case follows by using operator norms for the readout matrices and contributes only the corresponding output\-dimension factor\.

#### Step 1: OSF mean\-squared error\.

By Theorem 3\.1 of\[[1](https://arxiv.org/html/2608.05416#bib.bib23)\],

AT:=1T​∑t=1T‖yt−y^tOSF‖2≤O~​\(Q⋆2​log⁡\(Q⋆\)​log7⁡TT\)\.A\_\{T\}:=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\\\|^\{2\}\\leq\\widetilde\{O\}\\\!\\left\(\\frac\{Q\_\{\\star\}^\{2\}\\log\(Q\_\{\\star\}\)\\log^\{7\}T\}\{\\sqrt\{T\}\}\\right\)\.\(A\.1\)

#### Step 2: SpectraLDS filter approximation\.

Run Algorithm 2 of\[[17](https://arxiv.org/html/2608.05416#bib.bib24)\]onΦ1:k\\Phi\_\{1:k\}to obtain modesα1,…,αh\\alpha\_\{1\},\\dots,\\alpha\_\{h\}and a transformation matrixF∈ℝk×hF\\in\\mathbb\{R\}^\{k\\times h\}\(denotedMfM\_\{f\}in\[[17](https://arxiv.org/html/2608.05416#bib.bib24)\]\)\. We write

ϕ~j=∑i=1hF​\(j,i\)​μL​\(αi\),j=1,…,k\.\\widetilde\{\\phi\}\_\{j\}=\\sum\_\{i=1\}^\{h\}F\(j,i\)\\,\\mu\_\{L\}\(\\alpha\_\{i\}\),\\qquad j=1,\\dots,k\.Define the rowwise approximation error

ηk,h​\(L\):=maxj≤k⁡‖ϕj−ϕ~j‖1\.\\eta\_\{k,h\}\(L\):=\\max\_\{j\\leq k\}\\\|\\phi\_\{j\}\-\\widetilde\{\\phi\}\_\{j\}\\\|\_\{1\}\.\(A\.2\)By Theorem 1 of\[[17](https://arxiv.org/html/2608.05416#bib.bib24)\], on its high\-probability event,

ηk,h​\(L\)≲h​λmax​\(F\)​exp⁡\(−klog⁡L\),\\eta\_\{k,h\}\(L\)\\;\\lesssim\\;h\\,\\lambda\_\{\\max\}\(F\)\\,\\exp\\\!\\left\(\-\\frac\{k\}\{\\log L\}\\right\),\(A\.3\)whereλmax​\(F\)\\lambda\_\{\\max\}\(F\)is the largest singular value of the Moore–Penrose pseudoinverse of theh×kh\\times kcoefficient matrix produced by Algorithm 2 \(calledMMin\[[17](https://arxiv.org/html/2608.05416#bib.bib24)\], withF=M\+F=M^\{\+\}\)\. As shown in\[[17](https://arxiv.org/html/2608.05416#bib.bib24), Appendix A\.2\], the producth​λmax​\(F\)h\\,\\lambda\_\{\\max\}\(F\)is empirically bounded for the sampling distribution overαi\\alpha\_\{i\}used by Algorithm 2\.

#### Step 3: Filter error to prediction error\.

Lety^tdist\\hat\{y\}\_\{t\}^\{\\mathrm\{dist\}\}be the predictor obtained by replacing eachϕj\\phi\_\{j\}in the OSF predictor withϕ~j\\widetilde\{\\phi\}\_\{j\}\. As the finite\-memory AR terms∑rJr​ut−r\\sum\_\{r\}J\_\{r\}u\_\{t\-r\}and∑rPr​yt−r\\sum\_\{r\}P\_\{r\}y\_\{t\-r\}in \([2\.7](https://arxiv.org/html/2608.05416#S2.E7)\) are unchanged by the filter substitution and so do not contribute to the error, we have

y^tOSF−y^tdist=∑j=1kσj1/4​Mj​⟨ϕj−ϕ~j,ut−1:t−L⟩\+∑j=1kσj1/4​Nj​⟨ϕj−ϕ~j,yt−1:t−L⟩\.\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{dist\}\}=\\sum\_\{j=1\}^\{k\}\\sigma\_\{j\}^\{1/4\}M\_\{j\}\\left\\langle\\phi\_\{j\}\-\\widetilde\{\\phi\}\_\{j\},u\_\{t\-1:t\-L\}\\right\\rangle\+\\sum\_\{j=1\}^\{k\}\\sigma\_\{j\}^\{1/4\}N\_\{j\}\\left\\langle\\phi\_\{j\}\-\\widetilde\{\\phi\}\_\{j\},y\_\{t\-1:t\-L\}\\right\\rangle\.Using‖ut‖,‖yt‖≤R\\\|u\_\{t\}\\\|,\\\|y\_\{t\}\\\|\\leq Rand‖σj1/4​Mj‖op,‖σj1/4​Nj‖op≤D\\\|\\sigma\_\{j\}^\{1/4\}M\_\{j\}\\\|\_\{\\mathrm\{op\}\},\\\|\\sigma\_\{j\}^\{1/4\}N\_\{j\}\\\|\_\{\\mathrm\{op\}\}\\leq D, we have

‖⟨ϕj−ϕ~j,ut−1:t−L⟩‖≤R​‖ϕj−ϕ~j‖1≤R​ηk,h​\(L\),\\left\\\|\\left\\langle\\phi\_\{j\}\-\\widetilde\{\\phi\}\_\{j\},u\_\{t\-1:t\-L\}\\right\\rangle\\right\\\|\\leq R\\\|\\phi\_\{j\}\-\\widetilde\{\\phi\}\_\{j\}\\\|\_\{1\}\\leq R\\eta\_\{k,h\}\(L\),and the same bound holds for the observation\-history term\. Therefore,

‖y^tOSF−y^tdist‖≤2​k​D​R​ηk,h​\(L\)\.\\\|\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{dist\}\}\\\|\\leq 2kDR\\,\\eta\_\{k,h\}\(L\)\.Averaging over time gives the bound

1T​∑t=1T‖y^tOSF−y^tdist‖2≤4​k2​D2​R2​ηk,h​\(L\)2\.\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{dist\}\}\\\|^\{2\}\\leq 4k^\{2\}D^\{2\}R^\{2\}\\,\\eta\_\{k,h\}\(L\)^\{2\}\.\(A\.4\)
Substituting \([A\.3](https://arxiv.org/html/2608.05416#A1.E3)\) withL=TL=Tyields

1T​∑t=1T‖y^tOSF−y^tdist‖2≤O~​\(k2​D2​R2​h2​λmax​\(F\)2​exp⁡\(−2​klog⁡T\)\)\.\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{dist\}\}\\\|^\{2\}\\leq\\widetilde\{O\}\\\!\\left\(k^\{2\}D^\{2\}R^\{2\}h^\{2\}\\lambda\_\{\\max\}\(F\)^\{2\}\\exp\\\!\\left\(\-\\frac\{2k\}\{\\log T\}\\right\)\\right\)\.\(A\.5\)
The predictory^tdist\\hat\{y\}\_\{t\}^\{\\mathrm\{dist\}\}is realized exactly by an LDS: each geometric responseμL​\(αi\)\\mu\_\{L\}\(\\alpha\_\{i\}\)is implemented by a scalar recurrent mode, and the AR terms are implemented by finite shift registers\. We denote this exact recurrent realization byy^tLDS\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}\.

#### Step 4: Combine\.

For everytt,

yt−y^tLDS=\(yt−y^tOSF\)\+\(y^tOSF−y^tdist\)\.y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}=\(y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\)\+\(\\hat\{y\}\_\{t\}^\{\\mathrm\{OSF\}\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{dist\}\}\)\.Using

‖a\+b‖2≤2​‖a‖2\+2​‖b‖2,\\\|a\+b\\\|^\{2\}\\leq 2\\\|a\\\|^\{2\}\+2\\\|b\\\|^\{2\},together with \([A\.1](https://arxiv.org/html/2608.05416#A1.E1)\) and \([A\.4](https://arxiv.org/html/2608.05416#A1.E4)\), gives

1T​∑t=1T‖yt−y^tLDS‖2≤2​AT\+8​k2​D2​R2​ηk,h​\(L\)2\.\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}\\\|^\{2\}\\leq 2A\_\{T\}\+8k^\{2\}D^\{2\}R^\{2\}\\,\\eta\_\{k,h\}\(L\)^\{2\}\.Finally, substituting the SpectraLDS high\-probability bound \([A\.3](https://arxiv.org/html/2608.05416#A1.E3)\) withL=TL=Tgives the explicit estimate

1T​∑t=1T‖yt−y^tLDS‖2≤O~​\(Q⋆2​log⁡\(Q⋆\)​log7⁡TT\+k2​D2​R2​h2​λmax​\(F\)2​exp⁡\(−2​klog⁡T\)\)\.\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}\\\|^\{2\}\\leq\\widetilde\{O\}\\\!\\left\(\\frac\{Q\_\{\\star\}^\{2\}\\log\(Q\_\{\\star\}\)\\log^\{7\}T\}\{\\sqrt\{T\}\}\+k^\{2\}D^\{2\}R^\{2\}h^\{2\}\\lambda\_\{\\max\}\(F\)^\{2\}\\exp\\\!\\left\(\-\\frac\{2k\}\{\\log T\}\\right\)\\right\)\.
Sincek=Θ​\(log2⁡T\)k=\\Theta\(\\log^\{2\}T\), the factork2k^\{2\}is polylogarithmic inTTand is absorbed byO~​\(⋅\)\\widetilde\{O\}\(\\cdot\)\. The constantsDDandRR, together withh=O​\(k\)h=O\(k\), are also absorbed byO~​\(⋅\)\\widetilde\{O\}\(\\cdot\)\. Thus the same bound can be written in the cleaner form

1T​∑t=1T‖yt−y^tLDS‖2≤O~​\(Q⋆2​log⁡\(Q⋆\)​log7⁡TT\+λmax​\(F\)2​exp⁡\(−2​klog⁡T\)\)\.\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\\|y\_\{t\}\-\\hat\{y\}\_\{t\}^\{\\mathrm\{LDS\}\}\\\|^\{2\}\\leq\\widetilde\{O\}\\\!\\left\(\\frac\{Q\_\{\\star\}^\{2\}\\log\(Q\_\{\\star\}\)\\log^\{7\}T\}\{\\sqrt\{T\}\}\+\\lambda\_\{\\max\}\(F\)^\{2\}\\exp\\\!\\left\(\-\\frac\{2k\}\{\\log T\}\\right\)\\right\)\.This proves the theorem\.□\\square

## Appendix BAssumption: Observable Discretized Linear Lifting

We recall the discretized lifting assumption used in\[[1](https://arxiv.org/html/2608.05416#bib.bib23)\]\. For simplicity, we state the autonomous version and consider:

xt\+1=f​\(xt\),yt=g​\(xt\),x\_\{t\+1\}=f\(x\_\{t\}\),\\qquad y\_\{t\}=g\(x\_\{t\}\),
Fix a discretization scaleε\>0\\varepsilon\>0, and letS=\{s\(1\),…,s\(N\)\}S=\\\{s^\{\(1\)\},\\dots,s^\{\(N\)\}\\\}be anε\\varepsilon\-net of the radius\-RRball in state space\. Thus every bounded statexxhas a representativeπ​\(x\)∈S\\pi\(x\)\\in Swith

‖x−π​\(x\)‖≤ε,N=\|S\|≤\(2​Rε\)dX\.\\\|x\-\\pi\(x\)\\\|\\leq\\varepsilon,\\qquad N=\|S\|\\leq\\left\(\\frac\{2R\}\{\\varepsilon\}\\right\)^\{d\_\{X\}\}\.Define a finite\-state approximation to the nonlinear dynamics by

T​\(s\)=π​\(f​\(s\)\)\.T\(s\)=\\pi\(f\(s\)\)\.Equivalently, define a transition matrixA′∈ℝN×NA^\{\\prime\}\\in\\mathbb\{R\}^\{N\\times N\}by

Ai​j′=𝟏​\{T​\(s\(j\)\)=s\(i\)\}\.A^\{\\prime\}\_\{ij\}=\\mathbf\{1\}\\\{T\(s^\{\(j\)\}\)=s^\{\(i\)\}\\\}\.The lifted statezt∈ℝNz\_\{t\}\\in\\mathbb\{R\}^\{N\}is the one\-hot vector representing the current grid point\. Its evolution is linear:

zt\+1=A′​zt\.z\_\{t\+1\}=A^\{\\prime\}z\_\{t\}\.Define the observation matrixC′∈ℝdY×NC^\{\\prime\}\\in\\mathbb\{R\}^\{d\_\{Y\}\\times N\}by setting itsjjth column to the observation at the corresponding grid point:

C′​ej=g​\(s\(j\)\)\.C^\{\\prime\}e\_\{j\}=g\(s^\{\(j\)\}\)\.The discretized linear lifting is therefore the finite\-dimensional LDS

zt\+1=A′​zt,yt′=C′​zt\.z\_\{t\+1\}=A^\{\\prime\}z\_\{t\},\\qquad y^\{\\prime\}\_\{t\}=C^\{\\prime\}z\_\{t\}\.
By Lemma 5\.5 of\[[1](https://arxiv.org/html/2608.05416#bib.bib23)\], ifffandggare11\-Lipschitz and the trajectory is bounded, then this lifted LDS approximates the nonlinear output sequence over horizonTT:

∑t≤T‖yt′−yt‖≤T22​ε\.\\sum\_\{t\\leq T\}\\\|y^\{\\prime\}\_\{t\}\-y\_\{t\}\\\|\\leq\\frac\{T^\{2\}\}\{2\}\\varepsilon\.Thus, choosingε\\varepsilonsufficiently small produces a high\-dimensional LDS whose outputs track the nonlinear trajectory\. The dimensionN≤\(2​R/ε\)dXN\\leq\(2R/\\varepsilon\)^\{d\_\{X\}\}can be very large, but OSF is an improper learner: its runtime and parameter count do not depend on explicitly constructing or learning this lifted state\.

The additional assumption in Theorem[3\.1](https://arxiv.org/html/2608.05416#S3.Thmtheorem1)is that the pair\(A′,C′\)\(A^\{\\prime\},C^\{\\prime\}\)is observable\. That is, the observability matrix

𝒪​\(A′,C′\)=\[C′C′​A′C′​\(A′\)2⋮C′​\(A′\)N−1\]∈ℝN​dY×N\\mathcal\{O\}\(A^\{\\prime\},C^\{\\prime\}\)=\\begin\{bmatrix\}C^\{\\prime\}\\\\ C^\{\\prime\}A^\{\\prime\}\\\\ C^\{\\prime\}\(A^\{\\prime\}\)^\{2\}\\\\ \\vdots\\\\ C^\{\\prime\}\(A^\{\\prime\}\)^\{N\-1\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{Nd\_\{Y\}\\times N\}has full column rank:

rank⁡𝒪​\(A′,C′\)=N\.\\operatorname\{rank\}\\mathcal\{O\}\(A^\{\\prime\},C^\{\\prime\}\)=N\.Equivalently, no nonzero lifted state direction is invisible to all future linear observations:

C′​\(A′\)ℓ​v=0for all​ℓ=0,…,N−1⟹v=0\.C^\{\\prime\}\(A^\{\\prime\}\)^\{\\ell\}v=0\\quad\\text\{for all \}\\ell=0,\\dots,N\-1\\quad\\Longrightarrow\\quad v=0\.
This observability condition is used only in the comparison argument\. It ensures that a Luenberger observer can be formed for the lifted LDS and that the observer complexityQ⋆Q\_\{\\star\}appearing in the OSF bound is finite\.

## Appendix CLinear LDS data\-generation and training protocol

This appendix documents the data\-generation procedure, predictor initializations, and training protocol used in the linear\-LDS experiments of Section[4\.1](https://arxiv.org/html/2608.05416#S4.SS1)\.

### C\.1Data generation

Each ground\-truth LDS\(A,B,C\)\(A,B,C\)is sampled once per \(seed, task\) pair\.

- •SymmetricAA: drawQ∈ℝdstate×dstateQ\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{state\}\}\\times d\_\{\\mathrm\{state\}\}\}from the QR decomposition of a Gaussian matrix and eigenvaluesλi∼i\.i\.d\.Uniform​\(−ρ,ρ\)\\lambda\_\{i\}\\stackrel\{\{\\scriptstyle\\text\{i\.i\.d\.\}\}\}\{\{\\sim\}\}\\mathrm\{Uniform\}\(\-\\rho,\\rho\); setA=Q​diag​\(λ\)​Q⊤A=Q\\,\\mathrm\{diag\}\(\\lambda\)\\,Q^\{\\top\}\.
- •AsymmetricAA: drawA~i​j∼i\.i\.d\.𝒩​\(0,1/dstate\)\\tilde\{A\}\_\{ij\}\\stackrel\{\{\\scriptstyle\\text\{i\.i\.d\.\}\}\}\{\{\\sim\}\}\\mathcal\{N\}\(0,1/d\_\{\\mathrm\{state\}\}\)and rescale to the target spectral radius viaA=ρ​A~/maxi⁡\|λi​\(A~\)\|A=\\rho\\,\\tilde\{A\}/\\max\_\{i\}\|\\lambda\_\{i\}\(\\tilde\{A\}\)\|\. The eigenvalues are typically complex, producing oscillatory dynamics\.
- •B∈ℝdstate×dinB\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{state\}\}\\times d\_\{\\mathrm\{in\}\}\}andC∈ℝdout×dstateC\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{state\}\}\}have entries𝒩​\(0,1/din\)\\mathcal\{N\}\(0,1/d\_\{\\mathrm\{in\}\}\)and𝒩​\(0,1/dstate\)\\mathcal\{N\}\(0,1/d\_\{\\mathrm\{state\}\}\), respectively\.

Inputs are sampled i\.i\.d\. asut∼𝒩​\(0,Idin\)u\_\{t\}\\sim\\mathcal\{N\}\(0,I\_\{d\_\{\\mathrm\{in\}\}\}\), and the state evolves fromx0=0x\_\{0\}=0viaxt\+1=A​xt\+B​utx\_\{t\+1\}=Ax\_\{t\}\+Bu\_\{t\}with outputyt=C​xty\_\{t\}=Cx\_\{t\}\. We generateTtrain=10,000T\_\{\\mathrm\{train\}\}=10\{,\}000training time steps andTtest=2,000T\_\{\\mathrm\{test\}\}=2\{,\}000test time steps; the test inputs are drawn with a different RNG seed from the same distribution, so train and test share the same dynamics but no shared input samples\.

The predictor sees a length\-mmsliding window of past observations and predicts the current output\. From a trajectory of lengthTTthis yieldsN=T−mN=T\-mexamples; the input at indexiiis\(ui\+1,…,ui\+m\)\(u\_\{i\+1\},\\dots,u\_\{i\+m\}\)together with, in the\+y\+yvariant, the corresponding past outputs\(yi,…,yi\+m−1\)\(y\_\{i\},\\dots,y\_\{i\+m\-1\}\), and the target isyi\+my\_\{i\+m\}\. So the training set containsNtr=Ttrain−mN\_\{\\mathrm\{tr\}\}=T\_\{\\mathrm\{train\}\}\-msamples and the held\-out test set containsTtest−mT\_\{\\mathrm\{test\}\}\-msamples\.

### C\.2Predictor initializations

All predictors are strictly linear; the only nonlinearity in any of the parameterizations is thetanh\\tanhon Diagonal LDS eigenvalues\. We use default PyTorch dtypes \(float32 throughout training\) and the random initialization given here, applied with the same seed used for the ground\-truth LDS so that any seed\-to\-seed variance reflects a single jointly resampled configuration of data and model\.

- •STU\(withkkspectral filters\): the spectral filtersϕ∈ℝm×k\\phi\\in\\mathbb\{R\}^\{m\\times k\}are computed deterministically from them×mm\\times mHankel matrix and held fixed\. The mixerM∈ℝk×din×doutM\\in\\mathbb\{R\}^\{k\\times d\_\{\\mathrm\{in\}\}\\times d\_\{\\mathrm\{out\}\}\}is initialized i\.i\.d\.𝒩​\(0,1/\(k⋅din\)\)\\mathcal\{N\}\(0,1/\(k\\cdot d\_\{\\mathrm\{in\}\}\)\)\. For the experiments withm=512m=512, we choose SpectraLDSh=160h=160, and otherwise we chooseh=120h=120\.
- •Diagonal LDSwith hidden dimensionhh: the raw eigenvalue parameter isλ~∼Uniform​\(−2,2\)h\\tilde\{\\lambda\}\\sim\\mathrm\{Uniform\}\(\-2,2\)^\{h\}, and the trained eigenvalues areλ=tanh⁡\(λ~\)⋅0\.999\\lambda=\\tanh\(\\tilde\{\\lambda\}\)\\cdot 0\.999, which keeps everyλi∈\(−0\.999,0\.999\)\\lambda\_\{i\}\\in\(\-0\.999,0\.999\)for the entire training run and so guarantees stability without needing projection or clipping\. The input matrixB∈ℝh×dinB\\in\\mathbb\{R\}^\{h\\times d\_\{\\mathrm\{in\}\}\}has entries𝒩​\(0,1/din\)\\mathcal\{N\}\(0,1/d\_\{\\mathrm\{in\}\}\)and the output matrixC∈ℝdout×hC\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times h\}has entries𝒩​\(0,1/h\)\\mathcal\{N\}\(0,1/h\)\. We train the Diagonal LDS model convolutionally for maximal efficiency\.
- •General LDSwith hidden dimensionhh:A∈ℝh×hA\\in\\mathbb\{R\}^\{h\\times h\}is drawn from PyTorch’snn\.init\.orthogonal\_and then rescaled in\-place byρinit/ρ​\(A\)\\rho\_\{\\mathrm\{init\}\}/\\rho\(A\)withρinit=0\.95\\rho\_\{\\mathrm\{init\}\}=0\.95so thatρ​\(A\)=0\.95\\rho\(A\)=0\.95at initialization\. After this,AAis an unconstrainednn\.Parameter; we do not project, clip, or otherwise normalizeAAduring training, and rely only on the per\-step gradient\-norm clip described below to bound updates\. The input and readout matrices use the same scheme as the Diagonal LDS\.
- •Linear FIR: the per\-tap kernelW∈ℝm×dout×din\(±y\)W\\in\\mathbb\{R\}^\{m\\times d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}^\{\(\\pm y\)\}\}has entries𝒩​\(0,1/\(m​din\(±y\)\)\)\\mathcal\{N\}\(0,1/\(m\\,d\_\{\\mathrm\{in\}\}^\{\(\\pm y\)\}\)\), wheredin\(±y\)=din\+doutd\_\{\\mathrm\{in\}\}^\{\(\\pm y\)\}=d\_\{\\mathrm\{in\}\}\+d\_\{\\mathrm\{out\}\}in the\+y\+yvariant anddind\_\{\\mathrm\{in\}\}in the−y\-yvariant\. The factor ofmmin the variance is necessary to keep the predicted output’s varianceO​\(1\)O\(1\)at initialization; without it, themm\-tap sum makes early\-epoch predictions blow up by roughlym\\sqrt\{m\}and Adam’s first updates are catastrophically large\.

### C\.3Optimization and hyperparameter sweep

#### First\-order baseline \(Adam\)\.

All Adam\-trained models are run for 100 epochs at batch size 256 on the small\-scale and high\-spectral\-radius experiments and batch size 64 at the large scale \(where the longer windows and larger hidden dimensions make the larger batch run out of memory\)\. After each backward pass we applytorch\.nn\.utils\.clip\_grad\_norm\_\(model\.parameters\(\), 1\.0\)\. Adam’s other hyperparameters are PyTorch defaults \(β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999,ε=10−8\\varepsilon=10^\{\-8\}, weight decay0\)\. For each \(model,±y\\pm y\) we sweep the learning rate at seed0over the gridηAdam∈\{10−1,10−2,10−3,10−4,10−5\}\\eta\_\{\\mathrm\{Adam\}\}\\in\\\{10^\{\-1\},10^\{\-2\},10^\{\-3\},10^\{\-4\},10^\{\-5\}\\\}, select the value minimizing the geometric mean of final\-epoch test MSE across the symmetric and asymmetric tasks, and rerun at seeds\{0,1,2,3,4\}\\\{0,1,2,3,4\\\}with that LR\.

#### Second\-order learner for STU \(ONS, used in the high\-spectral\-radius experiment\)\.

The STU is linear in its parameters\. Collecting the spectral features into the flattened vectorg∈ℝk⋅din\(±y\)g\\in\\mathbb\{R\}^\{k\\cdot d\_\{\\mathrm\{in\}\}^\{\(\\pm y\)\}\}whereg=vec​\(gf​u​l​l\)g=\\text\{vec\}\(g\_\{full\}\)forgf​u​l​l​\[n,i\]=∑txrev​\[t,i\]​ϕ​\[t,n\]g\_\{full\}\[n,i\]=\\sum\_\{t\}x\_\{\\mathrm\{rev\}\}\[t,i\]\\,\\phi\[t,n\], the predictor reduces toy^=M⊤​g\\hat\{y\}=M^\{\\top\}gfor a flattened mixerMM\. Thus, for MSE loss, Online Newton Step\[[6](https://arxiv.org/html/2608.05416#bib.bib45)\]attainsO​\(log⁡T\)O\(\\log T\)regret\. We use the standard ONS update

M←M−η​At−1​gt​\(M⊤​gt−yt\)⊤,M\\leftarrow M\-\\eta\\,A\_\{t\}^\{\-1\}\\,g\_\{t\}\\,\(M^\{\\top\}g\_\{t\}\-y\_\{t\}\)^\{\\top\},with the inverse Hessian estimate maintained by Sherman–Morrison fromAt=α​I\+∑s≤tgs​gs⊤A\_\{t\}=\\alpha I\+\\sum\_\{s\\leq t\}g\_\{s\}g\_\{s\}^\{\\top\}\. We choseη=1\\eta=1and fixedα=1\\alpha=1after observing instability with our earlierα=1×10−3\\alpha=1\\times 10^\{\-3\}default\. Thus the final choice is\(η,α\)=\(1,1\)\(\\eta,\\alpha\)=\(1,1\)\.

### C\.4Per\-experiment dimensions

The two cross\-scale settings and the high\-spectral\-radius setting share the protocol above; only the dimensions differ\.

Table 3:Dimensions used in each experiment\. The number of spectral filterskkis fixed at4848throughout; for the parameter\-matched Diagonal and General LDS variants, the hidden dimensionhhis chosen so that the predictor matches STU’s parameter count \(k​\(din\(±y\)\)​doutk\(d\_\{\\mathrm\{in\}\}^\{\(\\pm y\)\}\)\\,d\_\{\\mathrm\{out\}\}\) at each scale\.In all settings the training set hasNtr=Ttrain−mN\_\{\\mathrm\{tr\}\}=T\_\{\\mathrm\{train\}\}\-msamples and the test set hasTtest−mT\_\{\\mathrm\{test\}\}\-msamples; withTtrain=10,000T\_\{\\mathrm\{train\}\}=10\{,\}000andTtest=2,000T\_\{\\mathrm\{test\}\}=2\{,\}000, the longest\-window cases \(m=512m=512: cross\-scale large and high spectral radius\) leave9,4889\{,\}488training windows and1,4881\{,\}488test windows\.

### C\.5Hyperparameter sweep

In this section, we include the results of the learning rate hyperparameter sweep for the learning linear dynamical systems experiments in Section[4](https://arxiv.org/html/2608.05416#S4)\.

Table 4:Learning\-rate sweep at spectral radiusρ=0\.999\\rho=0\.999\(din=dout=16d\_\{\\mathrm\{in\}\}\{=\}d\_\{\\mathrm\{out\}\}\{=\}16,dstate=64d\_\{\\mathrm\{state\}\}\{=\}64,m=512m\{=\}512, 100 epochs, seed 0\)\. Final test MSE on symmetric and asymmetric tasks\. For STU\+ONS, we fixedη=1\\eta=1and testedα∈\{1,1−3\}\\alpha\\in\\\{1,1^\{\-3\}\\\}, choosingα=1\\alpha=1\.Table 5:Learning\-rate sweep at the small scale \(din=dout=4d\_\{\\mathrm\{in\}\}\{=\}d\_\{\\mathrm\{out\}\}\{=\}4,dstate=8d\_\{\\mathrm\{state\}\}\{=\}8,m=256m\{=\}256\) and the large scale \(din=dout=dstate=64d\_\{\\mathrm\{in\}\}\{=\}d\_\{\\mathrm\{out\}\}\{=\}d\_\{\\mathrm\{state\}\}\{=\}64,m=512m\{=\}512\)\. We report final test MSE after100100\-epochs of training at seed0with Adam\. Bold marks the best LR per model and scale\.

## Appendix DMuJoCo behavior\-cloning data and training protocol

This appendix documents the teacher source, dataset construction, shared architecture, sequence\-mixing\-block details, training hyperparameters, and evaluation protocol used in the MuJoCo experiments of Section[4\.2](https://arxiv.org/html/2608.05416#S4.SS2)\.

### D\.1Environments

We evaluate the train\-then\-distill pipeline on six MuJoCo\-v5tasks spanning a range of locomotion and manipulation difficulties\. Table[6](https://arxiv.org/html/2608.05416#A4.T6)lists each environment alongside its observation and action dimensions, along with the history lengthmm, embedding dimensiondd, and number of spectral filterskkused by the student\. For SpectraLDS distillation, we useh=80h=80\.

Table 6:The six MuJoCo environments used in our behavior\-cloning experiments, listed with their observation dimensiondod\_\{o\}, action dimensiondad\_\{a\}, history lengthmm, and the used embedding dimensiondd\. Each student receives the concatenation\[ot,at−1\]\[o\_\{t\},a\_\{t\-1\}\]at each time step, so the input dimension to the encoder isdo\+dad\_\{o\}\+d\_\{a\}\. Environment descriptions are adapted from the Minari documentation\[[23](https://arxiv.org/html/2608.05416#bib.bib44)\]of the MuJoCo environments\[[20](https://arxiv.org/html/2608.05416#bib.bib42)\]\.
### D\.2Teachers and trajectory data

For each of the six environments we distill a single Soft Actor\-Critic\[[4](https://arxiv.org/html/2608.05416#bib.bib43)\]teacher policy trained by the Minari team\[[23](https://arxiv.org/html/2608.05416#bib.bib44)\]\.

Each teacher is rolled out forT=2,000,000T=2\{,\}000\{,\}000time steps in its environment using a fixed seed and the deterministic action mode\. The resulting trajectory\(ot,at\)t=0T−1\(o\_\{t\},a\_\{t\}\)\_\{t=0\}^\{T\-1\}is the only data the student ever sees: no rollouts of the student itself appear in the training distribution\. From this trajectory we construct a sliding\-window supervised dataset

𝒟=\{\(\(ot−m\+1:t,at−m:t−1\),at\):t=m,…,T−1\},\\mathcal\{D\}\\;=\\;\\big\\\{\\,\\big\(\(o\_\{t\-m\+1:t\},\\,a\_\{t\-m:t\-1\}\),\\;a\_\{t\}\\big\)\\,:\\,t=m,\\dots,T\-1\\,\\big\\\},in which the student observes a length\-mmwindow of past observations and previous actions at each time step and is trained, with mean\-squared error against the teacher’s actionata\_\{t\}, to predict the next action\. We do not subsample or shuffle the trajectory, so the entireT−mT\-msamples are used as training data for every student\.

The “test” set in this experiment is the live environment itself: trained students are deployed deterministically and evaluated onNeval=100N\_\{\\mathrm\{eval\}\}=100episodes per checkpoint\. We report best\-checkpoint reward, defined as the mean episode reward of the checkpoint that achieves the highest mean reward over its1,0001\{,\}000deployment episodes\. We save checkpoints every 5 epochs\. Each episode runs to the environment’s default time limit or natural termination under the standard\-v5reward function\.

### D\.3Shared encoder, head, and action squash

To isolate the effect of the sequence\-mixing block from the surrounding architecture, all three students share an identical encoder and decoder architecture with embedding dimensionddthat depends on the environment, \(see Table[6](https://arxiv.org/html/2608.05416#A4.T6)\)\.

- •Encoder\(ℝdo\+da→ℝd\\mathbb\{R\}^\{d\_\{o\}\+d\_\{a\}\}\\to\\mathbb\{R\}^\{d\}\): a two\-layer GELU MLP applied independently at each time step\. Both hidden layers have widthdd, no normalization layers, and GELU is applied between the two linear layers\. The encoder concatenates the current observation with the previous action before applying the MLP, so each time step in the sliding window contributes a singledd\-dimensional embedding to the sequence\-mixing block’s input\.
- •Sequence\-mixing block\(ℝm×d→ℝd\\mathbb\{R\}^\{m\\times d\}\\to\\mathbb\{R\}^\{d\}\): the only component that varies across the three students\. It consumes the length\-mmsequence of encoder embeddings and produces a singledd\-dimensional summary\. The three variants are detailed in Appendix[D\.4](https://arxiv.org/html/2608.05416#A4.SS4)\.
- •Decoder\(ℝd→ℝda\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\_\{a\}\}\): a three\-layer GELU MLP with first layerdd, second layerd/2d/2, and outputdad\_\{a\}, with GELU between layers and a finaltanhnormalization and scaling to ensure the student’s predictions live in the valid action box\.

The only architectural variation across the three students lies in the sequence\-mixing block\.

### D\.4Sequence\-mixing\-block variants

The three students differ only in their sequence\-mixing block\.

#### STU \(the train\-then\-distill variant\)\.

At training time, the block is a tensor\-dot Spectral Transform Unit\[[10](https://arxiv.org/html/2608.05416#bib.bib37)\]withk=16k=16spectral filters\. We use the tensor\-dot variant for parameter efficiency atd=128d=128; Appendix[F](https://arxiv.org/html/2608.05416#A6)compares it against the STU on the linear\-LDS task\. At evaluation time, the trained STU is converted to a Diagonal LDS via SpectraLDS, and the converted Diagonal LDS is what we deploy\. To allow for easier reuse of the published SpectraLDS conversions\[[17](https://arxiv.org/html/2608.05416#bib.bib24)\], we use the spectral filters computed withm=8192m=8192andk=48k=48and at8080\-dimensions, truncate, and consider the relevant subset ofk=16k=16filters\. The STU runs infloat64precision due to historical reasons, while other models are in the nativefloat32\. Based on later experimentation, we do not expect this contributes to any differences\.

#### Diagonal LDS,CC\-matched \(h=dh=d\)\.

A directly\-trained recurrencext\+1=diag​\(λ\)​xt\+B​ut,yt=C​xtx\_\{t\+1\}=\\mathrm\{diag\}\(\\lambda\)\\,x\_\{t\}\+Bu\_\{t\},\\ y\_\{t\}=Cx\_\{t\}, withλ=tanh⁡\(λ~\)​\(1−10−3\)\\lambda=\\tanh\(\\tilde\{\\lambda\}\)\\,\(1\-10^\{\-3\}\)to keep eigenvalues strictly inside the unit disk\.

#### Diagonal LDS, parameter\-matched\.

Identical to theCC\-matched variant in form, but withhhchosen so thath​\(1\+2​d\)=2​k​d\+d2h\(1\+2d\)=2kd\+d^\{2\}, giving exactly the same parameter count as STU’s tensor\-dot block\. This is the parameter\-matched baseline for the train\-then\-distill comparison\.

### D\.5Per\-student parameter counts

Table[7](https://arxiv.org/html/2608.05416#A4.T7)reports the total trainable parameters of each student alongside the SAC teacher’s parameter count\. All three students are smaller than the SAC actor used at inference time on every environment, and roughly an order of magnitude smaller than the full SAC checkpoint \(which carries two critic networks and their target copies in addition to the actor\)\. Within the students, STU and the parameter\-matched Diagonal LDS have essentially identical parameter counts by construction, while theCC\-matched Diagonal LDS is approximately twenty to twenty\-five percent larger because theh=dh=dchoice inflates its sequence\-mixing block\.

Table 7:Total trainable parameters per student model and the corresponding SAC teacher across the six MuJoCo environments\. The ”SAC actor” column gives the parameter count of the policy network used at inference time, while ”SAC full” adds the critic and target\-critic networks carried in the SAC checkpoint and so reflects the size of the full SAC artifact\. The SpectraLDS and parameter\-matched LDS students use fewer parameters than the SAC actor on every environment, and within the students theCC\-matched Diagonal LDS is uniformly larger than the STU pipeline by roughly twenty to twenty\-five percent\.
### D\.6Optimization and training protocol

All three students are trained for100100epochs on theT−mT\-m\-sample dataset using Adam at a fixed learning rate of3×10−43\\times 10^\{\-4\}and batch size1,0241\{,\}024\. We use PyTorch’s default Adam hyperparameters otherwise \(β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999,ε=10−8\\varepsilon=10^\{\-8\}, weight decay0\)\. We do not sweep the learning rate in MuJoCo; we use3×10−43\\times 10^\{\-4\}for every \(environment, student\) pair, and verified on a small set of pilot runs that all three students train stably at this rate\. We run five seeds\{0,1,2,3,4\}\\\{0,1,2,3,4\\\}per environment\-student pair\.

Checkpoints are saved every five training epochs and at the end of training\. Each checkpoint is evaluated on100100episodes, and the reward we report in Table[2](https://arxiv.org/html/2608.05416#S4.T2)is from taking the best checkpoint and re\-evaluating it on1,0001\{,\}000episodes\.

## Appendix ESpectraLDS distillation fidelity

A key step in our evaluation pipeline is converting the trained STU predictor into a SpectraLDS \(diagonal\-LDS\) predictor\. The filter fit is approximate, so we verify here that this approximation does not meaningfully change the test MSE\.

Figure[4](https://arxiv.org/html/2608.05416#A5.F4)overlays the STU test MSE \(from the original five\-seed runs\) and the SpectraLDS\-distilled test MSE \(from independent five\-seed runs on the same seeds\), for both the Adam and ONS optimizers on the high\-spectral\-radius setting \(ρ=0\.999\\rho=0\.999,d=16d=16,dstate=64d\_\{\\mathrm\{state\}\}=64,m=512m=512\)\. Each curve is the geometric mean over five seeds with geometric\-standard\-deviation bands; the solid line is the distilled model and the dashed line is the raw STU\.

The two curves are visually indistinguishable in every panel\. Quantitatively, the maximum relative error between the raw and distilled test MSEs, taken over all epochs and all seeds, is below1\.5%1\.5\\%in all conditions\.

![Refer to caption](https://arxiv.org/html/2608.05416v1/images/spectralds_fidelity.png)Figure 4:Test MSE versus epoch for the STU predictor \(dashed\) and the SpectraLDS\-distilled predictor \(solid\), atρ=0\.999\\rho=0\.999\. Each curve is the geometric mean over five seeds with geometric\-standard\-deviation bands; theyy\-axis is log\-scaled\. The solid and dashed curves overlap almost exactly in every panel, confirming that SpectraLDS distillation introduces negligible error\.
## Appendix FSTU vs\. tensor\-dot STU

Both STU variants begin from the same intermediate object\. Given an input windowx∈ℝm×dinx\\in\\mathbb\{R\}^\{m\\times d\_\{\\mathrm\{in\}\}\}and the fixed spectral filtersϕ∈ℝm×k\\phi\\in\\mathbb\{R\}^\{m\\times k\}, the spectrally\-filtered representation of the window is the matrix

Z=ϕ⊤​x∈ℝK×din,Zk,i=⟨ϕ:,k,x:,i⟩\.Z\\;=\\;\\phi^\{\\top\}x\\;\\in\\;\\mathbb\{R\}^\{K\\times d\_\{\\mathrm\{in\}\}\},\\qquad Z\_\{k,i\}\\;=\\;\\big\\langle\\phi\_\{:,k\},\\;x\_\{:,i\}\\big\\rangle\.ZZhas one row per spectral filter and one column per input channel; all of the time\-window information that the predictor will use lives in thisk×dink\\times d\_\{\\mathrm\{in\}\}matrix\. Both STU variants then linearly mapZZto adoutd\_\{\\mathrm\{out\}\}\-dimensional predictionyy\.

#### STU\.

The STU uses an unconstrained linear map fromℝk×din\\mathbb\{R\}^\{k\\times d\_\{\\mathrm\{in\}\}\}toℝdout\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\}, parameterized by a tensorM∈ℝk×din×doutM\\in\\mathbb\{R\}^\{k\\times d\_\{\\mathrm\{in\}\}\\times d\_\{\\mathrm\{out\}\}\}:

y=∑k,iMk,i,:​Zk,i\.y\\;=\\;\\sum\_\{k,i\}M\_\{k,i,:\}\\;Z\_\{k,i\}\.Every element ofMMis independently learned; the map hask​din​doutk\\,d\_\{\\mathrm\{in\}\}\\,d\_\{\\mathrm\{out\}\}free parameters\.

#### Tensor\-dot STU\.

The tensor\-dot variant replaces the single mixer tensorM∈ℝk×din×doutM\\in\\mathbb\{R\}^\{k\\times d\_\{\\mathrm\{in\}\}\\times d\_\{\\mathrm\{out\}\}\}with two matrices,

Min∈ℝk×din,Mout∈ℝdin×dout,M\_\{\\mathrm\{in\}\}\\in\\mathbb\{R\}^\{k\\times d\_\{\\mathrm\{in\}\}\},\\qquad M\_\{\\mathrm\{out\}\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\\times d\_\{\\mathrm\{out\}\}\},constraining the mixer to factorize asMk,i,o=Min,k,i⋅Mout,i,oM\_\{k,i,o\}=M\_\{\\mathrm\{in\},k,i\}\\cdot M\_\{\\mathrm\{out\},i,o\}\. Two benefits follow\.

First, the parameter count drops fromk​din​doutk\\,d\_\{\\mathrm\{in\}\}\\,d\_\{\\mathrm\{out\}\}tok​din\+din​doutk\\,d\_\{\\mathrm\{in\}\}\+d\_\{\\mathrm\{in\}\}\\,d\_\{\\mathrm\{out\}\}— at the MuJoCo dimensionsdin=dout=128d\_\{\\mathrm\{in\}\}=d\_\{\\mathrm\{out\}\}=128andk=16k=16this is a7\.5×7\.5\\timesreduction\.

Second, the factorization permits a cheaper convolutional implementation\. In the STU, predictingyyinvolvesk⋅dink\\cdot d\_\{\\mathrm\{in\}\}distinct filter convolutions\. The tensor\-dot factorization lets us instead combineMinM\_\{\\mathrm\{in\}\}with the spectral filters intodind\_\{\\mathrm\{in\}\}per\-input\-channel filtersf:,i=∑kMin,k,i​ϕ:,kf\_\{:,i\}=\\sum\_\{k\}M\_\{\\mathrm\{in\},k,i\}\\,\\phi\_\{:,k\}, then perform onlydind\_\{\\mathrm\{in\}\}time\-convolutions \(one per input channel\) before mixing throughMoutM\_\{\\mathrm\{out\}\}\. The number of convolutions thus drops fromk⋅dink\\cdot d\_\{\\mathrm\{in\}\}todind\_\{\\mathrm\{in\}\}\.

However, the tensor\-dot STU does have reduced representation capacity, which we analyze in the learning an LDS setting\.

#### Comparison\.

Figure[5](https://arxiv.org/html/2608.05416#A6.F5)reports test MSE versus epoch for both STU variants in\+y\+yand−y\-yform, at both the small \(dstate=8d\_\{\\mathrm\{state\}\}=8,m=256m=256\) and large \(dstate=64d\_\{\\mathrm\{state\}\}=64,m=512m=512\) scales\. Each curve is the geometric mean over five seeds with geometric\-standard\-deviation bands, on a log\-scaled vertical axis\. The STU outperforms the tensor\-dot approximation STU in each setting as expected, however, the tensor\-dot STU has approximatelyk×k\\timesfewer parameters\.

![Refer to caption](https://arxiv.org/html/2608.05416v1/images/stu_canonical_vs_tensordot.png)Figure 5:Test MSE versus epoch for STU \(blue\) and tensor\-dot STU \(red\), both withk=48k=48, and each in\+y\+y\(solid\) and−y\-y\(dashed\) variants, on the linear\-LDS benchmark of Section[4\.1](https://arxiv.org/html/2608.05416#S4.SS1)\. Top row: small scale \(d=4d=4,dstate=8d\_\{\\mathrm\{state\}\}=8,m=256m=256\)\. Bottom row: large scale \(d=64d=64,dstate=64d\_\{\\mathrm\{state\}\}=64,m=512m=512\)\. Curves are geometric means over five seeds with geometric\-standard\-deviation bands; the y\-axis is log\-scaled\.

Similar Articles