Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin Systems
Summary
This paper introduces Sparse Koopman Autoencoders (SKAEs) to identify local dynamical regimes in multibasin nonlinear systems, demonstrating superior forecasting performance and interpretable latent supports.
View Cached Full Text
Cached at: 09/01/26, 01:06 PM
# Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin Systems
Source: [https://arxiv.org/html/2608.29057](https://arxiv.org/html/2608.29057)
Aidan LiAffiliation:Chandar Research LabAffiliation:Mila – Quebec AI InstituteAffiliation:Université de MontréalEmail:[aidan\.li@mila\.quebec](mailto:)Mahan FathiAffiliation:NVIDIASarath ChandarAffiliation:Chandar Research LabAffiliation:Mila – Quebec AI InstituteAffiliation:Polytechnique MontréalAffiliation:Canada CIFAR AI ChairRoss GoroshinAffiliation:Mila – Quebec AI InstituteAffiliation:Université de Montréal
###### Abstract
Koopman autoencoders \(KAEs\) seek a higher\-dimensional latent representation in which nonlinear dynamics evolve linearly\. However, many interesting systems have multiple basins of attraction, and both theoretical and empirical work has shown these multibasin systems cannot generally admit a single finite\-dimensional global Koopman embedding under standard assumptions\. We posit that encoders with a sparsity\-inducing objective encouraging few active latent coefficients will provide latent supports as an inspectable basin\-modeling principle for Koopman autoencoders\. We use these encoders producing sparse latents in training Sparse Koopman Autoencoders \(SKAEs\) without basin labels or other regime annotations, and treat the learned latent supports as model\-produced regime variables after training\. Across a range of procedurally generated multibasin systems and chaotic flows, we show that SKAEs have superior forecasting performance compared to dense\-latent KAEs\. We also perform a mechanistic study that shows latent supports produced by SKAEs are both essential for the quality of the representation and useful for identifying basins on held\-out basin interior states, whereas dense\-latent KAEs collapse to an uninformative single family\. These results identify sparse latents and their corresponding supports as label\-free, interpretable regime variables for Koopman learning in nonlinear systems with multiple local dynamical laws\.
## 1Introduction
Learning representations for nonlinear dynamical systems remains a central problem in machine learning, scientific computing, and control\. A particularly useful goal is to embed nonlinear dynamics into coordinates in which the evolution is linear, or at least locally linear, because linear models can be combined with mature tools for prediction, estimation, and control, including model predictive control and linear quadratic regulation\[[28](https://arxiv.org/html/2608.29057#bib.bib28),[15](https://arxiv.org/html/2608.29057#bib.bib15),[20](https://arxiv.org/html/2608.29057#bib.bib20),[17](https://arxiv.org/html/2608.29057#bib.bib17),[27](https://arxiv.org/html/2608.29057#bib.bib27)\]\. This goal is also well motivated theoretically\. Classical dynamical\-systems results provide local topological linearizations near hyperbolic fixed points\[[14](https://arxiv.org/html/2608.29057#bib.bib14),[16](https://arxiv.org/html/2608.29057#bib.bib16)\], while Koopman eigenfunction constructions can yield linearizing coordinates on basins of attraction for stable equilibria and periodic orbits\[[23](https://arxiv.org/html/2608.29057#bib.bib23),[22](https://arxiv.org/html/2608.29057#bib.bib22)\]\. Thus, nonlinear systems may admit linear dynamics with a specific change of coordinates\.
These results establish when useful coordinates can exist, but they do not provide a finite\-data procedure for constructing such coordinates\. Koopman operator theory formalizes the linearization idea by lifting nonlinear state evolution to a linear operator acting on observables\[[18](https://arxiv.org/html/2608.29057#bib.bib18),[19](https://arxiv.org/html/2608.29057#bib.bib19),[29](https://arxiv.org/html/2608.29057#bib.bib29),[6](https://arxiv.org/html/2608.29057#bib.bib6)\]\. This viewpoint has inspired data\-driven approximations of Koopman dynamics, such as DMD and eDMD\[[33](https://arxiv.org/html/2608.29057#bib.bib33),[36](https://arxiv.org/html/2608.29057#bib.bib36),[37](https://arxiv.org/html/2608.29057#bib.bib37)\], and more recently to Koopman autoencoders, which learn an encoder from state space to a latent representation, a linear latent transition matrix, and a decoder back to state space\[[35](https://arxiv.org/html/2608.29057#bib.bib35),[26](https://arxiv.org/html/2608.29057#bib.bib26),[31](https://arxiv.org/html/2608.29057#bib.bib31),[3](https://arxiv.org/html/2608.29057#bib.bib3),[34](https://arxiv.org/html/2608.29057#bib.bib34),[10](https://arxiv.org/html/2608.29057#bib.bib10)\]\. These models provide an appealing route to practical learning: rather than hand\-designing observables, they learn embeddings directly from trajectory data\.
An important problem for Koopman approximations is that many practically interesting nonlinear systems typically contain multiple equilibria or basins of attraction; for example, a damped mechanical system can settle into different wells, a biological or chemical model can approach different stable equilibria, and a controlled system can move between regions with distinct local dynamics\. However, theory shows that these systems cannot generally admit a single finite\-dimensional global linearization with a continuous encoder\-decoder pair\[[23](https://arxiv.org/html/2608.29057#bib.bib23),[5](https://arxiv.org/html/2608.29057#bib.bib5),[32](https://arxiv.org/html/2608.29057#bib.bib32),[25](https://arxiv.org/html/2608.29057#bib.bib25)\], and empirical evidence supports this in practice\[[10](https://arxiv.org/html/2608.29057#bib.bib10)\]\. In such multistable systems, it is possible to linearize the dynamics within basins, but the linearizations needed in different basins of attraction are generally incompatible with each other\. Therefore, for multi\-attractor systems, we should not model all basins with one global linearization, and instead have distinct modeling of basins\.
This paper proposes inducing sparse latent structure as a practical and interpretable method for Koopman learning of multi\-attractor systems\. By inducing the state to activate only a small subset of coordinates in lifted latent space, we can incentivize different regions of state to occupy different active coordinates, since nearby states should have similar latent embeddings, and thus enable learning distinct local regimes\. In this sparse representation, nonzero coefficient values in a latent vector can provide continuous coordinates for prediction within a selected region, while the active indices provide a discrete support object that can be inspected across states and used to identify the appropriate local regime\.
Sparse coding and LISTA\-style encoders\[[13](https://arxiv.org/html/2608.29057#bib.bib13),[8](https://arxiv.org/html/2608.29057#bib.bib8),[1](https://arxiv.org/html/2608.29057#bib.bib1)\]provide the right inductive bias for this purpose\. They naturally induce piecewise\-linear, union\-of\-subspaces structure\[[30](https://arxiv.org/html/2608.29057#bib.bib30),[7](https://arxiv.org/html/2608.29057#bib.bib7),[9](https://arxiv.org/html/2608.29057#bib.bib9),[13](https://arxiv.org/html/2608.29057#bib.bib13)\], which is well\-matched to this setting where different attractors or basins require different local linearizations in theory\. Concurrent work on LpWorldModel \(LpWM\) arrives independently at a closely related view of sparse representations, finding that support can identify discrete dynamical regimes while feature magnitudes encode continuous within\-regime state\[[21](https://arxiv.org/html/2608.29057#bib.bib21)\]\. Whereas LpWM studies controlled, action\-conditioned JEPA world models for planning, we study multibasin dynamics through the Koopman framework, using sparse supports as label\-free regime variables under a learned linear latent evolution\. We hypothesize that sparse\-latent Koopman autoencoders will assign local dynamical regimes to distinct latent support families, yielding a learned regime variable without basin supervision\. Then, periodically decoding and re\-encoding the state can refresh the quality of the latent embedding over time by re\-selecting the correct local dynamics when trajectories move across the state space\[[10](https://arxiv.org/html/2608.29057#bib.bib10)\]\.
#### Contributions\.
We introduce Sparse Koopman Autoencoders \(SKAEs\), showing that sparse latent supports can serve as label\-free regime variables that reconcile basin\-specific Koopman linearizations with the practicality of a single trained model\. Across multibasin and chaotic benchmarks, SKAEs \(i\) improve long\-horizon forecasting while \(ii\) producing supports that are functionally necessary for accurate rollouts and \(iii\) interpretable as emergent local dynamical regimes\.
## 2Problem formulation
Figure 1:Top: procedurally generated multibasin systems where gray streamlines are the true vector field, colored dots mark support\-family/basin matches, and black dots mark support\-family/basin mismatches\. The learned LISTA support families are mapped post hoc to each family’s dominant evaluation basin\. Bottom: Dysts phase portraits with held\-out truth in gray and forecasts in color\. Red lines are predictions by the seed\-0 dense\-latent MLP, and blue lines are from the best seed\-0 LISTA variant\. The dense MLP makes significant mistakes in forecasting the Chua system, unlike LISTA\.The introduction motivates a multibasin version of Koopman learning\. Given sampled trajectories of a nonlinear system, the aim is to learn a state\-reconstructing latent representation whose evolution is linear, without being told which basin or local dynamical regime generated each observation\. We formulate the problem at the cadence of the stored benchmark observations\. For each system, let𝒳⊆ℝdx\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\_\{x\}\}be the observed state space and let
xk\+1=FΔt\(xk\),xk∈𝒳,x\_\{k\+1\}=F\_\{\\Delta t\}\(x\_\{k\}\),\\qquad x\_\{k\}\\in\\mathcal\{X\},\(1\)wherexk∈ℝdxx\_\{k\}\\in\\mathbb\{R\}^\{d\_\{x\}\}is the state at thekkth stored sample,FΔtF\_\{\\Delta t\}denotes the transformation to the next discrete observation step based on underlying dynamics, andΔt\\Delta tis the system\-specific observation interval\. Forecasting is therefore defined by repeatedly applying the mapFΔtF\_\{\\Delta t\}; a horizon ofhhmeanshhapplications ofFΔtF\_\{\\Delta t\}; we denotekkcompositions asFΔt\(k\)F\_\{\\Delta t\}^\{\\,\(k\)\}\. BecauseΔt\\Delta tmay differ across systems, equal values ofhhmay yield different physical time\.
#### Basins of attraction\.
Many systems of interest are multistable: different initial conditions can approach different long\-run behaviors\. A fixed point is a statex∗x^\{\\ast\}satisfyingFΔt\(x∗\)=x∗F\_\{\\Delta t\}\(x^\{\\ast\}\)=x^\{\\ast\}\. More generally, a trajectory may approach an attractorAA, such as a stable equilibrium, limit cycle, countable finite set, or chaotic invariant set\. The basin of attraction ofAAis the set of initial conditions whose trajectories converge toAA:
ℬ\(A\)=\{x0∈ℝdx:infa∈A‖FΔt\(k\)\(x0\)−a‖2→0ask→∞\}\.\\mathcal\{B\}\(A\)=\\left\\\{x\_\{0\}\\in\\mathbb\{R\}^\{d\_\{x\}\}:\\inf\_\{a\\in A\}\\left\\lVert F\_\{\\Delta t\}^\{\\,\(k\)\}\(x\_\{0\}\)\-a\\right\\rVert\_\{2\}\\\!\\to 0\\text\{ as \}k\\to\\infty\\right\\\}\.\(2\)When multiple attractors coexist, their basins partition the state space up to basin boundaries\. This defines the setup for this paper\. The task is to learn a representation that can support linear prediction where the appropriate linear law depends on where the state lies in the multibasin geometry\.
#### Koopman theory\.
The stored\-step mapFΔtF\_\{\\Delta t\}induces a Koopman operator𝒦\\mathcal\{K\}on scalar observablesg:𝒳→ℝg:\\mathcal\{X\}\\to\\mathbb\{R\}by composition,\(𝒦g\)\(x\)=g\(FΔt\(x\)\)\(\\mathcal\{K\}g\)\(x\)=g\(F\_\{\\Delta t\}\(x\)\)\. AlthoughFΔtF\_\{\\Delta t\}may be nonlinear,𝒦\\mathcal\{K\}is linear on the space of observables\. A finite\-dimensional Koopman approximation searches for a vector observableΦ:𝒳→ℝdz\\Phi:\\mathcal\{X\}\\to\\mathbb\{R\}^\{d\_\{z\}\}and a matrixKKsuch thatΦ\(FΔt\(x\)\)≈KΦ\(x\)\\Phi\(F\_\{\\Delta t\}\(x\)\)\\approx K\\Phi\(x\), while retaining enough information to reconstruct the state\.
#### Koopman autoencoders\.
We study Koopman autoencoders \(KAEs\) in this setting\. The model learns an encoderEθ:ℝdx→ℝdzE\_\{\\theta\}:\\mathbb\{R\}^\{d\_\{x\}\}\\rightarrow\\mathbb\{R\}^\{d\_\{z\}\}, a decoderDϕ:ℝdz→ℝdxD\_\{\\phi\}:\\mathbb\{R\}^\{d\_\{z\}\}\\rightarrow\\mathbb\{R\}^\{d\_\{x\}\}, and a latent linear transition matrixK∈ℝdz×dzK\\in\\mathbb\{R\}^\{d\_\{z\}\\times d\_\{z\}\}\. For observationxkx\_\{k\}at arbitrary timestepkk, the corresponding open\-loophh\-step forecast is
x^k\+h=Dϕ\(KhEθ\(xk\)\)\.\\hat\{x\}\_\{k\+h\}=D\_\{\\phi\}\\\!\\left\(K^\{h\}E\_\{\\theta\}\(x\_\{k\}\)\\right\)\.\(3\)The same matrixKKis applied at every step of the rollout, so any state\-dependent structure required for prediction must be encoded in the learned representation rather than in a supplied regime label\.
#### Training windows\.
Training uses windows of adjacent stored observations\. For a batch ofBBwindows of rollout lengthLL,\{xb,k,xb,k\+1,…,xb,k\+L\}b=1B\\\{x\_\{b,k\},x\_\{b,k\+1\},\\ldots,x\_\{b,k\+L\}\\\}\_\{b=1\}^\{B\}, wherebbindexes windows andℓ\\ellindexes offsets within a window, the model encodes each observed state, rolls out from the initial encoded state, and decodes both reconstructed observations and latent forecasts\. The encoded/reconstructed quantities are defined forℓ=0,…,L\\ell=0,\\ldots,L, while rollout/forecast quantities are defined forℓ=1,…,L\\ell=1,\\ldots,L:
zb,k\+ℓenc=Eθ\(xb,k\+ℓ\),zb,k\+ℓroll=Kℓzb,kenc,xb,k\+ℓrec=Dϕ\(zb,k\+ℓenc\),x^b,k\+ℓ=Dϕ\(zb,k\+ℓroll\)\.\\displaystyle z\_\{b,k\+\\ell\}^\{\\rm enc\}=E\_\{\\theta\}\(x\_\{b,k\+\\ell\}\),\\quad z\_\{b,k\+\\ell\}^\{\\rm roll\}=K^\{\\ell\}z\_\{b,k\}^\{\\rm enc\},\\quad x\_\{b,k\+\\ell\}^\{\\rm rec\}=D\_\{\\phi\}\(z\_\{b,k\+\\ell\}^\{\\rm enc\}\),\\quad\\hat\{x\}\_\{b,k\+\\ell\}=D\_\{\\phi\}\(z\_\{b,k\+\\ell\}^\{\\rm roll\}\)\.\(4\)The training losses described later are built from these reconstructed and rolled\-out quantities\. Each application ofKKis interpreted as one stored benchmark step, matching the observation map in[eq\.1](https://arxiv.org/html/2608.29057#S2.E1)\.
#### Model discovers basins unsupervised\.
In the intended training and deployment setting, the model is not given basin labels or trajectory\-to\-basin assignments\. When benchmark basin labels are available, they are reserved for post\-hoc evaluation; they are not used to choose the model, train the encoder or transition, or select supports for support\-conditioned predictions\.
## 3Method
### 3\.1Sparse supports as local regime variables
A finite\-dimensional Koopman representation for the multibasin systems in[Section2](https://arxiv.org/html/2608.29057#S2)should encode two complementary quantities: \(i\) continuous coordinates that linearize the dynamics within a basin, and \(ii\) indicate which local linearization is active\. We use sparsity to couple these two roles\. If the statexxis represented by a sparse latent codezz, then the support ofzzcan act as a regime variable, while the nonzero coefficient values provide coordinates for the corresponding local linear representation\.
Given an overcomplete dictionaryDD, sparse coding estimateszzby solving the Lasso objective
z⋆\(x\)=argminz∈ℝdz12‖x−Dz‖22\+λsc‖z‖1,λsc\>0\.z^\{\\star\}\(x\)=\\operatorname\*\{arg\\,min\}\_\{z\\in\\mathbb\{R\}^\{d\_\{z\}\}\}\\frac\{1\}\{2\}\\left\\lVert x\-Dz\\right\\rVert\_\{2\}^\{2\}\+\\lambda\_\{\\rm sc\}\\left\\lVert z\\right\\rVert\_\{1\},\\qquad\\lambda\_\{\\rm sc\}\>0\.\(5\)which finds a codezzwith few non\-zero coefficients to reconstruct the inputxx\. In other words, it encourages a sparse solution to the under\-specified problem \(classically,ℓ0\\ell\_\{0\}instead ofℓ1\\ell\_\{1\}will find the sparsest solution with the fewest non\-zero coefficients\)\.
The classical solver to this problem \(ISTA\[[4](https://arxiv.org/html/2608.29057#bib.bib4)\]\) can be unrolled into a depth\-LLnetwork of ISTA iterations, yielding learned ISTA \(LISTA\)\[[13](https://arxiv.org/html/2608.29057#bib.bib13)\]; in our setting, LISTA is an encoder yielding the sparse codez=Eθ\(x\)z=E\_\{\\theta\}\(x\), where the*support*is the set of activezzcoordinates\. With a linear decoder, the support determines which decoder columns participate in reconstructingxx, and the associated nonzero coefficients specify coordinates within that selected reconstruction subspace\. WhenDDandzzare learned jointly, as in sparse coding\[[30](https://arxiv.org/html/2608.29057#bib.bib30)\], this induces a data\-adaptive partition of the input state space where each region can be locally reconstructed with a code having the same support\.
We therefore posit that combining Koopman representations and sparse coding objectives in a sparse Koopman autoencoder \(SKAE\) implicitly captures multibasin attractor dynamics without explicit basin labels in training\. The Koopman objective encourages the continuous coefficients on each support to evolve linearly, while the sparsity objective encourages different local regimes to use different active coordinates\. Thus, the model can represent a multibasin system using a discrete support pattern to identify the active basin and continuous coefficients for its localized Koopman \(linear\) representation\. When ground\-truth basin labels are available in benchmarks, we use them only after training to assess support–basin alignment or to construct diagnostic interventions\.
### 3\.2Koopman autoencoder and sparse encoder
The predictor follows the KAE formulation in[Section2](https://arxiv.org/html/2608.29057#S2): an encoder, linear decoder, and latent transition, all of which are optimized jointly, produce the rollout in[eq\.3](https://arxiv.org/html/2608.29057#S2.E3)\. Like in sparse coding,\[[13](https://arxiv.org/html/2608.29057#bib.bib13)\]we constrain the columns of the decoderDϕ=DD\_\{\\phi\}=Dto be unit norm to avoid degenerate solutions\. We experiment withKKthat is dense, block diagonal, or softly block\-regularized\.
#### Sparse encoding\.
Our main sparse encoder is LISTA\[[13](https://arxiv.org/html/2608.29057#bib.bib13)\]\. It first computes a dense pre\-codect=fθ\(xt\)c\_\{t\}=f\_\{\\theta\}\(x\_\{t\}\), then applies learned shrinkage refinements forq=0,…,Q−1q=0,\\ldots,Q\-1:
ut\(0\)=shrink\(ct,τ\),ut\(q\+1\)=shrink\(Sut\(q\)\+ct,τ\),zt=ut\(Q\)\.u\_\{t\}^\{\(0\)\}=\\operatorname\{shrink\}\(c\_\{t\},\\tau\),\\qquad u\_\{t\}^\{\(q\+1\)\}=\\operatorname\{shrink\}\(Su\_\{t\}^\{\(q\)\}\+c\_\{t\},\\tau\),\\qquad z\_\{t\}=u\_\{t\}^\{\(Q\)\}\.\(6\)Here,shrink\(a,τ\)=sign\(a\)max\(\|a\|−τ,0\)\\operatorname\{shrink\}\(a,\\tau\)=\\operatorname\{sign\}\(a\)\\max\(\|a\|\-\\tau,0\)is applied element\-wise,SSis learned, andτ\\taucontrols the threshold scale for shrinkage\. The thresholding step creates exact zeros, so the encoder explicitly selects active latent coordinates while learning the selection rule from the prediction objective\.
#### Transition families\.
We vary the structure ofKKseparately from the encoder\. The dense transition leavesKKunrestricted\. The block\-diagonal transition partitions latent coordinates intoMMfixed groups and constrainsKblock=blkdiag\(K1,…,KM\)K\_\{\\rm block\}=\\operatorname\{blkdiag\}\(K\_\{1\},\\ldots,K\_\{M\}\)\. We also use a soft\-block LISTA ablation that keepsKKdense but penalizes cross\-group entries,ℒoffblock=∑i,j:g\(i\)≠g\(j\)\|Kij\|\\mathcal\{L\}\_\{\\rm offblock\}=\\sum\_\{i,j:\\,g\(i\)\\neq g\(j\)\}\|K\_\{ij\}\|, whereg\(i\)g\(i\)is the fixed architectural group containing coordinateii\.
### 3\.3Training objective
Training uses adjacent observation windows\. For a batch ofBBwindows with prediction horizonLL, the model produces reconstructed future statesxb,k\+ℓrecx\_\{b,k\+\\ell\}^\{\\rm rec\}, latent rolloutszb,k\+ℓrollz\_\{b,k\+\\ell\}^\{\\rm roll\}, encoded future latentszb,k\+ℓencz\_\{b,k\+\\ell\}^\{\\rm enc\}, and decoded rollout predictionsx^b,k\+ℓ\\hat\{x\}\_\{b,k\+\\ell\}as defined in[eq\.4](https://arxiv.org/html/2608.29057#S2.E4)\. The rollout is seeded by the initial state, and the following sums are over offsetsℓ=0,…,L\\ell=0,\\ldots,L:
ℒ¯pred\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{\\rm pred\}=1BLdx∑b,ℓ‖x^b,k\+ℓ−xb,k\+ℓ‖2,\\displaystyle=\\frac\{1\}\{BL\\sqrt\{d\_\{x\}\}\}\\sum\_\{b,\\ell\}\\left\\lVert\\hat\{x\}\_\{b,k\+\\ell\}\-x\_\{b,k\+\\ell\}\\right\\rVert\_\{2\},ℒ¯rec\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{\\rm rec\}=1BLdx∑b,ℓ‖xb,k\+ℓrec−xb,k\+ℓ‖2,\\displaystyle=\\frac\{1\}\{BL\\sqrt\{d\_\{x\}\}\}\\sum\_\{b,\\ell\}\\left\\lVert x\_\{b,k\+\\ell\}^\{\\rm rec\}\-x\_\{b,k\+\\ell\}\\right\\rVert\_\{2\},\(7\)ℒ¯lin\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{\\rm lin\}=1BL∑b,ℓ‖zb,k\+ℓroll−zb,k\+ℓenc‖2,\\displaystyle=\\frac\{1\}\{BL\}\\sum\_\{b,\\ell\}\\left\\lVert z\_\{b,k\+\\ell\}^\{\\rm roll\}\-z\_\{b,k\+\\ell\}^\{\\rm enc\}\\right\\rVert\_\{2\},ℒ¯sp\\displaystyle\\bar\{\\mathcal\{L\}\}\_\{\\rm sp\}=1BL∑b,ℓ‖zb,k\+ℓroll‖1\.\\displaystyle=\\frac\{1\}\{BL\}\\sum\_\{b,\\ell\}\\left\\lVert z\_\{b,k\+\\ell\}^\{\\rm roll\}\\right\\rVert\_\{1\}\.\(8\)The observation\-space terms are normalized bydx\\sqrt\{d\_\{x\}\}so that their scale is comparable across systems of different state dimension\. The latent terms are left in their native scale because the latent dimension is fixed within each comparison\. The total objective is
ℒ=\(λpredℒ¯pred\+λrecℒ¯rec\+λlinℒ¯lin\+λspℒ¯sp\)\+ℒstruct\.\\mathcal\{L\}=\\left\(\\lambda\_\{\\rm pred\}\\bar\{\\mathcal\{L\}\}\_\{\\rm pred\}\+\\lambda\_\{\\rm rec\}\\bar\{\\mathcal\{L\}\}\_\{\\rm rec\}\+\\lambda\_\{\\rm lin\}\\bar\{\\mathcal\{L\}\}\_\{\\rm lin\}\+\\lambda\_\{\\rm sp\}\\bar\{\\mathcal\{L\}\}\_\{\\rm sp\}\\right\)\+\\mathcal\{L\}\_\{\\rm struct\}\.\(9\)Hereℒstruct=0\\mathcal\{L\}\_\{\\rm struct\}=0for the dense and block\-diagonal transition variants\. For the soft\-block transition,ℒstruct=λoffℒoffblock\\mathcal\{L\}\_\{\\rm struct\}=\\lambda\_\{\\rm off\}\\mathcal\{L\}\_\{\\rm offblock\}, which probes whether a soft transition grouping is sufficient to induce useful support structure\. The prediction and latent\-linearity terms make the representation useful for Koopman rollout, the reconstruction term keeps the code tied to the observed state, and the sparsity term encourages selective activation\.
### 3\.4Support objects
In this subsection, we introduce several notions of support that are used at evaluation time\.
Letz=Eθ\(x\)∈ℝdzz=E\_\{\\theta\}\(x\)\\in\\mathbb\{R\}^\{d\_\{z\}\}\. Asupport rulefirst converts the latent codezzinto a Boolean active\-coordinate vectorm\(x\)=\{\|zi\|\>0\}∈\{0,1\}dzm\(x\)=\\\{\|z\_\{i\}\|\>0\\\}\\in\\\{0,1\\\}^\{d\_\{z\}\}, wheremi\(x\)=1m\_\{i\}\(x\)=1means that latent coordinateiiis active for statexx\. Equivalently, the exact support is the index setS\(x\)=\{i:mi\(x\)=1\}S\(x\)=\\\{i:m\_\{i\}\(x\)=1\\\}\. Exact supports defined by an arbitrary threshold can be unstable: a small perturbation inxx\-space can lead to a different support inzz, thus a single basin may be covered by several nearby or partially overlapping active\-coordinate vectors\. We therefore distinguish exact supports from*support families*, which group nearby Boolean support vectors and test whether these groups form basin\-scale objects\. Support families in general are more robust and lead to less basin fragmentation\.
We use two support constructions:
- •Absolute\-threshold support\.The absolute support vector is the Boolean mask denotedmabs\(x\)m\_\{\{\\rm abs\}\}\(x\), wheremabs,i\(x\)=𝟏\{\|zi\|\>10−3\}m\_\{\{\\rm abs\},i\}\(x\)=\\mathbf\{1\}\\\{\|z\_\{i\}\|\>10^\{\-3\}\\\}for indicesii\.Sabs\(x\)=\{i:mabs,i\(x\)=1\}S\_\{\\rm abs\}\(x\)=\\\{i:m\_\{\{\\rm abs\},i\}\(x\)=1\\\}is the equivalent active\-index set\.mabs\(x\)m\_\{\{\\rm abs\}\}\(x\)is the input to support families such asFabsF\_\{\\rm abs\}\(defined next\), and to canonical wrong\-support interventions\.
- •Absolute\-threshold support family\.FabsF\_\{\\rm abs\}is a group label for exact supports; every exact supportmabsm\_\{\{\\rm abs\}\}is assigned anFabsF\_\{\\rm abs\}label\.FabsF\_\{\\rm abs\}groups exact supports whose active\-index sets substantially overlap\. We define the overlap score as the Jaccard index, or intersection\-over\-union, between the active coordinates of two Boolean support vectorsmabs\(x\)m\_\{\{\\rm abs\}\}\(x\)\. On an evaluation collection𝒳\\mathcal\{X\}, the algorithm first counts all distinct Boolean vectors and then visits those distinct vectors from most to least frequent, following a leader\-style threshold clustering rule\. Each familyffstores one fixed representative vectorrfr\_\{f\}: the exact support vector that created that family\. For the next distinct vectormm, the algorithm computes the Jaccard index J\(m,rf\)=∑i𝟏\{mi=1andrf,i=1\}∑i𝟏\{mi=1orrf,i=1\},J\(m,r\_\{f\}\)=\\frac\{\\sum\_\{i\}\\mathbf\{1\}\\\{m\_\{i\}=1\\ \\text\{and\}\\ r\_\{f,i\}=1\\\}\}\{\\sum\_\{i\}\\mathbf\{1\}\\\{m\_\{i\}=1\\ \\text\{or\}\\ r\_\{f,i\}=1\\\}\},withJ\(0,0\)=1J\(0,0\)=1when both vectors are all zero\. If the best existing representative has Jaccard similarity at leastτ=0\.5\\tau=0\.5, every occurrence ofmabsm\_\{\\rm abs\}receives that family label\. Otherwise,mabsm\_\{\\rm abs\}starts a new family and becomes the representativerfr\_\{f\}for that family\. Pseudocode is provided in Appendix[A](https://arxiv.org/html/2608.29057#A1)\. We writeFabs\(x\)F\_\{\\rm abs\}\(x\)for the family assigned tomabs\(x\)m\_\{\\rm abs\}\(x\)\.
### 3\.5Periodic re\-encoding
If a support indicates the currently relevant local dynamics, a long rollout should be able to update that support\. We therefore evaluate periodic re\-encoding, following[Fathi et al\. \[10\]](https://arxiv.org/html/2608.29057#bib.bib10)\. Starting fromzk=Eθ\(xk\)z\_\{k\}=E\_\{\\theta\}\(x\_\{k\}\)and a re\-encoding periodmm, the model advances formmKoopman steps, decodes the predicted state, and encodes that prediction again:
z~k\+m=Kmzk,x~k\+m=Dϕ\(z~k\+m\),zk\+m=Eθ\(x~k\+m\)\.\\tilde\{z\}\_\{k\+m\}=K^\{m\}z\_\{k\},\\qquad\\tilde\{x\}\_\{k\+m\}=D\_\{\\phi\}\(\\tilde\{z\}\_\{k\+m\}\),\\qquad z\_\{k\+m\}=E\_\{\\theta\}\(\\tilde\{x\}\_\{k\+m\}\)\.\(10\)The same procedure is repeated over the rollout horizon\. Each rollout uses only the model’s predicted state; it does not use the true future state, a basin label, or an oracle for transitions between basins\. For each dynamical system, we tune a separate re\-encoding periodmmon the training set\.
## 4Experiments
Our experiments aim to evaluate if finite\-dimensional Koopman representations can be learned for multi\-attractor systems implicitly without resorting to side information such as basin labels\. We first discuss the quality of learned Koopman representations from sparse KAEs via a forecasting\-based evaluation\. Next, in the context of sparse inference, we evaluate the importance of the specific active set of coordinates for forecasting within a basin of attraction\. Finally, we evaluate alignment of support familiesFabsF\_\{\\rm abs\}with withheld basin labels\.
#### Benchmark systems\.
The main benchmark for our experiments contains 15 procedurally generated two\-dimensional multibasin systems\. These systems are constructed by specifying potential wells around known attractor coordinates, and superimposing dynamics that encourage transitions to different far\-away regimes; see Appendix[D](https://arxiv.org/html/2608.29057#A4)\. Because we specify the locations of the potential wells, the basin labels are well\-defined and used only for evaluation\. We also use a subset of 10 chaotic systems from the Dysts repository\[[11](https://arxiv.org/html/2608.29057#bib.bib11),[12](https://arxiv.org/html/2608.29057#bib.bib12)\]with the step sizedtdtmultiplied by a factor of 30 as an additional forecasting test\. Analytic expressions and other details are in Appendix[D](https://arxiv.org/html/2608.29057#A4),[E](https://arxiv.org/html/2608.29057#A5)\.
#### Training models\.
We compare six models shown in[Table1](https://arxiv.org/html/2608.29057#S4.T1)by varying the encoder and latent transition structure and studying their effects in this setting\. We test three LISTA sparse encoders \(where LISTA\-BD denotes block diagonal latent transitions and LISTA\-SB denotes soft\-block regularization\), two sparse\-latent MLP applying the sparsity penaltyℒ¯sp\\bar\{\\mathcal\{L\}\}\_\{\\rm sp\}\([8](https://arxiv.org/html/2608.29057#S3.E8)\) with dense\[[10](https://arxiv.org/html/2608.29057#bib.bib10)\]or block\-diagonal \(BD\) transitions, and a dense\-latent MLP control baseline with no sparsity incentive\. All models use a linear decoder with columns having unit norm constraint\.
Most of these architectures are biased toward sparse inference by varying degrees\. Our contribution is in studying how different ways of inducing sparsity can be useful\. For LISTA, encoded states are sparse by construction via thresholding, and rollout latents get sparsity pressure from \([8](https://arxiv.org/html/2608.29057#S3.E8)\)\. For sparse MLPs, the encoded states have induced sparsity through the ReLU activations andℓ1\\ell\_\{1\}regularization\[[38](https://arxiv.org/html/2608.29057#bib.bib38),[24](https://arxiv.org/html/2608.29057#bib.bib24)\]\. The dense MLP baseline has no intended zero\-producing mechanism\. Each model is trained for 15 random seeds on each system before evaluation\. Training details are reported in Appendix[B](https://arxiv.org/html/2608.29057#A2)\.
#### Statistical details\.
Point estimates are interquartile means \(IQM\) over seeds within each system to reduce seed uncertainty\[[2](https://arxiv.org/html/2608.29057#bib.bib2)\], then arithmetic means across systems\. Horizon\-trend lines use the same point estimate as the tables, with seed\-bootstrap bands that summarize uncertainty over seeds for an average system\. Statistical significance is rigorously tested with paired Wilcoxon signed\-rank tests and Holm corrections; details are provided in Appendix[C](https://arxiv.org/html/2608.29057#A3)\.
### 4\.1Forecasting experiments
Table 1:Forecasting performance and compact basin\-support diagnostics\. Forecasting values are MSEs; lower is better\. Each cell is an arithmetic mean across systems after summarizing seeds within each system by IQM\. Superscripts mark Holm\-corrected tests against the dense\-latent MLP baseline \(∗\\ast,p<0\.05p<0\.05\)\.Boldmarks the best entry in each directional column\.\(a\)15\-system multibasin benchmark\.\(b\)10\-system Dystsdt×30dt\{\\times\}30benchmark\.
Figure 2:Forecasting horizon trends; sparse models forecast better than dense models\. Lines show arithmetic means across per\-system seed\-IQM MSEs on a log scale\. Bands show fixed\-system seed\-bootstrap 95% intervals after system\-wise log\-relative normalization, anchored to the displayed mean, so they summarize seed uncertainty within an average system\.#### Sparse encoders learn high\-quality Koopman representations\.
The quality of the representation learned by a KAE is primarily measured by its ability to make accurate, multi\-step forecasts\. Thus, we evaluate the forecasting performance of trained sparse and dense\-latent KAEs on the two benchmarks spanning 25 systems\. Forecasting is evaluated on held\-out trajectories at the stored observation cadence and summarized by mean squared error \(MSE\)\.
The results \([Figures2](https://arxiv.org/html/2608.29057#S4.F2)and[1](https://arxiv.org/html/2608.29057#S4.T1)\) show that sparse models yield a significant reduction of MSE forecasting error at all horizons on both sets of systems\. In particular, at longer horizons, the dense MLP has 17\-27x larger MSE atH1000H1000on the multibasin systems and about 1\.5x larger MSE atH4000H4000on the Dysts systems compared to most of the sparse models \(the LISTA\-SB model outperforms the baseline but is noticeably worse than the rest\)\. On the other hand, there are marginal differences between MSE predictions among the sparse\-latent encoders\. These results suggest that having encoded sparsity is crucial for good performance over all horizon prediction lengths\.
#### Sparse supports are essential for good representations\.
If a support formed by sparse encoders is meant to carry basin\-relevant dynamical information, then changing only the initial active set should damage forecasts, even when the coefficient values are otherwise kept fixed\. We do an ablation study to show whether accurate forecasts within a basin of attraction require the correct coordinates to be active by isolating the selection of sparse coordinates from their values\. We do this by forcing a sparse KAE to use an incorrect active set with specific interventions, and evaluating 20\-step forecasting performance under these settings\. The interventions are the following:
1. 1\.Coordinate dropping:Force the coordinates ofzzwith thek∈\{1,…,10\}k\\in\\\{1,\\dots,10\\\}largest magnitude coefficients to zero before forecasting\.
2. 2\.Random support:the active coefficient values are preserved, but reassigned to randomly chosen inactive latent coordinates\.
\(a\)Dropping the largest active coordinates\.\(b\)Moving active values to inactive coordinates\.
Figure 3:Support\-coordinate interventions on a multibasin system; interventions severely increase forecasting errors compared to the standard rollout\.\(a\)The standard rollout is compared with cumulative dropping of the largest active coordinates of the initial sparse code; shaded bands are bootstrap 95% intervals over100100initial conditions\.\(b\)Active coefficient values are preserved but reassigned to randomly chosen inactive coordinates; the band is the interquartile range over shuffles\.Table 2:Accumulated MSE atH=21H=21for the support interventions; lower is better\.The results show that these interventions have a significant detrimental effect on forecast quality, even on very short horizons\. AtH=21H=21, the standard rollout has mean accumulated MSE 0\.0158, compared to 0\.508/1\.37/2\.08/3\.26/8\.51 for the rollouts after dropping the top 1/2/3/5/10 active coordinates, and random support shuffling is much more destructive as seen in[Figure3](https://arxiv.org/html/2608.29057#S4.F3)\. The trajectory panel \([Figure4\(a\)](https://arxiv.org/html/2608.29057#S4.F4.sf1)\) shows the qualitative effects of the intervention\. So, supports formed by sparse encoders do not merely correlate with basin labels; the active coordinates are key to the quality of the Koopman representation\.
### 4\.2Support\-basin alignment experiments
This mechanistic experiment examines the extent to which sparse support families fit after training can behave like basin\-interior regime variables\. States near separatrices are challenging and may confound the inferences: near these states, support changes can be induced by small perturbations inxx, making supports unreliable for basin identification at separatrices\. We therefore evaluate this diagnostic on held\-out states deep in basins far from basin boundaries\.
For a statexx, letd1\(x\)d\_\{1\}\(x\)andd2\(x\)d\_\{2\}\(x\)be its distances to the nearest and second\-nearest benchmark attractor centers\. Thebasin\-depth marginisd2\(x\)−d1\(x\)d\_\{2\}\(x\)\-d\_\{1\}\(x\)\. This support\-basin alignment experiment takes the top quartile of held\-out states ranked by basin\-depth margin within each benchmark basin, using ground\-truth benchmark geometry to identify each basin\. If a support familyFabsF\_\{\\rm abs\}identifies a basin\-specific regime used by the model, then on states deep inside a basin, it should leave little uncertainty about the withheld benchmark basin labelBB, measured by conditional entropyH\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\)as the main evaluation metric\.H\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\)should be small when support families are basin\-aligned\. Also, different basins should be adequately represented by distinct support families, so we count the number of distinct support families formed in a system, denoted as\|Fabs\|\|F\_\{\\rm abs\}\|, and compare it with the number of basins represented\.
#### Support familiesFabsF\_\{\\rm abs\}identify basin membership\.
As seen in[Table1](https://arxiv.org/html/2608.29057#S4.T1)and[Figure4\(b\)](https://arxiv.org/html/2608.29057#S4.F4.sf2), LISTAFabsF\_\{\\rm abs\}support families provide more information to identify a ground truth basin label than any other model architecture, as seen by the lowest conditional entropy of 0\.130\. Other LISTA models with structured latent transitions are close competitors, followed by the sparse\-latent MLPs, and then finally the dense\-latent MLP baseline\. LISTA\-based models also expose the most support families, which better matches the benchmark’s basin\-count mean/median 4\.20/4; ideally, if a basin can truly be recovered by a single support family cluster, then the basin count should be one\-to\-one with the number of support families,\|Fabs\|¯\\overline\{\|F\_\{\\rm abs\}\|\}\. In contrast, the dense\-latent MLP collapses the representation to one support family, so distinct basins have correlated learned latent dynamics\. This suggests that the support families induced by sparse coding may serve as a good proxy for ground\-truth basin labels; the sparse LISTA models may have the emergent ability to discover basins unsupervised\.
\(a\)Randomly shuffled support trajectories \(red\) compared to standard rollouts \(black\)\. Randomly shuffling the active coordinates leads to divergences\.\(b\)Distribution of conditional entropy results; Dense MLP cannot identify basins based on support families\. Dots show system–seed pairs, gray I\-bars mark first\-to\-third quartile ranges, and black bars mark IQM summaries forH\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\)\.
## 5Discussion
We show that sparse Koopman autoencoders can learn useful local regime structure in multibasin dynamical systems without observing basin labels\. Existing theory and empirical evidence indicate that multibasin systems generally should not be forced through one continuous finite\-dimensional global Koopman embedding\[[23](https://arxiv.org/html/2608.29057#bib.bib23),[5](https://arxiv.org/html/2608.29057#bib.bib5),[32](https://arxiv.org/html/2608.29057#bib.bib32),[10](https://arxiv.org/html/2608.29057#bib.bib10),[25](https://arxiv.org/html/2608.29057#bib.bib25)\]\. Our results show that sparse codes offer a practical alternative: a single trained modelcanfind good representations for multiple local Koopman regimes by using different latent supports\. This gives practitioners a simple diagnostic to inspect after training\.
We qualify some limitations\. First, our interpretability claims are evaluated on benchmark systems where basin geometry is known, and the support–basin alignment analysis is intentionally restricted to states deep inside basins\. Near separatrices, small state changes can imply very different long\-term outcomes, and support assignments become fragmented or unstable\. This is not a failure mode unique to sparse Koopman models, but it means that support families should be interpreted as reliable basin\-interior regime variables rather than perfect global basin classifiers\. Second, while the experiments include a diverse set of synthetic multibasin systems and external chaotic flows, they remain deterministic benchmarks\. Real\-world systems may contain observation noise, partial observability, nonstationarity, control inputs, or regime changes that are not cleanly basin\-like\.
More broadly, this work argues that interpretability in dynamical representation learning can come from the structure of the latent code\. Sparse supports provide an interface between continuous Koopman coordinates and discrete regime structure, making them a promising tool for multi\-regime forecasting, analysis, and eventually control\. By showing that these supports improve prediction, are necessary for accurate rollouts, and align with withheld basin information, this paper positions sparse latent structure as a practical step toward Koopman models for multibasin dynamics\.
## Acknowledgments and Disclosure of Funding
A\.L\. acknowledges support from the Fonds de recherche du Québec \(FRQ\) through a Master’s Research Scholarship \(DOI:[https://doi\.org/10\.69777/2008193](https://doi.org/10.69777/2008193)\)\. S\.C\. is supported by the Canada CIFAR AI Chairs program, the Canada Research Chair in Lifelong Machine Learning, and the NSERC Discovery Grant\. This research was enabled in part by compute resources, software and technical help provided by Mila \([mila\.quebec](https://mila.quebec/)\)\.
## References
- \[1\]Aviad Aberdam, Alona Golts, and Michael Elad\.Ada\-LISTA: Learned Solvers Adaptive to Varying Models\.*IEEE transactions on pattern analysis and machine intelligence*, 44\(12\):9222–9235, December 2022\.ISSN 1939\-3539\.doi:10\.1109/TPAMI\.2021\.3125041\.
- \[2\]Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare\.Deep reinforcement learning at the edge of the statistical precipice\.*Advances in Neural Information Processing Systems*, 34, 2021\.
- \[3\]Omri Azencot, N\. Benjamin Erichson, Vanessa Lin, and Michael Mahoney\.Forecasting Sequential Data Using Consistent Koopman Autoencoders\.In*Proceedings of the 37th International Conference on Machine Learning*, pages 475–485\. PMLR, November 2020\.
- \[4\]Amir Beck and Marc Teboulle\.A fast iterative shrinkage\-thresholding algorithm for linear inverse problems\.*SIAM journal on imaging sciences*, 2\(1\):183–202, 2009\.
- \[5\]Steven L\. Brunton, Bingni W\. Brunton, Joshua L\. Proctor, and J\. Nathan Kutz\.Koopman Invariant Subspaces and Finite Linear Representations of Nonlinear Dynamical Systems for Control\.*PLOS ONE*, 11\(2\), February 2016\.ISSN 1932\-6203\.doi:10\.1371/journal\.pone\.0150171\.URL[https://journals\.plos\.org/plosone/article?id=10\.1371/journal\.pone\.0150171](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0150171)\.
- \[6\]Steven L\. Brunton, Marko Budišić, Eurika Kaiser, and J\. Nathan Kutz\.Modern Koopman Theory for Dynamical Systems\.*SIAM Review*, 64\(2\):229–340, May 2022\.ISSN 0036\-1445\.doi:10\.1137/21M1401243\.URL[https://doi\.org/10\.1137/21M1401243](https://doi.org/10.1137/21M1401243)\.
- \[7\]Scott Shaobing Chen, David L\. Donoho, and Michael A\. Saunders\.Atomic Decomposition by Basis Pursuit\.*SIAM Journal on Scientific Computing*, 20\(1\):33–61, 1998\.doi:10\.1137/S1064827596304010\.URL[https://doi\.org/10\.1137/S1064827596304010](https://doi.org/10.1137/S1064827596304010)\.
- \[8\]Xiaohan Chen, Jialin Liu, Zhangyang Wang, and Wotao Yin\.Hyperparameter Tuning is All You Need for LISTA\.In*Advances in Neural Information Processing Systems*, volume 34, pages 11678–11689\. Curran Associates, Inc\., 2021\.
- \[9\]David L\. Donoho and Michael Elad\.Optimally sparse representation in general \(nonorthogonal\) dictionaries viaℓ1\\ell\_\{1\}minimization\.*Proceedings of the National Academy of Sciences*, 100\(5\):2197–2202, March 2003\.doi:10\.1073/pnas\.0437847100\.URL[https://doi\.org/10\.1073/pnas\.0437847100](https://doi.org/10.1073/pnas.0437847100)\.
- \[10\]Mahan Fathi, Clement Gehring, Jonathan Pilault, David Kanaa, Pierre\-Luc Bacon, and Ross Goroshin\.Course correcting koopman representations\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=A18gWgc5mi](https://openreview.net/forum?id=A18gWgc5mi)\.
- \[11\]William Gilpin\.Chaos as an interpretable benchmark for forecasting and data\-driven modelling\.In*Advances in Neural Information Processing Systems \(NeurIPS\) 2021*, 2021\.URL[http://arxiv\.org/abs/2110\.05266](http://arxiv.org/abs/2110.05266)\.
- \[12\]William Gilpin\.Model scale versus domain knowledge in statistical forecasting of chaotic systems\.*Physical Review Research*, 5\(4\):043252, December 2023\.doi:10\.1103/PhysRevResearch\.5\.043252\.URL[https://link\.aps\.org/doi/10\.1103/PhysRevResearch\.5\.043252](https://link.aps.org/doi/10.1103/PhysRevResearch.5.043252)\.
- \[13\]Karol Gregor and Yann LeCun\.Learning Fast Approximations of Sparse Coding\.In*Proceedings of the 27th International Conference on Machine Learning*, ICML’10, pages 399–406, Haifa, Israel, 2010\. Omnipress\.ISBN 978\-1\-60558\-907\-7\.doi:10\.5555/3104322\.3104374\.URL[https://dl\.acm\.org/doi/abs/10\.5555/3104322\.3104374](https://dl.acm.org/doi/abs/10.5555/3104322.3104374)\.
- \[14\]D\. M\. Grobman\.Homeomorphism of systems of differential equations\.*Doklady Akademii Nauk SSSR*, 128:880–881, 1959\.
- \[15\]Lars Grüne and Jürgen Pannek\.*Nonlinear Model Predictive Control*\.Communications and Control Engineering\. Springer International Publishing, Cham, 2017\.ISBN 978\-3\-319\-46023\-9 978\-3\-319\-46024\-6\.doi:10\.1007/978\-3\-319\-46024\-6\.URL[http://link\.springer\.com/10\.1007/978\-3\-319\-46024\-6](http://link.springer.com/10.1007/978-3-319-46024-6)\.
- \[16\]Philip Hartman\.A lemma in the theory of structural stability of differential equations\.*Proceedings of the American Mathematical Society*, 11\(4\):610–620, 1960\.doi:10\.1090/S0002\-9939\-1960\-0121542\-7\.
- \[17\]R\.E\. Kalman\.On the general theory of control systems\.*1st International IFAC Congress on Automatic and Remote Control, Moscow, USSR, 1960*, 1\(1\):491–502, August 1960\.ISSN 1474\-6670\.doi:10\.1016/S1474\-6670\(17\)70094\-8\.URL[https://www\.sciencedirect\.com/science/article/pii/S1474667017700948](https://www.sciencedirect.com/science/article/pii/S1474667017700948)\.
- \[18\]B\. O\. Koopman\.Hamiltonian Systems and Transformation in Hilbert Space\.*Proceedings of the National Academy of Sciences*, 17\(5\):315–318, May 1931\.doi:10\.1073/pnas\.17\.5\.315\.URL[https://www\.pnas\.org/doi/10\.1073/pnas\.17\.5\.315](https://www.pnas.org/doi/10.1073/pnas.17.5.315)\.
- \[19\]B\. O\. Koopman and J\. v\. Neumann\.Dynamical Systems of Continuous Spectra\.*Proceedings of the National Academy of Sciences*, 18\(3\):255–263, March 1932\.doi:10\.1073/pnas\.18\.3\.255\.URL[https://www\.pnas\.org/doi/abs/10\.1073/pnas\.18\.3\.255](https://www.pnas.org/doi/abs/10.1073/pnas.18.3.255)\.
- \[20\]Milan Korda and Igor Mezić\.Linear predictors for nonlinear dynamical systems: Koopman operator meets model predictive control\.*Automatica*, 93:149–160, July 2018\.ISSN 00051098\.doi:10\.1016/j\.automatica\.2018\.03\.046\.URL[http://arxiv\.org/abs/1611\.03537](http://arxiv.org/abs/1611.03537)\.
- \[21\]Yilun Kuang, Yash Dagade, Quentin Le Lidec, Lucas Maes, Randall Balestriero, and Yann LeCun\.LpWM: A Case for Sparse Representations in World Models, August 2026\.URL[http://arxiv\.org/abs/2608\.22764](http://arxiv.org/abs/2608.22764)\.
- \[22\]Matthew D\. Kvalheim and Shai Revzen\.Existence and uniqueness of global Koopman eigenfunctions for stable fixed points and periodic orbits\.*Physica D: Nonlinear Phenomena*, 425:132959, November 2021\.ISSN 0167\-2789\.doi:10\.1016/j\.physd\.2021\.132959\.URL[https://www\.sciencedirect\.com/science/article/pii/S0167278921001160](https://www.sciencedirect.com/science/article/pii/S0167278921001160)\.
- \[23\]Yueheng Lan and Igor Mezić\.Linearization in the large of nonlinear systems and Koopman operator spectrum\.*Physica D: Nonlinear Phenomena*, 242\(1\):42–53, January 2013\.ISSN 0167\-2789\.doi:10\.1016/j\.physd\.2012\.08\.017\.
- \[24\]Oliver W\. Layton, Siyuan Peng, and Scott T\. Steinmetz\.Relu, sparseness, and the encoding of optic flow in neural networks\.*Sensors*, 24\(23\), 2024\.ISSN 1424\-8220\.doi:10\.3390/s24237453\.
- \[25\]Zexiang Liu, Necmiye Ozay, and Eduardo D\. Sontag\.Properties of immersions for systems with multiple limit sets with implications to learning Koopman embeddings\.*Automatica*, 176:112226, June 2025\.ISSN 0005\-1098\.doi:10\.1016/j\.automatica\.2025\.112226\.URL[https://www\.sciencedirect\.com/science/article/pii/S0005109825001189](https://www.sciencedirect.com/science/article/pii/S0005109825001189)\.
- \[26\]Bethany Lusch, J\. Nathan Kutz, and Steven L\. Brunton\.Deep learning for universal linear embeddings of nonlinear dynamics\.*Nature Communications*, 9\(1\):4950, November 2018\.ISSN 2041\-1723\.doi:10\.1038/s41467\-018\-07210\-0\.
- \[27\]Giorgos Mamakoukas, Maria Castano, Xiaobo Tan, and Todd Murphey\.Local Koopman Operators for Data\-Driven Control of Robotic Systems\.In*Robotics: Science and Systems XV*\. Robotics: Science and Systems Foundation, June 2019\.ISBN 978\-0\-9923747\-5\-4\.doi:10\.15607/RSS\.2019\.XV\.054\.URL[http://www\.roboticsproceedings\.org/rss15/p54\.pdf](http://www.roboticsproceedings.org/rss15/p54.pdf)\.
- \[28\]D\. Q\. Mayne, J\. B\. Rawlings, C\. V\. Rao, and P\. O\. M\. Scokaert\.Constrained model predictive control: Stability and optimality\.*Automatica*, 36\(6\):789–814, June 2000\.ISSN 0005\-1098\.doi:10\.1016/S0005\-1098\(99\)00214\-9\.URL[https://www\.sciencedirect\.com/science/article/pii/S0005109899002149](https://www.sciencedirect.com/science/article/pii/S0005109899002149)\.
- \[29\]Igor Mezić\.Spectral properties of dynamical systems, model reduction and decompositions\.*Nonlinear Dynamics*, 41\(1–3\):309–325, 2005\.doi:10\.1007/s11071\-005\-2824\-x\.
- \[30\]Bruno A\. Olshausen and David J\. Field\.Emergence of simple\-cell receptive field properties by learning a sparse code for natural images\.*Nature*, 381\(6583\):607–609, June 1996\.ISSN 1476\-4687\.doi:10\.1038/381607a0\.URL[https://doi\.org/10\.1038/381607a0](https://doi.org/10.1038/381607a0)\.
- \[31\]Samuel E\. Otto and Clarence W\. Rowley\.Linearly recurrent autoencoder networks for learning dynamics\.*SIAM Journal on Applied Dynamical Systems*, 18\(1\):558–593, 2019\.doi:10\.1137/18M1177846\.
- \[32\]Shaowu Pan and Karthik Duraisamy\.On the lifting and reconstruction of nonlinear systems with multiple invariant sets\.*Nonlinear Dynamics*, 112\(12\):10157–10165, June 2024\.ISSN 1573\-269X\.doi:10\.1007/s11071\-024\-09581\-0\.URL[https://doi\.org/10\.1007/s11071\-024\-09581\-0](https://doi.org/10.1007/s11071-024-09581-0)\.
- \[33\]Peter J\. Schmid\.Dynamic mode decomposition of numerical and experimental data\.*Journal of Fluid Mechanics*, 656:5–28, 2010\.ISSN 0022\-1120\.doi:10\.1017/S0022112010001217\.
- \[34\]Haojie Shi and Max Q\.\-H\. Meng\.Deep Koopman Operator With Control for Nonlinear Systems\.*IEEE Robotics and Automation Letters*, 7\(3\):7700–7707, July 2022\.ISSN 2377\-3766\.doi:10\.1109/LRA\.2022\.3184036\.
- \[35\]Naoya Takeishi, Yoshinobu Kawahara, and Takehisa Yairi\.Learning koopman invariant subspaces for dynamic mode decomposition\.In*Advances in Neural Information Processing Systems 30*, volume 30, 2017\.URL[https://papers\.nips\.cc/paper/6713\-learning\-koopman\-invariant\-subspaces\-for\-dynamic\-mode\-decomposition](https://papers.nips.cc/paper/6713-learning-koopman-invariant-subspaces-for-dynamic-mode-decomposition)\.
- \[36\]Matthew O\. Williams, Ioannis G\. Kevrekidis, and Clarence W\. Rowley\.A Data–Driven Approximation of the Koopman Operator: Extending Dynamic Mode Decomposition\.*Journal of Nonlinear Science*, 25\(6\):1307–1346, December 2015\.ISSN 1432\-1467\.doi:10\.1007/s00332\-015\-9258\-5\.URL[https://doi\.org/10\.1007/s00332\-015\-9258\-5](https://doi.org/10.1007/s00332-015-9258-5)\.
- \[37\]Matthew O\. Williams, Clarence W\. Rowley, and Ioannis G\. Kevrekidis\.A kernel\-based method for data\-driven koopman spectral analysis\.*Journal of Computational Dynamics*, 2\(2\):247–265, May 2016\.ISSN 2158\-2491\.doi:10\.3934/jcd\.2015005\.URL[https://www\.aimsciences\.org/article/doi/10\.3934/jcd\.2015005](https://www.aimsciences.org/article/doi/10.3934/jcd.2015005)\.
- \[38\]Yuchen Zhang, Jason D\. Lee, and Michael I\. Jordan\.L1\-regularized neural networks are improperly learnable in polynomial time\.In Maria Florina Balcan and Kilian Q\. Weinberger, editors,*Proceedings of The 33rd International Conference on Machine Learning*, volume 48 of*Proceedings of Machine Learning Research*, pages 993–1001, New York, New York, USA, 20–22 Jun 2016\. PMLR\.URL[https://proceedings\.mlr\.press/v48/zhangd16\.html](https://proceedings.mlr.press/v48/zhangd16.html)\.
## Appendix ASupport definitions and sensitivity
[Section3\.4](https://arxiv.org/html/2608.29057#S3.SS4)defines the notation used in the main text\. This appendix records the implementation\-level objects behind that notation and explains how to read the corresponding diagnostics\. All support objects are computed after training from the learned encoder outputz=Eθ\(x\)z=E\_\{\\theta\}\(x\)\. Basin labels, basin counts, and trajectory\-to\-basin assignments are not used to train the model, define a support, or build a support family; when labels appear below they are evaluation\-only quantities used to score or display the learned objects\.
#### From a latent vector to an exact support key\.
The reducers first convert each latent vector into a Boolean maskm\(x\)∈\{0,1\}dzm\(x\)\\in\\\{0,1\\\}^\{d\_\{z\}\}\. The exact supportS\(x\)S\(x\)is the index set of the true entries of this mask, but the code compares exact supports by serializing the Boolean mask itself: two states have the same exact support if and only if alldzd\_\{z\}Boolean entries match\. The main support rule used in the paper is
mabs,i\(x\)\\displaystyle m\_\{\{\\rm abs\},i\}\(x\)=𝟏\{\|zi\|\>10−3\},\\displaystyle=\\mathbf\{1\}\\\{\|z\_\{i\}\|\>10^\{\-3\}\\\},Sabs\(x\)\\displaystyle S\_\{\\rm abs\}\(x\)=\{i:mabs,i\(x\)=1\},\\displaystyle=\\\{i:m\_\{\{\\rm abs\},i\}\(x\)=1\\\},\(11\)The inequalities are strict\.SabsS\_\{\\rm abs\}is a magnitude\-thresholded active set, so\|Sabs\|\|S\_\{\\rm abs\}\|is a measured state\-dependent statistic rather than a fixed sparsity budget\.
#### From exact supports to support families\.
For a chosen support rule and evaluation collection, the implementation counts distinct exact support keys, sorts them from most to least frequent, and uses the serialized key as a deterministic tie\-breaker\. It then performs the leader\-style greedy Jaccard merge described in[Section3\.4](https://arxiv.org/html/2608.29057#S3.SS4)\. Each family stores an actual Boolean support vector as its representative\. A new exact mask joins the existing representative with the largest Jaccard similarity if that similarity is at least0\.50\.5; otherwise it starts a new family\. Representatives are not optimized, averaged, or updated as centroids\. Family labels are therefore run\- and collection\-specific integer identifiers, and only their induced partition of states is meaningful\.
#### Support\-family creation algorithm\.
For completeness, the family construction used for bothFabsF\_\{\\rm abs\}andFtop8F\_\{\\rm top8\}is:
> Input:evaluation states𝒳\\mathcal\{X\}, a support ruleqq, and Jaccard thresholdτ=0\.5\\tau=0\.5\. Output:family labelFq\(x\)F\_\{q\}\(x\)for eachx∈𝒳x\\in\\mathcal\{X\}, and family representatives\{rf\}\\\{r\_\{f\}\\\}\. 1. 1\.Compute one exact Boolean maskmq\(x\)∈\{0,1\}dzm\_\{q\}\(x\)\\in\\\{0,1\\\}^\{d\_\{z\}\}for everyx∈𝒳x\\in\\mathcal\{X\}\. 2. 2\.Count the multiplicityc\(m\)c\(m\)of each distinct exact maskmm\. 3. 3\.Sort the distinct masks in decreasingc\(m\)c\(m\), using the serialized Boolean mask as a deterministic tie\-breaker\. 4. 4\.Initialize an empty list of families and representatives\. 5. 5\.For each distinct maskmmin the sorted order: 1. a\.If no family exists, create a new familyff, setrf←mr\_\{f\}\\leftarrow m, and assignmmtoff\. 2. b\.Otherwise computef⋆=argmaxfJ\(m,rf\)f^\{\\star\}=\\operatorname\*\{arg\\,max\}\_\{f\}J\(m,r\_\{f\}\)over existing family representatives\. 3. c\.IfJ\(m,rf⋆\)≥τJ\(m,r\_\{f^\{\\star\}\}\)\\geq\\tau, assign every occurrence ofmmto familyf⋆f^\{\\star\}\. 4. d\.IfJ\(m,rf⋆\)<τJ\(m,r\_\{f^\{\\star\}\}\)<\\tau, create a new familyff, setrf←mr\_\{f\}\\leftarrow m, and assign every occurrence ofmmtoff\. 6. 6\.For each statexx, returnFq\(x\)F\_\{q\}\(x\), the family assigned to its exact maskmq\(x\)m\_\{q\}\(x\)\.
The representative updaterf←mr\_\{f\}\\leftarrow moccurs only when a new family is created\. Existing representatives are never averaged, recomputed as majority masks, or moved toward later assigned masks\.
For top\-eight masks,J\(m,r\)≥0\.5J\(m,r\)\\geq 0\.5is equivalent to at least six shared active coordinates because both masks have size eight\. ForSabsS\_\{\\rm abs\}, the same threshold is adaptive to mask size through the intersection\-over\-union denominator\. A lower Jaccard threshold merges more exact supports and can collapse distinct basins; a higher threshold separates more masks and can makeH\(B∣F\)H\(B\\mid F\)small by over\-fragmenting each basin\. We therefore fix0\.50\.5rather than tuning it post hoc and interpret any entropy result together with the corresponding family count\.
#### Which support object is used where\.
SabsS\_\{\\rm abs\}\.This is the strict active\-coordinate object\. It is used to buildFabsF\_\{\\rm abs\}, to report exact\-support diagnostics such asH\(B∣Sabs\)H\(B\\mid S\_\{\\rm abs\}\),H\(Sabs∣B\)H\(S\_\{\\rm abs\}\\mid B\), exact\-support uniqueness, and\|Sabs\|\|S\_\{\\rm abs\}\|, and to form canonical masks for wrong\-support interventions\. In the intervention code, the canonical mask for basinbbis the most frequentSabsS\_\{\\rm abs\}key among candidate deep\-slice states from basinbb\. The “wrong” condition replaces it with canonical masks from other represented basins and holds that mask fixed during the masked latent rollout\.
FabsF\_\{\\rm abs\}\.This is the main basin\-support alignment object in[Table1](https://arxiv.org/html/2608.29057#S4.T1)\. It asks whether the many exactSabsS\_\{\\rm abs\}masks produced by a sparse encoder compress into basin\-scale families\. The primary directional metric isH\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\): low values mean that knowing the support family leaves little uncertainty about the withheld basin label\. The count\|Fabs\|¯\\overline\{\|F\_\{\\rm abs\}\|\}is deliberately non\-directional and should be compared with the number of basin labels represented in the evaluated slice\.
#### How to read the diagnostics\.
A lowH\(B∣Sabs\)H\(B\\mid S\_\{\\rm abs\}\)orH\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\)means that the support object predicts basin identity on the evaluated benchmark states\. It does not by itself imply a compact one\-support\-per\-basin code: a model can achieve lowH\(B∣Sabs\)H\(B\\mid S\_\{\\rm abs\}\)while assigning many different exact masks to the same basin\.H\(Sabs∣B\)H\(S\_\{\\rm abs\}\\mid B\), exact\-support uniqueness, and\|Sabs\|\|S\_\{\\rm abs\}\|diagnose this fragmentation, whileH\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\)and\|Fabs\|¯\\overline\{\|F\_\{\\rm abs\}\|\}ask whether the fragmented exact masks merge into a basin\-scale family structure\.
The takeaways are as follows\.FabsF\_\{\\rm abs\}is the object for the claim that sparse supports identify basin interiors\.SabsS\_\{\\rm abs\}is the intervention object for testing whether the active coordinates are functionally used\.
## Appendix BAdditional experimental details
This appendix records the paper\-facing experimental protocol\. The main experiments ask whether sparse Koopman autoencoders learn useful latent supports without basin supervision\. The model is never given basin labels, basin counts, or trajectory\-to\-basin assignments when training the encoder, decoder, latent transition, support rules, or periodic rollouts\. Benchmark basin metadata is used only after training to define evaluation slices, construct controlled diagnostic interventions, and score support–basin agreement\. The only training\-time exception is architectural: the block\-diagonal and soft\-block transition diagnostics use benchmark metadata to set the fixed number of transition groups, but they still do not receive trajectory labels or support labels\.
#### Benchmarks\.
The controlled benchmark contains1515two\-dimensional multibasin systems, listed with their equations and attractor metadata in[AppendixD](https://arxiv.org/html/2608.29057#A4)\. These systems are integrated to produce stored observation sequences; the learner observes only stored states, not vector fields, integration substeps, basin labels, or basin assignments\. The external forecasting stress test uses the1010three\-dimensional Dystsdt×30dt\{\\times\}30systems listed in[AppendixE](https://arxiv.org/html/2608.29057#A5)\. For Dysts, stored observations are generated at3030times each system’s native integration interval and coordinates are standardized after trajectory generation\.
#### Model rows\.
The six model rows in[Table1](https://arxiv.org/html/2608.29057#S4.T1)are LISTA, LISTA–BD, LISTA–SB, Sparse MLP, Sparse MLP–BD, and Dense MLP\. All use latent dimensiondz=256d\_\{z\}=256\. LISTA rows use shrinkage to produce exact zeros; MLP sparse rows use the sparsity penalty in[eq\.8](https://arxiv.org/html/2608.29057#S3.E8); Dense MLP removes the intended sparsity mechanism and sets the sparsity coefficient to zero\. The BD rows use a block\-diagonal latent transition\. The SB row keeps a dense transition but adds an off\-block penalty\. All rows use a linear decoder dictionary as the state decoder\. The controlled multibasin LISTA rows use threshold scaleα=0\.15\\alpha=0\.15, two shrinkage\-refinement loops, and a sign\-split final operation\. The Dystsdt×30dt\{\\times\}30LISTA rows use the same threshold scale but with one refinement loop and a ReLU final operation\. MLP encoders use two hidden layers of width6464; Dense MLP uses a tanh encoder without the sparse output gate\. LISTA uses MLPs of the same specification \(depth and width\) to produce the precode for refinement\.
#### Controlled multibasin training\.
For every controlled system, each model row is trained for seeds0,…,140,\\ldots,14\. Training uses rollout lengthL=8L=8windows, meaning99adjacent stored states per window, for200,000200\{,\}000optimization steps with minibatches of256256windows\. The optimizer is AdamW\. Encoder and decoder parameters use learning rate5×10−55\{\\times\}10^\{\-5\}, the latent transition uses learning rate5×10−65\{\\times\}10^\{\-6\}, and non\-transition parameters use weight decay10−410^\{\-4\}\. In[eq\.9](https://arxiv.org/html/2608.29057#S3.E9), the loss weights areλpred=1\\lambda\_\{\\rm pred\}=1,λrec=0\.03\\lambda\_\{\\rm rec\}=0\.03,λlin=1\\lambda\_\{\\rm lin\}=1, andλsp=3×10−3\\lambda\_\{\\rm sp\}=3\{\\times\}10^\{\-3\}for sparse rows; Dense MLP usesλsp=0\\lambda\_\{\\rm sp\}=0\. Soft\-block rows add the off\-blockL1L\_\{1\}transition penalty with weight10−410^\{\-4\}\.
All controlled rows use the same label\-free boundary\-emphasized reset distribution\. Before training,40964096candidate initial states are probed for3232steps\. Candidates are scored by short\-horizon perturbation sensitivity and residual late\-rollout motion, and the top10241024states are retained\. The scoring probe uses four perturbations, perturbation scale0\.040\.04, a late\-transient window of88steps, and transient weight0\.50\.5\. Half of training windows start from this retained pool with jitter scale0\.250\.25\. The retained pool is intended to expose boundary\-adjacent and transient regions that would otherwise be rare under ordinary resets; it is built without basin labels\.
#### Dystsdt×30dt\{\\times\}30training\.
Dysts rows are trained for the same seeds,0,…,140,\\ldots,14, with latent dimensiondz=256d\_\{z\}=256, minibatches of256256, rollout lengthL=10L=10windows, and100,000100\{,\}000optimization steps\. Dysts trajectories come from deterministic native caches with200200cached trajectories,30,00030\{,\}000stored observations per trajectory, and a2,0002\{,\}000\-step warm\-up\. The first cached trajectory starts from the Dysts default initial condition; remaining trajectories perturb that state with Gaussian noise at scale0\.20\.2times the coordinate\-wise Dysts standard deviation\. Train, validation, and test caches use separate deterministic namespaces, so held\-out test windows are disjoint from training windows and reusable across model seeds\. LISTA rows use the controlled\-benchmark learning rates\. MLP rows use learning rates10−410^\{\-4\}for encoder/decoder parameters and10−510^\{\-5\}for the latent transition\. Sparse Dysts rows useλsp=6×10−3\\lambda\_\{\\rm sp\}=6\{\\times\}10^\{\-3\}, while Dense MLP usesλsp=0\\lambda\_\{\\rm sp\}=0\.
#### Checkpoint selection\.
Checkpoint selection is fixed before test evaluation\. Every500500optimization steps, and again at the final step, the current model is rolled out for200200stored steps from1616fixed held\-out initial states using every\-step re\-encoding\. The checkpoint with the lowest final validation error under this proxy is used for reported evaluations\. Final tables aggregate completed seeds; they do not select the best seed\.
#### Training recipe selection\.
Exploratory sweeps varied LISTA depth, shrinkage scale, sparsity coefficient, sign\-split encodings, transition structure, and MLP sparsity controls\. These sweeps used shorter budgets or broader system lists and are not pooled with the final evidence\. The paper\-facing protocol is the fixed recipe in this appendix:1515seeds per retained system, the training budgets above, held\-out evaluation splits, and the statistical units in[AppendixC](https://arxiv.org/html/2608.29057#A3)\.
#### Forecasting experiments\.
Forecasting is evaluated on held\-out trajectories at the stored observation cadence\. A horizonhhalways meanshhstored observation steps, not a common physical time across systems\. The no\-reencoding rollout applies the learned latent transition for the full horizon before decoding\. The periodic rollout advances the model’s own predicted latent state, decodes the predicted state at a fixed periodmm, and re\-encodes that predicted state; it never uses the true future state\. Controlled multibasin forecasting uses100100held\-out initial conditions for each system and seed and reports the best score over the fixed period gridm∈\{10,25,50,100\}m\\in\\\{10,25,50,100\\\}\. Dysts long\-horizon forecasting uses100100held\-out test\-cache windows for each system and seed and reports the best score overm∈\{10,25,50,100,150,200\}m\\in\\\{10,25,50,100,150,200\\\}\. These grids are fixed before aggregation and are not tuned separately for each test trajectory\.
#### Support–basin alignment\.
Support alignment is computed after training from held\-out controlled\-system states\. For each statexx, the basin\-depth margin is the distance to the second\-nearest benchmark attractor center minus the distance to the nearest attractor center\. Within each represented benchmark basin, the evaluator keeps the top quartile of held\-out states by this margin\. This per\-basin deep slice uses ground\-truth geometry only to define the evaluation population\. It is the clean slice for testing basin\-support alignment because boundary states can legitimately have unstable or mixed supports\.
On this slice, the evaluator encodes states and constructsSabs\(x\)=\{i:\|zi\|\>10−3\}S\_\{\\rm abs\}\(x\)=\\\{i:\|z\_\{i\}\|\>10^\{\-3\}\\\}\. Exact Boolean masks are grouped intoFabsF\_\{\\rm abs\}support families by the deterministic greedy Jaccard rule in[Sections3\.4](https://arxiv.org/html/2608.29057#S3.SS4)and[A](https://arxiv.org/html/2608.29057#A1)with threshold0\.50\.5\. The main directional metric isH\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\), whereBBis the withheld benchmark basin label; lower values mean that knowing the learned support family leaves less uncertainty about basin identity\. The companion count\|Fabs\|¯\\overline\{\|F\_\{\\rm abs\}\|\}is descriptive and is interpreted relative to the number of represented basins, not as a monotone score to maximize\.
#### Support\-coordinate intervention figures\.
The representative coordinate\-intervention experiment isolates active\-coordinate identity from coefficient values\. Starting from a held\-out basin\-interior state, the model encodesz0=Eθ\(x0\)z\_\{0\}=E\_\{\\theta\}\(x\_\{0\}\), perturbs onlyz0z\_\{0\}, and then runs the ordinary learned transition and decoder without retraining\. In the coordinate\-dropping condition, the active entries are ranked by\|z0,i\|\|z\_\{0,i\}\|and the topk∈\{1,2,3,5,10\}k\\in\\\{1,2,3,5,10\\\}are set to zero before rollout\. In the random\-support condition, active coefficient values are preserved but reassigned to randomly chosen inactive coordinates\. The figure and table use100100held\-out basin\-interior starts; the random\-support condition uses2020random inactive\-coordinate reassignments for each start\. Errors are accumulated throughH=21H=21, with the full horizon curves shown in[Figure3](https://arxiv.org/html/2608.29057#S4.F3)and theH=21H=21summary in[Table2](https://arxiv.org/html/2608.29057#S4.T2)\.
#### Aggregation and significance\.
Point estimates in forecasting and support tables summarize seeds within each system by interquartile mean and then average system summaries across systems\. Multibasin forecasting, support\-alignment, and wrong\-support significance annotations use paired seed\-level effects within each controlled system against Dense MLP, with one\-sided alternatives fixed by the table direction and Holm correction across eligible systems\. Dysts forecasting uses each Dysts system as the independent unit after within\-system seed\-IQM summarization and applies a paired one\-sided exact sign test with Holm correction\. Support\-refresh and coordinate\-intervention panels are descriptive mechanism diagnostics\. Full statistical definitions and denominator conventions are in[AppendixC](https://arxiv.org/html/2608.29057#A3)\.
#### Hardware and resource details\.
Training was done on a SLURM cluster; RTX8000 was the main GPU being used\. Training jobs load CUDA12\.6\.012\.6\.0and request one cluster GPU,44CPU cores, and1616GB host memory per active training allocation\. Controlled multibasin training uses single\-run GPU allocations with a2424\-hour time limit\. Dystsdt×30dt\{\\times\}30training uses packed GPU allocations with up to1212runs per allocation, one GPU,44CPU cores,1616GB memory, and a7272\-hour time limit per packed allocation\. Dysts cache construction uses CPU allocations with44CPU cores,1616GB memory, and a2424\-hour time limit\. Dysts long\-horizon evaluation uses CPU allocations with44CPU cores,2424GB memory, and a66\-hour time limit for the paper\-facingdt×30dt\{\\times\}30evaluation queue\. Controlled support diagnostics and support\-routed analyses use CPU allocations with44CPU cores,1616–2424GB memory, and88–1212\-hour time limits, depending on the diagnostic\.
## Appendix CStatistical testing protocol
This appendix describes exactly what the significance annotations test\. Point estimates and statistical annotations answer different questions\. Point estimates in the forecasting and support\-diagnostic tables first summarize completed training seeds within each benchmark system by an interquartile mean \(IQM\), then average those system summaries across systems\. The main exception is the descriptive family\-count column\|Fabs\|¯\\overline\{\|F\_\{\\rm abs\}\|\}, which averages per\-system seed means because it is a count diagnostic rather than a directional performance metric\. Significance annotations instead ask whether a paired effect is reproducible under the independent unit implied by the experimental design\. We therefore do not pool seeds, trajectories, transfers, and systems into one sample\.
#### Common paired effects\.
For MSE comparisons against the dense\-latent MLP baseline in[Table1](https://arxiv.org/html/2608.29057#S4.T1), the paired effect for systemss, replicateuu, and horizonhhis
Δs,u,h=log10Es,u,hcand−log10Es,u,hDense=log10\(Es,u,hcand/Es,u,hDense\),\\Delta\_\{s,u,h\}=\\log\_\{10\}E^\{\\rm cand\}\_\{s,u,h\}\-\\log\_\{10\}E^\{\\rm Dense\}\_\{s,u,h\}=\\log\_\{10\}\\\!\\left\(E^\{\\rm cand\}\_\{s,u,h\}/E^\{\\rm Dense\}\_\{s,u,h\}\\right\),\(12\)whereEEis the held\-out best\-periodic MSE and the one\-sided alternative isΔ<0\\Delta<0\. Log ratios are used only for positive MSE ratios\. They match the multiplicative scientific question and reduce the influence of the heavy\-tailed seed\-level MSE scale\.
Support diagnostics use raw paired differences unless the displayed quantity is already an MSE ratio\. Mean\|Sabs\|\|S\_\{\\rm abs\}\|andH\(B∣Fabs\)H\(B\\mid F\_\{\\rm abs\}\)are tested against Dense MLP with alternative candidate<<baseline\.The family\-count diagnostic\|Fabs\|¯\\overline\{\|F\_\{\\rm abs\}\|\}is not assigned a one\-sided test because closeness to the represented basin count is not monotone in either direction\.
#### Holm correction\.
All Wilcoxon families use Holm step\-down correction atα=0\.05\\alpha=0\.05\. Ifp\(1\)≤⋯≤p\(m\)p\_\{\(1\)\}\\leq\\cdots\\leq p\_\{\(m\)\}are the eligible rawpp\-values in a test family, the adjusted value at rankiiis the running maximum ofmin\{1,\(m−j\+1\)p\(j\)\}\\min\\\{1,\(m\-j\+1\)p\_\{\(j\)\}\\\}forj≤ij\\leq i, then mapped back to the original cell order\. The implementation uses SciPy’s one\-sided Wilcoxon signed\-rank test with Wilcox zero method; all\-zero or otherwise degenerate paired differences are not counted as Holm passes\.
#### Controlled multibasin forecasting and support diagnostics\.
For the controlled multibasin columns of[Table1](https://arxiv.org/html/2608.29057#S4.T1)and the support\-diagnostic columns of[Table1](https://arxiv.org/html/2608.29057#S4.T1), the replicate unit is the training seed\. Candidate and baseline values are paired by the same benchmark system and seed\. For each candidate–metric–horizon–subset cell, the table\-generating scripts run a separate one\-sided paired Wilcoxon signed\-rank test within each system over the finite paired seed deltas\. MSE tests require positive finite MSEs before the log transform\. The paper\-facing table builders require at least four finite paired seeds for a system to be eligible for these within\-system tests\.
The raw per\-systempp\-values are Holm\-corrected across eligible systems for that cell\. A compact superscript∗\\astmeans the Holm\-corrected test passes at0\.050\.05in the prespecified direction\. The multibasin forecasting cells in the compact main table suppress the displayedK/NK/Nbecause every shown non\-Dense multibasin forecasting cell passes on all15/1515/15systems\.
#### Dysts forecasting superscripts\.
The Dysts columns in[Table1](https://arxiv.org/html/2608.29057#S4.T1)use benchmark systems, not seeds, as the independent inferential units\. For each Dysts system and horizon, seeds are first summarized within system:
Is,hcand=IQMr\{Es,r,hcand\},Is,hDense=IQMr\{Es,r,hDense\}\.I^\{\\rm cand\}\_\{s,h\}=\\operatorname\{IQM\}\_\{r\}\\\!\\left\\\{E^\{\\rm cand\}\_\{s,r,h\}\\right\\\},\\qquad I^\{\\rm Dense\}\_\{s,h\}=\\operatorname\{IQM\}\_\{r\}\\\!\\left\\\{E^\{\\rm Dense\}\_\{s,r,h\}\\right\\\}\.\(13\)Each system then contributes one paired log effect
ds,h=log10Is,hcand−log10Is,hDense\.d\_\{s,h\}=\\log\_\{10\}I^\{\\rm cand\}\_\{s,h\}\-\\log\_\{10\}I^\{\\rm Dense\}\_\{s,h\}\.\(14\)The compact Dysts table superscripts come from an exact one\-sided sign test on the number of systems withds,h<0d\_\{s,h\}<0\. Equivalently, fornnretained systems andwwsystems improved over Dense, the rawpp\-value is
Pr\{Binomial\(n,1/2\)≥w\}\.\\Pr\\\{\\operatorname\{Binomial\}\(n,1/2\)\\geq w\\\}\.These sign\-testpp\-values are Holm\-corrected across the non\-Dense model–horizon comparisons in the Dysts analysis family, and the compact table reads the corrected sign\-test values\. The signed\-rank andtt\-test values kept in the Dysts sidecar CSVs are audit diagnostics only; they do not determine the compact table superscripts\. With1010Dysts systems and this correction, a displayed Dysts superscript corresponds to improvement on all10/1010/10systems in that cell\.
#### Support\-coordinate interventions\.
[Figure3](https://arxiv.org/html/2608.29057#S4.F3)is also a mechanism diagnostic rather than a cross\-system significance table\. The coordinate\-dropping panel reports mean MSE curves over the selected held\-out starts, with95%95\\%bootstrap intervals from20002000bootstrap resamples of starts\. The random support\-shuffle panel preserves active coefficient values, moves them to random inactive coordinates, repeats the shuffle procedure, and displays the interquartile range, median, and mean over the shuffle outcomes\. These intervals describe intervention variability for the representative system and seed; they are not Holm\-corrected hypothesis\-test annotations\.
Figure 5:15 Multibasin systems: Per\-system effect sizes atH=1000H\{=\}1000versus the Dense MLP baseline\. Each point is one benchmark system; thexx\-coordinate is the per\-system median pairedlog10\\log\_\{10\}\-MSE difference \(candidate minus baseline\) over the common completed seeds, up to1515per system, with whiskers showing the95%95\\%bootstrap interval\. Filled markers clear Holm\-correctedα=0\.05\\alpha\{=\}0\.05\.
## Appendix DMultibasin benchmark inventory and system definitions
[Table3](https://arxiv.org/html/2608.29057#A4.T3)lists the1515two\-dimensional multibasin systems used in this paper\. Every row contributes to the retained\-system aggregate in[Table1](https://arxiv.org/html/2608.29057#S4.T1), and[Figure2\(a\)](https://arxiv.org/html/2608.29057#S4.F2.sf1)\. Four systems are additionally displayed as the main benchmark visual panels in[Figure1](https://arxiv.org/html/2608.29057#S2.F1)\. Claude Opus 4\.5 in Claude Code provided substantial help in procedurally generating candidates given initial specifications, from which these 15 systems were chosen\.
Basin labels are used only for benchmark annotation, controlled\-transfer construction, and evaluation\. Basin counts are benchmark metadata; for the structured\-transition ablation rows, they are also used only to set a fixed architecture group count, as described in[AppendixB](https://arxiv.org/html/2608.29057#A2), and not for support extraction, support\-family construction, routing, or rollout evaluation\.
Table 3:The1515multibasin benchmark systems\.#### Stored\-step convention\.
All retained systems are implemented as continuous\-time vector fieldsx˙=fs\(x\)\\dot\{x\}=f\_\{s\}\(x\)and integrated with fourth\-order Runge–Kutta to produce the stored observation mapFΔtF\_\{\\Delta t\}used in[eq\.1](https://arxiv.org/html/2608.29057#S2.E1)\. The training windows contain only stored statesxkx\_\{k\}, notfsf\_\{s\}, integration substeps, or basin assignments\.
#### Ground\-truth vector\-field visualizations\.
[Figure6](https://arxiv.org/html/2608.29057#A4.F6)shows the continuous\-time ground\-truth vector fields for all1515retained two\-dimensional multibasin systems\. Streamlines show the direction offs\(x\)f\_\{s\}\(x\), and black crosses mark the generator attractor centers or analytic equilibrium references used for benchmark annotation\. The streamline colors use a panel\-locallog10‖fs\(x\)‖2\\log\_\{10\}\\left\\lVert f\_\{s\}\(x\)\\right\\rVert\_\{2\}scale to reveal structure within each system; they should not be read as comparable speed scales across panels\. Detailed panels with per\-system colorbars are provided in[Figures7](https://arxiv.org/html/2608.29057#A4.F7),[8](https://arxiv.org/html/2608.29057#A4.F8),[9](https://arxiv.org/html/2608.29057#A4.F9),[10](https://arxiv.org/html/2608.29057#A4.F10)and[11](https://arxiv.org/html/2608.29057#A4.F11)\.
Figure 6:Ground\-truth vector fields for the retained1515\-system multibasin benchmark\. The models are trained only on stored state trajectories; these continuous\-time vector fields and attractor annotations are shown here only to document the benchmark geometry and are not provided to the learner\.\(a\)Local\-linear gates\.
\(b\)Transfer\-gated local\-linear\.
\(c\)Arrested spiral\.
Figure 7:Detailed ground\-truth vector\-field panels for the first three retained multibasin systems\. Colorbars report the panel\-locallog10‖fs\(x\)‖2\\log\_\{10\}\\left\\lVert f\_\{s\}\(x\)\\right\\rVert\_\{2\}scale\.\(a\)Asymmetric three\-well\.
\(b\)High\-cross three\-well\.
\(c\)Hexagonal six\-well\.
Figure 8:Detailed ground\-truth vector\-field panels for retained Gaussian\-well systems\. Colorbars report the panel\-locallog10‖fs\(x\)‖2\\log\_\{10\}\\left\\lVert f\_\{s\}\(x\)\\right\\rVert\_\{2\}scale\.\(a\)Octagonal eight\-well\.
\(b\)Pentagonal five\-well\.
\(c\)Square four\-well\.
Figure 9:Detailed ground\-truth vector\-field panels for retained polygonal Gaussian\-well systems\. Colorbars report the panel\-locallog10‖fs\(x\)‖2\\log\_\{10\}\\left\\lVert f\_\{s\}\(x\)\\right\\rVert\_\{2\}scale\.\(a\)Triple\-well Duffing\.
\(b\)SNIC multi\-attractor\.
\(c\)Transition\-routes four\-well\.
Figure 10:Detailed ground\-truth vector\-field panels for specialized retained systems\. Colorbars report the panel\-locallog10‖fs\(x\)‖2\\log\_\{10\}\\left\\lVert f\_\{s\}\(x\)\\right\\rVert\_\{2\}scale\.\(a\)Depth\-gradient four\-well\.
\(b\)Diamond four\-well\.
\(c\)L\-shaped five\-well\.
Figure 11:Detailed ground\-truth vector\-field panels for retained geometry variants\. Colorbars report the panel\-locallog10‖fs\(x\)‖2\\log\_\{10\}\\left\\lVert f\_\{s\}\(x\)\\right\\rVert\_\{2\}scale\.
#### Gaussian\-well family\.
Most retained catalog systems use a common two\-dimensional Gaussian potential with an independent rotational drift\. Forx=\(x1,x2\)∈ℝ2x=\(x\_\{1\},x\_\{2\}\)\\in\\mathbb\{R\}^\{2\}, letRx=\(x2,−x1\)Rx=\(x\_\{2\},\-x\_\{1\}\)and
Vs\(x\)=∑i=1Ms−as,iexp\(−‖x−cs,i‖222σs,i2\)\+γs\(x14\+x24\)\.V\_\{s\}\(x\)=\\sum\_\{i=1\}^\{M\_\{s\}\}\-a\_\{s,i\}\\exp\\\!\\left\(\-\\frac\{\\left\\lVert x\-c\_\{s,i\}\\right\\rVert\_\{2\}^\{2\}\}\{2\\sigma\_\{s,i\}^\{2\}\}\\right\)\+\\gamma\_\{s\}\(x\_\{1\}^\{4\}\+x\_\{2\}^\{4\}\)\.The implemented vector field is
x˙=−∇Vs\(x\)\+ωsRx\.\\dot\{x\}=\-\\nabla V\_\{s\}\(x\)\+\\omega\_\{s\}Rx\.\(15\)Definepi\(M\)\(r,φ\)=r\(cos\(2πi/M\+φ\),sin\(2πi/M\+φ\)\)p\_\{i\}^\{\(M\)\}\(r,\\varphi\)=r\(\\cos\(2\\pi i/M\+\\varphi\),\\sin\(2\\pi i/M\+\\varphi\)\),i=0,…,M−1i=0,\\ldots,M\-1\. The retained Gaussian\-well systems instantiate[eq\.15](https://arxiv.org/html/2608.29057#A4.E15)with the parameters in[Table4](https://arxiv.org/html/2608.29057#A4.T4)\.
Table 4:Parameters for retained Gaussian\-well systems\. Repeated pairs in the\(ai,σi\)\(a\_\{i\},\\sigma\_\{i\}\)column apply to every listed center\.The transition\-routes system uses the same Gaussian\-well form with centers
\{ci\}=\{\(−1\.8,1\.8\),\(1\.8,1\.8\),\(−1\.8,−1\.8\),\(1\.8,−1\.8\)\},\\\{c\_\{i\}\\\}=\\\{\(\-1\.8,1\.8\),\(1\.8,1\.8\),\(\-1\.8,\-1\.8\),\(1\.8,\-1\.8\)\\\},ai=3\.0a\_\{i\}=3\.0,σi=0\.6\\sigma\_\{i\}=0\.6,ω=1\.0\\omega=1\.0, andγ=0\.03\\gamma=0\.03, and adds a corridor\-dependent rotation term:
x˙=−∇Vs\(x\)\+\(1\.0\+0\.3Croute\(x2\)\)Rx,Croute\(x2\)=e−\(x2−1\.8\)2/0\.3\+e−\(x2\+1\.8\)2/0\.3\.\\dot\{x\}=\-\\nabla V\_\{s\}\(x\)\+\\bigl\(1\.0\+0\.3\\,C\_\{\\rm route\}\(x\_\{2\}\)\\bigr\)Rx,\\qquad C\_\{\\rm route\}\(x\_\{2\}\)=e^\{\-\(x\_\{2\}\-1\.8\)^\{2\}/0\.3\}\+e^\{\-\(x\_\{2\}\+1\.8\)^\{2\}/0\.3\}\.\(16\)
#### Native gated local\-linear systems\.
Both native gated systems place three centers on a circular layout, but with different radii and region definitions\. LetQ\(α\)=\(cosα−sinαsinαcosα\)Q\(\\alpha\)=\\begin\{pmatrix\}\\cos\\alpha&\-\\sin\\alpha\\\\ \\sin\\alpha&\\cos\\alpha\\end\{pmatrix\}andαb=2πb/3\\alpha\_\{b\}=2\\pi b/3,b∈\{0,1,2\}b\\in\\\{0,1,2\\\}\.
Forgated\_local\_linear,cb=1\.75\(cosαb,sinαb\)c\_\{b\}=1\.75\(\\cos\\alpha\_\{b\},\\sin\\alpha\_\{b\}\)\. Inside the radius\-1\.051\.05core of basinbb,
x˙=Ab\(x−cb\),Ab=Q\(αb\)A~bQ\(αb\)⊤,\\dot\{x\}=A\_\{b\}\(x\-c\_\{b\}\),\\qquad A\_\{b\}=Q\(\\alpha\_\{b\}\)\\widetilde\{A\}\_\{b\}Q\(\\alpha\_\{b\}\)^\{\\top\},with
A~0=\(−0\.9−1\.21\.2−0\.9\),A~1=\(−1\.350\.2−0\.3−0\.7\),A~2=\(−0\.7−0\.10\.5−1\.2\)\.\\widetilde\{A\}\_\{0\}=\\begin\{pmatrix\}\-0\.9&\-1\.2\\\\ 1\.2&\-0\.9\\end\{pmatrix\},\\quad\\widetilde\{A\}\_\{1\}=\\begin\{pmatrix\}\-1\.35&0\.2\\\\ \-0\.3&\-0\.7\\end\{pmatrix\},\\quad\\widetilde\{A\}\_\{2\}=\\begin\{pmatrix\}\-0\.7&\-0\.1\\\\ 0\.5&\-1\.2\\end\{pmatrix\}\.Outside the cores, the active sector is the nearest angular sectors\(x\)s\(x\)and the shared gate matrix
G=\(−1\.35−0\.90\.9−1\.35\)G=\\begin\{pmatrix\}\-1\.35&\-0\.9\\\\ 0\.9&\-1\.35\\end\{pmatrix\}drives the state asx˙=G\(x−cs\(x\)\)\\dot\{x\}=G\(x\-c\_\{s\(x\)\}\)\.
Forgated\_transfer\_linear,cb=1\.85\(cosαb,sinαb\)c\_\{b\}=1\.85\(\\cos\\alpha\_\{b\},\\sin\\alpha\_\{b\}\)\. The implemented radii and widths are
ρcore\\displaystyle\\rho\_\{\\rm core\}=0\.30,\\displaystyle=0\.30,ρsource\\displaystyle\\rho\_\{\\rm source\}=0\.80,\\displaystyle=0\.80,ρexit\\displaystyle\\rho\_\{\\rm exit\}=0\.60,\\displaystyle=0\.60,βexit\\displaystyle\\beta\_\{\\rm exit\}=0\.72,\\displaystyle=0\.72,wchan\\displaystyle w\_\{\\rm chan\}=0\.22,\\displaystyle=0\.22,δlane\\displaystyle\\delta\_\{\\rm lane\}=0\.28,\\displaystyle=0\.28,ρhand\\displaystyle\\rho\_\{\\rm hand\}=0\.45\.\\displaystyle=0\.45\.The core matrices use the same rotated form with templates
\(−1\.0−1\.11\.1−1\.0\),\(−1\.40\.2−0\.2−0\.7\),\(−0\.8−0\.30\.5−1\.3\)\.\\begin\{pmatrix\}\-1\.0&\-1\.1\\\\ 1\.1&\-1\.0\\end\{pmatrix\},\\quad\\begin\{pmatrix\}\-1\.4&0\.2\\\\ \-0\.2&\-0\.7\\end\{pmatrix\},\\quad\\begin\{pmatrix\}\-0\.8&\-0\.3\\\\ 0\.5&\-1\.3\\end\{pmatrix\}\.For each ordered pairp=\(s,t\)p=\(s,t\),s≠ts\\neq t, define
ds→t=ct−cs‖ct−cs‖2,ns→t=\(−ds→t,2,ds→t,1\),ot=ct‖ct‖2,d\_\{s\\to t\}=\\frac\{c\_\{t\}\-c\_\{s\}\}\{\\left\\lVert c\_\{t\}\-c\_\{s\}\\right\\rVert\_\{2\}\},\\qquad n\_\{s\\to t\}=\(\-d\_\{s\\to t,2\},d\_\{s\\to t,1\}\),\\qquad o\_\{t\}=\\frac\{c\_\{t\}\}\{\\left\\lVert c\_\{t\}\\right\\rVert\_\{2\}\},andχs,t=1\\chi\_\{s,t\}=1whensin\(αs−αt\)≥0\\sin\(\\alpha\_\{s\}\-\\alpha\_\{t\}\)\\geq 0, otherwiseχs,t=−1\\chi\_\{s,t\}=\-1\. The implemented channel entry and handoff points are
ep\\displaystyle e\_\{p\}=cs\+ρsourceds→t\+χs,tδlanens→t,\\displaystyle=c\_\{s\}\+\\rho\_\{\\rm source\}d\_\{s\\to t\}\+\\chi\_\{s,t\}\\delta\_\{\\rm lane\}n\_\{s\\to t\},hp\\displaystyle h\_\{p\}=ct\+ρhandot\+χs,tδlane\(−ot,2,ot,1\),\\displaystyle=c\_\{t\}\+\\rho\_\{\\rm hand\}o\_\{t\}\+\\chi\_\{s,t\}\\delta\_\{\\rm lane\}\(\-o\_\{t,2\},o\_\{t,1\}\),with channel directiondp=\(hp−ep\)/‖hp−ep‖2d\_\{p\}=\(h\_\{p\}\-e\_\{p\}\)/\\left\\lVert h\_\{p\}\-e\_\{p\}\\right\\rVert\_\{2\}, channel normalnp=\(−dp,2,dp,1\)n\_\{p\}=\(\-d\_\{p,2\},d\_\{p,1\}\), and channel lengthℓp=\(hp−ep\)⋅dp\\ell\_\{p\}=\(h\_\{p\}\-e\_\{p\}\)\\cdot d\_\{p\}\.
The vector field is piecewise analytic\. Core regions‖x−cb‖2≤ρcore\\left\\lVert x\-c\_\{b\}\\right\\rVert\_\{2\}\\leq\\rho\_\{\\rm core\}useAb\(x−cb\)A\_\{b\}\(x\-c\_\{b\}\)\. Source annuli use0\.65Ab\(x−cb\)0\.65A\_\{b\}\(x\-c\_\{b\}\), and other non\-channel background regions use0\.50Ab\(x−cb\)0\.50A\_\{b\}\(x\-c\_\{b\}\)for the nearest centercbc\_\{b\}\. Exit wedges forp=\(s,t\)p=\(s,t\)are source\-neighborhood states outside the core with‖x−cs‖2≥ρexit\\left\\lVert x\-c\_\{s\}\\right\\rVert\_\{2\}\\geq\\rho\_\{\\rm exit\}and\(x−cs\)⋅ds→t≥‖x−cs‖2cosβexit\(x\-c\_\{s\}\)\\cdot d\_\{s\\to t\}\\geq\\left\\lVert x\-c\_\{s\}\\right\\rVert\_\{2\}\\cos\\beta\_\{\\rm exit\}; they use
x˙=vexitds→t−λexit\(\(x−cs\)⋅ns→t\)ns→t,vexit=1\.0,λexit=1\.8;\\dot\{x\}=v\_\{\\rm exit\}d\_\{s\\to t\}\-\\lambda\_\{\\rm exit\}\\bigl\(\(x\-c\_\{s\}\)\\cdot n\_\{s\\to t\}\\bigr\)n\_\{s\\to t\},\\qquad v\_\{\\rm exit\}=1\.0,\\quad\\lambda\_\{\\rm exit\}=1\.8;and channel rectangles satisfying
0≤\(x−ep\)⋅dp≤ℓp,\|\(x−ep\)⋅np\|≤wchan0\\leq\(x\-e\_\{p\}\)\\cdot d\_\{p\}\\leq\\ell\_\{p\},\\qquad\|\(x\-e\_\{p\}\)\\cdot n\_\{p\}\|\\leq w\_\{\\rm chan\}use
x˙=vchandp−λchan\(\(x−ep\)⋅np\)np,vchan=1\.55,λchan=2\.8\.\\dot\{x\}=v\_\{\\rm chan\}d\_\{p\}\-\\lambda\_\{\\rm chan\}\\bigl\(\(x\-e\_\{p\}\)\\cdot n\_\{p\}\\bigr\)n\_\{p\},\\qquad v\_\{\\rm chan\}=1\.55,\\quad\\lambda\_\{\\rm chan\}=2\.8\.
#### Other specialized systems\.
The arrested spiral system has a background spiral flow plus four Gaussian traps and an origin trap:
x˙=−0\.3x\+2\.0Rx−4\.0∑i=14e−‖x−ci‖22/\(2⋅0\.25\)x−ci0\.25−1\.5e−‖x‖22/0\.3x0\.15−0\.02\(x13,x23\),\\dot\{x\}=\-0\.3x\+2\.0Rx\-4\.0\\sum\_\{i=1\}^\{4\}e^\{\-\\left\\lVert x\-c\_\{i\}\\right\\rVert\_\{2\}^\{2\}/\(2\\cdot 0\.25\)\}\\frac\{x\-c\_\{i\}\}\{0\.25\}\-1\.5e^\{\-\\left\\lVert x\\right\\rVert\_\{2\}^\{2\}/0\.3\}\\frac\{x\}\{0\.15\}\-0\.02\(x\_\{1\}^\{3\},x\_\{2\}^\{3\}\),withci∈\{\(1\.5,0\),\(0,1\.8\),\(−1,−1\),\(0\.8,−1\.5\)\}c\_\{i\}\\in\\\{\(1\.5,0\),\(0,1\.8\),\(\-1,\-1\),\(0\.8,\-1\.5\)\\\}\. Its fifth basin is the origin trap, matching the benchmark manifest count\.
The triple\-well Duffing system usesx=\(q,p\)x=\(q,p\)and
V\(q\)=q6/6−q4/2\+aq2,V′\(q\)=q5−2q3\+2aq,V\(q\)=q^\{6\}/6\-q^\{4\}/2\+aq^\{2\},\\qquad V^\{\\prime\}\(q\)=q^\{5\}\-2q^\{3\}\+2aq,witha=0\.3a=0\.3, dampingδ=0\.5\\delta=0\.5, rotation strengthω=1\.0\\omega=1\.0, and soft cubic confinement−ϵu3/s3\-\\epsilon u^\{3\}/s^\{3\}withϵ=0\.003\\epsilon=0\.003ands=4\.0s=4\.0:
q˙\\displaystyle\\dot\{q\}=\(1\+ω\)p−ϵq3/s3,\\displaystyle=\(1\+\\omega\)p\-\\epsilon q^\{3\}/s^\{3\},\(17\)p˙\\displaystyle\\dot\{p\}=−V′\(q\)−δp−ωq−ϵp3/s3\.\\displaystyle=\-V^\{\\prime\}\(q\)\-\\delta p\-\\omega q\-\\epsilon p^\{3\}/s^\{3\}\.
The SNIC system is implemented in polar form and then converted to Cartesian coordinates\. Withεnum=10−8\\varepsilon\_\{\\rm num\}=10^\{\-8\},r2=x12\+x22\+εnumr^\{2\}=x\_\{1\}^\{2\}\+x\_\{2\}^\{2\}\+\\varepsilon\_\{\\rm num\},r=r2r=\\sqrt\{r^\{2\}\},θ=atan2\(x2,x1\)\\theta=\\operatorname\{atan2\}\(x\_\{2\},x\_\{1\}\),cθ=x1/\(r\+εnum\)c\_\{\\theta\}=x\_\{1\}/\(r\+\\varepsilon\_\{\\rm num\}\), andsθ=x2/\(r\+εnum\)s\_\{\\theta\}=x\_\{2\}/\(r\+\\varepsilon\_\{\\rm num\}\),
r˙=r\(1−r2\)−0\.3cos\(3θ\),θ˙=1−1\.2cos\(3θ\)\.\\dot\{r\}=r\(1\-r^\{2\}\)\-0\.3\\cos\(3\\theta\),\\qquad\\dot\{\\theta\}=1\-1\.2\\cos\(3\\theta\)\.The Cartesian vector field is
x˙1\\displaystyle\\dot\{x\}\_\{1\}=r˙cθ−rθ˙sθ−0\.01x1\(x12\+x22\)\+0\.5x2,\\displaystyle=\\dot\{r\}c\_\{\\theta\}\-r\\dot\{\\theta\}s\_\{\\theta\}\-0\.01x\_\{1\}\(x\_\{1\}^\{2\}\+x\_\{2\}^\{2\}\)\+0\.5x\_\{2\},x˙2\\displaystyle\\dot\{x\}\_\{2\}=r˙sθ\+rθ˙cθ−0\.01x2\(x12\+x22\)−0\.5x1\.\\displaystyle=\\dot\{r\}s\_\{\\theta\}\+r\\dot\{\\theta\}c\_\{\\theta\}\-0\.01x\_\{2\}\(x\_\{1\}^\{2\}\+x\_\{2\}^\{2\}\)\-0\.5x\_\{1\}\.These analytic generators define the benchmark trajectories\. Held\-out evaluation annotations are produced by the benchmark label helpers or by endpoint\-rollout basin labeling, and the learned models observe only stored states\.
## Appendix EDysts benchmark inventory and system definitions
This appendix lists the ten Dysts systems used in the paper\-facingdt×30dt\{\\times\}30forecasting benchmark\. The systems are drawn from Dysts\[[11](https://arxiv.org/html/2608.29057#bib.bib11),[12](https://arxiv.org/html/2608.29057#bib.bib12)\]\. We used the continuous\-time right\-hand sides and metadata from the pinned package version resolved in the project environment,dysts0\.960\.96\. We use this package under the original license of the repository: Apache\-2\.0\. All ten systems are three\-dimensional autonomous flows\. The equations below are written in unstandardized native Dysts coordinates\(x,y,z\)\(x,y,z\); the training cache standardizes coordinates only after trajectories are generated\.
Table 5:The ten Dysts systems used in thedt×30dt\{\\times\}30forecasting benchmark\. The “stored step” column is3030times the native Dysts integration interval\.#### Closed\-form convention\.
For each system, the stored observations are generated fromx˙=fs\(x\)\\dot\{x\}=f\_\{s\}\(x\)with observation intervalΔt=30dtnative\\Delta t=30\\,dt\_\{\\rm native\}, wheredtnativedt\_\{\\rm native\}is the system\-specific Dysts integration interval in[Table5](https://arxiv.org/html/2608.29057#A5.T5)\. The learner observes only the resulting stored states, not the vector field, continuous\-time solver, or substeps\. All ten systems have explicit closed\-form right\-hand sides in Dysts\. Most are polynomial vector fields\. The Chua and San–Um–Srisuchinwong systems include absolute\-value terms, so they are piecewise smooth rather than globally analytic at the corresponding switching surfaces\.
#### Chua\.
The Chua system uses the piecewise\-linear diode nonlinearity
h\(x\)=m1x\+12\(m0−m1\)\(\|x\+1\|−\|x−1\|\),h\(x\)=m\_\{1\}x\+\\frac\{1\}\{2\}\(m\_\{0\}\-m\_\{1\}\)\\bigl\(\|x\+1\|\-\|x\-1\|\\bigr\),with parameters
α=15\.6,β=28,m0=−1\.142857,m1=−0\.71429\.\\alpha=15\.6,\\qquad\\beta=28,\\qquad m\_\{0\}=\-1\.142857,\\qquad m\_\{1\}=\-0\.71429\.The vector field is
x˙\\displaystyle\\dot\{x\}=α\(y−x−h\(x\)\),\\displaystyle=\\alpha\\bigl\(y\-x\-h\(x\)\\bigr\),\(18\)y˙\\displaystyle\\dot\{y\}=x−y\+z,\\displaystyle=x\-y\+z,z˙\\displaystyle\\dot\{z\}=−βy\.\\displaystyle=\-\\beta y\.
#### Dadras\.
With parameters
c=2,e=9,o=2\.7,p=3,r=1\.7,c=2,\\qquad e=9,\\qquad o=2\.7,\\qquad p=3,\\qquad r=1\.7,the Dadras system is
x˙\\displaystyle\\dot\{x\}=y−px\+oyz,\\displaystyle=y\-px\+oyz,\(19\)y˙\\displaystyle\\dot\{y\}=ry−xz\+z,\\displaystyle=ry\-xz\+z,z˙\\displaystyle\\dot\{z\}=cxy−ez\.\\displaystyle=cxy\-ez\.
#### Dequan Li\.
With parameters
a=40,c=1\.833,d=0\.16,ε=0\.65,f=20,k=55,a=40,\\qquad c=1\.833,\\qquad d=0\.16,\\qquad\\varepsilon=0\.65,\\qquad f=20,\\qquad k=55,the Dequan Li system is
x˙\\displaystyle\\dot\{x\}=a\(y−x\)\+dxz,\\displaystyle=a\(y\-x\)\+dxz,\(20\)y˙\\displaystyle\\dot\{y\}=kx\+fy−xz,\\displaystyle=kx\+fy\-xz,z˙\\displaystyle\\dot\{z\}=cz\+xy−εx2\.\\displaystyle=cz\+xy\-\\varepsilon x^\{2\}\.
#### Hadley\.
With parameters
a=0\.2,b=4,f=9,g=1,a=0\.2,\\qquad b=4,\\qquad f=9,\\qquad g=1,the Hadley system is
x˙\\displaystyle\\dot\{x\}=−y2−z2−ax\+af,\\displaystyle=\-y^\{2\}\-z^\{2\}\-ax\+af,\(21\)y˙\\displaystyle\\dot\{y\}=xy−bxz−y\+g,\\displaystyle=xy\-bxz\-y\+g,z˙\\displaystyle\\dot\{z\}=bxy\+xz−z\.\\displaystyle=bxy\+xz\-z\.
#### Lu–Chen–Cheng\.
With parameters
a=−10,b=−4,c=18\.1,a=\-10,\\qquad b=\-4,\\qquad c=18\.1,the Lu–Chen–Cheng system is
x˙\\displaystyle\\dot\{x\}=−aba\+bx−yz\+c,\\displaystyle=\-\\frac\{ab\}\{a\+b\}x\-yz\+c,\(22\)y˙\\displaystyle\\dot\{y\}=ay\+xz,\\displaystyle=ay\+xz,z˙\\displaystyle\\dot\{z\}=bz\+xy\.\\displaystyle=bz\+xy\.
#### Qi–Chen\.
With parameters
a=38,b=2\.666,c=80,a=38,\\qquad b=2\.666,\\qquad c=80,the Qi–Chen system is
x˙\\displaystyle\\dot\{x\}=a\(y−x\)\+yz,\\displaystyle=a\(y\-x\)\+yz,\(23\)y˙\\displaystyle\\dot\{y\}=cx\+y−xz,\\displaystyle=cx\+y\-xz,z˙\\displaystyle\\dot\{z\}=xy−bz\.\\displaystyle=xy\-bz\.
#### Sakarya\.
With parameters
a=−1,b=1,c=1,h=1,p=1,q=0\.4,r=0\.3,s=1,a=\-1,\\qquad b=1,\\qquad c=1,\\qquad h=1,\\qquad p=1,\\qquad q=0\.4,\\qquad r=0\.3,\\qquad s=1,the Sakarya system is
x˙\\displaystyle\\dot\{x\}=ax\+hy\+syz,\\displaystyle=ax\+hy\+syz,\(24\)y˙\\displaystyle\\dot\{y\}=−by−px\+qxz,\\displaystyle=\-by\-px\+qxz,z˙\\displaystyle\\dot\{z\}=cz−rxy\.\\displaystyle=cz\-rxy\.
#### San–Um–Srisuchinwong\.
With parametera=2a=2, the San–Um–Srisuchinwong system is
x˙\\displaystyle\\dot\{x\}=y−x,\\displaystyle=y\-x,\(25\)y˙\\displaystyle\\dot\{y\}=−ztanh\(x\),\\displaystyle=\-z\\tanh\(x\),z˙\\displaystyle\\dot\{z\}=−a\+xy\+\|y\|\.\\displaystyle=\-a\+xy\+\|y\|\.
#### Shimizu–Morioka\.
With parameters
a=0\.85,b=0\.5,a=0\.85,\\qquad b=0\.5,the Shimizu–Morioka system is
x˙\\displaystyle\\dot\{x\}=y,\\displaystyle=y,\(26\)y˙\\displaystyle\\dot\{y\}=x−ay−xz,\\displaystyle=x\-ay\-xz,z˙\\displaystyle\\dot\{z\}=−bz\+x2\.\\displaystyle=\-bz\+x^\{2\}\.
#### Wang–Sun\.
With parameters
a=0\.2,b=−0\.01,d=−0\.4,e=−1,f=−1,q=1,a=0\.2,\\qquad b=\-0\.01,\\qquad d=\-0\.4,\\qquad e=\-1,\\qquad f=\-1,\\qquad q=1,the Wang–Sun system is
x˙\\displaystyle\\dot\{x\}=ax\+qyz,\\displaystyle=ax\+qyz,\(27\)y˙\\displaystyle\\dot\{y\}=bx\+dy−xz,\\displaystyle=bx\+dy\-xz,z˙\\displaystyle\\dot\{z\}=ez\+fxy\.\\displaystyle=ez\+fxy\.Similar Articles
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
This paper applies TopK Sparse Autoencoders to three EEG foundation models (SleepFM, REVE, LaBraM) to extract interpretable feature dictionaries and introduces a framework for concept steering, revealing representational failures and clinical entanglements.
Sparse Autoencoders for Interpretable Out-of-Distribution Detection
This paper introduces a novel method using sparse autoencoders (SAEs) to learn interpretable features from intermediate network activations for out-of-distribution (OOD) detection, achieving state-of-the-art performance and providing insights into how distribution shifts affect learned representations.
Learning the Koopman Operator using Attention Free Transformers
This paper introduces attention-free latent memory and dynamic re-encoding to improve long-horizon predictions in Koopman autoencoders, reducing error accumulation on benchmark dynamical systems.
MetaKoopman: Bayesian Meta-Learning of Koopman Operators for Modeling Structured Dynamics under Distribution Shifts
MetaKoopman proposes a Bayesian meta-learning framework for modeling nonlinear dynamics using linear latent representations via Koopman operators, enabling closed-form updates and uncertainty quantification. It is validated on autonomous truck and trailer systems under adverse winter conditions, outperforming prior methods in prediction accuracy and robustness.
WriteSAE: Sparse Autoencoders for Recurrent State
WriteSAE introduces the first sparse autoencoder that decomposes matrix cache writes in state-space and hybrid recurrent language models, enabling superior token-level interventions compared to existing methods.