Principled Koopman Representations with Kalman Inference for Efficient Time-Series Prediction
Summary
This paper presents K2SVD, a principled method for learning Koopman operator representations with Kalman inference to achieve efficient and accurate time-series prediction, demonstrating superior performance over existing state-of-the-art approaches.
View Cached Full Text
Cached at: 09/17/26, 08:52 AM
# Principled Koopman Representations with Kalman Inference for Efficient Time-Series Prediction
Source: [https://arxiv.org/html/2609.17815](https://arxiv.org/html/2609.17815)
###### Abstract
The Koopman operator has been widely used for time\-series prediction in dynamical systems\. However, prior work that learns latent “Koopman spaces” using neural networks often did not construct a valid Koopman space for forecasting, as these representations may be mathematically inconsistent with the operator\-theoretic formulation and fail to capture the intrinsic low\-rank structure of system dynamics\. To address this issue, we introduce K2SVD, a method that explicitly learns the leading singular functions of the Koopman operator by optimizing a Hilbert\-Schmidt objective\. This yields a well\-defined low\-rank approximation of the Koopman operator with an interpretable linear combination, featuring a compact latent space with less than10%10\\%of the dimensions used in previous work\. In the learned Koopman space, K2SVD further captures temporal evolution with a linear Gaussian state\-space model and performs inference via Kalman filtering, mitigating noise accumulation during multi\-step prediction\. Empirical results show that K2SVD outperforms state\-of\-the\-art methods across multiple datasets, with significantly faster prediction speeds and lower computational cost than previous efficiency\-focused models\. This highlights the benefits of principled low\-rank Koopman representations and opens up broader potential for applications\.
1Department of Computer Science, UC Santa Barbara
rli667@ucsb\.edu, buyuheng@ucsb\.edu
## 1Introduction
Time\-series forecasting is a fundamental problem across a wide range of scientific and industrial domains, often requiring models to handle highly nonlinear transformations under noisy perturbations\. Recent deep learning models, including transformers\([Li et al\. 2019](https://arxiv.org/html/2609.17815#bib.bib28)\), sequence\-to\-sequence architectures, and foundation models\([Das et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib29);[Ansari et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib2)\), have achieved strong empirical performance\. However, in domains such as economics, weather, and energy, where explainable and expressive latent spaces are crucial, purely black\-box prediction models are often insufficient\.
Koopman operator\([Lan and Mezić 2013](https://arxiv.org/html/2609.17815#bib.bib5)\)offers a compelling alternative to black\-box models by providing a mathematically grounded and expressive latent space\. By lifting data into an infinite\-dimensional space of measurement functions, the Koopman framework represents nonlinear dynamics through a linear operator that maps measurements to their next\-step evolution\. This perspective also enables spectral methods to analyze the widely observed low\-rank structure in system dynamics, bridging classical dynamical systems theory and modern machine learning\. In particular, recent Koopman singular value decomposition \(SVD\)\-based methods\([Jeong et al\. 2025](https://arxiv.org/html/2609.17815#bib.bib6);[Kostic et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib3);[Wu and Noé 2020](https://arxiv.org/html/2609.17815#bib.bib4)\)learn the leading singular functions of the operator, demonstrating the effectiveness of Koopman representations for modeling dynamical systems\.
Despite these advances, many existing approaches to probabilistic time\-series prediction\([Liu et al\. 2026](https://arxiv.org/html/2609.17815#bib.bib10);[Liu et al\. 2023](https://arxiv.org/html/2609.17815#bib.bib17)\)that claim to construct a Koopman space rely on heuristic or loosely defined latent representations, and do not rigorously adhere to the underlying operator\-theoretic framework\. In particular, the learned embeddings are often optimized for reconstruction loss solely rather than for satisfying the definition of Koopman\-invariant subspaces\. As a result, they may fail to preserve linear evolution in a well\-defined function space and may require redundant modules to handle nonlinearity\. Consequently, the resulting “Koopman spaces” may not admit a consistent operator interpretation, and their associated dynamics can deviate from the true spectral structure, especially the low\-rank structure, leading to parameter redundancy and inefficient prediction\.
Figure 1:Comparison of predictive performance, GPU usage, and inference time on the Electricity dataset\. K2SVD outperforms competing methods in time and space efficiency, as well as overall performance\. Circle size indicates peak allocated GPU memory during training\.To address these limitations, we propose K2SVD, a principled framework that explicitly enforces a low\-rank Koopman structure by learning approximations to its leading singular functions\. Rather than treating the latent space as an unconstrained embedding, our method constructs a theoretically grounded representation by optimizing a Hilbert\-Schmidt objective that directly targets the best rank\-kkapproximation of the operator\. This yields a latent space with well\-defined linear dynamics, enabling faithful spectral characterization and linear temporal evolution, which also reduces the computation dramatically\. In the Koopman latent space, temporal evolution is modeled as a linear transition with nonlinear perturbations\. Given a principled Koopman space, we posit that these nonlinear residuals have low magnitude and limited predictive information, and can therefore be treated as Gaussian noise to be filtered\. Furthermore, due to the iterative nature of time\-series prediction, we perform inference via Kalman filtering to mitigate noise accumulation during the forward pass\.
The overall model is trained using a two\-stage scheme that first establishes the Koopman subspace and then refines the dynamical and observation components using Kalman filter in an end\-to\-end manner\. This design combines theoretical rigor with practical scalability, yielding a model that captures low\-rank dynamics while maintaining strong forecasting performance and parameter efficiency\. As shown in Figure[1](https://arxiv.org/html/2609.17815#S1.F1), K2SVD achieves competitive prediction performance with substantially low GPU memory usage and shorter inference time than existing methods\. Our contribution is as follows:
- •We propose K2SVD, a principled Koopman forecasting framework that learns a low\-rank Koopman space by approximating the leading singular functions of the Koopman operator through a Hilbert\-Schmidt objective, yielding a mathematically grounded latent space\.
- •We integrate the learned Koopman representation with a linear Gaussian state\-space model and Kalman filtering, enabling efficient multi\-step inference that mitigates noise accumulation\.
- •We empirically show that K2SVD achieves competitive forecasting performance across multiple dynamic systems and real\-world time\-series benchmarks while requiring substantially fewer parameters and shorter inference time than existing methods, demonstrating the benefit of explicitly modeling low\-rank Koopman structure\.
## 2Related Works
Probabilistic time series forecasting\.Time series prediction has evolved from classical autoregressive models to deep generative frameworks\. Early neural approaches such as DeepAR\([Salinas et al\. 2020](https://arxiv.org/html/2609.17815#bib.bib22)\), N\-BEATS\([Oreshkin et al\. 2020](https://arxiv.org/html/2609.17815#bib.bib23)\), and transformer\-based models learn flexible sequence representations for multi\-horizon forecasting\([Lim et al\. 2021](https://arxiv.org/html/2609.17815#bib.bib24);[Zhou et al\. 2021](https://arxiv.org/html/2609.17815#bib.bib25);[Rasul et al\. 2021](https://arxiv.org/html/2609.17815#bib.bib19)\)\. More recent works incorporate structured dynamical priors to improve long\-term prediction and interpretability\.
Koopman operator learning\.The Koopman operator framework provides a principled approach to analyzing nonlinear dynamical systems by lifting them into a linear space\. Classical works establish their spectral foundations and connections to ergodic theory and modal decomposition\([Mezić 2005](https://arxiv.org/html/2609.17815#bib.bib7);[Mezić 2013](https://arxiv.org/html/2609.17815#bib.bib8);[Rowley et al\. 2009](https://arxiv.org/html/2609.17815#bib.bib9)\)\. Building on this theory, data\-driven approximations such as dynamic mode decomposition \(DMD\) and its extensions enable finite\-dimensional representations of the Koopman operator directly from data\([Schmid 2010](https://arxiv.org/html/2609.17815#bib.bib11);[Tu 2013](https://arxiv.org/html/2609.17815#bib.bib12);[Williams et al\. 2015a](https://arxiv.org/html/2609.17815#bib.bib13);[Williams et al\. 2015b](https://arxiv.org/html/2609.17815#bib.bib14)\)\. In particular, low\-rank approximations based on SVD play a central role in extracting dominant Koopman modes, and more recently, learning\-based approaches leverage neural networks to recover Koopman subspaces\([Wu and Noé 2020](https://arxiv.org/html/2609.17815#bib.bib4);[Mardt et al\. 2018](https://arxiv.org/html/2609.17815#bib.bib15);[Kostic et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib3)\)\. These methods highlight the importance of low\-rank Koopman representations and SVD approximations for high\-dimensional dynamical systems\.
Koopman\-Based Probabilistic Forecasting\.Koopman\-based models have gained increasing attention\. Koopman Neural Forecaster \(KNF\) integrates Koopman embeddings with distributional forecasting to handle temporal distribution shifts\([Wang et al\. 2022](https://arxiv.org/html/2609.17815#bib.bib26)\), while Koopa introduces hierarchical Koopman predictors to decompose non\-stationary dynamics into invariant and time\-varying components\([Liu et al\. 2023](https://arxiv.org/html/2609.17815#bib.bib17)\)\. Related generative approaches, such as Koopman VAEs, embed linear dynamics into latent representations for stable sequence generation\([Naiman et al\. 2023](https://arxiv.org/html/2609.17815#bib.bib27)\)\. Hybrid probabilistic models further combine Koopman theory with latent\-variable methods, including Kalman filtering\. These methods typically use Koopman lifting with Kalman state\-space updates: K\-ESKF\([Huang et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib31)\)improves state estimation during highly agile motion, while K2VAE\([Wu et al\. 2025](https://arxiv.org/html/2609.17815#bib.bib1)\)uses variational autoencoder training for long\-term prediction\. Together, these works show that combining Koopman theory with probabilistic deep learning offers a powerful framework for modeling complex temporal dynamics under uncertainty\.
## 3Preliminaries
Koopman theory\.Consider a stochastic discrete\-time dynamical system
𝐱t\+1=ξ\(F\(𝐱t\),ϵt\)\.\\mathbf\{x\}\_\{t\+1\}=\\xi\(F\(\\mathbf\{x\}\_\{t\}\),\\epsilon\_\{t\}\)\.\(1\)Here,F:𝒳→𝒳F:\\mathcal\{X\}\\rightarrow\\mathcal\{X\}is a possibly nonlinear mapping for a domain𝒳⊆ℝd\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\},ϵt∼𝒟\\epsilon\_\{t\}\\sim\\mathcal\{D\}denotes independent random noise\. Assuming thatFFand the noise distribution𝒟\\mathcal\{D\}are time\-invariant, the process becomes a time\-homogeneous Markov process with transition densityp\(𝐱′∣𝐱\)p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\), i\.e\.,Pr\(𝐗t\+1∈A\|𝐗t=𝐱\)=∫Ap\(𝐱′∣𝐱\)d𝐱′\\Pr\(\\mathbf\{X\}\_\{t\+1\} \\in A \\mid\\mathbf\{X\}\_t = \\mathbf\{x\}\)=\\int\_\{A\}p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\)\\,d\\mathbf\{x\}^\{\\prime\}for all measurable setsA⊆𝒳A\\subseteq\\mathcal\{X\}and allt≥0t\\geq 0\.
In the stochastic setup, the dynamics is fully captured by the transition densityp\(𝐱′∣𝐱\)p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\), which is induced by\(ξ∘F\)\(⋅\)\(\\xi\\circ F\)\(\\cdot\), and thus the problem becomes analyzingp\(𝐱′∣𝐱\)p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\)of a Markov chain\. Here, the Koopman operator becomes the conditional expectation operator, i\.e\., for an observableg:𝒳→ℝg:\\mathcal\{X\}\\rightarrow\\mathbb\{R\},
\(𝒦g\)\(𝐱\)≜𝔼p\(𝐱′∣𝐱\)\[g\(𝐱′\)\]\.\(\\mathcal\{K\}g\)\(\\mathbf\{x\}\)\\triangleq\\mathbb\{E\}\_\{p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\)\}\[g\(\\mathbf\{x\}^\{\\prime\}\)\]\.\(2\)By the Markov property, repeatedly applying the Koopman operator will correspond to themulti\-step predictionin terms of the posterior mean ofg\(𝐗t\)g\(\\mathbf\{X\}\_\{t\}\)given𝐗0=𝐱0\\mathbf\{X\}\_\{0\}=\\mathbf\{x\}\_\{0\}, i\.e\.,\(𝒦tg\)\(𝐱0\)=𝔼p\(𝐱t∣𝐱0\)\[g\(𝐱t\)\]\(\\mathcal\{K\}^\{t\}g\)\(\\mathbf\{x\}\_\{0\}\)=\\mathbb\{E\}\_\{p\(\\mathbf\{x\}\_\{t\}\\mid\\mathbf\{x\}\_\{0\}\)\}\[g\(\\mathbf\{x\}\_\{t\}\)\]\. We will assume that the operator is compact throughout this paper\.
Problem setting\.Given historical dataX=\(𝐱1,𝐱2,⋯,𝐱T\)∈ℝN×TX=\(\\mathbf\{x\}\_\{1\},\\mathbf\{x\}\_\{2\},\\cdots,\\mathbf\{x\}\_\{T\}\)\\in\\mathbb\{R\}^\{N\\times T\}, our goal is to model the stochastic process and predict the future sequenceY=\(𝐱1,𝐱2,⋯,𝐱T\+1,𝐱T\+2,⋯,𝐱T\+t\)Y=\(\\mathbf\{x\}\_\{1\},\\mathbf\{x\}\_\{2\},\\cdots,\\mathbf\{x\}\_\{T\+1\},\\mathbf\{x\}\_\{T\+2\},\\cdots,\\mathbf\{x\}\_\{T\+t\}\)sampling each data from their conditional distributionp\(𝐱t\|𝐱0\)p\(\\mathbf\{x\}\_\{t\}\|\\mathbf\{x\}\_\{0\}\)\. Suppose two data points at different time spots𝐱,𝐱′\\mathbf\{x\},\\mathbf\{x\}^\{\\prime\}, there marginal distributions areρ0,ρ1\\rho\_\{0\},\\rho\_\{1\}satisfiesρ1\(𝐱′\)≜𝔼ρ0\(𝐱\)\[p\(𝐱′\|𝐱\)\]\\rho\_\{1\}\(\\mathbf\{x\}^\{\\prime\}\)\\triangleq\\mathbb\{E\}\_\{\\rho\_\{0\}\(\\mathbf\{x\}\)\}\[p\(\\mathbf\{x\}^\{\\prime\}\|\\mathbf\{x\}\)\]\.
Let𝒳\\mathcal\{X\}be a measurable space and letρ0\\rho\_\{0\}andρ1\\rho\_\{1\}be \(finite\) measures on𝒳\\mathcal\{X\}representing the distributions of the current and future states, respectively\. For any measureρ\\rhoon𝒳\\mathcal\{X\}defineLρ2\(𝒳\)≜\{f:𝒳→ℂ∣∫\|f\(𝐱\)\|2ρ\(d𝐱\)<∞\}L^\{2\}\_\{\\rho\}\(\\mathcal\{X\}\)\\triangleq\\\{f:\\mathcal\{X\}\\rightarrow\\mathbb\{C\}\\mid\\int\|f\(\\mathbf\{x\}\)\|^\{2\}\\rho\(d\\mathbf\{x\}\)<\\infty\\\}\. Equipped with the inner product⟨f,g⟩ρ≜∫f\(𝐱\)g\(𝐱\)ρ\(𝑑𝐱\)\\langle f,g\\rangle\_\{\\rho\}\\triangleq\\int f\(\\mathbf\{x\}\)g\(\\mathbf\{x\}\)\\rho\(d\\mathbf\{x\}\),Lρ2\(𝒳\)L^\{2\}\_\{\\rho\}\(\\mathcal\{X\}\)is a Hilbert space\. The Koopman operator𝒦\\mathcal\{K\}is then a mapping fromLρ12\(𝒳\)L^\{2\}\_\{\\rho\_\{1\}\}\(\\mathcal\{X\}\)toLρ02\(𝒳\)L^\{2\}\_\{\\rho\_\{0\}\}\(\\mathcal\{X\}\)\.
SVD of Operators\.Given𝒦:Lρ12\(𝒳\)→Lρ02\(𝒳\)\\mathcal\{K\}:L^\{2\}\_\{\\rho\_\{1\}\}\(\\mathcal\{X\}\)\\rightarrow L^\{2\}\_\{\\rho\_\{0\}\}\(\\mathcal\{X\}\), the Singular Value decomposition\(SVD\) decomposes𝒦\\mathcal\{K\}into a countable sum of rank\-1 operators:
𝒦=∑i=1∞σifi⊗gi,\\mathcal\{K\}=\\sum\_\{i=1\}^\{\\infty\}\\sigma\_\{i\}\\,f\_\{i\}\\otimes g\_\{i\},\(3\)where the outer product operator acts as\(fi⊗gi\)h=⟨gi,h⟩ρ1fi\(f\_\{i\}\\otimes g\_\{i\}\)h=\\langle g\_\{i\},h\\rangle\_\{\\rho\_\{1\}\}f\_\{i\}\. The singular values and singular functions satisfy:
𝒦gi=σifi,𝒦∗fi=σigi,\\mathcal\{K\}g\_\{i\}=\\sigma\_\{i\}f\_\{i\},\\qquad\\mathcal\{K\}^\{\*\}f\_\{i\}=\\sigma\_\{i\}g\_\{i\},\(4\)with orthogonality conditions:
⟨gi,gj⟩ρ1=δij,⟨fi,fj⟩ρ0=δij,\\langle g\_\{i\},g\_\{j\}\\rangle\_\{\\rho\_\{1\}\}=\\delta\_\{ij\},\\qquad\\langle f\_\{i\},f\_\{j\}\\rangle\_\{\\rho\_\{0\}\}=\\delta\_\{ij\},\(5\)where\{gi\}⊂Lρ12\(𝒳\)\\\{g\_\{i\}\\\}\\subset L^\{2\}\_\{\\rho\_\{1\}\}\(\\mathcal\{X\}\)forms an orthonormal basis of the input space \(observables at current state𝐱t\\mathbf\{x\}\_\{t\}\), and\{fi\}⊂Lρ02\(𝒳\)\\\{f\_\{i\}\\\}\\subset L^\{2\}\_\{\\rho\_\{0\}\}\(\\mathcal\{X\}\)forms an orthonormal basis of the output space \(observables at future state𝐱t\+1\\mathbf\{x\}\_\{t\+1\}\), with singular valuesσ1≥σ2≥⋯≥0\\sigma\_\{1\}\\geq\\sigma\_\{2\}\\geq\\cdots\\geq 0\. In a real\-world scenario, the transition often processesLow\-Rank \(LoRA\) properties, which impliesσi\\sigma\_\{i\}will decline to00rapidly asiiincreases\.
Second\-moment matricesare key quantities in SVD\. We thus introduce the following shorthand notation for convenience\. Thesecond\-moment matrixof vector\-valued functions𝐡1\(⋅\)\\mathbf\{h\}\_\{1\}\(\\cdot\)and𝐡2\(⋅\)\\mathbf\{h\}\_\{2\}\(\\cdot\)with respect to a distributionρ\\rhois denoted as
𝖬ρ\[𝐡1,𝐡2\]≜𝔼ρ\(𝐱\)\[𝐡1\(𝐱\)𝐡2\(𝐱\)⊤\],\\mathsf\{M\}\_\{\\rho\}\[\\mathbf\{h\}\_\{1\},\\mathbf\{h\}\_\{2\}\]\\triangleq\\mathbb\{E\}\_\{\\rho\(\\mathbf\{x\}\)\}\[\\mathbf\{h\}\_\{1\}\(\\mathbf\{x\}\)\\mathbf\{h\}\_\{2\}\(\\mathbf\{x\}\)^\{\\top\}\],\(6\)and we write𝖬ρ\[𝐡\]≜𝖬ρ\[𝐡,𝐡\]\\mathsf\{M\}\_\{\\rho\}\[\\mathbf\{h\}\]\\triangleq\\mathsf\{M\}\_\{\\rho\}\[\\mathbf\{h\},\\mathbf\{h\}\]for shorthand\. Thejoint second\-moment matrixof𝐡1\(⋅\)\\mathbf\{h\}\_\{1\}\(\\cdot\)and𝐡2\(⋅\)\\mathbf\{h\}\_\{2\}\(\\cdot\)over the joint distributionρ0\(𝐱\)p\(𝐱′∣𝐱\)\\rho\_\{0\}\(\\mathbf\{x\}\)p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\)is defined and denoted as
𝖳\[𝐡1,𝐡2\]≜𝔼ρ0\(𝐱\)p\(𝐱′∣𝐱\)\[𝐡1\(𝐱\)𝐡2\(𝐱′\)⊤\],\\mathsf\{T\}\[\\mathbf\{h\}\_\{1\},\\mathbf\{h\}\_\{2\}\]\\triangleq\\mathbb\{E\}\_\{\\rho\_\{0\}\(\\mathbf\{x\}\)p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\)\}\[\\mathbf\{h\}\_\{1\}\(\\mathbf\{x\}\)\\mathbf\{h\}\_\{2\}\(\\mathbf\{x\}^\{\\prime\}\)^\{\\top\}\],\(7\)which satisfies the identities𝖳\[𝐡1,𝐡2\]=𝖬ρ0\[𝐡1,𝒦𝐡2\]=𝖬ρ1\[𝒦∗𝐡1,𝐡2\]\\mathsf\{T\}\[\\mathbf\{h\}\_\{1\},\\mathbf\{h\}\_\{2\}\]=\\mathsf\{M\}\_\{\\rho\_\{0\}\}\[\\mathbf\{h\}\_\{1\},\\mathcal\{K\}\\mathbf\{h\}\_\{2\}\]=\\mathsf\{M\}\_\{\\rho\_\{1\}\}\[\\mathcal\{K\}^\{\*\}\\mathbf\{h\}\_\{1\},\\mathbf\{h\}\_\{2\}\]\. In practice, these quantities are to be estimated with trajectories, and their empirical estimates are denoted with hats \(^\\hat\{\\phantom\{x\}\}\), i\.e\.,𝖬^ρ0\[𝐟\]\\hat\{\\mathsf\{M\}\}\_\{\\rho\_\{0\}\}\[\\mathbf\{f\}\],𝖬^ρ1\[𝐠\]\\hat\{\\mathsf\{M\}\}\_\{\\rho\_\{1\}\}\[\\mathbf\{g\}\], and𝖳^\[𝐟,𝐠\]\\hat\{\\mathsf\{T\}\}\[\\mathbf\{f\},\\mathbf\{g\}\]\.
## 4Proposed Method
Figure 2:Overview of the two\-stage K2SVD procedure for learning a low\-rank Koopman space and performing Kalman inference\.We propose K2SVD, a time\-series forecasting framework that combines Koopman representations with Kalman inference to model complex nonlinear dynamics through linear evolution in a latent space\. Specifically, K2SVD learns a neural representation that lifts observations into a low\-rank Koopman space, where state transitions are modeled as linear dynamics perturbed by Gaussian noise\. This enables efficient and principled inference via the Kalman filter\.
K2SVD is trained with a two\-stage algorithm: first, it learns the low\-rank Koopman space by optimizing the Hilbert–Schmidt objective; second, it freezes the learned Koopman basis and trains the Kalman inference and observation components for forecasting\. The predicted latent states are then mapped back to the observation space through a linear projection\. This design combines neural expressivity with the analytical tractability of linear dynamical systems\. The full procedure is illustrated in Figure[2](https://arxiv.org/html/2609.17815#S4.F2)and the pipeline is summarized in Appendix[B](https://arxiv.org/html/2609.17815#A2)\.
### 4\.1First Stage: Learning the Koopman Space
#### Encoder
We considermultivariate patchesas tokens to implicitly model the cross\-variable interaction during state transition\. We divide the context series
X=\[𝐱1,𝐱2,⋯,𝐱T\]∈ℝN×TX=\[\\mathbf\{x\}\_\{1\},\\mathbf\{x\}\_\{2\},\\cdots,\\mathbf\{x\}\_\{T\}\]\\in\\mathbb\{R\}^\{N\\times T\}into non\-overlapping patches:
XP=\[𝐱1P,𝐱2P,⋯,𝐱nP\]∈ℝN×s×n,X^\{P\}=\[\\mathbf\{x\}\_\{1\}^\{P\},\\mathbf\{x\}\_\{2\}^\{P\},\\cdots,\\mathbf\{x\}\_\{n\}^\{P\}\]\\in\\mathbb\{R\}^\{N\\times s\\times n\},wheres=T/ns=T/ndenotes the patch size,nnis the number of patches, and patch𝐱iP∈ℝN×s\\mathbf\{x\}\_\{i\}^\{P\}\\in\\mathbb\{R\}^\{N\\times s\}\. Similarly, we patchwise the horizon as well with the same patch sizessas follows:
YP=\[𝐱1P,𝐱2P,⋯,𝐱mP\]∈ℝN×s×m,Y^\{P\}=\[\\mathbf\{x\}\_\{1\}^\{P\},\\mathbf\{x\}\_\{2\}^\{P\},\\cdots,\\mathbf\{x\}\_\{m\}^\{P\}\]\\in\\mathbb\{R\}^\{N\\times s\\times m\},In the following, we capture the singular functions of the desired Koopman space with neural networks:𝐟\(XP\)=\[𝐟\(𝐱1P\),𝐟\(𝐱2P\),⋯,𝐟\(𝐱nP\)\]\\mathbf\{f\}\(X^\{P\}\)=\[\\mathbf\{f\}\(\\mathbf\{x\}\_\{1\}^\{P\}\),\\mathbf\{f\}\(\\mathbf\{x\}\_\{2\}^\{P\}\),\\cdots,\\mathbf\{f\}\(\\mathbf\{x\}\_\{n\}^\{P\}\)\], and for similar cases𝐠\(XP\)=\[𝐠\(𝐱1P\),𝐠\(𝐱2P\),⋯,𝐠\(𝐱nP\)\]\\mathbf\{g\}\(X^\{P\}\)=\[\\mathbf\{g\}\(\\mathbf\{x\}\_\{1\}^\{P\}\),\\mathbf\{g\}\(\\mathbf\{x\}\_\{2\}^\{P\}\),\\cdots,\\mathbf\{g\}\(\\mathbf\{x\}\_\{n\}^\{P\}\)\]where𝐟,𝐠\\mathbf\{f\},\\mathbf\{g\}are desired neural network corresponding to the intrinsic property of data\.
#### LoRA Loss Function
The best rank\-ddapproximation of𝒦\\mathcal\{K\}in the Hilbert\-Schmidt norm is:
𝒦≈∑i=1dσifi\(𝐱\)⊗gi\(𝐱′\)\\mathcal\{K\}\\approx\\sum\_\{i=1\}^\{d\}\\sigma\_\{i\}\\,f\_\{i\}\(\\mathbf\{x\}\)\\otimes g\_\{i\}\(\\mathbf\{x\}^\{\\prime\}\)\(8\)with the learnedfif\_\{i\}andgig\_\{i\}spans the optimal approximated rankddsubspaceℋd\\mathcal\{H\}^\{d\}of the output and input space\. To bypass the numerical instability and optimization challenges,\([Jeong et al\. 2025](https://arxiv.org/html/2609.17815#bib.bib6)\)propose to directly minimize the low\-rank approximation error‖𝒦−∑i=1dfi⊗gi‖HS2\\\|\\mathcal\{K\}\-\\sum\_\{i=1\}^\{d\}f\_\{i\}\\otimes g\_\{i\}\\\|^\{2\}\_\{\\text\{HS\}\}to find the singular functions of the low\-rank Koopman operator Hilbert space, where∥⋅∥HS\\\|\\cdot\\\|\_\{HS\}denotes the Hilbert–Schmidt loss\. Succinctly, the learning objective can be expressed as:
ℒlora\(𝐟~,𝐠~\)≜−2tr\(𝖳\[𝐟~,𝐠~\]\)\+tr\(𝖬ρ0\[𝐟~\]𝖬ρ1\[𝐠~\]\),\\mathcal\{L\}\_\{\\text\{lora\}\}\(\\tilde\{\\mathbf\{f\}\},\\tilde\{\\mathbf\{g\}\}\)\\triangleq\-2\\,\\text\{tr\}\(\\mathsf\{T\}\[\\tilde\{\\mathbf\{f\}\},\\tilde\{\\mathbf\{g\}\}\]\)\+\\text\{tr\}\(\\mathsf\{M\}\_\{\\rho\_\{0\}\}\[\\tilde\{\\mathbf\{f\}\}\]\\mathsf\{M\}\_\{\\rho\_\{1\}\}\[\\tilde\{\\mathbf\{g\}\}\]\),\(9\)where𝐟~=\[𝐟\(𝐱2P\),𝐟\(𝐱3P\),⋯,𝐟\(𝐱nP\)\]\\tilde\{\\mathbf\{f\}\}\\\!=\\\!\[\\mathbf\{f\}\(\\mathbf\{x\}\_\{2\}^\{P\}\),\\mathbf\{f\}\(\\mathbf\{x\}\_\{3\}^\{P\}\),\\cdots,\\mathbf\{f\}\(\\mathbf\{x\}\_\{n\}^\{P\}\)\], and𝐠~=\[𝐠\(𝐱1P\),𝐠\(𝐱2P\),⋯,𝐠\(𝐱n−1P\)\]\\tilde\{\\mathbf\{g\}\}\\\!=\\\!\[\\mathbf\{g\}\(\\mathbf\{x\}\_\{1\}^\{P\}\),\\mathbf\{g\}\(\\mathbf\{x\}\_\{2\}^\{P\}\),\\cdots,\\mathbf\{g\}\(\\mathbf\{x\}\_\{n\-1\}^\{P\}\)\]\. A detailed derivation of the loss is provided in Appendix[A](https://arxiv.org/html/2609.17815#A1)\. Noting that the LoRA loss does not require ground\-truth prediction for calculating the gradient, we can train it separately in a stage before inference to stabilize and well\-structure the Koopman space\.
For generality we denote the base of singular space as𝐛∈\{𝐟,𝐠\}\\mathbf\{b\}\\in\\\{\\mathbf\{f\},\\mathbf\{g\}\\\}, and𝐛P=\[𝐛1,𝐛2,⋯,𝐛n\]\\mathbf\{b\}^\{P\}=\[\\mathbf\{b\}\_\{1\},\\mathbf\{b\}\_\{2\},\\cdots,\\mathbf\{b\}\_\{n\}\]\. Since the derived singular functions span the best approximated rankddsubspace of the input and output spaceLρ2\(𝒳\)L^\{2\}\_\{\\rho\}\(\\mathcal\{X\}\), the linear combination of them𝐡\(XP\)=∑i=0dzi𝐛i\(𝐱iP\)\\mathbf\{h\}\(X^\{P\}\)=\\sum\_\{i=0\}^\{d\}z\_\{i\}\\mathbf\{b\}\_\{i\}\(\\mathbf\{x\}\_\{i\}^\{P\}\)is also an approximation inℋd\\mathcal\{H\}^\{d\}of its corresponding infinite dimensional counterpart\. Therefore, there exists a𝖪b:ℋd→ℋd\\mathsf\{K\}\_\{b\}:\\mathcal\{H\}^\{d\}\\to\\mathcal\{H\}^\{d\}such that
\(𝒦𝐛\)\(𝐱\)=𝔼p\(𝐱′∣𝐱\)\[𝐛\(𝐱′\)\]≈𝖪𝐛𝐛\(𝐱\)\.\(\\mathcal\{K\}\\mathbf\{b\}\)\(\\mathbf\{x\}\)=\\mathbb\{E\}\_\{p\(\\mathbf\{x\}^\{\\prime\}\\mid\\mathbf\{x\}\)\}\[\\mathbf\{b\}\(\\mathbf\{x\}^\{\\prime\}\)\]\\approx\\mathsf\{K\}\_\{\\mathbf\{b\}\}\\mathbf\{b\}\(\\mathbf\{x\}\)\.\(10\)
Ideally, if the data obeys the dynamics given by \([1](https://arxiv.org/html/2609.17815#S3.E1)\), i\.e\., has no temporal distribution shift, Markovian, and the latent dimension is sufficient for linearization, the prediction can be performed by a simple linear rollout\. Given the linear relation,𝖪b\\mathsf\{K\}\_\{b\}can be approximately solved with linear regression𝖪𝐛ols≜\(𝖬ρ0\[𝐛\]\)\+𝖳\[𝐛\]∈ℝk×k\\mathsf\{K\}^\{\\text\{ols\}\}\_\{\\mathbf\{b\}\}\\triangleq\(\\mathsf\{M\}\_\{\\rho\_\{0\}\}\[\\mathbf\{b\}\]\)^\{\+\}\\mathsf\{T\}\[\\mathbf\{b\}\]\\in\\mathbb\{R\}^\{k\\times k\}and we multiply it repeatedly to the last state in the training data until the prediction fillsm=\(t\+T\)/sm=\(t\+T\)/stime steps to get the linearly reconstructed base vectors of horizon data,
𝐛L=\[𝖪𝐛ols𝐛\(𝐱nP\),\(𝖪𝐛ols\)2𝐛\(𝐱nP\),⋯,\(𝖪𝐛ols\)m−n𝐛\(𝐱nP\)\]\.\\mathbf\{b\}^\{L\}\\\!=\\\!\[\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{ols\}\\mathbf\{b\}\(\\mathbf\{x\}^\{P\}\_\{n\}\),\(\\\!\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{ols\}\\\!\)^\{2\}\\mathbf\{b\}\(\\mathbf\{x\}^\{P\}\_\{n\}\),\\cdots,\(\\\!\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{ols\}\\\!\)^\{m\\\!\-\\\!n\}\\mathbf\{b\}\(\\mathbf\{x\}^\{P\}\_\{n\}\)\]\.\(11\)Noted that𝐛L\\mathbf\{b\}^\{L\}has the same length\(m−n\)\(m\-n\)as the prediction, and we treat it as a preliminary estimate\. As shown in the ablation study in Section[5\.3](https://arxiv.org/html/2609.17815#S5.SS3.SSSx3), it is not sufficiently accurate, motivating the use of a Kalman filter\.
### 4\.2Second Stage: Forecast with Kalman Inference
#### Kalman Filter for State Space Evolution
As described above, in real\-world datasets, nonlinear residuals can cause misalignment between the actual base vector in the state space and its linear expansion𝐛L\\mathbf\{b\}^\{L\}\. This noise can arise from dimension reduction, nonlinear or non\-Markovian dynamics, random perturbations in the data, and errors in the linear regression estimate𝖪bols\\mathsf\{K\}\_\{b\}^\{\\mathrm\{ols\}\}\. Despite these nonlinear residuals, we posit that the learned Koopman space is sufficiently representative so that the residual components carry limited predictive information and can be mitigated as noise during inference\. Considering the recursive nature of time\-series prediction, the Kalman filter provides a lightweight, closed\-form inference procedure that minimizes estimation error under linear Gaussian dynamics\.
Fork≥n\+1k\\geq n\+1, its processing model is:
𝐛k=A𝐛k−1\+𝐰k,\\mathbf\{b\}\_\{k\}=A\\mathbf\{b\}\_\{k\-1\}\+\\mathbf\{w\}\_\{k\},\(12\)whereA∈d×dA\\in d\\times dis the transition matrix, which we initiate with𝖪bols\\mathsf\{K\}\_\{b\}^\{ols\}by definition\.𝐰k∼N\(𝟎,Q\)\\mathbf\{w\}\_\{k\}\\sim N\(\\mathbf\{0\},Q\)is the process noise andQQis its covariance matrix\. Accordingly, the observation model is:
𝐨k=H𝐛k\+𝐯k\.\\mathbf\{o\}\_\{k\}=H\\mathbf\{b\}\_\{k\}\+\\mathbf\{v\}\_\{k\}\.\(13\)HereH∈d×dH\\in d\\times dis the observation matrix, which we initiate withII\.𝐨k\\mathbf\{o\}\_\{k\}is the observation and we use the expanded linear part𝐛L\[k\]\\mathbf\{b\}^\{L\}\[k\]in \([11](https://arxiv.org/html/2609.17815#S4.E11)\), and𝐯k∼N\(𝟎,R\)\\mathbf\{v\}\_\{k\}\\sim N\(\\mathbf\{0\},R\)\. Then we conduct the prediction Step and update Step iteratively\. The prediction step can be formulated as:
𝐛^k\|k−1\\displaystyle\\hat\{\\mathbf\{b\}\}\_\{k\|k\-1\}=A𝐛^k−1\|k−1,\\displaystyle=A\\hat\{\\mathbf\{b\}\}\_\{k\-1\|k\-1\},\(14\)Σk\|k−1\\displaystyle\\Sigma\_\{k\|k\-1\}=AΣk−1\|k−1AT\+Q,\\displaystyle=A\\Sigma\_\{k\-1\|k\-1\}A^\{T\}\+Q,\(15\)where𝐛^k\|k−1\\hat\{\\mathbf\{b\}\}\_\{k\|k\-1\}is the predicted state andΣk\|k−1\\Sigma\_\{k\|k\-1\}is the predicted covariance matrix of the process uncertainty\. The update step then uses the Kalman gainGkG\_\{k\}to balance the prediction and observation, thereby refining the latent state:
Gk\\displaystyle G\_\{k\}=Σk\|k−1HT\(HΣk\|k−1HT\+R\)−1,\\displaystyle=\\Sigma\_\{k\|k\-1\}H^\{T\}\(H\\Sigma\_\{k\|k\-1\}H^\{T\}\+R\)^\{\-1\},\(16\)𝐛^k\|k\\displaystyle\\hat\{\\mathbf\{b\}\}\_\{k\|k\}=𝐛^k\|k−1\+Gk\(𝐨k−H𝐛^k\|k−1\),\\displaystyle=\\hat\{\\mathbf\{b\}\}\_\{k\|k\-1\}\+G\_\{k\}\(\\mathbf\{o\}\_\{k\}\-H\\hat\{\\mathbf\{b\}\}\_\{k\|k\-1\}\),\(17\)Σk\|k\\displaystyle\\Sigma\_\{k\|k\}=\(I−GkH\)Σk\|k−1\.\\displaystyle=\(I\-G\_\{k\}H\)\\Sigma\_\{k\|k\-1\}\.\(18\)
We note thatA,H,Q,RA,H,Q,Rare trainable parameters in the above process\. Iteratively, we obtain the refined predicted singular vectors𝐛^=\[𝐛^1,𝐛^2,⋯,𝐛^m\]\\hat\{\\mathbf\{b\}\}=\[\\hat\{\\mathbf\{b\}\}\_\{1\},\\hat\{\\mathbf\{b\}\}\_\{2\},\\cdots,\\hat\{\\mathbf\{b\}\}\_\{m\}\]\.
#### Decoder
Given the learned base functions𝐛\\mathbf\{b\}and an observablehh, we compute the projection coefficient𝐳𝐛ols≜\(𝖬ρ0\[𝐛\]\)†⟨h,𝐛⟩ρ0∈ℝk×N×s\\mathbf\{z\}^\{ols\}\_\{\\mathbf\{b\}\}\\triangleq\(\\mathsf\{M\}\_\{\\rho\_\{0\}\}\[\\mathbf\{b\}\]\)^\{\\dagger\}\\langle h,\\mathbf\{b\}\\rangle\_\{\\rho\_\{0\}\}\\in\\mathbb\{R\}^\{k\\times N\\times s\}by solving the linear regressionh\(𝐱\)≈𝐳⊤𝐛\(𝐱\)h\(\\mathbf\{x\}\)\\approx\\mathbf\{z\}^\{\\top\}\\mathbf\{b\}\(\\mathbf\{x\}\)\. This coefficient is fixed throughout the forecasting process\. In our case, we takeh\(𝐱\)=𝐱h\(\\mathbf\{x\}\)=\\mathbf\{x\}, so that the predicted data can be decoded from the learned base functions as follows:
X^P=𝐛^⊤𝐳𝐛ols\.\\hat\{\{X\}\}^\{P\}=\\hat\{\\mathbf\{b\}\}^\{\\top\}\\mathbf\{z\}^\{ols\}\_\{\\mathbf\{b\}\}\.\(19\)After prediction, we de\-patch the outputs to the predictionX^\\hat\{X\}and optimize an MSE\-based reconstruction loss, together with an additional KL\-divergence term on the filter output distribution𝐛^\\hat\{\\mathbf\{b\}\}\. This regularization encourages a well\-structured latent representation and helps preserve the orthogonality of the learned base functions, leading to the following loss:
ℒe=∥Y−X^∥22\+DKL\(ℚ\(𝐛^∣X\)∥ℙ\(𝐛∣X\)\),\\mathcal\{L\}\_\{e\}=\\\|Y\-\\hat\{X\}\\\|\_\{2\}^\{2\}\+D\_\{KL\}\\big\(\\mathbb\{Q\}\(\\hat\{\\mathbf\{b\}\}\\mid X\)\\,\\\|\\,\\mathbb\{P\}\(\\mathbf\{b\}\\mid X\)\\big\),\(20\)whereOPENℙ\(𝐛∣X\)\)=𝒩\(0,𝐈\)\\mathbb\{P\}\(\\mathbf\{b\}\\mid X\)\\big\)=\\mathcal\{N\}\(0,\\mathbf\{I\}\), andX^\\hat\{X\},𝐛^\\hat\{\\mathbf\{b\}\}are induced by the trainable Kalman filter\.
### 4\.3Overall Algorithm
We adopt a two\-stage training procedure for learning the Koopman representation\.
Train encoders with LoRA objective\.First, we independently train the Koopman module using the LoRA objective to capture the low\-rank dynamics and establish a stable latent space for linear evolution\. We then freeze the learned functions, assuming that the low\-rank structure has been adequately captured, and estimate the Koopman matrix\.
Dynamically updated matrix\.We estimate the Koopman operator𝖪𝐛\\mathsf\{K\}\_\{\\mathbf\{b\}\}via linear regression and initialize the transition matrixAAof the Kalman filter using𝖪𝐛ols\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{ols\}, averaged over 128 training batches\. Because the patched training data provide too few samples for reliable regression, we instead treat the finite\-dimensional Koopman matrix𝖪θ\\mathsf\{K\}\_\{\\theta\}as trainable and optimize it jointly in an end\-to\-end manner using the forecasting objectiveℒe\\mathcal\{L\}\_\{e\}\. The Koopman matrix is initialized using the static least\-squares estimate𝖪θ←𝖪𝐛ols\\mathsf\{K\}\_\{\\theta\}\\leftarrow\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}\. The proposed K2SVD preserves the principled operator\-consistent latent space learned by LoRA, while allowing the transition dynamics to adapt to the downstream multi\-step forecasting objective, mitigating the finite\-sample noise engendered in the regression\. Since𝖪θ∈ℝd×d\\mathsf\{K\}\_\{\\theta\}\\in\\mathbb\{R\}^\{d\\times d\}is low\-dimensional, making it trainable adds only a small computational and parameter overhead\.
## 5Experiments
In this section, we present empirical results that validate the effectiveness and performance of K2SVD\. We first examine the singular\-value spectra of the learned Koopman spaces to characterize the intrinsic rank structure of the datasets\. We evaluate forecasting on both low\- and high\-dimensional datasets, with an emphasis on computationally demanding high\-dimensional cases\. Finally, we report efficiency results and ablation studies to assess the contribution of each module\.
### 5\.1Experiment Setup
We evaluate K2SVD against state\-of\-the\-art \(SOTA\) baselines on real\-world and dynamic forecasting tasks\. To make the comparison clear, we organize the datasets by their dimensionality\.
Datasets\.We conduct experiments on nine forecasting datasets: eight real\-world datasets from ProbTS\([Zhang et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib30)\), a comprehensive benchmark for probabilistic forecasting; and one dynamical\-system dataset following the experimental setting of\([Jeong et al\. 2025](https://arxiv.org/html/2609.17815#bib.bib6)\)\. The high\-dimensional group consists of Electricity\-L, Weather\-L, and Ordered\-MNIST\. Electricity\-L contains the consumption of electricity in kWh, Weather\-L contains multivariate climatological measurements, and Ordered\-MNIST represents image\-based dynamic observations\. With tens or hundreds of dimensions, these datasets better assess whether the learned low\-rank Koopman space improves scalability and efficiency\. The low\-dimensional group includes ETTh1\-L, ETTh2\-L, ETTm1\-L, ETTm2\-L, Exchange\-L, and ILI\-L\. These datasets contain relatively few variables and evaluate whether K2SVD remains competitive when the observation space is already compact\. Detailed dataset statistics are reported in Table[1](https://arxiv.org/html/2609.17815#S5.T1)\.
The context and horizon length for the main results are displayed in table[1](https://arxiv.org/html/2609.17815#S5.T1)and additional results for longer term time series prediction are shown in the Appendix[D\.1](https://arxiv.org/html/2609.17815#A4.SS1)\.
Table 1:Statistics and Description of the nine datasets\.Dimension GroupDataset\#var\.rangefreq\.ContextHorizonDescriptionHigh\-dimensionalElectricity\-L321ℝ\+\\mathbb\{R\}^\{\+\}H9696Electricity consumption \(Kwh\)Ordered\-MNIST28\*28\(0,255\)\(0,255\)NA4848Periodic change of handwritten figuresWeather\-L21ℝ\+\\mathbb\{R\}^\{\+\}10min9696Local climatological dataLow\-dimensionalETTh1/h2\-L7ℝ\+\\mathbb\{R\}^\{\+\}H9696Electricity transformer temperature per hourETTm1/m2\-L7ℝ\+\\mathbb\{R\}^\{\+\}15min9696Electricity transformer temperature every 15 minExchange\-L8ℝ\+\\mathbb\{R\}^\{\+\}Busi\. Day9696Daily exchange rates of 8 countriesILI\-L7\(0,1\)\(0,1\)W3624Ratio of patients seen with influenza\-like illness
Baselines\.We compare K2SVD with seven baselines designed for both real\-world and dynamic system data\. These include four point\-forecasting models,K2K^\{2\}VAE\([Wu et al\. 2025](https://arxiv.org/html/2609.17815#bib.bib1)\), FITS\([Xu et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib16)\), iTransformer\([Liu et al\. 2024](https://arxiv.org/html/2609.17815#bib.bib21)\), and Koopa\([Liu et al\. 2023](https://arxiv.org/html/2609.17815#bib.bib17)\), as well as three generative models, TSDiff\([Kollovieh et al\. 2023](https://arxiv.org/html/2609.17815#bib.bib18)\), GRU NVP\([Rasul et al\. 2021](https://arxiv.org/html/2609.17815#bib.bib19)\), and CSDI\([Tashiro et al\. 2021](https://arxiv.org/html/2609.17815#bib.bib20)\), with impleameantion details provided in Appendix[C](https://arxiv.org/html/2609.17815#A3)\.
Evaluation Metrics\.We evaluate probabilistic forecasts using NRMSE \(Normalized Root Mean Square Error\) as defined below\. We also report model computational complexity and inference time in the efficiency section to highlight the benefits of the LoRA structure recovered by the proposed K2SVD algorithm\.
NRMSE=MSE1N∑i=1N\|Yi\|,withMSE=1N∑i=1N\(Yi−Y^i\)2\.\\mathrm\{NRMSE\}=\\frac\{\\sqrt\{\\mathrm\{MSE\}\}\}\{\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\|Y\_\{i\}\|\},\\ \\text\{with\}\\ \\mathrm\{MSE\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\(Y\_\{i\}\-\\hat\{Y\}\_\{i\}\\right\)^\{2\}\\\!\.
### 5\.2Main Results
Figure 3:The singular values of the Koopman space exhibit low\-rank properties on multiple datasets\.Empirical evidence for the LoRA property\.We observe low\-rank structure across both low\-dimensional and high\-dimensional datasets, a property that most existing models do not explicitly capture or exploit\. As shown by the Koopman\-space spectrum in Figure[3](https://arxiv.org/html/2609.17815#S5.F3), the singular values decay toward zero as the rank increases across all scenarios\. Most datasets exhibit this decay within roughly 10\-30 dimensions\. These observations motivate an adaptive choice of latent dynamic dimension, where we use dynamic dimension 16 for most datasets instead of the commonly chosen 128 in baselines\. The ablation study further shows that increasing the latent dimension has only a minor effect on precision but substantially increases computational cost\. One exception is ILI\-L, which requires approximately 60 Koopman dimensions before the singular values vanish, indicating a relatively high\-rank structure despite its low observation dimension\. Results for ILI\-L are reported in the Appendix[D\.3](https://arxiv.org/html/2609.17815#A4.SS3)\.
Table 2:Results of models using the NRMSE metric on the high\-dimensional datasets and Exchange\-L\. Best results are shown inbold, and second\-best results areunderlined\. ’\-’ indicates that the run took too long to finish\.Prediction performance\.The prediction results in Table[2](https://arxiv.org/html/2609.17815#S5.T2)show that K2SVD performs strongly on Electricity\-L, Weather\-L, Ordered\-MNIST, and Exchange\-L\. For fairness, a CNN encoder will replace all the encoders when models are applied on MNIST dataset\. As described in the dataset setup and summarized in Table[1](https://arxiv.org/html/2609.17815#S5.T1), the former three datasets form the high\-dimensional group, with observations consisting of many interacting variables or image\-like dynamic states\. K2SVD is well\-suited to these settings because it directly identifies a compact transition space from the underlying dynamics, reducing the redundancy and instability of less principled latent embeddings while mitigating the effect of noise through Kalman inference\.
For the low\-dimensional group, K2SVD also remains competitive, with full results reported in Appendix[D\.1](https://arxiv.org/html/2609.17815#A4.SS1)\. Exchange\-L is included because it is a low\-dimensional but challenging financial dataset characterized by weak periodicity and nonstationary dynamics\. Together, our results show that the principled Koopman transition space is beneficial across different observation scales, with the efficiency advantage to be evaluated in the next section\.
Table 3:Comparison of inference time in seconds, and GPU usage in GBs on the high\-dimensional datasets\. For each metric, the best results are shown inbold, and the second\-best results areunderlined\. “\-” indicates that the run took too long to finish\.Efficiency and compactness\.From an efficiency perspective, K2SVD establishes SOTA performance in both inference time and GPU memory usage among the evaluated methods, with its advantage becoming particularly pronounced on high\-dimensional datasets\. All models are evaluated on the same GPU under an identical software environment, and the results are reported in Table[3](https://arxiv.org/html/2609.17815#S5.T3)\. K2SVD achieves the fastest inference time on all three datasets\. On the two most computationally demanding benchmarks, Electricity and MNIST, it is approximately three times faster than every competing baseline\. In terms of memory usage, K2SVD requires less than one\-third of the peak GPU memory allocated by the most efficient baseline models\. Notably, K2SVD substantially outperforms FITS in both inference speed and memory usage, despite FITS being designed as a low\-cost frequency\-domain model\. This comparison demonstrates that the efficiency of K2SVD arises from its low\-rank transition space and from concentrating iterative prediction within this space\. Consequently, K2SVD achieves both superior predictive accuracy and substantially lower inference cost, yielding a more favorable accuracy\-efficiency trade\-off\.
### 5\.3Ablation Study
#### Ablation on the Dimension of Dynamic Space
We conduct an ablation study to examine the sufficiency of the latent Koopman dimension\. Specifically, we vary the rank of the learned Koopman subspace from 16 to 64 to investigate whether reducing the rank harms performance\. As shown in Table[4](https://arxiv.org/html/2609.17815#S5.T4), performance in prediction precision does not improve significantly yet sometimes drops as the dimension increases, while the total computation measured as FLOPs \(floating\-point operations per second\) booms\. Although inference time does not visibly increase, we infer that it is limited by memory bandwidth rather than compute capacity, as the models require no more than 1 GFLOP while the GPU provides over 100 TFLOPS of compute\. These results suggest that the dominant system dynamics can be effectively captured in a compact, low\-rank Koopman subspace, making higher\-dimensional representations redundant\.
Table 4:Ablation study on the Koopman latent dimension\. Low\-dimensional latent spaces achieve competitive accuracy, indicating redundancy in higher\-rank representations\.
#### Ablation on LoRA Objective
We argue that optimizing the encoder with the LoRA objective provides an efficient and expressive Koopman space\. In comparison to end\-to\-end methods, the encoder trained with the HS objective and frozen in later steps creates a more solid and precise LoRA space, and as the inference length increases it better resists noise and results in more accurate reconstruction\. To support this claim, we set the contextt=96t=96, horizonT=720T=720, remove stage 1 and directly train the encoder, Koopman matrix𝖪θ\\mathsf\{K\}\_\{\\theta\}, and Kalman parameters in an end\-to\-end theme with onlyℒe\\mathcal\{L\}\_\{e\}\. As shown in Table[5](https://arxiv.org/html/2609.17815#S5.T5), the training theme with LoRA objective achieves better performance under the same hidden\-space dimension on the datasets MNIST, Electricity and Exchange, and comparable results under weather\.
Table 5:Ablation study on LoRA objective shows that the first training stage has advantages in prediction performance\.
#### Ablation on Kalman inference
As we argued in section[4\.2](https://arxiv.org/html/2609.17815#S4.SS2), the output of the Koopman module, the linear rollout of the base vector in prediction spaces, contains noise from various sources, and we applied Kalman inference to handle it\. To justify the necessity, we use the linear rollout as the predicted base vectors and directly decode it to get the prediction in data space\. The results show different degrees of decline in performance, which reveals the noise in linear prediction and emphasizes the importance of Kalman inference\.
Table 6:Ablation study on Kalman inference shows the noisy property of linear prediction and the advantage of filtering\.
## 6Conclusion
This work revisits time\-series forecasting from the perspective of operator structure rather than only predictive capacity\. The central observation is that many forecasting benchmarks contain a compact transition structure that can be exposed through the spectrum of the learned Koopman space\. By building the forecasting model around this structure, K2SVD reduces the latent dynamic dimension substantially while retaining competitive prediction accuracy and achieving SOTA inference efficiency, strongly exceeds baseline models\.
The experiments also show results with clear distinction in intrinsically different datasets, revealing information about the data and provides visible judge through the learned spectrum rather than empirical observations\.
More broadly, these results suggest that efficient forecasting should not only compress model parameters, but also identify which part of the temporal dynamics is worth modeling explicitly\. Future work can extend this direction by allowing the Koopman rank and transition operator to adapt across time, and by developing uncertainty mechanisms that separate random perturbations from genuine changes in the dynamics\.
## References
- Ansariet al\.\(2024\)A\. F\. Ansari, L\. Stella, C\. Turkmen, X\. Zhang, P\. Mercado, H\. Shen, O\. Shchur, S\. S\. Rangapuram, S\. P\. Arango, S\. Kapoor, J\. Zschiegner, D\. C\. Maddix, H\. Wang, M\. W\. Mahoney, K\. Torkkola, A\. G\. Wilson, M\. Bohlke\-Schneider, and Y\. WangChronos: Learning the Language of Time Series\.External Links:2403\.07815,[Link](https://arxiv.org/abs/2403.07815)Cited by:[§1](https://arxiv.org/html/2609.17815#S1.p1.1)\.
- Daset al\.\(2024\)A\. Das, W\. Kong, R\. Sen, and Y\. ZhouA Decoder\-Only Foundation Model for Time\-Series Forecasting\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 10148–10167\.External Links:[Link](https://proceedings.mlr.press/v235/das24c.html)Cited by:[§1](https://arxiv.org/html/2609.17815#S1.p1.1)\.
- Huanget al\.\(2024\)P\. Huang, K\. Zheng, and G\. P\. FettweisData\-Driven Koopman Operator\-Based Error\-State Kalman Filter for Enhanced State Estimation of Quadrotors in Agile Flight\.In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems \(IROS\),pp\. 10983–10989\.External Links:[Document](https://dx.doi.org/10.1109/IROS58592.2024.10802457)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p3.1)\.
- Jeonget al\.\(2025\)M\. Jeong, J\. J\. Ryu, S\. Yun, and G\. W\. WornellEfficient Parametric SVD of Koopman Operator for Stochastic Dynamical Systems\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=kL2pnzClyD)Cited by:[Appendix A](https://arxiv.org/html/2609.17815#A1.p1.1),[§1](https://arxiv.org/html/2609.17815#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.17815#S4.SS1.SSSx2.p1.2),[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p2.1)\.
- Kolloviehet al\.\(2023\)M\. Kollovieh, A\. F\. Ansari, M\. Bohlke\-Schneider, J\. Zschiegner, H\. Wang, and Y\. WangPredict, Refine, Synthesize: Self\-Guiding Diffusion Models for Probabilistic Time Series Forecasting\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 28341–28364\.Cited by:[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p4.1)\.
- Kosticet al\.\(2024\)V\. R\. Kostic, P\. Novelli, R\. Grazzi, K\. Lounici, and M\. PontilLearning Invariant Representations of Time\-Homogeneous Stochastic Dynamical Systems\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=twSnZwiOIm)Cited by:[§1](https://arxiv.org/html/2609.17815#S1.p2.1),[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Lan and Mezić \(2013\)Y\. Lan and I\. MezićLinearization in the Large of Nonlinear Systems and Koopman Operator Spectrum\.Physica D: Nonlinear Phenomena242\(1\),pp\. 42–53\.External Links:[Document](https://dx.doi.org/10.1016/j.physd.2012.08.017)Cited by:[§1](https://arxiv.org/html/2609.17815#S1.p2.1)\.
- Liet al\.\(2019\)S\. Li, X\. Jin, Y\. Xuan, X\. Zhou, W\. Chen, Y\. Wang, and X\. YanEnhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§1](https://arxiv.org/html/2609.17815#S1.p1.1)\.
- Limet al\.\(2021\)B\. Lim, S\. Ö\. Arık, N\. Loeff, and T\. PfisterTemporal Fusion Transformers for Interpretable Multi\-Horizon Time Series Forecasting\.International Journal of Forecasting37\(4\),pp\. 1748–1764\.External Links:[Document](https://dx.doi.org/10.1016/j.ijforecast.2021.03.012)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p1.1)\.
- Liuet al\.\(2026\)N\. Liu, S\. Liu, X\. T\. Tong, and L\. JiangEstimate of Koopman Modes and Eigenvalues with Kalman Filter\.SIAM Journal on Scientific Computing48\(2\),pp\. A512–A539\.External Links:[Document](https://dx.doi.org/10.1137/24M1684207)Cited by:[§1](https://arxiv.org/html/2609.17815#S1.p3.1)\.
- Liuet al\.\(2024\)Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. LongiTransformer: Inverted Transformers Are Effective for Time Series Forecasting\.InInternational Conference on Learning Representations,Cited by:[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p4.1)\.
- Liuet al\.\(2023\)Y\. Liu, C\. Li, J\. Wang, and M\. LongKoopa: Learning Non\-Stationary Time Series Dynamics with Koopman Predictors\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 12271–12290\.Cited by:[§D\.3](https://arxiv.org/html/2609.17815#A4.SS3.p2.1),[§1](https://arxiv.org/html/2609.17815#S1.p3.1),[§2](https://arxiv.org/html/2609.17815#S2.p3.1),[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p4.1)\.
- Mardtet al\.\(2018\)A\. Mardt, L\. Pasquali, H\. Wu, and F\. NoéVAMPnets for Deep Learning of Molecular Kinetics\.Nature Communications9\(1\),pp\. 5\.External Links:[Document](https://dx.doi.org/10.1038/s41467-017-02388-1)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Mezić \(2005\)I\. MezićSpectral Properties of Dynamical Systems, Model Reduction and Decompositions\.Nonlinear Dynamics41,pp\. 309–325\.External Links:[Document](https://dx.doi.org/10.1007/s11071-005-2824-x)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Mezić \(2013\)I\. MezićAnalysis of Fluid Flows via Spectral Properties of the Koopman Operator\.Annual Review of Fluid Mechanics45,pp\. 357–378\.External Links:[Document](https://dx.doi.org/10.1146/annurev-fluid-011212-140652)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Naimanet al\.\(2023\)I\. Naiman, N\. B\. Erichson, P\. Ren, M\. W\. Mahoney, and O\. AzencotGenerative Modeling of Regular and Irregular Time Series Data via Koopman VAEs\.External Links:2310\.02619,[Link](https://arxiv.org/abs/2310.02619)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p3.1)\.
- Oreshkinet al\.\(2020\)B\. N\. Oreshkin, D\. Carpov, N\. Chapados, and Y\. BengioN\-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p1.1)\.
- Rasulet al\.\(2021\)K\. Rasul, A\. Sheikh, I\. Schuster, U\. Bergmann, and R\. VollgrafMultivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p1.1),[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p4.1)\.
- Rowleyet al\.\(2009\)C\. W\. Rowley, I\. Mezić, S\. Bagheri, P\. Schlatter, and D\. S\. HenningsonSpectral Analysis of Nonlinear Flows\.Journal of Fluid Mechanics641,pp\. 115–127\.External Links:[Document](https://dx.doi.org/10.1017/S0022112009992059)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Salinaset al\.\(2020\)D\. Salinas, V\. Flunkert, J\. Gasthaus, and T\. JanuschowskiDeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks\.International Journal of Forecasting36\(3\),pp\. 1181–1191\.External Links:[Document](https://dx.doi.org/10.1016/j.ijforecast.2019.07.001)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p1.1)\.
- Schmid \(2010\)P\. J\. SchmidDynamic Mode Decomposition of Numerical and Experimental Data\.Journal of Fluid Mechanics656,pp\. 5–28\.External Links:[Document](https://dx.doi.org/10.1017/S0022112010001217)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Tashiroet al\.\(2021\)Y\. Tashiro, J\. Song, Y\. Song, and S\. ErmonCSDI: Conditional Score\-Based Diffusion Models for Probabilistic Time Series Imputation\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 24804–24816\.Cited by:[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p4.1)\.
- Tu \(2013\)J\. H\. TuDynamic Mode Decomposition: Theory and Applications\.Ph\.D\. Thesis,Princeton University\.Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Wanget al\.\(2022\)R\. Wang, Y\. Dong, S\. Ö\. Arik, and R\. YuKoopman Neural Forecaster for Time Series with Temporal Distribution Shifts\.External Links:2210\.03675,[Link](https://arxiv.org/abs/2210.03675)Cited by:[§D\.3](https://arxiv.org/html/2609.17815#A4.SS3.p2.1),[§2](https://arxiv.org/html/2609.17815#S2.p3.1)\.
- Williamset al\.\(2015a\)M\. O\. Williams, I\. G\. Kevrekidis, and C\. W\. RowleyA Data\-Driven Approximation of the Koopman Operator: Extending Dynamic Mode Decomposition\.Journal of Nonlinear Science25\(6\),pp\. 1307–1346\.External Links:[Document](https://dx.doi.org/10.1007/s00332-015-9258-5)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Williamset al\.\(2015b\)M\. O\. Williams, C\. W\. Rowley, and I\. G\. KevrekidisA Kernel\-Based Method for Data\-Driven Koopman Spectral Analysis\.Journal of Computational Dynamics2\(2\),pp\. 247–265\.External Links:[Document](https://dx.doi.org/10.3934/jcd.2015005)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Wu and Noé \(2020\)H\. Wu and F\. NoéVariational Approach for Learning Markov Processes from Time Series Data\.Journal of Nonlinear Science30\(1\),pp\. 23–66\.External Links:[Document](https://dx.doi.org/10.1007/s00332-019-09567-y)Cited by:[§1](https://arxiv.org/html/2609.17815#S1.p2.1),[§2](https://arxiv.org/html/2609.17815#S2.p2.1)\.
- Wuet al\.\(2025\)X\. Wu, X\. Qiu, H\. Gao, J\. Hu, B\. Yang, and C\. GuoK2K^\{2\}VAE: A Koopman\-Kalman Enhanced Variational AutoEncoder for Probabilistic Time Series Forecasting\.External Links:2505\.23017,[Link](https://arxiv.org/abs/2505.23017)Cited by:[Appendix C](https://arxiv.org/html/2609.17815#A3.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.17815#S2.p3.1),[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p4.1)\.
- Xuet al\.\(2024\)Z\. Xu, A\. Zeng, and Q\. XuFITS: Modeling Time Series with10k10kParameters\.InInternational Conference on Learning Representations,Cited by:[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p4.1)\.
- Zhanget al\.\(2024\)J\. Zhang, X\. Wen, Z\. Zhang, S\. Zheng, J\. Li, and J\. BianProbTS: Benchmarking Point and Distributional Forecasting across Diverse Prediction Horizons\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 48045–48082\.Cited by:[§5\.1](https://arxiv.org/html/2609.17815#S5.SS1.p2.1)\.
- Zhouet al\.\(2021\)H\. Zhou, S\. Zhang, J\. Peng, S\. Zhang, J\. Li, H\. Xiong, and W\. ZhangInformer: Beyond Efficient Transformer for Long Sequence Time\-Series Forecasting\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 11106–11115\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v35i12.17325)Cited by:[§2](https://arxiv.org/html/2609.17815#S2.p1.1)\.
## Appendix AProof to the LoRA Loss
Here, we provide the proof to justify the LoRA loss proposed in\([Jeong et al\. 2025](https://arxiv.org/html/2609.17815#bib.bib6)\)\.
Recall thatℒlora\(𝐟,𝐠\)≜−2tr\(T\[𝐟,𝐠\]\)\+tr\(Mρ0\[𝐟\]Mρ1\[𝐠\]\)\\mathcal\{L\}\_\{\\text\{lora\}\}\(\\mathbf\{f\},\\mathbf\{g\}\)\\triangleq\-2\\,\\text\{tr\}\(\\text\{T\}\[\\mathbf\{f\},\\mathbf\{g\}\]\)\+\\text\{tr\}\(\\text\{M\}\_\{\\rho\_\{0\}\}\[\\mathbf\{f\}\]\\text\{M\}\_\{\\rho\_\{1\}\}\[\\mathbf\{g\}\]\)\. To prove, first note that
‖𝒦−∑i=1kfi⊗gi‖HS2−∥𝒦∥HS2=−2∑i=1k⟨fi,𝒦gi⟩ρ0\+∑i=1k∑j=1k⟨fi,fj⟩ρ0⟨gi,gj⟩ρ1\.\\displaystyle\\left\\\|\\mathcal\{K\}\-\\sum\_\{i=1\}^\{k\}f\_\{i\}\\otimes g\_\{i\}\\right\\\|\_\{\\text\{HS\}\}^\{2\}\-\\\|\\mathcal\{K\}\\\|\_\{\\text\{HS\}\}^\{2\}=\-2\\sum\_\{i=1\}^\{k\}\\langle f\_\{i\},\\mathcal\{K\}g\_\{i\}\\rangle\_\{\\rho\_\{0\}\}\+\\sum\_\{i=1\}^\{k\}\\sum\_\{j=1\}^\{k\}\\langle f\_\{i\},f\_\{j\}\\rangle\_\{\\rho\_\{0\}\}\\langle g\_\{i\},g\_\{j\}\\rangle\_\{\\rho\_\{1\}\}\.\(21\)
Here, the first term can be rewritten as
∑i=1k⟨fi,𝒦gi⟩ρ0\\displaystyle\\sum\_\{i=1\}^\{k\}\\langle f\_\{i\},\\mathcal\{K\}g\_\{i\}\\rangle\_\{\\rho\_\{0\}\}=∑i=1k𝔼ρ0\(𝐱\)p\(𝐱′\|𝐱\)\[fi\(𝐱\)gi\(𝐱′\)\]\\displaystyle=\\sum\_\{i=1\}^\{k\}\\mathbb\{E\}\_\{\\rho\_\{0\}\(\\mathbf\{x\}\)p\(\\mathbf\{x\}^\{\\prime\}\|\\mathbf\{x\}\)\}\\bigl\[f\_\{i\}\(\\mathbf\{x\}\)g\_\{i\}\(\\mathbf\{x\}^\{\\prime\}\)\\bigr\]\(22\)=tr\(𝔼ρ0\(𝐱\)p\(𝐱′\|𝐱\)\[𝐟\(𝐱\)𝐠\(𝐱′\)⊤\]\)=tr\(T\[𝐟,𝐠\]\)\.\\displaystyle=\\text\{tr\}\\\!\\left\(\\mathbb\{E\}\_\{\\rho\_\{0\}\(\\mathbf\{x\}\)p\(\\mathbf\{x\}^\{\\prime\}\|\\mathbf\{x\}\)\}\\bigl\[\\mathbf\{f\}\(\\mathbf\{x\}\)\\mathbf\{g\}\(\\mathbf\{x\}^\{\\prime\}\)^\{\\top\}\\bigr\]\\right\)=\\text\{tr\}\(\\text\{T\}\[\\mathbf\{f\},\\mathbf\{g\}\]\)\.\(23\)
Likewise, the second term is
∑i=1k∑j=1k⟨fi,fj⟩ρ0⟨gi,gj⟩ρ1\\displaystyle\\sum\_\{i=1\}^\{k\}\\sum\_\{j=1\}^\{k\}\\langle f\_\{i\},f\_\{j\}\\rangle\_\{\\rho\_\{0\}\}\\langle g\_\{i\},g\_\{j\}\\rangle\_\{\\rho\_\{1\}\}=tr\(𝔼ρ0\(𝐱\)\[𝐟\(𝐱\)𝐟\(𝐱\)⊤\]𝔼ρ1\(𝐱′\)\[𝐠\(𝐱′\)𝐠\(𝐱′\)⊤\]\)\\displaystyle=\\text\{tr\}\\\!\\left\(\\mathbb\{E\}\_\{\\rho\_\{0\}\(\\mathbf\{x\}\)\}\\bigl\[\\mathbf\{f\}\(\\mathbf\{x\}\)\\mathbf\{f\}\(\\mathbf\{x\}\)^\{\\top\}\\bigr\]\\mathbb\{E\}\_\{\\rho\_\{1\}\(\\mathbf\{x\}^\{\\prime\}\)\}\\bigl\[\\mathbf\{g\}\(\\mathbf\{x\}^\{\\prime\}\)\\mathbf\{g\}\(\\mathbf\{x\}^\{\\prime\}\)^\{\\top\}\\bigr\]\\right\)\(24\)=tr\(Mρ0\[𝐟\]Mρ1\[𝐠\]\)\.\\displaystyle=\\text\{tr\}\(\\text\{M\}\_\{\\rho\_\{0\}\}\[\\mathbf\{f\}\]\\text\{M\}\_\{\\rho\_\{1\}\}\[\\mathbf\{g\}\]\)\.\(25\)
This concludes the proof\.
## Appendix BAlgorithm
Algorithm 1Two\-Stage K2SVD Training and ForecastingInput:Training pairs
\(X,Y\)\(X,Y\); patch length
ss; Koopman rank
dd; variant
ν∈\{static,dynamic\}\\nu\\in\\\{\\text\{static\},\\text\{dynamic\}\\\}
Output:Encoders
𝐟,𝐠\\mathbf\{f\},\\mathbf\{g\}; Koopman matrix
𝖪\\mathsf\{K\}; Kalman parameters
A,H,Q,RA,H,Q,R; decoder coefficients
𝐳𝐛ols\\mathbf\{z\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}
Stage 1: Learn a low\-rank Koopman space;
for*each training batchXX*do
Split
XXinto non\-overlapping patches
XP=\[𝐱1P,…,𝐱nP\]X^\{P\}=\[\\mathbf\{x\}\_\{1\}^\{P\},\\ldots,\\mathbf\{x\}\_\{n\}^\{P\}\];
Compute shifted Koopman features
𝐟~=\[𝐟\(𝐱2P\),…,𝐟\(𝐱nP\)\]\\tilde\{\\mathbf\{f\}\}=\[\\mathbf\{f\}\(\\mathbf\{x\}\_\{2\}^\{P\}\),\\ldots,\\mathbf\{f\}\(\\mathbf\{x\}\_\{n\}^\{P\}\)\]and
𝐠~=\[𝐠\(𝐱1P\),…,𝐠\(𝐱n−1P\)\]\\tilde\{\\mathbf\{g\}\}=\[\\mathbf\{g\}\(\\mathbf\{x\}\_\{1\}^\{P\}\),\\ldots,\\mathbf\{g\}\(\\mathbf\{x\}\_\{n\-1\}^\{P\}\)\];
Evaluate
ℒLoRA←−2tr\(𝖳\[𝐟~,𝐠~\]\)\+tr\(𝖬ρ0\[𝐟~\]𝖬ρ1\[𝐠~\]\)\\mathcal\{L\}\_\{\\mathrm\{LoRA\}\}\\leftarrow\-2\\,\\mathrm\{tr\}\(\\mathsf\{T\}\[\\tilde\{\\mathbf\{f\}\},\\tilde\{\\mathbf\{g\}\}\]\)\+\\mathrm\{tr\}\(\\mathsf\{M\}\_\{\\rho\_\{0\}\}\[\\tilde\{\\mathbf\{f\}\}\]\\mathsf\{M\}\_\{\\rho\_\{1\}\}\[\\tilde\{\\mathbf\{g\}\}\]\);
Update
𝐟,𝐠\\mathbf\{f\},\\mathbf\{g\}by minimizing
ℒLoRA\\mathcal\{L\}\_\{\\mathrm\{LoRA\}\};
Initialize finite\-dimensional dynamics;
Choose the learned basis
𝐛∈\{𝐟,𝐠\}\\mathbf\{b\}\\in\\\{\\mathbf\{f\},\\mathbf\{g\}\\\}used for forecasting;
Estimate the regression\-induced Koopman matrix
𝖪𝐛ols←\(𝖬^ρ0\[𝐛\]\)†𝖳^\[𝐛\]\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}\\leftarrow\\bigl\(\\hat\{\\mathsf\{M\}\}\_\{\\rho\_\{0\}\}\[\\mathbf\{b\}\]\\bigr\)^\{\\dagger\}\\hat\{\\mathsf\{T\}\}\[\\mathbf\{b\}\]using training batches;
Initialize
A←𝖪𝐛olsA\\leftarrow\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\},
H←IH\\leftarrow I, and noise covariances
Q,RQ,R;
if*ν=dynamic\\nu=\\text\{dynamic\}*then
Initialize trainable
𝖪θ←𝖪𝐛ols\\mathsf\{K\}\_\{\\theta\}\\leftarrow\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}and set
𝖪←𝖪θ\\mathsf\{K\}\\leftarrow\\mathsf\{K\}\_\{\\theta\};
else
Set
𝖪←𝖪𝐛ols\\mathsf\{K\}\\leftarrow\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}and keep
𝖪\\mathsf\{K\}fixed;
Compute decoder coefficients
𝐳𝐛ols←\(𝖬ρ0\[𝐛\]\)†⟨h,𝐛⟩ρ0\\mathbf\{z\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}\\leftarrow\(\\mathsf\{M\}\_\{\\rho\_\{0\}\}\[\\mathbf\{b\}\]\)^\{\\dagger\}\\langle h,\\mathbf\{b\}\\rangle\_\{\\rho\_\{0\}\}with
h\(𝐱\)=𝐱h\(\\mathbf\{x\}\)=\\mathbf\{x\};
Stage 2: Train Kalman forecasting module;
Freeze
𝐟,𝐠\\mathbf\{f\},\\mathbf\{g\};
for*each training batch\(X,Y\)\(X,Y\)*do
Encode the final context patch to obtain the initial latent state
𝐛^n\|n\\hat\{\\mathbf\{b\}\}\_\{n\|n\};
Roll out the regression\-induced observation sequence
𝐛L\\mathbf\{b\}^\{L\}using
𝖪𝐛ols\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\};
for*each forecast patch indexk=n\+1,…,n\+mk=n\+1,\\ldots,n\+m*do
Predict latent state and covariance:
𝐛^k\|k−1←A𝐛^k−1\|k−1\\hat\{\\mathbf\{b\}\}\_\{k\|k\-1\}\\leftarrow A\\hat\{\\mathbf\{b\}\}\_\{k\-1\|k\-1\},
Σk\|k−1←AΣk−1\|k−1A⊤\+Q\\Sigma\_\{k\|k\-1\}\\leftarrow A\\Sigma\_\{k\-1\|k\-1\}A^\{\\top\}\+Q;
Set observation
𝐨k←𝐛L\[k\]\\mathbf\{o\}\_\{k\}\\leftarrow\\mathbf\{b\}^\{L\}\[k\]and compute Kalman gain
Gk←Σk\|k−1H⊤\(HΣk\|k−1H⊤\+R\)−1G\_\{k\}\\leftarrow\\Sigma\_\{k\|k\-1\}H^\{\\top\}\(H\\Sigma\_\{k\|k\-1\}H^\{\\top\}\+R\)^\{\-1\};
Update latent state and covariance:
𝐛^k\|k←𝐛^k\|k−1\+Gk\(𝐨k−H𝐛^k\|k−1\)\\hat\{\\mathbf\{b\}\}\_\{k\|k\}\\leftarrow\\hat\{\\mathbf\{b\}\}\_\{k\|k\-1\}\+G\_\{k\}\(\\mathbf\{o\}\_\{k\}\-H\\hat\{\\mathbf\{b\}\}\_\{k\|k\-1\}\),
Σk\|k←\(I−GkH\)Σk\|k−1\\Sigma\_\{k\|k\}\\leftarrow\(I\-G\_\{k\}H\)\\Sigma\_\{k\|k\-1\};
Decode patch forecasts
X^P←𝐛^⊤𝐳𝐛ols\\hat\{X\}^\{P\}\\leftarrow\\hat\{\\mathbf\{b\}\}^\{\\top\}\\mathbf\{z\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}and de\-patch to obtain
X^\\hat\{X\};
Compute
ℒe←∥Y−X^∥22\+DKL\(ℚ\(𝐛^∣X\)∥ℙ\(𝐛∣X\)\)\\mathcal\{L\}\_\{e\}\\leftarrow\\\|Y\-\\hat\{X\}\\\|\_\{2\}^\{2\}\+D\_\{KL\}\\\!\\left\(\\mathbb\{Q\}\(\\hat\{\\mathbf\{b\}\}\\mid X\)\\,\\\|\\,\\mathbb\{P\}\(\\mathbf\{b\}\\mid X\)\\right\);
if*ν=dynamic\\nu=\\text\{dynamic\}*then
Jointly update
𝖪θ\\mathsf\{K\}\_\{\\theta\}and
A,H,Q,RA,H,Q,Rusing
ℒe\\mathcal\{L\}\_\{e\};
else
Update
A,H,Q,RA,H,Q,Rby minimizing
ℒe\\mathcal\{L\}\_\{e\};
Algorithm[1](https://arxiv.org/html/2609.17815#algorithm1)summarizes the two\-stage training procedure described in the method section\. The first stage learns an operator\-consistent low\-rank Koopman space using the LoRA objective\. The second stage freezes this learned Koopman basis and trains the Kalman forecasting module\. The static variant K2SVD\-S keeps the regression\-induced Koopman matrix fixed, while K2SVD further updates the finite\-dimensional Koopman matrix during downstream forecasting training\.
## Appendix CExperimental Settings
All experiments are conducted on a single GPU with a global random seed of \(1, 2, 42, 123, 2026\) to ensure reproducibility and for constructing error bar\. The model is trained for up to 50 epochs, with each epoch consisting of 100 training batches\. Detailed setting to bu found in code\. We use a gradient clipping threshold of 0\.5 to ensure training stability and set the learning rate to1×10−31\\times 10^\{\-3\}\. Model checkpoints and logs are saved to a local results directory, and validation is performed at the end of every epoch\.
##### Two\-Stage Training\.
We employ a two\-stage training curriculum managed by a dedicated callback\. The first stage spans 15 epochs, during which the model undergoes a warm\-up phase to effectively initialize the Koopman operator\. The remaining 35 epochs constitute the second stage, in which the full model is optimized end\-to\-end\.
##### Model Configuration\.
The data preprocessing follows the K2VAE architecture\([Wu et al\. 2025](https://arxiv.org/html/2609.17815#bib.bib1)\), as input sequence is divided into non\-overlapping patches of length 24\. The Koopman latent space has a dynamic dimension of 16, initialized using a dynamic Koopman scheme, with multistep prediction enabled\. The encoder and decoder each consist of 3 hidden layers of width 256\. The VAE regularization weight is set toβ=0\.01\\beta=0\.01, imposing a light KL penalty to preserve the expressiveness of the latent space\. During inference, the model draws 100 samples per input to construct the predictive distribution, which is summarized over 20 quantiles\. The sampling schedule runs for 20 steps\.
##### Data Pipeline\.
Training and evaluation use batch sizes of 64 and 32, respectively, with 8 parallel data\-loading workers\. All input features are normalized using standard scaling \(zero mean, unit variance\) prior to model input\. A held\-out validation split is partitioned from the training set to monitor generalization during training\.
## Appendix DAdditional Empirical Results
### D\.1Additional Results on Long\-Term Time Series Prediction
Tables[8](https://arxiv.org/html/2609.17815#A4.T8)and[9](https://arxiv.org/html/2609.17815#A4.T9)summarize the NRMSE results for prediction lengths 96 and 720\. Table[8](https://arxiv.org/html/2609.17815#A4.T8)reports the low\-dimensional datasets, including Exchange\-L and the ETT series\. These datasets have compact observation spaces, so the main question is whether the learned Koopman transition remains stable when the prediction horizon increases\. K2SVD performs competitively across this group and obtains the best results on Exchange\-L for both prediction lengths, showing that the learned operator\-induced transition space is useful even when the raw data dimension is small\. On the ETT datasets, K2SVD is strongest on ETTh1\-L at prediction length 96 and on ETTh1\-L, ETTh2\-L, and ETTm2\-L at prediction length 720, whileK2K^\{2\}VAE and Koopa remain strong competitors on several shorter\-horizon or lower\-variance cases\.
Table[9](https://arxiv.org/html/2609.17815#A4.T9)reports the high\-dimensional datasets, including Electricity\-L, Ordered\-MNIST, and Weather\-L\. These datasets are the most relevant for evaluating the computational motivation of K2SVD, because the observations contain many coupled variables or image\-like dynamic states\. K2SVD achieves the best results on Electricity\-L and Ordered\-MNIST at prediction length 96 and remains competitive at prediction length 720\. On Weather\-L,K2K^\{2\}VAE attains the best NRMSE, while K2SVD remains close to the strongest baselines\. Together with the efficiency results in the main text, these results suggest that K2SVD provides a favorable accuracy\-efficiency trade\-off: it is not uniformly best on every dataset, but it is especially effective when a compact Koopman transition captures the dominant dynamics\.
### D\.2Additional Results on Long Prediction Efficiency
We further evaluate long\-horizon inference efficiency to examine whether the computational advantage of K2SVD persists when prediction requires many recursive rollout steps\. The comparison includes FITS, iTransformer, Koopa, andK2K^\{2\}VAE because they represent the main classes of efficiency\-competitive baselines used in this work\. This set therefore compares K2SVD against both general\-purpose long\-sequence predictors and methods that are structurally closest to it\.
Table[7](https://arxiv.org/html/2609.17815#A4.T7)shows that the advantage of K2SVD becomes more pronounced under long prediction lengths\. K2SVD achieves the fastest inference on all three datasets, reducing inference time by about2\.4×2\.4\\times–5\.5×5\.5\\timescompared with the strongest non\-K2SVD baseline\. Its GPU usage is also consistently the lowest, especially on MNIST, where the compact Koopman transition avoids carrying the high\-dimensional observation space through each rollout step\. The FLOPs comparison further highlights the source of this efficiency: after the encoder maps observations into the learned low\-rank Koopman space, long\-range prediction is performed by lightweight linear state evolution and Kalman inference rather than repeated high\-dimensional nonlinear transformations\. These results support the central design of K2SVD: by explicitly learning a compact operator\-induced transition space, the model improves not only predictive accuracy but also the time, memory, and arithmetic cost of long\-horizon forecasting\.
Table 7:Comparison of inference time in seconds, GPU usage in GBs, and quantification of computation by meta FLOPs on the selected high and low dimensional datasets when inference length isT=480T=480for MNIST andT=720T=720for the rest\. For each metric, the best results are shown inbold, and the second\-best results areunderlined\. “\-” indicates that the run took too long to finish\.Table 8:Results of models using the NRMSE metric on low\-dimensional datasets with prediction lengths 96 and 720\. Best results are shown inbold, and second\-best results areunderlined\. ’\*’ indicates that the model encounters invalid computation due to instability\.Table 9:Results of models using the NRMSE metric on high\-dimensional datasets with prediction lengths 96 and 720\. Best results are shown inbold, and second\-best results areunderlined\. ’\-’ indicates that the model takes too long to finish training\.
### D\.3Discussion of High\-rank Dataset
Although ILI\-L is low\-dimensional in its observation space, Figure[4](https://arxiv.org/html/2609.17815#A4.F4)shows that its Koopman spectrum decays more slowly than the other datasets\. We therefore report it separately as a high\-rank case\. Table[10](https://arxiv.org/html/2609.17815#A4.T10)shows that K2SVD remains comparable but does not dominate this dataset\. This behavior is consistent with the spectral evidence: when the effective Koopman rank is high, a compact latent space necessarily discards more non\-negligible singular directions, so the resulting finite\-dimensional transition cannot preserve as much of the predictive dynamics as it does on low\-rank datasets\.
ILI\-L also exhibits properties that are less aligned with the assumptions behind the compact Koopman approximation used in K2SVD\. Influenza\-like illness time series can undergo abrupt changes across epidemic phases, changes in dominant temporal patterns, and distribution shifts between context and forecast windows\. Prior Koopman forecasting work has emphasized that temporal distribution shifts and non\-stationary dynamics are central challenges in real\-world time\-series prediction\([Wang et al\. 2022](https://arxiv.org/html/2609.17815#bib.bib26);[Liu et al\. 2023](https://arxiv.org/html/2609.17815#bib.bib17)\)\. In such cases, a single low\-dimensional transition operator is harder to estimate reliably: the regression\-induced operator may be biased toward the context\-window dynamics, and recursive long\-horizon rollout can amplify this mismatch\. Methods with stronger local adaptation, generative flexibility, or larger effective latent capacity can therefore be more favorable on ILI\-L\. We view this result as an informative limitation rather than a contradiction of the main findings: K2SVD is designed to exploit datasets with dominant low\-rank Koopman structure, while ILI\-L represents a high\-rank and distribution\-shifted regime where that structural advantage is weaker\.
Table 10:Results of models using the NRMSE metric on ILI\-L with prediction lengths 96 and 720\. Best results are shown inbold, and second\-best results areunderlined\.
Figure 4:The singular values of the Koopman space exhibit higher rank properties on ILI\-L\.
### D\.4Additional Ablation study
#### Ablation on estimation of Koopman
As described in the method section, K2SVD applied a trainable Koopman matrix as the LoRA estimation as operator\. To show the benefit, we ablate the trainable matrix while initializing and fixing the Koopman matrix as the least\-squares estimate𝖪𝐛ols\\mathsf\{K\}\_\{\\mathbf\{b\}\}^\{\\mathrm\{ols\}\}and then optimizes only the Kalman filter parameters\. We denote the static matrix theme as K2SVD\-S and shows its results on multiple datasets\. Table[11](https://arxiv.org/html/2609.17815#A4.T11)shows that K2SVD consistently improves over K2SVD\-S on the evaluated datasets\. This improvement is consistent with the motivation in the overall algorithm: the static regression\-induced matrix is obtained from finite\-sample one\-step latent regression, which can be affected by noise and model mismatch\.
Table 11:Ablation study on the two training schemes using the NRMSE metric\. K2SVD\-S uses the regression\-induced Koopman matrix, while K2SVD dynamically updates the Koopman matrix during downstream forecasting training\.Similar Articles
MetaKoopman: Bayesian Meta-Learning of Koopman Operators for Modeling Structured Dynamics under Distribution Shifts
MetaKoopman proposes a Bayesian meta-learning framework for modeling nonlinear dynamics using linear latent representations via Koopman operators, enabling closed-form updates and uncertainty quantification. It is validated on autonomous truck and trailer systems under adverse winter conditions, outperforming prior methods in prediction accuracy and robustness.
Federated Low-Rank Koopman Learning for Multivariate Time-Series Anomaly Detection in IoT Systems
Proposes FedKAD, a federated Koopman anomaly detection framework for multivariate time series in IoT systems, using lightweight sliding-window Koopman representations and a Stiefel-ADMM algorithm for efficient communication and inference.
Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin Systems
This paper introduces Sparse Koopman Autoencoders (SKAEs) to identify local dynamical regimes in multibasin nonlinear systems, demonstrating superior forecasting performance and interpretable latent supports.
Learning the Koopman Operator using Attention Free Transformers
This paper introduces attention-free latent memory and dynamic re-encoding to improve long-horizon predictions in Koopman autoencoders, reducing error accumulation on benchmark dynamical systems.
Global Explanations for Multivariate Time Series Forecasting Models via $K$-Order Markov Approximations
This paper proposes KARMA, a method for explaining multivariate time series forecasting models by constructing a K-order Markov surrogate model that captures temporal dependencies, offering a five-level global explanation hierarchy.