Seasonal and Quantum-inspired Models for Neutron Monitor Time Series Forecasting

arXiv cs.LG Papers

Summary

The paper presents a comparative study of forecasting models for neutron monitor time series, including seasonal baselines and quantum-inspired architectures, highlighting that simple and functional models perform well on periodic scientific data.

arXiv:2609.30281v1 Announce Type: new Abstract: We present a focused and reproducible study of multi-horizon forecasting on the Lomnicky Stit neutron monitor (LMKS) time series. Our evaluation suite covers simple seasonal baselines, modern deep sequence models, and functional and quantum-inspired architectures, including Seasonal Naive, Long Short-Term Memory (LSTM), Temporal Convolutional Network (TCN), N-BEATS, Kolmogorov-Arnold Networks (KAN), and two quantum-inspired variants, QiLSTM and QiKAN. We describe the dataset characteristics, diagnostic analysis, preprocessing pipeline, and training procedures, and report aggregate point-forecast performance using mean absolute error (MAE) and root mean squared error (RMSE) for all evaluated models. Our quick-run results indicate that the quantum-inspired KAN variant, QiKAN, achieves the lowest aggregate forecasting error among the evaluated configurations, while the simple Seasonal Naive baseline remains remarkably competitive. These results suggest that, for highly periodic scientific monitoring time series, models incorporating strong seasonal or low-dimensional functional priors can match or outperform substantially more complex sequence architectures. The findings motivate further investigation of parsimonious and decomposable function approximators for forecasting periodic scientific signals.
Original Article
View Cached Full Text

Cached at: 09/29/26, 09:33 AM

# Seasonal and Quantum-Inspired Models for Neutron-Monitor Time-Series Forecasting
Source: [https://arxiv.org/html/2609.30281](https://arxiv.org/html/2609.30281)
###### Abstract

We present a focused, reproducible study of multi\-horizon forecasting on the Lomnický Štít neutron monitor \(LMKS\) time series\. The evaluation suite spans simple seasonal baselines, modern deep sequence models and functional/quantum\-inspired architectures: Seasonal Naive, LSTM, Temporal Convolutional Network \(TCN\), N\-BEATS, Kolmogorov–Arnold Networks \(KAN\) and two quantum\-inspired variants \(QiLSTM, QiKAN\)\. We describe dataset diagnostics, preprocessing and training recipes, and report aggregate point\-forecast performance \(MAE, RMSE\) for every model\. Our results show that methods encoding strong seasonal or low\-dimensional functional structure perform particularly well on LMKS: Our quick\-run results indicate that a quantum\-inspired KAN variant \(QiKAN\) attains the lowest aggregate error in the evaluated configurations, while the simple Seasonal Naive baseline remains remarkably competitive\. These findings highlight that, for highly periodic scientific monitoring series, parsimonious priors or decomposable function approximators can match or outperform more complex sequence models\.

††institute:Quantum AI Lab, Fractal AI Research, India,
and University College London, Gower Street, London, UKKeywords\.neutron monitor, time\-series forecasting, Kolmogorov–Arnold network, quantum\-inspired models, LMKS

## 1Introduction

Accurate and reliable time\-series forecasting is a foundational capability for scientific monitoring systems\. Ground\-based neutron monitors such as the Lomnický Štít \(LMKS\) station produce continuous records of secondary cosmic\-ray neutrons that are used in space\-weather research, radiation\-environment assessment, atmospheric studies and related operational tasks\. These records are typically available at hourly \(and sometimes sub\-hourly\) resolution and combine pronounced periodic behaviour \(diurnal and seasonal cycles\), long\-range persistence, and intermittent transient excursions driven by solar and geomagnetic activity\. Together, these characteristics make neutron monitor series an instructive and practically important testbed for multi\-horizon forecasting methods\.

From an operational perspective, reliable forecasts of neutron counts serve several concrete purposes: short\- to medium\-horizon prediction can support early warning systems for elevated radiation conditions, guide instrument scheduling and maintenance, and provide inputs to downstream decision\-support algorithms\. From a scientific perspective, forecasts and calibrated predictive intervals aid in separating predictable, quasi\-periodic components from unusual transients that merit further investigation\. In both use cases, models must balance sensitivity to transient events with robustness to regular seasonal structure, and they must supply well\-calibrated uncertainty estimates to enable principled monitoring and alarm thresholds\.

This work studies these challenges empirically by evaluating a diverse set of modelling approaches on the LMKS hourly record\. We compare simple, interpretable baselines that explicitly encode seasonality with a range of learned sequence models that embody different inductive biases: memory\-based recurrent models \(LSTM\), convolutional temporal models \(TCN\), interpretable basis\-expansion networks \(N\-BEATS\), Kolmogorov–Arnold inspired functional decompositions \(KAN\), and quantum\-inspired variants that alter internal algebraic structure \(QiLSTM, QiKAN\)\. This diversity is intended to reveal which inductive biases most effectively capture the LMKS dynamics under limited tuning budgets and typical data\-quality conditions\.

#### Contributions\.

The main contributions of this work are:

- •A reproducible multi\-horizon forecasting benchmark on the Lomnický Štít \(LMKS\) hourly neutron\-monitor series, with preprocessing and training recipes\.
- •A comparative evaluation of seasonal baselines, LSTM, TCN, N\-BEATS, KAN and quantum\-inspired variants \(QiLSTM, QiKAN\) under a common training regime\.
- •Empirical evidence that parsimonious seasonality\-aware and functional\-decomposition models can match or outperform more complex sequence models on highly periodic monitoring series\.

Note on terminology: by “quantum\-inspired” we do not mean that these models were executed on quantum hardware\. Instead the term denotes classical network variants that adopt algebraic motifs inspired by quantum\-mechanical concepts \(complex\-valued states, phase\-modulated recombination, approximate unitary constraints\)\. These are classical, implementable architectures; any future hardware\-based quantum implementation is left to future work\.

## 2Related work

Time\-series forecasting has matured along several complementary directions, spanning classical statistical methods, tree\-based learners\[[11](https://arxiv.org/html/2609.30281#bib.bib11),[10](https://arxiv.org/html/2609.30281#bib.bib10)\]and modern deep\-learning architectures\[[19](https://arxiv.org/html/2609.30281#bib.bib19)\]\. Recurrent networks such as LSTM and GRU remain fundamental for modelling temporal dependencies and gated memory\[[3](https://arxiv.org/html/2609.30281#bib.bib3)\], while convolutional sequence models \(notably the Temporal Convolutional Network\) exploit dilated causal convolutions to achieve very large receptive fields with efficient parallelism\[[4](https://arxiv.org/html/2609.30281#bib.bib4)\]\. Attention\-based Transformers were introduced to capture flexible long\-range interactions\[[5](https://arxiv.org/html/2609.30281#bib.bib5)\]and are widely used in NLP; for example BERT\[[20](https://arxiv.org/html/2609.30281#bib.bib20)\], and have been successfully adapted to forecasting problems where pairwise temporal relationships are important\[[5](https://arxiv.org/html/2609.30281#bib.bib5)\]\. Architectures such as N\-BEATS\[[6](https://arxiv.org/html/2609.30281#bib.bib6)\]adopt a different tack by using stacked fully\-connected blocks with explicit backcast/forecast projections to provide an interpretable basis\-expansion approach to multi\-horizon forecasting\.

A parallel strand of research studies architectures and representations motivated by classical function theory\. The Kolmogorov–Arnold representation theorem provides a theoretical basis for decomposing multivariate functions into sums of univariate components; Kolmogorov–Arnold Networks \(KAN\) operationalize this idea by combining learned linear projections with banks of univariate nonlinear subnetworks, a design that can be particularly parameter\-efficient when the target mapping admits low\-rank or decomposable structure\[[1](https://arxiv.org/html/2609.30281#bib.bib1),[2](https://arxiv.org/html/2609.30281#bib.bib2),[7](https://arxiv.org/html/2609.30281#bib.bib7)\]\. Inspired by alternative algebraic structures, several recent works explore “quantum\-inspired” modifications \(e\.g\., QKAN\[[8](https://arxiv.org/html/2609.30281#bib.bib8)\], QLSTM\[[9](https://arxiv.org/html/2609.30281#bib.bib9)\]\) to classical architectures \(for example complex\-valued internal states, unitary\-like transforms or phase\-coupling mechanisms\)\. These variants aim to introduce different inductive biases that can improve expressivity in certain regimes, though they are often more sensitive to optimisation and hyperparameter choices\.

On the uncertainty\-quantification side, heteroscedastic Gaussian output parameterizations\[[12](https://arxiv.org/html/2609.30281#bib.bib12)\]enable direct maximum\-likelihood training of per\-step variances\[[12](https://arxiv.org/html/2609.30281#bib.bib12)\], and deep ensembles provide a practical approach to capturing epistemic uncertainty\[[13](https://arxiv.org/html/2609.30281#bib.bib13)\]\. Conformal prediction complements model\-native uncertainty by offering distribution\-free recalibration procedures with finite\-sample guarantees; recent work\[[22](https://arxiv.org/html/2609.30281#bib.bib22)\],\[[23](https://arxiv.org/html/2609.30281#bib.bib23)\],\[[21](https://arxiv.org/html/2609.30281#bib.bib21)\]has extended conformal ideas to multivariate and multi\-horizon forecasting settings\[[14](https://arxiv.org/html/2609.30281#bib.bib14),[15](https://arxiv.org/html/2609.30281#bib.bib15)\]\. Proper scoring rules such as CRPS, negative log\-likelihood and classic point metrics \(MAE, RMSE\) remain standard for evaluating predictive quality and calibration\[[16](https://arxiv.org/html/2609.30281#bib.bib16)\]\.

Finally, for domain\-specific application to particle accelerators and related scientific instrumentation, recent reviews synthesise methodological choices and practical considerations\. In particular, Li and Adelmann \(2022\) provide a focused review of time\-series forecasting techniques and analyse their suitability for accelerator diagnostics and control tasks, highlighting practical trade\-offs between interpretability, latency, and robustness in operational settings\[[18](https://arxiv.org/html/2609.30281#bib.bib18)\]\. Their survey situates the present study within the practical demands of scientific monitoring: short\-latency forecasting, calibrated uncertainty for alarm thresholds, and robustness to instrumentation anomalies\.

## 3Dataset: LMKS \(Lomnický Štít Neutron Monitor\)

The Lomnický Štít neutron monitor \(LMKS\) provides one of the longest continuous ground\-based records of cosmic\-ray induced secondary neutrons\. Its hourly series exhibits a combination of clear diurnal and seasonal cycles, long\-range autocorrelations and occasional transient disturbances caused by solar and geomagnetic activity\. These characteristics make LMKS an excellent testbed for benchmarking forecasting methods that must simultaneously capture regular periodicity and remain sensitive to rare events\.

For this study we used the Lomnický Štít \(LMKS\) neutron\-monitor release archived on Zenodo\[[17](https://arxiv.org/html/2609.30281#bib.bib17)\]\(Institute of Experimental Physics, Slovak Academy of Sciences; DOI: 10\.5281/zenodo\.10790916\)\. Following the protocol implemented in our code, we first performed an audit of the raw data, removed non\-numeric metadata fields, and applied forward/backward filling to handle missing hourly counts\. All models were trained on standardized series segments to ensure comparability across architectures\.

To transform the series into a supervised learning problem, we adopted a sliding\-window approach: each training instance consists of a history ofTTpast values paired with the nextHHtarget values\. The quick\-run configuration reported in this paper usesT=20T=20hours of history to forecastH=10H=10future hours, providing a balanced setting for assessing multi\-horizon forecasting capability\. This design yields a benchmark dataset that stresses both short\-term predictive responsiveness and the ability to leverage longer\-term seasonal structure\.

We began with a feature audit and summary statistics\. The LMKS series displays strong autocorrelations at short lags and pronounced seasonal structure at daily and longer scales\. Missing observations are sporadic and were imputed with forward\-fill followed by backward\-fill; any leftover windows containing irrecoverable NaNs were dropped\. For models requiring standardized inputs we computed per\-channel training means and standard deviations and applied z\-score normalization to train/validation/test splits\. Figure[1](https://arxiv.org/html/2609.30281#S3.F1)shows the hourly LMKS series used in our experiments\.

Figure 1:LMKS hourly neutron monitor time series\. The series shows clear daily/seasonal patterns and transient variation that forecasting models must capture\.
## 4Models and implementation details

To ensure a reproducible and fair comparison, all models were implemented in PyTorch and trained under a common regime\. Unless stated otherwise, we used the Adam optimiser with an initial learning rate of10−310^\{\-3\}, batch size 64, gradient clipping at 1\.0, and early stopping on validation mean squared error with a patience of eight epochs\. Dropout between 0\.1 and 0\.3 was employed to regularise most models, and z\-score normalisation based on training\-set statistics was applied to all input sequences\. Forecast quality was measured using both mean absolute error \(MAE\) and root mean squared error \(RMSE\)\. The overall pipeline is illustrated in Figure[2](https://arxiv.org/html/2609.30281#S4.F2)\.

![Refer to caption](https://arxiv.org/html/2609.30281v1/Data_Flow_Pipeline.png)Figure 2:The data flow pipeline for the LMKS dataset\. We analyse and forecast using different temporally designed models### 4\.1Seasonal Naive

The Seasonal Naive predictor is an intentionally simple, fully deterministic baseline that directly copies observed seasonal structure into the forecast\. Let\{xt\}\\\{x\_\{t\}\\\}denote the univariate time series sampled hourly and letSSdenote the seasonal period \(for the LMKS dataset,S=24S=24hours to reflect the daily cycle\)\. For a forecast origin at timettand horizonh∈\{1,…,H\}h\\in\\\{1,\\dots,H\\\}the seasonal naive forecast is

x^t\+hSN=xt\+h−S\.\\hat\{x\}\_\{t\+h\}^\{\\text\{SN\}\}=x\_\{t\+h\-S\}\.\(1\)i\.e\. each predicted value is taken from the same hour in the previous daily cycle\. Despite its simplicity, this rule is a strong benchmark for datasets with pronounced diurnal structure: it requires no parameter estimation, is entirely reproducible, and directly exposes the amount of residual structure that any learning model must explain beyond seasonality replication\.

### 4\.2Long\-Short Term Memory \(LSTM\)

The Long Short–Term Memory \(LSTM\) model is used in an encoder–decoder sequence\-to\-sequence formulation to transform a history window of lengthTTinto a multi\-step forecast of lengthHH\. Denote the input history by𝐱t−T\+1:t=\(xt−T\+1,…,xt\)\\mathbf\{x\}\_\{t\-T\+1:t\}=\(x\_\{t\-T\+1\},\\dots,x\_\{t\}\)\. The encoder consumes this sequence and compresses temporal dependencies into a final latent state𝐳t\\mathbf\{z\}\_\{t\}, while the decoder unfolds this representation into the horizon𝐱^t\+1:t\+H\\hat\{\\mathbf\{x\}\}\_\{t\+1:t\+H\}\. Internally the LSTM cell updates can be written \(in standard real\-valued form\) as

𝐢u\\displaystyle\\mathbf\{i\}\_\{u\}=σ⁡\(Wi​\[𝐡u−1,xu\]\+bi\),\\displaystyle=\\sigma\(W\_\{i\}\[\\mathbf\{h\}\_\{u\-1\},x\_\{u\}\]\+b\_\{i\}\),\(2\)𝐟u\\displaystyle\\mathbf\{f\}\_\{u\}=σ⁡\(Wf​\[𝐡u−1,xu\]\+bf\),\\displaystyle=\\sigma\(W\_\{f\}\[\\mathbf\{h\}\_\{u\-1\},x\_\{u\}\]\+b\_\{f\}\),𝐨u\\displaystyle\\mathbf\{o\}\_\{u\}=σ⁡\(Wo​\[𝐡u−1,xu\]\+bo\),\\displaystyle=\\sigma\(W\_\{o\}\[\\mathbf\{h\}\_\{u\-1\},x\_\{u\}\]\+b\_\{o\}\),𝐜~u\\displaystyle\\tilde\{\\mathbf\{c\}\}\_\{u\}=tanh⁡\(Wc​\[𝐡u−1,xu\]\+bc\),\\displaystyle=\\tanh\(W\_\{c\}\[\\mathbf\{h\}\_\{u\-1\},x\_\{u\}\]\+b\_\{c\}\),𝐜u\\displaystyle\\mathbf\{c\}\_\{u\}=𝐟u⊙𝐜u−1\+𝐢u⊙𝐜~u,\\displaystyle=\\mathbf\{f\}\_\{u\}\\odot\\mathbf\{c\}\_\{u\-1\}\+\\mathbf\{i\}\_\{u\}\\odot\\tilde\{\\mathbf\{c\}\}\_\{u\},𝐡u\\displaystyle\\mathbf\{h\}\_\{u\}=𝐨u⊙tanh⁡\(𝐜u\)\.\\displaystyle=\\mathbf\{o\}\_\{u\}\\odot\\tanh\(\\mathbf\{c\}\_\{u\}\)\.withσ⁡\(⋅\)\\sigma\(\\cdot\)the logistic sigmoid and⊙\\odotthe elementwise product\. In the encoder–decoder variant, the encoder produces a summary state that initializes the decoder; the decoder then produces outputs either autoregressively \(feeding predicted values back as inputs\) or with scheduled teacher forcing \(mixing ground\-truth and model predictions during training to stabilise learning\)\. For deterministic forecasting the decoder’s hidden states are mapped to scalar forecasts via a linear readout and trained under mean\-squared error \(MSE\):

ℒMSE=1H​∑h=1H\(xt\+h−x^t\+h\)2\.\\mathcal\{L\}\_\{\\text\{MSE\}\}=\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\bigl\(x\_\{t\+h\}\-\\hat\{x\}\_\{t\+h\}\\bigr\)^\{2\}\.\(3\)For probabilistic forecasting the model can be extended to predict heteroscedastic Gaussian parameters, producing a meanμt\+h\\mu\_\{t\+h\}and a \(positive\) scale parameterσt\+h\\sigma\_\{t\+h\}; training then minimises the Gaussian negative log\-likelihood,

ℒNLL=12​∑h=1H\[log⁡\(σt\+h2\)\+\(xt\+h−μt\+h\)2σt\+h2\]\+const\.\\mathcal\{L\}\_\{\\text\{NLL\}\}=\\frac\{1\}\{2\}\\sum\_\{h=1\}^\{H\}\\left\[\\log\\\!\\big\(\\sigma\_\{t\+h\}^\{2\}\\big\)\+\\frac\{\\bigl\(x\_\{t\+h\}\-\\mu\_\{t\+h\}\\bigr\)^\{2\}\}\{\\sigma\_\{t\+h\}^\{2\}\}\\right\]\+\\text\{const\}\.\(4\)where the variance is enforced positive via a smooth transform such as a softplus\. Standard regularisation and optimisation practices — dropout between recurrent layers to reduce overfitting, and gradient clipping to avoid exploding gradients — support stable training and generalisation, while choices such as latent size or number of recurrent layers control the model’s capacity to capture longer\-term temporal patterns\.

### 4\.3Temporal Convolutional Network \(TCN\)

The Temporal Convolutional Network \(TCN\) offers a convolutional, non\-recurrent approach to sequence modelling that retains causal structure and allows for very large receptive fields through dilated convolutions\[[4](https://arxiv.org/html/2609.30281#bib.bib4)\], a design popularized in WaveNet\[[25](https://arxiv.org/html/2609.30281#bib.bib25)\]\. Let the convolutional kernel width bekkand suppose a stack ofLLresidual blocks is used with dilation factorsdℓd\_\{\\ell\}\(commonlydℓ=2ℓ−1d\_\{\\ell\}=2^\{\\ell\-1\}\)\. A single dilated causal convolution at layerℓ\\ellcomputes

yt\(ℓ\)=∑m=0k−1wm\(ℓ\)​xt−dℓ​m\.y^\{\(\\ell\)\}\_\{t\}=\\sum\_\{m=0\}^\{k\-1\}w^\{\(\\ell\)\}\_\{m\}\\,x\_\{t\-d\_\{\\ell\}m\}\.\(5\)guaranteeing that outputs at timettdepend only on past inputs\. The effective receptive field of an exponentially dilated stack is

R=1\+\(k−1\)​∑ℓ=0L−1dℓ\.R=1\+\(k\-1\)\\sum\_\{\\ell=0\}^\{L\-1\}d\_\{\\ell\}\.\(6\)which grows exponentially with depth and allows the model to cover the entire input windowTTwith relatively few layers\. Each residual block typically contains two causal convolutions, nonlinearities \(e\.g\. ReLU\), dropout, layer normalisation and a residual connection that adds the block input to its output; this architecture preserves stable gradient flow and enables parallel prediction of the full horizon by mapping the final feature representation to theHH\-step forecast in one shot\. The TCN is therefore well suited to efficient multi\-step forecasting when the dependencies to be modelled are temporally local after dilation, and when autoregressive decoding is undesirable for latency reasons\.

### 4\.4N\-BEATS

N\-BEATS is a purely feed\-forward, block\-stacked architecture that iteratively explains the input signal through repeated backcast subtraction while simultaneously contributing to the forecast\. Given an input window𝐱t−T\+1:t\\mathbf\{x\}\_\{t\-T\+1:t\}, each blockbbimplements a function that outputs a backcast𝐛\(b\)\\mathbf\{b\}^\{\(b\)\}\(an approximation of the portion of the input it explains\) and a forecast contribution𝐟\(b\)\\mathbf\{f\}^\{\(b\)\}:

\(𝐛\(b\),𝐟\(b\)\)\\displaystyle\\big\(\\mathbf\{b\}^\{\(b\)\},\\mathbf\{f\}^\{\(b\)\}\\big\)=ℬ\(b\)​\(𝐫\(b\)\),\\displaystyle=\\mathcal\{B\}^\{\(b\)\}\\big\(\\mathbf\{r\}^\{\(b\)\}\\big\),\(7\)𝐫\(b\+1\)\\displaystyle\\mathbf\{r\}^\{\(b\+1\)\}=𝐫\(b\)−𝐛\(b\)\.\\displaystyle=\\mathbf\{r\}^\{\(b\)\}\-\\mathbf\{b\}^\{\(b\)\}\.with the residual𝐫\(1\)=𝐱t−T\+1:t\\mathbf\{r\}^\{\(1\)\}=\\mathbf\{x\}\_\{t\-T\+1:t\}and the final forecast produced by additive aggregation𝐱^t\+1:t\+H=∑b𝐟\(b\)\\hat\{\\mathbf\{x\}\}\_\{t\+1:t\+H\}=\\sum\_\{b\}\\mathbf\{f\}^\{\(b\)\}\. In the generic\-block instantiation eachℬ\(b\)\\mathcal\{B\}^\{\(b\)\}is a fully connected MLP that implicitly learns basis functions over the input window; the iterative backcast subtraction encourages blocks to specialise \(for example into trend, seasonality or local corrections\) even without an explicit decomposition\. Training minimises the aggregated MSE over the horizon, and architectural choices such as block width, depth and number of blocks determine the richness of the learned basis expansions\. The model’s design affords interpretability of block\-level contributions while remaining flexible enough to approximate a wide class of temporal patterns\.

### 4\.5Kolmogorov–Arnold Networks \(KAN\)

Kolmogorov–Arnold Networks \(KAN\) operationalise the classical Kolmogorov–Arnold representation theorem, which states that a continuous multivariate function can be expressed as a finite sum of univariate nonlinear functions applied to linear combinations of the inputs\. Abstractly, for a mappingF:ℝn→ℝF:\\mathbb\{R\}^\{n\}\\to\\mathbb\{R\}there exist univariate functionsϕq\\phi\_\{q\}and linear projection weightsaqa\_\{q\}such that

F⁡\(𝐱\)≈∑q=1Qϕq​\(⟨𝐚q,𝐱⟩\)\.F\(\\mathbf\{x\}\)\\approx\\sum\_\{q=1\}^\{Q\}\\phi\_\{q\}\\big\(\\langle\\mathbf\{a\}\_\{q\},\\mathbf\{x\}\\rangle\\big\)\.\(8\)A KAN model leverages this decomposition by first projecting theTT\-length history \(viewed either as aTT\-dimensional vector or via a learned embedding\) into a collection of scalar componentsuq=⟨𝐚q,𝐱t−T\+1:t⟩u\_\{q\}=\\langle\\mathbf\{a\}\_\{q\},\\mathbf\{x\}\_\{t\-T\+1:t\}\\rangle, then passing each scalaruqu\_\{q\}through an independent univariate subnetworkϕq\\phi\_\{q\}and finally recombining the outputs linearly to form the multi\-step forecast\. In practice this induces a strong low\-rank inductive bias: the function from high\-dimensional history to forecast is modelled as a sum of simpler univariate nonlinearities, which yields parameter efficiency and can be particularly effective on quasi\-periodic signals where dominant projections capture most of the explanatory variance\. Regularisation, selection ofQQ, and the capacity of eachϕq\\phi\_\{q\}control the trade\-off between expressivity and overfitting\.

### 4\.6QiLSTM

Quantum\-inspired LSTM \(QiLSTM\) models extend the standard recurrent paradigm by permitting internal representations to carry phase information and by encouraging linear state transforms to behave like norm\-preserving rotations\. Concretely, the hidden state is treated as complex\-valued,𝐡u∈ℂd\\mathbf\{h\}\_\{u\}\\in\\mathbb\{C\}^\{d\}, and recurrent updates are constructed so that the principal linear operators satisfy an approximate unitarity constraintW†​W≈IW^\{\\dagger\}W\\approx I\. Gates and nonlinearities may be implemented either by operating separately on real and imaginary parts or by decomposing complex scalars into magnitude and phase and applying tailored operations to each component; final, real\-valued predictions are obtained by mapping complex outputs to the real domain via the real part, the magnitude, or a learned linear readout that consumes concatenated real and imaginary channels\. Practically, a structured front end that produces compact, phase\-aware embeddings of the input history is often used: local scalar encodings are combined through a chain of small linear cores whose inputs are formed by the outer\-product between the current bond state and the local encoding \(flattened and passed through the core\), and the resulting final bond state is pooled and projected to form the LSTM input\. Because QiLSTM aims to represent sustained oscillations and phase\-locked phenomena, training typically includes stability priors such as a unitarity regularizer

ℛunit\(W\)=∥W†W−I∥F2\.\\mathcal\{R\}\_\{\\mathrm\{unit\}\}\(W\)=\\bigl\\lVert W^\{\\dagger\}W\-I\\bigr\\rVert\_\{F\}^\{2\}\.\(9\)and optional amplitude or phase\-drift penalties; alternatively, one may parameterize certain linear blocks to be exactly unitary \(e\.g\., via exponentiation of skew\-Hermitian generators or structured products of Givens rotations\) when exact norm preservation is required\. Together, the complex\-valued state, phase\-aware nonlinearities, and norm\-preserving linear components give QiLSTM a compact mechanism for encoding rotations, phase delays, and interference\-like combination rules that are especially effective when the target dynamics contain prominent frequency and phase structure\.

### 4\.7QiKAN

The quantum\-inspired Kolmogorov–Arnold Network \(QiKAN\) adapts the KAN decomposition to support phase and interference by combining many localized univariate transforms with a small number of learned aggregators\. Given an input vectorx=\(x1,…,xn\)x=\(x\_\{1\},\\dots,x\_\{n\}\), QiKAN applies per\-coordinate univariate subnetworksϕj:ℝ→ℂm\\phi\_\{j\}:\\mathbb\{R\}\\to\\mathbb\{C\}^\{m\}\(orℝm\\mathbb\{R\}^\{m\}when a real formulation is preferred\) and formsKKaggregator channels whose inputs are linear combinations of these univariate encodings:

uk\(x\)=∑j=1nak​jϕj\(xj\),k=1,…,K\.u\_\{k\}\(x\)=\\sum\_\{j=1\}^\{n\}a\_\{kj\}\\,\\phi\_\{j\}\(x\_\{j\}\),\\quad k=1,\\dots,K\.\(10\)where the coefficientsak​ja\_\{kj\}are learned and may be complex to permit constructive and destructive interference across coordinates\. Each aggregator outputuku\_\{k\}is then mapped through a small nonlinear maphkh\_\{k\}to produce a contribution, and the final prediction is the sum of these contributions,

f^​\(x\)=∑k=1Khk​\(uk​\(x\)\)\.\\hat\{f\}\(x\)=\\sum\_\{k=1\}^\{K\}h\_\{k\}\\big\(u\_\{k\}\(x\)\\big\)\.\(11\)This architecture realizes a low\-rank, compositional recombination of localized features: when cross\-coordinate interactions are structured or effectively low\-dimensional, a modest number of aggregators suffices to capture global nonlinear effects while remaining statistically efficient\. In the complex\-valued variant, relative phases inϕj​\(xj\)\\phi\_\{j\}\(x\_\{j\}\)and in the coefficientsak​ja\_\{kj\}enable interference patterns that can succinctly represent oscillatory coupling across inputs; to preserve numerical stability one therefore typically couples the QiKAN with the same regularization motifs used in QiLSTM \(unitarity or approximate norm\-preservation on linear recombination layers, amplitude/phase penalties, and spectral normalization\)\. Training objectives combine the usual predictive loss \(e\.g\., MSE or MAE\) with these structural regularizers, and practical implementations often fit any dimensionality\-reducing transforms on training data only, checkpoint on validation performance, and tune the strength of unitary/phase penalties so as to stabilize dynamics without removing useful amplitude modulation\.

### 4\.8Calibration and evaluation metrics

Model performance was primarily assessed using point\-forecast errors, measured by Mean Absolute Error \(MAE\) and Root Mean Squared Error \(RMSE\)\. For a dataset ofNNforecast instances, each with a horizon of lengthHH, and predictionsy^i,t\\hat\{y\}\_\{i,t\}against ground truthyi,ty\_\{i,t\}, the metrics are defined as

MAE=1N​H​∑i=1N∑t=1H\|yi,t−y^i,t\|\.\\text\{MAE\}\\;=\\;\\frac\{1\}\{NH\}\\sum\_\{i=1\}^\{N\}\\sum\_\{t=1\}^\{H\}\\bigl\|y\_\{i,t\}\-\\hat\{y\}\_\{i,t\}\\bigr\|\.\(12\)
RMSE=1N​H​∑i=1N∑t=1H\(yi,t−y^i,t\)2\.\\text\{RMSE\}\\;=\\;\\sqrt\{\\frac\{1\}\{NH\}\\sum\_\{i=1\}^\{N\}\\sum\_\{t=1\}^\{H\}\\bigl\(y\_\{i,t\}\-\\hat\{y\}\_\{i,t\}\\bigr\)^\{2\}\}\\,\.\(13\)
These definitions match the implementation used in our codebase, where errors are first computed element\-wise across all horizon steps and then averaged across the full evaluation set\.

## 5Results

Table[1](https://arxiv.org/html/2609.30281#S5.T1)summarises aggregate forecasting performance \(MAE and RMSE\) for all evaluated models on the LMKS dataset\. The numbers are aggregated across the chosen horizon and were produced under the execution configuration described above\.

Table 1:LMKS forecasting results: aggregate MAE and RMSE across the forecast horizon \(quick\-run snapshot\)\.Figure 3:Results for the LMKS dataset\. In the quick\-run evaluation reported here, the quantum\-inspired KAN variant \(QiKAN\) achieved the lowest aggregate point forecast error among the configurations tested\.### Analysis

The quantum\-inspired KAN variant \(QiKAN\) achieves the lowest aggregate errors in this quick\-run snapshot\. This outcome indicates that strong periodic priors and low\-dimensional functional decompositions are particularly effective on the LMKS series, which exhibits pronounced periodicity and slowly varying behavior\. Classical sequence models \(LSTM\) and functional deep models \(N\-BEATS, KAN\) produce similar error magnitudes, while the TCN underperformed in our default quick\-run hyperparameter choices — likely because its receptive\-field design and convolutional inductive bias require careful kernel/dilation tuning to match the LMKS seasonality scale\. Quantum\-inspired LSTM \(QiLSTM\) shows higher error in the quick\-run setting, suggesting optimisation sensitivity that merits more exhaustive tuning\. Summary metrics are plotted in Figure[3](https://arxiv.org/html/2609.30281#S5.F3)\.

#### Limitations\.

This study used a focused, quick\-run experimental protocol with a modest hyperparameter search budget; as a consequence, some models \(notably the Qi\-variants and TCN\) exhibited sensitivity to initialization and regularisation choices\. Results reported here should therefore be interpreted as a reproducible snapshot under the stated defaults rather than exhaustive performance bounds\. We evaluated a single dataset snapshot \(LMKS hourly record from 1981\-12\-01 through 2023\-07\-10\); generalisation to other neutron\-monitor series or different sampling resolutions may require further tuning and validation\. Future work will expand hyperparameter sweeps, report full seed\-averaged statistics, and provide the full code and checkpoint artifacts for independent verification\.

## 6Discussion and Conclusion

Several practical lessons emerge from the LMKS experiments\. First, when a time series shows a strong, regular seasonal pattern, simple seasonality\-aware baselines—seasonal naïve forecasts, classical decomposition \(trend \+ seasonal \+ residual\), or lightweight parametric seasonality models—often perform very competitively\. These methods exploit dominant periodicity directly, remain stable over medium\-length horizons, and require minimal hyperparameter tuning or large training sets; for monitoring tasks with clear periodicity they are reliable, low\-cost baselines\.

Second, modern deep sequence models provide considerable representational flexibility but can be fragile with respect to architecture and receptive\-field choices\. Temporal Convolutional Networks \(TCNs\), Transformers, and related architectures can outperform classical methods, but only when receptive fields, positional encodings, dilation patterns, attention windows, and other design choices are well matched to the data’s periodicity and the forecast horizon\. In our LMKS experiments—where seasonality is strong and horizons are moderate—TCNs and Transformers required careful configuration to consistently beat simpler baselines, underscoring the need for targeted architecture search and thorough validation\.

Third, quantum\-inspired and quantum\-hybrid architectures are an intriguing exploratory direction\. Our quick\-run results show promise but indicate these models are not yet plug\-and\-play: they demand careful optimisation, appropriate inductive biases, and architecture\-specific regularisation to reliably surpass classical baselines\. Practical considerations—parameter initialisation, constrained parameter counts, robustness to noise, and training schedules—strongly influence whether a quantum\-inspired variant can deliver its theoretical advantages in finite\-data, noisy settings\.

These conclusions come from a focused evaluation on the LMKS neutron monitor dataset\. By comparing seasonal baselines, recurrent and convolutional sequence models, interpretable basis\-expansion networks \(functional decomposition\), and quantum\-inspired variants, we present a practical landscape of model performance for scientific monitoring\. The clearest takeaway is the enduring strength of seasonality\-aware baselines and functional decomposition: they yield dependable results with modest complexity, while more complex models require careful tuning to justify their use\.

Future work will implement and evaluate Variational Quantum Circuit \(VQC\)–based hybrid architectures—specifically a quantum\-augmented LSTM and a quantum KAN—by replacing or augmenting key linear transforms with compact VQC modules and by exploring encodings \(angle, amplitude, IQP\-style\) suited to periodic signals\. Experiments will proceed from high\-fidelity simulators to hardware runs, employ staged training and parameter\-efficient ansätze, and use architecture\-specific regularisation; benchmarks, ablations, and reproducible code, seeds, and configurations will clarify whether VQC hybrids offer consistent expressivity or robustness benefits for scientific monitoring\.

## References

- \(1\)A\. N\. Kolmogorov\.On the representation of continuous functions of several variables by superpositions of continuous functions of one variable and addition\.Doklady Akademii Nauk SSSR, 108:179–182, 1957\.
- \(2\)V\. I\. Arnold\.On functions of three variables\.Doklady Akademii Nauk SSSR, 114:679–681, 1957\.
- \(3\)S\. Hochreiter and J\. Schmidhuber\.Long short\-term memory\.Neural Computation, 9\(8\):1735–1780, 1997\.
- \(4\)S\. Bai, J\. Z\. Kolter, and V\. Koltun\.An empirical evaluation of generic convolutional and recurrent networks for sequence modeling\.arXiv preprint arXiv:1803\.01271, 2018\.
- \(5\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin\.Attention is all you need\.InAdvances in Neural Information Processing Systems \(NeurIPS\), pages 5998–6008, 2017\.
- \(6\)B\. N\. Oreshkin, D\. Carpov, N\. Chapados, and Y\. Bengio\.N\-BEATS: Neural basis expansion analysis for interpretable time series forecasting\.arXiv preprint arXiv:1905\.10437, 2019 \(ICLR 2020\)\.
- \(7\)Z\. Liu, Y\. Wang, S\. Vaidya, F\. Rühle, J\. Halverson, M\. Soljačić, T\. Y\. Hou, and M\. Tegmark\.KAN: Kolmogorov–Arnold networks\.arXiv preprint arXiv:2404\.19756, 2024\. \(Accepted at ICLR 2025\.\)
- \(8\)P\. Ivashkov, P\.\-W\. Huang, K\. Koor, L\. Pira, and P\. Rebentrost\.QKAN: Quantum Kolmogorov–Arnold networks\.arXiv preprint arXiv:2410\.04435, 2024\.
- \(9\)S\. Y\.\-C\. Chen, S\. Yoo, and Y\.\-L\. L\. Fang\.Quantum Long Short\-Term Memory \(QLSTM\)\.arXiv preprint arXiv:2009\.01783, 2020\.
- \(10\)T\. Chen and C\. Guestrin\.XGBoost: A scalable tree boosting system\.InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining \(KDD\), pages 785–794, 2016\.
- \(11\)L\. Breiman\.Random forests\.Machine Learning, 45\(1\):5–32, 2001\.
- \(12\)D\. A\. Nix and A\. S\. Weigend\.Estimating the mean and variance of the target probability distribution\.InProceedings of the IEEE International Conference on Neural Networks \(ICNN\), pages 55–60, 1994\.
- \(13\)B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell\.Simple and scalable predictive uncertainty estimation using deep ensembles\.InAdvances in Neural Information Processing Systems \(NeurIPS\), 2017\.
- \(14\)V\. Vovk, A\. Gammerman, and G\. Shafer\.Algorithmic Learning in a Random World\.Springer, 2005\.
- \(15\)A\. N\. Angelopoulos and S\. Bates\.A gentle introduction to conformal prediction and distribution\-free uncertainty quantification\.arXiv preprint arXiv:2107\.07511, 2021\.
- \(16\)T\. Gneiting and A\. E\. Raftery\.Strictly proper scoring rules, prediction, and estimation\.Journal of the American Statistical Association, 102\(477\):359–378, 2007\.
- \(17\)Institute of Experimental Physics \(Slovak Academy of Sciences\), S\. Mackovjak, R\. Langer, I\. Strhársky, S\. Štefánik, J\. Kubančák, I\. Kisvárdai, F\. Štempel, and L\. Randuška\.Measurements of Neutron Monitor at Lomnicky Stit \(LMKS\), dataset \(hourly and minute resolution\)\.Zenodo dataset, March 6 2024\. DOI: 10\.5281/zenodo\.10790916\.
- \(18\)S\. Li, A\. Adelmann\. Review of Time Series Forecasting Methods and Their Applications to Particle Accelerators\. arXiv:2209\.10705, 2022\.
- \(19\)I\. Goodfellow, Y\. Bengio, and A\. Courville\. Deep Learning\. MIT Press, 2016\.
- \(20\)J\. Devlin, M\.\-W\. Chang, K\. Lee, and K\. Toutanova\. BERT: Pre\-training of deep bidirectional transformers for language understanding\. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics \(NAACL\), pages 4171–4186, 2019\.
- \(21\)D\. P\. Kingma and M\. Welling\. Auto\-Encoding Variational Bayes\. arXiv preprint arXiv:1312\.6114, 2014\.
- \(22\)Y\. Gal and Z\. Ghahramani\. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning\. In Proceedings of the 33rd International Conference on Machine Learning \(ICML\), 2016\.
- \(23\)C\. E\. Rasmussen and C\. K\. I\. Williams\. Gaussian Processes for Machine Learning\. MIT Press, 2006\.
- \(24\)R\. J\. Hyndman and G\. Athanasopoulos\. Forecasting: Principles and Practice\. OTexts \(online textbook\), 2nd edition, 2018\.
- \(25\)A\. van den Oord, S\. Dieleman, H\. Zen, K\. Simonyan, O\. Vinyals, A\. Graves, N\. Kalchbrenner, A\. Senior, and K\. Kavukcuoglu\. WaveNet: A generative model for raw audio\. arXiv preprint arXiv:1609\.03499, 2016\.

Similar Articles