Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning
Summary
This paper introduces Gated QKAN-FWP, a scalable quantum-inspired sequence learning framework that combines Fast Weight Programmers with Kolmogorov-Arnold Networks using single-qubit data re-uploading circuits.
View Cached Full Text
Cached at: 05/11/26, 06:46 AM
# Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning
Source: [https://arxiv.org/html/2605.06734](https://arxiv.org/html/2605.06734)
Kuo\-Chung Peng\*,Department of Physics and Center for Theoretical Physics, National Taiwan University, Taipei, TaiwanNational Center for High\-Performance Computing, National Institutes of Applied Research, Hsinchu, TaiwanJiun\-Cheng JiangDepartment of Physics and Center for Theoretical Physics, National Taiwan University, Taipei, TaiwanNVIDIA AI Technology Center, NVIDIA Corp\., Taipei, TaiwanCenter for Quantum Science and Engineering, National Taiwan University, Taipei, TaiwanChen\-Yu LiuGraduate Institute of Applied Physics, National Taiwan University, Taipei, TaiwanEn\-Jui KuoDepartment of Electrophysics, National Yang Ming Chiao Tung University, Hsinchu, TaiwanYun\-Yuan WangNVIDIA AI Technology Center, NVIDIA Corp\., Taipei, TaiwanPrayag TiwariSchool of Information Technology, Halmstad University, SwedenAndrea CeschiniDepartment of Information Engineering, Electronics and Telecommunications \(DIET\), University of Rome “La Sapienza”, Rome, ItalyChi\-Sheng ChenBeth Israel Deaconess Medical Center & Harvard Medical School, Boston, MA, USAYu\-Chao HsuNational Center for High\-Performance Computing, National Institutes of Applied Research, Hsinchu, TaiwanCross College Elite Program, National Cheng Kung University, Tainan, TaiwanChun\-Hua LinDepartment of Physics and Center for Theoretical Physics, National Taiwan University, Taipei, TaiwanNational Center for High\-Performance Computing, National Institutes of Applied Research, Hsinchu, TaiwanTai\-Yue LiNational Center for High\-Performance Computing, National Institutes of Applied Research, Hsinchu, TaiwanAntonello RosatoDepartment of Information Engineering, Electronics and Telecommunications \(DIET\), University of Rome “La Sapienza”, Rome, ItalyMassimo PanellaDepartment of Information Engineering, Electronics and Telecommunications \(DIET\), University of Rome “La Sapienza”, Rome, ItalySimon SeeNVIDIA AI Technology Center, NVIDIA Corp\., Singapore, SingaporeSaif Al\-KuwariQatar Center for Quantum Computing, College of Science and Engineering, Hamad Bin Khalifa University, Doha, QatarKuan\-Cheng Chen†\\dagger,Qatar Center for Quantum Computing, College of Science and Engineering, Hamad Bin Khalifa University, Doha, QatarNan\-Yow Chen†\\dagger,National Center for High\-Performance Computing, National Institutes of Applied Research, Hsinchu, TaiwanHsi\-Sheng Goan†\\dagger,Department of Physics and Center for Theoretical Physics, National Taiwan University, Taipei, TaiwanNVIDIA AI Technology Center, NVIDIA Corp\., Taipei, TaiwanGraduate Institute of Applied Physics, National Taiwan University, Taipei, TaiwanPhysics Division, National Center for Theoretical Sciences, Taipei, Taiwan
###### Abstract
Fast Weight Programmers \(FWPs\) encode temporal dependencies through dynamically updated parameters rather than recurrent hidden states\. Quantum FWPs \(QFWPs\) extend this idea with variational quantum circuits \(VQCs\), but existing implementations rely on multi\-qubit architectures that are difficult to scale on noisy intermediate\-scale quantum \(NISQ\) devices and expensive to simulate classically\. We propose gated QKAN\-FWP, a fast\-weight framework that integrates FWP with Quantum\-inspired Kolmogorov–Arnold Network \(QKAN\) using single\-qubit data re\-uploading circuits as learnable nonlinear activation, known as DatA Re\-Uploading ActivatioN \(DARUAN\)\. We further introduce a scalar\-gated fast\-weight update rule that stabilizes parameter evolution, supported by a theoretical analysis of its adaptive memory kernel, geometric boundedness, and parallelizable gradient paths\. We evaluate the framework across time\-series benchmarks, MiniGrid reinforcement learning, and highlight real\-world solar cycle forecasting as our main practical result\. In the long\-horizon setting with 528\-month input window and 132\-month forecast horizon, our 12\.5k\-parameter model achieves lower scaled Mean Square Error \(MSE\), peak amplitude error, and peak timing error than a suite of classical recurrent baselines with up to 13× more parameters, including Long Short\-Term Memory \(LSTM\) networks \(25\.9k–89\.1k parameters\), WaveNet\-LSTM \(167k\), Vanilla recurrent neural network \(11\.5k\), and a Modified Echo State Network \(132k\)\. To validate NISQ compatibility, we further deploy the trained fast programmer on IonQ and IBM Quantum processors, recovering forecasting accuracy within 0\.1% relative MSE of the noiseless simulator at 1024 shots\. These results position gated QKAN\-FWP as a scalable, parameter\-efficient, and NISQ\-compatible approach to quantum\-inspired sequence modeling\.
Keywords:fast weight programming, quantum machine learning, Kolmogorov–Arnold networks, sequence modeling, reinforcement learning
††footnotetext:The views expressed in this article are those of the authors and do not represent the views of Wells Fargo\. This article is for informational purposes only\. Nothing contained in this article should be construed as investment advice\. Wells Fargo makes no express or implied warranties and expressly disclaims all legal, tax, and accounting implications related to this article\.## 1Introduction
Modeling long\-range temporal dependencies remains a central challenge in sequence learning and sequential decision making[L\+25b](https://arxiv.org/html/2605.06734#bib.bib58);[L\+25d](https://arxiv.org/html/2605.06734#bib.bib60); CCTW \([24](https://arxiv.org/html/2605.06734#bib.bib21)\)\. In quantum machine learning \(QML\), this challenge is amplified by noisy intermediate\-scale quantum \(NISQ\) hardware limitationsPre \([18](https://arxiv.org/html/2605.06734#bib.bib79)\)\. Consequently, deep, highly entangled quantum neural networks \(QNNs\) are difficult to execute reliably[A\+23b](https://arxiv.org/html/2605.06734#bib.bib2), costly to simulateC\+\([25](https://arxiv.org/html/2605.06734#bib.bib16)\), and hard to trainMBS\+\([18](https://arxiv.org/html/2605.06734#bib.bib70)\);[L\+25a](https://arxiv.org/html/2605.06734#bib.bib57), especially within recurrent or long\-horizon pipelinesC\+\([21](https://arxiv.org/html/2605.06734#bib.bib13)\); CVH\+\([22](https://arxiv.org/html/2605.06734#bib.bib34)\); B\+\([25](https://arxiv.org/html/2605.06734#bib.bib6)\)\. While hybrid variational quantum algorithms \(VQAs\)B\+\([22](https://arxiv.org/html/2605.06734#bib.bib4)\)have achieved breakthroughs in static domains like classificationBMB\+\([22](https://arxiv.org/html/2605.06734#bib.bib10)\);[L\+25e](https://arxiv.org/html/2605.06734#bib.bib61);[L\+25f](https://arxiv.org/html/2605.06734#bib.bib62); LPC\+\([25](https://arxiv.org/html/2605.06734#bib.bib66)\); CCL \([19](https://arxiv.org/html/2605.06734#bib.bib17)\); JHS\+\([25](https://arxiv.org/html/2605.06734#bib.bib51)\); CT \([25](https://arxiv.org/html/2605.06734#bib.bib33)\); CCT \([26](https://arxiv.org/html/2605.06734#bib.bib20)\); S\+\([22](https://arxiv.org/html/2605.06734#bib.bib83)\), generative modeling[H\+25b](https://arxiv.org/html/2605.06734#bib.bib43); S\+\([21](https://arxiv.org/html/2605.06734#bib.bib82)\); LW \([18](https://arxiv.org/html/2605.06734#bib.bib69)\);[CK25b](https://arxiv.org/html/2605.06734#bib.bib29);[CK25a](https://arxiv.org/html/2605.06734#bib.bib28)and mathematical problem\-solvingKPE \([21](https://arxiv.org/html/2605.06734#bib.bib56)\); PKAY \([22](https://arxiv.org/html/2605.06734#bib.bib78)\), extending them to sequential frameworks poses a severe computational bottleneck\. Quantum recurrent neural networks \(QRNNs\) require repeated circuit evaluations and backpropagation through time \(BPTT\) alongside expensive quantum gradient estimationWIWL \([22](https://arxiv.org/html/2605.06734#bib.bib96)\);[A\+23a](https://arxiv.org/html/2605.06734#bib.bib1)\. As sequence length \(window\-size\) grows, this training cost becomes prohibitiveBau \([20](https://arxiv.org/html/2605.06734#bib.bib7)\)\. Quantum Fast Weight Programmers \(QFWPs\)[Che24b](https://arxiv.org/html/2605.06734#bib.bib27)mitigate this burden by replacing hidden\-state dynamics with parameter dynamics\. In QFWP, a classical slow programmer generates the parameters of a fast quantum model at each time step, thereby avoiding explicit quantum gradient computation inside a recurrent loop\. However, existing QFWPs still rely on multi\-qubit circuits, limiting practical scalability in the NISQ era\. Recognizing these limitations, we shift our focus to aquantum\-inspiredparadigm that inherently bypasses the hardware constraints\. We proposegated QKAN\-FWP, integrating Quantum\-inspired Kolmogorov–Arnold Network \(QKAN\)JHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)into the fast\-weight programming framework\. QKAN utilizes single\-qubit data re\-uploading circuits as learnable nonlinear activations known as DatA Re\-Uploading ActivatioN \(DARUAN\)JHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\); SSM \([21](https://arxiv.org/html/2605.06734#bib.bib91)\); PSCLGFL \([20](https://arxiv.org/html/2605.06734#bib.bib80)\), circumventing multi\-qubit entanglement to provide expressive, hardware\-friendly, and simulation\-efficient modelingJHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)\. To further stabilize parameter evolution, we introduce a gated fast\-weight update rule\. By completely avoiding multi\-qubit entanglement bottlenecks, our architecture bridges the gap between quantum concepts and classical execution\. Therefore, we emphasize evaluating our model against classical baselines on practical tasks, specifically real\-world long\-horizon direct multi\-step forecasting—a capability that remains largely out of reach for prior quantum models constrained by NISQ limits\.
The main contributions of this work are as follows:
1. 1\.We propose gated QKAN\-FWP, a quantum\-inspired framework integrating QKAN modules with fast\-weight programming for efficient sequence modeling\.
2. 2\.We introduce a scalar\-gated fast\-weight mechanism that adaptively balances memory retention and new updates, with theoretical support through adaptive memory kernels, geometric bounds, and a parallelizable unrolled recursion that yields shallower gradient paths than general recurrent neural networks \(RNNs\)\.
3. 3\.We demonstrate strong empirical performance on real\-world multi\-step solar cycle forecasting, where our 12\.5k\-parameter model outperforms classical recurrent baselines spanning 11\.5k to 167k parameters \(up to 13× our model’s size\)\. We also evaluate comprehensively across time\-series benchmarks, and MiniGrid reinforcement learning \(RL\)\.
4. 4\.We validate NISQ compatibility by executing the trained fast programmer on two quantum processing units \(QPUs\), recovering forecasting performance within10−310^\{\-3\}relative Mean Square Error \(MSE\) of the noiseless simulator\.
## 2Related Work
##### Quantum sequence modeling and reinforcement learning
For sequential modeling, QRNNs and Quantum Long Short\-Term Memory \(QLSTM\) variants have been introduced to adapt quantum neural architectures for temporally dependent tasksBau \([20](https://arxiv.org/html/2605.06734#bib.bib7)\); CRP \([24](https://arxiv.org/html/2605.06734#bib.bib31)\); WG \([26](https://arxiv.org/html/2605.06734#bib.bib95)\); CYF \([22](https://arxiv.org/html/2605.06734#bib.bib35)\); CFD\+\([22](https://arxiv.org/html/2605.06734#bib.bib23)\); HCL\+\([25](https://arxiv.org/html/2605.06734#bib.bib44)\);[CCLL25b](https://arxiv.org/html/2605.06734#bib.bib19)\. Parallel to these developments, early quantum reinforcement learning \(QRL\) formulations assumed fully quantum environmentsDCLT \([08](https://arxiv.org/html/2605.06734#bib.bib38)\)\. Recent approaches instead utilize variational quantum circuits \(VQCs\) in classical environments with discrete or continuous observations[CCLL25a](https://arxiv.org/html/2605.06734#bib.bib18); SJD \([22](https://arxiv.org/html/2605.06734#bib.bib88)\); CYQ\+\([20](https://arxiv.org/html/2605.06734#bib.bib36)\); LS \([20](https://arxiv.org/html/2605.06734#bib.bib67)\); P\+\([24](https://arxiv.org/html/2605.06734#bib.bib75)\); D\+\([25](https://arxiv.org/html/2605.06734#bib.bib37)\)\. Furthermore, to overcome the limitations of partially observable environments, where agents must inherently track historical states, recent works have integrated QRNNs into RL policies[Che23b](https://arxiv.org/html/2605.06734#bib.bib25);[Che24a](https://arxiv.org/html/2605.06734#bib.bib26)\.
##### Fast\-weight programming and quantum extensions
Fast Weight Programmers \(FWPs\)Sch \([92](https://arxiv.org/html/2605.06734#bib.bib85),[93](https://arxiv.org/html/2605.06734#bib.bib86)\)replace recurrent hidden\-state evolution with dynamical evolution in parameter space\. A slow network updates the parameters of a fast network, enabling memory\-like behavior without explicit recurrence\. Subsequent classical work has combined FWPs with RNNsSS \([17](https://arxiv.org/html/2605.06734#bib.bib90)\)and established analogies to linear TransformersSIS \([21](https://arxiv.org/html/2605.06734#bib.bib87)\); ISCS \([21](https://arxiv.org/html/2605.06734#bib.bib47)\)\. QFWPs extend this paradigm by utilizing a parameterized quantum circuit as the fast programmer[Che24b](https://arxiv.org/html/2605.06734#bib.bib27)\. In QFWP, a classical slow network generates quantum circuit parameters on the fly, eliminating explicit quantum gradient computation inside the temporal loop\. To further reduce the parameter size, QT\-QFWP[L\+25c](https://arxiv.org/html/2605.06734#bib.bib59)uses a generative QNN to synthesize the slow programmer’s weights, leveraging quantum expressivity to address the scalability bottlenecks of classical slow networks\.
##### KAN and QKAN architectures
Kolmogorov–Arnold Networks \(KANs\) replace fixed activation functions in multilayer perceptrons \(MLPs\) with learnable univariate functions, yielding interpretable and parameter\-efficient nonlinear modeling[L\+25g](https://arxiv.org/html/2605.06734#bib.bib63); K\+\([24](https://arxiv.org/html/2605.06734#bib.bib54)\); LTM\+\([25](https://arxiv.org/html/2605.06734#bib.bib68)\); L\+\([26](https://arxiv.org/html/2605.06734#bib.bib64)\); S\+\([25](https://arxiv.org/html/2605.06734#bib.bib84)\); NWLDM \([25](https://arxiv.org/html/2605.06734#bib.bib73)\); YW \([25](https://arxiv.org/html/2605.06734#bib.bib100)\)\. This efficiency has motivated adaptation for temporal sequence modeling tasksHZLB \([25](https://arxiv.org/html/2605.06734#bib.bib45)\); J\+\([25](https://arxiv.org/html/2605.06734#bib.bib48)\); VRBPC \([24](https://arxiv.org/html/2605.06734#bib.bib93)\); XCW \([24](https://arxiv.org/html/2605.06734#bib.bib97)\); Liv \([24](https://arxiv.org/html/2605.06734#bib.bib65)\); YLZP \([25](https://arxiv.org/html/2605.06734#bib.bib99)\)\. QKAN extends the KAN architecture by implementing the edge functions with DARUANJHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)\. The resulting quantum\-inspired activations offer rich spectral expressivity while remaining lightweight and easily simulable\. Prior work[H\+25a](https://arxiv.org/html/2605.06734#bib.bib42)embeds QKAN inside the gates of a Long Short\-Term Memory \(LSTM\) cell to form QKAN\-LSTM\. Because its computation depends on the recurrent hidden stateht−1h\_\{t\-1\}, execution across the time dimension remains strictly sequential, and BPTT must traverse a chain ofTThidden\-state Jacobians\. In contrast, we deploy QKAN within a fast\-weight programmer\. Since the fast\-parameter updatesΔWk\\Delta W\_\{k\}depend solely on the inputxkx\_\{k\}rather than previous parametersWk−1W\_\{k\-1\}, we bypass the recurrent bottleneck, yielding shallower gradient paths \([Section˜5](https://arxiv.org/html/2605.06734#S5)\)\. This positions QKAN as a building block for fast\-weight programming, distinct from its nonlinear recurrent\-gate role in[H\+25a](https://arxiv.org/html/2605.06734#bib.bib42)\.
## 3Preliminaries
### 3\.1Quantum\-inspired Kolmogorov–Arnold Networks and Hybrid QKAN architecture
QKAN extends the KAN paradigm by replacing classical spline\-based edge functions with quantum\-inspired univariate functions realized by DARUANJHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\);[L\+25g](https://arxiv.org/html/2605.06734#bib.bib63)\. For an inputxx, each activation is defined as
ϕθ\(x\)=⟨0\|U†\(x;θ\)O^U\(x;θ\)\|0⟩,\\phi\_\{\\theta\}\(x\)=\\langle 0\|U^\{\\dagger\}\(x;\\theta\)\\hat\{O\}U\(x;\\theta\)\|0\\rangle,\(1\)whereU\(x;θ\)U\(x;\\theta\)is a parameterized single\-qubit data re\-uploading unitary andO^\\hat\{O\}is a measurement observable\. Repeating data re\-uploading induces a rich Fourier spectrum, enabling QKAN to represent highly nonlinear mappings with relatively few trainable parametersJHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)\. QKAN scales efficiently on CPUs, GPUs and HPC clusters, a property empirically validated by its use in large language models \(LLMs\)JHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)\. Beyond classical simulation efficiency, the strictly single\-qubit paradigm is compatible with current NISQ hardware, where state\-of\-the\-art platforms achieve single\-qubit error rates of10−510^\{\-5\}–10−710^\{\-7\}W\+\([25](https://arxiv.org/html/2605.06734#bib.bib94)\); R\+\([24](https://arxiv.org/html/2605.06734#bib.bib81)\); SLM\+\([25](https://arxiv.org/html/2605.06734#bib.bib89)\)\. In[Section˜6\.2\.1](https://arxiv.org/html/2605.06734#S6.SS2.SSS1), we confirm this compatibility by deploying our trained model on IonQ and IBM QPUs\.
We adopt the Hybrid QKAN \(HQKAN\) instantiation of the Jiang–Huang–Chen–Goan network \(JHCG Net\) first introduced inJHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)\. HQKAN has an encoder–processor–decoder structure: a classical encoder maps the input into a latent representation, a QKAN block performs nonlinear transformation in the latent space, and a decoder maps the transformed features to the output as illustrated in[Figure˜1](https://arxiv.org/html/2605.06734#S3.F1)\. Within our framework, HQKAN acts as a drop\-in programmer network\. When used as the slow programmer, it generates fast\-parameter updates from the current input\. When used as the fast programmer, its DARUAN parameters are dynamically updated by the slow programmer\.
Figure 1:HQKAN programmer architecture adapted fromJHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)\. The model consists of a classical encoder, a latent QKAN processor, and a decoder\. In this paper, HQKAN is used as a compact nonlinear programmer network inside the fast\-weight framework\.
### 3\.2Fast\-weight programming
##### FWP
FWPs model sequential data through dynamical evolution in parameter space rather than hidden\-state recurrence\. Letxtx\_\{t\}be the input at time steptt,S\(⋅\)S\(\\cdot\)the slow programmer, andF\(⋅;Wt\)F\(\\cdot;\\,W\_\{t\}\)the fast programmer with time\-dependent parametersWtW\_\{t\}\. The fast network producesyt=F\(xt;Wt\),y\_\{t\}=F\(x\_\{t\};W\_\{t\}\),while the slow programmer generates an updateΔWt=S\(xt\)\.\\Delta W\_\{t\}=S\(x\_\{t\}\)\.The fast parameters then evolve according to
Wt\+1=Wt\+ΔWt\.W\_\{t\+1\}=W\_\{t\}\+\\Delta W\_\{t\}\.\(2\)Temporal dependencies are therefore encoded in the trajectory of the fast parameters\{Wt\}\\\{W\_\{t\}\\\}\.
##### QFWP
In QFWP[Che24b](https://arxiv.org/html/2605.06734#bib.bib27), the fast programmer is a VQC\. A classical encoder mapsxtx\_\{t\}to two vectorsLt∈ℝlL\_\{t\}\\in\\mathbb\{R\}^\{l\}andQt∈ℝnQ\_\{t\}\\in\\mathbb\{R\}^\{n\}, corresponding to the number of circuit layerslland qubitsnn, respectively\. The update is formed as an outer product
ΔΘt=Lt⊗Qt,\(ΔΘt\)ij=Lt,iQt,j,\\Delta\\Theta\_\{t\}=L\_\{t\}\\otimes Q\_\{t\},\\qquad\(\\Delta\\Theta\_\{t\}\)\_\{ij\}=L\_\{t,i\}Q\_\{t,j\},which updates the quantum parametersΘt∈ℝl×n\\Theta\_\{t\}\\in\\mathbb\{R\}^\{l\\times n\}:Θt\+1=Θt\+ΔΘt\.\\Theta\_\{t\+1\}=\\Theta\_\{t\}\+\\Delta\\Theta\_\{t\}\.The model output is the expectation value of the fast VQC,
yt=⟨0\|U†\(Θt,xt\)O^U\(Θt,xt\)\|0⟩\.y\_\{t\}=\\langle 0\|U^\{\\dagger\}\(\\Theta\_\{t\},x\_\{t\}\)\\,\\hat\{O\}\\,U\(\\Theta\_\{t\},x\_\{t\}\)\|0\\rangle\.
## 4Methods
### 4\.1Gated fast\-weight update
A central contribution of this work is a gated update rule that stabilizes the evolution of the fast parameters\. At each time step, the slow programmer outputs the update components together with a scalar gategt∈\[0,1\]g\_\{t\}\\in\[0,1\]through a sigmoid nonlinearity\. The gate interpolates between the previously stored fast parameters and the newly generated update\. This mechanism is mathematically analogous to the “write\-strength” utilized in linear transformersSIS \([21](https://arxiv.org/html/2605.06734#bib.bib87)\); Y\+\([23](https://arxiv.org/html/2605.06734#bib.bib98)\), where a data\-dependent weight adaptively blends previous attention values with new updates\. While inSS \([17](https://arxiv.org/html/2605.06734#bib.bib90)\), a gated fast\-weight architecture was introduced for RNNs using an element\-wise matrix gate, our framework introduces a scalar gating mechanism\. This scalar approach ensures uniform parameter scaling, making it parameter\-efficient and naturally scalable\. For the fast parametersWtW\_\{t\}, our gated update is formulated as:
Wt\+1=gtWt\+\(1−gt\)ΔWt,gt∈\[0,1\]\.W\_\{t\+1\}=g\_\{t\}W\_\{t\}\+\(1\-g\_\{t\}\)\\Delta W\_\{t\},\\qquad g\_\{t\}\\in\[0,1\]\.\(3\)Intuitively, whengt→1g\_\{t\}\\to 1, the model retains its previously stored fast parameters, whereasgt→0g\_\{t\}\\to 0forces the model to rely entirely on the newly generated update\. We analyze these dynamics theoretically in[Section˜5](https://arxiv.org/html/2605.06734#S5)\.
Figure 2:Architectures of the proposed gated fast\-weight programmers\.\(a\) GQKAN\-FWP:An HQKAN slow programmer dynamically generates the parameters of a classical linear fast programmer following a gated update rule\.\(b\) GQKAN\-QKANFWP:Both programmers utilize HQKAN, with the slow programmer generating the DARUAN parameters for the fast module under the same gated mechanism\.
### 4\.2Model variants
To systematically evaluate our framework, we investigate the ablation variants summarized in[Table˜1](https://arxiv.org/html/2605.06734#S4.T1)\. For variants utilizing a classical fast programmer, the slow programmer produces update vectorsLt∈ℝlL\_\{t\}\\in\\mathbb\{R\}^\{l\},Dt∈ℝnD\_\{t\}\\in\\mathbb\{R\}^\{n\}, andBt∈ℝnB\_\{t\}\\in\\mathbb\{R\}^\{n\}\. In the ungated setting \(e\.g\., FWP\), the fast\-weightWtW\_\{t\}and biasbtb\_\{t\}are computed as:
Wt\+1=Wt\+Lt⊗Dtandbt\+1=bt\+Bt,W\_\{t\+1\}=W\_\{t\}\+L\_\{t\}\\otimes D\_\{t\}\\quad\\text\{and\}\\quad b\_\{t\+1\}=b\_\{t\}\+B\_\{t\},yielding the output:
yt=xtWt\+bt\.y\_\{t\}=x\_\{t\}W\_\{t\}\+b\_\{t\}\.Conversely, the gated variants \(e\.g\., G\-FWP, GQKAN\-FWP\) update these parameters according to[Equation˜3](https://arxiv.org/html/2605.06734#S4.E3)\. For models employing HQKAN as the fast programmer \(e\.g\., G\-QKANFWP, GQKAN\-QKANFWP\), letϕt\\phi\_\{t\}denote the fast\-parameters\. At each time step, the slow programmer generates the parameter updateΔϕt\\Delta\\phi\_\{t\}alongside the gategtg\_\{t\}\. The fast parameters then evolve via the gated mechanism:
ϕt\+1=gtϕt\+\(1−gt\)Δϕt,\\phi\_\{t\+1\}=g\_\{t\}\\phi\_\{t\}\+\(1\-g\_\{t\}\)\\Delta\\phi\_\{t\},and the prediction is produced by the fast HQKAN programmer:
yt=fHQKAN\(xt;ϕt\)\.y\_\{t\}=f\_\{\\mathrm\{HQKAN\}\}\(x\_\{t\};\\phi\_\{t\}\)\.
Structural illustrations of GQKAN\-FWP and GQKAN\-QKANFWP are presented in[Figure˜2](https://arxiv.org/html/2605.06734#S4.F2)\(a\) and \(b\), respectively\.
Table 1:Summary of model variants and their architectural components\.
## 5Theoretical Analysis
We provide a theoretical interpretation of the gated fast\-weight update which is the mechanism introduced in[Equation˜3](https://arxiv.org/html/2605.06734#S4.E3)\. This update is motivated as a way to interpolate between the previously stored fast parameters and the newly generated update\. The same analysis below also applies to the gated variants, after replacingWtW\_\{t\}by the corresponding fast parameters\. For comparison, the ungated fast\-weight recursion is given by[Equation˜2](https://arxiv.org/html/2605.06734#S3.E2), which accumulates all past updates additively\.
##### Unrolled form and adaptive memory kernel
By recursively expanding[Equation˜3](https://arxiv.org/html/2605.06734#S4.E3), we obtain
Wt\+1=\(∏s=1tgs\)W1\+∑k=1t\(1−gk\)\(∏s=k\+1tgs\)ΔWk\.W\_\{t\+1\}=\\left\(\\prod\_\{s=1\}^\{t\}g\_\{s\}\\right\)W\_\{1\}\+\\sum\_\{k=1\}^\{t\}\(1\-g\_\{k\}\)\\left\(\\prod\_\{s=k\+1\}^\{t\}g\_\{s\}\\right\)\\Delta W\_\{k\}\.\(4\)Therefore, the current fast parameters are a weighted aggregation of all past proposed fast states\{ΔWk\}k=1t\\\{\\Delta W\_\{k\}\\\}\_\{k=1\}^\{t\}, together with a decayed contribution from the initializationW1W\_\{1\}\. Define
β0,t:=∏s=1tgs,βk,t:=\(1−gk\)∏s=k\+1tgs,k=1,…,t\.\\beta\_\{0,t\}:=\\prod\_\{s=1\}^\{t\}g\_\{s\},\\quad\\beta\_\{k,t\}:=\(1\-g\_\{k\}\)\\prod\_\{s=k\+1\}^\{t\}g\_\{s\},\\quad k=1,\\dots,t\.\(5\)Sincegt∈\[0,1\]g\_\{t\}\\in\[0,1\], we haveβk,t≥0\\beta\_\{k,t\}\\geq 0for allkk, and one may verify by induction that
β0,t\+∑k=1tβk,t=1\.\\beta\_\{0,t\}\+\\sum\_\{k=1\}^\{t\}\\beta\_\{k,t\}=1\.\(6\)Hence[Equation˜4](https://arxiv.org/html/2605.06734#S5.E4)can be written as
Wt\+1=β0,tW1\+∑k=1tβk,tΔWk,W\_\{t\+1\}=\\beta\_\{0,t\}W\_\{1\}\+\\sum\_\{k=1\}^\{t\}\\beta\_\{k,t\}\\Delta W\_\{k\},\(7\)which shows that the gated dynamics implement an input\-dependent temporal kernel in parameter space\. This interpretation highlights a key distinction from the ungated update in[Equation˜2](https://arxiv.org/html/2605.06734#S3.E2), where every past update enters with a coefficient of11lacking a forgetting mechanism\. In contrast, under[Equation˜3](https://arxiv.org/html/2605.06734#S4.E3), the contribution ofΔWk\\Delta W\_\{k\}at timet\+1t\+1is weighted by
βk,t=\(1−gk\)∏s=k\+1tgs,\\beta\_\{k,t\}=\(1\-g\_\{k\}\)\\prod\_\{s=k\+1\}^\{t\}g\_\{s\},\(8\)which decays according to the subsequent gates\. Thus, the gated recursion supports both long\-memory and short\-memory behavior: when the subsequent gates remain close to11, older proposals are retained for many steps; when the gates are small, older proposals are rapidly forgotten\. In the special casegt≡gg\_\{t\}\\equiv g,[Equation˜8](https://arxiv.org/html/2605.06734#S5.E8)reduces to
βk,t=\(1−g\)gt−k,\\beta\_\{k,t\}=\(1\-g\)g^\{\\,t\-k\},\(9\)which is an exponential memory kernel\.
##### Geometric boundedness
A second useful consequence of[Equation˜7](https://arxiv.org/html/2605.06734#S5.E7)is thatWt\+1W\_\{t\+1\}lies in the convex hull of the set
𝒮t:=\{W1,ΔW1,…,ΔWt\}\.\\mathcal\{S\}\_\{t\}:=\\\{W\_\{1\},\\Delta W\_\{1\},\\dots,\\Delta W\_\{t\}\\\}\.\(10\)Therefore, for any norm∥⋅∥\\\|\\cdot\\\|,
‖Wt\+1‖\\displaystyle\\\|W\_\{t\+1\}\\\|≤β0,t‖W1‖\+∑k=1tβk,t‖ΔWk‖\\displaystyle\\leq\\beta\_\{0,t\}\\\|W\_\{1\}\\\|\+\\sum\_\{k=1\}^\{t\}\\beta\_\{k,t\}\\\|\\Delta W\_\{k\}\\\|≤max\{‖W1‖,‖ΔW1‖,…,‖ΔWt‖\}\.\\displaystyle\\leq\\max\\bigl\\\{\\\|W\_\{1\}\\\|,\\\|\\Delta W\_\{1\}\\\|,\\dots,\\\|\\Delta W\_\{t\}\\\|\\bigr\\\}\.\(11\)This provides a simple geometric boundedness property that the gated update cannot move the fast parameters outside the convex hull generated by the initialization and the historical proposals\. By contrast, the ungated recursion in[Equation˜2](https://arxiv.org/html/2605.06734#S3.E2)admits only the crude estimate
‖Wt\+1‖≤‖W1‖\+∑k=1t‖ΔWk‖,\\\|W\_\{t\+1\}\\\|\\leq\\\|W\_\{1\}\\\|\+\\sum\_\{k=1\}^\{t\}\\\|\\Delta W\_\{k\}\\\|,\(12\)which can grow linearly with the sequence length \(window\-size\) in the worst case\. Hence, whereas the ungated dynamics perform unconstrained additive accumulation, the gated dynamics replace this behavior by adaptive convex aggregation, yielding a built\-in forgetting mechanism together with a norm bound controlled by the historical proposals\.
##### Parallelizable parameter evolution and shallow gradient path
A further consequence of the unrolled form[Equation˜4](https://arxiv.org/html/2605.06734#S5.E4)is computational\. Since the slow programmer producesΔWk\\Delta W\_\{k\}and gategkg\_\{k\}directly fromxkx\_\{k\}alone, independent ofWk−1W\_\{k\-1\}, the sets\{\(ΔWk,gk\)\}k=1T\\\{\(\\Delta W\_\{k\},g\_\{k\}\)\\\}\_\{k=1\}^\{T\}for a sequence of lengthTTcan be computed in a single parallel pass\.
Observe that[Equation˜3](https://arxiv.org/html/2605.06734#S4.E3)is affine inWtW\_\{t\}with a*scalar*multiplier\. Writingat:=gt∈\[0,1\]a\_\{t\}:=g\_\{t\}\\in\[0,1\]andbt:=\(1−gt\)ΔWtb\_\{t\}:=\(1\-g\_\{t\}\)\\Delta W\_\{t\}, the recursion becomes
Wt\+1=atWt\+bt\.W\_\{t\+1\}=a\_\{t\}W\_\{t\}\+b\_\{t\}\.\(13\)The pairs\(gk,\(1−gk\)ΔWk\)\(g\_\{k\},\(1\-g\_\{k\}\)\\Delta W\_\{k\}\)compose under the associative rule
\(a′,b′\)∘\(a,b\)=\(a′a,a′b\+b′\),\(a^\{\\prime\},b^\{\\prime\}\)\\circ\(a,b\)=\(a^\{\\prime\}a,\\;a^\{\\prime\}b\+b^\{\\prime\}\),\(14\)so the trajectory\{Wt\}t=1T\\\{W\_\{t\}\\\}\_\{t=1\}^\{T\}can be resolved by a parallel prefix scanBle \([90](https://arxiv.org/html/2605.06734#bib.bib9)\); MC \([18](https://arxiv.org/html/2605.06734#bib.bib71)\)with
O\(Tp\+logp\)O\\\!\\left\(\\frac\{T\}\{p\}\+\\log p\\right\)\(15\)scan time onppprocessorsBle \([90](https://arxiv.org/html/2605.06734#bib.bib9)\), reducing toO\(logT\)O\(\\log T\)depth whenp=Θ\(T\)p=\\Theta\(T\), in contrast to theΩ\(T\)\\Omega\(T\)sequential depthMC \([18](https://arxiv.org/html/2605.06734#bib.bib71)\)of general nonlinear recurrent hidden\-state evolution\. Moreover, eachΔWk\\Delta W\_\{k\}factors through an independent forward pass of the slow programmer, so BPTT composes through products of scalar gates rather than a chain ofTTdense hidden\-state Jacobians as in QKAN\-LSTM[H\+25a](https://arxiv.org/html/2605.06734#bib.bib42)\.
##### Implication
The above analyses suggest that the gate plays three complementary roles: it induces an adaptive memory kernel, guarantees geometric boundedness of the fast parameters, and preserves the parallel, hidden\-state\-free structure of the FWP recursion\. Together, these properties help explain the empirically improved stability of the gated variants relative to their ungated counterparts\.
## 6Experimental Results
We evaluate the proposed framework on single\-step time\-series prediction, multi\-step real\-world forecasting and RL tasks\. To ensure robust and unbiased evaluation, all models across every experiment are independently trained and tested over five random seeds\. Furthermore, to provide a fair comparison of representational capacity, all quantum baselines are executed on classical simulators utilizing exact gradients computed via BPTT, without simulated hardware noise or finite measurement shots\. All quantum\-circuit simulation experiments are implemented using PennyLaneBIS\+\([18](https://arxiv.org/html/2605.06734#bib.bib8)\), PyTorchP\+\([19](https://arxiv.org/html/2605.06734#bib.bib74)\), and an open\-source QKAN implementation adapted fromJia \([25](https://arxiv.org/html/2605.06734#bib.bib52)\)111Available at[https://github\.com/Jim137/qkan](https://github.com/Jim137/qkan)\. To accelerate the QKAN framework, we adopt the PyTorch\-based efficient quantum\-circuit solver,FlashQKAN, introduced inJia \([25](https://arxiv.org/html/2605.06734#bib.bib52)\)\. By representing each QKAN layer as a tensor network and leveragingcuQuantumB\+\([23](https://arxiv.org/html/2605.06734#bib.bib5)\)to optimize the tensor\-contraction path, whilecuTileNVI \([25](https://arxiv.org/html/2605.06734#bib.bib72)\)is used for fused operator execution and block tiling to improve GPU throughput\. For the quantum hardware experiments in[Section˜6\.2\.1](https://arxiv.org/html/2605.06734#S6.SS2.SSS1), we execute the trained fast programmer on IonQ’sForte\-1trapped\-ion systemC\+\([24](https://arxiv.org/html/2605.06734#bib.bib15)\)using NVIDIA CUDA\-QK\+\([23](https://arxiv.org/html/2605.06734#bib.bib53)\)—a unified programming platform enabling seamless access to QPUs across modalities—with Amazon BraketAma \([20](https://arxiv.org/html/2605.06734#bib.bib3)\)as the access provider, and on the IBM Quantum superconducting Heron r3 processoribm\_aachenIBM \([26](https://arxiv.org/html/2605.06734#bib.bib46)\)via QiskitJA\+\([24](https://arxiv.org/html/2605.06734#bib.bib49)\)\.
### 6\.1Time\-series prediction
We evaluate the models on four benchmark datasets used in[Che24b](https://arxiv.org/html/2605.06734#bib.bib27)—Damped Simple Harmonic Motion \(SHM\), the Bessel function, and Nonlinear Auto\-Regressive Moving Average \(NARMA5 and NARMA10\)—and two additional datasets related to quantum dynamics: Delayed Quantum Control \(DQC\) and open quantum system Jaynes\-Cummings \(JC\) dynamics\. Across all tasks, we frame next\-step prediction as a sequential modeling problem\. Given an input sequence of sliding window\-size \(sequence\-length\)NNprevious observations\[xt−N,xt−N\+1,…,xt−1\]\[x\_\{t\-N\},x\_\{t\-N\+1\},\\dots,x\_\{t\-1\}\], the model processes each elementxτx\_\{\\tau\}one at a time forτ∈\[t−N,t−1\]\\tau\\in\[t\-N,t\-1\]\. After processing the full sequence, the model outputs a predictionyty\_\{t\}at the final time step, which is evaluated against the ground truthxtx\_\{t\}using MSE\. Each dataset is normalized to the range\[−1,1\]\[\-1,1\]and chronologically split into 80% training and 20% test data\. Each model is trained for 50 epochs with a batch size of 4 and a learning rate of1×10−31\\times 10^\{\-3\}\.[Table˜2](https://arxiv.org/html/2605.06734#S6.T2)summarizes each model’s trainable parameter counts for[Section˜6\.1](https://arxiv.org/html/2605.06734#S6.SS1)and[Section˜6\.3](https://arxiv.org/html/2605.06734#S6.SS3)\.
We evaluate the models in two stages\. Stage I fixes the input window\-size toN=16N=16as an ablation study to rank all variants under a common setting\. Stage II evaluates the top\-performing models across variable input window\-sizesN∈\{8,16,32,64\}N\\in\\\{8,16,32,64\\\}to test the models’ capacity to retain memory and capture both short\- and long\-range temporal dependencies\.
#### 6\.1\.1Datasets
Damped SHM\.Damped SHM is a standard benchmark for nonlinear function approximation\. We model the angular velocityθ˙\\dot\{\\theta\}of a damped pendulum governed by
d2θdt2\+bmdθdt\+gLsinθ=0,\\frac\{d^\{2\}\\theta\}\{dt^\{2\}\}\+\\frac\{b\}\{m\}\\frac\{d\\theta\}\{dt\}\+\\frac\{g\}\{L\}\\sin\\theta=0,whereg=9\.81g=9\.81,b=0\.15b=0\.15,L=1L=1, andm=1m=1, with initial conditionsθ\(0\)=0\\theta\(0\)=0andθ˙\(0\)=3\\dot\{\\theta\}\(0\)=3\.
Bessel function\.Bessel functions arise in many physical applications, such as wave propagation and heat conduction in cylindrical geometries\. The target is the second\-order Bessel function of the first kind,J2\(x\)J\_\{2\}\(x\), which satisfies
x2d2ydx2\+xdydx\+\(x2−α2\)y=0x^\{2\}\\frac\{d^\{2\}y\}\{dx^\{2\}\}\+x\\frac\{dy\}\{dx\}\+\(x^\{2\}\-\\alpha^\{2\}\)y=0with the series representation:
Jα\(x\)=∑m=0∞\(−1\)mm\!Γ\(m\+α\+1\)\(x2\)2m\+α\.J\_\{\\alpha\}\(x\)=\\sum\_\{m=0\}^\{\\infty\}\\frac\{\(\-1\)^\{m\}\}\{m\!\\,\\Gamma\(m\+\\alpha\+1\)\}\\left\(\\frac\{x\}\{2\}\\right\)^\{2m\+\\alpha\}\.
NARMA\.We use the standard NARMA5 \(n0=5n\_\{0\}=5\) and NARMA10 \(n0=10n\_\{0\}=10\) benchmarks following[Che24b](https://arxiv.org/html/2605.06734#bib.bib27)withM=300M=300timesteps generated from the recurrence:
yt\+1=αyt\+βyt∑j=0n0−1yt−j\+γut−n0\+1ut\+δ,y\_\{t\+1\}=\\alpha y\_\{t\}\+\\beta y\_\{t\}\\sum\_\{j=0\}^\{n\_\{0\}\-1\}y\_\{t\-j\}\+\\gamma u\_\{t\-n\_\{0\}\+1\}u\_\{t\}\+\\delta,where\(α,β,γ,δ\)=\(0\.3,0\.05,1\.5,0\.1\)\(\\alpha,\\beta,\\gamma,\\delta\)=\(0\.3,0\.05,1\.5,0\.1\)\. The input sequence is:
ut=0\.1\[sin\(2πα¯tT\)sin\(2πβ¯tT\)sin\(2πγ¯tT\)\+1\],u\_\{t\}=0\.1\\left\[\\sin\\left\(\\frac\{2\\pi\\bar\{\\alpha\}t\}\{T\}\\right\)\\sin\\left\(\\frac\{2\\pi\\bar\{\\beta\}t\}\{T\}\\right\)\\sin\\left\(\\frac\{2\\pi\\bar\{\\gamma\}t\}\{T\}\\right\)\+1\\right\],where\(α¯,β¯,γ¯,T\)=\(2\.11,3\.73,4\.11,100\)\(\\bar\{\\alpha\},\\bar\{\\beta\},\\bar\{\\gamma\},T\)=\(2\.11,3\.73,4\.11,100\)\.
DQC\.To evaluate the model’s capacity for long\-term temporal dependencies, we consider a non\-MarkovianFCB \([18](https://arxiv.org/html/2605.06734#bib.bib40)\)system of a two\-level atom \(qubit\) coupled to a semi\-infinite waveguide terminated by a mirror, inducing delayed quantum feedback via a bound state in the continuumCFBC \([19](https://arxiv.org/html/2605.06734#bib.bib22)\); TCK \([13](https://arxiv.org/html/2605.06734#bib.bib92)\)\. FollowingCYF \([22](https://arxiv.org/html/2605.06734#bib.bib35)\), we predict the output field intensityx\(t\)x\(t\), modeled as a sequence of localized pulses with decaying amplitude:
x\(t\)=∑n=010exp\[−10\(t−2n\)2\]exp\(−t16\)x\(t\)=\\sum\_\{n=0\}^\{10\}\\exp\\left\[\-10\(t\-2n\)^\{2\}\\right\]\\exp\\left\(\-\\frac\{t\}\{16\}\\right\)fort∈\[−2,20\]t\\in\[\-2,20\]\. This formulation captures the structured, non\-stationary nature of the delayed feedback response\.
Open Quantum System JC Dynamics\.To incorporate realistic environmental noise, we simulate the dynamics of an open quantum system based on the JC model using CUDA\-Q Dynamics, the open quantum system simulation backend of CUDA\-QK\+\([23](https://arxiv.org/html/2605.06734#bib.bib53)\)\. The system Hamiltonian is:
ℋ=ωca†a\+ωqσ\+σ−\+g\(σ−a†\+σ\+a\),\\mathcal\{H\}=\\omega\_\{c\}a^\{\\dagger\}a\+\\omega\_\{q\}\\sigma\_\{\+\}\\sigma\_\{\-\}\+g\(\\sigma\_\{\-\}a^\{\\dagger\}\+\\sigma\_\{\+\}a\),wherea†,aa^\{\\dagger\},aare bosonic creation and annihilation operators,σ\+,σ−\\sigma\_\{\+\},\\sigma\_\{\-\}are qubit raising and lowering operators,ωc=ωq=2π\\omega\_\{c\}=\\omega\_\{q\}=2\\pidenotes the resonant cavity and qubit frequencies, andg=πg=\\piis the coupling strength\. To capture non\-unitary evolution, we apply a collapse operatorC=γaC=\\sqrt\{\\gamma\}a\(decay rateγ=0\.05\\gamma=0\.05\) representing photon loss\. While the coupling ratiog/ω=0\.5g/\\omega=0\.5exceeds typical experimental values, it is chosen for simplicity and to yield a demanding benchmark curve that combines rapid oscillations with dissipative decay\. The system is initialized with the qubit in its ground state and a single cavity photon \(ρ0=\|g,1⟩⟨g,1\|\\rho\_\{0\}=\|g,1\\rangle\\langle g,1\|\), the target signal is the qubit excitation expectation probability⟨σ\+σ−⟩\(t\)\\langle\\sigma\_\{\+\}\\sigma\_\{\-\}\\rangle\(t\), evaluated over 3,000 discrete time steps fromt=0t=0tot=50t=50\.
Table 2:Trainable parameter counts for each model in the time\-series prediction and reinforcement\-learning experiments\.
#### 6\.1\.2Stage I: fixed window\-size evaluation
[Table˜3](https://arxiv.org/html/2605.06734#S6.T3)reports the final test loss at input window\-sizeN=16N=16\. HQKAN\-based gated models provide better overall balance across datasets\. In particular, GQKAN\-QKANFWP attains the best result on three of the six datasets, while GQKAN\-FWP and G\-QKANFWP each rank among the top two in multiple tasks\. In contrast, QFWP achieves the best result on NARMA10\. We advance GQKAN\-QKANFWP, G\-QKANFWP, GQKAN\-FWP and QFWP to Stage II\.
#### 6\.1\.3Stage II: variable window\-size evaluation
[Tables˜4](https://arxiv.org/html/2605.06734#S6.T4),[5](https://arxiv.org/html/2605.06734#S6.T5)and[6](https://arxiv.org/html/2605.06734#S6.T6)report the results, from which four trends emerge\. First, GQKAN\-QKANFWP exhibits the greatest robustness to varyingNN, attaining the lowest prediction error in 10 of the 24 settings\. Second, G\-QKANFWP dominates NARMA5 and NARMA10 at longer windows, indicating that an HQKAN fast programmer is well suited to long\-range nonlinear dependencies\. Third, GQKAN\-FWP leads on the quantum\-dynamics datasets \(DQC and JC\), where an HQKAN\-based slow programmer captures delayed feedback and dissipative noise effectively\. Fourth, while the QFWP attains an MSE of2×10−62\\times 10^\{\-6\}on NARMA10 atN=16N\{=\}16, it degrades to1\.3×10−41\.3\\times 10^\{\-4\}atN∈\{32,64\}N\\in\\\{32,64\\\}, a roughly60×60\\timescollapse\. Similar degradation trends are also shown in the quantum dynamics datasets[Table˜6](https://arxiv.org/html/2605.06734#S6.T6)\. Taken together, the results indicate that the gated QKAN\-FWP variants yield the most favorable balance between accuracy and stability across window sizes\. Among them, GQKAN\-QKANFWP exhibits the most consistent behavior across all three dataset families: it is best on every window\-size for the smooth\-dynamics benchmarks, best or second\-best on three of four window sizes for the quantum\-dynamics benchmarks, and on the NARMA benchmarks remains close to the leading variants\. This stability is consistent with the theoretical properties of the gated memory mechanism \([Section˜5](https://arxiv.org/html/2605.06734#S5)\) and the spectral expressivity of the HQKAN architecture \([Section˜3\.1](https://arxiv.org/html/2605.06734#S3.SS1)\), and is precisely the property required for real\-world forecasting, which motivates our selection of GQKAN\-QKANFWP for the real\-world forecasting study in[Section˜6\.2](https://arxiv.org/html/2605.06734#S6.SS2)\.
Table 3:Final test loss \(MSE, mean±\\pmstd over 5 seeds\) with window\-sizeN=16N=16\. Best/second\-best results are shown inbold/underlined\.Table 4:Final test loss \(MSE, mean±\\pmstd over 5 seeds\) on theBessel functionandDamped SHMdatasets\. Best/second\-best results are shown inbold/underlined\.Table 5:Final test loss \(MSE, mean±\\pmstd over 5 seeds\) on theNARMA5andNARMA10datasets\. Best/second\-best results are shown inbold/underlined\.Table 6:Final test loss \(MSE, mean±\\pmstd over 5 seeds\) on theDelayed Quantum ControlandJaynes\-Cummingsdatasets\. Best/second\-best results are shown inbold/underlined\.
### 6\.2Real\-World Direct Multi\-Step Prediction
While[Section˜6\.1](https://arxiv.org/html/2605.06734#S6.SS1)demonstrates our models’ ability to capture synthetic dynamics via single\-step prediction, practical applications often demand long\-horizon, multi\-step forecasting in complex, non\-stationary environments\. To evaluate this, we apply GQKAN\-QKANFWP to solar cycle forecasting, a well\-known challenge in solar physicsBN \([18](https://arxiv.org/html/2605.06734#bib.bib11)\); Pet \([20](https://arxiv.org/html/2605.06734#bib.bib77)\); Pes \([08](https://arxiv.org/html/2605.06734#bib.bib76)\), and conclude this section with an inference test on available real quantum hardware[Section˜6\.2\.1](https://arxiv.org/html/2605.06734#S6.SS2.SSS1)\. We use 3,326 monthly averaged sunspot records spanning 1749–2026 from the World Data Center SILSO222Data available at[http://sidc\.be/silso/datafiles](http://sidc.be/silso/datafiles)\.CL \([16](https://arxiv.org/html/2605.06734#bib.bib30)\)\. FollowingBPP\+\([20](https://arxiv.org/html/2605.06734#bib.bib12)\), we frame this as a univariate multi\-step forecasting task: a sliding window maps a 528\-month input \(roughly four solar cycles\) to a 132\-month forecast horizon \(one cycle\)\. The output at the final time step serves as our prediction\. To prioritize the accurate prediction of solar cycle maximaBN \([18](https://arxiv.org/html/2605.06734#bib.bib11)\); Pet \([20](https://arxiv.org/html/2605.06734#bib.bib77)\), we optimize a peak\-aware MSE loss:
ℒ=1B∑\(𝐲−𝐲^\)2\(1\+α𝐲\),\\mathcal\{L\}=\\frac\{1\}\{B\}\\sum\(\\mathbf\{y\}\-\\mathbf\{\\hat\{y\}\}\)^\{2\}\(1\+\\alpha\\mathbf\{y\}\),whereBBis the batch size andα=1\.0\\alpha=1\.0scales the penalty for peak values\. We also report two additional metrics to quantify peak prediction accuracy in terms of absolute amplitude difference and temporal displacementBN \([18](https://arxiv.org/html/2605.06734#bib.bib11)\): peak amplitude error, which measures the absolute difference in sunspot numbers:
PAE=1M∑i=1M\|max\(𝐲\(i\)\)−max\(𝐲^\(i\)\)\|\\mathrm\{PAE\}=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}\\left\|\\max\(\\mathbf\{y\}^\{\(i\)\}\)\-\\max\(\\mathbf\{\\hat\{y\}\}^\{\(i\)\}\)\\right\|and Peak Timing Error, which measures the temporal displacement in months:
PTE=1M∑i=1M\|argmax\(𝐲\(i\)\)−argmax\(𝐲^\(i\)\)\|,\\mathrm\{PTE\}=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}\\left\|\\operatorname\{argmax\}\(\\mathbf\{y\}^\{\(i\)\}\)\-\\operatorname\{argmax\}\(\\mathbf\{\\hat\{y\}\}^\{\(i\)\}\)\\right\|,whereMMis the number of test sequences, and𝐲\(i\),𝐲^\(i\)\\mathbf\{y\}^\{\(i\)\},\\mathbf\{\\hat\{y\}\}^\{\(i\)\}are the ground\-truth and predicted sequences, respectively\.
We benchmark GQKAN\-QKANFWP against the WaveNet\-LSTM and LSTM baselines fromBPP\+\([20](https://arxiv.org/html/2605.06734#bib.bib12)\)\(LSTM\-L withH=132H=132, LSTM\-S withH=64H=64\), as well as the Vanilla RNN and Modified Echo State Network \(MESN\) baselines fromEFGC\+\([23](https://arxiv.org/html/2605.06734#bib.bib39)\)\. To isolate architectural capacity from training\-protocol confounds, we re\-run all baselines under a single standardized protocol rather than strictly replicating prior setups\. Using the raw monthly data normalized to\[0,1\]\[0,1\], a 528\-month input window, and a 132\-month forecast horizon, we chronologically partition the dataset into 80% training, 10% validation, and 10% test sets\. All baselines except MESN use batch normalization, 30% dropout, a batch size of 32, and 100 epochs, evaluated across five random seeds\. For these models, learning rates are individually tuned from\{1,2,2\.5\}×10−3\\\{1,2,2\.5\\\}\\times 10^\{\-3\}on the validation split, and the checkpoint with the lowest average validation loss is used for testing\. For MESN, we adopt the reservoir configuration ofEFGC\+\([23](https://arxiv.org/html/2605.06734#bib.bib39)\)unchanged, except for setting the output horizon to 132 instead of 129 to match our forecast window\. Its input delay embedding is drawn from the tail of the same 528\-month input window fed to the other baselines, so the test split is identical across all models\. Because MESN is fit in one pass by closed\-form weighted ridge regression, it has no learning rate or epoch budget\. Because our protocol evaluates long, raw input sequences rather than the 13\-month smoothed data and variable windows used inEFGC\+\([23](https://arxiv.org/html/2605.06734#bib.bib39)\), the baseline metrics reported here reflect performance under stricter conditions\. Similarly, our WaveNet\-LSTM reproduction differs quantitatively fromBPP\+\([20](https://arxiv.org/html/2605.06734#bib.bib12)\)due to our peak\-aware loss and our chronological, strictly unseen test split rather than their 5\-fold cross\-validation scheme\.
Figure 3:Forecasting Solar Cycle 23 \(test set\)\.\(a\)Mean forecasts and shaded±1σ\\pm 1\\sigmabands across 5 random seeds for each model\. While the ground truth \(black\) exhibits substantial month\-to\-month variability, the GQKAN\-QKANFWP±1σ\\pm 1\\sigmaenvelope \(orange shading\) contains the ground truth throughout the rising, peak, and descending phases of the cycle\. Among the baselines, only LSTM\-L produces a mean prediction that overlaps the GQKAN\-QKANFWP envelope, yet uses approximately 7×\\timesmore parameters \(see[Table˜7](https://arxiv.org/html/2605.06734#S6.T7)\)\. The remaining baselines either systematically under\-predict the peak or fail to produce a coherent cycle\.\(b\)Full context: four preceding solar cycles followed by the forecast window \(shaded\)\.Table 7:Solar cycle forecasting performance across models\. Baseline architectures followBPP\+\([20](https://arxiv.org/html/2605.06734#bib.bib12)\)\(LSTM, WaveNet\-LSTM\) andEFGC\+\([23](https://arxiv.org/html/2605.06734#bib.bib39)\)\(Vanilla RNN, MESN\)\. Gradient\-based baselines use their individually tuned learning rates; MESN is fit by closed\-form weighted ridge regression and has no learning rate\. Best results inbold, second\-bestunderlined\. Values are mean±\\pmstd over 5 seeds\.Figure 4:Solar cycle forecasting results for GQKAN\-QKANFWP\.Orange markers represent continuous one\-step\-ahead forecasts stacking the first step of the 132\-step horizon across overlapping sliding windows, while the red curves denote full 132\-step predictions on Solar Cycle 22 and the ongoing Solar Cycle 25 generated from a single input window\. The “Test Split” line marks the beginning of the test set\. Additionally, the “History Ends” line indicates the separation between the available historical data and the model’s future prediction\.As summarized in[Table˜7](https://arxiv.org/html/2605.06734#S6.T7), GQKAN\-QKANFWP attains the lowest scaled MSE, PAE, and PTE among all evaluated models\. This suggests that its advantage reflects a joint improvement in overall fit and peak prediction rather than a trade\-off between them\. This is achieved with only 12\.5k parameters, roughly77–13×13\\timesfewer than the competitive baselines \(LSTM\-L: 89k; MESN: 132k; WaveNet\-LSTM: 167k\)\. The model also exhibits the lowest seed\-to\-seed variance on scaled MSE \(±0\.0016\\pm 0\.0016\) and the second\-lowest variance on PAE and PTE \(after LSTM\-S\), indicating the reported gains are stable rather than seed\-dependent\.[Figure˜3](https://arxiv.org/html/2605.06734#S6.F3)visualizes the per\-model performance on Solar Cycle 23 \(SC23\) as seed\-averaged forecasts with±1σ\\pm 1\\sigmabands\. GQKAN\-QKANFWP’s±1σ\\pm 1\\sigmaenvelope \(orange shading\) contains the ground truth throughout the rising, peak, and descending phases of the cycle\. LSTM\-L is the only baseline whose mean prediction overlaps this envelope throughout the cycle but has approximately 7×\\timesmore parameters\. The remaining baselines either systematically under\-predict the solar maximum or fail to form a coherent cycle\.[Figure˜4](https://arxiv.org/html/2605.06734#S6.F4)illustrates the overall forecasting behavior: orange markers denote continuous one\-step\-ahead forecasts obtained by stacking the first value of each 132\-month horizon across overlapping sliding windows, while red curves show the full 132\-step predictions for SC22 and the ongoing SC25, each generated from a single input window in the test set\. The model captures the macroscopic cycle structure on SC22 and produces a stable projection for SC25\. Overall, these results suggest that GQKAN\-QKANFWP can process long input sequences \(528 months\) and produce direct multi\-step forecasts over a long horizon \(132 months\) while maintaining low overall MSE and preserving the amplitude and timing of the cycle maxima, a regime in which substantially larger classical recurrent baselines, under the same training protocol, tend to degrade on at least one of these axes\.
Table 8:Execution of GQKAN\-QKANFWP’s fast programmer on QPUs\. Relative MSE is the MSE with respect to the noiseless simulator output horizon\.#### 6\.2\.1Execution on real quantum hardware
To validate NISQ compatibility, we deploy the fast programmer of a trained GQKAN\-QKANFWP onto IonQ’sForte\-1and IBM’sibm\_aachenQPUs\. The slow programmer and the gated fast\-parameter recursion are evaluated classically, while the fast programmer runs on the QPUs, isolating the DARUAN module to quantify noise impact without retraining\. The model utilizes 200 single\-qubit DARUAN circuits, which execute in parallel\. On the 156\-qubitibm\_aachen\(Heron r3\), we selected 100 qubits via a composite calibration score \(readout error∈\[2\.6,9\.0\]×10−3\\in\[2\.6,9\.0\]\\times 10^\{\-3\}, SX error∈\[0\.8,6\.3\]×10−4\\in\[0\.8,6\.3\]\\times 10^\{\-4\}\), applying Qiskit’soptimization\_level=1without further error mitigation\. For the 36\-qubitForte\-1, we uniformly pinned the first 20 qubits as Braket’s device\-level aggregates \(single\-qubit randomized\-benchmarking fidelity≈0\.9998\\approx 0\.9998, SPAM fidelity≈0\.9937\\approx 0\.9937\) preclude meaningful per\-qubit ranking\. We evaluate the full shot sweepN∈\{1,16,64,256,1024\}N\\in\\\{1,16,64,256,1024\\\}onibm\_aachento characterize the convergence\-with\-shots behavior, and report the converged high\-shot settingN=1024N=1024onForte\-1as an independent cross\-platform check\. We report scaled MSE, PAE, PTE, and the relative MSE with respect to the noiseless simulator\. As shown in[Table˜8](https://arxiv.org/html/2605.06734#S6.T8)and[Figure˜5](https://arxiv.org/html/2605.06734#S6.F5), GQKAN\-QKANFWP’s forecasts on both QPUs converge within∼8\.5×10−4\\sim 8\.5\\times 10^\{\-4\}of relative MSE atN=1024N=1024shots—saturating the𝒪\(1/N\)∼10−3\\mathcal\{O\}\(1/N\)\\sim 10^\{\-3\}statistical floor imposed by shot\-noise\-limited expectation\-value estimationGLM \([04](https://arxiv.org/html/2605.06734#bib.bib41)\); KOS \([07](https://arxiv.org/html/2605.06734#bib.bib55)\)\.
Figure 5:Forecasting Solar Cycle 24 from GQKAN\-QKANFWP’s fast programmer executed on QPUs\.\(a\)Forte\-1atN=1024N=1024shots\. \(b\)ibm\_aachenacross shot countsN∈\{1,16,64,256,1024\}N\\in\\\{1,16,64,256,1024\\\}; forecasts converge to the noiseless simulator asNNincreases, recovering cycle shape and peak within∼10−3\\sim 10^\{\-3\}relative MSE atN=1024N\{=\}1024\.
### 6\.3Reinforcement learning
Following the benchmark of[Che24b](https://arxiv.org/html/2605.06734#bib.bib27), we evaluate RL agents on the MiniGrid\-Empty environmentC\+\([23](https://arxiv.org/html/2605.06734#bib.bib14)\)trained with asynchronous advantage actor\-critic \(A3C\)[Che23a](https://arxiv.org/html/2605.06734#bib.bib24)\. Since QFWP has been shown to converge substantially faster than QLSTM[Che24b](https://arxiv.org/html/2605.06734#bib.bib27); CRPC \([26](https://arxiv.org/html/2605.06734#bib.bib32)\), we restrict the comparison to the FWP variants\. At each step the agent receives a 147\-dimensional observation \(7×7×37\{\\times\}7\{\\times\}3flattened local viewport\) and selects among seven discrete actions, with sparse rewardR=1−0\.9stepsmax stepsR=1\-0\.9\\,\\frac\{\\text\{steps\}\}\{\\text\{max steps\}\}on reaching the goal \(max steps=4n2\\text\{max steps\}=4n^\{2\}for ann×nn\{\\times\}ngrid\) and zero otherwise\. We evaluaten∈\{5,6,8,16\}n\\in\\\{5,6,8,16\\\}as shown in[Figure˜6](https://arxiv.org/html/2605.06734#S6.F6)\. Each model is trained for 10,000 episodes with 80 workers, learning rate1×10−41\{\\times\}10^\{\-4\},β1=0\.92\\beta\_\{1\}\{=\}0\.92,β2=0\.999\\beta\_\{2\}\{=\}0\.999, rollout lengthL=5L\{=\}5, and a discount factorγ=0\.9\\gamma\{=\}0\.9\. Reported rewards are smoothed by a 100\-episode moving average and averaged across workers and seeds\.
Figure 6:MiniGrid\-Empty environments\.The agents are evaluated across environments of increasing scale: \(a\) 5×\\times5, \(b\) 6×\\times6, \(c\) 8×\\times8, and \(d\) 16×\\times16\.[Figure˜7](https://arxiv.org/html/2605.06734#S6.F7)\(a\) shows that adding the gate generally improves both convergence stability and final performance\. Final rewards across all environment sizes are reported in[Table˜9](https://arxiv.org/html/2605.06734#S6.T9), with full learning curves in[Figure˜7](https://arxiv.org/html/2605.06734#S6.F7)\(b\)–\(e\)\. Ungated QFWP degrades substantially as the grid grows, while the gated variants remain stable\. Among them, G\-QKANFWP attains the highest or second\-highest final reward in the larger environments \(6×66\{\\times\}6,8×88\{\\times\}8, and16×1616\{\\times\}16\) despite slightly slower early convergence, indicating that the HQKAN fast programmer retains capacity in more complex state spaces\. The fully HQKAN\-based GQKAN\-QKANFWP reaches competitive rewards with substantially fewer parameters: on the16×1616\{\\times\}16task it achieves0\.974±0\.0010\.974\\pm 0\.001with only 1,114 trainable parameters, versus0\.975±0\.0010\.975\\pm 0\.001for G\-FWP at 2,665 parameters, a∼\\sim58% reduction at essentially matched performance\.[Figure˜7](https://arxiv.org/html/2605.06734#S6.F7)\(e\) further shows that GQKAN\-FWP converges slightly faster than classical G\-FWP on the16×1616\{\\times\}16grid, suggesting a training\-efficiency benefit from the quantum\-inspired slow programmer even when the fast programmer remains classical\.
Figure 7:Model performance on MiniGrid\-Empty environments\.The curves show mean episodic reward with shaded regions denoting standard deviation across 5 seeds\.\(a\)Gated\-vs\-ungated ablation on the5×55\{\\times\}5grid: gated architectures yield higher stability and asymptotic rewards than their ungated counterparts\.\(b\)–\(e\)Scaling across5×55\{\\times\}5,6×66\{\\times\}6,8×88\{\\times\}8, and16×1616\{\\times\}16grids for the top\-performing variants\.Table 9:Final reward on MiniGrid\-Empty tasks \(mean±\\pmstd over five seeds\)\. Higher is better\. Best/second\-best results in each column are shown inbold/underlined\.
## 7Conclusion
We presented gated QKAN\-FWP, a quantum\-inspired sequence learning framework that mitigates the scalability and execution bottlenecks inherent to the NISQ era\. By relying exclusively on HQKAN—modules empirically shown capable of scaling to LLMsJHCG \([25](https://arxiv.org/html/2605.06734#bib.bib50)\)—our framework inherits strong scalability while circumventing the costs of multi\-qubit entanglement\. A scalar\-gated fast\-weight update further stabilizes parameter evolution, a property we interpret theoretically through adaptive memory kernels, geometric boundedness, and a parallel\-scan\-compatible recursion\. Empirical evaluations highlight the architecture’s versatility and parameter efficiency\. In time\-series prediction, HQKAN\-based gated variants exhibited the greatest robustness over extended input windows\. On real\-world solar cycle forecasting, our 12\.5k\-parameter GQKAN\-QKANFWP achieved lower scaled MSE, PAE, and PTE than a suite of classical recurrent baselines spanning 11\.5k to 167k parameters—up to 13× larger than ours—including LSTM\-L, WaveNet\-LSTM, and MESN\. Furthermore, deploying the trained fast programmer on IonQ’sForte\-1and IBM’sibm\_aachenrecovered forecasting accuracy within∼10−3\\sim 10^\{\-3\}relative MSE of the simulator at 1024 shots, confirming the NISQ compatibility of the single\-qubit design\. In MiniGrid RL task, the framework achieved competitive performance with a 58% parameter reduction relative to prior baselines\.
Although original KAN architectures face optimization challenges in ultra\-large\-scale scenariosNWLDM \([25](https://arxiv.org/html/2605.06734#bib.bib73)\); YW \([25](https://arxiv.org/html/2605.06734#bib.bib100)\), our framework structurally mitigates this burden\. Processing sequences autoregressively keeps the HQKAN input dimension independent of sequence length, and operating within HQKAN’s reduced\-dimensional latent space compresses computational overhead; for tasks with ultra\-large input/output dimensions, this overhead can be further managed via structural groupingYW \([25](https://arxiv.org/html/2605.06734#bib.bib100)\); Jia \([25](https://arxiv.org/html/2605.06734#bib.bib52)\)\. Pushing these dimensional limits will therefore drive our future work, alongside analyzing optimization dynamics and extending execution on physical quantum hardware beyond inference\.
## Acknowledgment
K\.\-C\. Peng, J\.\-C\. Jiang, Y\.\-C\. Hsu and C\.\-H\. Lin thank the National Center for High\-Performance Computing \(NCHC\), National Institutes of Applied Research \(NIAR\), Taiwan, for providing computational and storage resources supported by the National Science and Technology Council \(NSTC\), Taiwan, under Grants No\. NSTC 114\-2119\-M\-007\-013\. H\.\-S\. Goan acknowledges support from the NSTC, Taiwan, under Grants No\. NSTC 113\-2112\-M\-002\-022\-MY3, No\. NSTC 113\-2119\-M\-002\-021, No\. NSTC 114\-2119\-M\-002\-018, No\. NSTC 114\-2119\-M\-002\-017\-MY3, and from the National Taiwan University under Grants No\. NTU\-CC\-115L8937, No\. NTU\-CC\-115L893704 and No\. NTU\-CC\-115L8512\. H\.\-S\. Goan is also grateful for the support of the “Center for Advanced Computing and Imaging in Biomedicine \(NTU\-115L900702\)” through the Featured Areas Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education \(MOE\), Taiwan, the support of Taiwan Semiconductor Research Institute \(TSRI\) through the Joint Developed Project \(JDP\) and the support from the Physics Division, National Center for Theoretical Sciences, Taiwan\. EJK acknowledges financial support from the NSTC of Taiwan under Grant No\. NSTC 114\-2112\-M\-A49\-036\-MY3\. The authors acknowledge the National Taiwan University–IBM Quantum Hub \(NTU–IBM Q Hub\) and Cloud Computing Center for Quantum Science & Technology at National Cheng Kung University for providing IBM Q system and Amazon Braket platforms\.
## References
- \[1\]Amira Abbas et al\.On quantum backpropagation, information reuse, and cheating measurement collapse\.Advances in Neural Information Processing Systems, 36:44792–44819, 2023\.
- \[2\]Amira Abbas et al\.Quantum optimization: Potential, challenges, and the path forward\.arXiv preprint arXiv:2312\.02279, 2023\.
- Ama \[20\]Amazon Web Services\.Amazon Braket\.[https://aws\.amazon\.com/braket/](https://aws.amazon.com/braket/), 2020\.Accessed: 2026\-04\-22\.
- B\+\[22\]Kishor Bharti et al\.Noisy intermediate\-scale quantum algorithms\.Reviews of Modern Physics, 94\(1\):015004, 2022\.
- B\+\[23\]Harun Bayraktar et al\.cuQuantum SDK: A High\-Performance Library for Accelerating Quantum Science\.In2023 IEEE International Conference on Quantum Computing and Engineering \(QCE\), volume 01, pages 1050–1061, 2023\.
- B\+\[25\]Ryan Babbush et al\.The grand challenge of quantum applications\.arXiv preprint arXiv:2511\.09124, 2025\.
- Bau \[20\]Johannes Bausch\.Recurrent quantum neural networks\.Advances in neural information processing systems, 33:1368–1379, 2020\.
- BIS\+\[18\]Ville Bergholm, Josh Izaac, Maria Schuld, Christian Gogolin, Shahnawaz Ahmed, Vishnu Ajith, M Sohaib Alam, Guillermo Alonso\-Linaje, B AkashNarayanan, Ali Asadi, et al\.Pennylane: Automatic differentiation of hybrid quantum\-classical computations\.arXiv preprint arXiv:1811\.04968, 2018\.
- Ble \[90\]Guy E Blelloch\.Prefix sums and their applications\.1990\.
- BMB\+\[22\]Denis Bokhan, Alena S Mastiukova, Aleksey S Boev, Dmitrii N Trubnikov, and Aleksey K Fedorov\.Multiclass classification using quantum convolutional neural networks with hybrid quantum\-classical learning\.Frontiers in Physics, 10:1069985, 2022\.
- BN \[18\]Prantika Bhowmik and Dibyendu Nandy\.Prediction of the strength and timing of sunspot cycle 25 reveal decadal\-scale space environmental conditions\.Nature communications, 9\(1\):5209, 2018\.
- BPP\+\[20\]B Benson, WD Pan, A Prasad, GA Gary, and Q Hu\.Forecasting solar cycle 25 using deep neural networks\.Solar Physics, 295\(5\):65, 2020\.
- C\+\[21\]Marco Cerezo et al\.Variational quantum algorithms\.Nature Reviews Physics, 3\(9\):625–644, 2021\.
- C\+\[23\]Maxime Chevalier\-Boisvert et al\.Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal\-oriented tasks\.InAdvances in Neural Information Processing Systems 36, New Orleans, LA, USA, December 2023\.
- C\+\[24\]Jwo\-Sy Chen et al\.Benchmarking a trapped\-ion quantum computer with 30 qubits\.Quantum, 8:1516, 2024\.
- C\+\[25\]Marco Cerezo et al\.Does provable absence of barren plateaus imply classical simulability?Nature Communications, 16\(1\):7907, 2025\.
- CCL \[19\]Iris Cong, Soonwon Choi, and Mikhail D Lukin\.Quantum convolutional neural networks\.Nature Physics, 15\(12\):1273–1278, 2019\.
- \[18\]Kuan\-Cheng Chen, Samuel Yen\-Chi Chen, Chen\-Yu Liu, and Kin K Leung\.Quantum\-train\-based distributed multi\-agent reinforcement learning\.In2025 IEEE Symposium for Multidisciplinary Computational Intelligence Incubators \(MCII Companion\), pages 1–5\. IEEE, 2025\.
- \[19\]Kuan\-Cheng Chen, Samuel Yen\-Chi Chen, Chen\-Yu Liu, and Kin K Leung\.Toward large\-scale distributed quantum long short\-term memory with modular quantum computers\.In2025 International Wireless Communications and Mobile Computing \(IWCMC\), pages 337–342\. IEEE, 2025\.
- CCT \[26\]Chi\-Sheng Chen, Samuel Yen\-Chi Chen, and Hsin\-Hsiung Tseng\.Exploring the potential of QEEGNet for cross\-task and cross\-dataset electroencephalography encoding with quantum machine learning\.Journal of Signal Processing Systems, 98:5, 2026\.
- CCTW \[24\]Chi\-Sheng Chen, Samuel Yen\-Chi Chen, Aidan Hung\-Wen Tsai, and Chun\-Shu Wei\.QEEGNet: Quantum machine learning for enhanced electroencephalography encoding\.In2024 IEEE Workshop on Signal Processing Systems \(SiPS\), pages 153–158\. IEEE, 2024\.
- CFBC \[19\]Giuseppe Calajó, Yao\-Lung L Fang, Harold U Baranger, and Francesco Ciccarello\.Exciting a bound state in the continuum through multiphoton scattering plus delayed quantum feedback\.Physical review letters, 122\(7\):073601, 2019\.
- CFD\+\[22\]Samuel Yen\-Chi Chen, Daniel Fry, Amol Deshmukh, Vladimir Rastunkov, and Charlee Stefanski\.Reservoir computing via quantum recurrent neural networks\.arXiv preprint arXiv:2211\.02612, 2022\.
- \[24\]Samuel Yen\-Chi Chen\.Asynchronous training of quantum reinforcement learning\.Procedia Computer Science, 222:321–330, 2023\.
- \[25\]Samuel Yen\-Chi Chen\.Quantum deep recurrent reinforcement learning\.InICASSP 2023\-2023 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\), pages 1–5\. IEEE, 2023\.
- \[26\]Samuel Yen\-Chi Chen\.Efficient quantum recurrent reinforcement learning via quantum reservoir computing\.InICASSP 2024\-2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\), pages 13186–13190\. IEEE, 2024\.
- \[27\]Samuel Yen\-Chi Chen\.Learning to program variational quantum circuits with fast weights\.In2024 International Joint Conference on Neural Networks \(IJCNN\), pages 1–9\. IEEE, 2024\.
- \[28\]Chi\-Sheng Chen and En\-Jui Kuo\.Quantum\-enhanced natural language generation: A multi\-model framework with hybrid quantum\-classical architectures\.arXiv preprint arXiv:2508\.21332, 2025\.
- \[29\]Chi\-Sheng Chen and En\-Jui Kuo\.Quantum reinforcement learning\-guided diffusion model for image synthesis via hybrid quantum\-classical generative model architectures\.arXiv preprint arXiv:2509\.14163, 2025\.
- CL \[16\]Frédéric Clette and Laure Lefèvre\.The new sunspot number: assembling all corrections\.Solar Physics, 291\(9\):2629–2651, 2016\.
- CRP \[24\]Andrea Ceschini, Antonello Rosato, and Massimo Panella\.A variational approach to quantum gated recurrent units\.Journal of Physics Communications, 8\(8\):085004, 2024\.
- CRPC \[26\]Andrea Ceschini, Antonello Rosato, Massimo Panella, and Samuel Yen\-Chi Chen\.Quantum fast weight programming for time series prediction\.InICASSP 2026\-2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\), pages 22032–22036\. IEEE, 2026\.
- CT \[25\]Chi\-Sheng Chen and Aidan Hung\-Wen Tsai\.Quantum adaptive self\-attention for financial rebalancing: An empirical study on automated market makers in decentralized finance\.arXiv preprint arXiv:2509\.16955, 2025\.
- CVH\+\[22\]Marco Cerezo, Guillaume Verdon, Hsin\-Yuan Huang, Lukasz Cincio, and Patrick J Coles\.Challenges and opportunities in quantum machine learning\.Nature computational science, 2\(9\):567–576, 2022\.
- CYF \[22\]Samuel Yen\-Chi Chen, Shinjae Yoo, and Yao\-Lung L Fang\.Quantum long short\-term memory\.InIcassp 2022\-2022 IEEE international conference on acoustics, speech and signal processing \(ICASSP\), pages 8622–8626\. IEEE, 2022\.
- CYQ\+\[20\]Samuel Yen\-Chi Chen, Chao\-Han Huck Yang, Jun Qi, Pin\-Yu Chen, Xiaoli Ma, and Hsi\-Sheng Goan\.Variational quantum circuits for deep reinforcement learning\.IEEE access, 8:141007–141024, 2020\.
- D\+\[25\]Shuhong Dai et al\.Quantum reinforcement learning for qos\-aware real\-time job scheduling in cloud systems\.IEEE Systems Journal, 2025\.
- DCLT \[08\]Daoyi Dong, Chunlin Chen, Hanxiong Li, and Tzyh\-Jong Tarn\.Quantum reinforcement learning\.IEEE Transactions on Systems, Man, and Cybernetics, Part B \(Cybernetics\), 38\(5\):1207–1220, 2008\.
- EFGC\+\[23\]Aleix Espuña Fontcuberta, Anubhab Ghosh, Saikat Chatterjee, Dhrubaditya Mitra, and Dibyendu Nandy\.Forecasting solar cycle 25 with physical model\-validated recurrent neural networks\.Solar Physics, 298\(1\):8, 2023\.
- FCB \[18\]Yao\-Lung L Fang, Francesco Ciccarello, and Harold U Baranger\.Non\-Markovian dynamics of a qubit due to single\-photon scattering in a waveguide\.New Journal of Physics, 20\(4\):043035, 2018\.
- GLM \[04\]Vittorio Giovannetti, Seth Lloyd, and Lorenzo Maccone\.Quantum\-enhanced measurements: beating the standard quantum limit\.Science, 306\(5700\):1330–1336, 2004\.
- \[42\]Yu\-Chao Hsu et al\.QKAN\-LSTM: Quantum\-inspired Kolmogorov\-Arnold long short\-term memory\.arXiv preprint arXiv:2512\.05049, 2025\.
- \[43\]Hsin\-Yuan Huang et al\.Generative quantum advantage for classical and quantum problems\.arXiv preprint arXiv:2509\.09033, 2025\.
- HCL\+\[25\]Yu\-Chao Hsu, Nan\-Yow Chen, Tai\-Yu Li, Po\-Heng Henry Lee, and Kuan\-Cheng Chen\.Quantum kernel\-based long short\-term memory for climate time\-series forecasting\.In2025 International Conference on Quantum Communications, Networking, and Computing \(QCNC\), pages 421–426\. IEEE, 2025\.
- HZLB \[25\]Songtao Huang, Zhen Zhao, Can Li, and Lei Bai\.Timekan: Kan\-based frequency decomposition learning architecture for long\-term time series forecasting\.arXiv preprint arXiv:2502\.06910, 2025\.
- IBM \[26\]IBM Quantum\.IBM Quantum\.[https://quantum\.cloud\.ibm\.com/](https://quantum.cloud.ibm.com/), 2026\.Accessed: 2026\-04\-22\.
- ISCS \[21\]Kazuki Irie, Imanol Schlag, Róbert Csordás, and Jürgen Schmidhuber\.Going beyond linear transformers with recurrent fast weight programmers\.Advances in neural information processing systems, 34:7703–7717, 2021\.
- J\+\[25\]Imen Jarraya et al\.SOH\-KLSTM: A hybrid Kolmogorov\-Arnold network and LSTM model for enhanced lithium\-ion battery health monitoring\.Journal of Energy Storage, 122:116541, 2025\.
- JA\+\[24\]Ali Javadi\-Abhari et al\.Quantum computing with Qiskit, 2024\.
- JHCG \[25\]Jiun\-Cheng Jiang, Yu\-Chao Huang, Tianlong Chen, and Hsi\-Sheng Goan\.Quantum variational activation functions empower Kolmogorov\-Arnold networks\.arXiv preprint arXiv:2509\.14026, 2025\.
- JHS\+\[25\]Mingrui Jing, Erdong Huang, Xiao Shi, Shengyu Zhang, and Xin Wang\.Quantum recurrent embedding neural network\.arXiv preprint arXiv:2506\.13185, 2025\.
- Jia \[25\]Jiun\-Cheng Jiang\.QKAN: Quantum\-inspired Kolmogorov\-Arnold network, 2025\.
- K\+\[23\]Jin\-Sung Kim et al\.Cuda quantum: The platform for integrated quantum\-classical computing\.In2023 60th ACM/IEEE Design Automation Conference \(DAC\), pages 1–4\. IEEE, 2023\.
- K\+\[24\]Akash Kundu et al\.KANQAS: Kolmogorov\-Arnold network for quantum architecture search\.EPJ Quantum Technology, 11\(1\):76, 2024\.
- KOS \[07\]Emanuel Knill, Gerardo Ortiz, and Rolando D Somma\.Optimal quantum measurements of expectation values of observables\.Physical Review A—Atomic, Molecular, and Optical Physics, 75\(1\):012328, 2007\.
- KPE \[21\]Oleksandr Kyriienko, Annie E Paine, and Vincent E Elfving\.Solving nonlinear differential equations with differentiable quantum circuits\.Physical Review A, 103\(5\):052416, 2021\.
- \[57\]Martin Larocca et al\.Barren plateaus in variational quantum computing\.Nature Reviews Physics, 7\(4\):174–189, 2025\.
- \[58\]Shengsheng Lin et al\.Segrnn: Segment recurrent neural network for long\-term time series forecasting\.IEEE Internet of Things Journal, 2025\.
- \[59\]Chen\-Yu Liu et al\.Programming variational quantum circuits with quantum\-train agent\.In2025 International Conference on Quantum Communications, Networking, and Computing \(QCNC\), pages 544–548\. IEEE, 2025\.
- \[60\]Chen\-Yu Liu et al\.Quantum\-enhanced parameter\-efficient learning for typhoon trajectory forecasting\.In2025 IEEE International Conference on Quantum Computing and Engineering \(QCE\), volume 1, pages 2046–2056\. IEEE, 2025\.
- \[61\]Chen\-Yu Liu et al\.Quantum\-train: Rethinking hybrid quantum\-classical machine learning in the model compression perspective\.Quantum Machine Intelligence, 7\(2\):80, 2025\.
- \[62\]Hongfeng Liu et al\.Neural quantum embedding via deterministic quantum computation with one qubit\.Physical Review Letters, 135\(8\):080603, 2025\.
- \[63\]Ziming Liu et al\.KAN: Kolmogorov–Arnold networks\.InThe Thirteenth International Conference on Learning Representations, 2025\.
- L\+\[26\]Jin Lee et al\.KANO: Kolmogorov–Arnold neural operator\.InThe Fourteenth International Conference on Learning Representations, 2026\.
- Liv \[24\]Ioannis E Livieris\.C\-KAN: A new approach for integrating convolutional layers with kolmogorov–Arnold networks for time\-series forecasting\.Mathematics, 12\(19\):3022, 2024\.
- LPC\+\[25\]Chen\-Yu Liu, Leonardo Placidi, Kuan\-Cheng Chen, Samuel Yen\-Chi Chen, and Gabriel Matos\.You only measure once: On designing single\-shot quantum machine learning models\.arXiv preprint arXiv:2509\.20090, 2025\.
- LS \[20\]Owen Lockwood and Mei Si\.Reinforcement learning with quantum variational circuit\.InProceedings of the AAAI conference on artificial intelligence and interactive digital entertainment, volume 16, pages 245–251, 2020\.
- LTM\+\[25\]Ziming Liu, Max Tegmark, Pingchuan Ma, Wojciech Matusik, and Yixuan Wang\.Kolmogorov–Arnold networks meet science\.Physical Review X, 15\(4\):041051, 2025\.
- LW \[18\]Seth Lloyd and Christian Weedbrook\.Quantum generative adversarial learning\.Physical review letters, 121\(4\):040502, 2018\.
- MBS\+\[18\]Jarrod R McClean, Sergio Boixo, Vadim N Smelyanskiy, Ryan Babbush, and Hartmut Neven\.Barren plateaus in quantum neural network training landscapes\.Nature communications, 9\(1\):4812, 2018\.
- MC \[18\]Eric Martin and Chris Cundy\.Parallelizing linear recurrent neural nets over sequence length\.InInternational Conference on Learning Representations, 2018\.
- NVI \[25\]NVIDIA\.NVIDIA CUDA Tile, 2025\.
- NWLDM \[25\]Amir Noorizadegan, Sifan Wang, Leevan Ling, and Juan P Dominguez\-Morales\.A practitioner’s guide to Kolmogorov\-Arnold networks\.arXiv preprint arXiv:2510\.25781, 2025\.
- P\+\[19\]Adam Paszke et al\.Pytorch: An imperative style, high\-performance deep learning library\.Advances in Neural Information Processing Systems, 32, 2019\.
- P\+\[24\]Yash J Patel et al\.Curriculum reinforcement learning for quantum architecture search under hardware errors\.arXiv preprint arXiv:2402\.03500, 2024\.
- Pes \[08\]William Dean Pesnell\.Predictions of solar cycle 24\.Solar Physics, 252\(1\):209–220, 2008\.
- Pet \[20\]Kristóf Petrovay\.Solar cycle prediction\.Living Reviews in Solar Physics, 17\(1\):2, 2020\.
- PKAY \[22\]Taylor L Patti, Jean Kossaifi, Anima Anandkumar, and Susanne F Yelin\.Variational quantum optimization with multibasis encodings\.Physical Review Research, 4\(3\):033142, 2022\.
- Pre \[18\]John Preskill\.Quantum computing in the NISQ era and beyond\.Quantum, 2:79, 2018\.
- PSCLGFL \[20\]Adrián Pérez\-Salinas, Alba Cervera\-Lierta, Elies Gil\-Fuster, and José I Latorre\.Data re\-uploading for a universal quantum classifier\.Quantum, 4:226, 2020\.
- R\+\[24\]David A Rower et al\.Suppressing counter\-rotating errors for fast single\-qubit gates with fluxonium\.PRX Quantum, 5\(4\):040342, 2024\.
- S\+\[21\]Samuel A Stein et al\.Qugan: A quantum state fidelity based generative adversarial network\.In2021 IEEE international conference on quantum computing and engineering \(QCE\), pages 71–81\. IEEE, 2021\.
- S\+\[22\]Samuel A Stein et al\.Quclassi: A hybrid deep neural network architecture based on quantum state fidelity\.Proceedings of Machine Learning and Systems, 4:251–264, 2022\.
- S\+\[25\]Shriyank Somvanshi et al\.A survey on kolmogorov–Arnold network\.ACM Computing Surveys, 58\(2\):1–35, 2025\.
- Sch \[92\]Jürgen Schmidhuber\.Learning to control fast\-weight memories: An alternative to dynamic recurrent networks\.Neural Computation, 4\(1\):131–139, 1992\.
- Sch \[93\]Jürgen Schmidhuber\.Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets\.InInternational Conference on Artificial Neural Networks, pages 460–463\. Springer, 1993\.
- SIS \[21\]Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber\.Linear transformers are secretly fast weight programmers\.InInternational conference on machine learning, pages 9355–9366\. PMLR, 2021\.
- SJD \[22\]Andrea Skolik, Sofiene Jerbi, and Vedran Dunjko\.Quantum agents in the gym: a variational quantum algorithm for deep q\-learning\.Quantum, 6:720, 2022\.
- SLM\+\[25\]Molly C Smith, Aaron D Leu, Koichiro Miyanishi, Mario F Gely, and David M Lucas\.Single\-qubit gates with errors at the 10\-7 level\.Physical Review Letters, 134\(23\):230601, 2025\.
- SS \[17\]Imanol Schlag and Jürgen Schmidhuber\.Gated fast weights for on\-the\-fly neural program generation\.InNIPS Metalearning Workshop, 2017\.
- SSM \[21\]Maria Schuld, Ryan Sweke, and Johannes Jakob Meyer\.Effect of data encoding on the expressive power of variational quantum\-machine\-learning models\.Physical Review A, 103\(3\):032430, 2021\.
- TCK \[13\]Tommaso Tufarelli, Francesco Ciccarello, and MS Kim\.Dynamics of spontaneous emission in a single\-end photonic waveguide\.Physical Review A—Atomic, Molecular, and Optical Physics, 87\(1\):013820, 2013\.
- VRBPC \[24\]Cristian J Vaca\-Rubio, Luis Blanco, Roberto Pereira, and Màrius Caus\.Kolmogorov\-Arnold networks \(kans\) for time series analysis\.In2024 IEEE Globecom Workshops \(GC Wkshps\), pages 1–6\. IEEE, 2024\.
- W\+\[25\]Yi\-Hsien Wu et al\.Simultaneous high\-fidelity single\-qubit gates in a spin qubit array, 2025\.
- WG \[26\]Tzong\-Daw Wu and Hsi\-Sheng Goan\.Quantum recurrent unit: A parameter\-efficient quantum neural network architecture for NISQ devices\.arXiv preprint arXiv:2601\.18164, 2026\.
- WIWL \[22\]David Wierichs, Josh Izaac, Cody Wang, and Cedric Yen\-Yu Lin\.General parameter\-shift rules for quantum gradients\.Quantum, 6:677, 2022\.
- XCW \[24\]Kunpeng Xu, Lifei Chen, and Shengrui Wang\.Kolmogorov\-Arnold networks for time series: Bridging predictive power and interpretability\.arXiv preprint arXiv:2406\.02496, 2024\.
- Y\+\[23\]Songlin Yang et al\.Gated linear attention transformers with hardware\-efficient training\.arXiv preprint arXiv:2312\.06635, 2023\.
- YLZP \[25\]Peter Tettey Yamak, Yujian Li, Ting Zhang, and Muhammad Salman Pathan\.Kolmogorov\-Arnold networks for time series forecasting: a comprehensive review\.Cluster Computing, 28\(14\):929, 2025\.
- YW \[25\]Xingyi Yang and Xinchao Wang\.Kolmogorov\-Arnold transformer\.InThe Thirteenth International Conference on Learning Representations, 2025\.
## Appendix AParallel Evaluation of the Gated Fast\-Weight Recursion
This appendix establishes that the gated fast\-weight recursion admits efficient parallel evaluation in the forward pass\. We first reformulate the update[Equation˜3](https://arxiv.org/html/2605.06734#S4.E3)as an affine recurrence with a scalar multiplier and show that the induced composition operator on parameter pairs is associative\. This structure allows us to express the trajectory\{Wt\}t=1T\\\{W\_\{t\}\\\}\_\{t=1\}^\{T\}as a prefix product, which can be evaluated using a work\-efficient parallel scan\. We then analyze the resulting complexity, showing that the full trajectory can be computed inO\(logT\)O\(\\log T\)depth on an unbounded\-processor PRAM and inO\(T/p\+logp\)O\(T/p\+\\log p\)time onppprocessors\. Finally, we contrast this behavior with theΩ\(T\)\\Omega\(T\)sequential depth required for general nonlinear recurrences, highlighting the structural source of parallelism in the gated update\.
### A\.1Affine Reformulation and Associativity
Throughout this appendix we work with the gated fast\-weight recursion
Wt\+1=gtWt\+\(1−gt\)ΔWt,gt∈\[0,1\],t=1,2,…,T,W\_\{t\+1\}\\;=\\;g\_\{t\}\\,W\_\{t\}\\,\+\\,\(1\-g\_\{t\}\)\\,\\Delta W\_\{t\},\\qquad g\_\{t\}\\in\[0,1\],\\quad t=1,2,\\dots,T,\(16\)with initial fast parametersW1∈ℝm×nW\_\{1\}\\in\\mathbb\{R\}^\{m\\times n\}\. The gategt∈\[0,1\]g\_\{t\}\\in\[0,1\]and proposalΔWt∈ℝm×n\\Delta W\_\{t\}\\in\\mathbb\{R\}^\{m\\times n\}are produced by the slow programmer from the inputxtx\_\{t\}alone:
ΔWt=SΔ\(xt\),gt=σ\(Sg\(xt\)\),\\Delta W\_\{t\}\\;=\\;S\_\{\\Delta\}\(x\_\{t\}\),\\qquad g\_\{t\}\\;=\\;\\sigma\\\!\\bigl\(S\_\{g\}\(x\_\{t\}\)\\bigr\),\(17\)so that\{\(ΔWt,gt\)\}t=1T\\\{\(\\Delta W\_\{t\},g\_\{t\}\)\\\}\_\{t=1\}^\{T\}depend only on the input sequence\{xt\}t=1T\\\{x\_\{t\}\\\}\_\{t=1\}^\{T\}and slow\-programmer weights, never on previous fast parametersW<tW\_\{<t\}\. This independence is what makes the entire set of pairs computable in a single parallel pass over time\.
The remaining task — and the technical content of this appendix — is to show that the parameter trajectory\{Wt\}t=1T\+1\\\{W\_\{t\}\\\}\_\{t=1\}^\{T\+1\}*itself*can be resolved in parallel, despite the apparent sequential coupling in[Equation˜16](https://arxiv.org/html/2605.06734#A1.E16)\.
Define
at:=gt∈\[0,1\],bt:=\(1−gt\)ΔWt∈ℝm×n\.a\_\{t\}\\;:=\\;g\_\{t\}\\in\[0,1\],\\qquad b\_\{t\}\\;:=\\;\(1\-g\_\{t\}\)\\,\\Delta W\_\{t\}\\in\\mathbb\{R\}^\{m\\times n\}\.\(18\)Substituting into[Equation˜16](https://arxiv.org/html/2605.06734#A1.E16)yields
Wt\+1=atWt\+bt,W\_\{t\+1\}\\;=\\;a\_\{t\}\\,W\_\{t\}\\,\+\\,b\_\{t\},\(19\)which is affine inWtW\_\{t\}with a*scalar*multiplierat∈\[0,1\]a\_\{t\}\\in\[0,1\]acting by ordinary scalar multiplication onWt∈ℝm×nW\_\{t\}\\in\\mathbb\{R\}^\{m\\times n\}\.
The scalar nature ofata\_\{t\}is essential\. If, instead, the multiplier were a matrixAt∈ℝm×mA\_\{t\}\\in\\mathbb\{R\}^\{m\\times m\}acting by left multiplication \(as in element\-wise gated RNNs\), composition would still be associative, but each composed element would be a densem×mm\\times mmatrix; the prefix scan would carryO\(m2\)O\(m^\{2\}\)payload per node, and gradient composition would require multiplyingTTdense matrices, reintroducing the conditioning issues we wish to avoid\. The single scalar gategtg\_\{t\}keeps both the scan payload dimension and the gradient chain trivial\.
Consider the set of*affine pairs*𝒜:=ℝ×ℝm×n\\mathcal\{A\}:=\\mathbb\{R\}\\times\\mathbb\{R\}^\{m\\times n\}, where the first component is a scalar multiplier and the second is a matrix offset\. Define the binary operator∘:𝒜×𝒜→𝒜\\circ:\\mathcal\{A\}\\times\\mathcal\{A\}\\to\\mathcal\{A\}by
\(a′,b′\)∘\(a,b\):=\(a′a,a′b\+b′\)\.\(a^\{\\prime\},b^\{\\prime\}\)\\circ\(a,b\)\\;:=\\;\\bigl\(a^\{\\prime\}a,\\;a^\{\\prime\}b\+b^\{\\prime\}\\bigr\)\.\(20\)Geometrically,\(a,b\)\(a,b\)represents the affine mapW↦aW\+bW\\mapsto aW\+b, and[Equation˜20](https://arxiv.org/html/2605.06734#A1.E20)expresses the composition of two such maps in the conventional left\-to\-right order:
\[\(a′,b′\)∘\(a,b\)\]\(W\)=a′\(aW\+b\)\+b′=\(a′a\)W\+\(a′b\+b′\)\.\\bigl\[\(a^\{\\prime\},b^\{\\prime\}\)\\circ\(a,b\)\\bigr\]\(W\)\\;=\\;a^\{\\prime\}\(aW\+b\)\+b^\{\\prime\}\\;=\\;\(a^\{\\prime\}a\)W\+\(a^\{\\prime\}b\+b^\{\\prime\}\)\.\(21\)
The associativity of∘\\circis a specialization of the standard recurrence\-to\-scan reduction for first\-order linear recurrences\[[9](https://arxiv.org/html/2605.06734#bib.bib9), §4\.1\]to the case where the multiplier is a scalar inℝ\\mathbb\{R\}and the offset is a matrix inℝm×n\\mathbb\{R\}^\{m\\times n\}\. We record the proof in our setting both for completeness and to fix notation for the gradient analysis in[Appendix˜B](https://arxiv.org/html/2605.06734#A2)\.
###### Lemma A\.1\(Associativity of∘\\circ\)\.
For any\(a,b\),\(a′,b′\),\(a′′,b′′\)∈𝒜\(a,b\),\\,\(a^\{\\prime\},b^\{\\prime\}\),\\,\(a^\{\\prime\\prime\},b^\{\\prime\\prime\}\)\\in\\mathcal\{A\},
\(\(a′′,b′′\)∘\(a′,b′\)\)∘\(a,b\)=\(a′′,b′′\)∘\(\(a′,b′\)∘\(a,b\)\)\.\\bigl\(\(a^\{\\prime\\prime\},b^\{\\prime\\prime\}\)\\circ\(a^\{\\prime\},b^\{\\prime\}\)\\bigr\)\\circ\(a,b\)\\;=\\;\(a^\{\\prime\\prime\},b^\{\\prime\\prime\}\)\\circ\\bigl\(\(a^\{\\prime\},b^\{\\prime\}\)\\circ\(a,b\)\\bigr\)\.
###### Proof\.
We expand both sides directly from the definition[Equation˜20](https://arxiv.org/html/2605.06734#A1.E20)\.
*Left\-hand side\.*First compute the inner composition:
\(a′′,b′′\)∘\(a′,b′\)=\(a′′a′,a′′b′\+b′′\)\.\(a^\{\\prime\\prime\},b^\{\\prime\\prime\}\)\\circ\(a^\{\\prime\},b^\{\\prime\}\)\\;=\\;\(a^\{\\prime\\prime\}a^\{\\prime\},\\;a^\{\\prime\\prime\}b^\{\\prime\}\+b^\{\\prime\\prime\}\)\.Then compose with\(a,b\)\(a,b\):
\(a′′a′,a′′b′\+b′′\)∘\(a,b\)\\displaystyle\\bigl\(a^\{\\prime\\prime\}a^\{\\prime\},\\;a^\{\\prime\\prime\}b^\{\\prime\}\+b^\{\\prime\\prime\}\\bigr\)\\circ\(a,b\)=\(\(a′′a′\)a,\(a′′a′\)b\+\(a′′b′\+b′′\)\)\\displaystyle=\\bigl\(\(a^\{\\prime\\prime\}a^\{\\prime\}\)\\,a,\\;\\;\(a^\{\\prime\\prime\}a^\{\\prime\}\)\\,b\+\(a^\{\\prime\\prime\}b^\{\\prime\}\+b^\{\\prime\\prime\}\)\\bigr\)=\(a′′a′a,a′′a′b\+a′′b′\+b′′\)\.\\displaystyle=\\bigl\(a^\{\\prime\\prime\}a^\{\\prime\}a,\\;\\;a^\{\\prime\\prime\}a^\{\\prime\}b\+a^\{\\prime\\prime\}b^\{\\prime\}\+b^\{\\prime\\prime\}\\bigr\)\.
*Right\-hand side\.*First compute the inner composition:
\(a′,b′\)∘\(a,b\)=\(a′a,a′b\+b′\)\.\(a^\{\\prime\},b^\{\\prime\}\)\\circ\(a,b\)\\;=\\;\(a^\{\\prime\}a,\\;a^\{\\prime\}b\+b^\{\\prime\}\)\.Then compose with\(a′′,b′′\)\(a^\{\\prime\\prime\},b^\{\\prime\\prime\}\):
\(a′′,b′′\)∘\(a′a,a′b\+b′\)\\displaystyle\(a^\{\\prime\\prime\},b^\{\\prime\\prime\}\)\\circ\\bigl\(a^\{\\prime\}a,\\;a^\{\\prime\}b\+b^\{\\prime\}\\bigr\)=\(a′′\(a′a\),a′′\(a′b\+b′\)\+b′′\)\\displaystyle=\\bigl\(a^\{\\prime\\prime\}\(a^\{\\prime\}a\),\\;\\;a^\{\\prime\\prime\}\(a^\{\\prime\}b\+b^\{\\prime\}\)\+b^\{\\prime\\prime\}\\bigr\)=\(a′′a′a,a′′a′b\+a′′b′\+b′′\)\.\\displaystyle=\\bigl\(a^\{\\prime\\prime\}a^\{\\prime\}a,\\;\\;a^\{\\prime\\prime\}a^\{\\prime\}b\+a^\{\\prime\\prime\}b^\{\\prime\}\+b^\{\\prime\\prime\}\\bigr\)\.
The two expressions are identical, completing the proof\. ∎
Three remarks are in order\.
##### Remark 1 \(no constancy or smoothness assumption on the gates\)\.
The proof of[Lemma˜A\.1](https://arxiv.org/html/2605.06734#A1.Thmlemma1)treatsa,a′,a′′a,a^\{\\prime\},a^\{\\prime\\prime\}as arbitrary scalars andb,b′,b′′b,b^\{\\prime\},b^\{\\prime\\prime\}as arbitrary matrices\. Associativity therefore holds pointwise for every triple of pairs, regardless of whether the scalars are constants, time\-varying, or input\-dependent\. In particular, no assumption such asgt≡gg\_\{t\}\\equiv gor smoothness ofgtg\_\{t\}inttis required\. The gates produced by the slow programmer in[Equation˜17](https://arxiv.org/html/2605.06734#A1.E17)satisfy this trivially\.
##### Remark 2 \(identity element\)\.
The pair\(1,0m×n\)\(1,\\,0\_\{m\\times n\}\)is a two\-sided identity for∘\\circ:
\(1,0\)∘\(a,b\)=\(a,b\)=\(a,b\)∘\(1,0\),\(1,0\)\\circ\(a,b\)\\;=\\;\(a,b\)\\;=\\;\(a,b\)\\circ\(1,0\),which lets us extend prefix products to indices≤0\\leq 0when convenient by padding with\(1,0\)\(1,0\)\.
##### Remark 3 \(non\-commutativity\)\.
The operator∘\\circis*not*commutative: in general,
\(a′,b′\)∘\(a,b\)=\(a′a,a′b\+b′\)≠\(aa′,ab′\+b\)=\(a,b\)∘\(a′,b′\),\(a^\{\\prime\},b^\{\\prime\}\)\\circ\(a,b\)\\;=\\;\(a^\{\\prime\}a,\\,a^\{\\prime\}b\+b^\{\\prime\}\)\\;\\neq\\;\(aa^\{\\prime\},\\,ab^\{\\prime\}\+b\)\\;=\\;\(a,b\)\\circ\(a^\{\\prime\},b^\{\\prime\}\),since the offset components differ\. This is consistent with the order\-sensitive nature of sequence modeling: the proposal at timekkshould influenceWt\+1W\_\{t\+1\}via gatesgk\+1,…,gtg\_\{k\+1\},\\dots,g\_\{t\}, not via gatesg1,…,gk−1g\_\{1\},\\dots,g\_\{k\-1\}\. Associativity, but not commutativity, is what enables parallel reassociation while preserving sequence order\.
### A\.2Trajectory as a Prefix Product
We now show that one step of the recursion[Equation˜19](https://arxiv.org/html/2605.06734#A1.E19)is exactly one application of∘\\circ, and consequently that the entire trajectory is a prefix product\.
###### Proposition 1\(Prefix\-product form\)\.
LetW1∈ℝm×nW\_\{1\}\\in\\mathbb\{R\}^\{m\\times n\},at∈ℝa\_\{t\}\\in\\mathbb\{R\}, andbt∈ℝm×nb\_\{t\}\\in\\mathbb\{R\}^\{m\\times n\}fort=1,…,Tt=1,\\dots,T, and defineWt\+1=atWt\+btW\_\{t\+1\}=a\_\{t\}W\_\{t\}\+b\_\{t\}as in[Equation˜19](https://arxiv.org/html/2605.06734#A1.E19)\. Define the running prefix product
Pt:=\(at,bt\)∘\(at−1,bt−1\)∘⋯∘\(a1,b1\)∘\(1,W1\),t=1,…,T\.P\_\{t\}\\;:=\\;\(a\_\{t\},b\_\{t\}\)\\circ\(a\_\{t\-1\},b\_\{t\-1\}\)\\circ\\cdots\\circ\(a\_\{1\},b\_\{1\}\)\\circ\(1,W\_\{1\}\),\\qquad t=1,\\dots,T\.\(22\)Then for everytt,
Pt=\(αt,Wt\+1\),αt=∏s=1tas,P\_\{t\}\\;=\\;\\Bigl\(\\alpha\_\{t\},\\;W\_\{t\+1\}\\Bigr\),\\qquad\\alpha\_\{t\}\\;=\\;\\prod\_\{s=1\}^\{t\}a\_\{s\},\(23\)i\.e\., the second component of thett\-th prefix product is exactly the fast\-parameter state at timet\+1t\+1\.
###### Proof\.
We proceed by induction ontt\.
*Base case \(t=1t=1\)\.*By direct computation,
P1=\(a1,b1\)∘\(1,W1\)=\(a1⋅1,a1W1\+b1\)=\(a1,W2\)\.P\_\{1\}\\;=\\;\(a\_\{1\},b\_\{1\}\)\\circ\(1,W\_\{1\}\)\\;=\\;\\bigl\(a\_\{1\}\\cdot 1,\\;\\;a\_\{1\}W\_\{1\}\+b\_\{1\}\\bigr\)\\;=\\;\(a\_\{1\},\\;W\_\{2\}\)\.HenceP1=\(α1,W2\)P\_\{1\}=\(\\alpha\_\{1\},W\_\{2\}\)withα1=a1\\alpha\_\{1\}=a\_\{1\}, as claimed\.
*Inductive step\.*AssumePt=\(αt,Wt\+1\)P\_\{t\}=\(\\alpha\_\{t\},W\_\{t\+1\}\)withαt=∏s=1tas\\alpha\_\{t\}=\\prod\_\{s=1\}^\{t\}a\_\{s\}\. Then
Pt\+1\\displaystyle P\_\{t\+1\}=\(at\+1,bt\+1\)∘Pt\\displaystyle\\;=\\;\(a\_\{t\+1\},b\_\{t\+1\}\)\\circ P\_\{t\}=\(at\+1,bt\+1\)∘\(αt,Wt\+1\)\\displaystyle\\;=\\;\(a\_\{t\+1\},b\_\{t\+1\}\)\\circ\(\\alpha\_\{t\},W\_\{t\+1\}\)=\(at\+1αt,at\+1Wt\+1\+bt\+1\)\\displaystyle\\;=\\;\\bigl\(a\_\{t\+1\}\\alpha\_\{t\},\\;\\;a\_\{t\+1\}W\_\{t\+1\}\+b\_\{t\+1\}\\bigr\)=\(αt\+1,Wt\+2\),\\displaystyle\\;=\\;\\bigl\(\\alpha\_\{t\+1\},\\;\\;W\_\{t\+2\}\\bigr\),where the last equality usesαt\+1=at\+1αt\\alpha\_\{t\+1\}=a\_\{t\+1\}\\alpha\_\{t\}and the recursionWt\+2=at\+1Wt\+1\+bt\+1W\_\{t\+2\}=a\_\{t\+1\}W\_\{t\+1\}\+b\_\{t\+1\}\. HencePt\+1P\_\{t\+1\}also has the claimed form, completing the induction\. ∎
##### Worked example forT=3T=3\.
To make[Proposition˜1](https://arxiv.org/html/2605.06734#Thmproposition1)concrete, we unroll the prefix product forT=3T=3\. Composing left\-to\-right starting from\(1,W1\)\(1,W\_\{1\}\):
P1=\(a1,b1\)∘\(1,W1\)\\displaystyle P\_\{1\}\\;=\\;\(a\_\{1\},b\_\{1\}\)\\circ\(1,W\_\{1\}\)=\(a1,a1W1\+b1\)=\(a1,W2\),\\displaystyle\\;=\\;\\bigl\(a\_\{1\},\\;\\;a\_\{1\}W\_\{1\}\+b\_\{1\}\\bigr\)\\;=\\;\(a\_\{1\},\\,W\_\{2\}\),P2=\(a2,b2\)∘P1\\displaystyle P\_\{2\}\\;=\\;\(a\_\{2\},b\_\{2\}\)\\circ P\_\{1\}=\(a2,b2\)∘\(a1,W2\)\\displaystyle\\;=\\;\(a\_\{2\},b\_\{2\}\)\\circ\(a\_\{1\},W\_\{2\}\)=\(a2a1,a2W2\+b2\)=\(a2a1,W3\),\\displaystyle\\;=\\;\\bigl\(a\_\{2\}a\_\{1\},\\;\\;a\_\{2\}W\_\{2\}\+b\_\{2\}\\bigr\)\\;=\\;\(a\_\{2\}a\_\{1\},\\,W\_\{3\}\),P3=\(a3,b3\)∘P2\\displaystyle P\_\{3\}\\;=\\;\(a\_\{3\},b\_\{3\}\)\\circ P\_\{2\}=\(a3,b3\)∘\(a2a1,W3\)\\displaystyle\\;=\\;\(a\_\{3\},b\_\{3\}\)\\circ\(a\_\{2\}a\_\{1\},W\_\{3\}\)=\(a3a2a1,a3W3\+b3\)=\(a3a2a1,W4\)\.\\displaystyle\\;=\\;\\bigl\(a\_\{3\}a\_\{2\}a\_\{1\},\\;\\;a\_\{3\}W\_\{3\}\+b\_\{3\}\\bigr\)\\;=\\;\(a\_\{3\}a\_\{2\}a\_\{1\},\\,W\_\{4\}\)\.Substitutingat=gta\_\{t\}=g\_\{t\}andbt=\(1−gt\)ΔWtb\_\{t\}=\(1\-g\_\{t\}\)\\Delta W\_\{t\}recovers, in the second component ofP3P\_\{3\},
W4=g3g2g1W1\+g3g2\(1−g1\)ΔW1\+g3\(1−g2\)ΔW2\+\(1−g3\)ΔW3,W\_\{4\}\\;=\\;g\_\{3\}g\_\{2\}g\_\{1\}\\,W\_\{1\}\+g\_\{3\}g\_\{2\}\(1\-g\_\{1\}\)\\Delta W\_\{1\}\+g\_\{3\}\(1\-g\_\{2\}\)\\Delta W\_\{2\}\+\(1\-g\_\{3\}\)\\Delta W\_\{3\},which matches term\-by\-term the unrolled form[Equation˜4](https://arxiv.org/html/2605.06734#S5.E4)of the main text\. This verifies the equivalence between the recursive and prefix\-product views\.
##### Tree\-shaped reassociation\.
Crucially, by associativity \([Lemma˜A\.1](https://arxiv.org/html/2605.06734#A1.Thmlemma1)\), the same prefix productP3P\_\{3\}can be evaluated by any parenthesization of the operands, including the balanced tree
P3=\(\(a3,b3\)∘\(a2,b2\)\)∘\(\(a1,b1\)∘\(1,W1\)\),P\_\{3\}\\;=\\;\\bigl\(\(a\_\{3\},b\_\{3\}\)\\circ\(a\_\{2\},b\_\{2\}\)\\bigr\)\\circ\\bigl\(\(a\_\{1\},b\_\{1\}\)\\circ\(1,W\_\{1\}\)\\bigr\),in which the two inner compositions are independent and can be performed concurrently\. This is the foundation of the parallel scan in[Section˜A\.3](https://arxiv.org/html/2605.06734#A1.SS3)\.
### A\.3Parallel Evaluation via Associative Scan
Any associative binary operator admits a work\-efficient parallel prefix scan\[[9](https://arxiv.org/html/2605.06734#bib.bib9)\]\. We briefly recall the standard up\-sweep / down\-sweep algorithm for completeness, then quantify its depth on the operator∘\\circ\.
##### Up\-sweep\.
Place the inputs\{\(ak,bk\)\}k=1T\\\{\(a\_\{k\},b\_\{k\}\)\\\}\_\{k=1\}^\{T\}at the leaves of a balanced binary tree of heighth:=⌈log2T⌉h:=\\lceil\\log\_\{2\}T\\rceil\. At each internal node, combine the two children under∘\\circ\. Because all combinations at a given level are independent, each level is performed in parallel on an unbounded\-processor PRAM inO\(1\)O\(1\)time, givingO\(logT\)O\(\\log T\)time in total\. After the up\-sweep, each internal nodevvholds the reduction of all leaves in its subtree, and the root holds the full reductionPTP\_\{T\}\.
##### Down\-sweep\.
A second pass propagates partial prefixes back down the tree\. Initialize the root with the identity\(1,0\)\(1,0\)\. At each internal node, the left child receives the value passed down from the parent, and the right child receives that value composed \(via∘\\circ\) with the up\-sweep value of the left child\. After the down\-sweep reaches the leaves, leafkkholds the exclusive prefix\(ak−1,bk−1\)∘⋯∘\(a1,b1\)\(a\_\{k\-1\},b\_\{k\-1\}\)\\circ\\cdots\\circ\(a\_\{1\},b\_\{1\}\); composing with the leaf’s own value yields the inclusive prefixPk=\(ak,bk\)∘⋯∘\(a1,b1\)∘\(1,W1\)P\_\{k\}=\(a\_\{k\},b\_\{k\}\)\\circ\\cdots\\circ\(a\_\{1\},b\_\{1\}\)\\circ\(1,W\_\{1\}\), whose second component isWk\+1W\_\{k\+1\}by[Proposition˜1](https://arxiv.org/html/2605.06734#Thmproposition1)\. The down\-sweep also takesO\(logT\)O\(\\log T\)time on unbounded processors\.
##### Depth and work\.
The total depth isO\(logT\)O\(\\log T\)and the total work \(number of∘\\circoperations\) isO\(T\)O\(T\), matching the sequential cost up to a constant factor\. Each∘\\circoperation, by[Equation˜20](https://arxiv.org/html/2605.06734#A1.E20), requires one scalar multiplication, one scalar\-times\-matrix scaling, and one matrix addition, i\.e\.,O\(mn\)O\(mn\)scalar operations forW∈ℝm×nW\\in\\mathbb\{R\}^\{m\\times n\}\. The parallel scan therefore performsO\(Tmn\)O\(Tmn\)scalar operations in total, identical \(up to constants\) to theTTscalar\-matrix\-add steps of the sequential recursion, but withO\(logT\)O\(\\log T\)rather thanO\(T\)O\(T\)depth\.
On a realistic machine withp≪Tp\\ll Tprocessors, the standard three\-phase implementation of an associative scan\[[9](https://arxiv.org/html/2605.06734#bib.bib9)\]achieves
Tscan\(T,p\)=O\(Tp\+logp\),T\_\{\\text\{scan\}\}\(T,p\)\\;=\\;O\\\!\\left\(\\frac\{T\}\{p\}\+\\log p\\right\),\(24\)with total workO\(T\)O\(T\)\. The three phases are:
1. 1\.Local reduction\.Partition theTTinputs intoppcontiguous blocks of size⌈T/p⌉\\lceil T/p\\rceil\. Each processor sequentially reduces its block under∘\\circinO\(T/p\)O\(T/p\)time, producingppblock reductions\.
2. 2\.Tree\-based scan over block reductions\.Apply the up\-sweep / down\-sweep scan of[Section˜A\.3](https://arxiv.org/html/2605.06734#A1.SS3)to theppblock reductions, yielding the prefix of block reductions inO\(logp\)O\(\\log p\)depth\.
3. 3\.Local propagation\.Each processor takes its received prefix and sequentially propagates it across its block, completing the per\-leaf prefixes inO\(T/p\)O\(T/p\)time\.
Summing the three phases gives[Equation˜24](https://arxiv.org/html/2605.06734#A1.E24)\. Whenp=Θ\(T\)p=\\Theta\(T\), theT/pT/pterm isO\(1\)O\(1\)and[Equation˜24](https://arxiv.org/html/2605.06734#A1.E24)reduces to the unbounded\-processor depthO\(logT\)O\(\\log T\)\. Whenp≪Tp\\ll T, theT/pT/pterm dominates and the parallel speedup over sequential evaluation is linear inpp\.
Martin and Cundy\[[71](https://arxiv.org/html/2605.06734#bib.bib71)\]apply this primitive to feature\-space linear recurrences of the formht=λt⊙ht−1\+xth\_\{t\}=\\lambda\_\{t\}\\odot h\_\{t\-1\}\+x\_\{t\}inside neural\-network cells \(SRU, QRNN, GILR\-LSTM\), demonstrating practical speedups of up to9×9\\timesover serial linear\-RNN evaluation on modern GPUs\. Our analysis above establishes that the same primitive applies to the parameter\-space trajectory\{Wt\}\\\{W\_\{t\}\\\}of a gated fast\-weight programmer, complementing prior work on parallel scans for hidden\-state recurrences\[[71](https://arxiv.org/html/2605.06734#bib.bib71)\]\. Crucially, in our setting the scan is performed once over the parameter sequence and produces the entire trajectory\{Wt\}t=1T\\\{W\_\{t\}\\\}\_\{t=1\}^\{T\}for a sequence of lengthTT, in contrast to nonlinear recurrent architectures, which must serialize overttin both forward and backward passes\.
### A\.4Sequential Depth Lower Bound for Nonlinear Recurrences
For a recurrence of the general form
ht\+1=f\(ht,xt\),h\_\{t\+1\}\\;=\\;f\(h\_\{t\},x\_\{t\}\),\(25\)withffnonlinear inhth\_\{t\}, no analogous reassociation is available\. Specifically, the two\-step composition
ht\+2=f\(f\(ht,xt\),xt\+1\)h\_\{t\+2\}\\;=\\;f\\bigl\(f\(h\_\{t\},x\_\{t\}\),\\;x\_\{t\+1\}\\bigr\)is not, in general, expressible as a single application of an operator that depends only on\{xt,xt\+1\}\\\{x\_\{t\},x\_\{t\+1\}\\\}and a fixed\-size summary of the past\. As a consequence, computinghTh\_\{T\}fromh0h\_\{0\}requiresΩ\(T\)\\Omega\(T\)sequential applications offf, regardless of available parallelism, and no associative\-scan acceleration is possible\[[71](https://arxiv.org/html/2605.06734#bib.bib71)\]\.
This contrast is precisely what differentiates the gated fast\-weight framework from recurrent integrations of HQKAN such as QKAN\-LSTM\[[42](https://arxiv.org/html/2605.06734#bib.bib42)\], in which an HQKAN block sits inside an LSTM cell and depends on the recurrent hidden stateht−1h\_\{t\-1\}\. Such architectures inherit theΩ\(T\)\\Omega\(T\)sequential depth of LSTMs and require BPTT to traverse a chain ofTTdense hidden\-state Jacobians\. The gated fast\-weight framework decouples parameter\-space evolution from any hidden\-state loop: eachΔWk\\Delta W\_\{k\}is computed fromxkx\_\{k\}alone, and the trajectory\{Wt\}\\\{W\_\{t\}\\\}is resolved by an associative scan inO\(logT\)O\(\\log T\)depth\.
##### Summary\.
The gated fast\-weight recursion admits a parallel evaluation strategy because its affine structure induces an associative composition law\. This allows the parameter trajectory to be computed via a prefix scan in logarithmic depth, in contrast to the inherently sequential evaluation of general nonlinear recurrences\. This parallel structure is a direct consequence of the scalar gating mechanism and will also play a central role in simplifying gradient composition, as we show next\.
## Appendix BGradient Composition in the Gated Fast\-Weight Recursion
We now turn to the backward pass and show that the same affine structure underlying the parallel scan also governs gradient composition\. In particular, we derive the sensitivity ofWt\+1W\_\{t\+1\}to an earlier proposalΔWk\\Delta W\_\{k\}and show that the resulting Jacobian reduces to a single scalar multiplier\. This leads to a fundamental simplification: instead of propagating gradients through a chain of dense Jacobians as in general recurrent architectures, temporal dependencies are mediated entirely by products of scalar gates\. As a result, gradient propagation along the time axis is both shallower and better conditioned\.
### B\.1Unrolled form revisited
Recursively expanding[Equation˜16](https://arxiv.org/html/2605.06734#A1.E16)gives the unrolled form \(also stated as[Equation˜4](https://arxiv.org/html/2605.06734#S5.E4)in the main text\):
Wt\+1=\(∏s=1tgs\)W1\+∑k=1t\(1−gk\)\(∏s=k\+1tgs\)ΔWk,W\_\{t\+1\}\\;=\\;\\left\(\\prod\_\{s=1\}^\{t\}g\_\{s\}\\right\)W\_\{1\}\\;\+\\;\\sum\_\{k=1\}^\{t\}\(1\-g\_\{k\}\)\\\!\\left\(\\prod\_\{s=k\+1\}^\{t\}g\_\{s\}\\right\)\\\!\\Delta W\_\{k\},\(26\)with the convention∏s=k\+1tgs=1\\prod\_\{s=k\+1\}^\{t\}g\_\{s\}=1whenk=tk=t\. Define the scalar memory coefficients
β0,t:=∏s=1tgs,βk,t:=\(1−gk\)∏s=k\+1tgs,k=1,…,t\.\\beta\_\{0,t\}\\;:=\\;\\prod\_\{s=1\}^\{t\}g\_\{s\},\\qquad\\beta\_\{k,t\}\\;:=\\;\(1\-g\_\{k\}\)\\\!\\prod\_\{s=k\+1\}^\{t\}g\_\{s\},\\quad k=1,\\dots,t\.\(27\)By induction one verifiesβ0,t\+∑k=1tβk,t=1\\beta\_\{0,t\}\+\\sum\_\{k=1\}^\{t\}\\beta\_\{k,t\}=1andβk,t∈\[0,1\]\\beta\_\{k,t\}\\in\[0,1\]for allkk, soWt\+1W\_\{t\+1\}lies in the convex hull of\{W1,ΔW1,…,ΔWt\}\\\{W\_\{1\},\\Delta W\_\{1\},\\dots,\\Delta W\_\{t\}\\\}, recovering the geometric boundedness property of[Section˜5](https://arxiv.org/html/2605.06734#S5)\.
### B\.2Sensitivity to a single proposal
We compute∂Wt\+1/∂ΔWk\\partial W\_\{t\+1\}/\\partial\\Delta W\_\{k\}entrywise\. Fixk∈\{1,…,t\}k\\in\\\{1,\\dots,t\\\}, and letWt\+1\(ij\)W\_\{t\+1\}^\{\(ij\)\}andΔWk\(i′j′\)\\Delta W\_\{k\}^\{\(i^\{\\prime\}j^\{\\prime\}\)\}denote arbitrary entries ofWt\+1W\_\{t\+1\}andΔWk\\Delta W\_\{k\}, respectively\. From[Equation˜26](https://arxiv.org/html/2605.06734#A2.E26), the entryWt\+1\(ij\)W\_\{t\+1\}^\{\(ij\)\}depends onΔWk\\Delta W\_\{k\}only through the termβk,tΔWk\\beta\_\{k,t\}\\Delta W\_\{k\}, and within that term it depends onΔWk\(i′j′\)\\Delta W\_\{k\}^\{\(i^\{\\prime\}j^\{\\prime\}\)\}only when\(i′,j′\)=\(i,j\)\(i^\{\\prime\},j^\{\\prime\}\)=\(i,j\), with derivativeβk,t\\beta\_\{k,t\}\. Hence
∂Wt\+1\(ij\)∂ΔWk\(i′j′\)=βk,tδii′δjj′,\\frac\{\\partial W\_\{t\+1\}^\{\(ij\)\}\}\{\\partial\\Delta W\_\{k\}^\{\(i^\{\\prime\}j^\{\\prime\}\)\}\}\\;=\\;\\beta\_\{k,t\}\\,\\delta\_\{ii^\{\\prime\}\}\\delta\_\{jj^\{\\prime\}\},\(28\)whereδ\\deltais the Kronecker delta\. Vectorizing both sides, the Jacobian is
∂vec\(Wt\+1\)∂vec\(ΔWk\)=βk,tImn,\\frac\{\\partial\\mathrm\{vec\}\(W\_\{t\+1\}\)\}\{\\partial\\mathrm\{vec\}\(\\Delta W\_\{k\}\)\}\\;=\\;\\beta\_\{k,t\}\\,I\_\{mn\},\(29\)a scalar multiple of themn×mnmn\\times mnidentity\. The dense Jacobian collapses to a*single scalar*multiplierβk,t\\beta\_\{k,t\}propagated independently along each parameter coordinate\.
##### Worked example:t=3t=3,k=1k=1\.
ForT=3T=3andk=1k=1,[Equation˜26](https://arxiv.org/html/2605.06734#A2.E26)reads
W4=β0,3W1\+β1,3ΔW1\+β2,3ΔW2\+β3,3ΔW3,W\_\{4\}\\;=\\;\\beta\_\{0,3\}W\_\{1\}\+\\beta\_\{1,3\}\\Delta W\_\{1\}\+\\beta\_\{2,3\}\\Delta W\_\{2\}\+\\beta\_\{3,3\}\\Delta W\_\{3\},with
β0,3=g1g2g3,β1,3=\(1−g1\)g2g3,β2,3=\(1−g2\)g3,β3,3=\(1−g3\)\.\\beta\_\{0,3\}=g\_\{1\}g\_\{2\}g\_\{3\},\\qquad\\beta\_\{1,3\}=\(1\-g\_\{1\}\)g\_\{2\}g\_\{3\},\\qquad\\beta\_\{2,3\}=\(1\-g\_\{2\}\)g\_\{3\},\\qquad\\beta\_\{3,3\}=\(1\-g\_\{3\}\)\.Therefore
∂vec\(W4\)∂vec\(ΔW1\)=\(1−g1\)g2g3Imn,\\frac\{\\partial\\mathrm\{vec\}\(W\_\{4\}\)\}\{\\partial\\mathrm\{vec\}\(\\Delta W\_\{1\}\)\}\\;=\\;\(1\-g\_\{1\}\)g\_\{2\}g\_\{3\}\\,I\_\{mn\},which is bounded by11in absolute value becausegs∈\[0,1\]g\_\{s\}\\in\[0,1\]for allssand1−g1∈\[0,1\]1\-g\_\{1\}\\in\[0,1\]\. One verifies directly thatβ0,3\+β1,3\+β2,3\+β3,3=1\\beta\_\{0,3\}\+\\beta\_\{1,3\}\+\\beta\_\{2,3\}\+\\beta\_\{3,3\}=1, recovering the convex\-combination structure\.
### B\.3Backpropagation through the fast parameters
Letℒ\\mathcal\{L\}denote a downstream loss whose dependence onΔWk\\Delta W\_\{k\}flows throughWt\+1W\_\{t\+1\}for varioust≥kt\\geq k\(typically through outputsyt=F\(xt;Wt\)y\_\{t\}=F\(x\_\{t\};W\_\{t\}\)fort\>kt\>k\)\. The chain rule gives
∂ℒ∂vec\(ΔWk\)=∑t≥k∂ℒ∂vec\(Wt\+1\)∂vec\(Wt\+1\)∂vec\(ΔWk\)=∑t≥kβk,t∂ℒ∂vec\(Wt\+1\)\.\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\mathrm\{vec\}\(\\Delta W\_\{k\}\)\}\\;=\\;\\sum\_\{t\\geq k\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\mathrm\{vec\}\(W\_\{t\+1\}\)\}\\,\\frac\{\\partial\\mathrm\{vec\}\(W\_\{t\+1\}\)\}\{\\partial\\mathrm\{vec\}\(\\Delta W\_\{k\}\)\}\\;=\\;\\sum\_\{t\\geq k\}\\beta\_\{k,t\}\\,\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\mathrm\{vec\}\(W\_\{t\+1\}\)\}\.\(30\)The gradient flowing from timet\+1t\+1back to stepkkis therefore the upstream gradient∂ℒ/∂vec\(Wt\+1\)\\partial\\mathcal\{L\}/\\partial\\mathrm\{vec\}\(W\_\{t\+1\}\)*scalar\-rescaled*byβk,t\\beta\_\{k,t\}, with no matrix multiplication along the temporal direction\.
Since the slow programmer producesΔWk\\Delta W\_\{k\}fromxkx\_\{k\}alone via an independent forward pass \(cf\.[Equation˜17](https://arxiv.org/html/2605.06734#A1.E17)\), the gradient with respect to slow\-programmer weightsθS\\theta\_\{S\}further factors as
∂ℒ∂θS=∑k=1T∂ℒ∂vec\(ΔWk\)∂vec\(ΔWk\)∂θS,\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\theta\_\{S\}\}\\;=\\;\\sum\_\{k=1\}^\{T\}\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\mathrm\{vec\}\(\\Delta W\_\{k\}\)\}\\,\\frac\{\\partial\\mathrm\{vec\}\(\\Delta W\_\{k\}\)\}\{\\partial\\theta\_\{S\}\},where each term∂vec\(ΔWk\)/∂θS\\partial\\mathrm\{vec\}\(\\Delta W\_\{k\}\)/\\partial\\theta\_\{S\}traverses only the depth of one slow\-programmer forward pass, not a chain ofTTrecurrent steps\. The gradient depth across time is therefore controlled entirely by the scalar productβk,t\\beta\_\{k,t\}, while the per\-step gradient depth is the \(constant\) depth of one slow\-programmer evaluation\.
### B\.4Bounded, non\-explosive gradient magnitudes
The scalar coefficientsβk,t\\beta\_\{k,t\}are products of factors in\[0,1\]\[0,1\]:
0≤βk,t=\(1−gk\)∏s=k\+1tgs≤1\.0\\;\\leq\\;\\beta\_\{k,t\}\\;=\\;\(1\-g\_\{k\}\)\\\!\\prod\_\{s=k\+1\}^\{t\}g\_\{s\}\\;\\leq\\;1\.\(31\)Two consequences follow\.
##### No explosion\.
The norm of the temporal Jacobian is bounded above by11for every\(k,t\)\(k,t\):
‖∂vec\(Wt\+1\)/∂vec\(ΔWk\)‖2=βk,t≤1\.\\Bigl\\\|\\,\\partial\\mathrm\{vec\}\(W\_\{t\+1\}\)/\\partial\\mathrm\{vec\}\(\\Delta W\_\{k\}\)\\,\\Bigr\\\|\_\{2\}\\;=\\;\\beta\_\{k,t\}\\;\\leq\\;1\.Repeated composition of these factors across time can only contract gradients; it cannot amplify them\. This rules out exploding gradients along the time axis by construction, without recourse to gradient clipping or spectral regularization\.
##### Possible vanishing\.
Ifgsg\_\{s\}is consistently small fors∈\{k\+1,…,t\}s\\in\\\{k\+1,\\dots,t\\\}, the product∏s=k\+1tgs\\prod\_\{s=k\+1\}^\{t\}g\_\{s\}shrinks geometrically and gradients to earlyΔWk\\Delta W\_\{k\}may vanish\. This is the standard adaptive\-memory trade\-off: gates near11retain long\-range gradient flow, while gates near0implement aggressive forgetting\. Crucially, the gatesgsg\_\{s\}are produced by the slow programmer fromxsx\_\{s\}alone, so their values are learned per input rather than fixed, and the model can in principle learn to retain long\-range dependencies where the data warrant it\.
### B\.5Comparison with dense recurrent Jacobians
For a general nonlinear recurrence[Equation˜25](https://arxiv.org/html/2605.06734#A1.E25), the analogous sensitivity is a product of dense cell\-Jacobians,
∂ht\+1∂hk=∏s=k\+1tJs,Js:=∂f\(hs,xs\)∂hs∈ℝd×d,\\frac\{\\partial h\_\{t\+1\}\}\{\\partial h\_\{k\}\}\\;=\\;\\prod\_\{s=k\+1\}^\{t\}J\_\{s\},\\qquad J\_\{s\}\\;:=\\;\\frac\{\\partial f\(h\_\{s\},x\_\{s\}\)\}\{\\partial h\_\{s\}\}\\in\\mathbb\{R\}^\{d\\times d\},\(32\)whereddis the hidden\-state dimension\. The product of such matrices is well known to be the source of both exploding and vanishing gradients\[[71](https://arxiv.org/html/2605.06734#bib.bib71)\]: its spectral norm is bounded only by∏s‖Js‖2\\prod\_\{s\}\\\|J\_\{s\}\\\|\_\{2\}, which grows or shrinks geometrically witht−kt\-kand depends on every entry of every intermediate Jacobian\. Backpropagation through time therefore requires \(i\) storing allTTactivations to recompute or query theJsJ\_\{s\}, and \(ii\) sequentially multiplyingTTdensed×dd\\times dmatrices on the backward pass, costingO\(Td3\)O\(Td^\{3\}\)work andΩ\(T\)\\Omega\(T\)depth\.
The gated fast\-weight framework replaces the dense product∏sJs\\prod\_\{s\}J\_\{s\}with the scalar productβk,t\\beta\_\{k,t\}\. The asymptotic comparison is summarized in[Table˜10](https://arxiv.org/html/2605.06734#A2.T10)\.
Table 10:Comparison of temporal gradient composition between the gated fast\-weight framework and a general nonlinear recurrence on hidden stateht∈ℝdh\_\{t\}\\in\\mathbb\{R\}^\{d\}\(vs\. fast parametersWt∈ℝm×nW\_\{t\}\\in\\mathbb\{R\}^\{m\\times n\}\)\.##### Summary\.
Gradient propagation through the gated fast\-weight recursion reduces to the composition of scalar coefficients in\[0,1\]\[0,1\], rather than products of dense Jacobians\. This yields two key advantages: \(i\) bounded, non\-explosive gradient magnitudes by construction, and \(ii\) reduced effective depth of temporal gradient paths, which can be evaluated via parallel reductions\. Compared to general nonlinear recurrences, this structure leads to both improved conditioning and lower computational complexity in the backward pass\.
## Appendix CConvergence Analysis on Time\-Series Benchmarks
This section provides a detailed, close\-up view of the learning behavior of GQKAN\-QKANFWP and the standard QFWP in time series benchmark task introduced in[Section˜6\.1](https://arxiv.org/html/2605.06734#S6.SS1)\. While the main text reports aggregate metrics and qualitative comparisons up to 50 training epochs, here we extend the analysis to 100 epochs and visualize the full convergence trajectories at the most demanding window\-size settingN=64N=64\. This extended view allows us to distinguish between optimization speed and representational limits, and to assess whether performance gaps persist or close with additional training\. For each task we display the model predictions at four representative training epochs, with solid lines denoting the seed\-averaged mean across five independent random seed initializations and shaded bands indicating the corresponding±1σ\\pm 1\\sigmaenvelope\. This visualization complements the aggregate metrics reported in Tables[4](https://arxiv.org/html/2605.06734#S6.T4)–[6](https://arxiv.org/html/2605.06734#S6.T6)by exposing temporal qualitative behavior—amplitude tracking, phase alignment, convergence speed, and seed\-to\-seed stability—that scalar MSE values alone cannot fully capture\. Across all six tasks, two recurring trends emerge\. First, GQKAN\-QKANFWP attains a near\-perfect overlap with the ground truth substantially earlier in training than QFWP, and in the smooth\-dynamics and quantum\-dynamics tasks this overlap is already reached by epoch 15\. Second, the GQKAN\-QKANFWP±1σ\\pm 1\\sigmaenvelope remains tight from early epochs onward and barely broadens in the test region, whereas the QFWP envelope is consistently wider and tends to expand past the train/test boundary, indicating both higher seed\-to\-seed variability and weaker out\-of\-sample generalization at long input windows\. Crucially, the extended training to epoch 100 shows that these gaps do not close with additional training\.
##### Damped SHM\.
[Figure˜8](https://arxiv.org/html/2605.06734#A3.F8)illustrates the forecasting trajectories on the Damped SHM dataset\. The target is a smooth, weakly\-damped oscillation with a slowly\-decaying amplitude envelope and a mildly amplitude\-dependent period, jointly testing amplitude tracking, the retention of a slow envelope, and sensitivity to nonlinearity\. The GQKAN\-QKANFWP produces predictions that are visually indistinguishable from the ground truth from epoch 15 onward, with a±1σ\\pm 1\\sigmaband so narrow it remains within the line thickness of the mean curve across both the training and test regions\. In contrast, the QFWP baseline systematically under\-predicts the oscillation amplitude at every displayed epoch and fails to preserve the damping envelope; its mean prediction stabilizes as a low\-amplitude oscillation whose phase progressively drifts relative to the ground truth, and its variance band visibly broadens in the test region\. The contrast does not narrow as training progresses: even at epoch 100 the QFWP retains a clear amplitude deficit, indicating that this is a representational limit rather than an optimization gap\. The qualitative behavior is consistent with the roughly three\-orders\-of\-magnitude reduction in test MSE at epoch 50 achieved by GQKAN\-QKANFWP atN=64N=64on this dataset as shown in[Section˜6\.1](https://arxiv.org/html/2605.06734#S6.SS1)\.
##### Bessel function\.
[Figure˜9](https://arxiv.org/html/2605.06734#A3.F9)reports the learning trajectories on the second\-order Bessel function of the first kind,J2\(x\)J\_\{2\}\(x\), whose envelope decays as a power law \(∼x−1/2\\sim x^\{\-1/2\}\) rather than exponentially and whose local period drifts mildly withxx, in contrast to the strict periodicity of Damped SHM\. At epoch 15 the GQKAN\-QKANFWP already tracks both the period and the slowly\-decaying amplitude envelope, with a tight variance band along the entire window\. The QFWP captures the dominant frequency but consistently underestimates the early\-time amplitude and exhibits a small but persistent phase offset that accumulates over later cycles\. From epoch 30 onward GQKAN\-QKANFWP refines its amplitude estimate so as to overlay the ground truth, whereas the QFWP mean curve plateaus at a reduced amplitude and continues to drift in phase, particularly past the train/test split\. The seed\-averaged shaded bands further reveal that the QFWP exhibits noticeably greater run\-to\-run variability throughout training, while the GQKAN\-QKANFWP envelope remains essentially invisible at the displayed scale\. These observations align with two orders of magnitude reduction in test MSE at epoch 50 reported quantitatively, and together suggest that a single fixed\-depth additive update rule is insufficient to track multi\-scale oscillatory structure at long input windows\.
##### NARMA5\.
NARMA5 \([Figure˜10](https://arxiv.org/html/2605.06734#A3.F10)\) is a nonlinear autoregressive sequence of ordern0=5n\_\{0\}=5whose target contains sharp, irregular peaks driven by the order\-55autoregressive feedback in the recurrence\. The qualitative behavior atN=64N=64is informative for two reasons\. First, both models predict a near\-constant trajectory at epoch 15, reflecting the difficulty of identifying the underlying nonlinear dependence solely from a 64\-step input window\. Second, the two models diverge sharply thereafter: the GQKAN\-QKANFWP begins to recover the peak structure around epoch 30 and progressively sharpens its peaks through epoch 100, producing a mean curve that approximately tracks the ground\-truth maxima and minima with a moderate but tightening variance band\. The QFWP, by contrast, remains essentially flat across all four displayed epochs and exhibits a wide, weakly\-informative variance band that barely contracts during training\. This pattern is consistent with our broader observation in[Section˜6\.1](https://arxiv.org/html/2605.06734#S6.SS1)that QFWP undergoes substantial degradation at long input windows on the NARMA family, whereas the gated HQKAN\-based variants maintain stable predictive behavior\.
##### NARMA10\.
NARMA10 \([Figure˜11](https://arxiv.org/html/2605.06734#A3.F11)\) doubles the autoregressive order ton0=10n\_\{0\}=10and therefore amplifies the difficulty of capturing the autoregressive structure\. The qualitative behavior mirrors that of NARMA5 but with a more pronounced gap\. By epoch 50 the GQKAN\-QKANFWP resolves the larger peaks of the target, and by epoch 100 it tracks both the major and minor variations with a tight variance envelope\. The QFWP captures only a smoothed approximation of the trend, missing most of the peak\-to\-valley structure, and its variance band remains broadest in the early epochs and only modestly tightens over training\. The persistently high run\-to\-run variability of the QFWP at this longer\-memory setting indicates that the additive update rule has difficulty stabilizing across seeds, whereas the seed\-robustness of GQKAN\-QKANFWP is consistent with the geometric boundedness property of the gated update derived in[Section˜5](https://arxiv.org/html/2605.06734#S5)the fast parameters are constrained to the convex hull of the historical proposals, which prevents the unbounded additive accumulation that destabilizes QFWP at longNN\.
##### Delayed Quantum Control\.
The DQC task \([Figure˜12](https://arxiv.org/html/2605.06734#A3.F12)\) consists of localized pulses with a decaying envelope, and therefore requires the model to retain temporal structure across multiple delay intervals\. The GQKAN\-QKANFWP reproduces both the pulse shape and the decaying amplitude envelope from epoch 15 onward, with a variance band so narrow it remains within the mean curve\. The QFWP qualitatively tracks the dominant pulse structure in the training region but visibly degrades past the train/test split: the mean curve undershoots the pulse peaks, and the variance band broadens\. Although both models capture the gross periodicity of the signal, the quantitative gap between them spans roughly two orders of magnitude in test MSE at epoch 50, and this gap is most evident in the unseen test region\. This supports the claim that the gated update rule preserves long\-range temporal structure that the additive QFWP loses at long input windows—precisely the regime in which non\-Markovian feedback through the bound\-state\-in\-the\-continuum mechanism makes accurate forecasting most demanding\.
##### Jaynes–Cummings dynamics\.
The JC dataset \([Figure˜13](https://arxiv.org/html/2605.06734#A3.F13)\) combines rapid cavity\-qubit oscillations with dissipative photon\-loss decay, producing the highest\-frequency target in our benchmark suite\. Already at epoch 15 the GQKAN\-QKANFWP overlays the ground truth across the full sequence, capturing both the carrier oscillation and the slowly\-decaying amplitude envelope, with a±1σ\\pm 1\\sigmaband so narrow it remains within the line thickness of the mean curve\. Subsequent epochs \(30, 50, 100\) preserve this alignment with no visible drift, indicating that the model converges early and stably on this task\. The QFWP, in contrast, exhibits a persistent amplitude deficit at every displayed epoch—its mean curve resolves the carrier frequency but underestimates its amplitude by roughly half, and the variance band visibly broadens past the train/test split, with extrapolation errors growing toward the end of the sequence\. Although QFWP slowly recovers some amplitude through training, even at epoch 100 it fails to match the ground\-truth envelope, confirming that this is a representational rather than an optimization gap\. The visual gap between the two models is the most dramatic in our benchmark suite, mirroring the largest quantitative gap as well: the QFWP test MSE on this dataset atN=64N=64exceeds the corresponding GQKAN\-QKANFWP value by roughly three orders of magnitude\. We interpret this gap as a joint consequence of \(i\) the spectral expressivity of the HQKAN\-based fast programmer, which provides a rich Fourier basis well\-suited to high\-frequency dynamics, and \(ii\) the gated update rule, which prevents the additive accumulation of irrelevant high\-frequency parameter drift that the standard QFWP cannot suppress at long input windows\.
##### Summary of qualitative trends\.
Across all six tasks atN=64N=64, three consistent conclusions emerge\. First, GQKAN\-QKANFWP achieves an early\-epoch alignment with the ground truth that QFWP never matches on smooth\-dynamics and quantum\-dynamics tasks and reaches only partially on the NARMA family\. Second, the seed\-to\-seed variability of GQKAN\-QKANFWP remains negligible throughout training, whereas QFWP exhibits persistently wide variance bands that often expand beyond the train/test boundary\. Third, and most importantly, extending training to 100 epochs does not close the performance gap: the QFWP mean predictions stabilize at qualitatively incorrect amplitudes or frequencies, indicating a representational limitation rather than an optimization delay\. These extended\-horizon observations reinforce the quantitative results reported in[Section˜6\.1](https://arxiv.org/html/2605.06734#S6.SS1)and provide direct visual evidence that the gated update rule and HQKAN\-based fast\-weight programming framework enable stable, accurate, and seed\-robust long\-window forecasting in regimes where the standard QFWP fails to converge\.
Figure 8:Forecasting performance on the Damped SHM dataset \(Window\-size N=64\)\.Panels\(a\),\(b\),\(c\), and\(d\)illustrate the model predictions at training epochs 15, 30, 50, and 100, respectively\. Solid lines denote the mean prediction across five independent random seed initializations for the proposed GQKAN\-QKANFWP and the QFWP baseline\. The shaded region represents the±1σ\\pm 1\\sigmavariance envelope for each model, demonstrating the model’s stability across initializations\. The ground truth dynamics are shown in dashed charcoal, and the vertical dotted line indicates the boundary between the training/validation phase and the unseen test set\.Figure 9:Forecasting performance on the Bessel function dataset \(Window\-size N=64\)\.Panels\(a\),\(b\),\(c\), and\(d\)illustrate the model predictions at training epochs 15, 30, 50, and 100, respectively\. Solid lines denote the mean prediction across five independent random seed initializations for the proposed GQKAN\-QKANFWP and the QFWP baseline\. The shaded region represents the±1σ\\pm 1\\sigmavariance envelope for each model, demonstrating the model’s stability across initializations\. The ground truth dynamics are shown in dashed charcoal, and the vertical dotted line indicates the boundary between the training/validation phase and the unseen test set\.Figure 10:Forecasting performance on the NARMA5 dataset \(Window\-size N=64\)\.Panels\(a\),\(b\),\(c\), and\(d\)illustrate the model predictions at training epochs 15, 30, 50, and 100, respectively\. Solid lines denote the mean prediction across five independent random seed initializations for the proposed GQKAN\-QKANFWP and the QFWP baseline\. The shaded region represents the±1σ\\pm 1\\sigmavariance envelope for each model, demonstrating the model’s stability across initializations\. The ground truth dynamics are shown in dashed charcoal, and the vertical dotted line indicates the boundary between the training/validation phase and the unseen test set\.Figure 11:Forecasting performance on the NARMA10 dataset \(Window\-size N=64\)\.Panels\(a\),\(b\),\(c\), and\(d\)illustrate the model predictions at training epochs 15, 30, 50, and 100, respectively\. Solid lines denote the mean prediction across five independent random seed initializations for the proposed GQKAN\-QKANFWP and the QFWP baseline\. The shaded region represents the±1σ\\pm 1\\sigmavariance envelope for each model, demonstrating the model’s stability across initializations\. The ground truth dynamics are shown in dashed charcoal, and the vertical dotted line indicates the boundary between the training/validation phase and the unseen test set\.Figure 12:Forecasting performance on the Delayed Quantum Control dataset \(Window\-size N=64\)\.Panels\(a\),\(b\),\(c\), and\(d\)illustrate the model predictions at training epochs 15, 30, 50, and 100, respectively\. Solid lines denote the mean prediction across five independent random seed initializations for the proposed GQKAN\-QKANFWP and the QFWP baseline\. The shaded region represents the±1σ\\pm 1\\sigmavariance envelope for each model, demonstrating the model’s stability across initializations\. The ground truth dynamics are shown in dashed charcoal, and the vertical dotted line indicates the boundary between the training/validation phase and the unseen test set\.Figure 13:Forecasting performance on the Jaynes\-Cummings dataset \(Window\-size N=64\)\.Panels\(a\),\(b\),\(c\), and\(d\)illustrate the model predictions at training epochs 15, 30, 50, and 100, respectively\. Solid lines denote the mean prediction across five independent random seed initializations for the proposed GQKAN\-QKANFWP and the QFWP baseline\. The shaded region represents the±1σ\\pm 1\\sigmavariance envelope for each model, demonstrating the model’s stability across initializations\. The ground truth dynamics are shown in dashed charcoal, and the vertical dotted line indicates the boundary between the training/validation phase and the unseen test set\.Similar Articles
Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning
# Paper page - Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning Source: [https://huggingface.co/papers/2605.06734](https://huggingface.co/papers/2605.06734) Authors: , , , , , , , , , , , , , , , , , ## Abstract Quantum\-inspired fast\-weight programming framework using single\-qubit circuits achieves superior forecasting performance with reduced parameters compared to classical recurrent models while maintaining NISQ device compatibility\. [Fast Weight Programmers](https://huggingfac
Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting
This paper introduces Complementary Matrix Gating (CMG) for QKAN-based fast-weight programmers, enabling coordinate-wise memory control for quantum dynamics forecasting. The method shows consistent improvements and low mean-squared errors on quantum simulation benchmarks.
Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
This paper proposes quantum-inspired recurrent models (QKAN-FWPs) for traffic-matrix forecasting, demonstrating superior accuracy with fewer parameters compared to LSTM baselines.
Ultrafast machine learning on FPGAs via Kolmogorov-Arnold Networks
This post explains the author's Master's thesis on using Kolmogorov-Arnold Networks (KANs) for ultrafast machine learning on FPGAs, achieving sub-microsecond inference and online learning via custom hardware architectures. It references two accepted papers: KANELÉ for LUT-based evaluation (FPGA 2026 Best Paper) and a method for on-FPGA online learning (ICML 2026).
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning
This paper introduces a hybrid quantum-inspired Kolmogorov-Arnold network for privacy-aware federated learning of ECG data, demonstrating reduced parameters and communication costs while improving classification metrics compared to traditional MLP.