Titans-QFWP: A Regime-Aware Hybrid Quantum Fast Weight Programmer for Portfolio Optimization
Summary
Titans-QFWP is a hybrid reinforcement learning architecture integrating a Quantum Fast Weight Programmer with Titans-style memory for adaptive portfolio optimization, achieving strong performance on S&P 500 stocks.
View Cached Full Text
Cached at: 09/01/26, 01:07 PM
# A Regime-Aware Hybrid Quantum Fast Weight Programmer for Portfolio Optimization
Source: [https://arxiv.org/html/2608.29093](https://arxiv.org/html/2608.29093)
## Titans\-QFWP: A Regime\-Aware Hybrid Quantum Fast Weight Programmer for Portfolio OptimizationThanks:The views expressed in this article are those of the authors and do not represent the views of Wells Fargo\. This article is for informational purposes only\. Nothing contained in this article should be construed as investment advice\. Wells Fargo makes no express or implied warranties and expressly disclaims all legal, tax, and accounting implications related to this article\.Thanks:†These authors contributed equally to this work\.Thanks:\*Corresponding author\.
[https://orcid.org/0009-0002-5141-7909](https://orcid.org/0009-0002-5141-7909)Ming\-Kai Hung†Affiliation:Dept\. of Technology Application and Human Resource Development National Taiwan Normal University Taipei, Taiwan[https://orcid.org/0009-0000-3296-1824](https://orcid.org/0009-0000-3296-1824)Jun\-Hao Chen†Affiliation:Dept\. of Bioenvironmental Systems Engineering National Taiwan University Taipei, Taiwan[https://orcid.org/0000-0003-0114-4826](https://orcid.org/0000-0003-0114-4826)Samuel Yen\-Chi Chen\*Affiliation:Wells Fargo USAAffiliation:
###### Abstract
We propose Titans\-QFWP, a hybrid reinforcement learning architecture integrating a Quantum Fast Weight Programmer with Titans\-style memory \(Persistence, Surprise, and Forgetting\) for adaptive portfolio optimization\. To address high\-dimensional market features, we introduce an enhanced A3C2framework with Hungarian\-alignedKK\-means clustering and scaled log\-return rewards\. Evaluated on 468 S&P 500 stocks under an Equal\-Parameter\-Count \(EPC\) benchmark with approximately 3,000 trainable parameters, Titans\-QFWP achieves strong performance \(median ARR0\.42600\.4260, Calmar8\.55048\.5504, IR0\.84270\.8427\)\. Ablation results reveal that quantum gating fundamentally reshapes memory component roles, with Persistence supporting drawdown control, Surprise contributing to return generation, and Forgetting providing additional stabilization\. By stabilizing these quantum representations, the model enables defensive allocation during market drawdowns while preserving upside potential\.
###### Index Terms:
Deep Reinforcement Learning, Fast Weight Programmer, Portfolio Optimization, Quantum Machine Learning, Regime\-Aware Architecture
## IIntroduction
Financial markets are inherently time\-varying\[[1](https://arxiv.org/html/2608.29093#bib.bib1)\]and non\-stationary\[[2](https://arxiv.org/html/2608.29093#bib.bib2)\], driven by regime shifts and secular changes\[[3](https://arxiv.org/html/2608.29093#bib.bib3)\]that undermine traditional portfolio optimization\[[4](https://arxiv.org/html/2608.29093#bib.bib5),[5](https://arxiv.org/html/2608.29093#bib.bib4)\]\. Deep Reinforcement Learning \(DRL\) addresses these challenges through sequential policy optimization\[[6](https://arxiv.org/html/2608.29093#bib.bib6),[7](https://arxiv.org/html/2608.29093#bib.bib7),[8](https://arxiv.org/html/2608.29093#bib.bib8)\], but it often struggles with high computational costs\[[9](https://arxiv.org/html/2608.29093#bib.bib9)\]and high\-dimensional state spaces\[[10](https://arxiv.org/html/2608.29093#bib.bib10)\]\. To manage such expansive spaces, static clustering is typically employed; however, it remains fundamentally incapable of capturing abrupt market transitions\[[11](https://arxiv.org/html/2608.29093#bib.bib11)\]\. In the quantum domain, Variational Quantum Circuits \(VQC\) can exploit high\-dimensional Hilbert spaces\[[12](https://arxiv.org/html/2608.29093#bib.bib12)\], yet they lack adaptability due to fixed parameterization\[[13](https://arxiv.org/html/2608.29093#bib.bib13),[14](https://arxiv.org/html/2608.29093#bib.bib14),[15](https://arxiv.org/html/2608.29093#bib.bib15)\]\. Although the Quantum Fast Weight Programmer \(QFWP\) introduces dynamic weights\[[16](https://arxiv.org/html/2608.29093#bib.bib16),[17](https://arxiv.org/html/2608.29093#bib.bib17)\]to overcome this, its update mechanisms often degrade memory stability\. To address these challenges, we propose a dynamically adaptive hybrid architecture with two key contributions:
1. 1\.Titans\-QFWP Architecture\.Upgrading the original Q\-A3C2\[[18](https://arxiv.org/html/2608.29093#bib.bib20)\], we replace the VQC with Titans\-QFWP\. This architecture integrates QFWP with Titans\-style memory \(Persistence, Surprise, Forgetting\)\[[19](https://arxiv.org/html/2608.29093#bib.bib18)\]\.
2. 2\.Enhanced A3C2Framework\.We redesign the A3C2reinforcement learning environment\[[18](https://arxiv.org/html/2608.29093#bib.bib20)\], incorporating Hungarian\-alignedKK\-means clustering, a defensive cash action, and a scaled log\-return reward\[[20](https://arxiv.org/html/2608.29093#bib.bib19)\]\.
## IIMethodology
Fig\. 1:System architecture of the hybrid Titans\-QFWP agent integrated with the Enhanced A3C2continuous\-reward portfolio simulation environment\.### II\-AState Representation and Temporal Alignment
Using a 60\-day lookback, we rebalance everyD=20D=20market days\. For each assetiiat timett, the 5\- and 20\-day moving averages and 20\-day volatility of log\-returns \(Ri,t\(5\),Ri,t\(20\),Vi,t\(20\)R\_\{i,t\}^\{\(5\)\},R\_\{i,t\}^\{\(20\)\},V\_\{i,t\}^\{\(20\)\}\) are cross\-sectionallyzz\-scored to form𝐅t∈ℝN×3\\mathbf\{F\}\_\{t\}\\in\\mathbb\{R\}^\{N\\times 3\}\(N=468N=468\), and then partitioned intoK=10K=10clusters viaKK\-means\. To maintain temporal consistency, the Hungarian algorithm aligns centroids by minimizingL2L\_\{2\}drift:
πt∗=argmin∑k=0K−1π‖𝐜k,t−1−𝐜~π\(k\),t‖2\.\\pi\_\{t\}^\{\*\}=\\arg\\min\_\{\\pi\}\\sum\_\{k=0\}^\{K\-1\}\\\|\\mathbf\{c\}\_\{k,t\-1\}\-\\tilde\{\\mathbf\{c\}\}\_\{\\pi\(k\),t\}\\\|\_\{2\}\.\(1\)To filter short\-term market noise for the next time step, the matched centroids are smoothed with EMA \(λ=0\.8\\lambda=0\.8\) via𝐜k,t=λ𝐜k,t−1\+\(1−λ\)𝐜~πt∗\(k\),t\\mathbf\{c\}\_\{k,t\}=\\lambda\\mathbf\{c\}\_\{k,t\-1\}\+\(1\-\\lambda\)\\tilde\{\\mathbf\{c\}\}\_\{\\pi\_\{t\}^\{\*\}\(k\),t\}, initialized based on the 20\-day log\-return componentck,0\(R\(20\)\)c\_\{k,0\}^\{\(R^\{\(20\)\}\)\}\. Each cluster𝒞k\\mathcal\{C\}\_\{k\}yields five features: meansck,t\(X\)c\_\{k,t\}^\{\(X\)\}forX∈\{R\(5\),R\(20\),V\(20\)\}X\\in\\\{R^\{\(5\)\},R^\{\(20\)\},V^\{\(20\)\}\\\}, relative size\|𝒞k\|/N\|\\mathcal\{C\}\_\{k\}\|/N, and momentum dispersion\(\|𝒞k\|−1∑i∈𝒞k\(Ri,t\(20\)−ck,t\(R\(20\)\)\)2\)1/2\(\|\\mathcal\{C\}\_\{k\}\|^\{\-1\}\\sum\_\{i\\in\\mathcal\{C\}\_\{k\}\}\(R\_\{i,t\}^\{\(20\)\}\-c\_\{k,t\}^\{\(R^\{\(20\)\}\)\}\)^\{2\}\)^\{1/2\}\. Concatenating these5K5Kfeatures with three S&P 500 index features \(Rspx,t\(5\),Rspx,t\(20\),Vspx,t\(20\)R\_\{spx,t\}^\{\(5\)\},R\_\{spx,t\}^\{\(20\)\},V\_\{spx,t\}^\{\(20\)\}\) forms the 53\-dimensional state𝐬t\\mathbf\{s\}\_\{t\}\(empty clusters zero\-padded\)\.
### II\-BPortfolio Construction and Reward Function
At each 20\-day steptt, let𝐰target,t−1\\mathbf\{w\}\_\{\\mathrm\{target\},t\-1\}be the prior target weight vector and𝐆t=exp\(∑τ=120𝐫τ\)\\mathbf\{G\}\_\{t\}=\\exp\(\\sum\_\{\\tau=1\}^\{20\}\\mathbf\{r\}\_\{\\tau\}\)the compound gross return vector derived from daily log\-returns𝐫τ\\mathbf\{r\}\_\{\\tau\}\. The drifted weights are calculated as𝐰drift,t=\(𝐰target,t−1⊙𝐆t\)/\(𝐰target,t−1⊤𝐆t\)\\mathbf\{w\}\_\{\\mathrm\{drift\},t\}=\(\\mathbf\{w\}\_\{\\mathrm\{target\},t\-1\}\\odot\\mathbf\{G\}\_\{t\}\)/\(\\mathbf\{w\}\_\{\\mathrm\{target\},t\-1\}^\{\\top\}\\mathbf\{G\}\_\{t\}\), where⊙\\odotdenotes the element\-wise multiplication\. The new weights are then assigned following an inverse volatility strategy,𝐰target,t=𝝈t−1/‖𝝈t−1‖1\\mathbf\{w\}\_\{\\mathrm\{target\},t\}=\\bm\{\\sigma\}\_\{t\}^\{\-1\}/\\\|\\bm\{\\sigma\}\_\{t\}^\{\-1\}\\\|\_\{1\}, using the 60\-day historical volatility vector𝝈t\\bm\{\\sigma\}\_\{t\}\. If the cash actionat=Ka\_\{t\}=Kis chosen,𝐰target,t=𝟎\\mathbf\{w\}\_\{\\mathrm\{target\},t\}=\\mathbf\{0\}\. Rebalancing to these target weights at a cost ratec=0\.0015c=0\.0015yields a net survival fractionμt∈\(0,1\]\\mu\_\{t\}\\in\(0,1\], which is solved by successive substitution:μt=1−c‖𝐰target,t−μt𝐰drift,t‖1\\mu\_\{t\}=1\-c\\\|\\mathbf\{w\}\_\{\\mathrm\{target\},t\}\-\\mu\_\{t\}\\mathbf\{w\}\_\{\\mathrm\{drift\},t\}\\\|\_\{1\}\. With the gross portfolio returnRportfolio,t=𝐰target,t−1⊤\(𝐆t−𝟏\)R\_\{\\mathrm\{portfolio\},t\}=\\mathbf\{w\}\_\{\\mathrm\{target\},t\-1\}^\{\\top\}\(\\mathbf\{G\}\_\{t\}\-\\mathbf\{1\}\), the net realized return is given byRrealized,t=μt\(1\+Rportfolio,t\)−1R\_\{\\mathrm\{realized\},t\}=\\mu\_\{t\}\(1\+R\_\{\\mathrm\{portfolio\},t\}\)\-1\. To ensure stability, the reward floors this return at10−610^\{\-6\}and scales it by 100, treated as a hyperparameter:
rt=100⋅ln\(max\(1\+Rrealized,t,10−6\)\)\.r\_\{t\}=100\\cdot\\ln\\Bigl\(\\max\\bigl\(1\+R\_\{\\mathrm\{realized\},t\},\\,10^\{\-6\}\\bigr\)\\Bigr\)\.\(2\)
### II\-CTitans\-QFWP Architecture and Optimization
As illustrated in Fig\.[1](https://arxiv.org/html/2608.29093#S2.F1), the policy fuses Titans\-style memory with the Quantum Fast\-Weight Programmer \(QFWP\)\. First, we apply twotanh\\tanhactivations to compress the state𝐬t∈ℝ53\\mathbf\{s\}\_\{t\}\\in\\mathbb\{R\}^\{53\}into an encoded vector𝐱enc,t∈ℝnq\\mathbf\{x\}\_\{\\mathrm\{enc\},t\}\\in\\mathbb\{R\}^\{n\_\{q\}\}\(nq=8n\_\{q\}=8\)\. This is augmented by a trainable matrix𝐏∈ℝnq×np\\mathbf\{P\}\\in\\mathbb\{R\}^\{n\_\{q\}\\times n\_\{p\}\}\(np=4n\_\{p\}=4\) through an attention mechanism\. The attention weights are computed as𝐚t=softmax\(𝐖attn𝐱enc,t\)∈ℝnp\\mathbf\{a\}\_\{t\}=\\operatorname\{softmax\}\(\\mathbf\{W\}\_\{\\mathrm\{attn\}\}\\mathbf\{x\}\_\{\\mathrm\{enc\},t\}\)\\in\\mathbb\{R\}^\{n\_\{p\}\}with𝐖attn∈ℝnp×nq\\mathbf\{W\}\_\{\\mathrm\{attn\}\}\\in\\mathbb\{R\}^\{n\_\{p\}\\times n\_\{q\}\}, yielding the augmented representation:
𝐱in,t=𝐱enc,t\+𝐏𝐚t\.\\mathbf\{x\}\_\{\\mathrm\{in\},t\}=\\mathbf\{x\}\_\{\\mathrm\{enc\},t\}\+\\mathbf\{P\}\\mathbf\{a\}\_\{t\}\.\(3\)A slow\-program layer projects𝐱in,t\\mathbf\{x\}\_\{\\mathrm\{in\},t\}to𝐮t∈ℝhslow\\mathbf\{u\}\_\{t\}\\in\\mathbb\{R\}^\{h\_\{\\mathrm\{slow\}\}\}\(hslow=23h\_\{\\mathrm\{slow\}\}=23\)\. From𝐮t\\mathbf\{u\}\_\{t\}, three parallel softmax heads generate layer \(𝐩t\(l\)∈ℝL\\mathbf\{p\}\_\{t\}^\{\(l\)\}\\in\\mathbb\{R\}^\{L\},L=2L=2\), qubit \(𝐩t\(q\)∈ℝnq\\mathbf\{p\}\_\{t\}^\{\(q\)\}\\in\\mathbb\{R\}^\{n\_\{q\}\}\), and axis \(𝐩t\(a\)∈ℝ3\\mathbf\{p\}\_\{t\}^\{\(a\)\}\\in\\mathbb\{R\}^\{3\}for PauliX,Y,ZX,Y,Z\) distributions, whose normalized outer product forms the Surprise Tensor𝒮t=𝚵t/\(‖𝚵t‖F\+ε\)\\mathcal\{S\}\_\{t\}=\\bm\{\\Xi\}\_\{t\}/\(\\\|\\bm\{\\Xi\}\_\{t\}\\\|\_\{F\}\+\\varepsilon\), where𝚵t=𝐩t\(l\)⊗𝐩t\(q\)⊗𝐩t\(a\)\\bm\{\\Xi\}\_\{t\}=\\mathbf\{p\}\_\{t\}^\{\(l\)\}\\otimes\\mathbf\{p\}\_\{t\}^\{\(q\)\}\\otimes\\mathbf\{p\}\_\{t\}^\{\(a\)\}andε=10−8\\varepsilon=10^\{\-8\}\. Conditioned on𝐱in,t\\mathbf\{x\}\_\{\\mathrm\{in\},t\}, learned sigmoid gates for memory retention \(1−αt1\-\\alpha\_\{t\}\), momentum \(ηt\\eta\_\{t\}\), and surprise theta \(θt\\theta\_\{t\}\) dictate the evolution of fast\-weight memory from𝐌0=𝚯0=𝟎\\mathbf\{M\}\_\{0\}=\\bm\{\\Theta\}\_\{0\}=\\mathbf\{0\}:
𝐌t=ηt𝐌t−1−θt𝒮t,𝚯t=\(1−αt\)𝚯t−1\+𝐌t\.\\mathbf\{M\}\_\{t\}=\\eta\_\{t\}\\mathbf\{M\}\_\{t\-1\}\-\\theta\_\{t\}\\mathcal\{S\}\_\{t\},\\quad\\bm\{\\Theta\}\_\{t\}=\(1\-\\alpha\_\{t\}\)\\bm\{\\Theta\}\_\{t\-1\}\+\\mathbf\{M\}\_\{t\}\.\(4\)To serve as valid rotation angles, the raw memory is bounded via𝚯^t=clamp\(𝚯t,−π,π\)∈\[−π,π\]L×nq×3\\hat\{\\bm\{\\Theta\}\}\_\{t\}=\\operatorname\{clamp\}\(\\bm\{\\Theta\}\_\{t\},\-\\pi,\\pi\)\\in\[\-\\pi,\\pi\]^\{L\\times n\_\{q\}\\times 3\}to parameterize the variational quantum circuit\. Starting from uniform superposition\|ψ0⟩=H⊗nq\|0⟩⊗nq\|\\psi\_\{0\}\\rangle=H^\{\\otimes n\_\{q\}\}\|0\\rangle^\{\\otimes n\_\{q\}\}, a feature encoding layer\|ψenc⟩=⨂q=1nqRY\(πxin,t,q\)\|ψ0⟩\|\\psi\_\{\\mathrm\{enc\}\}\\rangle=\\bigotimes\_\{q=1\}^\{n\_\{q\}\}R\_\{Y\}\\bigl\(\\pi\\,x\_\{\\mathrm\{in\},t,q\}\\bigr\)\|\\psi\_\{0\}\\rangleprecedesLLlayers of CNOT entanglersUentU\_\{\\mathrm\{ent\}\}and parameterized rotationsUrot\(l\)\(𝚯^t\)=⨂q=1nq\[RZ\(Θ^t,l,q,2\)RY\(Θ^t,l,q,1\)RX\(Θ^t,l,q,0\)\]U\_\{\\mathrm\{rot\}\}^\{\(l\)\}\\bigl\(\\hat\{\\bm\{\\Theta\}\}\_\{t\}\\bigr\)=\\bigotimes\_\{q=1\}^\{n\_\{q\}\}\\\!\\Bigl\[R\_\{Z\}\\bigl\(\\hat\{\\Theta\}\_\{t,l,q,2\}\\bigr\)R\_\{Y\}\\bigl\(\\hat\{\\Theta\}\_\{t,l,q,1\}\\bigr\)R\_\{X\}\\bigl\(\\hat\{\\Theta\}\_\{t,l,q,0\}\\bigr\)\\Bigr\]\. Finally, the output state\|ψout⟩=∏l=1L\(Urot\(l\)Uent\)\|ψenc⟩\|\\psi\_\{\\mathrm\{out\}\}\\rangle=\\prod\_\{l=1\}^\{L\}\(U\_\{\\mathrm\{rot\}\}^\{\(l\)\}U\_\{\\mathrm\{ent\}\}\)\|\\psi\_\{\\mathrm\{enc\}\}\\rangleyields the expectation vector𝐳t∈\[−1,1\]nq\\mathbf\{z\}\_\{t\}\\in\[\-1,1\]^\{n\_\{q\}\}, composed of the Pauli\-Z expectation values⟨ψout\|Zq\|ψout⟩\\langle\\psi\_\{\\text\{out\}\}\|Z\_\{q\}\|\\psi\_\{\\text\{out\}\}\\ranglefor each qubit\. This vector is subsequently processed through a ReLU activation function to generate actor and critic heads for cluster selection, with the constituent stocks weighted by inverse volatility\.
## IIIExperimental Setup
Dataset\.Daily log\-returns of S&P 500 stocks \(Jan\. 2015–Apr\. 2026\) are used\. Assets with more than 10% missing observations are excluded, yielding 468 stocks\. All features are constructed using only the information available at timett\. The data are divided into 2,262 training days \(2015\-01\-06 to 2023\-12\-29\), 252 validation days \(2024\-01\-02 to 2024\-12\-31\), and 332 test days \(2025\-01\-02 to 2026\-04\-30\)\.
Equal\-Parameter\-Count \(EPC\) Benchmark\.For fair comparison, all 11 architectures are limited to approximately 3,000 trainable parameters\. Inspired by Titans \(Memory\-as\-Context\)\[[19](https://arxiv.org/html/2608.29093#bib.bib18)\], our proposed Titans\-QFWP \(Full\) is evaluated alongside its variants \(allocations detailed in Table[I](https://arxiv.org/html/2608.29093#S3.T1)\)\.
TABLE I:Parameter configuration for the EPC benchmark\.ArchitectureTotalQuantumClassicalClassical ArchitecturesTitans \(Memory\-as\-Context\)2,99702,997FWP3,00003,000Titans\-FWP \(Full\)2,99902,999Titans\-FWP \(−\-Forgetting\)2,99002,990Titans\-FWP \(−\-Persistence\)2,93102,931Titans\-FWP \(−\-Surprise\)2,98102,981Quantum ArchitecturesQFWP3,0012612,740Titans\-QFWP \(Full\)3,0001,2531,747Titans\-QFWP \(−\-Forgetting\)2,9911,2531,738Titans\-QFWP \(−\-Persistence\)2,9321,2531,679Titans\-QFWP \(−\-Surprise\)2,9821,2531,729
## IVExperimental Results and Discussion
### IV\-AOverall Analysis
Results over 10 random seeds \(1–10\), reported as median \[IQR\] in Table[II](https://arxiv.org/html/2608.29093#S4.T2), are compared with the market benchmark \(S&P 500 Buy & Hold\)\. Titans\-QFWP \(Full\) achieves higher Annualized Rate of Return \(ARR;0\.42600\.4260vs\.0\.25280\.2528\), Calmar Ratio \(Calmar;8\.55048\.5504vs\.7\.16387\.1638\), and the highest Information Ratio \(IR;0\.84270\.8427\)\. In contrast, the market benchmark exhibits lower Maximum Drawdown \(MDD;−0\.0353\-0\.0353\) and higher Sortino Ratio \(Sortino;8\.54058\.5405\), partly due to the exclusion of transaction costs\. These results demonstrate that Titans\-QFWP effectively captures market trends under the strong bull\-market regime\.
TABLE II:S&P 500 out\-of\-sample performance across asset allocation baselines \(median \[IQR\], seeds 1–10\)\.ArchitectureARRMDDSortinoCalmarIRMarket BenchmarkS&P 500 \(Buy & Hold\)0\.25280\.2528−0\.0353\-0\.03538\.54058\.54057\.16387\.1638N/ABaseline & Classical ArchitecturesTitans \(Memory\-as\-Context\)0\.2557\[0\.2649\]0\.2557\\;\\;\[0\.2649\]−0\.0512\[0\.0355\]\\mathbf\{\-0\.0512\}\\;\\;\[\\mathbf\{0\.0355\}\]3\.3645\[4\.4568\]3\.3645\\;\\;\[4\.4568\]3\.9077\[7\.2651\]3\.9077\\;\\;\[7\.2651\]0\.1374\[1\.8125\]0\.1374\\;\\;\[1\.8125\]FWP0\.3423\[0\.1904\]\\mathbf\{0\.3423\}\\;\\;\[\\mathbf\{0\.1904\}\]−0\.0918\[0\.1046\]\-0\.0918\\;\\;\[0\.1046\]3\.3498\[6\.1209\]3\.3498\\;\\;\[6\.1209\]3\.9429\[7\.7964\]3\.9429\\;\\;\[7\.7964\]0\.4447\[1\.3117\]\\mathbf\{0\.4447\}\\;\\;\[\\mathbf\{1\.3117\}\]Titans\-FWP \(Full\)0\.3175\[0\.2013\]0\.3175\\;\\;\[0\.2013\]−0\.0606\[0\.0648\]\-0\.0606\\;\\;\[0\.0648\]4\.9382\[3\.6665\]\\mathbf\{4\.9382\}\\;\\;\[\\mathbf\{3\.6665\}\]6\.2776\[5\.6673\]\\mathbf\{6\.2776\}\\;\\;\[\\mathbf\{5\.6673\}\]0\.3891\[0\.9531\]0\.3891\\;\\;\[0\.9531\]Titans\-FWP \(−\-Forgetting\)0\.2752\[0\.1713\]0\.2752\\;\\;\[0\.1713\]−0\.0790\[0\.0453\]\-0\.0790\\;\\;\[0\.0453\]3\.5281\[3\.4122\]3\.5281\\;\\;\[3\.4122\]3\.4519\[4\.0856\]3\.4519\\;\\;\[4\.0856\]0\.1993\[1\.0045\]0\.1993\\;\\;\[1\.0045\]Titans\-FWP \(−\-Persistence\)0\.2841\[0\.2500\]0\.2841\\;\\;\[0\.2500\]−0\.0702\[0\.0540\]\-0\.0702\\;\\;\[0\.0540\]4\.1309\[2\.8475\]4\.1309\\;\\;\[2\.8475\]4\.5385\[3\.4245\]4\.5385\\;\\;\[3\.4245\]0\.1927\[1\.2757\]0\.1927\\;\\;\[1\.2757\]Titans\-FWP \(−\-Surprise\)0\.1907\[0\.1834\]0\.1907\\;\\;\[0\.1834\]−0\.0774\[0\.0731\]\-0\.0774\\;\\;\[0\.0731\]2\.8587\[3\.8253\]2\.8587\\;\\;\[3\.8253\]2\.5238\[6\.5771\]2\.5238\\;\\;\[6\.5771\]−0\.4547\[1\.1375\]\-0\.4547\\;\\;\[1\.1375\]Quantum ArchitecturesQFWP0\.1806\[0\.1437\]0\.1806\\;\\;\[0\.1437\]−0\.1017\[0\.0887\]\-0\.1017\\;\\;\[0\.0887\]1\.8849\[1\.1011\]1\.8849\\;\\;\[1\.1011\]1\.5299\[2\.2969\]1\.5299\\;\\;\[2\.2969\]−0\.4002\[0\.8815\]\-0\.4002\\;\\;\[0\.8815\]Titans\-QFWP \(Full\)0\.4260\[0\.4197\]\\mathbf\{0\.4260\}\\;\\;\[\\mathbf\{0\.4197\}\]−0\.0500\[0\.0212\]\\mathbf\{\-0\.0500\}\\;\\;\[\\mathbf\{0\.0212\}\]6\.9018\[6\.5930\]\\mathbf\{6\.9018\}\\;\\;\[\\mathbf\{6\.5930\}\]8\.5504\[6\.0697\]\\mathbf\{8\.5504\}\\;\\;\[\\mathbf\{6\.0697\}\]0\.8427\[1\.3992\]\\mathbf\{0\.8427\}\\;\\;\[\\mathbf\{1\.3992\}\]Titans\-QFWP \(−\-Forgetting\)0\.1965\[0\.1276\]0\.1965\\;\\;\[0\.1276\]−0\.0520\[0\.0415\]\-0\.0520\\;\\;\[0\.0415\]3\.0186\[1\.3646\]3\.0186\\;\\;\[1\.3646\]3\.3316\[2\.0293\]3\.3316\\;\\;\[2\.0293\]−0\.3887\[0\.9337\]\-0\.3887\\;\\;\[0\.9337\]Titans\-QFWP \(−\-Persistence\)0\.1929\[0\.1297\]0\.1929\\;\\;\[0\.1297\]−0\.1010\[0\.0400\]\-0\.1010\\;\\;\[0\.0400\]2\.0905\[1\.1939\]2\.0905\\;\\;\[1\.1939\]2\.1663\[1\.7055\]2\.1663\\;\\;\[1\.7055\]−0\.3136\[0\.6757\]\-0\.3136\\;\\;\[0\.6757\]Titans\-QFWP \(−\-Surprise\)0\.1586\[0\.1749\]0\.1586\\;\\;\[0\.1749\]−0\.0782\[0\.0318\]\-0\.0782\\;\\;\[0\.0318\]2\.6949\[1\.5092\]2\.6949\\;\\;\[1\.5092\]2\.3728\[1\.3064\]2\.3728\\;\\;\[1\.3064\]−0\.5138\[1\.1810\]\-0\.5138\\;\\;\[1\.1810\]
### IV\-BAblation Diagnostics
#### IV\-B1Classical Ablation \(Titans\-FWP\)
In classical variants \(Table[III](https://arxiv.org/html/2608.29093#S4.T3)\), removing Surprise produces the largest degradation in ARR, Sortino, Calmar, and IR\. Removing Forgetting results in the largest deterioration in MDD\. Removing Persistence leads to smaller performance changes across most metrics\.
TABLE III:Performance metrics for Titans\-FWP ablation variants \(median \[IQR\], seeds 1–10\)\.ArchitectureARRMDDSortinoCalmarIRClassical Ablation Study: Titans\-FWPFWP \(Baseline\)0\.3423\[0\.1904\]\\mathbf\{0\.3423\}\\;\\;\[\\mathbf\{0\.1904\}\]−0\.0918\[0\.1046\]\-0\.0918\\;\\;\[0\.1046\]3\.3498\[6\.1209\]3\.3498\\;\\;\[6\.1209\]3\.9429\[7\.7964\]3\.9429\\;\\;\[7\.7964\]0\.4447\[1\.3117\]\\mathbf\{0\.4447\}\\;\\;\[\\mathbf\{1\.3117\}\]Titans\-FWP \(Full\)0\.3175\[0\.2013\]0\.3175\\;\\;\[0\.2013\]−0\.0606\[0\.0648\]\\mathbf\{\-0\.0606\}\\;\\;\[\\mathbf\{0\.0648\}\]4\.9382\[3\.6665\]\\mathbf\{4\.9382\}\\;\\;\[\\mathbf\{3\.6665\}\]6\.2776\[5\.6673\]\\mathbf\{6\.2776\}\\;\\;\[\\mathbf\{5\.6673\}\]0\.3891\[0\.9531\]0\.3891\\;\\;\[0\.9531\]Titans\-FWP \(−\-Forgetting\)0\.2752\[0\.1713\]0\.2752\\;\\;\[0\.1713\]−0\.0790\[0\.0453\]\-0\.0790\\;\\;\[0\.0453\]3\.5281\[3\.4122\]3\.5281\\;\\;\[3\.4122\]3\.4519\[4\.0856\]3\.4519\\;\\;\[4\.0856\]0\.1993\[1\.0045\]0\.1993\\;\\;\[1\.0045\]Titans\-FWP \(−\-Persistence\)0\.2841\[0\.2500\]0\.2841\\;\\;\[0\.2500\]−0\.0702\[0\.0540\]\-0\.0702\\;\\;\[0\.0540\]4\.1309\[2\.8475\]4\.1309\\;\\;\[2\.8475\]4\.5385\[3\.4245\]4\.5385\\;\\;\[3\.4245\]0\.1927\[1\.2757\]0\.1927\\;\\;\[1\.2757\]Titans\-FWP \(−\-Surprise\)0\.1907\[0\.1834\]0\.1907\\;\\;\[0\.1834\]−0\.0774\[0\.0731\]\-0\.0774\\;\\;\[0\.0731\]2\.8587\[3\.8253\]2\.8587\\;\\;\[3\.8253\]2\.5238\[6\.5771\]2\.5238\\;\\;\[6\.5771\]−0\.4547\[1\.1375\]\-0\.4547\\;\\;\[1\.1375\]
#### IV\-B2Quantum Ablation \(Titans\-QFWP\)
In quantum variants \(Table[IV](https://arxiv.org/html/2608.29093#S4.T4)\), removing Persistence causes the largest degradation in the median MDD, Sortino, and Calmar, while removing Surprise leads to the lowest median ARR and IR\. These patterns suggest that quantum gating alters memory roles, with greater expressive capacity favoring Persistence over Forgetting for representation preservation\.
TABLE IV:Performance metrics for Titans\-QFWP ablation variants \(median \[IQR\], seeds 1–10\)\.ArchitectureARRMDDSortinoCalmarIRQuantum Ablation Study: Titans\-QFWPQFWP \(Baseline\)0\.1806\[0\.1437\]0\.1806\\;\\;\[0\.1437\]−0\.1017\[0\.0887\]\-0\.1017\\;\\;\[0\.0887\]1\.8849\[1\.1011\]1\.8849\\;\\;\[1\.1011\]1\.5299\[2\.2969\]1\.5299\\;\\;\[2\.2969\]−0\.4002\[0\.8815\]\-0\.4002\\;\\;\[0\.8815\]Titans\-QFWP \(Full\)0\.4260\[0\.4197\]\\mathbf\{0\.4260\}\\;\\;\[\\mathbf\{0\.4197\}\]−0\.0500\[0\.0212\]\\mathbf\{\-0\.0500\}\\;\\;\[\\mathbf\{0\.0212\}\]6\.9018\[6\.5930\]\\mathbf\{6\.9018\}\\;\\;\[\\mathbf\{6\.5930\}\]8\.5504\[6\.0697\]\\mathbf\{8\.5504\}\\;\\;\[\\mathbf\{6\.0697\}\]0\.8427\[1\.3992\]\\mathbf\{0\.8427\}\\;\\;\[\\mathbf\{1\.3992\}\]Titans\-QFWP \(−\-Forgetting\)0\.1965\[0\.1276\]0\.1965\\;\\;\[0\.1276\]−0\.0520\[0\.0415\]\-0\.0520\\;\\;\[0\.0415\]3\.0186\[1\.3646\]3\.0186\\;\\;\[1\.3646\]3\.3316\[2\.0293\]3\.3316\\;\\;\[2\.0293\]−0\.3887\[0\.9337\]\-0\.3887\\;\\;\[0\.9337\]Titans\-QFWP \(−\-Persistence\)0\.1929\[0\.1297\]0\.1929\\;\\;\[0\.1297\]−0\.1010\[0\.0400\]\-0\.1010\\;\\;\[0\.0400\]2\.0905\[1\.1939\]2\.0905\\;\\;\[1\.1939\]2\.1663\[1\.7055\]2\.1663\\;\\;\[1\.7055\]−0\.3136\[0\.6757\]\-0\.3136\\;\\;\[0\.6757\]Titans\-QFWP \(−\-Surprise\)0\.1586\[0\.1749\]0\.1586\\;\\;\[0\.1749\]−0\.0782\[0\.0318\]\-0\.0782\\;\\;\[0\.0318\]2\.6949\[1\.5092\]2\.6949\\;\\;\[1\.5092\]2\.3728\[1\.3064\]2\.3728\\;\\;\[1\.3064\]−0\.5138\[1\.1810\]\-0\.5138\\;\\;\[1\.1810\]
### IV\-CPortfolio Dynamics and Asset Selection
As shown in Fig\.[2](https://arxiv.org/html/2608.29093#S4.F2)\(Seed 1\), Titans\-QFWP achieves a cumulative return of \+30\.33%, slightly outperforming the S&P 500 \(\+28\.46%\), driven by allocations to higher\-return clusters \(e\.g\., Clusters 7 and 9\)\. The model adaptively reallocates to cash during market downturns\.
Averaged across 10 seeds \(1–10\), the top\-20 selected stocks exhibit concentrated outperformance in technology sub\-sectors, including data storage, optical networking, photonics, and memory semiconductors \(Fig\.[3](https://arxiv.org/html/2608.29093#S4.F3)\)\. Measured by active return \(excess simple return over the S&P 500\), WDC and CIEN outperform the benchmark at every time step, with peak active returns of \+42\.7% and \+51\.6%\. In addition, COHR, LITE, and MU also contribute substantial excess returns\.
Fig\. 2:Dynamic cluster selection of Titans\-QFWP vs\. the S&P 500 index\.Fig\. 3:Top\-20 selected stocks of Titans\-QFWP vs\. the S&P 500 index\.
## VLimitations
Limitations include survivorship bias, a short out\-of\-sample period, seed variability, simplified transaction costs, and evaluation of quantum components usinglightning\.qubitquantum simulation\. Risk\-adjusted metrics are reported descriptively rather than as statistically conclusive evidence\.
## VIConclusion
Integrating QFWP with Titans\-style memory provides a robust framework for addressing market non\-stationarity\. Crucially, the expressive capacity of quantum gating fundamentally alters learning dynamics: it shifts the mechanism of representation preservation, relying more on Persistence rather than Forgetting for drawdown control\. Although frictionless passive benchmarks exhibit a marginal downside advantage during bull markets, Titans\-QFWP mitigates this through regime\-aware allocation, dynamically transitioning between defensive and high\-performing asset clusters\. These findings illustrate how quantum\-classical memory architectures reorganize representation learning, advancing adaptive financial reinforcement learning\.
## References
- \[1\]A\. M\. Ozbayoglu, M\. U\. Gudelek, and O\. B\. Sezer\(2020\)Deep learning for financial applications: a survey\.Applied Soft Computing93,pp\. 106384\.External Links:[Document](https://dx.doi.org/10.1016/j.asoc.2020.106384)Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[2\]J\. D\. Hamilton\(1989\)A new approach to the economic analysis of nonstationary time series and the business cycle\.Econometrica: Journal of the Econometric Society57\(2\),pp\. 357–384\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[3\]A\. Ang and A\. Timmermann\(2012\)Regime changes and financial markets\.Annual Review of Financial Economics4\(1\),pp\. 313–337\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[4\]F\. J\. Fabozzi, H\. M\. Markowitz, and F\. Gupta\(2008\)Portfolio selection\.InHandbook of finance,Vol\.2,pp\. 3–13\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[5\]M\. Lopez de Prado\(2016\)Building diversified portfolios that outperform out\-of\-sample\.The Journal of Portfolio Management42\(4\),pp\. 59–69\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[6\]D\. Cao, Y\. Liu, Y\. Wang, Q\. Zhang, and W\. Hu\(2025\)Deep reinforcement learning in power systems resilience: a review\.IEEE Transactions on Reliability74\(4\),pp\. 5356–5370\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[7\]K\. Arulkumaran, M\. P\. Deisenroth, M\. Brundage, and A\. A\. Bharath\(2017\)Deep reinforcement learning: a brief survey\.IEEE Signal Processing Magazine34\(6\),pp\. 26–38\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[8\]P\. Henderson, R\. Islam, P\. Bachman, J\. Pineau, D\. Precup, and D\. Meger\(2018\)Deep reinforcement learning that matters\.InProceedings of the 32nd AAAI Conference on Artificial Intelligence,pp\. 3207–3214\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[9\]V\. Mnih, A\. P\. Badia, M\. Mirza, A\. Graves, T\. Lillicrap, T\. Harley, D\. Silver, and K\. Kavukcuoglu\(2016\)Asynchronous Methods for Deep Reinforcement Learning\.In33rd International Conference on Machine Learning \(ICML\),pp\. 1928–1937\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[10\]H\. Yang, X\. Liu, S\. Zhong, and A\. Walid\(2020\)Deep reinforcement learning for automated stock trading: an ensemble strategy\.In1st ACM International Conference on AI in Finance \(ICAIF\),pp\. 1–8\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[11\]S\. Aghabozorgi, A\. S\. Shirkhorshidi, and T\. Y\. Wah\(2015\)Time\-series clustering–a decade review\.Information Systems53,pp\. 16–38\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[12\]M\. Schuld and N\. Killoran\(2019\)Quantum machine learning in feature Hilbert spaces\.Physical Review Letters122\(4\),pp\. 040504\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[13\]K\. Bharti, A\. Cervera\-Lierta, T\. H\. Kyaw, T\. Haug, S\. Alperin\-Lea, A\. Anand, M\. Degroote, H\. Heimonen, J\. S\. Kottmann, T\. Menke, W\. Mok, S\. Sim, L\. Kwek, and A\. Aspuru\-Guzik\(2022\)Noisy intermediate\-scale quantum algorithms\.Reviews of Modern Physics94\(1\),pp\. 015004\.External Links:[Document](https://dx.doi.org/10.1103/RevModPhys.94.015004),[Link](https://doi.org/10.1103/RevModPhys.94.015004)Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[14\]M\. Cerezo, A\. Arrasmith, R\. Babbush, S\. C\. Benjamin, S\. Endo, K\. Fujii, J\. R\. McClean, K\. Mitarai, X\. Yuan, L\. Cincio, and P\. J\. Coles\(2021\)Variational quantum algorithms\.Nature Reviews Physics3\(9\),pp\. 625–644\.External Links:[Document](https://dx.doi.org/10.1038/s42254-021-00348-9),[Link](https://doi.org/10.1038/s42254-021-00348-9)Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[15\]A\. Peruzzo, J\. McClean, P\. Shadbolt, M\. Yung, X\. Zhou, P\. J\. Love, A\. Aspuru\-Guzik, and J\. L\. O’brien\(2014\)A variational eigenvalue solver on a photonic quantum processor\.Nature communications5\(1\),pp\. 4213\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[16\]S\. Y\. C\. Chen\(2024\)Learning to program variational quantum circuits with fast weights\.In2024 International Joint Conference on Neural Networks \(IJCNN\),pp\. 1–9\.Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[17\]J\. Chen, M\. Hung, Y\. Tsai, and S\. Y\. Chen\(2026\)Batched training for QLSTM vs\. QFWP: a system\-oriented approach to EPC\-aware RMSE\-DA\.InIEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Cited by:[§I](https://arxiv.org/html/2608.29093#S1.p1.1)\.
- \[18\]Y\. Liu, Y\. Tsai, and S\. Y\. Chen\(2026\)Q\-A3C2\{\}^\{2\}: quantum reinforcement learning with time\-series dynamic clustering for adaptive ETF stock selection\.InIEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 22012–22016\.Cited by:[item 1](https://arxiv.org/html/2608.29093#S1.I1.i1.p1.1),[item 2](https://arxiv.org/html/2608.29093#S1.I1.i2.p1.1)\.
- \[19\]A\. Behrouz, P\. Zhong, and V\. Mirrokni\(2025\)Titans: learning to memorize at test time\.arXiv preprint arXiv:2501\.00663\.External Links:2501\.00663Cited by:[item 1](https://arxiv.org/html/2608.29093#S1.I1.i1.p1.1),[§III](https://arxiv.org/html/2608.29093#S3.p2.1)\.
- \[20\]Z\. Jiang, D\. Xu, and J\. Liang\(2017\)A deep reinforcement learning framework for the financial portfolio management problem\.arXiv preprint arXiv:1706\.10059\.External Links:1706\.10059Cited by:[item 2](https://arxiv.org/html/2608.29093#S1.I1.i2.p1.1)\.Similar Articles
Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
This paper proposes quantum-inspired recurrent models (QKAN-FWPs) for traffic-matrix forecasting, demonstrating superior accuracy with fewer parameters compared to LSTM baselines.
Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning
# Paper page - Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning Source: [https://huggingface.co/papers/2605.06734](https://huggingface.co/papers/2605.06734) Authors: , , , , , , , , , , , , , , , , , ## Abstract Quantum\-inspired fast\-weight programming framework using single\-qubit circuits achieves superior forecasting performance with reduced parameters compared to classical recurrent models while maintaining NISQ device compatibility\. [Fast Weight Programmers](https://huggingfac
Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting
This paper introduces Complementary Matrix Gating (CMG) for QKAN-based fast-weight programmers, enabling coordinate-wise memory control for quantum dynamics forecasting. The method shows consistent improvements and low mean-squared errors on quantum simulation benchmarks.
Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning
This paper introduces Gated QKAN-FWP, a scalable quantum-inspired sequence learning framework that combines Fast Weight Programmers with Kolmogorov-Arnold Networks using single-qubit data re-uploading circuits.
A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management
A three-phase deep reinforcement learning system for personalized portfolio management that addresses ticker lock-in, monolithic objectives, and static user models, using a cross-asset encoder pretrained with self-supervised learning and the Chronos time series foundation model, fine-tuned with Mixture of Experts and PPO, and personalized via LoRA.