DualCast: A Dual-Path Language Model for Bimodal Financial Time-Series Forecasting
Summary
DualCast is a dual-path LLM framework that extends a frozen Qwen3-8B backbone with a discrete financial vocabulary for bimodal financial time-series forecasting, combining a fast numerical path with a news-conditioned slow revision path optimized via group relative policy optimization.
View Cached Full Text
Cached at: 10/02/26, 09:48 AM
# A Dual-Path Language Model for Bimodal Financial Time-Series Forecasting
Source: [https://arxiv.org/html/2609.38197](https://arxiv.org/html/2609.38197)
Journal:Nuclear Physics BWentao Zhao‡Affiliation:Department of Automation, Tsinghua University, Beijing, 100084, ChinaAffiliation:School of Software, Tsinghua University, Beijing, 100084, ChinaHongqiang Wu‡Affiliation:School of Computer Science, Northwestern Polytechnical University, Xi’an, 710129, ChinaAffiliation:Inspur Yunzhou Industrial Internet Co\. Ltd, Jinan, 250101, ChinaZhaochen ZanAffiliation:Department of Automation, Tsinghua University, Beijing, 100084, ChinaYu Zhang†Affiliation:School of Computer Science, Northwestern Polytechnical University, Xi’an, 710129, ChinaBiqing Huang†Affiliation:Department of Automation, Tsinghua University, Beijing, 100084, China
###### Abstract
Financial time\-series forecasting must capture price dynamics across heterogeneous assets while incorporating news available at prediction time\. We introduce DualCast, a dual\-path framework that extends a frozen language model with a discrete financial vocabulary\. Each log\-return patch is represented by a learned summary token and three residual shape tokens, preserving local drift and volatility while allowing shape patterns to be shared across assets\. To improve codebook utilization, we develop adaptive frequency\-equalizing residual vector quantization, which rebalances overloaded codewords without compromising reconstruction accuracy\. The fast path trains only the new financial\-token embeddings and output heads on a frozen Qwen3\-8B backbone\. A toggleable LoRA adapter enables a slow path that conditions on the fast forecast and news available at the forecast origin to produce a revised prediction\. The reviser is initialized by supervised fine\-tuning and further optimized with a return\-space group relative policy optimization objective that rewards improvements over the fast forecast\. In zero\-shot evaluations covering equities and energy prices at five\-minute, daily, and weekly resolutions, the slow path achieves the lowest mean absolute percentage error among the compared methods in 8 of 12 dataset–horizon settings, including every longest\-horizon setting\. News ablations indicate additional gains in most tested settings, although their magnitude varies across markets\. DualCast thus combines a fast numerical forecaster with an optional text\-conditioned revision mechanism\.
###### Keywords:
time\-series forecasting , large language models , financial forecasting , vector quantization , multimodal learning , reinforcement learning
## 1Introduction
Despite recent progress in applying large language models \(LLMs\) to time\-series forecasting, financial markets remain particularly challenging\. First, financial data are highly heterogeneous, exhibit low signal\-to\-noise ratios, and evolve under rapidly changing volatility regimes, creating a substantial gap from conventional forecasting benchmarks and making stable pattern learning difficult\. Second, financial forecasting requires not only modeling historical price dynamics but also understanding and reasoning over external information such as news\. While LLMs excel at language understanding and knowledge integration, they are not naturally designed for numerical prediction and time\-series modeling\. Bridging these two requirements remains a central challenge for financial foundation models\.
Most modern forecasting models, from patch\-based Transformers to large\-scale time\-series foundation models, are designed for signals that are smooth, strongly periodic, and governed by relatively stable dynamics, such as electricity load, traffic, and weather\. Financial markets differ fundamentally: they are noisy, weakly autocorrelated, and characterized by rapidly changing volatility regimes\. Moreover, financial time series are highly heterogeneous across assets, with volatility levels often spanning multiple orders of magnitude\. As a result, representations learned under one scale regime frequently fail to generalize to others, while models trained on high\-variance assets tend to overlook the subtle signals present in lower\-variance markets\. Bridging this gap requires a representation that captures transferable market patterns while remaining robust to large differences in scale and volatility\.
A substantial fraction of price\-moving information originates outside historical price trajectories\. Earnings announcements, policy decisions, macroeconomic releases, and the collective sentiment reflected in financial news and social media continuously shape market behavior\. Consequently, even highly expressive numerical forecasting models remain structurally blind to information that has already been disclosed through text\. While recent large language models offer powerful language understanding and knowledge integration capabilities, they are not naturally optimized for numerical forecasting\. Existing foundation models therefore tend to specialize in either numerical modeling or language reasoning, but rarely both\. Closing this gap requires a bimodal forecasting framework that jointly reasons over historical market dynamics and contemporaneous textual information\.
Large language models \(LLMs\) provide a promising foundation for addressing both challenges simultaneously: they already encode rich semantic knowledge and reasoning capabilities over text, while their autoregressive architecture offers a natural framework for sequential prediction\. Recent efforts have adapted LLMs to time\-series forecasting through three main paradigms: reprogramming continuous signals into the embedding space of a frozen LLM\[[11](https://arxiv.org/html/2609.38197#bib.bib1),[36](https://arxiv.org/html/2609.38197#bib.bib3),[23](https://arxiv.org/html/2609.38197#bib.bib10)\], converting numerical values into textual representations and leveraging next\-token prediction directly\[[9](https://arxiv.org/html/2609.38197#bib.bib2)\], and quantizing time series into discrete tokens so that forecasting becomes language modeling over an extended vocabulary\[[2](https://arxiv.org/html/2609.38197#bib.bib4),[24](https://arxiv.org/html/2609.38197#bib.bib13)\]\. For financial forecasting, we argue that the quantization paradigm is particularly appealing\. By representing market dynamics as discrete tokens, it places price trajectories and textual information within a unified autoregressive sequence over a shared vocabulary, enabling joint modeling of historical prices and external news\. More importantly, quantization transforms forecasting into a discrete prediction problem, producing verifiable targets that naturally support reinforcement learning with outcome\-based rewards\. Recent studies have shown that such verifiable rewards are effective for eliciting deliberate reasoning in forecasting tasks\[[37](https://arxiv.org/html/2609.38197#bib.bib28),[20](https://arxiv.org/html/2609.38197#bib.bib27)\]\. In contrast, continuous reprogramming does not naturally provide an explicit discrete target space, while textual stringification may fragment numerical structure and incur substantial token overhead\. These advantages make the quantize\-and\-adapt paradigm a natural foundation for building a bimodal financial forecasting model that integrates both numerical dynamics and textual reasoning\.
Committing to the quantize\-and\-adapt paradigm turns the two challenges above into three design requirements\. \(1\) Discretization must preserve scale information\. Financial assets exhibit volatility levels spanning multiple orders of magnitude, and naive quantization either lets high\-volatility assets dominate the codebook or collapses low\-volatility series into near\-degenerate representations\. While normalization alleviates this issue, it also removes scale information that is often predictive\. \(2\) Acquiring numerical forecasting ability must not come at the expense of textual reasoning\. Fine\-tuning an entire LLM on tokenized time series risks shifting model capacity toward numerical patterns and degrading the language understanding needed to interpret news and justify forecast revisions\. \(3\) Fast forecasting and slow reasoning should coexist within a single model\. Most predictions require only a lightweight history\-to\-future mapping, while text\-conditioned deliberation is mainly valuable during regime shifts\. A practical financial foundation model should therefore support both efficient forecasting and optional reasoning without maintaining separate systems or relearning the forecasting process\.
We presentDualCast, a dual\-path framework built on a simple principle: a finite\-capacity backbone should acquire numerical forecasting ability without sacrificing the pretrained linguistic and reasoning competence it later needs to interpret news\. To preserve scale information across heterogeneous assets, the*fast mode*employs a decoupled scale–shape tokenizer that represents normalized price dynamics and market scale as separate discrete tokens, enabling robust forecasting across diverse volatility regimes\. These time\-series tokens are introduced through a fully frozen Qwen3\-8B backbone\[[32](https://arxiv.org/html/2609.38197#bib.bib18)\], with only the newly added token embeddings trained, thereby preserving the model’s native language capabilities by construction\. To support both efficient forecasting and optional deliberation, the*slow mode*adds a toggleable adapter\[[10](https://arxiv.org/html/2609.38197#bib.bib19)\]on the same frozen backbone\. When enabled, it conditions on news and the fast\-mode prediction to produce a revised forecast together with a reasoning trace, optimized via reinforcement learning in return space\. Because both modes share the same backbone and forecasting format, reasoning can be invoked only when additional context is valuable, while retaining the efficiency of the fast forecasting path\.
Figure[1](https://arxiv.org/html/2609.38197#S1.F1)summarizes the framework\. The scale–shape tokenizer converts each return patch into one summary token and three residual shape tokens\. Fast mode learns to forecast this financial\-token sequence with a frozen language\-model backbone, whereas slow mode activates a lightweight adapter that conditions on causally available text and the fast forecast to generate a revised path\.
Figure 1:Overview ofDualCast\. The scale–shape tokenizer separates scale statistics from residual shape patterns and maps both to a hierarchical financial\-token sequence\. Fast mode predicts future financial tokens using the shared frozen backbone\. Slow mode activates a LoRA adapter and uses causally available text together with the fast forecast to produce a news\-conditioned revision\. The adapter is initialized by supervised fine\-tuning and subsequently optimized with the return\-space GRPO objective described in Section[3](https://arxiv.org/html/2609.38197#S3)\.#### Contributions
- 1\.Scale\-aware financial tokenization\.We represent each log\-return patch with one token from a learned summary codebook and three residual shape tokens\. Adaptive frequency\-equalizing residual vector quantization \(AFE\-RVQ\) rebalances assignments within the shape codebooks, reducing codeword usage imbalance while maintaining reconstruction accuracy\.
- 2\.Parameter\-efficient dual\-path adaptation\.We extend a frozen Qwen3\-8B backbone with financial tokens and train their embeddings and output heads for the fast forecasting path\. A toggleable LoRA adapter enables a slow revision path on the same backbone\. Across the two training stages, 48\.9 million parameters, or 0\.60% of the model, are trainable\.
- 3\.News\-conditioned revision with return\-space feedback\.The slow path conditions on the fast forecast and news available at the forecast origin to produce a revised prediction\. We initialize the reviser through supervised fine\-tuning and further optimize it with group relative policy optimization, using return\-space rewards for improvement over the fast forecast and cross\-sectional consistency\.
## 2Related Work
### 2\.1LLMs for Time\-Series Forecasting
Adapting language models to numerical forecasting has produced three methodological families\.*Reprogramming*approaches keep the LLM frozen and learn a thin interface that projects patched series into the model’s embedding space: Time\-LLM\[[11](https://arxiv.org/html/2609.38197#bib.bib1)\]aligns patches with text prototypes, GPT4TS\[[36](https://arxiv.org/html/2609.38197#bib.bib3)\]fine\-tunes only the normalization and projection layers of a frozen GPT\-2, and TEST\[[23](https://arxiv.org/html/2609.38197#bib.bib10)\]contrasts time\-series and text embeddings to activate the LLM\. These methods preserve the backbone but operate on continuous embeddings, which makes them incompatible with the discrete, RL\-optimizable targets we require\.*Stringification*methods such as LLMTime\[[9](https://arxiv.org/html/2609.38197#bib.bib2)\]render numbers as digit strings and forecast by next\-token sampling; they are simple and training\-free but waste sequence length on digits and fracture numerical semantics\.*Quantization*methods discretize the signal into a finite vocabulary and train a language model over it\. Chronos\[[2](https://arxiv.org/html/2609.38197#bib.bib4)\]bins scaled values into a fixed vocabulary, while TOTEM\[[24](https://arxiv.org/html/2609.38197#bib.bib13)\]learns a VQ\-VAE codebook of temporal patterns; both demonstrate that “time series as language” transfers the mature machinery of sequence modelling to forecasting\. In parallel, a line of domain\-specific foundation models—TimesFM\[[6](https://arxiv.org/html/2609.38197#bib.bib5)\], Moirai\[[27](https://arxiv.org/html/2609.38197#bib.bib6)\], MOMENT\[[8](https://arxiv.org/html/2609.38197#bib.bib7)\], Timer\[[18](https://arxiv.org/html/2609.38197#bib.bib8)\], Time\-MoE\[[22](https://arxiv.org/html/2609.38197#bib.bib9)\], and Sundial\[[17](https://arxiv.org/html/2609.38197#bib.bib38)\]—pretrains Transformers from scratch on large numerical corpora\. KTD\-Fin\[[38](https://arxiv.org/html/2609.38197#bib.bib37)\]finds no persistent positive stock\-selection alpha among the LLM trading agents evaluated under leakage\-controlled conditions\. These models are strong unimodal forecasters but discard the pretrained text and reasoning priors that bimodal financial forecasting depends on\.
Our tokenizer sits in the quantization family but departs from it in two ways\. First, following the patch\-wise consensus of PatchTST\[[19](https://arxiv.org/html/2609.38197#bib.bib11)\]and PITS\[[14](https://arxiv.org/html/2609.38197#bib.bib14)\]and per\-instance normalization of RevIN\[[12](https://arxiv.org/html/2609.38197#bib.bib12)\], we explicitly*factor*each patch into a normalized shape and a removed scale, then re\-quantize the scale with a learned codebook rather than discarding it or binning it coarsely\. Second, we adopt residual vector quantization from neural audio coding\[[35](https://arxiv.org/html/2609.38197#bib.bib16)\]with the rotation\-trick straight\-through estimator\[[7](https://arxiv.org/html/2609.38197#bib.bib17)\]and EMA codebooks\[[25](https://arxiv.org/html/2609.38197#bib.bib15)\], and add a frequency\-equalization step so the resulting LLM vocabulary is densely utilized\.
### 2\.2Text–Time\-Series Fusion for Financial Forecasting
Combining textual signals with market data has a long history\. StockNet\[[31](https://arxiv.org/html/2609.38197#bib.bib22)\]pairs tweets with prices for movement prediction and remains a standard benchmark; FinBERT\[[3](https://arxiv.org/html/2609.38197#bib.bib23)\]adapts BERT to financial sentiment; and BloombergGPT\[[28](https://arxiv.org/html/2609.38197#bib.bib24)\]and FinGPT\[[33](https://arxiv.org/html/2609.38197#bib.bib25)\]build finance\-specialized LLMs, though their temporal modelling is weak\. A second thread designs explicit fusion architectures—attention gating, cross\-attention between text and price, and contrastive alignment—typified by market\-guided transformers such as MASTER\[[15](https://arxiv.org/html/2609.38197#bib.bib26)\]\. These systems generally treat text and series as two encoders fused by a bespoke module, and the numerical channel is a regression head rather than a generative sequence\. We instead place price tokens and news tokens in a*single*autoregressive sequence over a shared vocabulary, so fusion is performed by the backbone’s native attention and the model can both read news and*generate*a refined numerical forecast\.
### 2\.3Reasoning and Reinforcement Learning for Time Series
A recent direction reframes forecasting as explicit reasoning rather than a one\-shot mapping\. “Slow\-thinking” studies\[[37](https://arxiv.org/html/2609.38197#bib.bib28),[5](https://arxiv.org/html/2609.38197#bib.bib29)\]show that prompting an LLM to deliberate before predicting can outperform direct prediction, and surveys document a rapid shift toward reasoning\- and agent\-centric time\-series systems\[[13](https://arxiv.org/html/2609.38197#bib.bib30)\]; reinforcement learning over discrete token outputs has emerged as the standard way to elicit such chain\-of\-thought via verifiable rewards\[[20](https://arxiv.org/html/2609.38197#bib.bib27)\]\. We follow this philosophy—reasoning is*elicited by RL with verifiable rewards*\[[21](https://arxiv.org/html/2609.38197#bib.bib20),[34](https://arxiv.org/html/2609.38197#bib.bib21)\]rather than distilled from synthetic traces—but our contribution is orthogonal to how the reasoning trace itself is produced, and lies instead in*where*and*how cheaply*reasoning is added\. First, we keep the backbone permanently frozen and adapt only a compact set of new embeddings, so the model’s native text and reasoning ability is preserved for the refinement stage instead of being diluted by domain fine\-tuning\. Second, reasoning is an*optional*second pass: a single toggleable adapter recovers the fast single\-forward forecaster when disabled and a slow, text\-conditioned reasoner when enabled\. Third, the refinement sequence regenerates the forecast in a numerical context, so the history\-to\-target attention learned during pretraining is reused at refinement time rather than relearned\. Together these make the reasoning path inexpensive to add and inexpensive to skip, which is essential when a fast estimate already suffices and deliberation is warranted only at regime shifts\.
## 3Method
### Overview and Problem Setup
Letp1:C\+H=\{p1,…,pC\+H\}p\_\{1:C\+H\}=\\\{p\_\{1\},\\ldots,p\_\{C\+H\}\\\}denote a univariate positive price series, wherep1:Cp\_\{1:C\}is the historical context andpC\+1:C\+Hp\_\{C\+1:C\+H\}is the forecasting target\. In the multimodal setting, the forecast origin is also associated with a timestamped text stream𝒯≤C\\mathcal\{T\}\_\{\\leq C\}, such as news headlines, filings, or social media posts, whose timestamps are no later than the forecast origin\. The task is to predict the future price path while respecting this no\-leakage constraint\.
DualCastdecomposes the task into a fast numerical path and a slow text\-conditioned refinement path\. Stage 1 learns a time\-series language: prices are converted into discrete tokens and a frozen LLM is trained to autoregressively forecast the next tokens\. Stage 2 keeps the same frozen backbone and Stage 1 time\-series vocabulary, attaches a small adapter, and asks the model to revise the Stage 1 forecast using recent text\. This design separates two capabilities that are both needed in financial forecasting: robust pattern matching over noisy price histories and language\-based reasoning over exogenous events\.
### 3\.1Shape–Scale Tokenization with AFE\-RVQ
A local price segment contains two fundamentally different types of information: its movement pattern and its scale\. The former describes relative temporal dynamics, such as rallies, reversals, or consolidation behaviors, while the latter reflects local drift and volatility\. Quantizing both factors within a single representation space forces the tokenizer to model heterogeneous phenomena simultaneously and reduces token efficiency\. We therefore decompose each patch into a*shape*component and a*scale*component, which are discretized separately and later combined into a unified token stream\.
We represent the price trajectory by log\-returns rather than raw prices,
rt=logpt−logpt−1,r\_\{t\}=\\log p\_\{t\}\-\\log p\_\{t\-1\},\(1\)which removes dependence on the absolute price level and makes the representation invariant to multiplicative rescaling\. The return series is partitioned into non\-overlapping patches of lengthLL, and each patch is independently normalized using reversible instance normalization \(RevIN\)\[[12](https://arxiv.org/html/2609.38197#bib.bib12)\]\. The normalized patch serves as the shape signal, while the removed normalization statistics are retained as patch\-level scale information\.
To bridge continuous financial signals and discrete language modeling, we convert each normalized patch into a small set of discrete shape tokens\. An encoder first maps the normalized patchr~\(w\)\\tilde\{r\}^\{\(w\)\}into a latent representation, which is then discretized using residual vector quantization \(RVQ\)\[[35](https://arxiv.org/html/2609.38197#bib.bib16)\]\. By representing each patch as a composition of codewords drawn from multiple codebooks, RVQ provides a compact yet expressive vocabulary of recurring return patterns\. The resulting code indices form the shape tokens, allowing numerically different but structurally similar price movements to share a common symbolic representation\.
Because every RVQ codeword becomes one trainable embedding row in the extended LLM vocabulary, its assignment frequency directly determines how often that row receives a gradient\. Plain RVQ can concentrate a disproportionate fraction of assignments on a small number of codewords, producing a peaked financial\-token distribution even when most entries are active\. We therefore introduce*Adaptive Frequency\-Equalizing RVQ*\(AFE\-RVQ\), a periodic rebalancing operator applied independently at each residual level during tokenizer training\.
Let𝒞=\{ci\}i=1K\\mathcal\{C\}=\\\{c\_\{i\}\\\}\_\{i=1\}^\{K\}be one codebook and lethih\_\{i\}denote the number of latent vectors assigned tocic\_\{i\}over a pool𝒵\\mathcal\{Z\}of recent batches\. With the uniform expected loadh¯=\|𝒵\|/K\\bar\{h\}=\|\\mathcal\{Z\}\|/K, AFE\-RVQ identifies overloaded and starved entries as
𝒪=\{i:hi\>τhih¯\},𝒟=\{i:hi<τloh¯\}\.\\mathcal\{O\}=\\\{i:h\_\{i\}\>\\tau\_\{\\mathrm\{hi\}\}\\bar\{h\}\\\},\\qquad\\mathcal\{D\}=\\\{i:h\_\{i\}<\\tau\_\{\\mathrm\{lo\}\}\\bar\{h\}\\\}\.\(2\)For an overloaded entryii, the latent vectors assigned tocic\_\{i\}are split with a smallkk\-means problem\. One sub\-center replacescic\_\{i\}, while the remaining sub\-centers are written into available entries in𝒟\\mathcal\{D\}\. The exponential\-moving\-average accumulators of every relocated entry are reset so that it is not pulled back toward its previous location, and a small anti\-symmetric perturbation separates sibling sub\-centers that would otherwise re\-merge\. In our deployed tokenizer, assignment counts are collected fromP=64P=64recent batches and rebalancing is applied every 500 steps withτlo=0\.5\\tau\_\{\\mathrm\{lo\}\}=0\.5andτhi=2\.0\\tau\_\{\\mathrm\{hi\}\}=2\.0\. The operation adds no parameters and is disabled at inference time\. As shown in Section[4\.5](https://arxiv.org/html/2609.38197#S4.SS5), it reduces peak codeword load while preserving reconstruction fidelity; we therefore use it to obtain a better\-balanced vocabulary rather than claim additional representational capacity\.
Since the normalization used to isolate shape patterns removes local drift and volatility, the discarded scale information is represented separately using a discrete summary token\.
The retained statistics\(μw,σw\)\(\\mu\_\{w\},\\sigma\_\{w\}\)are discretized using a learned summary vocabulary constructed from observed\(μw,logσw\)\(\\mu\_\{w\},\\log\\sigma\_\{w\}\)pairs in the training set\. Each summary token corresponds to a frequently occurring local market regime characterized by a particular combination of drift and volatility\. Unlike manually designed discretization grids, the learned vocabulary allocates capacity according to the empirical distribution of market states and therefore uses tokens more efficiently\.
Each patch is ultimately represented by one summary tokenSSfollowed by three residual shape tokens,
\[S,L0,L1,L2\]\.\[S,L\_\{0\},L\_\{1\},L\_\{2\}\]\.\(3\)Placing the summary token first allows subsequent shape tokens to be interpreted under the corresponding local market regime\. The resulting token sequence forms a discrete financial language that can be modeled directly by a standard autoregressive LLM\.
### 3\.2Fast Mode: Financial Language Modeling
Fast mode treats the tokenized price sequence as a discrete financial language and trains a frozen LLM to autoregressively model its dynamics\. Given a sequence of summary and shape tokens produced by AFE\-RVQ, the model predicts future tokens using only historical price information\. This mode serves as the efficient forecasting path ofDualCastand provides the baseline prediction that is later refined by the slow mode\.
We extend the Qwen3\-8B vocabulary with the financial tokens introduced in Section[3\.1](https://arxiv.org/html/2609.38197#S3.SS1), including one summary vocabulary and three shape vocabularies\. Rather than random initialization, the newly added embeddings inherit the geometry of the learned tokenizers, preserving similarity relationships among both market regimes and local price patterns\. Throughout fast\-mode training, the pretrained LLM backbone remains frozen\. The resulting token stream follows a deterministic grammar that uniquely determines the type of the next token from its position, enabling a factorized prediction objective in which the softmax is restricted to the valid token subset associated with the next token type rather than competing over the entire vocabulary\. This design eliminates interference from irrelevant text tokens and better aligns optimization with the structure of the discrete financial language\.
The overall training objective is
ℒfast=λsumℒsumCE\+∑k=02λLkℒLkCE\.\\mathcal\{L\}\_\{\\mathrm\{fast\}\}=\\lambda\_\{\\textsc\{sum\}\}\\,\\mathcal\{L\}\_\{\\textsc\{sum\}\}^\{\\mathrm\{CE\}\}\+\\sum\_\{k=0\}^\{2\}\\lambda\_\{L\_\{k\}\}\\,\\mathcal\{L\}\_\{L\_\{k\}\}^\{\\mathrm\{CE\}\}\.\(4\)where later residual levels receive smaller weights due to their higher entropy and lower predictability\. The first patch of each sequence is used only as context, while all subsequent tokens contribute autoregressive prediction losses\.
Fast mode learns an interface between the discrete financial language and the pretrained LLM\. Only the newly introduced financial\-token embeddings and prediction heads are optimized, while all original Qwen parameters remain fixed\. This design allows the model to acquire forecasting ability without altering its native language and reasoning capabilities, enabling the same backbone to be reused by both the fast and slow modes ofDualCast\.
### 3\.3Slow Mode: Text\-Conditioned Forecast Revision
While the fast mode captures numerical regularities from historical prices, it is intentionally blind to textual information\. The slow mode therefore formulates multimodal forecasting as a*forecast revision*problem, refining a strong fast\-mode baseline using news and recent market context rather than generating a forecast from scratch\. Such a revision\-based formulation mirrors how investors typically reason about markets: they first extrapolate from observed trends and then revise their expectations in response to newly arriving information\.
Formally, given historical pricesp1:Cp\_\{1:C\}, related news𝒯≤C\\mathcal\{T\}\_\{\\leq C\}, and a fast\-mode forecast𝐲fast\\mathbf\{y\}\_\{\\mathrm\{fast\}\}, the slow mode produces a refined forecast𝐲slow\\mathbf\{y\}\_\{\\mathrm\{slow\}\}by conditioning on all three sources of information\.
To realize this revision process, we express each example in the native chat format of the base language model\. The prompt contains three explicitly separated information sources: historical prices, related news, and the fast\-mode forecast, while the model is instructed to generate a refined forecast for the target horizon together with a reasoning trace\.
The slow mode is first initialized through supervised fine\-tuning on automatically constructed revision examples\. For each price–news pair, the base model is prompted to analyze the potential market implications of the news and produce a rationale followed by a refined forecast\. To better align textual reasoning with the newly introduced time\-series vocabulary, we augment a subset of training examples with an explicit trend description derived from recent price movements at the beginning of the reasoning trace\. This lightweight grounding procedure encourages the model to associate time\-series tokens with semantic concepts such as upward, downward, or sideways market behavior\.
Supervision is applied primarily to the refined forecast tokens, while prompt tokens are masked\. The trend\-grounding examples serve only as an auxiliary alignment signal and do not alter the forecasting target\. As a result, the model learns both the revision behavior and the structured output format without requiring manually annotated financial reasoning traces\.
### 3\.4Slow Mode: Utility Optimization via Revision Feedback
While supervised fine\-tuning enables the model to imitate revision behavior from automatically constructed examples, it primarily optimizes token\-level likelihood and does not directly reflect forecasting utility\. In particular, it does not explicitly encourage the model to improve upon the fast\-mode baseline in a decision\-relevant manner\. As a result, the model may produce plausible revisions without consistently improving predictive performance\.
To address this limitation, we optimize the slow mode using group\-relative policy optimization \(GRPO\), directly operating in the forecast \(return\) space\. The training objective is defined as a composite reward:
R=wimpRimp\+wICRIC\+wfmtRfmt,R=w\_\{\\mathrm\{imp\}\}R\_\{\\mathrm\{imp\}\}\+w\_\{\\mathrm\{IC\}\}R\_\{\\mathrm\{IC\}\}\+w\_\{\\mathrm\{fmt\}\}R\_\{\\mathrm\{fmt\}\},\(5\)whereRimpR\_\{\\mathrm\{imp\}\}measures improvement over the fast\-mode baseline,RICR\_\{\\mathrm\{IC\}\}captures cross\-sectional directional consistency, andRfmtR\_\{\\mathrm\{fmt\}\}enforces valid structured output\.
Letρ\(x\)\\rho\(x\)denote the cumulative return implied by forecastxx, and letρ⋆\\rho^\{\\star\}denote the realized cumulative return\. We define a volatility\-normalized error:
e\(x\)=\|ρ\(x\)−ρ⋆\|s,s=2σCH,e\(x\)=\\frac\{\|\\rho\(x\)\-\\rho^\{\\star\}\|\}\{s\},\\qquad s=2\\sigma\_\{C\}\\sqrt\{H\},\(6\)whereσC\\sigma\_\{C\}is estimated from the historical context andHHis the forecast horizon\. The improvement reward is then defined as:
Rimp=clip\(e\(𝐲fast\)−e\(𝐲slow\)max\(e\(𝐲fast\),Δ\),−1,1\)\.R\_\{\\mathrm\{imp\}\}=\\operatorname\{clip\}\\\!\\left\(\\frac\{e\(\\mathbf\{y\}\_\{\\mathrm\{fast\}\}\)\-e\(\\mathbf\{y\}\_\{\\mathrm\{slow\}\}\)\}\{\\max\(e\(\\mathbf\{y\}\_\{\\mathrm\{fast\}\}\),\\Delta\)\},\-1,1\\right\)\.\(7\)
This formulation ensures that the slow mode is rewarded only when it strictly improves upon the fast\-mode forecast, explicitly aligning optimization with the goal of forecast revision rather than standalone prediction\.
WhileRimpR\_\{\\mathrm\{imp\}\}measures absolute improvement, financial decisions are often driven by correct cross\-sectional ranking rather than point\-wise accuracy\. We therefore introduce a rank\-based directional reward computed over a rollout batch\.
For each sampleiiin a batch, we compute volatility\-normalized returns:
zi=ρ\(𝐲i\)2σiHi,zi⋆=ρi⋆2σiHi\.z\_\{i\}=\\frac\{\\rho\(\\mathbf\{y\}\_\{i\}\)\}\{2\\sigma\_\{i\}\\sqrt\{H\_\{i\}\}\},\\qquad z\_\{i\}^\{\\star\}=\\frac\{\\rho\_\{i\}^\{\\star\}\}\{2\\sigma\_\{i\}\\sqrt\{H\_\{i\}\}\}\.\(8\)
We define a normalized rank operatorℛ\(⋅\)\\mathcal\{R\}\(\\cdot\)that maps values within a batch to\[0,1\]\[0,1\]with tie averaging:
z~i=ℛ\(zi\),z~i⋆=ℛ\(zi⋆\)\.\\tilde\{z\}\_\{i\}=\\mathcal\{R\}\(z\_\{i\}\),\\qquad\\tilde\{z\}\_\{i\}^\{\\star\}=\\mathcal\{R\}\(z\_\{i\}^\{\\star\}\)\.\(9\)
We define a cross\-sectional rank\-consistency reward using the normalized within\-batch ranks of the predicted and realized volatility\-normalized returns:
RIC,i=4\(z~i−12\)\(z~i⋆−12\),R\_\{\\mathrm\{IC\},i\}=4\\left\(\\tilde\{z\}\_\{i\}\-\\frac\{1\}\{2\}\\right\)\\left\(\\tilde\{z\}\_\{i\}^\{\\star\}\-\\frac\{1\}\{2\}\\right\),\(10\)Here, \(z~i\\tilde\{z\}\_\{i\}\) and \(z~i⋆\\tilde\{z\}\_\{i\}^\{\\star\}\) denote the predicted and realized ranks, respectively\. The reward is positive when the two ranks fall on the same side of the rank midpoint and negative when they fall on opposite sides\. Aggregated over a rollout batch, it encourages agreement in the relative ordering of returns across assets\. Because it depends only on ranks, this reward does not directly measure the sign or numerical magnitude of an individual predicted return; forecast accuracy is addressed separately by the return\-space improvement reward\.
We include a lightweight format rewardRfmtR\_\{\\mathrm\{fmt\}\}to enforce valid structured generation, including separation between reasoning traces and forecast tokens\. It additionally penalizes language collapse in the reasoning trace, discouraging degenerate repetition and the leakage of time\-series tokens into the textual rationale\. This term primarily acts as a constraint and contributes little gradient once training stabilizes\.
The reinforcement learning objective aligns the slow mode with the goal of forecast revision: improving upon a strong numerical baseline while remaining sensitive to cross\-sectional market structure\. This ensures that optimization is driven by forecasting utility rather than language modeling likelihood\.
## 4Experiments
We evaluateDualCaston financial forecasting across multiple markets, sampling frequencies, and modalities to assess both its forecasting accuracy and its ability to generalize beyond the distributions seen during pretraining\. We first compareDualCastagainst strong time\-series foundation models, multimodal forecasters, and general\-purpose LLMs under a strict zero\-shot protocol\. We then investigate its ability to transfer across heterogeneous assets and unseen temporal resolutions, and finally analyze the contribution of the proposed scale–shape tokenizer and deliberative revision mechanism\.
### 4\.1Experimental Setup
#### Datasets
We evaluate on three benchmark sources covering two financial markets, three temporal resolutions, and both price\-only and multimodal forecasting settings\. Together, they assess generalization across markets, sampling frequencies, and input modalities\.MTBench\-short\[[4](https://arxiv.org/html/2609.38197#bib.bib31)\]provides event\-aligned five\-minute U\.S\. equity windows\.FinMultiTime\[[30](https://arxiv.org/html/2609.38197#bib.bib32)\]provides daily OHLC price series paired with timestamped financial news for U\.S\. equities \(SP500\) and Chinese equities \(HS300\)\. We use only text available at the forecast origin, ensuring strict temporal causality\. Table[1](https://arxiv.org/html/2609.38197#S4.T1)summarizes the evaluation datasets, markets, temporal resolutions, and input modalities\.Time\-MMD \(Energy\)\[[16](https://arxiv.org/html/2609.38197#bib.bib33)\]consists of weekly U\.S\. gasoline retail prices paired with textual market reports, extending evaluation beyond equity markets to commodity price forecasting\.
#### Window construction
All settings use a lookback ofC=96C\{=\}96observations and horizonsH∈\{16,32,64\}H\\\!\\in\\\!\\\{16,32,64\\\}\. We construct the evaluation origins using the maximum horizonHmax=64H\_\{\\max\}\{=\}64and then reuse exactly the same origins for the shorter horizons by truncating the target to its first 16 or 32 observations\. All baselines follow the same construction\. Table[2](https://arxiv.org/html/2609.38197#S4.T2)summarizes the parameter\-efficient training footprint ofDualCast; the backbone remains frozen throughout both training stages\.
Table 1:Evaluation datasets, covering markets, sampling frequencies, and input modalities\. All datasets are evaluated zero\-shot at horizonsH∈\{16,32,64\}H\\in\\\{16,32,64\\\}\.Table 2:Parameter\-efficient training footprint ofDualCast\. The 8\.2B backbone remains frozen; 48\.9M parameters \(0\.60%\) are trained across the two stages\.Table 3:Forecasting MAPE \(%,↓\\downarrow\) across markets and frequencies\. MTBench\-short is five\-minute and evaluated with price only;SP500/HS300are daily, and Energy is weekly\.DualCastuses no dataset\-specific fine\-tuning\. Best per column inbold\.
#### Baselines
We compare against three families: \(1\)*time\-series foundation models*, in their latest released versions, Chronos\-2\[[1](https://arxiv.org/html/2609.38197#bib.bib34)\], TimesFM\-2\.5\[[6](https://arxiv.org/html/2609.38197#bib.bib5)\], Timer\[[18](https://arxiv.org/html/2609.38197#bib.bib8)\], and Time\-MoE\[[22](https://arxiv.org/html/2609.38197#bib.bib9)\], and Moirai\-2\.0\[[27](https://arxiv.org/html/2609.38197#bib.bib6)\]; \(2\)*dedicated multimodal forecasters*that natively consume text, ChatTime\[[26](https://arxiv.org/html/2609.38197#bib.bib35)\]and Aurora\[[29](https://arxiv.org/html/2609.38197#bib.bib36)\]; and \(3\) an LLM\-prompting baseline that feeds the same recent prices and news to a general\-purpose Qwen3\-8B\[[32](https://arxiv.org/html/2609.38197#bib.bib18)\]as text and parses its numeric forecast\. All baselines run in their released zero\-shot configurations on the same windows and horizons\.
#### Metrics
Because financial series span several orders of magnitude in scale and absolute errors are not comparable across assets, we report mean absolute percentage error \(MAPE, %\) at horizonsH∈\{16,32,64\}H\\\!\\in\\\!\\\{16,32,64\\\}\. ForDualCastwe report the fast forecaster \(Fast\), the SFT reviser \(SFT\), and the RL\-tuned slow reviser \(Slow\);Slowis the default\.
### 4\.2Main Results
DualCast \(Slow\) achieves the lowest MAPE in 8 of 12 dataset–horizon settings as Table[3](https://arxiv.org/html/2609.38197#S4.T3)shows, including H=64 on all four evaluation subsets\. On five\-minute MTBench, whose resolution is entirely unseen during pretraining,Slowobtains the best results across all horizons\. This demonstrates that the scale–shape tokenizer is able to extract market patterns that are agnostic to sampling frequency\. Despite keeping the entire backbone frozen, our method achieves competitive or superior performance compared with cutting\-edge time\-series foundation models, including Chronos\-2, TimesFM\-2\.5, and Moirai\-2\.0, especially at long forecasting horizons critical for real\-world financial analysis\. Consistent with baseline protocols, DualCast follows the same price\-only pretraining paradigm and only updates merely0\.6%0\.6\\%of the total parameters\.DualCastclearly outperforms ChatTime and Aurora; feeding long news text into ChatTime even*inflates*its error on Energy \(≥15%\\geq\\\!15\\%\), as serializing text disrupts its numeric scale anchoring, whereas our method keeps the model anchored to the price dynamics\. Figures[4](https://arxiv.org/html/2609.38197#S4.F4)and[2](https://arxiv.org/html/2609.38197#S4.F2)visualize stable forecasts across markets and the Fast–SFT–Slow progression on identical windows, respectively\.
### 4\.3Refinement: Fast vs\. Slow
DualCastprovides three forecasting modes based on the same frozen backbone: the fast numerical forecaster \(Fast\), a supervised fine\-tuned reviser \(SFT\), and a reviser further optimized by reinforcement learning \(Slow\)\. Table 4 compares these modes on identical evaluation windows\. MAPE decreases or remains unchanged from Fast to SFT to Slow in every dataset–horizon setting; SFT and Slow tie on MTBench\-short at H=16 and H=64\. At H=64, MAPE decreases from 18\.06% to 12\.43% and then 9\.19% on HS300, and from 15\.26% to 14\.03% and then 9\.80% on SP500\. These comparisons show that revision improves the fast forecast across the evaluated settings, with larger gains at longer horizons\. Figure[2](https://arxiv.org/html/2609.38197#S4.F2)illustrates this progression on representative windows\.
Table 4:Ablation of the threeDualCastmodes on identical windows\. MAPE \(%,↓\\downarrow\); best per row group inbold\. MTBench\-short is five\-minute and price\-only;SP500/HS300are daily, and Energy is weekly\.Table 5:News ablation: point error \(MAPE, %,↓\\downarrow\) and directional accuracy \(↑\\uparrow\) of theSlowreviser with real news vs\. with news removed, on identical date\-aligned windows\. Better of the two conditions per column inbold\.
\(a\) MTBench \(5\-minute\)

\(b\) SPIP,SP500\(daily\)
Figure 2:Representative Fast–SFT–Slow forecasts on identical windows\. Gray denotes the observed history, black the realized future, blue the fast Stage\-1 forecast, green the SFT reviser, and orange the RL\-tuned slow reviser\. The MTBench example shows a small refinement when the fast estimate is already reasonable; the SP500 example shows the two refinement stages correcting a larger fast\-mode deviation\. These cases are illustrative; aggregate improvements are reported in Table[4](https://arxiv.org/html/2609.38197#S4.T4)\.Table[3](https://arxiv.org/html/2609.38197#S4.T3)also shows that directly prompting Qwen3\-8B with recent prices and news as text yields substantially higher MAPE thanDualCast\(Slow\) on the daily equity subsets\. On SP500, Qwen3\-8B records MAPE of 25\.66–31\.15% across the three horizons, compared with 4\.81–9\.80% forDualCast\(Slow\)\. On HS300, the corresponding ranges are 15\.64–22\.51% and 7\.08–9\.19%\. Thus, direct text prompting is a weaker baseline in these evaluations\.
Table 6:Cost of deliberation: per\-forecast latency \(median\) and number of generated tokens onSP500\(single H800, batch 1, greedy decoding\)\. The fast path decodes only the forecast tokens; the slow path additionally generates a reasoning trace and regenerates the forecast\.Table[6](https://arxiv.org/html/2609.38197#S4.T6)quantifies the price of reasoning onSP500\(single H800, batch 1, greedy decoding\)\. The fast path emits only the1616forecast tokens and completes in0\.370\.37s per window, whereas the slow path autoregressively generates a∼167\\sim\\\!167\-token reasoning trace before regenerating the forecast, costing5\.855\.85s—about16×16\\timesmore\. Because the adapter is toggleable and both modes share the frozen backbone, disabling deliberation recovers the fast forecaster exactly, with no separate model to maintain\. The reasoning pass is therefore something to spend only when text is likely to matter, while the large majority of forecasts are served by the cheap fast path\.
### 4\.4Does Text Help? A News Ablation
The revision\-based slow mode is designed to fold textual information into a numerical forecast, but MAPE alone cannot reveal whether the text is actually used: a lower error may simply reflect the RL stage sharpening the numerical baseline\. To isolate the contribution of text, we re\-run theSlowreviser on strictly date\-aligned windows under two conditions–with the real news at the forecast origin, and with the news removed–holding everything else fixed\. To assess whether news affects forecast direction as well as point accuracy, we also report directional accuracy\.
Table[5](https://arxiv.org/html/2609.38197#S4.T5)shows that the text channel is useful, and we evaluate performance via two complementary metrics: point\-prediction MAPE and directional accuracy\. Forpoint error, real news lowers MAPE in eight of the nine dataset–horizon comparisons\. The gains are largest and most consistent onHS300\(−0\.16/−0\.23/−0\.27\-0\.16/\-0\.23/\-0\.27\), remain positive onSP500\(−0\.04/−0\.06/−0\.08\-0\.04/\-0\.06/\-0\.08\), and appear at the longer Energy horizons \(−0\.05/−0\.21\-0\.05/\-0\.21forH=32/64H\{=\}32/64\); the only exception is a negligible\+0\.01\+0\.01increase on Energy atH=16H\{=\}16\. Fordirection, Energy improves at every horizon \(\+0\.06/\+0\.12/\+0\.06\+0\.06/\+0\.12/\+0\.06\), andSP500shows smaller consistent gains \(\+0\.02/\+0\.03/\+0\.02\+0\.02/\+0\.03/\+0\.02\)\.HS300is mixed: news improves point accuracy at all horizons but directional accuracy only atH=64H\{=\}64\. These ablations indicate that real news improves point accuracy in most tested settings, while its effect varies across datasets, horizons, and metrics\. They also show that news conditioning is not uniformly beneficial for every metric and horizon, motivating selective activation of the slow path rather than unconditional deliberation\.

\(a\) HPE,H=32H\{=\}32

\(b\) BMRN,H=64H\{=\}64
Figure 3:Additional SP500 revision examples\. The slow path \(orange\) conditions on the fast forecast \(blue\) and causally available text\. In both examples it removes a substantial portion of the fast path’s scale or local\-trajectory error, although it does not reproduce every short\-term fluctuation in the realized path \(black\)\. These examples complement the paired aggregate news ablation in Table[5](https://arxiv.org/html/2609.38197#S4.T5)\.
\(a\) MTBench \(5\-minute\)

\(b\)SP500\(daily\)

\(c\)HS300\(daily\)
Figure 4:Representative slow\-path forecasts \(orange dashed\) against the realized future \(black\) following the observed history \(gray\); the dotted vertical line marks the forecast origin\. The three examples cover five\-minute MTBench and daily SP500 and HS300 data\.
### 4\.5Tokenizer Analysis
Two design choices underlie the scale–shape tokenizer\. First, each RVQ codeword maps to one embedding row in the extended vocabulary, so highly imbalanced assignments produce correspondingly imbalanced training of the added vocabulary\. Our Adaptive Frequency\-Equalizing RVQ \(AFE\-RVQ\) periodically splits overloaded latent clusters and relocates their sub\-centers into starved slots\.
Under the controlled3×10243\\times 1024setting, AFE\-RVQ preserves reconstruction performance \(→0\.07810\.0783\\\!\\to\\\!0\.0781\), while achieving a modest improvement in finite\-budget active\-code fraction \(→0\.8540\.839\\\!\\to\\\!0\.854\)\.
The primary benefit of AFE\-RVQ lies in mitigating skewed codeword usage\. Peak load — the maximum assignment count relative to uniform expectation — drops drastically across all residual layers\. ForL0L\_\{0\}, the worst\-case codeword load falls from7\.9×7\.9\\timesto3\.9×3\.9\\times; forL1L\_\{1\}from5\.0×5\.0\\timesto2\.6×2\.6\\times; forL2L\_\{2\}from3\.4×3\.4\\timesto2\.1×2\.1\\times\. Extreme peak load means a small subset of codewords dominate training updates, harming embedding learning for under\-utilized entries\.
### 4\.6Qualitative Analysis
Figure[4](https://arxiv.org/html/2609.38197#S4.F4)shows stable slow\-path forecasts across five\-minute MTBench and daily U\.S\. and Chinese equities\. Figure[2](https://arxiv.org/html/2609.38197#S4.F2)separately visualizes the progression from Fast to SFT and RL\. Figure[3](https://arxiv.org/html/2609.38197#S4.F3)adds two SP500 windows in which slow revision reduces a substantial fast\-path error\. Figure[5](https://arxiv.org/html/2609.38197#S4.F5)presents two directional correction cases with accompanying news and model\-generated chain\-of\-thought rationales\. These examples illustrate individual revisions; aggregate effects on forecast accuracy and direction are reported in Tables[4](https://arxiv.org/html/2609.38197#S4.T4)and[5](https://arxiv.org/html/2609.38197#S4.T5)\.
COPX \(Copper Miners ETF\) • U\.S\. market • bullish \(H=64H\{=\}64\)Fast−3\.3%\-3\.3\\%⟶\\longrightarrowSlow\+5\.9%\+5\.9\\%⟶\\longrightarrowReal\+5\.2%\+5\.2\\%News\.An approximately $131\.4M weekly inflow into the Copper Miners ETF \(COPX\)—a4\.9%4\.9\\%increase in outstanding units—while several underlying miners are slightly lower on the day\.Reasoning\. \(1\) The large ETF inflow signals rising institutional interest in copper miners, despite mixed short\-term moves in individual constituents\. \(2\) Net inflows indicate positive sentiment; if demand fundamentals hold, gradual upward pressure should continue\. \(3\) I expect COPX to rise moderately over the next 64 trading days, roughly\+5%\+5\\%to\+10%\+10\\%\.Outcome\.Fastwrongly predicts a slight decline; the news\-grounded revision flips the forecast upward and nearly matches the realized\+5\.2%\+5\.2\\%\.
Kuang\-Chi Technologies \(002625\.SZ\) • Chinese market • bearish \(H=64H\{=\}64\)Fast\+27\.6%\+27\.6\\%⟶\\longrightarrowSlow−3\.0%\-3\.0\\%⟶\\longrightarrowReal−6\.4%\-6\.4\\%News\.The controlling shareholder and its concert parties announce a plan to reduce their stake by up to2%2\\%of shares outstanding\.Reasoning\. \(1\) A controlling\-shareholder sell\-down raises concerns about valuation and outlook; with the price already in decline, it can intensify selling pressure\. \(2\) Share reductions are generally read as a negative signal, driving near\-term outflows and weaker sentiment\. \(3\) Over the next 64 trading days I expect continued downward pressure, about−5%\-5\\%to−7%\-7\\%\.Outcome\.Fastconfidently predicts a large rise; reading the insider sell\-down as bearish, the revision corrects the sign and tracks the realized−6\.4%\-6\.4\\%\.
Figure 5:Interpretable revisions: two representative cases where theSlowreviser reads the news, reasons explicitly, and corrects the direction \(and sign\) of theFastforecast while approximately matching the realized magnitude\. Such directional corrections are illustrative rather than universal; consistent with the news ablation \(Table[5](https://arxiv.org/html/2609.38197#S4.T5)\), text helps where it carries information beyond the price path\.
## 5Limitations
This work has two main limitations\. First, the beneficial effect of textual news is heterogeneous across different market domains and forecasting settings\. While news information effectively improves forecasting performance in most scenarios, the magnitude of its contribution varies across asset types and prediction horizons, leading to marginal or mixed improvements under certain conditions\.
Second, the high\-performance slow deliberation mode suffers from limited inference speed and increased computational overhead\. Compared with the efficient fast forecasting path, the slow path requires additional reasoning token generation, which substantially increases inference latency\. This limits the deployment efficiency of the slow mode in time\-sensitive real\-world scenarios\. Optimizing the reasoning efficiency of the deliberative forecasting pipeline remains a promising direction for future work\.
## 6Conclusion
We introducedDualCast, a dual\-path language\-model framework for financial time\-series forecasting\. The framework represents each return patch via disentangled summary and residual shape tokens, and leverages AFE\-RVQ to balance codeword usage while maintaining high reconstruction fidelity\. Based on a frozen Qwen3\-8B backbone, DualCast supports two complementary inference modes: a fast financial token decoder for efficient numerical forecasting, and a LoRA\-enhanced slow revision module that incorporates causal news information and refines fast predictions\. The slow path is initialized via supervised fine\-tuning and further optimized with return\-space GRPO to improve revision quality\. Extensive experiments across minute\-level U\.S\. equities, daily U\.S\. and Chinese equities, and weekly energy markets demonstrate that the slow revision mode achieves superior performance, especially over long forecasting horizons\. Ablation studies validate the effectiveness of news\-guided revision and reinforcement learning optimization\. News inputs consistently reduce point prediction errors across most dataset\-horizon settings and deliver prominent directional improvements on energy assets\.
Empirical results verify our core design paradigm: multimodal financial forecasting can be effectively formulated as an optional news\-based revision upon strong numerical base predictions\. Future work will explore uncertainty\-aware evaluation, generalized market adaptation, and learnable routing strategies to further balance prediction accuracy and inference efficiency\.
## 7Acknowledgement
This work was supported by the Science and Technology Program of Qingdao Municipality under Grant 25\-1\-1\-gjgg\-32\-gx\.
## References
- \[1\]A\. F\. Ansari, O\. Shchur, J\. Küken, A\. Auer, B\. Han, P\. Mercado, S\. S\. Rangapuram, H\. Shen, L\. Stella, X\. Zhang,et al\.\(2025\)Chronos\-2: from univariate to universal forecasting\.arXiv preprint arXiv:2510\.15821\.Cited by:[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[2\]A\. F\. Ansari, L\. Stella, C\. Turkmen, X\. Zhang, P\. Mercado, H\. Shen, O\. Shchur, S\. S\. Rangapuram, S\. P\. Arango, S\. Kapoor,et al\.\(2024\)Chronos: learning the language of time series\.arXiv preprint arXiv:2403\.07815\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[3\]D\. Araci\(2019\)Finbert: financial sentiment analysis with pre\-trained language models\.arXiv preprint arXiv:1908\.10063\.Cited by:[§2\.2](https://arxiv.org/html/2609.38197#S2.SS2.p1.1)\.
- \[4\]J\. Chen, A\. Feng, Z\. Zhao, J\. Garza, G\. Nurbek, C\. Qin, A\. Maatouk, L\. Tassiulas, Y\. Gao, and R\. Ying\(2025\)Mtbench: a multimodal time series benchmark for temporal reasoning and question answering\.arXiv preprint arXiv:2503\.16858\.Cited by:[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px1.p1.1.3),[Table 1](https://arxiv.org/html/2609.38197#S4.T1.2.2.1.1)\.
- \[5\]M\. Cheng, J\. Wang, D\. Wang, X\. Tao, Q\. Liu, and E\. Chen\(2026\)Can slow\-thinking llms reason over time? empirical studies in time series forecasting\.InProceedings of the Nineteenth ACM International Conference on Web Search and Data Mining,pp\. 99–110\.Cited by:[§2\.3](https://arxiv.org/html/2609.38197#S2.SS3.p1.1)\.
- \[6\]A\. Das, W\. Kong, R\. Sen, and Y\. Zhou\(2023\)A decoder\-only foundation model for time\-series forecasting\.arXiv preprint arXiv:2310\.10688\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[7\]C\. Fifty, R\. Junkins, D\. Duan, A\. Iyengar, J\. Liu, E\. Amid, S\. Thrun, and C\. Ré\(2025\)Restructuring vector quantization with the rotation trick\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 19153–19188\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p2.1)\.
- \[8\]M\. Goswami, K\. Szafer, A\. Choudhry, Y\. Cai, S\. Li, and A\. Dubrawski\(2024\)Moment: a family of open time\-series foundation models\.arXiv preprint arXiv:2402\.03885\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[9\]N\. Gruver, M\. Finzi, S\. Qiu, and A\. G\. Wilson\(2023\)Large language models are zero\-shot time series forecasters\.Advances in neural information processing systems36,pp\. 19622–19635\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[10\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2021\)Lora: low\-rank adaptation of large language models\.arXiv preprint arXiv:2106\.09685\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p6.1)\.
- \[11\]M\. Jin, S\. Wang, L\. Ma, Z\. Chu, J\. Zhang, X\. Shi, P\. Chen, Y\. Liang, Y\. Li, S\. Pan,et al\.\(2024\)Time\-llm: time series forecasting by reprogramming large language models\.InInternational conference on learning representations,Vol\.2024,pp\. 23857–23880\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[12\]T\. Kim, J\. Kim, Y\. Tae, C\. Park, J\. Choi, and J\. Choo\(2021\)Reversible instance normalization for accurate time\-series forecasting against distribution shift\.InInternational conference on learning representations,Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p2.1),[§3\.1](https://arxiv.org/html/2609.38197#S3.SS1.p2.2)\.
- \[13\]Y\. Kong, Y\. Yang, S\. Wang, C\. Liu, Y\. Liang, M\. Jin, S\. Zohren, D\. Pei, Y\. Liu, and Q\. Wen\(2025\)Achieving time series reasoning requires rethinking model design, tasks formulation, and evaluation\.arXiv preprint arXiv:2502\.01477\.Cited by:[§2\.3](https://arxiv.org/html/2609.38197#S2.SS3.p1.1)\.
- \[14\]S\. Lee, T\. Park, and K\. Lee\(2024\)Learning to embed time series patches independently\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 27599–27620\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p2.1)\.
- \[15\]T\. Li, Z\. Liu, Y\. Shen, X\. Wang, H\. Chen, and S\. Huang\(2024\)Master: market\-guided stock transformer for stock price forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 162–170\.Cited by:[§2\.2](https://arxiv.org/html/2609.38197#S2.SS2.p1.1)\.
- \[16\]H\. Liu, S\. Xu, Z\. Zhao, L\. Kong, H\. Kamarthi, A\. B\. Sasanur, M\. Sharma, J\. Cui, Q\. Wen, C\. Zhang,et al\.\(2024\)Time\-mmd: multi\-domain multimodal dataset for time series analysis\.Advances in Neural Information Processing Systems37,pp\. 77888–77933\.Cited by:[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px1.p1.1.7),[Table 1](https://arxiv.org/html/2609.38197#S4.T1.2.5.1.1)\.
- \[17\]Y\. Liu, G\. Qin, Z\. Shi, Z\. Chen, C\. Yang, X\. Huang, J\. Wang, and M\. Long\(2025\)Sundial: a family of highly capable time series foundation models\.arXiv preprint arXiv:2502\.00816\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[18\]Y\. Liu, H\. Zhang, C\. Li, X\. Huang, J\. Wang, and M\. Long\(2024\)Timer: generative pre\-trained transformers are large time series models\.arXiv preprint arXiv:2402\.02368\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[19\]Y\. Nie, N\. H\. Nguyen, P\. Sinthong, and J\. Kalagnanam\(2022\)A time series is worth 64 words: long\-term forecasting with transformers\.arXiv preprint arXiv:2211\.14730\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p2.1)\.
- \[20\]F\. Parker, N\. Chan, C\. Zhang, and K\. Ghobadi\(2025\)Eliciting chain\-of\-thought reasoning for time series analysis using reinforcement learning\.arXiv preprint arXiv:2510\.01116\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.38197#S2.SS3.p1.1)\.
- \[21\]Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§2\.3](https://arxiv.org/html/2609.38197#S2.SS3.p1.1)\.
- \[22\]X\. Shi, S\. Wang, Y\. Nie, D\. Li, Z\. Ye, Q\. Wen, and M\. Jin\(2024\)Time\-MoE: billion\-scale time series foundation models with mixture of experts\.arXiv preprint arXiv:2409\.16040\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[23\]C\. Sun, H\. Li, Y\. Li, and S\. Hong\(2024\)TEST: text prototype aligned embedding to activate llm’s ability for time series\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 37854–37881\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[24\]S\. Talukder, Y\. Yue, and G\. Gkioxari\(2024\)Totem: tokenized time series embeddings for general time series analysis\.arXiv preprint arXiv:2402\.16412\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[25\]A\. Van Den Oord O\. Vinyalset al\.\(2017\)Neural discrete representation learning\.Advances in neural information processing systems30\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p2.1)\.
- \[26\]C\. Wang, Q\. Qi, J\. Wang, H\. Sun, Z\. Zhuang, J\. Wu, L\. Zhang, and J\. Liao\(2025\)Chattime: a unified multimodal time series foundation model bridging numerical and textual data\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 12694–12702\.Cited by:[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[27\]G\. Woo, C\. Liu, A\. Kumar, C\. Xiong, S\. Savarese, and D\. Sahoo\(2024\)Unified training of universal time series forecasting transformers\.arXiv preprint arXiv:2402\.02592\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[28\]S\. Wu, O\. Irsoy, S\. Lu, V\. Dabravolski, M\. Dredze, S\. Gehrmann, P\. Kambadur, D\. Rosenberg, and G\. Mann\(2023\)Bloomberggpt: a large language model for finance\.arXiv preprint arXiv:2303\.17564\.Cited by:[§2\.2](https://arxiv.org/html/2609.38197#S2.SS2.p1.1)\.
- \[29\]X\. Wu, J\. Jin, W\. Qiu, P\. Chen, Y\. Shu, B\. Yang, and G\. Guo\(2026\)Aurora: towards universal generative multimodal time series forecasting\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 103239–103277\.Cited by:[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[30\]W\. Xu, D\. Xiang, Y\. Liu, X\. Wang, Y\. Ma, L\. Zhang, S\. Hu, C\. Xu, and J\. Zhang\(2025\)Finmultitime: a four\-modal bilingual dataset for financial time\-series analysis\.arXiv preprint arXiv:2506\.05019\.Cited by:[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px1.p1.1.5),[Table 1](https://arxiv.org/html/2609.38197#S4.T1.2.3.1.1.1),[Table 1](https://arxiv.org/html/2609.38197#S4.T1.2.4.1.1.1)\.
- \[31\]Y\. Xu and S\. B\. Cohen\(2018\)Stock movement prediction from tweets and historical prices\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1970–1979\.Cited by:[§2\.2](https://arxiv.org/html/2609.38197#S2.SS2.p1.1)\.
- \[32\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p6.1),[§4\.1](https://arxiv.org/html/2609.38197#S4.SS1.SSS0.Px3.p1.1.1)\.
- \[33\]H\. Yang, X\. Liu, and C\. D\. Wang\(2023\)Fingpt: open\-source financial large language models\.arXiv preprint arXiv:2306\.06031\.Cited by:[§2\.2](https://arxiv.org/html/2609.38197#S2.SS2.p1.1)\.
- \[34\]Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.\(2026\)Dapo: an open\-source llm reinforcement learning system at scale\.Advances in Neural Information Processing Systems38,pp\. 113222–113244\.Cited by:[§2\.3](https://arxiv.org/html/2609.38197#S2.SS3.p1.1)\.
- \[35\]N\. Zeghidour, A\. Luebs, A\. Omran, J\. Skoglund, and M\. Tagliasacchi\(2021\)Soundstream: an end\-to\-end neural audio codec\.IEEE/ACM Transactions on Audio, Speech, and Language Processing30,pp\. 495–507\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p2.1),[§3\.1](https://arxiv.org/html/2609.38197#S3.SS1.p3.1)\.
- \[36\]T\. Zhou, P\. Niu, L\. Sun, R\. Jin,et al\.\(2023\)One fits all: power general time series analysis by pretrained lm\.Advances in neural information processing systems36,pp\. 43322–43355\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.
- \[37\]Y\. Zhou, Y\. Luo, M\. Cheng, Q\. Liu, J\. Wang, D\. Wang, and E\. Chen\(2025\)Time series forecasting as reasoning: a slow\-thinking approach with reinforced llms\.arXiv preprint arXiv:2506\.10630\.Cited by:[§1](https://arxiv.org/html/2609.38197#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.38197#S2.SS3.p1.1)\.
- \[38\]T\. Zhu, W\. Zhao, R\. Sun, B\. Luan, J\. Lu, S\. Wang, J\. Li, D\. Jiang, Y\. He, and Z\. Bai\(2026\)From knowing to doing: a memory\-controlled benchmark for llm trading agents on stock markets\.arXiv preprint arXiv:2605\.28359\.Cited by:[§2\.1](https://arxiv.org/html/2609.38197#S2.SS1.p1.1)\.Similar Articles
CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series Forecasting
CryptoL is a unified framework for cryptocurrency multivariate time series forecasting that mitigates scale dominance and enforces physical constraints to enhance prediction accuracy and financial validity.
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
CastFSR is a Fast–Slow–Reflect agentic reasoning framework that leverages LLMs for context-aware time series forecasting, combining fast lightweight forecasters, slow deliberative reasoning, and reflective evaluation to improve forecasting accuracy and consistency.
TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training Split
The paper presents TW3Cast, a time-series forecasting system that uses a frozen router of lightly fine-tuned foundation models to achieve top performance on the GIFT-Eval benchmark without agents or language models.
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
This paper introduces the Bicameral Model, which couples two frozen language models through a trainable neural interface on their intermediate hidden states to enable continuous, concurrent coordination without serialized text exchanges. The approach demonstrates significant improvements in arithmetic and logic tasks by allowing an auxiliary model to operate tools in parallel with a primary model.
A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis
This paper presents a unified multi-modal framework integrating reinforcement learning, high-frequency trading, game-theoretic approaches, and cross-modal sentiment analysis for intelligent financial systems, claiming significant improvements over single-domain systems.