PALM: Point-in-Time Adaptation for Financial Language Models

arXiv cs.LG Papers

Summary

This paper proposes PALM, a low-rank adapter method to adapt financial language models to specific time points without full retraining, reducing look-ahead bias and computational costs in financial backtests.

arXiv:2609.30316v1 Announce Type: new Abstract: Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time (PIT) language models are pretrained on chronologically filtered corpora and released as one checkpoint per calendar year, each with a documented cutoff. However, each additional year costs a full pretraining run, and whether that run is necessary has never been tested. In this paper, we show that the annual pretraining run is not necessary. We instead compare each checkpoint against the newer one that replaced it, and find that the newer checkpoint scores no better on the same evaluation window. Motivated by this observation, we propose PALM (Point-in-time Adaptation for financial Language Models), a simple yet effective alternative to annual pretraining that fits a low-rank adapter on text published before the decision date without modifying any pretrained weight. We further find that a small adapter is enough to add a new period to the knowledge an old checkpoint already encodes, and that this outperforms continued pretraining. We validate PALM on a decade of financial news and on various families of PIT models, whose cutoffs span two decades and whose sizes range from 1.3 to 4.2B. Code is available at: https://github.com/seunghan96/palm.
Original Article
View Cached Full Text

Cached at: 09/29/26, 09:34 AM

# PALM: Point-in-Time Adaptation for Financial Language Models
Source: [https://arxiv.org/html/2609.30316](https://arxiv.org/html/2609.30316)
CCS:Computing methodologies Natural language processingCCS:Applied computing Economics,Jun SeoAffiliation:LG AI Research,Seoul,Republic of Korea,Jaehoon LeeAffiliation:LG AI Research,Seoul,Republic of Korea,Junhyeok KangAffiliation:LG AI Research,Seoul,Republic of Korea,Sangjun HanAffiliation:LG AI Research,Seoul,Republic of Korea,Sungdong YooAffiliation:LG AI Research,Seoul,Republic of Korea,Minjae KimAffiliation:LG AI Research,Seoul,Republic of Korea,Tae Yoon LimAffiliation:LG AI Research,Seoul,Republic of Korea,Dongwan KangAffiliation:LG AI Research,Seoul,Republic of Korea,Hwanil ChoiAffiliation:LG AI Research,Seoul,Republic of Korea,Soonyoung LeeAffiliation:LG AI Research,Seoul,Republic of KoreaandWonbin AhnAffiliation:LG AI Research,Seoul,Republic of Korea

###### Abstract\.

Language models used in financial backtests suffer fromlook\-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict\. To handle this issue,point\-in\-time \(PIT\) language modelsare pretrained on chronologically filtered corpora and released as one checkpoint per calendar year, each with a documented cutoff\. However, each additional year costs a full pretraining run, and whether that run is necessary has never been tested\. In this paper, we show thatthe annual pretraining run is not necessary\. We instead compare each checkpoint against the newer one that replaced it, and find that the newer checkpoint scores no better on the same evaluation window\. Motivated by this observation, we propose PALM \(Point\-in\-timeAdaptation for financialLanguageModels\), a simple yet effective alternative to annual pretraining that fits a low\-rank adapter on text published before the decision date without modifying any pretrained weight\. We further find that a small adapter is enough to add a new period to the knowledge an old checkpoint already encodes, and that this outperforms continued pretraining\. We validate PALM on a decade of financial news and on various families of PIT models, whose cutoffs span two decades and whose sizes range from 1\.3 to 4\.2B\. Code is available at:[https://github\.com/seunghan96/palm](https://github.com/seunghan96/palm)\.

###### Keywords:

Large language models, Point\-in\-time modeling, Look\-ahead bias, Low\-rank adaptation, Financial news

## 1\.Introduction

Large language models \(LLMs\) have become a standard tool for extracting information from unstructured text, and finance is among the domains that have adopted them fastest\([Wu et al\., 2023](https://arxiv.org/html/2609.30316#bib.bib26);[Li et al\., 2023](https://arxiv.org/html/2609.30316#bib.bib18)\)\. Financial text arrives continuously as news articles, earnings calls, and regulatory filings, and prior work\([Ke et al\., 2019](https://arxiv.org/html/2609.30316#bib.bib5);[Dong et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib12);[Wu et al\., 2025](https://arxiv.org/html/2609.30316#bib.bib21)\)shows that this text carries information about future returns\. A common setup is to score each article with an LLM, rank assets on that score to form a cross\-sectional portfolio, and evaluate the portfolio on historical data\([Iacovides et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib19);[Vamvourellis and Mehta, 2025](https://arxiv.org/html/2609.30316#bib.bib20)\)\.

Figure 1\.Naive LLM vs\. PIT LLM vs\. PALM \(Ours\)\. A naive LLM has read the evaluation period and is ineligible\. A PIT suite avoids this with one pretraining run per year, whereas PALM needs one model and one small adapter per year\.A timeline from 2017 to 2029 with three rows\. The naive model reads text past the decision date at 2026 and its overrun is hatched as look\-ahead bias\. The point\-in\-time suite is four bars, one per calendar year, each stopping before the decision date\. Our row is one bar of the same pretrained text followed by four small adapter blocks, one per year\.However, a backtest is trustworthy only if the model has not already read the period it is tested on\. As shown in Figure[1](https://arxiv.org/html/2609.30316#S1.F1), a model trained through 2026 has observed how the market responded to news published in 2022, and its score for that news recalls the outcome rather than forecasting it\. This is known aslook\-ahead bias, and it inflates backtested performance\([Glasserman and Lin, 2023](https://arxiv.org/html/2609.30316#bib.bib15);[Lopez\-Lira and Tang, 2026](https://arxiv.org/html/2609.30316#bib.bib16);[Li et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib30)\)\. The bias grows with every release, since a newer model has read more of whatever period a study tests, and its size cannot be measured, as a released model rarely documents either its corpus or its cutoff\([Cheng et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib6)\)\.

To eliminate this bias by construction, recent work has introduced point\-in\-time \(PIT\) language models\([Sarkar and Vafa, 2024](https://arxiv.org/html/2609.30316#bib.bib23)\)\. Each model is pretrained only on text published before its own cutoff, and one such checkpoint, orvintage, is released per calendar year with that cutoff documented\([He et al\., 2025a](https://arxiv.org/html/2609.30316#bib.bib1);[He et al\., 2025b](https://arxiv.org/html/2609.30316#bib.bib2);[Kelly et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib3);[Yan et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib4)\)\. A backtest run at datettuses a vintage whose cutoff precedestt, and that vintage has therefore never seen the period it is tested on\.

Although this removes look\-ahead, the guarantee is expensive to maintain, since each additional year of coverage requires a full pretraining run\. Each such family of vintages, orsuite, now spans two decades of yearly pretraining runs\. These runs are justified by a single premise, that a checkpoint decays with itsstaleness, the gap between its cutoff and the date of the backtest\. However, this premise has not been tested, since each vintage is evaluated on the period that follows its own cutoff, which moves the checkpoint and the period together\. As a result,a vintage is never compared against the vintage that replaced it\.

In this paper, we find thatstaleness does not degrade downstream performance in any suite\(i\.e\., the newer vintage is not the better one\)\. To this end, we proposePALM\(Point\-in\-timeAdaptation for financialLanguageModels\), a simple yet effective alternative to pretraining a new checkpoint, which fits a low\-rank adapter on the text the current one has not yet read\. All it needs is text and the language modeling objective, and it leaves every pretrained weight frozen\. The adapter reads no text from the evaluation window, and an adapted checkpoint therefore remains eligible whenever the original checkpoint is\. The main contributions are:

- •To the best of our knowledge, we are the first to evaluate the necessity of the pretraining run behind a PIT suite, and we find that staleness does not degrade downstream performance\.
- •We propose PALM, a simple yet effective plug\-in method that avoids look\-ahead bias by training a low\-rank adapter on eligible text instead of pretraining a new checkpoint\.
- •We conduct extensive experiments across various families of PIT models, and further show that a small adapter outperforms continued pretraining of the same checkpoint\.

## 2\.Related Works

Table 1\.Positioning of PALM\. Point\-in\-Time means that no text published after the decision date was read\.LLMs for financial text\.Textual sources such as news, earnings calls, and filings carry information about future returns, and extracting it has moved from lexicon counting to supervised models\([Ke et al\., 2019](https://arxiv.org/html/2609.30316#bib.bib5)\)and now to LLMs\([Lopez\-Lira and Tang, 2026](https://arxiv.org/html/2609.30316#bib.bib16);[Wu et al\., 2023](https://arxiv.org/html/2609.30316#bib.bib26);[Li et al\., 2023](https://arxiv.org/html/2609.30316#bib.bib18);[Lee et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib17)\)\. The standard evaluation scores each document, ranks assets cross\-sectionally, and reports an information coefficient\([Iacovides et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib19);[Iacovides et al\., 2025](https://arxiv.org/html/2609.30316#bib.bib27);[Vamvourellis and Mehta, 2025](https://arxiv.org/html/2609.30316#bib.bib20);[Wu et al\., 2025](https://arxiv.org/html/2609.30316#bib.bib21)\)\. These studies are run on historical data with models whose cutoff is unknown, and PIT models were introduced to remove that uncertainty by construction\.

PIT language models\.Pretraining on date\-filtered text has produced several public suites, including ChronoBERT and ChronoGPT\([He et al\., 2025a](https://arxiv.org/html/2609.30316#bib.bib1)\), PIT\([Kelly et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib3)\), and DatedGPT\([Yan et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib4)\)\. They are evaluated on whether the resulting signal is comparable to that of an unconstrained model, but not on whether the vintage that replaced it is any better\. A complementary line estimates the cutoff of a released model from dated text\([Cheng et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib6)\), which is necessary when it is not documented\. We instead treat the documented cutoff as the experimental variable, varying it across vintages while the evaluation window is held fixed\.

Temporal degradation of LLMs\.Language models are known to degrade on text drawn from after their training window\([Lazaridou et al\., 2021](https://arxiv.org/html/2609.30316#bib.bib7);[Dhingra et al\., 2022](https://arxiv.org/html/2609.30316#bib.bib8);[Luu et al\., 2022](https://arxiv.org/html/2609.30316#bib.bib9)\), an effect measured on perplexity and on knowledge probes\([Loureiro et al\., 2022](https://arxiv.org/html/2609.30316#bib.bib28);[Zhu et al\., 2025](https://arxiv.org/html/2609.30316#bib.bib29)\)\. That line of work evaluates the model rather than the decision, and proposes continued pretraining as the remedy\. We instead measure it on the decision itself, scoring financial news and reporting the information coefficient a practitioner would act on\.

Parameter\-efficient adaptation\.Low\-rank adaptation\([Hu et al\., 2022](https://arxiv.org/html/2609.30316#bib.bib13);[Dettmers et al\., 2023](https://arxiv.org/html/2609.30316#bib.bib25)\)adds a small trainable module to a frozen model, mainly to reduce the cost of supervised finetuning\([Houlsby et al\., 2019](https://arxiv.org/html/2609.30316#bib.bib14);[Li and Liang, 2021](https://arxiv.org/html/2609.30316#bib.bib24);[Iacovides et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib19)\)\. A cheaper alternative is to leave the weights untouched and supply the missing text at inference time, as in retrieval augmentation\([Lewis et al\., 2020](https://arxiv.org/html/2609.30316#bib.bib10);[Izacard et al\., 2023](https://arxiv.org/html/2609.30316#bib.bib11);[Choi et al\., 2025](https://arxiv.org/html/2609.30316#bib.bib22)\), which under a point\-in\-time constraint needs an index restricted to eligible documents\. We instead use an adapter to keep a checkpoint current rather than to finetune it cheaply, and the binding constraint is not the parameter budget but the publication date of the corpus\.

Positioning of our work\.An ideal method of keeping a model current would have three properties at once: 1\) point\-in\-time eligibility, 2\) a focus on financial text, and 3\) an update to the model itself\. As summarized in Table[1](https://arxiv.org/html/2609.30316#S2.T1), no line of work has all three\. PALM combines them by fitting a parameter\-efficient update on financial text the checkpoint was already allowed to read\.

Figure 2\.Key concepts and notation\.A timeline from 2017 to 2029\. Four horizontal bars, bracketed on the left as one point\-in\-time suite, each covering the text one vintage has read up to its own cutoff\. Three bars stop before the decision date at 2026 and are marked eligible, and one runs past it into the shaded evaluation window and is marked ineligible\. A dotted line marks the cutoff of the topmost vintage and a double arrow from it to the decision date is labelled staleness\.Figure 3\.Overall framework of PALM\. \(a\)Signal Extraction: A dated article is turned into a signal by contrasting upward against downward continuations in a single forward pass, and the signals are ranked within a date\. The blue square marks the position whose next\-token distribution is read\. \(b\)Point\-in\-Time Adaptation: Every pretrained weight is frozen and a low\-rank adapter is fitted on the text published between the checkpoint cutoff and the decision date\.Two boxed panels\. The left panel takes a dated news article, ends the prompt where a direction word would follow, reads the log\-odds between upward and downward continuations out of one forward pass, and ranks the assets of that date\. The right panel freezes the pretrained weights and trains a rank\-r adapter on the text published between the cutoff and the decision date, with an arrow showing that the adapter reads only that interval\.
## 3\.Proposed Method: PALM

### 3\.1\.Problem Setup

Vintages\.Let𝒟=\{\(xi,ti\)\}i=1N\\mathcal\{D\}=\\\{\(x\_\{i\},t\_\{i\}\)\\\}\_\{i=1\}^\{N\}be a corpus of documentsxix\_\{i\}with publication datestit\_\{i\}, and let𝒟≤tc=\{\(xi,ti\)∈𝒟:ti≤tc\}\\mathcal\{D\}\_\{\\leq t\_\{c\}\}=\\\{\(x\_\{i\},t\_\{i\}\)\\in\\mathcal\{D\}:t\_\{i\}\\leq t\_\{c\}\\\}denote its restriction to text published on or before a cutofftct\_\{c\}\. A PIT suite is a family\{ℳtc\}tc∈𝒞\\\{\\mathcal\{M\}\_\{t\_\{c\}\}\\\}\_\{t\_\{c\}\\in\\mathcal\{C\}\}of checkpoints produced by a single training pipeline, whereℳtc\\mathcal\{M\}\_\{t\_\{c\}\}carries parametersθtc\\theta\_\{t\_\{c\}\}obtained by minimizing the language modeling loss on the truncated corpus,

\(1\)θtc=arg⁡minθ​𝔼x∼𝒟≤tc​\[ℓLM​\(x,θ\)\]\.\\theta\_\{t\_\{c\}\}=\\arg\\min\_\{\\theta\}\\;\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\\leq t\_\{c\}\}\}\\left\[\\ell\_\{\\mathrm\{LM\}\}\(x;\\theta\)\\right\]\.Producing one vintage therefore costs one full pretraining run, and\|𝒞\|\|\\mathcal\{C\}\|grows by one every calendar year\.

Eligibility and staleness\.A study conducted at a decision datetdt\_\{d\}may useℳtc\\mathcal\{M\}\_\{t\_\{c\}\}only whentc≤tdt\_\{c\}\\leq t\_\{d\}, since otherwise the model has read text published after the decision it informs\. We call such a vintageeligiblefortdt\_\{d\}, and define itsstalenessass=td−tcs=t\_\{d\}\-t\_\{c\}, both of which Figure[2](https://arxiv.org/html/2609.30316#S2.F2)draws\. The prevailing convention is to select the eligible vintage of smallest staleness,

\(2\)tc⋆=max⁡\{tc∈𝒞:tc≤td\},t\_\{c\}^\{\\star\}=\\max\\\{t\_\{c\}\\in\\mathcal\{C\}:t\_\{c\}\\leq t\_\{d\}\\\},which is precisely what makes each additional year of coverage worth the run in Eq\. \([1](https://arxiv.org/html/2609.30316#S3.E1)\)\.

Comparison design\.Testing whether a largertct\_\{c\}yields a better model requires holding everything else constant\. However, existing evaluations move the window withtct\_\{c\}, and the two effects cannot be separated\. We instead fix a decision datetdt\_\{d\}, score every eligible vintage on that window, and regress performance on stalenessss\.

### 3\.2\.Signal Extraction

As shown in Figure[3](https://arxiv.org/html/2609.30316#S2.F3), the method has two halves, one that reads a signal out of a checkpoint and one that keeps it current\.

Scoring\.A checkpoint is judged by the signal it extracts from dated text\. Letxj​tx\_\{jt\}be a document concerning assetjjand datedtt, and letπ⁡\(x\)\\pi\(x\)denote a prompt that stops just before the word naming the direction of the move\. The signal assigned to that document is read from a single forward pass as the log\-odds between two sets of continuations,

\(3\)fθ\(x\)=log∑w∈𝒲\+pθ\(w∣π\(x\)\)−log∑w∈𝒲−pθ\(w∣π\(x\)\),f\_\{\\theta\}\(x\)=\\log\\\!\\\!\\sum\_\{w\\in\\mathcal\{W\}^\{\+\}\}\\\!\\\!p\_\{\\theta\}\(w\\mid\\pi\(x\)\)\\;\-\\;\\log\\\!\\\!\\sum\_\{w\\in\\mathcal\{W\}^\{\-\}\}\\\!\\\!p\_\{\\theta\}\(w\\mid\\pi\(x\)\),where𝒲\+\\mathcal\{W\}^\{\+\}and𝒲−\\mathcal\{W\}^\{\-\}hold upward and downward wordings\.

Measuring the signal\.Signals are compared within a date rather than across dates, since the level offθf\_\{\\theta\}carries no meaning while its cross\-sectional ordering does\. Writerj,t\+1r\_\{j,t\+1\}for the realized next\-period return andρ\\rhofor the Spearman correlation\. The quality of the signal over a windowTTis then the information coefficient

\(4\)IC=1\|T\|​∑t∈Tρ⁡\(\{fθ​\(xj​t\)\}j,\{rj,t\+1\}j\)\.\\mathrm\{IC\}=\\frac\{1\}\{\|T\|\}\\sum\_\{t\\in T\}\\rho\\big\(\\\{f\_\{\\theta\}\(x\_\{jt\}\)\\\}\_\{j\},\\;\\\{r\_\{j,t\+1\}\\\}\_\{j\}\\big\)\.Subtracting the mean within each date and taking the sign gives an equally weighted long\-short portfolio, and we report its return and Sharpe ratio\. The three measures are defined in detail in Section[4\.1](https://arxiv.org/html/2609.30316#S4.SS1)\.

### 3\.3\.Point\-in\-Time Adaptation

PALM brings an eligible vintage up to the decision datetdt\_\{d\}without changing a single pretrained weight\. Let

\(5\)𝒟\(tc,td\]=\{\(xi,ti\)∈𝒟:tc<ti≤td\}\\mathcal\{D\}\_\{\(t\_\{c\},t\_\{d\}\]\}=\\\{\(x\_\{i\},t\_\{i\}\)\\in\\mathcal\{D\}:t\_\{c\}<t\_\{i\}\\leq t\_\{d\}\\\}be the text published after the cutoff but before the decision date\. The vintage has never read it, and a practitioner standing attdt\_\{d\}already holds all of it\.

Adapter\.For every attention and feed\-forward projectionW∈ℝdout×dinW\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}ofθtc\\theta\_\{t\_\{c\}\}, PALM adds a low\-rank update\([Hu et al\., 2022](https://arxiv.org/html/2609.30316#bib.bib13)\)beside it as

\(6\)h=W​x\+αr​B​A​x,A∈ℝr×din,B∈ℝdout×r,h=Wx\+\\frac\{\\alpha\}\{r\}BAx,\\qquad A\\in\\mathbb\{R\}^\{r\\times d\_\{\\mathrm\{in\}\}\},\\;B\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times r\},with rankr≪min⁡\(din,dout\)r\\ll\\min\(d\_\{\\mathrm\{in\}\},d\_\{\\mathrm\{out\}\}\)and a fixed scaleα\\alpha\. BothAAandBBare trained, and we initializeB=0B=0, which leaves the adapted model identical to the vintage before any update is applied\.

Objective\.Writingϕ=\{\(Aℓ,Bℓ\)\}ℓ\\phi=\\\{\(A\_\{\\ell\},B\_\{\\ell\}\)\\\}\_\{\\ell\}for the adapter parameters across layers, PALM solves

\(7\)ϕ⋆=arg⁡minϕ​𝔼x∼𝒟\(tc,td\]​\[ℓLM​\(x,θtc,ϕ\)\]\\phi^\{\\star\}=\\arg\\min\_\{\\phi\}\\;\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\(t\_\{c\},t\_\{d\}\]\}\}\\left\[\\ell\_\{\\mathrm\{LM\}\}\(x;\\theta\_\{t\_\{c\}\},\\phi\)\\right\]withθtc\\theta\_\{t\_\{c\}\}held frozen, and returns the adapted modelℳtc→td\\mathcal\{M\}\_\{t\_\{c\}\\to t\_\{d\}\}\. Eq\. \([7](https://arxiv.org/html/2609.30316#S3.E7)\) is the same objective as Eq\. \([1](https://arxiv.org/html/2609.30316#S3.E1)\), differing only in the interval the corpus is drawn from and in which parameters are free\. The adapter is far smaller, with\|ϕ\|≪\|θtc\|\|\\phi\|\\ll\|\\theta\_\{t\_\{c\}\}\|\.

Two properties follow\. First,ℳtc→td\\mathcal\{M\}\_\{t\_\{c\}\\to t\_\{d\}\}is eligible fortdt\_\{d\}wheneverℳtc\\mathcal\{M\}\_\{t\_\{c\}\}is, since neither Eq\. \([1](https://arxiv.org/html/2609.30316#S3.E1)\) nor Eq\. \([7](https://arxiv.org/html/2609.30316#S3.E7)\) reads text published aftertdt\_\{d\}\. Second, PALM is a plug\-in method, as the update in Eq\. \([6](https://arxiv.org/html/2609.30316#S3.E6)\) sits inside existing linear layers and does not change the architecture\.

## 4\.Experiments

### 4\.1\.Experimental Setup

Datasets\.From FNSPID\([Dong et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib12)\)we construct a panel of 252,000 article\-day observations over 1,099 tickers between 2014 and 2023, sampled evenly across months\. An article is credited to a trading day only if its timestamp precedes that day’s close, and the target is the next session’s adjusted\-close return\. The decision date is the end of November 2020, and every eligible vintage is scored on one window running from December 2020 through December 2023\.

Table 2\.Experimental setup: Data, Models, Evaluation\.Table 3\.Model checkpoints\. Every suite releases one vintage per calendar year without gaps, and the counts differ only through where each suite starts and where it stops\.SuiteScoredEligibleExcludedcutoffsnnCutoffsnnCutoffsChronoGPT261999–2024211999–20192020 tolatestChronoGPT\-Instruct212004–2024162004–2019DatedGPT122013–202472013–2019Total59—44—15Backbones\.We use the public PIT suites that train each vintage separately, and apply PALM to each without modifying the backbone\. As shown in Table[3](https://arxiv.org/html/2609.30316#S4.T3), these span three training pipelines and two decades of cutoffs\. One further suite, PIT\-4B\([Kelly et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib3)\), releases its vintages as snapshots of one continuous run\. We hold it out and treat it in Section[5\.2](https://arxiv.org/html/2609.30316#S5.SS2)as the pretraining PALM replaces\.

Evaluation metrics\.We report three measures, each of which reads the same daily series in a different way\.

- •a\)Information coefficient\.How well the score orders the cross\-section on a given day, before any position is taken\.
- •b\)Annualized return\.What that ordering is worth once it is turned into an equally weighted long\-short book\.
- •c\)Sharpe ratio\.What the same book earns per unit of risk, the number a desk is held to\.

The first is Eq\. \([4](https://arxiv.org/html/2609.30316#S3.E4)\), and for the other two we writewj​t=sign⁡\(fθ​\(xj​t\)−f¯t\)w\_\{jt\}=\\operatorname\{sign\}\(f\_\{\\theta\}\(x\_\{jt\}\)\-\\bar\{f\}\_\{t\}\)for the position taken on namejjat datett, withf¯t\\bar\{f\}\_\{t\}the mean score on that date\. The portfolio return is thenrt=∑jwj​t​rj,t\+1/∑j\|wj​t\|r\_\{t\}=\\sum\_\{j\}w\_\{jt\}r\_\{j,t\+1\}/\\sum\_\{j\}\|w\_\{jt\}\|, with meanr¯\\bar\{r\}and standard deviationσr\\sigma\_\{r\}over the window\. We report252​r¯252\\,\\bar\{r\}as the annualized return and252​r¯/σr\\sqrt\{252\}\\,\\bar\{r\}/\\sigma\_\{r\}as the Sharpe ratio\. Unless otherwise stated, the IC is in units of10−310^\{\-3\}and the annualized return in percentage points\. Wins is the share of checkpoints whose measure the update raises, with the same checkpoint scored before and after\. Both arms of a comparison are scored on the same span of the panel\. An IC near0\.010\.01is not small in economic terms\. The fundamental law of active management givesIR=IC​N\\mathrm\{IR\}=\\mathrm\{IC\}\\sqrt\{N\}forNNindependent bets, and our window carries roughly 86 names per day, orN≈21,672N\\approx 21,672\{\}per year\.

Implementation details\.Scores come from Eq\. \([3](https://arxiv.org/html/2609.30316#S3.E3)\) throughout\. Scoring involves no sampling, and every backbone is loaded at the settings of its original release, which leaves the adapter as the only thing we vary\. Unless stated otherwise, PALM uses rankr=2r=2with scaleα=2​r\\alpha=2r, and trains for 600 steps at a learning rate of10−410^\{\-4\}on sequences of 256 tokens\.

Table 4\.Effect of staleness\. No suite carries the negative slope the annual pretraining run is meant to prevent, meaning thata further year leaves the signal where it found it\.Table 5\.Effectiveness of PALM across backbones\. The same adapter raises the IC on every suite, showing thatthe gain does not depend on which model PALM is attached to\.Figure 4\.Effect of staleness and of PALM\. \(a\) Information coefficient against staleness on the fixed evaluation window, which leaves the checkpoint as the only quantity varying between points\. Points are demeaned within suite, matching the specification of Table[4](https://arxiv.org/html/2609.30316#S4.T4)\. \(b\) The same checkpoints before and after PALM, paired on the days both systems scored\.Two panels\. The left one scatters the information coefficient against staleness for every eligible checkpoint of three suites, with a fitted line that does not slope downward\. The right one connects each checkpoint to itself before and after the adapter, and almost every connector rises\.Figure 5\.Effect of the adapter capacity\. The gain in the IC is positive at every adapter rankrr, and a small rank already reaches it, showing thata few directions carry the update\.The gain in the information coefficient plotted against adapter rank on a log scale from one to thirty\-two, with a bootstrap interval at each rank\. The curve is positive throughout, rises to a peak at small rank, and falls towards zero at the largest rank\.
### 4\.2\.Effect of Staleness

To see whether an older checkpoint scores worse on the decision it is used for, we regress the IC on staleness within each suite, holding the evaluation window fixed\. Slopes are per year of staleness, over the eligible checkpoints of Table[3](https://arxiv.org/html/2609.30316#S4.T3)\. As shown in Figure[4](https://arxiv.org/html/2609.30316#S4.F4)\(a\) and Table[4](https://arxiv.org/html/2609.30316#S4.T4), no suite carries a slope distinguishable from zero\. The slopes even lean positive, which is the opposite of the improvement the annual pretraining run is designed for\. The results show thatthe more recent checkpoint is not the better one\.

### 4\.3\.Effectiveness of PALM

To see whether an adapter can make effective use of the new text where the annual pretraining run cannot, we score each checkpoint with and without PALM and compare the two\. As shown in Figure[4](https://arxiv.org/html/2609.30316#S4.F4)\(b\) and Table[5](https://arxiv.org/html/2609.30316#S4.T5), PALM raises the IC on every backbone and on nine checkpoints in ten, and raises all three measures once the backbones are pooled\. A Wilcoxon signed\-rank test over the 44 paired checkpoints givesp<0\.001p<0\.001, and the two suites with enough vintages to test on their own reachp<0\.001p<0\.001andp=0\.004p=0\.004\. Note that the three suites share no training pipeline, no tokenizer and no architecture family\. A gain appearing in all of them is therefore a property of the update rather than of any one release, anda single adapter setting is enough to recover it everywhere\.

## 5\.Analysis

### 5\.1\.Effect of the Adapter Capacity

Table 6\.Continued pretraining against PALM, on PIT\-4B\([Kelly et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib3)\), whose vintages come from one run\. Six more years of the run leave the score where they found it while the adapter moves it, showing thatcompute is not the missing ingredient\.We vary the adapter rank over the same 44 checkpoints of Table[5](https://arxiv.org/html/2609.30316#S4.T5), which asks how much capacity the update needs\. The curve averages over all three backbones, and the rank sets how many independent directions the update may use\. As shown in Figure[5](https://arxiv.org/html/2609.30316#S4.F5), the gain is positive at every rank, rising to a peak at moderate rank and falling beyond it\. A small rank cannot carry the update, and a large one overwrites what the checkpoint already holds\. We adoptr=2r=2, the smallest rank that is not worse than the peak\.

### 5\.2\.Comparison with Continued Pretraining

While the vintages of Section[4\.3](https://arxiv.org/html/2609.30316#S4.SS3)are trained from scratch, another option iscontinued pretraining, which extends one run rather than starting a new one\. PIT\-4B\([Kelly et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib3)\)is built this way, with its vintages taken as monthly snapshots of one continuous pretraining run over time\-ordered text\. The gap between two snapshots is therefore a controlled comparison, with the run as the only treatment\. Continued pretraining isthe procedure PALM is proposed to replace, and this suite lets us compare the two on the same vintages\.

Table[6](https://arxiv.org/html/2609.30316#S5.T6)compares three things\. The first row \(w/o update\) averages the snapshots the run produced between 2013 and 2018, the second \(w/ continued pretraining\) is the snapshot it produced in 2019, and the third \(w/ PALM111We user=16r=16here, since this backbone is 4096 wide against 1536 to 2048 elsewhere\.\) is the first row again with an adapter attached\. Here, the continued pretraining moves every weight in the model, while the adapter moves a few directions and trains only on what the vintage is missing\. The table shows that six more years of the run \(w/ continued pretraining\) move the IC by\+0\.10\+0\.10, while the adapter \(w/ PALM\) moves the same checkpoints by\+1\.37\+1\.37\. The older snapshots carrying an adapter overtake the one the run produced six years later\. This demonstrates thatan adapter costing a fraction of a percent of the weights outperforms six more years of continued pretraining at four billion parameters\.

### 5\.3\.Persistence of the Gain

Table 7\.How often a suite must be retrained\. The adapted score does not fall as the base ages, and the oldest group adapted outscores the newest group unadapted, showing thatno retraining horizon appears within two decades of cutoffs\.The daily difference in information coefficient between the adapted and unadapted checkpoints, accumulated over the evaluation window\. The curve climbs steadily from the start of the window to its end rather than jumping in one stretch\.Figure 6\.Cumulative gain over the evaluation window\. The daily difference accumulates through the window rather than in one stretch, showing thatthe gain is not one episode\.Age of the base checkpoint\.We ask whether the gain shrinks as the base checkpoint ages, since a gain that faded with age would mean an old vintage eventually has to be replaced rather than adapted\. Table[6](https://arxiv.org/html/2609.30316#S5.F6)groups the same 44 paired checkpoints by the age of the base, and no such replacement age appears within the two decades of cutoffs we test\. An old base is adapted as well as a new one, and the adapted score of the oldest checkpoints exceeds the unadapted score of the newest ones\. This says more than that the annual pretraining run is unnecessary, sinceage is not by itself a reason to replace a checkpoint\.

Position in the evaluation window\.To check that the gain is spread across the window rather than concentrated in part of it, we add up the daily difference between the adapted and the unadapted score\. As shown in Figure[6](https://arxiv.org/html/2609.30316#S5.F6), the total climbs steadily from the first day to the last rather than jumping once\. The gain is positive in every calendar year of the window, and no year gives it back\.

### 5\.4\.Origin of the Gain

Table 8\.Articles whose score moves most under PALM, for a 2015 checkpoint\. Most of them concern a firm or an event the vintage had never encountered, showing thatthe update lands on the material the vintage could not have seen\.A horizontal bar chart of the words whose score moves most under the adapter\. Terms are listed on the vertical axis and the mean shift in standard deviations on the horizontal one, with roughly equal numbers moving up and down\.Figure 7\.The vocabulary the update moves\. Every term shown moves the same way on four checkpoints and came into use only after their cutoffs, and the update is repricing language the vintage could not have read, showing thatwhat a stale checkpoint lacks is how later coverage is written\.The articles the update moves\.To see what kind of article the adapter changes the score of, we put the same articles through one checkpoint twice, with and without PALM, and compare the two scores\. Specifically, we take the ChronoGPT vintage cut at the end of 2015 and standardize the scores in each arm to mean zero and unit variance\. Table[7](https://arxiv.org/html/2609.30316#S5.F7)reports the six largest shifts \(scores w/ PALM−\-scores w/o PALM\), and most of them are about a firm or an event the vintage had never encountered\.

Figure[7](https://arxiv.org/html/2609.30316#S5.F7)repeats the test on words222Words are filtered on two conditions: 1\) the shift has the same sign on all four checkpoints, ruling out chance, and 2\) the word is at least twice as common in the panel after the cutoff as before, leaving only what the vintage could not have read\.rather than on six articles, scoring each word by the average shift of the articles that contain it\. The update pushes these words in both directions rather than merely inflating them, indicating that this is not a familiarity effect, where reading a word more often makes the model score it higher whatever the word means\. The largest shifts concern electric vehicle makers, a commodity rally driven by the war in Ukraine, and a clinical\-trial setback at a firm the panel did not cover in 2015\.

What the gain needs, and what it could reach\.We run two artificial interventions on the adaptation corpus\. First, we scramble the word order of the corpus, which destroys what the text says while leaving its words and its size untouched\. This asks whether the gain came from reading the new text at all, or merely from moving the weights\. As shown in Table[9](https://arxiv.org/html/2609.30316#S5.T9)\(a\), the scrambled corpus keeps\+0\.48\+0\.48of the\+1\.26\+1\.26IC gain while the return falls to−0\.24\-0\.24, and rank 8 turns every measure negative\. The gain is therefore not a by\-product of the training, sinceit needs what the text says\.

Second, we let the adapter read past the decision date, which adds text the eligibility rule forbids while leaving the checkpoint and the evaluation window untouched\. This asks whether that rule holds the update back, or whether the eligible text already carries what the adapter needs\. As shown in Table[9](https://arxiv.org/html/2609.30316#S5.T9)\(b\), doing so raises the IC by\+1\.05\+1\.05against\+1\.26\+1\.26for staying eligible, and loses to it atp=0\.008p=0\.008\. The eligibility rule therefore costs nothing measurable, sincethe forbidden text holds nothing the eligible text does not\.

Table 9\.What the gain needs, and what it could reach\.The gain needs the eligible text, and nothing beyond it\.Table 10\.Sensitivity analysis 1 \- Training budget\. The gain appears only once the adapter has been trained enough to fit the corpus, showing thatthe update has to be trained and not merely attached to the model\.
### 5\.5\.Sensitivity Analysis

To check that the gain comes from the update rather than from the settings we report, we vary one choice at a time and pair each variant against the same checkpoints of Table[5](https://arxiv.org/html/2609.30316#S4.T5)\. Every other choice is held at the reported setting, and we examine four aspects:

- •Training budget: How much training the update needs\.
- •Decision date: Whether moving the date we fixed matters\.
- •Adapted projections: Which part of the block carries it\.
- •Share of the text: How much of the corpus it must read\.

Training budget\.To check how much training the adapter needs, we vary the number of steps and the learning rate\. Table[10](https://arxiv.org/html/2609.30316#S5.T10)reports the gain at each setting\. Cutting the steps to a third, or the rate by a factor of five, removes the gain, while doubling the steps changes little\. This demonstrates thatthe update has to be trained past a floor, above which more training adds little\.

Table 11\.Sensitivity analysis 2 \- Decision date\. Moving the date by two years in either direction leaves the gain in place, showing thatthe result does not depend on our window\.Table 12\.Sensitivity analysis 3 \- Adapted projections\. Neither half of the block reproduces the gain on its own, showing thatthe update is not carried by one part of the layer\(FFN: feed\-forward projections\)\.Table 13\.Sensitivity analysis 4 \- Share of eligible text\. A tenth of the corpus carries four fifths of what all of it delivers, showing thatthe gain sits in the text nearest the decision date\.Decision date\.To check whether the result depends on the date we fixed, we move it two years in each direction\. Table[11](https://arxiv.org/html/2609.30316#S5.T11)reports the gain at each date, and the pool moves with it, as a later date admits vintages an earlier one excludes\. The gain survives both, showing thatthe gain does not come from where we placed it\.

Adapted projections\.To see which part of the block carries the gain, we adapt the attention and the feed\-forward projections one at a time\. Table[12](https://arxiv.org/html/2609.30316#S5.T12)reports the gain for each of the two, and for both together as we report them elsewhere\. Adapting either half by itself recovers roughly half of what adapting both recovers, which shows thatneither part carries the gain on its own\.

Share of the eligible text\.To see how much text the update needs, we fit the adapter on a fraction of the eligible corpus rather than on all of it, keeping the most recent slice in each case\. Table[13](https://arxiv.org/html/2609.30316#S5.T13)reports the gain at each share, from a tenth of the corpus up to the whole of it\. A tenth already carries four fifths of what the full corpus delivers, and the remaining ninety percent of the text buys the rest, which shows thatthe gain is concentrated in the text that sits closest to the decision date\.

## 6\.Limitations

Our evidence comes from one asset class over a single decade, and from the point\-in\-time suites that are public today\. Whether the same picture holds for other markets, other prediction horizons, or backbones an order of magnitude larger is beyond what a study of this scope settles\. The signal we measure is weak in absolute terms, as signals extracted from news generally are, and conclusions about stronger signals do not follow from it\.

The gain is also established under one protocol, from how a score is read out of the model to how that score becomes a position\. We report what that protocol measures, and we do not claim that the size of the gain is invariant to those choices\. Separating which of them the gain rests on would take a study designed around that question rather than around the annual pretraining run\.

## 7\.Conclusion

In this paper, we tested whether the annual pretraining run of a PIT suite is necessary\. Holding the evaluation window fixed, and comparing each vintage against the vintage that replaced it, we found that staleness does not degrade downstream performance\. Motivated by this observation, we proposed PALM, a simple yet effective plug\-in method that fits a low\-rank adapter on eligible text without modifying any pretrained weight\.

Our results suggest that a PIT suite should be maintained byupdating the checkpoint it already hasrather than by producing a new one each calendar year\. They also suggest that the update can be small, since an adapter of a few directions already recovers the gain\. This preserves the chronological guarantee the suite is built to provide, and frees the budget the calendar consumes for the coverage that is actually missing, which is more architectures rather than more years of one\.

## References

- Chenget al\.\(2024\)J\. Cheng, M\. Marone, O\. Weller, D\. Lawrie, D\. Khashabi, and B\. Van DurmeDated data: tracing knowledge cutoffs in large language models\.InConference on Language Modeling \(COLM\),Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p2.1),[§2](https://arxiv.org/html/2609.30316#S2.p2.1)\.
- Choiet al\.\(2025\)C\. Choi, J\. Kwon, J\. Ha, H\. Choi, C\. Kim, Y\. Lee, J\. Sohn, and A\. Lopez\-LiraFinDER: financial dataset for question answering and evaluating retrieval\-augmented generation\.InProceedings of the 6th ACM International Conference on AI in Finance,Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p4.1)\.
- Dettmerset al\.\(2023\)T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. ZettlemoyerQLoRA: efficient finetuning of quantized llms\.InNeurIPS,Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p4.1)\.
- Dhingraet al\.\(2022\)B\. Dhingra, J\. R\. Cole, J\. M\. Eisenschlos, D\. Gillick, J\. Eisenstein, and W\. W\. CohenTime\-aware language models as temporal knowledge bases\.Transactions of the Association for Computational Linguistics10,pp\. 257–273\.Cited by:[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.4.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p3.1)\.
- Donget al\.\(2024\)Z\. Dong, X\. Fan, and Z\. PengFNSPID: a comprehensive financial news dataset in time series\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,Cited by:[Table 14](https://arxiv.org/html/2609.30316#A1.T14.4.3.2),[Appendix A](https://arxiv.org/html/2609.30316#A1.p1.1),[§1](https://arxiv.org/html/2609.30316#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.30316#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2609.30316#S4.T2.2.3.2.1)\.
- Glasserman and Lin \(2023\)P\. Glasserman and C\. LinAssessing look\-ahead bias in stock return predictions generated by GPT sentiment analysis\.arXiv:2309\.17322\.Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p2.1)\.
- Heet al\.\(2025a\)S\. He, L\. Lv, A\. Manela, and J\. WuChronologically consistent large language models\.arXiv:2502\.21206\.Cited by:[1st item](https://arxiv.org/html/2609.30316#A2.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.30316#S1.p3.1),[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.3.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p2.1),[Table 2](https://arxiv.org/html/2609.30316#S4.T2.2.8.2.1)\.
- Heet al\.\(2025b\)S\. He, L\. Lv, A\. Manela, and J\. WuInstruction tuning chronologically consistent language models\.arXiv:2510\.11677\.Cited by:[1st item](https://arxiv.org/html/2609.30316#A2.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.30316#S1.p3.1),[Table 2](https://arxiv.org/html/2609.30316#S4.T2.2.8.2.1)\.
- Houlsbyet al\.\(2019\)N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. de Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. GellyParameter\-efficient transfer learning for NLP\.InICML,Cited by:[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.5.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p4.1)\.
- Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InICLR,Cited by:[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.5.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p4.1),[§3\.3](https://arxiv.org/html/2609.30316#S3.SS3.p2.1)\.
- Iacovideset al\.\(2024\)G\. Iacovides, T\. Konstantinidis, M\. Xu, and D\. MandicFinLlama: llm\-based financial sentiment analysis for algorithmic trading\.InProceedings of the 5th ACM International Conference on AI in Finance,Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1),[§2](https://arxiv.org/html/2609.30316#S2.p4.1)\.
- Iacovideset al\.\(2025\)G\. Iacovides, W\. Zhou, and D\. MandicFinDPO: financial sentiment analysis for algorithmic trading through preference optimization of llms\.InProceedings of the 6th ACM International Conference on AI in Finance,Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Izacardet al\.\(2023\)G\. Izacard, P\. Lewis, M\. Lomeli, L\. Hosseini, F\. Petroni, T\. Schick,et al\.Atlas: few\-shot learning with retrieval augmented language models\.Journal of Machine Learning Research24\(251\),pp\. 1–43\.Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p4.1)\.
- Keet al\.\(2019\)Z\. T\. Ke, B\. T\. Kelly, and D\. XiuPredicting returns with text data\.NBER Working Paper 26186\.Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p1.1),[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.2.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Kellyet al\.\(2026\)B\. Kelly, S\. Malamud, J\. Schwab, and T\. A\. XuScaling point\-in\-time language models\.arXiv:2607\.11889\.Cited by:[2nd item](https://arxiv.org/html/2609.30316#A2.I1.i2.p1.1),[§1](https://arxiv.org/html/2609.30316#S1.p3.1),[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.3.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p2.1),[§4\.1](https://arxiv.org/html/2609.30316#S4.SS1.p2.1),[§5\.2](https://arxiv.org/html/2609.30316#S5.SS2.p1.1),[Table 6](https://arxiv.org/html/2609.30316#S5.T6)\.
- Lazaridouet al\.\(2021\)A\. Lazaridou, A\. Kuncoro, E\. Gribovskaya, D\. Agrawal, A\. Liska, T\. Terzi, M\. Gimenez,et al\.Mind the gap: assessing temporal generalization in neural language models\.InNeurIPS,Cited by:[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.4.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p3.1)\.
- Leeet al\.\(2026\)S\. Lee, J\. Seo, J\. Lee, S\. Yoo, M\. Kim, T\. Y\. Lim, D\. Kang, H\. Choi, S\. Lee, and W\. AhnFinSTaR: towards financial reasoning with time series reasoning models\.InKDD Workshop on SciSoc Agents and LLMs,Cited by:[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.2.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal,et al\.Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.InNeurIPS,Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p4.1)\.
- Liet al\.\(2026\)W\. W\. Li, M\. Wang, and T\. MaSummoning the oracle to slay it: mitigating look\-ahead bias in financial backtesting with large language models\.arXiv:2605\.24564\.Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p2.1)\.
- Li and Liang \(2021\)X\. L\. Li and P\. LiangPrefix\-tuning: optimizing continuous prompts for generation\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics,Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p4.1)\.
- Liet al\.\(2023\)Y\. Li, S\. Wang, H\. Ding, and H\. ChenLarge language models in finance: a survey\.InProceedings of the Fourth ACM International Conference on AI in Finance,Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Lopez\-Lira and Tang \(2026\)A\. Lopez\-Lira and Y\. TangCan ChatGPT forecast stock price movements? return predictability and large language models\.Journal of Financial Economics\.Note:ForthcomingCited by:[§1](https://arxiv.org/html/2609.30316#S1.p2.1),[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.2.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Loshchilov and Hutter \(2019\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.InICLR,Cited by:[Appendix C](https://arxiv.org/html/2609.30316#A3.p2.1)\.
- Loureiroet al\.\(2022\)D\. Loureiro, F\. Barbieri, L\. Neves, L\. Espinosa Anke, and J\. Camacho\-ColladosTimeLMs: diachronic language models from twitter\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations,Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p3.1)\.
- Luuet al\.\(2022\)K\. Luu, D\. Khashabi, S\. Gururangan, K\. Mandyam, and N\. A\. SmithTime waits for no one\! analysis and challenges of temporal misalignment\.InProceedings of NAACL\-HLT,pp\. 5944–5958\.Cited by:[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.4.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p3.1)\.
- Sarkar and Vafa \(2024\)S\. K\. Sarkar and K\. VafaLookahead bias in pretrained language models\.SSRN Working Paper 4754678\.Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p3.1)\.
- Vamvourellis and Mehta \(2025\)D\. Vamvourellis and D\. MehtaReasoning or overthinking: evaluating large language models on financial sentiment analysis\.InProceedings of the 6th ACM International Conference on AI in Finance,Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Wuet al\.\(2023\)S\. Wu, O\. Irsoy, S\. Lu, V\. Dabravolski, M\. Dredze, S\. Gehrmann, P\. Kambadur, D\. Rosenberg, and G\. MannBloombergGPT: a large language model for finance\.arXiv:2303\.17564\.Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Wuet al\.\(2025\)Y\. Wu, E\. M\. Akin, C\. Martineau, V\. Grégoire, and A\. VenerisExtracting the structure of press releases for predicting earnings announcement returns\.InProceedings of the 6th ACM International Conference on AI in Finance,Cited by:[§1](https://arxiv.org/html/2609.30316#S1.p1.1),[§2](https://arxiv.org/html/2609.30316#S2.p1.1)\.
- Yanet al\.\(2026\)Y\. Yan, R\. Tang, Z\. Gao, W\. Jiang, and Y\. LuDatedGPT: preventing lookahead bias in large language models with time\-aware pretraining\.arXiv:2603\.11838\.Cited by:[1st item](https://arxiv.org/html/2609.30316#A2.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.30316#S1.p3.1),[Table 1](https://arxiv.org/html/2609.30316#S2.T1.2.3.1.1),[§2](https://arxiv.org/html/2609.30316#S2.p2.1),[Table 2](https://arxiv.org/html/2609.30316#S4.T2.2.9.2.1)\.
- Zhuet al\.\(2025\)C\. Zhu, N\. Chen, Y\. Gao, Y\. Zhang, P\. Tiwari, and B\. WangIs your llm outdated? a deep look at temporal generalization\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics,Cited by:[§2](https://arxiv.org/html/2609.30316#S2.p3.1)\.

## Appendix ADataset

The panel is built from FNSPID\([Dong et al\., 2024](https://arxiv.org/html/2609.30316#bib.bib12)\), a public corpus of dated financial news\. We sample it evenly across months, which keeps any single year from dominating the pool\. Each row is one article, and a ticker carrying several articles on one day contributes several rows rather than one average, which is why Section[4\.1](https://arxiv.org/html/2609.30316#S4.SS1)counts the panel in article\-days rather than in firm\-days of the kind a daily study would use\.

An article joins a trading day only when its timestamp precedes that day’s close, taken conservatively at 20:00 UTC\. A story published after the close therefore belongs to the next trading day and not to the session it could not have moved\. The target is the return of thefollowingsession, computed from adjusted closes, and never the session the article itself belongs to\. Both rules keep the panel from scoring an article against a move that came before it\. The decision date then splits the panel into the two halves the method needs, one an adapter may read and one every checkpoint is scored on\. Table[14](https://arxiv.org/html/2609.30316#A1.T14)gives the size of each half and the rest of what the panel holds\.

Table 14\.Details of FNSPID\. Fewer than a third of the observations sit in the evaluation window, and every eligible vintage is scored on all of them, showing thatthe comparison rests on the same articles whichever checkpoint is being read\.
## Appendix BBackbones

Four public point\-in\-time suites are used, and they divide into two kinds\.

- •ChronoGPT\([He et al\., 2025a](https://arxiv.org/html/2609.30316#bib.bib1)\), ChronoGPT\-Instruct\([He et al\., 2025b](https://arxiv.org/html/2609.30316#bib.bib2)\)and DatedGPT\([Yan et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib4)\)train each vintage separately from scratch, which makes a vintage a self\-contained model rather than a stage of anything\.
- •PIT\-4B\([Kelly et al\., 2026](https://arxiv.org/html/2609.30316#bib.bib3)\)publishes monthly snapshots of one continuous run, which is why Section[5\.2](https://arxiv.org/html/2609.30316#S5.SS2)holds it out as the pretraining PALM replaces\.

Every checkpoint is loaded at the settings of its original release, in bfloat16 and in evaluation mode, and nothing about the architecture is changed\. One scoring routine reads all four suites, which keeps the comparison between them free of any difference in how a score was taken\. Table[15](https://arxiv.org/html/2609.30316#A2.T15)lists what each suite releases and how much of it is eligible for our decision date\.

Table 15\.The four suites\. The three suites we pair release one vintage per calendar year without gaps and their eligible counts differ only through where each suite starts, while PIT\-4B publishes monthly checkpoints from which we take one per calendar year, showing thatthe paired pool is set by release history rather than by any choice of ours\.
## Appendix CImplementation Details

Every pretrained parameter is frozen before a single adapter is created, since injecting a module does not by itself freeze the layers it was not placed in, and leaving them trainable would make the run a full finetune rather than a low\-rank update\. The four suites do not name their attention and feed\-forward layers alike, and the loader therefore holds one list of names per architecture\. InitializingBBat zero leaves the adapted model identical to the vintage it was attached to until the first optimizer step is taken\.

We do not tune the optimization, since a result that needed a tuned schedule would be a result about the schedule\. The optimizer is AdamW\([Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.30316#bib.bib31)\), and each batch is drawn with replacement from the eligible interval\. Section[J](https://arxiv.org/html/2609.30316#A10)reports what happens on either side of the rank, the number of steps and the learning rate, one at a time\.

Reading a score involves no trained parameter and no sampling\. The same input therefore returns the same score on every run\. Table[16](https://arxiv.org/html/2609.30316#A3.T16)collects the settings that hold everywhere in the paper unless a sensitivity arm states otherwise\.

Table 16\.The settings that hold throughout\. Everything except the rank, the steps and the rate is fixed once and never tuned per suite, showing thatthe gain is not the product of a search over configurations\.
## Appendix DThree Reported Measures

All three measures are read off one daily series\. Each trading day of the evaluation window yields a cross\-section of scores and the returns those scores were meant to order, and contributes one number to each of the three\. They differ in what they do with a day rather than in the data they read, which is why a variant can move one of them while leaving the other two where it found them\.

The unit of the cross\-section is thearticleand not the firm, since a ticker carrying several articles on one day enters the ranking several times with the same realized return, and the median day carries 99 articles over 86 distinct tickers\. A day is scored only when it holds at least eight articles and four distinct scores, which removes 6 of the 773 dates in the window\. Both conventions hold everywhere in the paper\.

The tables report these measures in fixed units\. The IC is given in units of10−310^\{\-3\}and the annualized return in percentage points, and Wins is the share of paired checkpoints a variant improves\. Bold marks the best entry among the arms a row or a pair compares\.

### D\.1\.Information Coefficient

The information coefficient asks whether the ordering the model produced on a given day matched the ordering the market delivered\. Writingρ\\rhofor the Spearman rank correlation,fθ​\(xj​t\)f\_\{\\theta\}\(x\_\{jt\}\)for the score of articlejjon daytt, andrj,t\+1r\_\{j,t\+1\}for the return of the following session, Eq\. \([4](https://arxiv.org/html/2609.30316#S3.E4)\) averages the daily correlation over the windowTT,

IC=1\|T\|​∑t∈Tρ⁡\(\{fθ​\(xj​t\)\}j,\{rj,t\+1\}j\)\.\\mathrm\{IC\}=\\frac\{1\}\{\|T\|\}\\sum\_\{t\\in T\}\\rho\\big\(\\\{f\_\{\\theta\}\(x\_\{jt\}\)\\\}\_\{j\},\\;\\\{r\_\{j,t\+1\}\\\}\_\{j\}\\big\)\.The correlation is computedwithina day, which removes the move common to every name that session and prevents a market direction from being read as a signal\. It is arankcorrelation, which makes the measure invariant to the units offθf\_\{\\theta\}and keeps a single large return from setting the value on its own\. The daily correlations are then averaged with equal weight, which keeps a day with many articles from counting for more than a thin one\. Each of the three choices keeps the measure from reading something other than the ordering it is meant to score\.

On the worked example of Section[I](https://arxiv.org/html/2609.30316#A9), the 91 articles of 13 December 2021 giveICt=\+0\.0067\\mathrm\{IC\}\_\{t\}=\+0\.0067, and averaging over the 767 scored days gives0\.006610\.00661for that checkpoint, which the tables print as6\.616\.61in units of10−310^\{\-3\}\. A value of that size looks negligible against the maximum of one, and the fundamental law of active management is what puts it in scale\. WithIR=IC​N\\mathrm\{IR\}=\\mathrm\{IC\}\\sqrt\{N\}and roughly 86 names a day over 252 sessions,N≈21,672N\\approx 21\{,\}672andN≈147\\sqrt\{N\}\\approx 147, which turns an IC of0\.00660\.0066into an information ratio near one\.

### D\.2\.Annualized Return

The information coefficient scores an ordering, and the return scores what an ordering is worth once positions are taken\. The score is demeaned within the day, using that day’s mean scoref¯t\\bar\{f\}\_\{t\}, and its sign becomes the position, which gives a long\-short book that is equally weighted by construction,

wj​t=sign⁡\(fθ​\(xj​t\)−f¯t\),rt=∑jwj​t​rj,t\+1∑j\|wj​t\|,Return=252​r¯\.w\_\{jt\}=\\operatorname\{sign\}\\big\(f\_\{\\theta\}\(x\_\{jt\}\)\-\\bar\{f\}\_\{t\}\\big\),\\qquad r\_\{t\}=\\frac\{\\sum\_\{j\}w\_\{jt\}\\,r\_\{j,t\+1\}\}\{\\sum\_\{j\}\|w\_\{jt\}\|\},\\qquad\\text\{Return\}=252\\,\\bar\{r\}\.Taking the sign rather than the score itself is what separates this measure from the IC\. The IC uses the whole ordering, while the book only asks which side of the day’s average an article fell on, and it therefore discards the magnitude of the score\. That is why an adapter can raise the IC while leaving the book unmoved, and Section[E](https://arxiv.org/html/2609.30316#A5)reports one case where the two disagree in sign\.

The book is not exactly dollar neutral\. The sign of a demeaned score splits a day into two sides, but the split need not be even, and on 13 December 2021 it was 55 long against 36 short\. The denominator∑j\|wj​t\|\\sum\_\{j\}\|w\_\{jt\}\|normalizes by the number of positions rather than by the smaller side, and the reported return is therefore the average return of a position taken that day\. No transaction cost, borrowing cost or slippage is deducted, which makes every return here a gross figure to be read as a difference between arms rather than as a tradable one\.

For the same checkpoint,rt=−0\.33%r\_\{t\}=\-0\.33\\%on the example day, and252​r¯=9\.19%252\\,\\bar\{r\}=9\.19\\%over the window\. That annualization multiplies a daily mean by 252 and does not compound, which keeps the measure additive across arms and lets the difference between two rows be read directly as a difference in annualized terms rather than in daily ones\.

### D\.3\.Sharpe Ratio

The Sharpe ratio divides the same daily series by its own volatility, which is what a desk is held to rather than the size of the return,

Sharpe=252​r¯σr,\\text\{Sharpe\}=\\sqrt\{252\}\\;\\frac\{\\bar\{r\}\}\{\\sigma\_\{r\}\},withr¯\\bar\{r\}andσr\\sigma\_\{r\}the mean and the standard deviation ofrtr\_\{t\}over the window\. No risk\-free rate is subtracted, which is standard for a long\-short book that holds no net cash position, and the252\\sqrt\{252\}scales a daily ratio to an annual one under the usual assumption that the daily returns carry no serial correlation from one session to the next\.

The Sharpe ratio and the return move together but not identically, since a variant can raise the mean while raising the dispersion by more\. For the checkpoint above, a daily mean of0\.00040\.0004against a daily standard deviation of0\.00350\.0035gives1\.651\.65\. The tables of Section[J](https://arxiv.org/html/2609.30316#A10)therefore report the three together, since a change that lifts all three at once isa different kind of evidencefrom one that lifts only the IC\.

## Appendix EStaleness on All Three Measures

Table[4](https://arxiv.org/html/2609.30316#S4.T4)reports only the information coefficient, and Table[17](https://arxiv.org/html/2609.30316#A5.T17)adds the two portfolio measures, which do not agree with it\. The disagreement is carried by one suite, since ChronoGPT supplies half the pool and is the only suite whose portfolio slopes are significant\. The portfolio also discards the magnitude the IC uses, and we therefore report the IC in the main text and treat the portfolio as a check on it\. Table[18](https://arxiv.org/html/2609.30316#A5.T18)does the same for Section[5\.3](https://arxiv.org/html/2609.30316#S5.SS3), and slopes are per year of staleness with suite fixed effects\.

The gain can also be read with the trading day as the unit, which gives up the pairing behind the paired test\. Averaging the 44 checkpoints within each date and resampling in monthly blocks returns the same\+1\.26\+1\.26, with a 95 percent interval of\[−0\.02,\+2\.43\]\[\-0\.02,\+2\.43\]andp=0\.027p=0\.027one\-sided\. Read together,a newer vintage ranks no better and may trade slightly worse, and PALM improves both\.

Table 17\.Effect of staleness on all three measures\. The IC slope is positive in every suite while the portfolio slopes fall, and the only significant portfolio slopes belong to the one suite that supplies half the pool, showing thatthe disagreement between the two kinds of score rests on a single release history\.Table 18\.The age of the base checkpoint, as regressions rather than as groups\. Neither the adapted score nor the gain over the unadapted checkpoint declines with staleness, showing thatthe grouping of Table[6](https://arxiv.org/html/2609.30316#S5.F6)survives a continuous test\.
## Appendix FEvery Eligible Checkpoint

Table[19](https://arxiv.org/html/2609.30316#A6.T19)lists the paired result for each of the 44 checkpoints the main text averages over, which lets a reader see the spread the averages are drawn from rather than only their mean\. The cutoff of a vintage is the year through which it was trained\. The four checkpoints the update does not improve sit in all three suites and at cutoffs fifteen years apart, andneither a suite nor an era accounts for them\.

Table 19\.Every eligible checkpoint, without and with PALM\. The update raises the IC on 40 of the 44, and the four it does not are spread across all three suites, showing thatthe average is not carried by a handful of checkpoints\.
## Appendix GAdapter Capacity

Figure[5](https://arxiv.org/html/2609.30316#S4.F5)plots the gain in the information coefficient against the adapter rank\. Table[20](https://arxiv.org/html/2609.30316#A7.T20)gives the same curve as numbers and adds the two portfolio measures, which agree on its shape\. Section[5\.2](https://arxiv.org/html/2609.30316#S5.SS2)then adapts the widest suite at a larger rank than the main text uses elsewhere, and Table[21](https://arxiv.org/html/2609.30316#A7.T21)is the grid on which that choice rests\. Every entry of that grid is a gain against the same checkpoint without an adapter\.

Table 20\.Effect of the adapter capacity, on all three reported metrics\. All three peak at rank three and fall away from it, and the portfolio measures turn negative beyond rank twelve while the IC stays positive, showing thatcapacity past a few directions costs the portfolio measures more than it buys them\.Table 21\.Adapter capacity against backbone width\. The best rank is four or less on the narrow suites \(ChronoGPT, ChronoGPT\-Instruct\) and sixteen on the wide ones \(DatedGPT, PIT\-4B\), and the rank we report is therefore too small for the widest suite rather than ineffective, showing thatthe capacity of the update should be read against the width of the backbone\.
## Appendix HRobustness to the Direction Words

The signal of Eq\. \([3](https://arxiv.org/html/2609.30316#S3.E3)\) contrasts two sets of continuations, and the sets the reported results use are named in Section[I](https://arxiv.org/html/2609.30316#A9)\. Those words are a choice rather than a measurement, and we therefore repeat the whole comparison with two disjoint replacements, changing nothing else\. Table[22](https://arxiv.org/html/2609.30316#A8.T22)reports the result on the 19 checkpoints for which all three wordings were run, which is a smaller pool than the main text uses because the replacements were run only on two of the three suites\.

The unadapted checkpoint is sensitive to the wording, scoring between4\.814\.81and6\.656\.65depending on which words name the direction\. The adapted checkpoint is not, landing within0\.050\.05of the same value under all three\. The gain therefore survives the replacement and is larger under both alternatives than under the wording we report, which makesthe reported setting the least favorable of the three\.

Table 22\.The same comparison read through three disjoint sets of direction words\. The unadapted score moves with the wording while the adapted score does not, showing thatthe gain belongs to the update and not to the words that read it out\.
## Appendix IA Worked Example of the Pipeline

This section follows one article through every stage of Section[4\.1](https://arxiv.org/html/2609.30316#S4.SS1), from the string the model reads to the numbers the tables report\. The checkpoint is the ChronoGPT vintage cut at the end of 2019, which is eligible for the decision date at the end of November 2020, and the article is drawn from 13 December 2021, inside the evaluation window\.

Stage 1: The prompt\.Scoring uses one template throughout\. The ticker and the article text are substituted into it, the text is truncated at 1,500 characters, and the string stops immediately before the word that would name the direction\.

> News about AAL: Why Airline Stocks Are Falling Today\. Shares of American Airlines Group \(NASDAQ: AAL\), United Airlines Holdings \(NASDAQ: UAL\), Delta Air Lines \(NYSE: DAL\), and Southwest Airlines \(NYSE: LUV\) all fell as much as 5% on Monday, a weak day for markets overall\.…\\ldots Following this news, the stock price of AAL will

The publication date never enters the prompt\. It is used outside the model to assign the article to a trading day, to define the cross\-section it is ranked within, and to select the return it is scored against\. The prompt itself therefore carriesno marker of when the article was written\.

Stage 2: The direction words\.The next\-token distribution at the final position is read once, and the probability mass on two disjoint sets of continuations is compared\. The reported results use

𝒲\+\\displaystyle\\mathcal\{W\}^\{\+\}=\{rise,increase,go up,gain\},\\displaystyle=\\\{\\text\{\\ rise\},\\ \\text\{\\ increase\},\\ \\text\{\\ go up\},\\ \\text\{\\ gain\}\\\},𝒲−\\displaystyle\\mathcal\{W\}^\{\-\}=\{fall,decrease,go down,drop\}\.\\displaystyle=\\\{\\text\{\\ fall\},\\ \\text\{\\ decrease\},\\ \\text\{\\ go down\},\\ \\text\{\\ drop\}\\\}\.Each set is compared at the first token position at which its members differ, since a tokenizer that emits the leading space separately would otherwise return the same identifier for every candidate and collapse the contrast to zero\. Substituting Eq\. \([3](https://arxiv.org/html/2609.30316#S3.E3)\) givesfθ=−1\.272f\_\{\\theta\}=\-1\.272for the article above, a negative value that places it at the bottom of its date\.

Stage 3: The cross\-section\.All 91 articles dated 13 December 2021 are ranked byfθf\_\{\\theta\}\. Table[23](https://arxiv.org/html/2609.30316#A9.T23)lists the three highest and the three lowest\. The model is right about the airline story and wrong about the dividend announcement, which is the ordinary case at this signal strength\.

Stage 4: The two daily numbers\.The Spearman correlation between the 91 scores and the 91 realized next\-session returns isICt=\+0\.0067\\mathrm\{IC\}\_\{t\}=\+0\.0067\. Taking the sign of the score demeaned within the date and holding each name at equal weight gives a long\-short return ofrt=−0\.33%r\_\{t\}=\-0\.33\\%on that day\. AveragingICt\\mathrm\{IC\}\_\{t\}over the 767 scored days of the window gives6\.616\.61in units of10−310^\{\-3\}for this checkpoint, and annualizingrtr\_\{t\}gives the return and the Sharpe ratio of the same row\.

Table 23\.One scored date, 13 December 2021, for the ChronoGPT vintage cut at the end of 2019\. Signals are compared within the date, and the individual outcomes disagree with the ordering as often as they agree, showing thatthe signal lives in the average over many dates rather than in any one of them\.Stage 5: What PALM changes\.Nothing above involves a return label or any trained parameter\. PALM leaves Stages 1 to 4 untouched and changes only the weights that producefθf\_\{\\theta\}, by fitting the adapter of Eq\. \([7](https://arxiv.org/html/2609.30316#S3.E7)\) on the text published between the end of 2019 and the decision date\. The same 91 articles are then scored again by the adapted checkpoint, and the difference between the two rows of every table is the difference between those two sets of scores\.

## Appendix JSensitivity of the Update

The main text asks each sensitivity question in the paragraph that raises it and tabulates a few rows of the answer\. This section gives the full sweep behind each of those tables, together with the questions the main text does not tabulate at all\. Every entry is a gain against the same checkpoint without an adapter\. The pool holds 44 checkpoints except in Table[25](https://arxiv.org/html/2609.30316#A10.T25), where moving the date changes which vintages are eligible\.

Adaptation schedule\.The adapter can read the eligible interval once as a single corpus, or once per calendar year in sequence\. Table[24](https://arxiv.org/html/2609.30316#A10.T24)compares the two, and the single pass is better on the IC while the two are close on the portfolio measures\. We therefore adopt the single pass, which is also the cheaper of the two\.

Table 24\.Adaptation schedule\. Reading the eligible interval once beats walking it year by year, showing thatthe update does not need to follow the order the text was published in\.Decision date\.The date we fix determines both the eligible pool and the evaluation window, and moving it changes the experiment rather than a setting within it\. Table[25](https://arxiv.org/html/2609.30316#A10.T25)runs the whole comparison at seven dates\. The gain is positive at every one of them on all three measures, and the two weakest dates are the two most recent, where the eligible pool is largest and the window is shortest\. At those two dates the share of checkpoints the update improves falls below half on the portfolio measures, so the average there is not carried by the median checkpoint: with the window down to two years and to one, a return and a Sharpe ratio are read off too few days to separate the two arms\. We therefore read this sweep as a statement about the IC, which is positive and improves most checkpoints at all seven dates, rather than as evidence that the portfolio gain is steady across them\.

Table 25\.Decision date\. The IC gain is positive at every date, showing thatthe reading of the IC does not depend on where the window sits\.Share of the eligible text\.The adapter is fitted on a fraction of the eligible corpus, keeping the slice nearest the decision date in each case\. Table[26](https://arxiv.org/html/2609.30316#A10.T26)reports six shares from a twentieth of the corpus to all of it\. A twentieth already carries three quarters of what the whole corpus delivers on the IC, and the portfolio measures are largest at the smallest share rather than at the largest\.

Table 26\.Share of the eligible text\. A small slice nearest the decision date carries most of the gain, showing thatwhat the update needs is the text closest to the date rather than the volume of it\.Training budget\.How long the adapter is trained and how fast matter more than any other choice we varied\. Table[27](https://arxiv.org/html/2609.30316#A10.T27)sweeps the number of steps and the learning rate one at a time\. At three hundred steps, or at a rate of5×10−55\\times 10^\{\-5\}, the update moves the IC by less than a third of what the reported setting moves it, and beyond twelve hundred steps the portfolio measures turn negative while the IC is still positive\. The reported setting sits in the interval where all three agree\. That setting is a single choice applied to every suite, every checkpoint and every decision date, and nothing in the paper is tuned per backbone, as Table[16](https://arxiv.org/html/2609.30316#A3.T16)states\. The sweep is therefore a description of the surface around that point rather than the search that produced it, and what it says is that the update has to be trained far enough to fit the corpus and not so far that it overwrites what the checkpoint already holds\.

Table 27\.Training budget\. The gain appears only once the adapter has been trained past a floor and fades once it is trained well beyond it, showing thatthe update has to be trained rather than merely attached\.Tokens read per article\.Each article is truncated before it reaches the adapter, and the cut fixes how much of a long article the update sees\. Table[28](https://arxiv.org/html/2609.30316#A10.T28)varies the cut over a range of sixteen\. The gain is flat across the whole range, and the shortest setting is as good as the longest\.

Table 28\.Tokens read per article\. Sixty\-four tokens deliver what a thousand do, showing thatthe update is carried by the opening of an article rather than by its length\.Random seed\.The adapter is initialized and the corpus is shuffled from a seed, and a result that moved with it would not be a result\. Table[29](https://arxiv.org/html/2609.30316#A10.T29)repeats the reported setting under seven seeds\. Every seed improves at least four checkpoints out of every five\.

Table 29\.Random seed\. Seven seeds place the IC gain between\+1\.08\+1\.08and\+1\.35\+1\.35, showing thatthe result is not an artifact of the one initialization we happened to report\.

Similar Articles

Scaling Point-in-Time Language Models

arXiv cs.CL

This paper demonstrates that scaling point-in-time language models—trained exclusively on text available up to each calendar date—can substantially narrow the performance gap with unrestricted models, enabling valid backtests and causal inference in finance and social sciences. The authors train decoder-only transformers up to 4B parameters on 1 trillion chronologically filtered tokens and release the full pipeline.

PorTAL: Portable Task Adapters for LLMs (3 minute read)

TLDR AI

PorTAL is a novel architecture that decouples task fine-tuning from specific base model weights, enabling portable task adapters that can be transferred to new models with minimal retraining. It achieves ~98% of LoRA's accuracy gain on unseen models using only half the calibration data.