Scaling Point-in-Time Language Models
Summary
This paper demonstrates that scaling point-in-time language models—trained exclusively on text available up to each calendar date—can substantially narrow the performance gap with unrestricted models, enabling valid backtests and causal inference in finance and social sciences. The authors train decoder-only transformers up to 4B parameters on 1 trillion chronologically filtered tokens and release the full pipeline.
View Cached Full Text
Cached at: 07/15/26, 04:21 AM
# Scaling Point-in-Time Language Models
Source: [https://arxiv.org/html/2607.11889](https://arxiv.org/html/2607.11889)
Bryan Kelly, Semyon Malamud, Johannes Schwab, and Teng Andrea XuBryan Kelly is at Yale School of Management, AQR Capital Management, and NBER;[www\.bryankellyacademic\.org](https://arxiv.org/html/2607.11889v1/www.bryankellyacademic.org)\. Johannes Schwab is at the Swiss Finance Institute, EPFL\. Semyon Malamud is at the Swiss Finance Institute, EPFL, and CEPR, and is a consultant to AQR\. Teng Andrea Xu is at AQR Capital Management\. Semyon Malamud gratefully acknowledges the financial support of the Swiss Finance Institute and the Swiss National Science Foundation, Grant 100018\-228042\. AQR Capital Management is a global investment management firm that may or may not apply similar investment techniques or methods of analysis as described herein\. The views expressed here are those of the authors and not necessarily those of AQR\. This work was supported by a grant from the Swiss National Supercomputing Centre \(CSCS\) under project ID lp46 on Alps\. This paper was written with assistance from Claude, an AI assistant by Anthropic\.
###### Abstract
Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences\. Point\-in\-time language models—trained exclusively on text available up to each calendar date—eliminate this leakage by construction, but existing efforts typically produce models that lag substantially behind their unconstrained counterparts\.
We show that this performance gap can be substantially narrowed through scale\. Training decoder\-only transformers with up to 4 billion parameters on 1 trillion chronologically filtered tokens from FineWeb, we construct a sequence of monthly model checkpoints spanning 2013–2024\. Across a range of common\-sense reasoning and language understanding benchmarks, our models approach the performance of leading open\-weight models of comparable size \(e\.g\., Gemma\-3\-4B and LLaMA\-7B\) trained on temporally unrestricted data, although a performance gap remains on several tasks\. Instruction fine\-tuning via LoRA further improves downstream usability\.
We release the complete pipeline—including dataset construction, training infrastructure, and evaluation code—to enable reproducible point\-in\-time language modeling and to support research applications that require strict temporal validity\.
## 1Introduction
Large language models now permeate economics, finance, and the social sciences, enabling researchers to extract nuanced signals from unstructured text\(Gentzkowet al\.,[2019](https://arxiv.org/html/2607.11889#bib.bib50); Horton,[2023](https://arxiv.org/html/2607.11889#bib.bib43); Korinek,[2023](https://arxiv.org/html/2607.11889#bib.bib45); Hoberg and Manela,[2025](https://arxiv.org/html/2607.11889#bib.bib5)\)\. Yet models trained on temporally unrestricted internet corpora inevitably absorb information that was unavailable at the time of the events they are used to study\. This*lookahead bias*—also called training leakage—can invalidate backtests, distort measures of risk and return, and undermine causal inference\(Glasserman and Lin,[2023](https://arxiv.org/html/2607.11889#bib.bib6); Lopez\-Lira and Tang,[2023](https://arxiv.org/html/2607.11889#bib.bib54); Sarkar and Vafa,[2024](https://arxiv.org/html/2607.11889#bib.bib44); Levy,[2024](https://arxiv.org/html/2607.11889#bib.bib7); Ludwiget al\.,[2025](https://arxiv.org/html/2607.11889#bib.bib8); Huanget al\.,[2026](https://arxiv.org/html/2607.11889#bib.bib1)\)\.
A natural remedy is to train*point\-in\-time*language models: models whose training data are restricted to text published on or before a given date\. Recent work has demonstrated the feasibility of this approach\(Heet al\.,[2025a](https://arxiv.org/html/2607.11889#bib.bib60),[b](https://arxiv.org/html/2607.11889#bib.bib180); Yanet al\.,[2026](https://arxiv.org/html/2607.11889#bib.bib10)\), producing chronologically consistent models \(ChronoBERT, ChronoGPT, and DatedGPT\) that outperform earlier no\-leakage baselines\. However, these models remain substantially smaller and weaker than leading open\-weight models, raising a practical question:*how much performance must one sacrifice to eliminate lookahead bias?*
We show that the answer is:*very little*\. By scaling both model size and training data—up to 4 billion parameters and 1 trillion tokens of chronologically filtered web text—we obtain point\-in\-time models whose common\-sense reasoning and language comprehension approach those of Gemma\-3\-4B\(Teamet al\.,[2025](https://arxiv.org/html/2607.11889#bib.bib37)\)and LLaMA\-7B\(Touvronet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib32)\), models trained on full, temporally unrestricted corpora\. Our results suggest that the performance gap attributed to temporal restrictions is largely a gap in scale\(Henighanet al\.,[2020](https://arxiv.org/html/2607.11889#bib.bib69); Kaplanet al\.,[2020](https://arxiv.org/html/2607.11889#bib.bib218); Hoffmannet al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib82)\), not an inherent limitation of the point\-in\-time paradigm\.
Contributions\.Our contributions are fourfold\.*First*, we advance the state of the art in point\-in\-time LLM training\. Starting from the GPT\-2\-based architecture ofJordanet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib64)\)used inHeet al\.\([2025a](https://arxiv.org/html/2607.11889#bib.bib60)\)andYanet al\.\([2026](https://arxiv.org/html/2607.11889#bib.bib10)\), we scale training data by approximately140×140\\times, increase model size from 1\.5B to 4B parameters, extend the context length from 1,536 to 2,048 tokens, and increase the embedding dimension from 768 to 4,096\. The resulting models achieve zero\-shot accuracy on standard benchmarks that is within a few percentage points of leading open\-weight models trained without temporal constraints\.*Second*, we instruction\-tune our models using LoRA\(Huet al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib23)\)and evaluate them on IFEval\(Zhouet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib17)\), a programmatically verifiable benchmark that avoids the well\-documented biases of LLM\-as\-a\-judge evaluation\(Zhenget al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib177); Liuet al\.,[2023b](https://arxiv.org/html/2607.11889#bib.bib172); Shenet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib19); Wanget al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib18)\)\.*Third*, we release the complete project code—covering dataset construction, model training, and evaluation—thereby substantially lowering the barrier to reproducible point\-in\-time LLM research\.111Project Github:[Point\-In\-Time\-LLM](https://github.com/DjoFE2021/Point-In-Time-LLM)\.*Fourth*, we assess the economic value of our models in an asset pricing application\. We construct point\-in\-time textual signals from news and use them to forecast returns and macroeconomic conditions\. This setting provides a stringent out\-of\-sample test of whether temporally consistent language models extract economically meaningful information rather than inadvertently exploiting look\-ahead bias\. We show that point\-in\-time models deliver robust predictive performance and economically significant gains relative to models trained on standard, non\-temporally constrained corpora, highlighting the importance of chronological consistency for financial applications\.
## 2Literature Review
Text\-based asset pricing\.A large body of literature studies how news text forecasts stock returns\. Early work relies on dictionary\-based or word\-count methods\(Tetlock,[2007](https://arxiv.org/html/2607.11889#bib.bib56); Tetlocket al\.,[2008](https://arxiv.org/html/2607.11889#bib.bib57)\), while supervised approaches extract sentiment signals tailored to return predictions\(Keet al\.,[2019](https://arxiv.org/html/2607.11889#bib.bib4)\)\. More recently, large language models have enabled richer representations of news content\.Chenet al\.\([2022](https://arxiv.org/html/2607.11889#bib.bib52)\)show that LLM embeddings of news articles strongly predict next\-day returns, andLopez\-Lira and Tang \([2023](https://arxiv.org/html/2607.11889#bib.bib54)\)demonstrate that prompting ChatGPT for headline sentiment yields robust trading signals\.Bybeeet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib51)\)use topic modeling on news text to construct interpretable macroeconomic factors\.Didisheimet al\.\([2026](https://arxiv.org/html/2607.11889#bib.bib2)\)build on these advances by decomposing news embeddings into predictable and surprise components, uncovering a “pure news” anomaly that exceed many previously documented market anomalies\. Crucially, they verify their findings using chronologically consistent LLMs fromHeet al\.\([2025a](https://arxiv.org/html/2607.11889#bib.bib60)\), confirming that the anomaly is not an artifact of lookahead bias—precisely the kind of application that motivates our work\.
Point\-in\-time language models\.Our work builds on and extends the point\-in\-time LLM framework introduced byHeet al\.\([2025a](https://arxiv.org/html/2607.11889#bib.bib60)\)andHeet al\.\([2025b](https://arxiv.org/html/2607.11889#bib.bib180)\)\. Those authors train ChronoBERT and ChronoGPT, a suite of chronologically consistent models with knowledge cutoffs from 1999, and demonstrate their competitive performance on NLP benchmarks and financial forecasting\. We demonstrate that this paradigm can be pushed substantially further through increased training data, model scaling, and cutting edge training techniques\. Specifically, we scale training from the original GPT\-2\-based ChronoGPT to a 4B\-parameter decoder\-only transformer trained on 1 trillion tokens, extending the context length from 1,536 to 2,048 tokens and the embedding dimension from 768 to 4,096 to match architectures such as Mistral\-7B\(Jianget al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib181)\)\.
## 3Methodology
The current literature suggests that large language models should first be pre\-trained on large\-scale, general\-purpose corpora to acquire broad linguistic and world knowledge\(Brownet al\.,[2020](https://arxiv.org/html/2607.11889#bib.bib98); Touvronet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib32)\), and subsequently fine\-tuned on instruction\-following datasets to improve their generalization and zero\-shot performance on unseen tasks\(Weiet al\.,[2021](https://arxiv.org/html/2607.11889#bib.bib24); Sanhet al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib166); Ouyanget al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib164)\)\. In this work, we closely follow the curriculum training procedure described inLambertet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib167)\), with an new initial stage devoted to producing a strong pre\-trained point\-in\-time base model\. Our primary motivation is twofold: first, to demonstrate that point\-in\-time models can achieve common\-sense reasoning and language comprehension on par with models trained on full, temporally unrestricted corpora; and second, to provide a fully reproducible training pipeline, including code and data configurations, to facilitate future research\.
Model\.We train decoder\-only large language models with 1\.5B and 4B parameters \(henceforth PIT\-1\.5B and PIT\-4B, respectively\), adapting the implementation ofJordanet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib64)\)to suit our requirements\. The architecture builds upon the GPT framework\(Radfordet al\.,[2018](https://arxiv.org/html/2607.11889#bib.bib30)\)and incorporates several recent advances in optimization and scaling, including matrix function\-based preconditioning\(Higham,[2008](https://arxiv.org/html/2607.11889#bib.bib490); Schulz,[1933](https://arxiv.org/html/2607.11889#bib.bib491)\), learning rate modernization\(Bernstein and Newhouse,[2024](https://arxiv.org/html/2607.11889#bib.bib492)\), distributed Shampoo optimization\(Guptaet al\.,[2018](https://arxiv.org/html/2607.11889#bib.bib493); Anilet al\.,[2020](https://arxiv.org/html/2607.11889#bib.bib494)\), scaling strategies\(Hägele and others,[2024](https://arxiv.org/html/2607.11889#bib.bib495)\), and architectural refinements such as value residual learning\(Zhou and others,[2024](https://arxiv.org/html/2607.11889#bib.bib496)\), as demonstrated in Gemma 2\(Team and others,[2024](https://arxiv.org/html/2607.11889#bib.bib497)\)\. The model is trained using the standard next\-token prediction objective\(Radfordet al\.,[2018](https://arxiv.org/html/2607.11889#bib.bib30)\)\. Under this autoregressive formulation, the joint probability of a sequence\(x1,…,xT\)\(x\_\{1\},\\ldots,x\_\{T\}\)factorizes as
pθ\(x1,…,xT\)=∏t=1Tpθ\(xt∣x1,…,xt−1\),p\_\{\\theta\}\(x\_\{1\},\\ldots,x\_\{T\}\)=\\prod\_\{t=1\}^\{T\}p\_\{\\theta\}\(x\_\{t\}\\mid x\_\{1\},\\ldots,x\_\{t\-1\}\),in practice, the model minimizes the negative log\-likelihood \(cross\-entropy\) of the next token:
ℒ\(θ\)=−𝔼x∼𝒟\[∑t=1Tlogpθ\(xt∣x1,…,xt−1\)\]\.\\mathcal\{L\}\(\\theta\)=\-\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\}\\left\[\\sum\_\{t=1\}^\{T\}\\log p\_\{\\theta\}\(x\_\{t\}\\mid x\_\{1\},\\ldots,x\_\{t\-1\}\)\\right\]\.The model is trained on a temporally ordered stream of tokens, directly aligning our setup with the incremental and continual learning literature\(Guptaet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib62); Keet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib22); Parmaret al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib29); Chenet al\.,[2025](https://arxiv.org/html/2607.11889#bib.bib61)\)\. We checkpoint the model weights on a monthly basis\. A complete description of the architecture and hyperparameters is provided in Appendix[A](https://arxiv.org/html/2607.11889#A1)\.
Pre\-Training\.The selection of a pre\-training corpus is a critical design decision, as an inappropriate data composition may lead to catastrophic forgetting\(McCloskey and Cohen,[1989](https://arxiv.org/html/2607.11889#bib.bib74); Kirkpatricket al\.,[2017](https://arxiv.org/html/2607.11889#bib.bib75); Luoet al\.,[2025](https://arxiv.org/html/2607.11889#bib.bib13); Kothaet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib15)\)\. We use the FineWeb dataset\(Penedoet al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib162)\), which provides a high\-quality, filtered and deduplicated snapshot of internet text spanning from 2013 to 2025\. The dataset comprises 15 trillion English tokens drawn from 96 Common Crawl\(Common Crawl,[2024](https://arxiv.org/html/2607.11889#bib.bib58)\)snapshots, processed through language identification and quality filtering pipelines\. Models trained on FineWeb have been shown to consistently outperform those trained on other publicly available web\-scale corpora\(Raffelet al\.,[2020a](https://arxiv.org/html/2607.11889#bib.bib84); Penedoet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib85); Soldainiet al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib86); Sobolevaet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib87); Ortiz Suárezet al\.,[2019](https://arxiv.org/html/2607.11889#bib.bib88); Gaoet al\.,[2020](https://arxiv.org/html/2607.11889#bib.bib89)\)across a range of language understanding, reasoning, and knowledge benchmarks\.
Instruction Fine\-Tuning\.Pre\-training with a self\-supervised next\-token prediction objective on massive unlabeled corpora induces broad linguistic and world knowledge, but does not directly optimize performance on downstream tasks\(Raffelet al\.,[2020b](https://arxiv.org/html/2607.11889#bib.bib26)\)\. Fine\-tuning addresses this limitation and serves as a central mechanism of modern transfer learning in NLP, enabling pre\-trained models to adapt to new domains, tasks, or instructions using relatively small amounts of labeled data while preserving the capabilities acquired during pre\-training\(Weiet al\.,[2021](https://arxiv.org/html/2607.11889#bib.bib24)\)\. In particular, instruction fine\-tuning adapts a pre\-trained model using a collection of datasets formatted as natural language instructions\. Extensive work has demonstrated that this procedure consistently improves generalization to unseen tasks\(Weiet al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib163); Ouyanget al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib164); Chunget al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib165); Sanhet al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib166); Lambertet al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib167)\)\. WhileLambertet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib167)\)perform a full\-parameter fine\-tune of the base model, we diverge from their approach and instead employ low\-rank adaptation \(LoRA\)\(Huet al\.,[2022](https://arxiv.org/html/2607.11889#bib.bib23)\)\. In short, given a pre\-trained weight matrixW0∈ℝdout×dinW\_\{0\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}, LoRA freezesW0W\_\{0\}and parameterizes the update as a low\-rank decomposition
W=W0\+BA,B∈ℝdout×r,A∈ℝr×din,r≪min\{dout,din\},W=W\_\{0\}\+BA,\\qquad B\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times r\},\\quad A\\in\\mathbb\{R\}^\{r\\times d\_\{\\mathrm\{in\}\}\},\\quad r\\ll\\min\\\{d\_\{\\mathrm\{out\}\},d\_\{\\mathrm\{in\}\}\\\},so that only the factorsAAandBBare learned\.222We setr=16r=16in all experiments\.We adopt LoRA over full\-parameter fine\-tuning for three reasons\. First,Bidermanet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib20)\)demonstrate that LoRA acts as a strong implicit regularizer: it adapts the model to new tasks while incurring substantially less representational drift than full fine\-tuning, thereby mitigating catastrophic forgetting of previously acquired capabilities\. Second, LoRA is significantly more compute\-efficient: with rankr=16r=16, the number of trainable parameters per layer reduces fromdin×doutd\_\{\\mathrm\{in\}\}\\times d\_\{\\mathrm\{out\}\}tor\(din\+dout\)r\(d\_\{\\mathrm\{in\}\}\+d\_\{\\mathrm\{out\}\}\), a reduction of over two orders of magnitude for typical hidden dimensions\. Finally, prior work shows that LoRA can match the performance of full\-parameter fine\-tuning within 1–2% when properly tuned, while preserving its primary advantages of reduced memory usage and faster training\(Leeet al\.,[2026](https://arxiv.org/html/2607.11889#bib.bib11); Zhaoet al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib12)\)\.
Table 1:Datasets used across training stages\. The FineWeb corpus is indexed by publication timestamp; all fine\-tuning datasets are temporally filtered to remove references to real\-world events beyond the target cutoff \(Appendix[B](https://arxiv.org/html/2607.11889#A2)\)\.StageDatasetPenedoet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib162)\)TokensFilteredPTHuggingFaceFW/fineweb170B / 1T–StageDatasetLambertet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib167)\); Xuet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib9)\)ExamplesFilteredSFTai2\-adapt\-dev/evol\_codealpaca\_heval\_decontaminated106,790100,114ai2\-adapt\-dev/personahub\_code\_v2\_3499934,94334,748ai2\-adapt\-dev/tulu\_v3\.9\_open\_math\_2\_gsm8k\_50k50,00049,772ai2\-adapt\-dev/numinamath\_tir\_math\_decontaminated64,19163,910ai2\-adapt\-dev/personahub\_ifdata\_manual\_seed\_v3\_2998029,82725,293argilla/ifeval\-like\-data456,304386,584Data\.Table[1](https://arxiv.org/html/2607.11889#S3.T1)summarizes the datasets used in each training stage\. While the FineWeb corpus is already indexed with publication timestamps allowing us to enforce a temporal cutoff directly, the remaining fine\-tuning datasets required additional curation to ensure that no instruction\-response pairs reference real\-world events beyond the target time horizon\. We describe our temporal filtering procedure in Appendix[B](https://arxiv.org/html/2607.11889#A2)\. In addition, we remove examples exceeding the model’s context length and filter out non\-English instances to maintain a consistent linguistic distribution across all training stages\. We cap the Argilla/IFEval\-like\(Xuet al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib9)\)data at 270,000 examples to maintain an approximate balance whereby half of the dataset comprises coding and mathematical problems, while the remaining half focuses on rigorous adherence to user instructions\.
Evaluation\.We rely on the widely usedGaoet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib14)\)library to evaluate both our models and the benchmark models, supporting the reproducibility of our results\. In short,Gaoet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib14)\)provides the research community with a common library that offers standardized implementations for evaluating LLM performance across a wide range of end\-to\-end benchmarks\.
## 4Results
### 4\.1Common Sense Reasoning & Language Comprehension
Figure 1:HellaSwag accuracy \(%\) over time for our PIT\-1\.5B \(170B tokens\) and PIT\-4B \(1T tokens\) models trained on chronologically ordered FineWeb data\. Horizontal lines denote published baselines: Gemma3\-4B \(77%\)Teamet al\.\([2025](https://arxiv.org/html/2607.11889#bib.bib37)\), Llama\-7B \(76%\)Touvronet al\.\([2023](https://arxiv.org/html/2607.11889#bib.bib32)\), GPT2\-XL \(50\.9%\)Radfordet al\.\([2019](https://arxiv.org/html/2607.11889#bib.bib93)\)\(as reported inWuet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib27)\)\), Gemma3\-1B \(62\.3%\)Teamet al\.\([2025](https://arxiv.org/html/2607.11889#bib.bib37)\), ChronoGPT 2024 \(44%\)Heet al\.\([2025a](https://arxiv.org/html/2607.11889#bib.bib60)\), DateGPT \(53\.2%\)Yanet al\.\([2026](https://arxiv.org/html/2607.11889#bib.bib10)\), and random guess \(25%\)\.Table 2:Zero\-shot accuracy \(%\) on standard common sense reasoning benchmarks\.FollowingTouvronet al\.\([2023](https://arxiv.org/html/2607.11889#bib.bib32)\), we first evaluate our pre\-trained models on common\-sense reasoning tasks that assess whether language models have internalized broad world knowledge and can perform implicit reasoning beyond surface\-level pattern matching\. We report zero\-shot accuracy on seven widely used benchmarks: BoolQ\(Clarket al\.,[2019](https://arxiv.org/html/2607.11889#bib.bib498)\), PIQA\(Bisket al\.,[2020](https://arxiv.org/html/2607.11889#bib.bib499)\), HellaSwag\(Zellerset al\.,[2019](https://arxiv.org/html/2607.11889#bib.bib59)\), WinoGrande\(Sakaguchiet al\.,[2021](https://arxiv.org/html/2607.11889#bib.bib501)\), ARC \(Easy and Challenge\)\(Clarket al\.,[2018](https://arxiv.org/html/2607.11889#bib.bib502)\), and OpenBookQA\(Mihaylovet al\.,[2018](https://arxiv.org/html/2607.11889#bib.bib503)\)\. All evaluations are conducted in the zero\-shot setting, following standard practice in the language modeling literature\.
Table[2](https://arxiv.org/html/2607.11889#S4.T2)reports the results\.333At the time of writing, we were unable to locate the DatedGPT weights and therefore report the numbers from the original paperYanet al\.\([2026](https://arxiv.org/html/2607.11889#bib.bib10)\)\.PIT\-4B outperforms every other point\-in\-time LLM, including its 1B\-parameter counterpart PIT\-1B\_2024, ChronoGPT\_2024\(Heet al\.,[2025a](https://arxiv.org/html/2607.11889#bib.bib60)\), and DatedGPTYanet al\.\([2026](https://arxiv.org/html/2607.11889#bib.bib10)\), with particularly large gains on HellaSwag \(\+28\.3pp and \+19pp over ChronoGPT\_2024 and DatedGPT\_2024, respectively\)\. More importantly, PIT\-4B closes much of the gap to LLaMA\-7B and Gemma\-3\-4B—models trained without any temporal restriction and, in the case of LLaMA\-7B, with nearly twice the parameter count\. On PIQA and WinoGrande, the gap to LLaMA\-7B and Gemma\-3\-4B is nearly negligible\. These results demonstrate that chronological consistency need not come at the cost of strong language understanding\.
### 4\.2Instruction Following
Figure 2:IFEval instruction\-following accuracy \(%\) for our instruction\-tuned PIT models\. Models are fine\-tuned with LoRA on chronologically consistent data\.Evaluation with IFEval\.A common approach to evaluating instruction\-tuned models is to use an LLM as a judge\(Liuet al\.,[2023a](https://arxiv.org/html/2607.11889#bib.bib174); Fuet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib168); Chiang and Lee,[2023](https://arxiv.org/html/2607.11889#bib.bib169)\)\. For instance,Heet al\.\([2025b](https://arxiv.org/html/2607.11889#bib.bib180)\)evaluate their instruction\-tuned ChronoGPT using AlpacaEval\(Duboiset al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib178)\), reporting length\-controlled win rates against Qwen1\.5\-1\.8B\-Chat\(Baiet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib179)\)\. However, LLM\-as\-a\-judge metrics are reference\-free and suffer from well\-documented biases: self\-bias, where the judge prefers its own outputs\(Liuet al\.,[2023b](https://arxiv.org/html/2607.11889#bib.bib172)\); positional bias, where the judge favors responses based on presentation order\(Wanget al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib173)\); and verbosity bias, where the judge prefers longer answers\(Zhenget al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib171)\)\.Zhenget al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib177)\)demonstrate that a “null model” outputting constant, prompt\-irrelevant responses can achieve top scores on AlpacaEval, Arena\-Hard\-Auto\(Liet al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib176)\), and MT\-Bench\(Baiet al\.,[2024](https://arxiv.org/html/2607.11889#bib.bib175)\)\. We refer the reader toYeet al\.\([2024](https://arxiv.org/html/2607.11889#bib.bib170)\)for a comprehensive discussion of these failure modes\.
To avoid these pitfalls, we evaluate instruction following using IFEval\(Zhouet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib17)\), a benchmark that tests compliance with verifiable constraints—length limits, required keywords, formatting rules—checked programmatically rather than by an LLM judge\.
Every model is evaluated using the defaultIFEvalsettings to maximize reproducibility, following the best practices of the public leaderboard444[https://huggingface\.co/spaces/open\-llm\-leaderboard/open\_llm\_leaderboard\#/](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard#/): temperature=0=0, maximum generation tokens=1280=1280, anddo\_sample=False\.
### 4\.3Economic Gains
A key motivation for point\-in\-time language models is their use in financial applications where lookahead bias can invalidate results\. FollowingChenet al\.\([2022](https://arxiv.org/html/2607.11889#bib.bib52)\)andDidisheimet al\.\([2026](https://arxiv.org/html/2607.11889#bib.bib2)\), we evaluate the economic value of our models by extracting news embeddings and using them to predict stock returns\. Because our models are trained exclusively on text available at each point in time, any return predictability we document is free of training leakage by construction\.
To quantify economic performance, we evaluate portfolios using the Sharpe ratio, the standard metric of risk\-adjusted return, which measures average return per unit of volatility\. This provides a natural analogue to downstream evaluation in machine learning: a model is useful insofar as it produces signals that translate into high Sharpe ratio portfolios out\-of\-sample\.
We constructppbase portfolios from the embeddings and denote their returns at timettby
𝑭t=\(Ft,1,…,Ft,p\)⊤∈ℝp\.\\boldsymbol\{F\}\_\{t\}=\(F\_\{t,1\},\\dots,F\_\{t,p\}\)^\{\\top\}\\in\\mathbb\{R\}^\{p\}\.We then combine these base portfolios using penalized Maximum Sharpe Ratio Regression \(MSRR\)Kelly and Xiu \([2023](https://arxiv.org/html/2607.11889#bib.bib187)\), which directly estimates portfolio weights to maximize out\-of\-sample Sharpe ratio while controlling overfitting\. Given a rolling window ofWWperiods, the weights𝝀∈ℝp\\boldsymbol\{\\lambda\}\\in\\mathbb\{R\}^\{p\}are obtained as
𝝀^\(z\)=argmin𝝀∈ℝp1W∑u=t−Wt−1\(1−𝝀⊤𝐅u\+1\)2\+z‖𝝀‖22\.\\hat\{\\boldsymbol\{\\lambda\}\}\(z\)=\\arg\\min\_\{\\boldsymbol\{\\lambda\}\\in\\mathbb\{R\}^\{p\}\}\\frac\{1\}\{W\}\\sum\_\{u=t\-W\}^\{t\-1\}\\left\(1\-\\boldsymbol\{\\lambda\}^\{\\top\}\\mathbf\{F\}\_\{u\+1\}\\right\)^\{2\}\+z\\\|\\boldsymbol\{\\lambda\}\\\|\_\{2\}^\{2\}\.\(1\)The resulting portfolio return is
rt\+1MSRR=𝝀^\(z\)⊤𝐅t\+1\.r\_\{t\+1\}^\{\\text\{MSRR\}\}=\\hat\{\\boldsymbol\{\\lambda\}\}\(z\)^\{\\top\}\\mathbf\{F\}\_\{t\+1\}\.This objective is equivalent to maximizing a regularized Sharpe ratio over the span of base portfolios\.
Embedding ConstructionLetdhd\_\{h\}denote the dimension of the final hidden representation of the language model \(e\.g\.,dh=4096d\_\{h\}=4096for PIT\-4B\)\. For an articleiiabout stocksspublished on daydd, let
𝐡d,s\(i\)∈ℝdh\\mathbf\{h\}\_\{d,s\}^\{\(i\)\}\\in\\mathbb\{R\}^\{d\_\{h\}\}be the final\-layer hidden state of the last token\. We define the article embedding as
𝐞d,s\(i\):=𝐡d,s\(i\)\.\\mathbf\{e\}\_\{d,s\}^\{\(i\)\}:=\\mathbf\{h\}\_\{d,s\}^\{\(i\)\}\.We aggregate embeddings in two steps\. First, we construct daily stock\-level embeddings:
𝐞¯d,s=1Nd,s∑i=1Nd,s𝐞d,s\(i\)\.\\bar\{\\mathbf\{e\}\}\_\{d,s\}=\\frac\{1\}\{N\_\{d,s\}\}\\sum\_\{i=1\}^\{N\_\{d,s\}\}\\mathbf\{e\}\_\{d,s\}^\{\(i\)\}\.Second, we aggregate to the monthly level:
𝐞~t,s=1\|𝒟t\|∑d∈𝒟t𝐞¯d,s,\\tilde\{\\mathbf\{e\}\}\_\{t,s\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{t\}\|\}\\sum\_\{d\\in\\mathcal\{D\}\_\{t\}\}\\bar\{\\mathbf\{e\}\}\_\{d,s\},where𝒟t\\mathcal\{D\}\_\{t\}denotes the set of trading days in monthtt\. To remove market\-wide components, we demean embeddings cross\-sectionally:
𝐞^t,s=𝐞~t,s−1\|𝒮t\|∑s′∈𝒮t𝐞~t,s′\.\\hat\{\\mathbf\{e\}\}\_\{t,s\}=\\tilde\{\\mathbf\{e\}\}\_\{t,s\}\-\\frac\{1\}\{\|\\mathcal\{S\}\_\{t\}\|\}\\sum\_\{s^\{\\prime\}\\in\\mathcal\{S\}\_\{t\}\}\\tilde\{\\mathbf\{e\}\}\_\{t,s^\{\\prime\}\}\.
We then match𝐞^t,s\\hat\{\\mathbf\{e\}\}\_\{t,s\}with next\-month returnsRt\+1,sR\_\{t\+1,s\}, yielding a panel dataset\{\(𝐞^t,s,Rt\+1,s\)\}\\\{\(\\hat\{\\mathbf\{e\}\}\_\{t,s\},R\_\{t\+1,s\}\)\\\}\.
Portfolio ConstructionWe consider two alternatives to build managed portfolios from the embedding dataset:
- •Linear portfolios\.We define 𝐅t\+1lin=∑s∈𝒮t𝐞^t,sRt\+1,s∈ℝdh\.\\mathbf\{F\}\_\{t\+1\}^\{\\mathrm\{lin\}\}=\\sum\_\{s\\in\\mathcal\{S\}\_\{t\}\}\\hat\{\\mathbf\{e\}\}\_\{t,s\}\\,R\_\{t\+1,s\}\\in\\mathbb\{R\}^\{d\_\{h\}\}\.\(2\)Each embedding dimension corresponds to one portfolio, yielding weights that are linear in characteristics as inBrandtet al\.\([2009](https://arxiv.org/html/2607.11889#bib.bib379)\)\.
- •Random\-feature portfolios\.To capture nonlinearities, we apply a random feature map: ϕ\(𝐞^t,s\)=ReLU\(𝐞^t,sΓ\),Γ∈ℝdh×P,\\phi\(\\hat\{\\mathbf\{e\}\}\_\{t,s\}\)=\\mathrm\{ReLU\}\(\\hat\{\\mathbf\{e\}\}\_\{t,s\}\\Gamma\),\\quad\\Gamma\\in\\mathbb\{R\}^\{d\_\{h\}\\times P\},withΓjk∼𝒩\(0,1\)\\Gamma\_\{jk\}\\sim\\mathcal\{N\}\(0,1\)andP=7000P=7000\. The corresponding portfolios are 𝐅t\+1RF=∑s∈𝒮tϕ\(𝐞^t,s\)Rt\+1,s∈ℝP\.\\mathbf\{F\}\_\{t\+1\}^\{\\mathrm\{RF\}\}=\\sum\_\{s\\in\\mathcal\{S\}\_\{t\}\}\\phi\(\\hat\{\\mathbf\{e\}\}\_\{t,s\}\)\\,R\_\{t\+1,s\}\\in\\mathbb\{R\}^\{P\}\.\(3\)
Data and Point\-in\-Time implementationOur Dow Jones News dataset spans the period from June 1979 to May 2020\. Because the earliest available point\-in\-time \(PIT\) language model is timestamped December 2013, we use this model to generate embeddings for all articles dated prior to 2014\. Starting in 2014, we implement a strictly out\-of\-sample embedding procedure: for each calendar yeartt, we use the PIT model available at the end of the previous year \(Decembert−1t\-1\) to embed all articles published during yeartt\. For example, the December 2013 model is used for 2014 articles, the December 2014 model for 2015 articles, and so on\. This rolling scheme ensures that embeddings are constructed using only information available at each point in time, thereby avoiding look\-ahead bias\.
Full\-sample benchmarksTo assess whether the economic gains of PIT models come at the cost of reduced signal quality, we introduce two benchmark variants based on the same architectures as PIT\-4B and PIT\-4B\-FT\. Specifically,Full\-4BandFull\-4B\-FTcorrespond to the final training checkpoints of these models, obtained after training on the full corpus\.
Because these final checkpoints are trained on all available data, including text that occurs after the prediction date, the resulting embeddings may incorporate future information and are therefore subject to look\-ahead bias\. We use these variants as benchmarks to evaluate whether enforcing point\-in\-time training leads to a loss in economic significance relative to models that exploit the full dataset\.
Estimation and EvaluationWe estimate portfolio weights using the Maximum Sharpe Ratio Regression \(MSRR\) in \([1](https://arxiv.org/html/2607.11889#S4.E1)\) on a rolling window of lengthT=360T=360months\. At each rebalancing date, we solve the MSRR problem over a grid of shrinkage parametersz∈\{10i\}i=−66z\\in\\\{10^\{i\}\\\}\_\{i=\-6\}^\{6\}, rescaled according to
zeff=pTztr\(Σ\),z\_\{\\text\{eff\}\}=\\frac\{p\}\{T\}\\,z\\,\\mathrm\{tr\}\(\\Sigma\),whereppdenotes the number of base portfolios andtr\(Σ\)\\mathrm\{tr\}\(\\Sigma\)is the trace of the covariance matrix ofFtF\_\{t\}estimated over the rolling window\. For each value ofzz, we compute the corresponding out\-of\-sample Sharpe ratio, and report results averaged across the grid\. All performance statistics are computed over the post\-December 2013 period, which constitutes the true out\-of\-sample evaluation\. We also report results by constructing portfolios on subsamples of stocks sorted by market capitalization\. FollowingJensenet al\.\([2023](https://arxiv.org/html/2607.11889#bib.bib205)\), we classify stocks into mega, large, small, and micro categories using the 80th, 50th, 20th, and 1st percentiles of NYSE market capitalization, respectively\. Stocks below the 1st percentile are excluded\.
Figure[3](https://arxiv.org/html/2607.11889#S4.F3)reports out\-of\-sample annualized Sharpe ratios of the MSRR portfolios across model variants, portfolio constructions, and stock size groups\. Three main findings emerge:
1. 1\.Point\-in\-time models produce economically meaningful signalseven under strict temporal constraints\. Across both portfolio constructions, PIT models generate positive Sharpe ratios in multiple segments, and perform comparably to their full\-sample counterparts \(Full\-4B and Full\-4B\-FT\)\. This indicates that return predictability from text is not solely an artifact of look\-ahead bias\. Instead, these results suggest that news embeddings contain genuine forward\-looking information that persists even when training leakage is eliminated\.
2. 2\.Scaling substantially improves performance, with economically large gains\. In the random\-feature specification for mega\-cap stocks, Sharpe ratios increase from 0\.56 and 0\.59 for ChronoGPT\-base and ChronoGPT\-instruct to 0\.80 and 0\.79 for PIT\-4B and PIT\-4B\-FT, representing improvements of roughly 35–45%\. In the linear specification, the gains are even larger: mega\-cap Sharpe ratios rise from 0\.18 and 0\.24 to 0\.78 and 0\.49, corresponding to increases of up to 3\-4×\\times\. Similar patterns hold across other size groups, indicating that scaling consistently enhances the economic value of the extracted signals\.
3. 3\.Fine\-tuning has heterogeneous effects: while instruction tuning preserves strong performance, it does not uniformly dominate the base PIT\-4B model across all segments\.
Overall, these results provide strong evidence that strict temporal validity does not eliminate economically useful information, and that scaling plays a central role in recovering and amplifying that information\.
Figure 3:Average out\-of\-sample \(post 2013\-12\) annualized Sharpe ratio across the regularization grid for each model, reported by size group\. Results are based on a rolling 360\-month training window\.Linearuses raw embeddings \([2](https://arxiv.org/html/2607.11889#S4.E2)\);Random Featuresapplies a ReLU random projection toP=7,000P\{=\}7\{,\}000features \([3](https://arxiv.org/html/2607.11889#S4.E3)\) prior to MSRR \([1](https://arxiv.org/html/2607.11889#S4.E1)\)\.
## 5Conclusion
This paper examines whether strict temporal validity in language model training necessarily entails a significant performance sacrifice\. Our results suggest that it does not\. By scaling point\-in\-time pre\-training to 1 trillion chronologically filtered tokens and 4 billion parameters, we obtain models that substantially outperform prior chronologically consistent baselines and approach the performance of strong open\-weight models trained on temporally unrestricted corpora\. The main message is that much of the apparent cost of eliminating lookahead bias reflects a scale deficit rather than an inherent limitation of the point\-in\-time paradigm\.
We show this both in standard NLP evaluations and in a financially meaningful downstream application\. On common\-sense reasoning and language comprehension benchmarks, PIT\-4B closes a large share of the gap to Gemma\-3\-4B and LLaMA\-7B despite operating under a strict chronological constraint\. After instruction fine\-tuning with LoRA on temporally filtered data, the model also exhibits materially improved instruction\-following performance on IFEval, indicating that temporal consistency can be preserved throughout the post\-training pipeline\. In an asset\-pricing application, embeddings extracted from our point\-in\-time models generate economically meaningful out\-of\-sample return predictability, with larger models delivering stronger Sharpe ratios, especially in information\-rich segments such as mega\-cap stocks\. These findings highlight the practical value of point\-in\-time models in empirical settings where leakage can otherwise invalidate inference\.
More broadly, our results suggest that point\-in\-time language modeling is now a viable foundation for research in finance, economics, and the social sciences\. Researchers no longer need to choose between temporal validity and competitive performance in language modeling\. By releasing the full training, filtering, and evaluation pipeline, we aim to make chronologically consistent language modeling reproducible and easier to adopt in applications where causal interpretation, historical fidelity, and backtest integrity are essential\.
Several limitations remain\. First, although scaling greatly narrows the gap to unrestricted models, it does not eliminate it completely, especially on some reasoning\-intensive tasks\. Second, our asset\-pricing application focuses on a single benchmark dataset and a single portfolio construction framework; broader evidence across tasks and domains remains an important direction for future work\. Third, while our fine\-tuning data are aggressively filtered to mitigate temporal leakage, stronger methods for temporally robust post\-training and preference alignment warrant further study\. Finally, our current evaluation is incomplete for the most recent period, and updated results for 2022–2025 will be reported in a subsequent version\.
Overall, the evidence in this paper points to a simple conclusion: point\-in\-time LLMs can be made both credible and useful at a modern scale\. This opens the door to a new generation of temporally grounded models for scientific measurement, historical analysis, and decision\-making under genuine information constraints\.
## References
- Scalable second order optimization for deep learning\.arXiv preprint arXiv:2002\.09018\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- G\. Bai, J\. Liu, X\. Bu, Y\. He, J\. Liu, Z\. Zhou, Z\. Lin, W\. Su, T\. Ge, B\. Zheng,et al\.\(2024\)Mt\-bench\-101: a fine\-grained benchmark for evaluating large language models in multi\-turn dialogues\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7421–7454\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- J\. Bai, S\. Bai, Y\. Chu, Z\. Cui, K\. Dang, X\. Deng, Y\. Fan, W\. Ge, Y\. Han, F\. Huang, B\. Hui, L\. Ji, M\. Li, J\. Lin, R\. Lin, D\. Liu, G\. Liu, C\. Lu, K\. Lu, J\. Ma, R\. Men, X\. Ren, X\. Ren, C\. Tan, S\. Tan, J\. Tu, P\. Wang, S\. Wang, W\. Wang, S\. Wu, B\. Xu, J\. Xu, A\. Yang, H\. Yang, J\. Yang, S\. Yang, Y\. Yao, B\. Yu, H\. Yuan, Z\. Yuan, J\. Zhang, X\. Zhang, Y\. Zhang, Z\. Zhang, C\. Zhou, J\. Zhou, X\. Zhou, and T\. Zhu \(2023\)Qwen technical report\.External Links:2309\.16609,[Link](https://arxiv.org/abs/2309.16609)Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- J\. Bernstein and L\. Newhouse \(2024\)Old optimizer, new norm: an anthology\.arXiv preprint arXiv:2409\.20325\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- D\. Biderman, J\. Portes, J\. J\. G\. Ortiz, M\. Paul, P\. Greengard, C\. Jennings, D\. King, S\. Havens, V\. Chiley, J\. Frankle,et al\.\(2024\)Lora learns less and forgets less\.arXiv preprint arXiv:2405\.09673\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p4.7)\.
- Y\. Bisk, R\. Zellers, R\. L\. Bras, J\. Gao, and Y\. Choi \(2020\)PIQA: reasoning about physical commonsense in natural language\.InAAAI,Cited by:[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p1.1)\.
- M\. W\. Brandt, P\. Santa\-Clara, and R\. Valkanov \(2009\)Parametric portfolio policies: exploiting characteristics in the cross\-section of equity returns\.The Review of Financial Studies22\(9\),pp\. 3411–3447\.Cited by:[1st item](https://arxiv.org/html/2607.11889#S4.I1.i1.p1.2)\.
- T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p1.1)\.
- L\. Bybee, B\. Kelly, A\. Manela, and D\. Xiu \(2024\)Business news and business cycles\.The Journal of Finance79\(5\),pp\. 3105–3147\.Cited by:[§2](https://arxiv.org/html/2607.11889#S2.p1.1)\.
- J\. Chen, Z\. Chen, J\. Wang, K\. Zhou, Y\. Zhu, J\. Jiang, Y\. Min, W\. X\. Zhao, Z\. Dou, J\. Mao,et al\.\(2025\)Towards effective and efficient continual pre\-training of large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5779–5795\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.3)\.
- Y\. Chen, B\. T\. Kelly, and D\. Xiu \(2022\)Expected returns and large language models\.Available at SSRN 4416687\.Cited by:[§2](https://arxiv.org/html/2607.11889#S2.p1.1),[§4\.3](https://arxiv.org/html/2607.11889#S4.SS3.p1.1)\.
- C\. Chiang and H\. Lee \(2023\)Can large language models be an alternative to human evaluations?\.arXiv preprint arXiv:2305\.01937\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- H\. W\. Chung, L\. Hou, S\. Longpre, B\. Zoph, Y\. Tay, W\. Fedus, Y\. Li, X\. Wang, M\. Dehghani, S\. Brahma, A\. Webson, S\. S\. Gu, Z\. Dai, M\. Suzgun, X\. Chen, A\. Chowdhery, A\. Castro\-Ros, M\. Pellat, K\. Robinson, D\. Valter, S\. Narang, G\. Mishra, A\. Yu, V\. Zhao, Y\. Huang, A\. Dai, H\. Yu, S\. Petrov, E\. H\. Chi, J\. Dean, J\. Devlin, A\. Roberts, D\. Zhou, Q\. V\. Le, and J\. Wei \(2022\)Scaling instruction\-finetuned language models\.External Links:2210\.11416,[Link](https://arxiv.org/abs/2210.11416)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- C\. Clark, K\. Lee, M\. Chang, T\. Kwiatkowski, M\. Collins, and K\. Toutanova \(2019\)BoolQ: exploring the surprising difficulty of natural yes/no questions\.NAACL\-HLT\.Cited by:[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p1.1)\.
- P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, O\. Tafjord, P\. Turney, and D\. Khashabi \(2018\)Think you have solved question answering? try arc, the ai2 reasoning challenge\.InArXiv preprint arXiv:1803\.05457,Cited by:[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p1.1)\.
- Common Crawl \(2024\)Common crawl corpus\.Note:[https://commoncrawl\.org](https://commoncrawl.org/)Accessed: January, 2026Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- A\. Didisheim, B\. Kelly, M\. Pourmohammadi, and H\. Tian \(2026\)The inefficient pricing of news shocks\.Working Paper\.Cited by:[§2](https://arxiv.org/html/2607.11889#S2.p1.1),[§4\.3](https://arxiv.org/html/2607.11889#S4.SS3.p1.1)\.
- Y\. Dubois, B\. Galambosi, P\. Liang, and T\. B\. Hashimoto \(2024\)Length\-controlled alpacaeval: a simple way to debias automatic evaluators\.arXiv preprint arXiv:2404\.04475\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- J\. Fu, S\. Ng, Z\. Jiang, and P\. Liu \(2023\)Gptscore: evaluate as you desire\.arXiv preprint arXiv:2302\.04166\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- L\. Gao, S\. Biderman, S\. Black, L\. Golding, T\. Hoppe, C\. Foster, J\. Phang, H\. He, A\. Thite, N\. Nabeshima, S\. Presser, and C\. Leahy \(2020\)The pile: an 800gb dataset of diverse text for language modeling\.External Links:2101\.00027,[Link](https://arxiv.org/abs/2101.00027)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- L\. Gao, J\. Tow, B\. Abbasi, S\. Biderman, S\. Black, A\. DiPofi, C\. Foster, L\. Golding, J\. Hsu, A\. Le Noac’h, H\. Li, K\. McDonell, N\. Muennighoff, C\. Ociepa, J\. Phang, L\. Reynolds, H\. Schoelkopf, A\. Skowron, L\. Sutawika, E\. Tang, A\. Thite, B\. Wang, K\. Wang, and A\. Zou \(2024\)The language model evaluation harness\.Zenodo\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.12608602),[Link](https://zenodo.org/records/12608602)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p6.1)\.
- M\. Gentzkow, B\. Kelly, and M\. Taddy \(2019\)Text as data\.Journal of Economic Literature57\(3\),pp\. 535–574\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- P\. Glasserman and C\. Lin \(2023\)Assessing look\-ahead bias in stock return predictions generated by GPT sentiment analysis\.arXiv preprint arXiv:2309\.17322\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- K\. Gupta, B\. Thérien, A\. Ibrahim, M\. L\. Richter, Q\. Anthony, E\. Belilovsky, I\. Rish, and T\. Lesort \(2023\)Continual pre\-training of large language models: how to \(re\) warm your model?\.arXiv preprint arXiv:2308\.04014\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.3)\.
- V\. Gupta, T\. Koren, and Y\. Singer \(2018\)Shampoo: preconditioned stochastic tensor optimization\.InProceedings of the 35th International Conference on Machine Learning,pp\. 1842–1850\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- A\. Hägeleet al\.\(2024\)Scaling laws and compute\-optimal training beyond fixed training durations\.arXiv preprint arXiv:2405\.18392\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- S\. He, L\. Lv, A\. Manela, and J\. Wu \(2025a\)Chronologically consistent large language models\.arXiv preprint arXiv:2502\.21206\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p2.1),[§1](https://arxiv.org/html/2607.11889#S1.p4.1),[§2](https://arxiv.org/html/2607.11889#S2.p1.1),[§2](https://arxiv.org/html/2607.11889#S2.p2.1),[Figure 1](https://arxiv.org/html/2607.11889#S4.F1),[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p2.1)\.
- S\. He, L\. Lv, A\. Manela, and J\. Wu \(2025b\)Instruction tuning chronologically consistent language models\.arXiv preprint arXiv:2510\.11677\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p2.1),[§2](https://arxiv.org/html/2607.11889#S2.p2.1),[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- T\. Henighan, J\. Kaplan, M\. Katz, M\. Chen, C\. Hesse, J\. Jackson, H\. Jun, T\. B\. Brown, P\. Dhariwal, S\. Gray,et al\.\(2020\)Scaling laws for autoregressive generative modeling\.arXiv preprint arXiv:2010\.14701\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p3.1)\.
- N\. J\. Higham \(2008\)Functions of matrices: theory and computation\.Society for Industrial and Applied Mathematics,Philadelphia, PA\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- G\. Hoberg and A\. Manela \(2025\)The natural language of finance\.Foundations and Trends in Finance\.Note:SSRN 5119322Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. d\. L\. Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark,et al\.\(2022\)Training compute\-optimal large language models\.arXiv preprint arXiv:2203\.15556\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p3.1)\.
- J\. J\. Horton \(2023\)Large language models as simulated economic agents: what can we learn from homo silicus?\.Technical reportNational Bureau of Economic Research\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.\(2022\)Lora: low\-rank adaptation of large language models\.\.Iclr1\(2\),pp\. 3\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p4.1),[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- W\. Huang, A\. J\. Menkveld, and S\. Yu \(2026\)AI “errors”\.Working Paper\.Note:SSRN 6408138Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- T\. I\. Jensen, B\. Kelly, and L\. H\. Pedersen \(2023\)Is there a replication crisis in finance?\.The Journal of Finance78\(5\),pp\. 2465–2518\.Cited by:[§4\.3](https://arxiv.org/html/2607.11889#S4.SS3.p9.6)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. Sayed \(2023\)Mistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§2](https://arxiv.org/html/2607.11889#S2.p2.1)\.
- K\. Jordan, J\. Bernstein, B\. Rappazzo, @fernbear\.bsky\.social, B\. Vlado, Y\. Jiacheng, F\. Cesista, B\. Koszarsky, and @Grad62304977 \(2024\)Modded\-nanogpt: speedrunning the nanogpt baseline\.External Links:[Link](https://github.com/KellerJordan/modded-nanogpt)Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p4.1),[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei \(2020\)Scaling laws for neural language models\.arXiv preprint arXiv:2001\.08361\.External Links:[Link](https://arxiv.org/abs/2001.08361)Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p3.1)\.
- Z\. T\. Ke, B\. T\. Kelly, and D\. Xiu \(2019\)Predicting returns with text data\.NBER Working Paper 26186\.Cited by:[§2](https://arxiv.org/html/2607.11889#S2.p1.1)\.
- Z\. Ke, Y\. Shao, H\. Lin, T\. Konishi, G\. Kim, and B\. Liu \(2023\)Continual pre\-training of language models\.arXiv preprint arXiv:2302\.03241\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.3)\.
- B\. Kelly and D\. Xiu \(2023\)Financial machine learning\.Foundations and Trends® in Finance13\(3\-4\),pp\. 205–363\.Cited by:[§4\.3](https://arxiv.org/html/2607.11889#S4.SS3.p3.4)\.
- J\. Kirkpatrick, R\. Pascanu, N\. Rabinowitz, J\. Veness, G\. Desjardins, A\. A\. Rusu, K\. Milan, J\. Quan, T\. Ramalho, A\. Grabska\-Barwinska,et al\.\(2017\)Overcoming catastrophic forgetting in neural networks\.Proceedings of the national academy of sciences114\(13\),pp\. 3521–3526\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- A\. Korinek \(2023\)Language models and cognitive automation for economic research\.Technical reportnational Bureau of economic Research\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- S\. Kotha, J\. M\. Springer, and A\. Raghunathan \(2023\)Understanding catastrophic forgetting in language models via implicit inference\.arXiv preprint arXiv:2309\.10105\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- N\. Lambert, J\. Morrison, V\. Pyatkin, S\. Huang, H\. Ivison, F\. Brahman, L\. J\. V\. Miranda, A\. Liu, N\. Dziri, S\. Lyu,et al\.\(2024\)Tulu 3: pushing frontiers in open language model post\-training\.arXiv preprint arXiv:2411\.15124\.Cited by:[Table 1](https://arxiv.org/html/2607.11889#S3.T1.1.3.3.2),[§3](https://arxiv.org/html/2607.11889#S3.p1.1),[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- Y\. Lee, C\. Ko, P\. Chen, and M\. Yeh \(2026\)Learning rate matters: vanilla lora may suffice for llm fine\-tuning\.arXiv preprint arXiv:2602\.04998\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p4.7)\.
- B\. Levy \(2024\)Caution ahead: numerical reasoning and look\-ahead bias in AI models\.Fama\-Miller Working Paper\.Note:SSRN 5082861Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- T\. Li, W\. Chiang, E\. Frick, L\. Dunlap, T\. Wu, B\. Zhu, J\. E\. Gonzalez, and I\. Stoica \(2024\)From crowdsourced data to high\-quality benchmarks: arena\-hard and benchbuilder pipeline\.arXiv preprint arXiv:2406\.11939\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. Zhu \(2023a\)G\-eval: nlg evaluation using gpt\-4 with better human alignment\.arXiv preprint arXiv:2303\.16634\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- Y\. Liu, N\. S\. Moosavi, and C\. Lin \(2023b\)LLMs as narcissistic evaluators: when ego inflates evaluation scores\.arXiv preprint arXiv:2311\.09766\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p4.1),[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- A\. Lopez\-Lira and Y\. Tang \(2023\)Can chatgpt forecast stock price movements? return predictability and large language models\.arXiv preprint arXiv:2304\.07619\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1),[§2](https://arxiv.org/html/2607.11889#S2.p1.1)\.
- J\. Ludwig, S\. Mullainathan, and A\. Rambachan \(2025\)Large language models: an applied econometric framework\.NBER Working Paper 33344\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- Y\. Luo, Z\. Yang, F\. Meng, Y\. Li, J\. Zhou, and Y\. Zhang \(2025\)An empirical study of catastrophic forgetting in large language models during continual fine\-tuning\.IEEE Transactions on Audio, Speech and Language Processing\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- M\. McCloskey and N\. J\. Cohen \(1989\)Catastrophic interference in connectionist networks: the sequential learning problem\.InPsychology of learning and motivation,Vol\.24,pp\. 109–165\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- T\. Mihaylov, P\. Clark, T\. Khot, and A\. Sabharwal \(2018\)Can a suit of armor conduct electricity? a new dataset for open book question answering\.InEMNLP,Cited by:[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p1.1)\.
- P\. J\. Ortiz Suárez, B\. Sagot, and L\. Romary \(2019\)Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures\.In7th Workshop on the Challenges in the Management of Large Corpora \(CMLC\-7\),Cardiff, United Kingdom\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. Christiano, J\. Leike, and R\. Lowe \(2022\)Training language models to follow instructions with human feedback\.External Links:2203\.02155,[Link](https://arxiv.org/abs/2203.02155)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p1.1),[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- J\. Parmar, S\. Satheesh, M\. Patwary, M\. Shoeybi, and B\. Catanzaro \(2024\)Reuse, don’t retrain: a recipe for continued pretraining of language models\.arXiv preprint arXiv:2407\.07263\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.3)\.
- G\. Penedo, H\. Kydlíček, A\. Lozhkov, M\. Mitchell, C\. A\. Raffel, L\. Von Werra, T\. Wolf,et al\.\(2024\)The fineweb datasets: decanting the web for the finest text data at scale\.Advances in Neural Information Processing Systems37,pp\. 30811–30849\.Cited by:[Table 1](https://arxiv.org/html/2607.11889#S3.T1.1.1.1.2),[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- G\. Penedo, Q\. Malartic, D\. Hesslow, R\. Cojocaru, A\. Cappelli, H\. Alobeidli, B\. Pannier, E\. Almazrouei, and J\. Launay \(2023\)The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only\.External Links:2306\.01116,[Link](https://arxiv.org/abs/2306.01116)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- A\. Radford, K\. Narasimhan, T\. Salimans, I\. Sutskever,et al\.\(2018\)Improving language understanding by generative pre\-training\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, I\. Sutskever,et al\.\(2019\)Language models are unsupervised multitask learners\.OpenAI blog1\(8\),pp\. 9\.Cited by:[Figure 1](https://arxiv.org/html/2607.11889#S4.F1)\.
- C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. Liu \(2020a\)Exploring the limits of transfer learning with a unified text\-to\-text transformer\.J\. Mach\. Learn\. Res\.21,pp\. 140:1–140:67\.External Links:[Link](https://jmlr.org/papers/v21/20-074.html)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. Liu \(2020b\)Exploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of machine learning research21\(140\),pp\. 1–67\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- K\. Sakaguchi, R\. L\. Bras, C\. Bhagavatula, and Y\. Choi \(2021\)WinoGrande: an adversarial winograd schema challenge at scale\.Communications of the ACM\.Cited by:[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p1.1)\.
- V\. Sanh, A\. Webson, C\. Raffel, S\. H\. Bach, L\. Sutawika, Z\. Alyafeai, A\. Chaffin, A\. Stiegler, T\. L\. Scao, A\. Raja, M\. Dey, M\. S\. Bari, C\. Xu, U\. Thakker, S\. S\. Sharma, E\. Szczechla, T\. Kim, G\. Chhablani, N\. Nayak, D\. Datta, J\. Chang, M\. T\. Jiang, H\. Wang, M\. Manica, S\. Shen, Z\. X\. Yong, H\. Pandey, R\. Bawden, T\. Wang, T\. Neeraj, J\. Rozen, A\. Sharma, A\. Santilli, T\. Fevry, J\. A\. Fries, R\. Teehan, T\. Bers, S\. Biderman, L\. Gao, T\. Wolf, and A\. M\. Rush \(2022\)Multitask prompted training enables zero\-shot task generalization\.External Links:2110\.08207,[Link](https://arxiv.org/abs/2110.08207)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p1.1),[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- S\. K\. Sarkar and K\. Vafa \(2024\)Lookahead bias in pretrained language models\.InICML 2025 Workshop on Reliable and Responsible Foundation Models,Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p1.1)\.
- G\. Schulz \(1933\)Iterative berechnung der reziproken matrix\.Zeitschrift für Angewandte Mathematik und Mechanik13,pp\. 57–59\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- C\. Shen, L\. Cheng, X\. Nguyen, Y\. You, and L\. Bing \(2023\)Large language models are not yet human\-level evaluators for abstractive summarization\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 4215–4233\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p4.1)\.
- D\. Soboleva, F\. Al\-Khateeb, R\. Myers, J\. R\. Steeves, J\. Hestness, and N\. Dey \(2023\)SlimPajama: a 627b token cleaned and deduplicated version of redpajama\.Note:[https://www\.cerebras\.ai/blog/slimpajama\-a\-627b\-token\-cleaned\-and\-deduplicated\-version\-of\-redpajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama)Accessed: 2026\-02\-24Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- L\. Soldaini, R\. Kinney, A\. Bhagia, D\. Schwenk, D\. Atkinson, R\. Authur, B\. Bogin, K\. Chandu, J\. Dumas, Y\. Elazar, V\. Hofmann, A\. Jha, S\. Kumar, L\. Lucy, X\. Lyu, N\. Lambert, I\. Magnusson, J\. Morrison, N\. Muennighoff, A\. Naik, C\. Nam, M\. Peters, A\. Ravichander, K\. Richardson, Z\. Shen, E\. Strubell, N\. Subramani, O\. Tafjord, E\. Walsh, L\. Zettlemoyer, N\. Smith, H\. Hajishirzi, I\. Beltagy, D\. Groeneveld, J\. Dodge, and K\. Lo \(2024\)Dolma: an open corpus of three trillion tokens for language model pretraining research\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 15725–15788\.External Links:[Link](https://aclanthology.org/2024.acl-long.840/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.840)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p3.1)\.
- G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière,et al\.\(2025\)Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p3.1),[Figure 1](https://arxiv.org/html/2607.11889#S4.F1)\.
- G\. Teamet al\.\(2024\)Gemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
- P\. C\. Tetlock, M\. Saar\-Tsechansky, and S\. Macskassy \(2008\)More than words: quantifying language to measure firms’ fundamentals\.The journal of finance63\(3\),pp\. 1437–1467\.Cited by:[§2](https://arxiv.org/html/2607.11889#S2.p1.1)\.
- P\. C\. Tetlock \(2007\)Giving content to investor sentiment: the role of media in the stock market\.The Journal of finance62\(3\),pp\. 1139–1168\.Cited by:[§2](https://arxiv.org/html/2607.11889#S2.p1.1)\.
- H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar,et al\.\(2023\)Llama: open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p3.1),[§3](https://arxiv.org/html/2607.11889#S3.p1.1),[Figure 1](https://arxiv.org/html/2607.11889#S4.F1),[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p1.1)\.
- P\. Wang, L\. Li, L\. Chen, Z\. Cai, D\. Zhu, B\. Lin, Y\. Cao, L\. Kong, Q\. Liu, T\. Liu,et al\.\(2024\)Large language models are not fair evaluators\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9440–9450\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p4.1)\.
- P\. Wang, L\. Li, L\. Chen, Z\. Cai, D\. Zhu, B\. Lin, Y\. Cao, Q\. Liu, T\. Liu, and Z\. Sui \(2023\)Large language models are not fair evaluators\.arXiv preprint arXiv:2305\.17926\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- J\. Wei, M\. Bosma, V\. Y\. Zhao, K\. Guu, A\. W\. Yu, B\. Lester, N\. Du, A\. M\. Dai, and Q\. V\. Le \(2021\)Finetuned language models are zero\-shot learners\.arXiv preprint arXiv:2109\.01652\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p1.1),[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- J\. Wei, M\. Bosma, V\. Y\. Zhao, K\. Guu, A\. W\. Yu, B\. Lester, N\. Du, A\. M\. Dai, and Q\. V\. Le \(2022\)Finetuned language models are zero\-shot learners\.External Links:2109\.01652,[Link](https://arxiv.org/abs/2109.01652)Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p4.2)\.
- M\. Wu, A\. Waheed, C\. Zhang, M\. Abdul\-Mageed, and A\. F\. Aji \(2024\)Lamini\-lm: a diverse herd of distilled models from large\-scale instructions\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 944–964\.Cited by:[Figure 1](https://arxiv.org/html/2607.11889#S4.F1)\.
- Z\. Xu, F\. Jiang, L\. Niu, Y\. Deng, R\. Poovendran, Y\. Choi, and B\. Y\. Lin \(2024\)Magpie: alignment data synthesis from scratch by prompting aligned llms with nothing\.arXiv preprint arXiv:2406\.08464\.Cited by:[Table 1](https://arxiv.org/html/2607.11889#S3.T1.1.3.3.2),[§3](https://arxiv.org/html/2607.11889#S3.p5.1)\.
- Y\. Yan, R\. Tang, Z\. Gao, W\. Jiang, and Y\. Lu \(2026\)DatedGPT: preventing lookahead bias in large language models with time\-aware pretraining\.arXiv preprint arXiv:2603\.11838\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p2.1),[§1](https://arxiv.org/html/2607.11889#S1.p4.1),[Figure 1](https://arxiv.org/html/2607.11889#S4.F1),[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p2.1),[footnote 3](https://arxiv.org/html/2607.11889#footnote3)\.
- J\. Ye, Y\. Wang, Y\. Huang, D\. Chen, Q\. Zhang, N\. Moniz, T\. Gao, W\. Geyer, C\. Huang, P\. Chen,et al\.\(2024\)Justice or prejudice? quantifying biases in llm\-as\-a\-judge\.arXiv preprint arXiv:2410\.02736\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. Choi \(2019\)Hellaswag: can a machine really finish your sentence?\.arXiv preprint arXiv:1905\.07830\.Cited by:[§4\.1](https://arxiv.org/html/2607.11889#S4.SS1.p1.1)\.
- J\. Zhao, T\. Wang, W\. Abid, G\. Angus, A\. Garg, J\. Kinnison, A\. Sherstinsky, P\. Molino, T\. Addair, and D\. Rishi \(2024\)Lora land: 310 fine\-tuned llms that rival gpt\-4, a technical report\.arXiv preprint arXiv:2405\.00732\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p4.7)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- X\. Zheng, T\. Pang, C\. Du, Q\. Liu, J\. Jiang, and M\. Lin \(2024\)Cheating automatic llm benchmarks: null models achieve high win rates\.arXiv preprint arXiv:2410\.07137\.Cited by:[§1](https://arxiv.org/html/2607.11889#S1.p4.1),[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p1.1)\.
- J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. Hou \(2023\)Instruction\-following evaluation for large language models\.arXiv preprint arXiv:2311\.07911\.Cited by:[Appendix B](https://arxiv.org/html/2607.11889#A2.p1.1),[§1](https://arxiv.org/html/2607.11889#S1.p4.1),[§4\.2](https://arxiv.org/html/2607.11889#S4.SS2.p2.1)\.
- Z\. Zhouet al\.\(2024\)Value residual learning for alleviating attention concentration in transformers\.arXiv preprint arXiv:2410\.17897\.Cited by:[§3](https://arxiv.org/html/2607.11889#S3.p2.1)\.
## Appendix ATraining Configuration
We train two GPT\-2\-style decoder\-only transformer models on chronologically ordered FineWeb data\. PIT\-1\.5B has 52 layers, 12 attention heads, and an embedding dimension of 1,536 \(head dim = 128\), totaling approximately 1\.5B parameters trained on 170B tokens\. PIT\-4B has 20 layers, 32 attention heads, and an embedding dimension of 4,096 \(head dim = 128\), totaling approximately 4\.2B parameters trained on 1T tokens\. Both models use a vocabulary of 50,304 tokens with a sequence length of 2,048\. Training employs a hybrid optimizer: Muon \(momentum = 0\.95\) for the transformer blocks and AdamW \(β1\\beta\_\{1\}= 0\.9,β2\\beta\_\{2\}= 0\.95\) for the language model head, with a peak learning rate of 0\.0009 and no weight decay\. The learning rate follows a trapezoidal schedule with linear warmdown over the final 50% of training steps\.
## Appendix BRemoving Temporal Leakage
We employ gpt\-5\-nano to identify and remove temporally sensitive examples from the datasets used in the supervised fine\-tuning stage\. We present each example from the SFT and DPO datasets to gpt\-5\-nano, prepending the prompt illustrated in Figure[4](https://arxiv.org/html/2607.11889#A2.F4), and retain only those examples classified as “timeless”—namely, examples that do not reference real\-world events, real individuals, specific dates, pop culture, or trending topics\. This filtering step is designed to ensure that the training data remains invariant to temporal context, thereby preventing the model from internalizing ephemeral or potentially outdated knowledge during fine\-tuning\. Notably, we observe that the majority of examples in both datasets are naturally retained through this process, as instruction\-following data predominantly comprises verifiable format constraints rather than world knowledge\[Zhouet al\.,[2023](https://arxiv.org/html/2607.11889#bib.bib17)\], while mathematical and coding problems are inherently grounded in abstract reasoning, formal logic, and counterfactual scenarios that are not subject to temporal drift\.
You are a binary classifier\. Determine if text contains TIME\-SENSITIVE FACTS that could become outdated or incorrect over time\.
Output ONLY "0" \(timeless\) or "1" \(time\-aware\)\.
Output "1" ONLY if the text states facts that could BECOME OUTDATED, such as: \-\- Who currently holds a political office \("Biden is the president"\) \-\- Recent or dated events \("The 2024 Olympics were held in Paris"\) \-\- Current statistics, prices, or rankings \("Tesla stock is at $X"\) \-\- A real person doing something specific at a specific time \("Elon Musk announced X in 2024"\) \-\- Laws, policies, or regulations tied to a specific time frame
Output "0" if the text is: \-\- Math problems or solutions \(even with character names like "Alice" or "John"\) \-\- Programming/coding tasks or tutorials \(even mentioning real tools: Python, Java, NetBeans, Eclipse, AWS, Docker, etc\.\) \-\- General educational content about real software, frameworks, or technologies \-\- Generic advice or best practices \(even about real products\) \-\- Hypothetical scenarios \(even with human names\) \-\- Instruction\-following tasks \-\- Scientific facts that don’t change \("water boils at 100C"\) \-\- Abstract reasoning or logic puzzles
KEY DISTINCTION: Mentioning a real tool, company, or product by name does NOT make text time\-aware\. Only FACTUAL CLAIMS THAT COULD BECOME OUTDATED do\.
Examples: \-\- "Use NetBeans profiler for CPU analysis" → 0 \(generic advice about a tool\) \-\- "Write a REST API using Flask and AWS Lambda" → 0 \(coding tutorial\) \-\- "Google announced Gemini 2\.0 in December 2024" → 1 \(dated event\) \-\- "The president of the United States is Joe Biden" → 1 \(will become outdated\) \-\- "Tesla’s market cap exceeded $1 trillion in 2024" → 1 \(time\-sensitive fact\) \-\- "Solve: 2x \+ 3 = 7" → 0 \(math\) \-\- "Anna bought 5 apples at $2 each" → 0 \(hypothetical\) \-\- "Python’s GIL prevents true multithreading" → 0 \(technical fact, stable\) \-\- "React 18 introduced concurrent rendering" → 0 \(historical tech fact, won’t change\) \-\- "As of 2024, React is the most popular framework" → 1 \(ranking changes over time\)
Figure 4:Prompt used for temporal leakage classification\.Similar Articles
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
This paper presents the first large-scale empirical analysis of 118 transformer models, revealing critical performance walls where success rates drop from 88.1% at 512 tokens to 0% at 2048 tokens, challenging prevailing scaling assumptions.
Scaling laws for neural language models
Foundational empirical study demonstrating power-law scaling relationships between language model performance and model size, dataset size, and compute budget, with implications for optimal training allocation and sample efficiency.
Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics
This paper introduces a method to predict best-of-N inference scaling gains for language models using cheap statistics from a single labeled validation-set sampling pass. A compact predictor with three core features achieves Spearman ρ=0.90 with actual gains, enabling screening of configurations before expensive reward-model scoring.
PALM: Point-in-Time Adaptation for Financial Language Models
This paper proposes PALM, a low-rank adapter method to adapt financial language models to specific time points without full retraining, reducing look-ahead bias and computational costs in financial backtests.
@_jasonwei: When language models first started using tools well, I was sympathetic to the narrative that instead of scaling up lang…
The author argues that while tool use allows smaller language models to perform tasks effectively, larger models remain crucial for speed, reliability, and internalized knowledge, emphasizing the ongoing need for scaling in AI.