思考的代价:测试时推理在LLM交易中是否奏效?

arXiv cs.AI 论文

摘要

本研究评估大型语言模型(如DeepSeek、GPT和Gemini)在交易中测试时推理是否提升净投资组合回报,发现在不同条件下额外的推理并不能可靠地增强经济结果。

arXiv:2609.30705v1 Announce Type: new Abstract: While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:41

# The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
Source: [https://arxiv.org/html/2609.30705](https://arxiv.org/html/2609.30705)
Jiayi Chenemail:[jc2693@njit\.edu](mailto:[email protected])Affiliation:Department of Computer Science,New Jersey Institute of Technology,Newark,New Jersey,USAGuiling Wangemail:[guiling\.wang@njit\.edu](mailto:[email protected])Affiliation:Department of Computer Science,New Jersey Institute of Technology,Newark,New Jersey,USA

© , 2026

###### Abstract\.

While inference\-time reasoning in large language models \(LLMs\) promises better decision making, its higher computational cost may not yield better economic outcomes\. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs\. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families\. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed\. Our evaluation covers a full year of U\.S\. equities under three input conditions: numerical, identifiable news, and masked news\. It includes more than 800,000 asset predictions and repeated model generations\. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns\. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic\. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar\. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment\.

###### Keywords:

large language models, test\-time compute, reasoning, financial decision making, algorithmic trading, backtesting

## 1\.Introduction

More reasoning is not necessarily better judgment\. A large language model \(LLM\) can produce a longer, more elaborate analysis and still make the same trade or a worse one\. This tension arises with*inference\-time reasoning*: extra internal computation requested after a model has been trained, without changing its weights or task input\. On benchmarks with verifiable answers, allocating more*test\-time compute*, the token budget spent on this extra processing, can improve accuracy by extending a reasoning trajectory, sampling more candidates, or invoking a*verifier*that checks candidate answers\([Wei et al\., 2022](https://arxiv.org/html/2609.30705#bib.bib1);[Wang et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib2);[Snell et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib5);[Muennighoff et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib6)\)\. Trading provides a harder test\. Correctness is revealed only later, through noisy prices and after costs, so the relevant question is not whether a response appears more thoughtful, but whether paying for more reasoning improves the resulting decision\.

Between the request for more reasoning and the realized investment outcome lies a sequence of transformations:

Reasoning effort→\\rightarrowstock scores→\\rightarrowholdings→\\rightarrownet returns

A*stock score*is the model’s numerical assessment of a stock’s expected return relative to the other candidates\. Each link in the sequence can weaken or reverse the apparent benefit of extra computation\. A large score change may leave the selected portfolio unchanged, while a small change near a selection boundary may replace a*long*position, which profits when its price rises, or a*short*position, which profits when its price falls\. Any gross improvement can then disappear after transaction costs\. The effect may also fail to repeat because each model call produces a*stochastic generation*, meaning one sampled output from an otherwise identical request\. More reasoning can therefore change scores and holdings without creating stable economic value\.

Existing financial LLM studies show that language models can extract information from news, answer financial questions, and participate in trading systems that use tools or multiple agents\([Lopez\-Lira and Tang, 2023](https://arxiv.org/html/2609.30705#bib.bib19);[Chen et al\., 2022](https://arxiv.org/html/2609.30705#bib.bib20);[Yu et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib21);[Zhang et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib22);[Yu et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib23);[Xiao et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib25)\)\. These studies establish that systems built around LLMs can trade, but they often change models, prompts, tools, memories, and decision procedures together\. They therefore do not isolate what is purchased by increasing reasoning inside an otherwise fixed trader\. Research on test\-time compute offers limited guidance because it focuses on mathematics, coding, and other tasks with immediate correctness signals\([Lightman et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib4);[Snell et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib5);[DeepSeek\-AI, 2025](https://arxiv.org/html/2609.30705#bib.bib7)\)\. Together, these literatures leave a narrower economic question unanswered:*when the model, information, prompt, required output, and trading mechanism are held fixed, does buying more reasoning improve portfolio outcomes?*

We answer this question with a controlled intervention on reasoning effort\. Within each model family, the reasoning level changes while the information available on each decision date, the prompt, the required output, the candidate stocks, and the portfolio rules remain fixed\. Figure[1](https://arxiv.org/html/2609.30705#S1.F1)shows this design as two reasoning paths entering the same scoring, portfolio, cost, and paired evaluation pipeline\. Our result is cautionary: added reasoning changes scores, selected assets, failure patterns, and operating cost, but it does not produce a reproducible improvement in return after costs in this setting\.

![An experimental pipeline sends numerical, identifiable news, or masked news inputs through fixed tasks, two reasoning levels, structured asset scoring, portfolio construction with long and short positions, returns after costs, and a paired estimator of treatment effects.](https://arxiv.org/html/2609.30705v1/plots/figure1_controlled_pipeline.png)Figure 1\.Controlled evaluation pipeline\. Within each backbone, reasoning level varies while information available by each formation date, prompts, output schema \(required fields, types, and ranges\), candidate assets, and portfolio rules remain fixed\. The information conditions are numerical inputs, identifiable news with identifying details retained, and masked news with identifying details removed\. Repeated calls on frozen formation dates assess the stability of scores, portfolios, and treatment effects\.An experimental pipeline sends numerical, identifiable news, or masked news inputs through fixed tasks, two reasoning levels, structured asset scoring, portfolio construction with long and short positions, returns after costs, and a paired estimator of treatment effects\.The study follows three predetermined model families over 241*formation dates*in 2024, where a formation date is a trading day on which new holdings are selected\. The common*low versus baseline comparison*raises reasoning from the baseline level to low\. The baseline is none for DeepSeek and GPT and minimal for Gemini, whose interface cannot guarantee that reasoning is fully disabled\. DeepSeek additionally provides a curve across none, low, high, and maximum reasoning\. Each model ranks the same 100 liquid U\.S\. equities under three information conditions:*numerical*features with identifiers removed;*identifiable news*, which retains company identity; and*masked news*, in which company names and other identifying details are removed\. The scores enter the same equal weight,*market neutral*portfolio, which buys the 10 highest ranked stocks, shorts the 10 lowest ranked stocks, and balances the two sides to remove broad market direction\. Each selection remains active for five days, creating overlapping*formation cohorts*, or batches of positions opened on different dates\. Trading incurs 10 basis points \(bps\) whenever capital is bought or sold; one bp is 0\.01 percentage points\.

This controlled design also determines how the evidence must be analyzed\. Although the experiment contains 808,800 asset predictions, the predictions and model tasks within a date are not independent economic observations\. Primary inference instead uses 241 paired daily portfolio return differences indexed by*return date*, a trading day on which active portfolios earn returns\. We apply a Newey–West correction with five lags, a heteroskedasticity and autocorrelation consistent \(HAC\) method that adjusts uncertainty for changing return variance and serial dependence\([Newey and West, 1987](https://arxiv.org/html/2609.30705#bib.bib45)\)\. A predeclared*economic threshold*of0\.250\.25bps/day defines the smallest gain considered worthwhile\. Invalid tasks receive a neutral score of 0 but remain in the comparison under an*intention\-to\-treat*rule\. To test whether an apparent effect survives another model call, we also conduct*frozen audits*: on 48 fixed formation dates, each backbone receives the same numerical or news task three times under unchanged rules\.

The results reveal a consistent difference between behavioral change and economic improvement\. None of the nine annual low versus baseline tests establishes a material return benefit, although the generally wide intervals do not show that every effect equals zero\. Repeated generations expose a sharper deployment risk\. In DeepSeek’s three generation audit with masked news, low reasoning reduces return by9\.0169\.016bps/day relative to none \(95% confidence interval \[CI\],−16\.248\-16\.248to−1\.784\-1\.784\)\. DeepSeek’s response across four reasoning levels is also nonmonotonic, so a favorable adjacent step does not imply that maximum reasoning outperforms no reasoning\. GPT and Gemini each change sign across three generations in at least one news condition\. Reasoning further changes*output structure reliability*, whether a response follows the required fields and format, in different directions across models\. Improved reliability therefore does not by itself imply improved return\.

The practical conclusion is asymmetric\. Added reasoning has a certain token, monetary, and latency cost, while its economic benefit remains unresolved in nine annual tests and is harmful in one replicated audit\. Baseline reasoning is therefore the conservative default unless the target task shows reproducible incremental value\. This conclusion is limited to the tested models, inputs, historical year, and portfolio mechanism\. It does not imply that reasoning never helps financial decisions, that the estimated effects are statistically equivalent to zero, or that the evaluated baseline portfolios are prospectively profitable\. Within that scope, this paper makes three contributions\.

1. \(1\)A controlled economic test of reasoning\.We isolate reasoning effort within each backbone while holding the surrounding financial workflow fixed, estimating the marginal value of inference\-time compute rather than differences among complete agent systems\.
2. \(2\)Evidence across the decision chain\.We connect structured scores to selected portfolios and returns after costs, then repeat generations to identify reasoning effects that appear favorable once but do not reproduce\.
3. \(3\)Accounting for deployment\.We jointly report return effects, portfolio stability, invalid outputs, token use, and inference cost, revealing when behavioral or reliability changes fail to create economic value\.

## 2\.Related Work

### 2\.1\.Financial Language Models and Return Prediction

Financial natural language processing \(NLP\) progressed from*encoders*adapted to finance, which turn financial text into numerical representations for sentiment and communications\([Araci, 2019](https://arxiv.org/html/2609.30705#bib.bib11);[Yang et al\., 2020](https://arxiv.org/html/2609.30705#bib.bib12)\), to generative models and instruction tuning frameworks for finance\([Wu et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib13);[Yang et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib14)\)\. Evaluation resources have expanded accordingly\. FinQA targets numerical reasoning over reports\([Chen et al\., 2021](https://arxiv.org/html/2609.30705#bib.bib15)\); PIXIU and FinBen cover broad financial understanding and decision tasks\([Xie et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib16);[Xie et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib17)\); and FinTradeBench combines fundamental information with trading signals\([Agrawal et al\., 2026](https://arxiv.org/html/2609.30705#bib.bib18)\)\. These works establish that finance requires both domain knowledge and quantitative reasoning, but benchmark accuracy does not itself show that additional inference\-time reasoning improves a portfolio\.

A second line of research asks whether representations learned by language models predict returns\. Lopez\-Lira and Tang find that ChatGPT labels of news headlines forecast subsequent stock movements\([Lopez\-Lira and Tang, 2023](https://arxiv.org/html/2609.30705#bib.bib19)\)\. Chen, Kelly, and Xiu use embeddings from LLMs to form signals that compare expected returns across stocks on the same date\([Chen et al\., 2022](https://arxiv.org/html/2609.30705#bib.bib20)\)\. These studies motivate our numerical and news conditions, but they also sharpen the distinction that guides our experiment\. Representation quality concerns the information encoded by a model; our target quantity, or*estimand*, is the incremental value of spending more compute while holding the representation, prompt, and data fixed\.

### 2\.2\.LLM Trading Agents and Financial Benchmarks

Recent systems embed LLMs in richer trading loops\. FinMem introduces layered memory and character design\([Yu et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib21)\); FinAgent combines multimodal inputs and tools\([Zhang et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib22)\); FinCon uses a hierarchy of managers and analysts with conceptual verbal reinforcement\([Yu et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib23)\); and FinRobot provides a general platform for financial agents\([Yang et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib24)\)\. TradingAgents organizes specialized agents into a simulated trading firm\([Xiao et al\., 2024](https://arxiv.org/html/2609.30705#bib.bib25)\)\. These systems demonstrate breadth, but their components interact: changing reflection, memory, tools, or agent roles changes more than the reasoning level\.

Benchmark work increasingly focuses on decision realism\. INVESTORBENCH spans financial products and agent tasks\([Li et al\., 2025a](https://arxiv.org/html/2609.30705#bib.bib26)\)\. StockBench evaluates sequential decisions while checking*contamination*, meaning possible exposure to evaluation information during training\([Chen et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib27)\)\. From Knowing to Doing masks identifiers to separate memorized market history from decision competence\([Zhu et al\., 2026](https://arxiv.org/html/2609.30705#bib.bib28)\), while Profit Mirage tests*leakage*\(future or evaluation information entering the task\) and performance after the model’s training data cutoff\([Li et al\., 2025b](https://arxiv.org/html/2609.30705#bib.bib30)\)\. Agentic Trading synthesizes audit concerns for financial agents\([Xia et al\., 2026](https://arxiv.org/html/2609.30705#bib.bib29)\)\. We adopt opaque identifiers and masked news for the same reason, while acknowledging that masking cannot prove that historical events were absent from pretraining\.

Several recent studies move closer to our question by testing reasoning in financial decisions\. Sodha compares direct and thinking LLMs in cross\-sectional stock ranking and finds no net portfolio advantage for the thinking model after costs\([Sodha, 2025](https://arxiv.org/html/2609.30705#bib.bib32)\)\. Cheng et al\. vary reasoning intensity and market speed in a simulator for tender selection and real time execution, exposing a tradeoff between decision quality and latency\([Cheng et al\., 2026](https://arxiv.org/html/2609.30705#bib.bib33)\)\. Sitjar toggles a chain\-of\-thought scaffold within the decision process for a trader of the S&P 500 exchange\-traded fund \(SPY\) and finds no main improvement in trading performance or*calibration*, the agreement between stated confidence and observed outcomes\([Sitjar, 2026](https://arxiv.org/html/2609.30705#bib.bib31)\)\. These studies show why a controlled reasoning comparison is needed, so we do not claim to be the first to make one\. Our contribution is a within backbone design across three model families, common numerical, identifiable news, and masked news tasks, a four level reasoning curve for one backbone, explicit transaction costs, and repeated generation evidence at the score, portfolio, and return levels\.

### 2\.3\.Test\-Time Compute and Adaptive Reasoning

*Chain\-of\-thought prompting*asks a model to produce intermediate reasoning steps and can improve tasks with multiple steps\([Wei et al\., 2022](https://arxiv.org/html/2609.30705#bib.bib1)\)\.*Self\-consistency*samples several such paths and aggregates their answers\([Wang et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib2)\), while Tree of Thoughts explicitly searches among intermediate possibilities\([Yao et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib3)\)\. Process supervision and stepwise verification score intermediate steps to select better trajectories\([Lightman et al\., 2023](https://arxiv.org/html/2609.30705#bib.bib4)\)\. More recent work studies scaling directly at inference\. Snell et al\. show that test\-time compute can substitute for parameter scale when allocation and verification are effective\([Snell et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib5)\); s1 demonstrates simple budget forcing\([Muennighoff et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib6)\); and DeepSeek\-R1 exemplifies models trained for extended reasoning\([DeepSeek\-AI, 2025](https://arxiv.org/html/2609.30705#bib.bib7)\)\.

The gains are not free or necessarily monotonic\. Think or Not? documents diminishing information gain in long traces and proposes*adaptive stopping*, ending reasoning when extra steps add little information\([Yong et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib8)\)\. Adaptive Model and Strategy Routing lets a system choose a model and reasoning strategy under a budget\([Pan et al\., 2026](https://arxiv.org/html/2609.30705#bib.bib9)\), and a recent survey organizes controllable and adaptive methods for test\-time compute\([Alomrani et al\., 2025](https://arxiv.org/html/2609.30705#bib.bib10)\)\. These approaches usually optimize accuracy against a verifier\. We instead treat realized portfolio return as the downstream outcome and test the more basic premise that a provider’s reasoning control has incremental economic value in a noisy decision pipeline with delayed feedback\.

### 2\.4\.Evaluation Under Financial Noise

Financial evidence requires stronger controls than a static accuracy benchmark\. News can move markets\([Tetlock, 2007](https://arxiv.org/html/2609.30705#bib.bib34)\), and machine learning can extract nonlinear return signals\([Gu et al\., 2020](https://arxiv.org/html/2609.30705#bib.bib35)\), but those relationships are*nonstationary*: they can change over time\.*Data snooping*, or repeatedly searching historical data until a favorable result appears, can invalidate conventional tests\([Lo and MacKinlay, 1990](https://arxiv.org/html/2609.30705#bib.bib36)\)\. Searches across many rules therefore require corrections within families to control false discoveries across the full set\([Sullivan et al\., 1999](https://arxiv.org/html/2609.30705#bib.bib37);[White, 2000](https://arxiv.org/html/2609.30705#bib.bib38)\), and the superior predictive ability test addresses comparison against many alternatives\([Hansen, 2005](https://arxiv.org/html/2609.30705#bib.bib39)\)\. A*backtest*is a simulation on historical data; selecting among many backtests raises the probability of overfitting and motivates the*deflated Sharpe ratio*, which discounts performance for selection and returns that are not normally distributed\([Bailey et al\., 2016](https://arxiv.org/html/2609.30705#bib.bib40);[Bailey and Lopez de Prado, 2014](https://arxiv.org/html/2609.30705#bib.bib41)\)\. The large set of published return predictors raises the same problem of multiple testing\([Harvey et al\., 2016](https://arxiv.org/html/2609.30705#bib.bib42)\), while combining many candidate signals can itself induce severe overfitting\([Novy\-Marx, 2015](https://arxiv.org/html/2609.30705#bib.bib43)\)\. Decay in anomalies after publication reinforces the need for evidence on unseen data and in future deployment\([McLean and Pontiff, 2016](https://arxiv.org/html/2609.30705#bib.bib44)\)\. These concerns determine the controls used in our evaluation\.

Accordingly, our paired design avoids treating the 100 asset predictions on a date as independent economic observations\. We estimate effects on daily portfolio returns with HAC inference\([Newey and West, 1987](https://arxiv.org/html/2609.30705#bib.bib45)\), disclose all tested contrasts between backbones and information conditions, adjust the three primary tests within each model family, and retain failed outputs under an intention\-to\-treat rule\. The resulting claim is deliberately narrow: we estimate the marginal economic value of reasoning effort inside a frozen financial decision pipeline rather than comparing complete agents or asserting that reasoning has no value in finance\.

Table 1\.Experimental design\. Each primary study uses 241 formation dates, meaning trading days when new holdings are selected\. Each audit generation is produced by a repeated call on the frozen audit formation dates\.

## 3\.Experimental Design

### 3\.1\.Research Question and Estimand

An*estimand*is the population quantity an experiment is designed to measure\. For backbonebb, information conditionaa, reasoning levelrr, and portfolio return datett, letRb,a,r,tnetR^\{\\mathrm\{net\}\}\_\{b,a,r,t\}denote the net return of the frozen portfolio implementation\. The common low versus baseline estimand is

\(1\)Δb,a=𝔼⁡\[Rb,a,low,tnet−Rb,a,base,tnet\],\\Delta\_\{b,a\}=\\mathbb\{E\}\\\!\\left\[R^\{\\mathrm\{net\}\}\_\{b,a,\\mathrm\{low\},t\}\-R^\{\\mathrm\{net\}\}\_\{b,a,\\mathrm\{base\},t\}\\right\],where base is none for DeepSeek and GPT and minimal for Gemini\. Positive values favor low reasoning\. The primary inferential observations are daily portfolio return differences indexed by return date\. Asset predictions and model tasks within a formation date are inputs to a single portfolio, not independent return observations\.

We predefine the economic thresholdϵ=0\.25\\epsilon=0\.25bps/day as the smallest improvement considered worthwhile\. Using a two\-sided 95% CI\[L,U\]\[L,U\]and a one\-sided 95% upper boundU\.95U\_\{\.95\}, we classify a contrast as: material benefit ifL\>ϵL\>\\epsilon; harm ifU<0U<0; no material benefit ifU\.95<ϵU\_\{\.95\}<\\epsilon; and inconclusive otherwise\. This asymmetric rule reflects the deployment question: added reasoning has known compute cost and must demonstrate sufficient incremental value\. An inconclusive result is not*equivalence*, which would require ruling out effects outside prespecified small bounds\.

### 3\.2\.Market Sample and Information Timing

The experiment uses liquid common equities listed in the United States\.*Information timing*means that each decision uses only information available by its formation date\. The initial snapshot, dated December 29, 2023, contains 120 eligible names, of which 100 are active on each evaluated formation date\. There are 241 formation dates from January 10 through December 23, 2024\. Each formation date’s*cross\-section*, meaning the 100 stocks compared on that date, is deterministically split into four model tasks of 25 assets each\. Assets use*opaque identifiers*, random codes that reveal no company name, and no future return enters an inference packet or repair decision\.

These dated decisions become overlapping portfolio observations\. A*formation cohort*is the portfolio opened on one formation date\. It enters at that session’s open, earns open\-to\-close return on that date, and then close\-to\-close return on the next four sessions\. A*return date*is a trading day on which active formation cohorts earn portfolio returns\. After the first four formation dates are removed, the realized path contains 241 return dates, from January 17 through December 30\.

The frozen stochastic audit applies the same timing logic to 48 clustered formation dates selected without reference to their eventual returns\. Because formation cohorts are held for five days and overlap, the audit portfolio path contains 68 return dates\. Repeated generations are averaged or compared within return date; they never multiply the number of market observations\.

### 3\.3\.Backbones and Reasoning Levels

Table[1](https://arxiv.org/html/2609.30705#S2.T1)gives the final analyzed panel\. All three backbones use numerical, identifiable news, and masked news inputs, but the available reasoning controls differ\. DeepSeek tests none, low, high, and maximum reasoning\. GPT tests none versus low\. Gemini tests minimal versus low because Gemini 3\.x could not guarantee a condition with reasoning fully disabled\. Only DeepSeek therefore supports a claim about the response across four reasoning levels; GPT and Gemini replicate the common low versus baseline comparison\.

The three evaluated backbones were DeepSeek\-Flash, GPT\-5\.6 Luna, and Gemini 3\.1 Flash\-Lite\. We refer to them as DeepSeek, GPT, and Gemini except where the exact model name is needed\. All were predetermined for the study\. Before each backbone was evaluated, its exact reasoning contrast, inputs, prompts, output contract, failure policy, portfolio mapping, and analysis plan were frozen\. Each backbone therefore supplies a separately controlled intervention within one model rather than an exercise in model selection\.

### 3\.4\.Information Conditions

The*numerical*condition supplies 16 daily*cross\-sectional rank features*\. Each raw measure is ranked across the broader fixed research panel available on that date, scaled to\[−1,1\]\[\-1,1\], and then retained for the 100 active candidate stocks\. The measures are return lagged one day; momentum from past returns over 5, 21, 63, 126, and 252 days; volatility over 5, 21, and 63 days; downside volatility over 21 days; a volume ratio over 21 days; log dollar volume over 21 days;*Amihud illiquidity*over 21 days, the amount of price movement per dollar traded; previous intraday return; previous trading range; and log previous close\. For each feature the packet includes its current value, mean over 5 days, mean over 20 days, and change over 20 days\.

The news conditions provide article information available by each formation date\.*Identifiable news*retains identifying details\.*Masked news*removes direct company, ticker, date, and event identifiers while preserving the text structure relevant to the task\. This manipulation reduces obvious historical recognition but cannot guarantee that a model has no latent memory of an event\.

### 3\.5\.Prompt, Output Contract, and Integrity Controls

Every model receives the same substantive instruction: estimate relative return over the next five trading sessions, score the strongest candidate near\+1\+1and the weakest near−1\-1, and return every opaque asset identifier exactly once\. The output*schema*is the set of required fields, types, and ranges that a machine can check\. It contains*score*in\[−1,1\]\[\-1,1\],*confidence*,*materiality*, and*uncertainty*in\[0,1\]\[0,1\], plus an*abstain*flag that lets the model decline to score an asset\. Prompts, asset packets, candidate order, schema, output allowance, and downstream mapping are frozen within each paired comparison\.

Responses undergo schema and identifier validation\. Retries and repairs under deterministic rules are*outcome blind*, meaning they are made without viewing subsequent returns; raw attempts, usage, and repair records are retained\. When a task remains invalid, all 25 rows are assigned a neutral score of 0 and kept under the primary intention\-to\-treat rule\. DeepSeek ended with complete task coverage after 17 deterministic contract repairs, affecting 15 of 404,400 asset rows by neutral imputation\. Across the three primary information conditions, GPT retained 57 invalid tasks \(0\.99%\) and Gemini retained 136 \(2\.35%\)\. Analyses that remove incomplete dates or match omitted task chunks are reported as robustness checks, not replacements for the frozen primary analysis\.

−10\-10−5\-50055101015152020Gemini maskedGemini identifiableGemini numericalGPT maskedGPT identifiableGPT numericalDeepSeek maskedDeepSeek identifiableDeepSeek numerical\+0\.25\+0\.25economic thresholdEffect of low relative to baseline reasoning on net return \(bps/day\)Figure 2\.Primary low versus baseline return effects across 241 return dates\. Circles denote numerical inputs, squares denote identifiable news, and diamonds denote masked news\. Horizontal bars are two\-sided 95% HAC CIs, which account for changing variance and serial dependence\. The solid line is zero and the dashed line is the\+0\.25\+0\.25bps/day economic threshold\. No interval establishes a material benefit\.A horizontal forest plot shows nine effects of low reasoning relative to baseline across numerical, identifiable news, and masked news inputs\. All confidence intervals cross zero and the positive economic threshold\.

## 4\.Portfolio Evaluation and Statistical Inference

### 4\.1\.Mapping Scores to Portfolios

The portfolio mapping converts each model output into one comparable trading signal\. For assetiion formation datedd, the downstream signal is

\(2\)si,d=\{0,if abstaining or imputed to​0,qi,d​\(0\.5\+0\.5​ci,d\),otherwise\.s\_\{i,d\}=\\begin\{cases\}0,&\\text\{if abstaining or imputed to \}0,\\\\ q\_\{i,d\}\\left\(0\.5\+0\.5c\_\{i,d\}\\right\),&\\text\{otherwise\}\.\\end\{cases\}Here,qqis the model score andccis model reported confidence\. The model reported*materiality*and*uncertainty*fields are preserved for diagnostics but do not alter position size\.

We then ranksi,ds\_\{i,d\}across the 100 names on each formation date, short the bottom 10, and buy the top 10\. Positions receive equal weights within each side\. Thus*gross exposure*is 2 \(one dollar long plus one dollar short per dollar of capital\) and*net exposure*is 0 \(long and short dollar amounts cancel\)\. If abstentions leave fewer than 20 eligible assets, the frozen fallback ranks all neutralized signals and breaks ties deterministically by asset identifier\.

Each formation cohort is held for five sessions\. The implemented daily portfolio is the simple average of active cohorts, so formation decisions overlap in calendar time\. All reasoning levels use the identical simulator and candidate universe\.

### 4\.2\.Trading and Inference Costs

The primary assumption for trading costs is 10 bps one way, with 5, 20, and 50 bps sensitivity analyses\. For a portfolio with capitalAA, incremental inference expenseCCcan be expressed as

\(3\)inference cost \(bps\)=10,000​C/A\.\\text\{inference cost \(bps\)\}=10\{,\}000\\,C/A\.We keep provider inference dollars separate from market returns because the same research workload can support different capital levels\. Deployment conclusions use the direction and uncertainty of the effect after trading costs first, then ask whether any supported gain could cover inference expense\.

### 4\.3\.Paired Inference

For every comparison between a backbone and information condition, we align the two portfolio paths by return date and computeDt=Rlow,tnet−Rbase,tnetD\_\{t\}=R^\{\\mathrm\{net\}\}\_\{\\mathrm\{low\},t\}\-R^\{\\mathrm\{net\}\}\_\{\\mathrm\{base\},t\}\. The mean and*standard error*, an estimate of sampling variability, use a Newey–West estimator with five lags to account for changing variance and the serial dependence induced by overlapping formation cohorts\([Newey and West, 1987](https://arxiv.org/html/2609.30705#bib.bib45)\)\. For the comparable annual low versus baseline results, all three backbones have three information conditions\. Their two\-sided*pp\-values*, which give the probability under an assumption of zero effect of obtaining a result at least this extreme, are adjusted with the*Holm method*, a stepwise correction that controls false positives across the three tests within each backbone\. This harmonized reporting layer does not replace DeepSeek’s predeclared sequential procedure: none to low, low to maximum when the first comparison is unresolved, and the total effect from none to maximum\. The collected high level is analyzed as part of a secondary curve across four reasoning levels\.

As supplementary uncertainty checks, we add two*bootstraps*, which estimate variability by repeatedly resampling the observed data\. They do not replace the primary HAC inference\. A paired*moving block bootstrap*\([Kunsch, 1989](https://arxiv.org/html/2609.30705#bib.bib46)\)resamples 50,000 sequences of consecutive return dates\. Its main block of five sessions matches the holding horizon, while blocks of 10 and 20 sessions test sensitivity\. A*pairs bootstrap over formation cohorts*instead resamples each cohort’s complete contribution to gross returns over five sessions together with its paired difference in transaction costs\. The cohort contributions reconcile exactly to the observed sum of net returns\. Confidence intervals use the 2\.5th and 97\.5th percentiles of the resampled effects, and two\-sided tests center the resampled differences at zero\. The moving block result is the stronger check of dependence over calendar time because the cohort bootstrap treats formation cohorts as interchangeable even though nearby cohorts can share shocks on the same return date\.

Together, these procedures produce the point estimates, two\-sided CIs, one\-sided economic bounds, and decisions reported in the paper\. A conventional*null hypothesis test*begins from a default assumption of zero effect; merely failing to reject that assumption is not treated as evidence that reasoning has no effect\.

We next ask whether an annual estimate is concentrated in one part of the year\. As an exploratory diagnostic, we partition each saved primary daily effect series by return date into calendar halves \(January–June and July–December\) and quarters\. We recompute HAC intervals with five lags within each subgroup and use HAC regressions for*heterogeneity tests*, which ask whether the effect differs across halves or quarters\. We adjust the resulting 18 tests with the Holm method\. We also remove each quarter in turn and measure its share of the sum of absolute quarterly contributions\. These partitions use no new model calls, do not change the annual estimands, and are not treated as independent replications\.

### 4\.4\.Evaluation Across Stochastic Generations

On the frozen audit formation dates we assess three layers of reproducibility: \(i\)*Spearman rank agreement*, which measures whether assets keep a similar ordering; \(ii\)*Jaccard overlap*, the fraction of names shared by the selected long and short sets; and \(iii\) the difference between low and baseline reasoning for each generation\. For every information condition, each backbone has the primary generation restricted to the audit dates plus two additional generations\. For GPT and Gemini we also form a portfolio from*ensemble scores*by averaging the three signals before ranking, alongside the average within each return date of the three return effects\. Because ranking and selection of the sets with the highest and lowest ranks are nonlinear, the ensemble score effect need not equal the arithmetic mean of the three return effects\.

## 5\.Results

### 5\.1\.RQ1: Does Low Reasoning Improve Annual Performance?

Figure[2](https://arxiv.org/html/2609.30705#S3.F2)and Table[2](https://arxiv.org/html/2609.30705#S5.T2)report the nine comparable annual low versus baseline effects\. None satisfies the rule for material benefit\. Six point estimates are positive and three are negative, but every 95% CI includes zero and the\+0\.25\+0\.25bps/day economic threshold\. DeepSeek’s numerical point estimate is the largest at\+6\.159\+6\.159bps/day, yet its interval ranges from−7\.243\-7\.243to\+19\.560\+19\.560\. The identifiable news estimates are\+0\.378\+0\.378for DeepSeek,−0\.626\-0\.626for GPT, and−0\.142\-0\.142for Gemini, and all three intervals likewise admit gains and losses\. Holm adjustment across the three information conditions within each backbone does not change any decision\. The annual evidence therefore supports*no reliable benefit detected*, not the stronger claim of a zero effect\.

Table 2\.Primary low versus baseline results at a one way trading cost of 10 bps\. Adjustedpp\-values use the Holm method across the three information conditions separately within each backbone\.These paired effects describe the value of additional reasoning, not the absolute attractiveness of either portfolio\. Appendix Table[7](https://arxiv.org/html/2609.30705#A6.T7)reports the underlying paths: 13 of the 18 portfolios at the baseline and low reasoning levels have negative mean net return\. The positive annualized returns are confined to Gemini identifiable news at both levels and Gemini masked news at minimal reasoning; the largest is2\.16%2\.16\\%\. A*compounded annualized return*is the equivalent annual growth rate\. A*Sharpe ratio*divides mean return by return volatility, and*maximum drawdown*is the largest portfolio loss from a peak to a trough\.

The conclusion also depends on the frozen portfolio construction\. Appendix Table[8](https://arxiv.org/html/2609.30705#A6.T8)varies the number of selected names\. Across 36 exploratory combinations of backbone, information condition, and width, 35 do not establish material benefit\. The exception is DeepSeek masked news with 20 names per side \(\+5\.769\+5\.769bps/day; HAC 95% CI\[0\.597,10\.942\]\[0\.597,10\.942\]\), which also has a positive interval from a moving block bootstrap with five sessions per block\. However, no HAC or moving blockpp\-value survives Holm correction across the 36 cells\. The exception does not alter the frozen primary estimand based on 10 long and 10 short names, but it bounds the conclusion to that tested construction\.

### 5\.2\.RQ2: Does More Reasoning Consistently Improve Returns?

Only DeepSeek tests a curve across four reasoning levels\. Figure[3](https://arxiv.org/html/2609.30705#S5.F3)a is a signed matrix of point estimates averaged across generations for portfolios formed on the audit sample of 48 formation dates, relative to none\. The printed values make the nonmonotonic patterns explicit without implying a smooth response between provider levels\. For numerical inputs, the effect declines from none through low, high, and maximum\. For identifiable news, low reasoning drops sharply, high changes little, and maximum recovers part of the loss but not all of it\. Masked news remains below none at every positive reasoning level\. These are descriptive level estimates; confidence intervals are attached to the predeclared contrasts rather than invented for cumulative points\.

The confirmatory chain is consistent with the visual result\. Low minus none is−0\.529\-0\.529bps/day for numerical inputs,−8\.525\-8\.525for identifiable news, and−9\.016\-9\.016for masked news\. The CI for masked news,\[−16\.248,−1\.784\]\[\-16\.248,\-1\.784\], establishes harm\. Maximum minus none remains inconclusive for numerical and identifiable news and is classified as no material benefit for masked news\. The only decisive adjacent secondary contrast is maximum minus high for identifiable news,\+4\.639\+4\.639bps/day \(95% CI\[\+0\.376,\+8\.903\]\[\+0\.376,\+8\.903\]\)\. Because maximum still fails to beat none, a favorable local step does not justify the claim that more reasoning is better\.

### 5\.3\.RQ3: Are Reasoning Effects Stable Across Generations?

Figure[3](https://arxiv.org/html/2609.30705#S5.F3)b begins with masked news, where the contrast between reproducible harm and unstable direction is clearest\. It shows low versus baseline estimates as three unconnected generation markers for each backbone, together with the average within each return date and its 95% HAC interval\. DeepSeek’s three individual audit generations are all negative \(−7\.490\-7\.490,−9\.327\-9\.327, and−10\.231\-10\.231bps/day\), producing the harm result after aggregation within each return date\. GPT instead changes sign across\+5\.800\+5\.800,−6\.665\-6\.665, and−2\.371\-2\.371bps/day\. Its average across three generations within each return date is−1\.079\-1\.079bps/day \(95% CI\[−5\.173,3\.015\]\[\-5\.173,3\.015\]\), while its ensemble score portfolio is likewise inconclusive at−0\.957\-0\.957bps/day \(95% CI\[−6\.814,4\.900\]\[\-6\.814,4\.900\]\)\.

Gemini also changes sign: its first two generations are positive \(\+1\.777\+1\.777and\+7\.918\+7\.918bps/day\), whereas the third is−0\.630\-0\.630bps/day\. The average within each return date is\+3\.022\+3\.022bps/day \(95% CI\[−3\.377,9\.421\]\[\-3\.377,9\.421\]\), and the ensemble score portfolio remains inconclusive at\+1\.609\+1\.609bps/day \(95% CI\[−6\.545,9\.764\]\[\-6\.545,9\.764\]\)\. Thus the stability result is not that all generations are negative\. It is that a favorable isolated generation does not establish reproducible economic value\. The third generations retain one of 384 GPT tasks and 13 of 384 Gemini tasks as neutral after invalid outputs; none is selectively rerun\. These audits do not replace the primary estimates from 241 return dates or add independent market observations\.

The other information conditions preserve the deployment concern without reproducing the same directions\. For identifiable news, GPT’s generation effects are\+0\.082\+0\.082,−0\.649\-0\.649, and−3\.881\-3\.881bps/day\. Their average within each return date is−1\.483\-1\.483bps/day, with a 95% CI from−5\.940\-5\.940to\+2\.974\+2\.974, while the score ensemble is−3\.604\-3\.604bps/day, with a CI from−8\.699\-8\.699to\+1\.491\+1\.491\. Gemini changes sign across−0\.688\-0\.688,\+1\.476\+1\.476, and\+8\.870\+8\.870bps/day\. Its average within each return date is\+3\.219\+3\.219bps/day, with a CI from−7\.032\-7\.032to\+13\.470\+13\.470, and its ensemble is\+3\.882\+3\.882bps/day, with a CI from−6\.561\-6\.561to\+14\.325\+14\.325\. None establishes a reliable benefit\. Appendix Table[10](https://arxiv.org/html/2609.30705#A8.T10)shows that similar overall scores still yield incomplete agreement in the selected portfolio tails\.

Numerical inputs complete the same three generation design\. GPT’s effects are−3\.972\-3\.972,−0\.927\-0\.927, and−1\.112\-1\.112bps/day\. Their average within each return date is−2\.004\-2\.004bps/day \(95% CI\[−6\.386,2\.379\]\[\-6\.386,2\.379\]\), and the score ensemble is−3\.866\-3\.866bps/day \(95% CI\[−9\.870,2\.138\]\[\-9\.870,2\.138\]\)\. Gemini’s generation effects are\+2\.750\+2\.750,\+0\.788\+0\.788, and\+2\.290\+2\.290bps/day\. Their average within each return date is\+1\.943\+1\.943bps/day \(95% CI\[−6\.161,10\.047\]\[\-6\.161,10\.047\]\), while the score ensemble is−0\.222\-0\.222bps/day \(95% CI\[−8\.158,7\.714\]\[\-8\.158,7\.714\]\)\. This sign reversal after score averaging illustrates that stable looking scores do not guarantee a stable selected portfolio\. None of the numerical audit summaries establishes a reliable benefit\. The two added GPT generations retain six of 768 tasks as neutral after invalid outputs; both added Gemini generations have complete valid coverage\.

![Panel a is a signed matrix of DeepSeek effects at none, low, high, and maximum reasoning for numerical, identifiable news, and masked news inputs. Panel b plots three separate generation markers for masked news and an aggregate confidence interval for each backbone. DeepSeek is consistently negative, while GPT and Gemini each change sign.](https://arxiv.org/html/2609.30705v1/figure3_effect_matrix_replication.png)Figure 3\.Reasoning levels and repeated call stability\. \(a\) DeepSeek effects averaged across generations for portfolios formed on 48 audit formation dates, relative to none; cells are descriptive because intervals for cumulative levels were not computed\. \(b\) Three unconnected generation markers per backbone; diamonds and bars show the average within each return date and the 95% HAC CI adjusted for serial dependence\. Repeated calls are not independent market samples\.Panel a is a signed matrix of DeepSeek effects at none, low, high, and maximum reasoning for numerical, identifiable news, and masked news inputs\. Panel b plots three separate generation markers for masked news and an aggregate confidence interval for each backbone\. DeepSeek is consistently negative, while GPT and Gemini each change sign\.
### 5\.4\.Robustness Across Resampling and Calendar Periods

We first test whether the conclusions depend on the primary HAC calculation\. The resampling checks preserve the headline decisions \(Appendix Table[6](https://arxiv.org/html/2609.30705#A4.T6)\)\. All nine low versus baseline contrasts across 241 return dates remain inconclusive under both the moving block bootstrap with five sessions per block and the bootstrap over formation cohorts\. For identifiable news, the moving block intervals are\[−9\.099,9\.287\]\[\-9\.099,9\.287\]for DeepSeek,\[−6\.020,4\.586\]\[\-6\.020,4\.586\]for GPT, and\[−7\.690,8\.497\]\[\-7\.690,8\.497\]bps/day for Gemini\. Changing the block length to 10 or 20 sessions does not produce a decision of material benefit\. In the DeepSeek audit of three generations using masked news, the moving block interval is\[−17\.929,−3\.477\]\[\-17\.929,\-3\.477\]and the interval from the formation cohort bootstrap is\[−15\.482,−2\.281\]\[\-15\.482,\-2\.281\]bps/day, so the harm classification survives both checks\.

We then test whether the annual directions are confined to one part of 2024 \(Appendix Table[11](https://arxiv.org/html/2609.30705#A9.T11)\)\. Only GPT numerical keeps the same sign in both halves of the year\. DeepSeek numerical changes from−10\.027\-10\.027bps/day in the first half to\+20\.687\+20\.687in the second \(raw HAC heterogeneityp=\.014p=\.014\), while Gemini masked news ranges from\+12\.083\+12\.083in Q3 to−17\.974\-17\.974in Q4 \(raw test of equality across quartersp=\.045p=\.045\)\. Neither result survives Holm correction across the 18 exploratory heterogeneity tests by half and quarter\. Seven of nine annual directions also flip when at least one quarter is omitted, although the largest absolute quarter contribution is at most 64\.5%\. The annual findings are therefore not reducible to one isolated quarter, but neither are they temporally uniform\. This analysis strengthens the conclusion that reasoning effects vary across market episodes within 2024; it does not substitute for another year\.

## 6\.Operational Reliability and Compute Economics

### 6\.1\.Output Structure Reliability

Reasoning affects task validity differently across backbones\. Fisher’s exact test, which compares two failure proportions, is used as a diagnostic\. GPT low reasoning increases numerical failures from 4 to 14 of 964 tasks \(0\.41% to 1\.45%;p=\.030p=\.030\), while failures for masked news are similar \(8 versus 9\)\. For identifiable news, GPT failures fall from 17 to 5 \(p=\.016p=\.016\), and Gemini failures fall from 54 to 14 \(p<\.000001p<\.000001\)\. Gemini low reasoning also reduces failures for masked news from 44 to 20 of 964 tasks \(4\.56% to 2\.07%;p=\.0032p=\.0032\), but neither reliability improvement produces a return benefit\. DeepSeek has no final invalid tasks after rare outcome blind repairs\. Reliability is therefore a separate deployment dimension, not a proxy for investment performance; Table[3](https://arxiv.org/html/2609.30705#S6.T3)retains the primary rates of invalid tasks\.

### 6\.2\.Tokens, Dollars, and the Deployment Decision

Because reliability and return can diverge, deployment also requires a separate accounting of compute\. Table[3](https://arxiv.org/html/2609.30705#S6.T3)reports compute, cost, and reliability for all evaluated tasks within each backbone\. These are backbone totals rather than efficiency comparisons at a matched workload\. DeepSeek dominates compute because it evaluates four reasoning levels across all three information conditions\. Its evaluation uses 269\.0 million reported reasoning tokens\. GPT uses 0\.37 million reasoning tokens, while Gemini uses 7\.93 million, illustrating that nominal level labels are not comparable measures of actual compute across providers\.

For GPT’s primary evaluation, low reasoning costs an additional $0\.375 for numerical inputs and $0\.092 for masked news, while identifiable news costs $0\.281 less because the low level produces fewer output tokens in this run\. These provider charges are modest relative to the return intervals, but a low dollar cost does not turn an unsupported return gain into evidence of benefit\. Latency also depends on the provider and delivery mode and is not used as a comparable headline metric\.

Table 3\.Compute, cost, and reliability for all evaluated tasks within each backbone\. Token counts are reported by providers and should not be compared as identical units of cognition\. Invalid primary tasks is the share of primary tasks still unresolved after the frozen validation and repair policy\. NA denotes not aggregated; token columns are not aggregated across providers, and total cost is summed before row values are rounded\.

## 7\.Discussion

### 7\.1\.Reasoning Is Not a Monotonic Economic Control

The results distinguish behavioral responsiveness from economic usefulness\. Reasoning effort clearly changes outputs: token use rises, rankings move, portfolio constituents change, and rates of invalid outputs sometimes shift\. Yet those changes do not form a stable monotonic relationship with returns after costs\. DeepSeek provides the strongest demonstration\. An adjacent increase from high to maximum helps identifiable news on the audit sample, but maximum still fails to outperform none\. Treating reasoning as a conventional scalar*hyperparameter*, or a tunable numerical setting for which more is often presumed better, is therefore misleading in this setting\.

This pattern is compatible with several mechanisms, none of which our experiment proves\. Longer traces may assign too much weight to weak narrative evidence, revise a useful first judgment, increase confidence without increasing information, or simply create another stochastic route through an underdetermined task\. Conversely, reasoning may improve local extraction or schema compliance while leaving the economically relevant rank order unchanged\. A causal mechanism study would require randomized interventions on trace content or a task with intermediate labels, not inspection of provider summaries alone\.

### 7\.2\.What the Negative Result Does and Does Not Mean

The most defensible statement is that additional low reasoning did not deliver a reliable improvement in this frozen 2024 pipeline\. It is stronger than saying that nine conventionalpp\-values exceed 0\.05 because it incorporates transaction costs, an economic threshold, repeated generations, and operational failures\. It is weaker than equivalence, which requires formal evidence that effects lie within small prespecified bounds, because most CIs still include economically meaningful gains and losses\. The study also does not establish that the baseline trading strategies are profitable, that reasoning cannot help other financial tasks, or that the evaluated provider controls correspond to the same quantity of compute\.

### 7\.3\.Implication for Governed Deployment

In an institutional process for model risk, increasing reasoning effort should be treated as a configuration change that requires downstream validation, not as a capability upgrade that is beneficial by default\. Approval evidence should match the deployed decision rule and include action stability, failures, costs, and performance across repeated generations\. Under that standard, an isolated favorable generation or an unresolved positive point estimate is insufficient to promote a more expensive reasoning level\. The conservative default in this study is therefore a governance decision under uncertainty, not a claim that baseline reasoning is intrinsically superior\.

## 8\.Limitations

The first limitation is temporal and model scope\. Evidence comes from one 2024 backtest of U\.S\. equities and the provider interfaces observed at execution\. The analysis by calendar period spans distinct episodes within the year and shows that no contrast is wholly driven by one quarter, but the frequent sign changes reinforce the need for replication across years or in a prospective study\. Masking reduces cues to recognizable events but cannot rule out latent historical knowledge, and future provider updates may change behavior\. Reasoning level support is also asymmetric: only DeepSeek exposes the four tested levels, so provider labels should not be interpreted as equal units of compute\.

The second limitation is precision and portfolio dependence\. Most primary intervals admit economically meaningful gains and losses; the 80% minimum detectable effects are 6\.200–19\.156 bps/day versus the 0\.25 bps/day economic threshold \(Appendix Table[9](https://arxiv.org/html/2609.30705#A7.T9)\)\. The result is therefore a conservative deployment decision, not an equivalence claim\. It is also tied to the frozen rule that selects 10 long and 10 short names: one of 36 exploratory width cells is positive before, but not after, correction across the family\.

The final limitation concerns deployment and mechanism\. The simulator models fixed transaction costs and overlapping holdings but not nonlinear*market impact*\(prices moving because of the strategy’s own trades\), borrow availability for short sales, capacity, latency, or uncertainty in live orders\. The experiment estimates the total effect of requesting more reasoning; it does not identify the hidden computation causing a score or trade change\. These omissions lead directly to the next validation step: prospective*shadow trading*, which runs the strategy on live inputs without placing orders, is required before any claim about live deployment\.

## 9\.Conclusion

We evaluate inference\-time reasoning as an economic intervention in a frozen LLM system for ranking equities\. Across DeepSeek, GPT, and Gemini, none of nine low versus baseline tests across 241 return dates establishes a material improvement in net portfolio return\. DeepSeek’s replicated result with masked news shows harm, its curve across four reasoning levels is nonmonotonic, and repeated generations reveal sign instability across the news conditions\. Reasoning can also improve or degrade output reliability without improving returns\. In this setting, extra thought changes decisions but does not reliably pay for itself\. For the primary implementation with 10 long and 10 short names, baseline reasoning is therefore the appropriate default until a prospective study of the target task demonstrates reproducible incremental value\. The exploratory exception for portfolio width shows why that recommendation should not be transferred unchanged to a different portfolio rule\.

## References

- Agrawalet al\.\(2026\)Y\. Agrawal, A\. Dutta, M\. M\. Hasan, S\. Karmaker, and A\. DuttaFinTradeBench: a financial reasoning benchmark for LLMs\.arXiv preprint arXiv:2603\.19225\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.19225),[Link](https://arxiv.org/abs/2603.19225)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Alomraniet al\.\(2025\)M\. A\. Alomrani, Y\. Zhang, D\. Li, Q\. Sun, S\. Pal, Z\. Zhang, Y\. Hu,et al\.Reasoning on a budget: a survey of adaptive and controllable test\-time compute in LLMs\.arXiv preprint arXiv:2507\.02076\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2507.02076),[Link](https://arxiv.org/abs/2507.02076)Cited by:[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p2.1)\.
- Araci \(2019\)D\. AraciFinBERT: financial sentiment analysis with pre\-trained language models\.arXiv preprint arXiv:1908\.10063\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1908.10063),[Link](https://arxiv.org/abs/1908.10063)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Baileyet al\.\(2016\)D\. H\. Bailey, J\. M\. Borwein, M\. Lopez de Prado, and Q\. J\. ZhuThe probability of backtest overfitting\.The Journal of Computational Finance20\(4\),pp\. 39–69\.External Links:[Document](https://dx.doi.org/10.21314/JCF.2016.322)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Bailey and Lopez de Prado \(2014\)D\. H\. Bailey and M\. Lopez de PradoThe deflated sharpe ratio: correcting for selection bias, backtest overfitting, and non\-normality\.The Journal of Portfolio Management40\(5\),pp\. 94–107\.External Links:[Document](https://dx.doi.org/10.3905/jpm.2014.40.5.094)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Chenet al\.\(2025\)Y\. Chen, Z\. Yao, Y\. Liu, A\. Xin, J\. Ye, J\. Yu, L\. Hou, and J\. LiStockBench: can LLM agents trade stocks profitably in real\-world markets?\.arXiv preprint arXiv:2510\.02209\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2510.02209),[Link](https://arxiv.org/abs/2510.02209)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p2.1)\.
- Chenet al\.\(2022\)Y\. Chen, B\. T\. Kelly, and D\. XiuExpected returns and large language models\.SSRN Electronic Journal\.External Links:[Document](https://dx.doi.org/10.2139/ssrn.4416687),[Link](https://ssrn.com/abstract=4416687)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p2.1)\.
- Chenet al\.\(2021\)Z\. Chen, W\. Chen, C\. Smiley, S\. Shah, I\. Borova, D\. Langdon, R\. Moussa, M\. Beane, T\. Huang, B\. Routledge, and W\. Y\. WangFinQA: a dataset of numerical reasoning over financial data\.InProceedings of EMNLP,External Links:[Link](https://arxiv.org/abs/2109.00122)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Chenget al\.\(2026\)I\. Cheng, M\. Granger, J\. Shi, and V\. StrelaAgents are not algorithms: the tradeoffs of decision\-time reasoning in AI trading\.SSRN Electronic Journal\.External Links:[Document](https://dx.doi.org/10.2139/ssrn.6713620),[Link](https://ssrn.com/abstract=6713620)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p3.1)\.
- DeepSeek\-AI \(2025\)DeepSeek\-AIDeepSeek\-R1: incentivizing reasoning capability in LLMs via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.12948),[Link](https://arxiv.org/abs/2501.12948)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p1.1)\.
- Guet al\.\(2020\)S\. Gu, B\. Kelly, and D\. XiuEmpirical asset pricing via machine learning\.The Review of Financial Studies33\(5\),pp\. 2223–2273\.External Links:[Document](https://dx.doi.org/10.1093/rfs/hhaa009)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Hansen \(2005\)P\. R\. HansenA test for superior predictive ability\.Journal of Business & Economic Statistics23\(4\),pp\. 365–380\.External Links:[Document](https://dx.doi.org/10.1198/073500105000000063)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Harveyet al\.\(2016\)C\. R\. Harvey, Y\. Liu, and H\. Zhu… And the cross\-section of expected returns\.The Review of Financial Studies29\(1\),pp\. 5–68\.External Links:[Document](https://dx.doi.org/10.1093/rfs/hhv059)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Kunsch \(1989\)H\. R\. KunschThe jackknife and the bootstrap for general stationary observations\.The Annals of Statistics17\(3\),pp\. 1217–1241\.External Links:[Document](https://dx.doi.org/10.1214/aos/1176347265)Cited by:[§4\.3](https://arxiv.org/html/2609.30705#S4.SS3.p2.1)\.
- Liet al\.\(2025a\)H\. Li, Y\. Cao, Y\. Yu, S\. R\. Javaji, Z\. Deng, Y\. He,et al\.INVESTORBENCH: a benchmark for financial decision\-making tasks with LLM\-based agent\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics,External Links:[Link](https://arxiv.org/abs/2412.18174)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p2.1)\.
- Liet al\.\(2025b\)X\. Li, Y\. Zeng, X\. Xing, J\. Xu, and X\. XuProfit mirage: revisiting information leakage in LLM\-based financial agents\.arXiv preprint arXiv:2510\.07920\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2510.07920),[Link](https://arxiv.org/abs/2510.07920)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p2.1)\.
- Lightmanet al\.\(2023\)H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. CobbeLet’s verify step by step\.arXiv preprint arXiv:2305\.20050\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2305.20050),[Link](https://arxiv.org/abs/2305.20050)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p1.1)\.
- Lo and MacKinlay \(1990\)A\. W\. Lo and A\. C\. MacKinlayData\-snooping biases in tests of financial asset pricing models\.The Review of Financial Studies3\(3\),pp\. 431–467\.External Links:[Document](https://dx.doi.org/10.1093/rfs/3.3.431)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Lopez\-Lira and Tang \(2023\)A\. Lopez\-Lira and Y\. TangCan ChatGPT forecast stock price movements? return predictability and large language models\.arXiv preprint arXiv:2304\.07619\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2304.07619),[Link](https://arxiv.org/abs/2304.07619)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p2.1)\.
- McLean and Pontiff \(2016\)R\. D\. McLean and J\. PontiffDoes academic research destroy stock return predictability?\.The Journal of Finance71\(1\),pp\. 5–32\.External Links:[Document](https://dx.doi.org/10.1111/jofi.12365)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Muennighoffet al\.\(2025\)N\. Muennighoff, Z\. Yang, W\. Shi, X\. L\. Li, L\. Fei\-Fei, H\. Hajishirzi, L\. Zettlemoyer, P\. Liang, E\. Candes, and T\. Hashimotos1: simple test\-time scaling\.arXiv preprint arXiv:2501\.19393\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.19393),[Link](https://arxiv.org/abs/2501.19393)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p1.1),[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p1.1)\.
- Newey and West \(1987\)W\. K\. Newey and K\. D\. WestA simple, positive semi\-definite, heteroskedasticity and autocorrelation consistent covariance matrix\.Econometrica55\(3\),pp\. 703–708\.External Links:[Document](https://dx.doi.org/10.2307/1913610)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p6.1),[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p2.1),[§4\.3](https://arxiv.org/html/2609.30705#S4.SS3.p1.1)\.
- Novy\-Marx \(2015\)R\. Novy\-MarxBacktesting strategies based on multiple signals\.Technical reportTechnical Report21329,National Bureau of Economic Research\.External Links:[Document](https://dx.doi.org/10.3386/w21329),[Link](https://www.nber.org/papers/w21329)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Panet al\.\(2026\)Z\. Pan, K\. Zhang, Y\. Zhao, and Y\. HanAdaptive model and strategy routing for cost\-efficient LLM services\.InProceedings of the ACM Web Conference,pp\. 5568–5578\.External Links:[Document](https://dx.doi.org/10.1145/3774904.3792556),[Link](https://doi.org/10.1145/3774904.3792556)Cited by:[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p2.1)\.
- Sitjar \(2026\)T\. N\. SitjarReasoning into confidence: a controlled single\-trader ablation of in\-decision chain\-of\-thought on LLM trading performance and calibration\.SSRN Electronic Journal\.External Links:[Document](https://dx.doi.org/10.2139/ssrn.6946379),[Link](https://ssrn.com/abstract=6946379)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p3.1)\.
- Snellet al\.\(2025\)C\. Snell, J\. Lee, K\. Xu, and A\. KumarScaling LLM test\-time compute optimally can be more effective than scaling parameters for reasoning\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2408.03314)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p1.1),[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p1.1)\.
- Sodha \(2025\)R\. H\. SodhaWhen reasoning fails: evaluating “thinking” LLMs for stock prediction\.arXiv preprint arXiv:2511\.08608\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.08608),[Link](https://arxiv.org/abs/2511.08608)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p3.1)\.
- Sullivanet al\.\(1999\)R\. Sullivan, A\. Timmermann, and H\. WhiteData\-snooping, technical trading rule performance, and the bootstrap\.The Journal of Finance54\(5\),pp\. 1647–1691\.External Links:[Document](https://dx.doi.org/10.1111/0022-1082.00163)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Tetlock \(2007\)P\. C\. TetlockGiving content to investor sentiment: the role of media in the stock market\.The Journal of Finance62\(3\),pp\. 1139–1168\.External Links:[Document](https://dx.doi.org/10.1111/j.1540-6261.2007.01232.x)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2203.11171)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p1.1),[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 24824–24837\.External Links:[Link](https://arxiv.org/abs/2201.11903)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p1.1),[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p1.1)\.
- White \(2000\)H\. WhiteA reality check for data snooping\.Econometrica68\(5\),pp\. 1097–1126\.External Links:[Document](https://dx.doi.org/10.1111/1468-0262.00152)Cited by:[§2\.4](https://arxiv.org/html/2609.30705#S2.SS4.p1.1)\.
- Wuet al\.\(2023\)S\. Wu, O\. Irsoy, S\. Lu, V\. Dabravolski, M\. Dredze, S\. Gehrmann, P\. Kambadur, D\. Rosenberg, and G\. MannBloombergGPT: a large language model for finance\.arXiv preprint arXiv:2303\.17564\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2303.17564),[Link](https://arxiv.org/abs/2303.17564)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Xiaet al\.\(2026\)Y\. Xia, P\. You, T\. Wang, F\. Liu, H\. Qi, X\. Wu, and S\. ZhangAgentic trading: when LLM agents meet financial markets\.arXiv preprint arXiv:2605\.19337\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.19337),[Link](https://arxiv.org/abs/2605.19337)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p2.1)\.
- Xiaoet al\.\(2024\)Y\. Xiao, E\. Sun, D\. Luo, and W\. WangTradingAgents: multi\-agents LLM financial trading framework\.arXiv preprint arXiv:2412\.20138\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2412.20138),[Link](https://arxiv.org/abs/2412.20138)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p1.1)\.
- Xieet al\.\(2024\)Q\. Xie, W\. Han, Z\. Chen, R\. Xiang, X\. Zhang,et al\.FinBen: a holistic financial benchmark for large language models\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Link](https://arxiv.org/abs/2402.12659)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Xieet al\.\(2023\)Q\. Xie, W\. Han, X\. Zhang, Y\. Lai, M\. Peng, A\. Lopez\-Lira, and J\. HuangPIXIU: a large language model, instruction data and evaluation benchmark for finance\.arXiv preprint arXiv:2306\.05443\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2306.05443),[Link](https://arxiv.org/abs/2306.05443)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Yanget al\.\(2023\)H\. Yang, X\. Liu, and C\. D\. WangFinGPT: open\-source financial large language models\.arXiv preprint arXiv:2306\.06031\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2306.06031),[Link](https://arxiv.org/abs/2306.06031)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Yanget al\.\(2024\)H\. Yang, B\. Zhang, N\. Wang, C\. Guo, X\. Zhang, L\. Lin, J\. Wang, T\. Zhou, M\. Guan, R\. Zhang, and C\. D\. WangFinRobot: an open\-source AI agent platform for financial applications using large language models\.arXiv preprint arXiv:2405\.14767\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2405.14767),[Link](https://arxiv.org/abs/2405.14767)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p1.1)\.
- Yanget al\.\(2020\)Y\. Yang, M\. C\. S\. Uy, and A\. HuangFinBERT: a pretrained language model for financial communications\.arXiv preprint arXiv:2006\.08097\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2006.08097),[Link](https://arxiv.org/abs/2006.08097)Cited by:[§2\.1](https://arxiv.org/html/2609.30705#S2.SS1.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, D\. Yu, J\. Zhao, I\. Shafran, T\. L\. Griffiths, Y\. Cao, and K\. NarasimhanTree of thoughts: deliberate problem solving with large language models\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://arxiv.org/abs/2305.10601)Cited by:[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p1.1)\.
- Yonget al\.\(2025\)X\. Yong, X\. Zhou, Y\. Zhang, J\. Li, Y\. Zheng, and X\. WuThink or not? exploring thinking efficiency in large reasoning models via an information\-theoretic lens\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Document](https://dx.doi.org/10.52202/085713-0098),[Link](https://papers.neurips.cc/paper_files/paper/2025/hash/04185b5ae2d450ef39bd53c0ec4802cb-Abstract-Conference.html)Cited by:[§2\.3](https://arxiv.org/html/2609.30705#S2.SS3.p2.1)\.
- Yuet al\.\(2023\)Y\. Yu, H\. Li, Z\. Chen, Y\. Jiang, Y\. Li, D\. Zhang, R\. Liu, J\. W\. Suchow, and K\. KhashanahFinMem: a performance\-enhanced LLM trading agent with layered memory and character design\.arXiv preprint arXiv:2311\.13743\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2311.13743),[Link](https://arxiv.org/abs/2311.13743)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p1.1)\.
- Yuet al\.\(2024\)Y\. Yu, Z\. Yao, H\. Li, Z\. Deng, Y\. Jiang, Y\. Cao, Z\. Chen,et al\.FinCon: a synthesized LLM multi\-agent system with conceptual verbal reinforcement for enhanced financial decision making\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2407.06567)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p1.1)\.
- Zhanget al\.\(2024\)W\. Zhang, L\. Zhao, H\. Xia, S\. Sun, J\. Sun, M\. Qin,et al\.A multimodal foundation agent for financial trading: tool\-augmented, diversified, and generalist\.arXiv preprint arXiv:2402\.18485\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2402.18485),[Link](https://arxiv.org/abs/2402.18485)Cited by:[§1](https://arxiv.org/html/2609.30705#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p1.1)\.
- Zhuet al\.\(2026\)T\. Zhu, W\. Zhao, R\. Sun, B\. Luan, J\. Lu, S\. Wang,et al\.From knowing to doing: a memory\-controlled benchmark for LLM trading agents on stock markets\.arXiv preprint arXiv:2605\.28359\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.28359),[Link](https://arxiv.org/abs/2605.28359)Cited by:[§2\.2](https://arxiv.org/html/2609.30705#S2.SS2.p2.1)\.

## Appendix APrompt and Structured Output Contract

The system instruction casts the model as a conservative cross\-sectional equity forecaster, treats supplied article text as untrusted data, restricts the response to the supplied inputs, and requires exactly one structured text object in JavaScript Object Notation \(JSON\)\. The user prompt asks for relative return over the next five trading sessions\. It instructs the model to favor persistent evidence across horizons and penalize reversal, unstable volatility, illiquidity, contradiction, and uncertainty\. In news tasks it distinguishes new material information about the target from repetition, commentary, price narration, legal advertising, and noise\.

Each of the 25 expected opaque identifiers must appear exactly once in:

```
{"items":[
  {"asset":"A0123456789",
   "score":0.0,
   "confidence":0.0,
   "materiality":0.0,
   "uncertainty":0.0,
   "abstain":false}
]}
```

The exact prompts are stored as compressed immutable artifacts with hashes\. Schema syntax differs by provider, but the fields and bounds are common\.

### A\.1\.Reasoning Intervention Settings

Table 4\.Provider controls defining the reasoning intervention\. Prompt, inputs, schema, candidate order, and output limit remain fixed within each backbone\.DeepSeek disables thinking for none and enables it with the requested effort otherwise; GPT variesreasoning\.effort; Gemini variesthinkingLevel\. No condition receives an additional reasoning instruction\. Provider labels and output limits are not directly comparable quantities of computation\. Reasoning summaries or traces are retained for auditing but are not treated as faithful records of hidden computation\.

### A\.2\.Numerical Data Provenance and Volume

The numerical condition is derived from the Center for Research in Security Prices \(CRSP\) daily stock files accessed through WRDS\. The source fields include daily prices, returns, trading volume, opening prices, high and low price fields, share codes, exchange codes, and delisting records\. Dated security name intervals restrict the sample to ordinary common shares listed on the NYSE, AMEX, or Nasdaq\. Reported returns are combined with delisting returns\. Under the frozen cleaning rule, when a delisting code from 500 to 599 has no reported delisting return, the return is set to−30%\-30\\%; these codes indicate a delisting related to poor performance\. Other missing delisting returns remain missing\. All model features end at the preceding trading session’s close\.

Raw CRSP files from 2020 through 2024 support the required histories\. The resulting feature construction panel fixes 1,000 liquid securities using information ending December 30, 2022, and contains 974,641 security and date rows over 1,005 trading dates from January 4, 2021, through December 31, 2024\. Feature ranks are computed within this broader panel before values are retained for the 100 active candidates\. Across the 241 formation dates in the experiment, the models receive 24,100 stock and date numerical records\. Each record contains 16 ranked features and four summaries per feature, for 64 values per stock and 1,542,400 values in the unique 2024 numerical input panel\. The same records are reused across reasoning levels and model families, so repeated calls do not create additional market observations\.

### A\.3\.News Data Provenance and Volume

The news conditions use the licensed EOD Historical Data \(EODHD\) news feed\. Articles were linked to the active CRSP security universe through unambiguous provider ticker mappings\. For each formation date, an article was eligible only if its provider timestamp fell from 4:00 p\.m\. Eastern Time on the preceding trading session up to, but not including, 9:00 a\.m\. Eastern Time on the formation date\. Duplicate query returns were normalized to stable article identifiers, and the three most recent eligible articles for each stock and formation date were retained\. Each model received the publisher, title, elapsed hours before the decision, and full article text, subject to fixed context limits of 9,000 characters per article and 18,000 characters per stock\. Identifiable and masked news used the same articles\. Masking replaced issuer names, tickers, explicit calendar dates and years, and the publisher label while otherwise preserving the supplied content\.

Across 241 formation dates and 100 stocks per date, 16,916 of 24,100 stock and date combinations \(70\.2%\) contained at least one eligible article, while 7,184 \(29\.8%\) contained none\. The retained packets contained 38,097 article references linked to stocks, representing 29,444 unique provider article identifiers and 13 publisher labels\. Among the 16,916 nonempty packets, 4,623 contained one article, 3,405 contained two, and 8,888 contained three\. A source article linked to more than one stock is counted once among unique identifiers but once for each stock among the 38,097 references\.

## Appendix BComplete DeepSeek Confirmatory Chain

Table 5\.DeepSeek audit across three generations, bps/day\. Return date remains the market sampling dimension\.The secondary adjacent contrasts \(point estimates in bps/day\) are: numerical, low–none−0\.529\-0\.529, high–low−2\.839\-2\.839, maximum–high−1\.507\-1\.507; identifiable news,−8\.525\-8\.525,−0\.167\-0\.167,\+4\.639\+4\.639; and masked news,−9\.016\-9\.016,\+1\.289\+1\.289,−0\.519\-0\.519\. Only maximum–high for identifiable news has a two\-sided 95% CI entirely above zero\.

## Appendix CSensitivity to Failure Handling

GPT’s 57 invalid primary tasks and Gemini’s 136 invalid primary tasks remain imputed as neutral in the intention\-to\-treat estimate\. Two frozen sensitivity analyses test this choice\. The*complete date analysis*removes every formation date containing any failure\. The*matched chunk analysis*removes the same chunk of 25 assets from both reasoning levels whenever either level fails\. For both GPT and Gemini, all three information conditions remain inconclusive under both analyses\. Thus failure handling is operationally important but does not determine the headline return conclusion\.

DeepSeek’s deterministic repair procedure affected 17 of 16,176 tasks\. Across the backbone evaluation, 15 omitted asset rows were imputed as neutral and seven asset identifier errors with an edit distance of one were corrected when the intended match was unique\. Repairs were completed without reading outcomes, and the raw attempts and amendments are retained\.

## Appendix DBootstrap Robustness

Table 6\.Paired bootstrap intervals at a one way cost of 10 bps\. MBB\-5 is the moving\-block bootstrap with five sessions per block; cohort resamples complete formation contributions\. Each uses 50,000 deterministic resamples\.These supplementary checks do not replace the primary HAC analysis\. All nine annual contrasts remain inconclusive, and the audit contrast remains harmful\. Moving\-block sensitivities at 10 and 20 sessions also leave every annual contrast inconclusive and the audit harmful\. Draws in a machine readable format, cluster contributions, exact centered bootstrappp\-values, and deterministic seeds are retained in the robustness artifact\.

## Appendix ESensitivity to Transaction Costs

The portfolio simulator replays every condition at one way costs of 5, 10, 20, and 50 bps\. GPT’s low versus baseline effects remain inconclusive at every cost\. Its numerical estimate ranges from\+3\.530\+3\.530bps/day at 5 bps cost to\+3\.619\+3\.619at 50 bps; its identifiable news estimate ranges from−0\.647\-0\.647to−0\.460\-0\.460; and its estimate for masked news ranges from\+0\.004\+0\.004to\+0\.332\+0\.332\. Gemini likewise remains inconclusive: numerical effects range from\+0\.721\+0\.721to\+0\.762\+0\.762, identifiable news ranges from−0\.060\-0\.060to−0\.803\-0\.803, and masked news ranges from−1\.574\-1\.574to−1\.066\-1\.066\. The small changes occur because reasoning levels induce different turnover\. Full confidence intervals and portfolio paths are provided in evaluation files that machines can read\.

## Appendix FAbsolute Performance and Sensitivity to Portfolio Width

Table[7](https://arxiv.org/html/2609.30705#A6.T7)supplies economic context for the paired effects\. Each cell reports baseline/low reasoning; Gemini’s baseline is minimal\. Net is mean return after trading costs, the Sharpe ratio scales mean return by return volatility, maximum drawdown \(MDD\) is the largest loss from a peak to a trough, and turnover is the average fraction of capital traded in one direction\. Absolute levels are descriptive historical backtests and do not establish prospective profitability\.

Table 7\.Absolute portfolio context at a one way cost of 10 bps\. Each pair is baseline/low reasoning; metrics are defined immediately above\.
Table[8](https://arxiv.org/html/2609.30705#A6.T8)reuses the frozen scores while varying only the number of selected names per side\. The result for 10 long and 10 short names remains primary\. One of 36 exploratory cells meets the unadjusted rule for material benefit: DeepSeek masked news at 20 names per side\. Its HAC 95% CI is\[0\.597,10\.942\]\[0\.597,10\.942\]bps/day and its interval from a moving\-block bootstrap with five sessions per block is\[0\.690,11\.223\]\[0\.690,11\.223\]\. Across the family of 36 cells, however, no HAC or moving\-blockpp\-value survives Holm adjustment; both adjustedpp\-values for this cell equal 1\.000\. The sweep therefore demonstrates sensitivity to portfolio width rather than a confirmatory reasoning benefit\.

Table 8\.Sensitivity to portfolio width: net return for low relative to baseline reasoning \(bps/day\) at a one way cost of 10 bps\. Columns give names selected per side\. An asterisk marks the sole unadjusted cell with material benefit\.

## Appendix GPrecision and Detectable Effects

Table[9](https://arxiv.org/html/2609.30705#A7.T9)reports a design diagnostic based on the observed paired HAC standard error\. The*minimum detectable effect*\(MDE\) is the smallest true effect that a design of this precision would detect with a chosen probability; the two\-sided 80% MDE is\(z\.975\+z\.80\)​SE^\(z\_\{\.975\}\+z\_\{\.80\}\)\\widehat\{\\mathrm\{SE\}\}\. This calculation summarizes the precision of the realized design under its estimated dependence structure; it is not an equivalence test and does not convert an inconclusive contrast into a zero effect\.

Table 9\.HAC precision diagnostics for the nine primary low versus baseline contrasts, in bps/day; MDE is minimum detectable effect\.

## Appendix HAgreement Across Repeated Generations

*Spearman rank correlation*\(ρ\\rho\) measures how similarly two generations order all assets\.*Tail Jaccard*is the fraction of names shared by their combined selected long and short sets\.*Identical tails*records whether both selected sets match exactly\.

Table 10\.Mean pairwise generation agreement across the three generation pairs\.
Each pairwise agreement calculation contains 4,800 asset rows per level and information condition\. No audited GPT or Gemini formation date reproduces the exact same tails in any generation pair\.

## Appendix IAnalysis by Calendar Period

Table[11](https://arxiv.org/html/2609.30705#A9.T11)partitions the nine primary daily effects of low relative to baseline using predefined calendar boundaries\. Values are descriptive subgroup means in bps/day; the annual estimates and inference remain primary\.

Table 11\.Effects by calendar period in 2024, in bps/day\. H1 is January–June, H2 is July–December, and Q1–Q4 are calendar quarters, assigned by portfolio return date\.
DeepSeek numerical has a raw heterogeneityp=\.014p=\.014for H2 minus H1, and Gemini masked news has a rawp=\.045p=\.045for equality across quarters; their values after Holm adjustment across the family of 18 tests are\.246\.246and\.757\.757\. No heterogeneity test survives the family correction\. The largest quarter accounts for 34\.2%–64\.5% of each contrast’s sum of absolute quarterly contributions, so no annual estimate is wholly produced by one quarter\. However, only GPT numerical retains the same sign in both halves, and seven annual directions flip when the estimate is recomputed after omitting at least one quarter\. These diagnostics show temporal breadth and instability within 2024, not independent evidence across market years\.

## Appendix JProtocol Integrity and Claim Boundaries

Table[12](https://arxiv.org/html/2609.30705#A10.T12)summarizes the fixed study status\. All three backbones and low versus baseline contrasts were predetermined, and each backbone evaluation froze its inputs, prompts, schema, failure policy, portfolio mapping, and estimand before outcome evaluation\.

Table 12\.Protocol status of the three completed backbone evaluations\. Operational amendments did not change reasoning levels or return analysis\.
Additional disclosures are required for interpretation:

- •GPT’s primary evaluation used 3,120 Batch tasks and 736 synchronous Responses tasks under a frozen delivery amendment made without viewing outcomes\.*Delivery mode*means how requests were sent to the provider; because it changes with calendar time, it is not a causal treatment\.
- •Every backbone has three generations for every information condition on the same audit sample of 48 formation dates\. The repeated calls measure output and portfolio stability without changing the primary annual estimands or the number of return dates\.
- •Gemini 3\.1 Flash\-Lite passed the frozen admission gate for identifier fidelity, a check that it reproduces each opaque asset code exactly, before backbone evaluation\.
- •Gemini amendments to state labels, response files, concurrency \(the number of simultaneous requests\), and orchestration \(request scheduling\) altered transport only\. Prompts, tasks, reasoning levels, failure handling, budgets, and the evaluator were unchanged\.

Artifact availability and responsible use\.The authors retain complete evaluation records and plan to release the evaluation code, prompts, result summaries, evaluator hashes, and artifact manifest in a public GitHub repository\. Licensed market and news data, together with provider outputs restricted by applicable terms, will not be redistributed\. This retrospective simulation placed no live trades and recruited no human participants\. It is not investment advice or evidence of prospective profitability\. Any public release will preserve the historical backtest label, documented claim boundaries, and applicable licensing restrictions\.

相似文章

自主交易系统能否为其智能行为买单?

arXiv cs.AI

本文介绍了TradeLens,一个基于轨迹的诊断工具包,用于评估基于LLM的自主交易系统是否将其推理和工具使用成本转化为可衡量的增量利润,并分析了DeepSeek-V3.2和GLM-4.7等模型的失败模式。