Semantic State Abstraction Interfaces for LLM-Augmented Portfolio Decisions: Multi-Axis News Decomposition and RL Diagnostics
Summary
This paper introduces Semantic State Abstraction Interfaces (SSAI) to separate representation hypotheses from optimization variance in LLM-augmented portfolio decisions. It concludes that SSAI's apparent advantage is largely a basket-selection effect, with dense encodings and principal components performing better empirically.
View Cached Full Text
Cached at: 05/11/26, 06:44 AM
# Semantic State Abstraction Interfaces for LLM-Augmented Portfolio Decisions: Multi-Axis News Decomposition and RL Diagnostics
Source: [https://arxiv.org/html/2605.06730](https://arxiv.org/html/2605.06730)
Likhita Yerra AIVANCITY School of AI & Data likhita\.yerra@aivancity\.education&Remi Uttejitha Allam AIVANCITY School of AI & Data remi\.allam@aivancity\.education
###### Abstract
We introduceSemantic State Abstraction Interfaces \(SSAI\): a methodological template for mapping sparse unstructured text intoKKauditable, named coordinates with neutral defaults on no\-news days, designed to separate representation hypotheses from optimisation variance in sequential decision systems\. Our contribution is the framework and its evaluation protocol—not a claim that SSAI outperforms denser alternatives\.
We instantiate SSAI withK=4K\{=\}4axes \(*sentiment*,*risk*,*confidence*,*volatility forecast*\) on a US\-equity panel \(30 NASDAQ\-100 names, FNSPID news, 2019–2023 test\), and evaluate it across three parallel estimators—direct factor portfolios \(SFP/SRF/SCW\), supervised ridge forecasters, and RL agents \(DP\-PPO, SAC\)—that share the same fixedϕ\\phiso that signal and optimiser effects can be read off separately\.
What the experiments show\.The four\-factor SFP reaches 307\.2% CR / Sharpe 1\.067 over 2019–2023\.*However*, this apparent advantage over buy\-and\-hold \(243\.6%\) does not survive its own controls: SFP underperforms stratum\-matched B&H in every coverage tercile \(Table[3](https://arxiv.org/html/2605.06730#S5.T3)\), reverses at≥0\.2\{\\geq\}0\.2% per\-trade costs, and the four\-factor vs\. sentiment\-only daily advantage is non\-significant \(Wilcoxonp=0\.556p\{=\}0\.556\)\. A data\-driven first principal component of the four axes \(PC1\-SFP\) achieves 433\.6% CR under the same portfolio rule—126pp above SSAI—and a FinBERT direct\-portfolio baseline achieves 386\.3% \(Table[5](https://arxiv.org/html/2605.06730#S5.T5)\)\. The honest summary is that the SSAI portfolio advantage is substantially a*basket\-selection composition effect*, and both PC1 and dense neural encodings are empirically stronger ranking signals in this setting\.
What the experiments also show\.Naively appending sparse LLM columns hurts ridge forecasting; high\-conviction semantic tilts and lexical baselines \(VADER, TF–IDF/SVD\) recover price\-only Sharpe\. The RL block is a diagnostic: DP\-PPO trails buy\-and\-hold; SAC with identical SSAI observations improves Sharpe \(1\.128 vs\. 1\.032\), isolating algorithm dependence rather than representation advantage\. Seed\-mean Sharpe differences across 21 DP\-PPO seeds are non\-significant \(Wilcoxonp≈0\.25p\{\\approx\}0\.25–0\.310\.31\)\.
Contribution framing\.We present SSAI as an*interpretability\-performance frontier*instrument and a*cautionary diagnostic*: it characterises the cost of maintaining auditable axes \(126pp CR vs\. PC1 over five years\), shows that RL performance is dominated by algorithm choice rather than representation, and provides a reusable template for separating semantic signal from optimisation noise in other sparse\-text decision settings\.
## 1Introduction
Deep Reinforcement Learning \(DRL\) is widely studied for sequential decision problems where observations mix numerical signals with unstructured evidence\(Mnih et al\.,[2015](https://arxiv.org/html/2605.06730#bib.bib12); Schulman et al\.,[2017](https://arxiv.org/html/2605.06730#bib.bib15); Liu et al\.,[2021](https://arxiv.org/html/2605.06730#bib.bib7)\)\. Large language models \(LLMs\) offer flexible encoders for text\(Brown et al\.,[2020](https://arxiv.org/html/2605.06730#bib.bib3); Liu et al\.,[2023](https://arxiv.org/html/2605.06730#bib.bib9); Wu et al\.,[2023](https://arxiv.org/html/2605.06730#bib.bib17)\), yet most pipelines obscure a basic design question: should textual evidence appear as a high\-dimensional latent vector, as a single scalar \(e\.g\., sentiment\), or as a small inventory of named semantic coordinates—especially when text arrives only on a sparse subset of decision dates?
Prior work has connected NLP with RL trading in limited ways\.Yang et al\. \([2020a](https://arxiv.org/html/2605.06730#bib.bib18)\)use technical indicators as the sole state representation;Benhenda \([2025](https://arxiv.org/html/2605.06730#bib.bib2)\)add a single sentiment score from DeepSeek\-V3; andLopez\-Lira and Tang \([2023](https://arxiv.org/html/2605.06730#bib.bib10)\)use ChatGPT to classify daily market sentiment as a binary signal\. A single sentiment score is information\-sparse: it conflates orthogonal concepts—whether a headline is positive differs fundamentally from whether the underlying situation is*risky*or whether the market is likely to be*volatile*\.
#### Contribution\.
We introduceSemantic State Abstraction Interfaces \(SSAI\): mapsϕ\\phifrom document sets𝒩s,d\\mathcal\{N\}\_\{s,d\}toKKauditable axes elicited by fixed prompts from a frozen LLM, with a neutral default when𝒩s,d=∅\\mathcal\{N\}\_\{s,d\}=\\emptyset\(Definition 1, Appendix[A](https://arxiv.org/html/2605.06730#A1)\)\. The core contribution is anevaluation design that separates representation from optimisation by construction: a fixed, sharedϕ\\phiis evaluated simultaneously by low\-variance estimators \(factor portfolios, ridge forecasters\) and high\-variance nonlinear estimators \(DP\-PPO, SAC\), making both effects separately observable\. SSAI axes are auditable and support prompt\-level ablation at the cost of predictive performance; this constraint is a methodological choice\. All empirical validation is on the US equity panel above; Definition 1 is domain\-agnostic \(Appendix[D](https://arxiv.org/html/2605.06730#A4)\)\.
#### Research questions\.
We organise the experiments around three concrete questions:RQ1: do multi\-axis LLM signals change policy behaviour relative to neutral\-masked inputs?RQ2: are any LLM axes aligned, even weakly, with realised market quantities?RQ3: is observed performance primarily a property of the semantic state representation, or of the learning algorithm used to exploit it? Results are conservative: four\-factor SFP has a 15pp CR edge over sentiment\-only \(non\-significant; composition artefact\); supervised LLM features help only as sparse high\-conviction tilts; RL masking is directional but underpowered; SAC substantially improves Sharpe over DP\-PPO under identical SSAI states\.
We make three contributions, each diagnostic rather than performance\-competitive in nature:
1. 1\.SSAI template \(reusable framework\)\.Definition 1 provides a domain\-agnostic, auditable interface for sparse text in sequential decision systems, instantiated here withK=4K\{=\}4integer\-valued axes and neutral defaults\. The template is designed for settings requiring axis\-level auditability, human\-in\-the\-loop override, or prompt\-perturbation testing—not for settings where maximising cumulative return is the sole objective\.
2. 2\.Diagnostic evaluation protocol with cautionary findings\.Three parallel estimators \(SFP/SRF/SCW, ridge forecasters, RL agents\) share the sameϕ\\phi, isolating representation from optimisation effects\. The main finding is*cautionary*: the 63pp SFP vs\. B&H gap is a composition artefact \(fails within\-stratum controls, reverses at≥0\.2%\{\\geq\}0\.2\\%costs\); a PC1 of the four axes outperforms SSAI by 126pp CR; a FinBERT direct\-portfolio baseline outperforms SSAI by 79pp CR\. These are the paper’s most important empirical results\.
3. 3\.Interpretability\-performance frontier characterisation\.Table[5](https://arxiv.org/html/2605.06730#S5.T5)places SSAI on the frontier alongside PC1 and FinBERT\-SFP, making the cost of interpretability constraints precisely measurable: 126pp CR over five years relative to the best alternative using the same daily top\-10 rule\. This frontier characterisation—not a performance claim—is the proposed contribution to the LLM\-augmented decision\-making literature\.
## 2Related Work
Liu et al\. \([2021](https://arxiv.org/html/2605.06730#bib.bib7)\)introduce FinRL with PPO and SAC baselines\(Yang et al\.,[2020b](https://arxiv.org/html/2605.06730#bib.bib19); Liu et al\.,[2022](https://arxiv.org/html/2605.06730#bib.bib8)\); risk\-constrained RL has used CVaR objectives\(Tamar et al\.,[2015](https://arxiv.org/html/2605.06730#bib.bib16)\)and Lagrangian relaxation\(Ray et al\.,[2019](https://arxiv.org/html/2605.06730#bib.bib13)\)\. FinGPT\(Liu et al\.,[2023](https://arxiv.org/html/2605.06730#bib.bib9)\)and BloombergGPT\(Wu et al\.,[2023](https://arxiv.org/html/2605.06730#bib.bib17)\)fine\-tune LLMs on financial corpora;Benhenda \([2025](https://arxiv.org/html/2605.06730#bib.bib2)\)add a scalar DeepSeek\-V3 sentiment score to RL state\. We extend prior scalar sentiment interfaces\(Benhenda,[2025](https://arxiv.org/html/2605.06730#bib.bib2); Lopez\-Lira and Tang,[2023](https://arxiv.org/html/2605.06730#bib.bib10)\)with SSAIs instantiated through zero\-shot multi\-axis prompting \(no encoder fine\-tuning\) and attach diagnostic evaluations across portfolios, supervised learners, and RL\. News data are from FNSPID\(Dong et al\.,[2024](https://arxiv.org/html/2605.06730#bib.bib4)\)\.
Positioning\.Unlike representation learning\(Bengio et al\.,[2013](https://arxiv.org/html/2605.06730#bib.bib1)\)or disentangled methods\(Higgins et al\.,[2017](https://arxiv.org/html/2605.06730#bib.bib5)\)that trainϕ\\phiend\-to\-end, SSAI*fixes*ϕ\\phiante\-hoc via prompt design, achieving named\-axis interpretability without encoder training\. Unlike post\-hoc methods \(LIME\(Ribeiro et al\.,[2016](https://arxiv.org/html/2605.06730#bib.bib14)\), SHAP\(Lundberg and Lee,[2017](https://arxiv.org/html/2605.06730#bib.bib11)\)\), SSAI requires no separate explanation step: axes are interpretable by construction\. The key gap SSAI fills is the*shared interface*: fixingϕ\\phiacross estimators of different complexity makes representation and optimisation effects separately observable\.
Interpretability senses\.We use “interpretability” in three distinct senses throughout, which we now make explicit\.*\(a\) Axis\-level human\-readable semantics*: each axis has a fixed natural\-language name and integer range, allowing a practitioner to read off “the model raised risk from 2 to 4 on this date” without post\-hoc processing\.*\(b\) Post\-hoc auditability via prompt perturbation*: fixingϕ\\phienables controlled counterfactual tests—one can rescore with a modified prompt and observe the effect on downstream decisions, something impossible with a jointly trained encoder\.*\(c\) Regulatory legibility*: integer scores on named axes are explainable to a compliance team in plain language; a PCA component loading is not\. SSAI delivers properties \(a\)–\(c\) by construction\. PC1 and FinBERT deliver none of \(a\)–\(c\)\. The interpretability\-performance analysis in Table[5](https://arxiv.org/html/2605.06730#S5.T5)and Section[5\.4](https://arxiv.org/html/2605.06730#S5.SS4)should be read as characterising the cost of properties \(a\)–\(c\) jointly\.
Baseline scope\.We isolate SSAI design choices with deterministic portfolios and matched supervised learners; lexical headline aggregates \(VADER; train\-fit TF–IDF/SVD\) now appear as dense ridge inputs \(Table[4](https://arxiv.org/html/2605.06730#S5.T4)\)\. Table[5](https://arxiv.org/html/2605.06730#S5.T5)adds a FinBERT direct\-portfolio baseline \(FinBERT\-SFP, 386\.3% CR\) and a PC1 baseline \(PC1\-SFP, 433\.6% CR\) under the same top\-10 rule\. What this submission does*not*cover: FinBERT or dense embeddings as RL state vectors, multi\-seed factor portfolio evaluation, or comparison against FinGPT/BloombergGPT full encoders—these remain open gaps \(see Limitations, Section[6](https://arxiv.org/html/2605.06730#S6)\)\.
## 3Method
Schematic of the news\-to\-policy pipeline is Figure[2](https://arxiv.org/html/2605.06730#A2.F2)\(Appendix[B](https://arxiv.org/html/2605.06730#A2)\)\.
### 3\.1Multi\-Signal LLM Scoring
Given a news article headline or summaryttabout stock tickerss, we issue a single structured prompt to an LLM and extract four integer scoresσ=\(σsent,σrisk,σconf,σvol\)∈\{1,…,5\}4\\sigma=\(\\sigma\_\{\\text\{sent\}\},\\sigma\_\{\\text\{risk\}\},\\sigma\_\{\\text\{conf\}\},\\sigma\_\{\\text\{vol\}\}\)\\in\\\{1,\\ldots,5\\\}^\{4\}\.
σsent\\sigma\_\{\\text\{sent\}\}Sentiment: 1 = very negative, 5 = very positive\.
σrisk\\sigma\_\{\\text\{risk\}\}Risk: 1 = very low company/market risk, 5 = very high risk\.
σconf\\sigma\_\{\\text\{conf\}\}Confidence: 1 = very uncertain outlook, 5 = very certain\.
σvol\\sigma\_\{\\text\{vol\}\}Volatility forecast: 1 = very calm expected price, 5 = highly volatile\.
We use explicit axes rather than dense article embeddings because the study is designed to audit a state interface, not merely maximize predictive capacity\. The integer coordinates support ticker\-day coverage checks, residualization against sentiment, leave\-one\-axis masking, and deployment\-facing interpretation of why a policy or ranking rule changes exposure\. Dense latent text encoders and price–news Transformers are natural stronger baselines, but they do not provide the same axis\-level causal and diagnostic handles without additional attribution machinery\.
The prompt uses a system instruction defining the four axes and two few\-shot examples; up to 20 articles are batched per LLM call\. We apply this protocol to 40,850 articles from the FNSPID NASDAQ subset, scoring 39,995 \(97\.9%\) successfully\.111LLM identity \(reproducibility\)\.Scoring was performed usinggpt\-4o\(OpenAI chat\-completions API, model snapshotgpt\-4o\-2024\-08\-06\) between August and October 2024 with temperature 0, top\-p 1\. The full prompt template, including per\-axis rubrics and two canonical few\-shot exemplars, is included verbatim in the artifact package \(prompts/score\_news\.txt\)\. Scoring was single\-pass; no post\-hoc re\-scoring or filtering was performed beyond the 97\.9% parse success rate reported above\.Score summaries appear in Table[14](https://arxiv.org/html/2605.06730#A6.T14)\(Appendix[F](https://arxiv.org/html/2605.06730#A6)\); risk and volatility are right\-skewed \(mean≈2\.5\\approx 2\.5\), consistent with adverse\-news skew in financial headlines\.
Signal coverage is uneven after aggregation to the trading panel: This sparsity is a central reason we treat the interface as a weak semantic factor rather than a dense predictive signal\. Importantly, the direct portfolio experiments already control for this: SFP uses signal deviations only on days with non\-neutral coverage, and outperforms equal\-weight buy\-and\-hold by 63\.6pp over 2019–2023—a gap that cannot be attributed to coverage volume alone since mean full\-scoring evaluation exceeds neutral\-masked evaluation of the same checkpoints across 21 seeds \(Table[8](https://arxiv.org/html/2605.06730#S5.T8)\)\.
Signal coverage, per\-axis IC, and a four\-panel internal validation \(score distributions, inter\-signal Spearman correlations\|ρ\|=0\.58\|\\rho\|\{=\}0\.58–0\.810\.81, lag\-1 autocorrelation0\.470\.47–0\.570\.57, and confidence ICp=0\.004p\{=\}0\.004\) are reported in Appendix[F](https://arxiv.org/html/2605.06730#A6)\(Figure[4](https://arxiv.org/html/2605.06730#A6.F4), Tables[15](https://arxiv.org/html/2605.06730#A6.T15),[16](https://arxiv.org/html/2605.06730#A6.T16)\)\. The key takeaway: scores are not white noise—they exhibit the directional cross\-axis structure expected from financial news and are temporally persistent at the ticker level\.*Effective dimensionality caveat:*PCA on the four axes over non\-neutral stock\-days \(N=5,579N\{=\}5\{,\}579\) finds PC1 explains82\.1%of variance \(PC1\+PC2: 91\.8%\), so the effective dimensionality of SSAI is approximately 1—the four axes primarily capture a single sentiment–risk polarity\. SRF’s residualisation addresses linear redundancy within this low\-dimensional space\.*Economic magnitude caveat:*the confidence ICp=0\.004p\{=\}0\.004is statistically significant, but atN=83,040N\{=\}83\{,\}040stock\-days this corresponds to a Spearman IC of approximately0\.0220\.022—economically small \(Information Ratio≈IC×252≈0\.35\\approx IC\\times\\sqrt\{252\}\\approx 0\.35; expected annualised alpha≈6\\approx 6–1010bp for a single axis under standard factor arithmetic\)\. The IC evidence motivates the interface design but should not be read as implying practically large alpha generation from the confidence axis alone\.
For each trading dayddand tickerss, we aggregate all articles published in the window\[d−3,d\]\[d\-3,d\]by computing the mean of each signal dimension\. Days with no matching articles retain a neutral value of 3\.0\.
#### Evaluation invariant\.
Becauseϕ\\phiis frozen and shared across all estimators, differences in outcome between estimators are attributable to algorithm choice, not representation \(see Appendix[A\.1](https://arxiv.org/html/2605.06730#A1.SS1)for formal statement\)\.
### 3\.2Trading Environment
We build on the FinRLStockTradingEnvgymnasium environment\. The state vector on dayddis:
𝐬d=\[cd,𝐩d,𝐡d,𝐟d,𝝈d\]∈ℝ1\+4N\+KN\\mathbf\{s\}\_\{d\}=\\bigl\[c\_\{d\},\\;\\mathbf\{p\}\_\{d\},\\;\\mathbf\{h\}\_\{d\},\\;\\mathbf\{f\}\_\{d\},\\;\\boldsymbol\{\\sigma\}\_\{d\}\\bigr\]\\in\\mathbb\{R\}^\{1\+4N\+KN\}\(1\)
wherecdc\_\{d\}is available cash,𝐩d∈ℝN\\mathbf\{p\}\_\{d\}\\in\\mathbb\{R\}^\{N\}are closing prices,𝐡d∈ℝN\\mathbf\{h\}\_\{d\}\\in\\mathbb\{R\}^\{N\}are current share holdings,𝐟d∈ℝKN\\mathbf\{f\}\_\{d\}\\in\\mathbb\{R\}^\{KN\}areK=7K=7technical indicators \(MACD, Bollinger upper/lower, RSI30, CCI30, ADX30, 30\-day SMA, 60\-day SMA\) for each ofN=30N=30stocks, and𝝈d∈ℝ4N\\boldsymbol\{\\sigma\}\_\{d\}\\in\\mathbb\{R\}^\{4N\}are the four LLM signals\. The total dimension is1\+2×30\+\(7\+4\)×30=4211\+2\\times 30\+\(7\+4\)\\times 30=421\.
Actionsad∈\[−1,1\]Na\_\{d\}\\in\[\-1,1\]^\{N\}are scaled by a maximum trade sizehmax=100h\_\{\\text\{max\}\}=100shares\. Transaction costs of 0\.1% per trade are applied\. A turbulence index\(Kritzman and Li,[2010](https://arxiv.org/html/2605.06730#bib.bib6)\)triggers a no\-buy constraint when it exceeds a threshold of 380, implementing a simple market\-regime filter\.
### 3\.3Drawdown\-Penalised PPO Reward
The per\-step reward applies*drawdown shaping*:
rt=ΔWtW0⋅λ−α⋅max\(0,Wpeak−WtWpeak\)2r\_\{t\}=\\frac\{\\Delta W\_\{t\}\}\{W\_\{0\}\}\\cdot\\lambda\-\\alpha\\cdot\\max\\\!\\left\(0,\\;\\frac\{W\_\{\\text\{peak\}\}\-W\_\{t\}\}\{W\_\{\\text\{peak\}\}\}\\right\)^\{2\}\(2\)
whereΔWt=Wt−Wt−1\\Delta W\_\{t\}=W\_\{t\}\-W\_\{t\-1\}is the change in portfolio value,λ=10−4\\lambda=10^\{\-4\}is the reward scale,WpeakW\_\{\\text\{peak\}\}is the running maximum portfolio value, andα=0\.1\\alpha=0\.1is the drawdown penalty coefficient\. The quadratic form penalises large drawdowns super\-linearly, encouraging the agent to protect gains rather than risk them for marginal upside \(Appendix[B\.6](https://arxiv.org/html/2605.06730#A2.SS6)interprets the shaping term\)\.
We train with a multi\-process PPO implementation \(OpenAI SpinningUp\), using 4 MPI workers, hidden layers\(512,512\)\(512,512\)with Tanh activations, learning rate3×10−43\\times 10^\{\-4\}, clip ratio 0\.2, and 30 epochs over the 2013–2018 training set \(≈\\approx45,300 daily observations across 30 stocks\)\.Naming\.Released checkpoints use theCPPOfilename prefix for historical compatibility; all tables label this agentDP\-PPO\(drawdown\-shaped reward, Sec\.[3\.3](https://arxiv.org/html/2605.06730#S3.SS3)\), not Lagrangian constrained optimisation\.
### 3\.4LLM Signal Integration
All four LLM channels enter𝝈d\\boldsymbol\{\\sigma\}\_\{d\}as described above\. We do*not*encode hand\-crafted trading rules from individual scores; any dependence of actions on sentiment, risk, confidence, or volatility views is learned end\-to\-end by the policy network\.
#### Semantic Factor Portfolio \(SFP\)\.
SFP fits a ridge model \(train 2013–2018,λ=10−3\\lambda\{=\}10^\{\-3\}\) from signal deviations\(σ−3\)\(\\sigma\{\-\}3\)to 5\-day returns; weights arefrozenfor the entire 2019–2023 test window\. Daily top\-10 long\-only portfolio; 0\.1% transaction cost\.*SRF*fits on residualsϵj=σj−aj−bjσsent\\epsilon\_\{j\}=\\sigma\_\{j\}\-a\_\{j\}\-b\_\{j\}\\sigma\_\{\\text\{sent\}\}, testing whether non\-sentiment axes add structure beyond recoded sentiment\.*SCW*applies a validation\-year softmax over factor scores for conviction\-weighted allocation\.
## 4Experimental Setup
#### Data\.
We use Yahoo Finance OHLCV for 30 liquid NASDAQ\-100 constituents \(tickers listed in Appendix[B\.4](https://arxiv.org/html/2605.06730#A2.SS4.SSS0.Px1)\)\. Training: 2013\-01\-02 to 2018\-12\-31 \(≈\\approx45,300 observations\)\. Out\-of\-sample testing: 2019\-01\-02 to 2023\-12\-29 \(1,258 trading days\)\.
News articles are sourced from the FNSPID dataset\(Dong et al\.,[2024](https://arxiv.org/html/2605.06730#bib.bib4)\)via Hugging Face \(Zihan1004/FNSPID\), filtered to the same 30 tickers\. 40,850 articles are collected; 39,995 \(97\.9%\) are successfully scored by the LLM via a hosted chat\-completions API\.
#### Baselines\.
SFP/SRF/SCW: ridge\-learned factor portfolios \(four axes / sentiment\-only / residual axes / conviction\-weighted\);Supervised ridge: 5\-day return forecasters with price, LLM semantics, VADER, TF–IDF/SVD, or high\-conviction tilt;PC1\-SFP / FinBERT\-SFP: principal\-component and dense\-encoder portfolio baselines \(Table[5](https://arxiv.org/html/2605.06730#S5.T5)\);SAC: off\-policy agent with same SSAI observations;DP\-PPO neutral: LLM signals fixed at 3\.0;Equal\-weight B&H, Momentum, Equal\-Vol: passive and rule\-based price benchmarks\.
#### Metrics\.
Cumulative Return \(CR\), Annual Return \(AR\), Sharpe, Sortino, Maximum Drawdown \(MDD\), Calmar \(=AR/\|MDD\|=AR/\|MDD\|\)\. Transformer and constrained\-risk baselines remain future work \(see Section[2](https://arxiv.org/html/2605.06730#S2)\)\.
## 5Results
### 5\.1Main Performance Comparison
Table[1](https://arxiv.org/html/2605.06730#S5.T1)reports all metrics on the 2019–2023 test period \(numeric cells matcheval\_harness\.pyexports\)\.
Table 1:Main Strategy Comparison \(2019–2023\)Figure 1:Normalized portfolio value on the 2019–2023 test window \(DP\-PPO with LLM signals versus baselines\), produced byeval\_harness\.py\.#### Key findings\.
Factor portfolios\(Table[5](https://arxiv.org/html/2605.06730#S5.T5)\): PC1\-SFP433\.6%/1\.256Sharpe; FinBERT\-SFP386\.3%/1\.206; 4\-axis SSAI\-SFP307\.2%/1\.067; B&H243\.6%/1\.032\. The 63pp SSAI–B&H gap fails within\-stratum controls \(all three coverage terciles; Table[3](https://arxiv.org/html/2605.06730#S5.T3)\), reverses at≥0\.2%\{\\geq\}0\.2\\%costs, and the 4\-axis vs\. sentiment\-only advantage is non\-significant \(Wilcoxonp=0\.556p\{=\}0\.556; Table[11](https://arxiv.org/html/2605.06730#A5.T11)\)\.Supervised\(Table[4](https://arxiv.org/html/2605.06730#S5.T4)\): naive LLM appending hurts ridge; VADER/TF–IDF and high\-conviction tilt recover Sharpe∼\\sim0\.674; FinBERT\-ridge 125\.0%/0\.649≈\\approxprice\-only\.RL: SAC \(7 seeds\) mean Sharpe1\.059±0\.1401\.059\\pm 0\.140vs\. DP\-PPO \(21 seeds\)0\.920±0\.0990\.920\\pm 0\.099, Mann\-Whitneyp=0\.027p\{=\}\\mathbf\{0\.027\}; CR not significant \(p=0\.604p\{=\}0\.604\); DP\-PPO 21\-seed Sharpe0\.920±0\.0990\.920\\pm 0\.099\(full\) vs\.0\.907±0\.0940\.907\\pm 0\.094\(masked\), Wilcoxonp≈0\.25p\{\\approx\}0\.25–0\.310\.31; removing*confidence*drops CR to 185\.1%; removing*volatility\_forecast*enlarges MDD to−\-43\.8%\.
### 5\.2Direct Semantic Factor Portfolio
Table[2](https://arxiv.org/html/2605.06730#S5.T2)learns semantic factor weights on 2013–2018 and applies a daily long\-only top\-10 rule in 2019–2023\. Four\-factor SFP and SCW improve cumulative return and Sharpe versus sentiment\-only and buy\-and\-hold; SRF preserves gains after residualizing non\-sentiment axes on the training panel \(linear diagnostic control\)\.
SFP/SRF/SCW share identical Sharpe because all three use the same top\-10 rule; basket membership is highly stable across weight configurations \(see Appendix[E\.7](https://arxiv.org/html/2605.06730#A5.SS7)\)\.
Table 2:Semantic Factor Portfolio variants\. Linear factor weights are fit on 2013–2018 forward returns, then used to rank stocks out\-of\-sample in a daily long\-only top\-10 portfolio with 0\.1% transaction costs\. SRF residualizes non\-sentiment axes against sentiment\. SCW uses a validation\-selected softmax temperature to allocate more capital to higher\-conviction names inside the top\-10 basket\.#### Coverage\-stratified analysis: key qualification \(Table[3](https://arxiv.org/html/2605.06730#S5.T3)\)\.
SFP underperforms within\-stratum B&H in all three coverage terciles\(Low: 248\.6% vs\. 300\.0%; Mid: 361\.9% vs\. 379\.6%; High: 232\.6% vs\. 352\.1%\)\. The aggregate 63pp gap is a cross\-universe*composition*effect: SFP’s top\-10 rule selects mid\-coverage names that delivered strong returns, not names with within\-stratum signal advantage\. Within the top\-10 highest\-coverage tickers alone \(NVDA, GOOGL, AVGO, …\), equal\-weight B&H returns 352\.1% while SFP \(k=5\) returns only 232\.6%—a 119pp*deficit*\. SFP’s outperformance over B&H further requires≤\\leq0\.1% per\-trade costs; at 0\.2% the excess reverses \(Table[12](https://arxiv.org/html/2605.06730#A5.T12), Appendix[E\.4](https://arxiv.org/html/2605.06730#A5.SS4)\)\.
Counterintuitively, the highest\-coverage tercile shows the*worst*SSAI relative performance \(232\.6% vs\. 352\.1% B&H\): high\-coverage names are the most efficiently priced, AI\-momentum drove those tickers precisely on high\-risk/negative\-sentiment news days, and 2013–2018 factor weights are structurally misspecified for the 2022–2023 AI\-cycle regime \(see Appendix[E\.6](https://arxiv.org/html/2605.06730#A5.SS6)for full analysis\)\.
Table 3:Coverage\-stratified SFP vs\. equal\-weight B&H \(2019–2023\)\. Tickers split into thirds by non\-neutral signal coverage fraction \(Low<<5%: MSFT/META/ISRG/COST/INTC/AMZN/AAPL/AMD/PANW/SNPS; Mid 5–15%: INTU/KLAC/ORCL/TXN/REGN/MCHP/LRCX/CDNS/TSLA/ADBE; High\>\>15%: MRVL/AMAT/NFLX/ADI/QCOM/ASML/MU/NVDA/AVGO/GOOGL\)\. Within each tercile, SFP \(k=5k=5\) is evaluated using the full\-panel fitted weights restricted to that universe\.SFP underperforms within\-tercile B&H in all three groups, indicating that the full\-portfolio outperformance \(307% vs\. 244% B&H over all 30 stocks\) is substantially a cross\-universe selection effect—SFP picks mid\-coverage names that happened to outperform the full equal\-weight basket—rather than purely a within\-coverage\-tier signal quality advantage\.
### 5\.3Supervised Forecasting Stress Test
Table[4](https://arxiv.org/html/2605.06730#S5.T4)fits ridge models predicting 5\-day returns from price/technical features with sparse LLM semantic columns, lexical headline dense features \(VADER; TF–IDF/SVD\), or both combined via high\-conviction semantic tilt\. Price\-only:126\.4%CR/Sharpe0\.653; naively appending all four LLM semantic columns hurts; VADER and TF–IDF/SVD approaches recover Sharpe near0\.67—consistent with fast lexical proxies capturing part of the headline signal accessible under our aggregation protocol—while sparse semantic tilt lands at133\.9%/0\.674\. FinBERT \(CR 125\.0%, Sharpe 0\.649\) performs comparably to price\-only, confirming that dense neural text features—like naive multi\-axis semantics—do not improve ridge forecasting without careful integration\.
Table 4:Supervised forecasting baselines\. Ridge models predict 5\-day forward returns from price/technical features alone, price plus LLM sentiment or four\-factor semantics, price plus lexical dense headlines \(VADER; train\-fit TF–IDF/SVD\), optional FinBERT tone when available, or a semantic tilt\. Ridge strength and the semantic tilt are selected on 2018 validation Sharpe, refit on 2013–2018, and evaluated once on 2019–2023 using the same daily top\-10 portfolio rule as SFP\.
### 5\.4Factor Comparison: PC1, Softmax, FinBERT, and SSAI
Table[5](https://arxiv.org/html/2605.06730#S5.T5)adds three baselines under the same top\-10 rule: PC1\-SFP \(first principal component of the four standardised axes; 81\.9% explained variance\), Softmax\-SFP \(equal\-weight mean of the same four standardised axes\), and FinBERT\-SFP \(ProsusAI/finbert sentiment\)\.
PC1 loadings are nearly equal\-weighted\.The PC1 eigenvector has absolute loadings of 0\.514/0\.520/0\.480/0\.485 across the four axes, normalising to 0\.257/0\.260/0\.240/0\.242 — within 2pp of uniform\. This motivated the Softmax\-SFP test: Softmax\-SFP achieves439\.8%CR \(Sharpe 1\.248\) vs\. PC1\-SFP 433\.6% \(Sharpe 1\.256\), a 6pp difference within noise\.*The axis decomposition itself costs nothing relative to the PC1 compression\.*
The 126pp gap is a portfolio\-rule effect, not an interpretability cost\.Since Softmax\-SFP≈\\approxPC1\-SFP, the full gap to 4\-axis SSAI\-SFP \(307\.2%\) traces to the ridge\-trained factor weights used in SSAI\-SFP, not to the named\-axis interface\. Practitioners who rank by equal\-weighted standardised SSAI scores retain full auditability \(axis\-level inspection, override, and prompt perturbation\) while recovering PC1\-level returns\. Named axis weights remain valuable for regulatory legibility and causal attribution; the four\-axis interface can be used with any weighting scheme\.
FinBERT\-SFP \(386\.3%\) outperforms SSAI\-SFP here yet underperformed in ridge forecasting \(Table[4](https://arxiv.org/html/2605.06730#S5.T4)\), confirming dense encodings are competitive direct ranking signals but not naive ridge features\.
Table 5:Factor portfolio comparison \(2019–2023, 30 tickers, daily top\-10 rebalancing\)\.PC1\-SFP: first principal component of four standardised semantic axes \(81\.9% explained variance; loadings≈\\approxequal at 0\.257/0\.260/0\.240/0\.242\)\.Softmax\-SFP: equal\-weight mean of the same four standardised axes \(fully auditable; no post\-hoc compression\)\.FinBERT\-SFP: ProsusAI/finbert sentiment as ranking signal\.4\-axis SSAI \(SFP\): ridge\-trained factor weights from Table[2](https://arxiv.org/html/2605.06730#S5.T2)\.*Key finding*: PC1\-SFP≈\\approxSoftmax\-SFP \(6pp gap\), confirming the axis decomposition itself is near\-free; the full 126pp gap PC1 vs\. SSAI\-SFP traces to the portfolio weighting rule, not the named\-axis interface\.
### 5\.5Algorithm Baseline: SAC
Table[6](https://arxiv.org/html/2605.06730#S5.T6)compares SAC \(7 seeds\) against DP\-PPO \(21 seeds\), both using the identical LLM\-enriched observation vector\.SAC achieves mean Sharpe1\.059±0\.1401\.059\\pm 0\.140vs\. DP\-PPO0\.920±0\.0990\.920\\pm 0\.099; Mann\-WhitneyUUtest givesp=0\.027p\{=\}0\.027\(two\-sided\), a statistically significant advantage at the 5% level\. CR means are not significantly different \(202\.3%±38\.9%202\.3\\%\\pm 38\.9\\%vs\.195\.0%±45\.3%195\.0\\%\\pm 45\.3\\%;p=0\.604p\{=\}0\.604\)\.
Since both agents share the same frozenϕ\\phi, the Sharpe gap is attributable to algorithm differences \(off\-policy replay, entropy regularisation\) rather than representation differences—directly answering RQ3 and confirming the SSAI evaluation invariant \(Appendix[A\.1](https://arxiv.org/html/2605.06730#A1.SS1)\)\. The observation that CR is not significantly different while Sharpe is confirms that SAC’s advantage is primarily in downside management \(reduced variance\), not in raw return\. RL results and factor portfolios occupy different performance regimes by design: RL agents operate under stochastic policy noise evaluated as diagnostics; factor portfolios use deterministic top\-10 rules\. Whether SAC with a FinBERT state would narrow the factor\-portfolio gap is open\.
Table 6:Algorithm baseline with identical four\-signal observations \(2019–2023\)\.SAC \(7 seeds\) and DP\-PPO \(21 seeds\) share the same LLM\-enriched state vector\. Mean±\\pmstd reported; Mann\-WhitneyUUtest \(two\-sided\) compares the two seed distributions\. SAC achieves significantly higher Sharpe \(p=0\.027p\{=\}0\.027\); CR gap is not significant \(p=0\.604p\{=\}0\.604\), consistent with the algorithm\-vs\-representation attribution: identical observations, different objectives\. See Section[5\.5](https://arxiv.org/html/2605.06730#S5.SS5)\.
### 5\.6Ablation: Signal Contribution
Table[7](https://arxiv.org/html/2605.06730#S5.T7)reports leave\-one\-signal\-out masking at*evaluation*time on fixed trained weights; it is not a replacement for retraining\-based causal attribution\. The table shows DP\-PPO neutral \(196\.8%/0\.926\) marginally exceeding DP\-PPO full signals \(190\.4%/0\.917\) in CR on the reported checkpoint\. This reversal is not paradoxical: \(i\) the difference is within the 21\-seed standard deviation \(±\\pm17pp CR\), confirmed by Wilcoxonp≈0\.25p\{\\approx\}0\.25–0\.310\.31; \(ii\) evaluation\-time masking changes the policy’s state distribution without retraining, so the value function’s calibration on full\-signal inputs is disrupted; \(iii\) the multi\-seed means preserve the expected direction \(full Sharpe0\.920±0\.096\>0\.920\\pm 0\.096\>masked0\.907±0\.0940\.907\\pm 0\.094, Table[8](https://arxiv.org/html/2605.06730#S5.T8)\)\. The single\-checkpoint CR reversal is a noise realisation, not evidence that SSAI signals harm the policy\.
Table 7:Signal Ablation Study — DP\-PPO Variants
### 5\.7Multi\-seed robustness
Table[8](https://arxiv.org/html/2605.06730#S5.T8)summarises mean±\\pmstd cumulative return and Sharpe across 21 seeds;*neutral eval*masks LLM coordinates at test time\. Bootstrap CIs and paired Wilcoxonpp\-values are reported in the table caption and Section[5](https://arxiv.org/html/2605.06730#S5)\.
Table 8:Multi\-seed robustness \(21 seeds, 2019–2023\)\. Mean±\\pmstd across seeds for cumulative return and Sharpe\.Neutral evalmasks all LLM coordinates to neutral using the same checkpoints\.Sub\-period and transaction\-cost sensitivity analyses are reported in Appendix[E](https://arxiv.org/html/2605.06730#A5)\(Tables[9](https://arxiv.org/html/2605.06730#A5.T9),[10](https://arxiv.org/html/2605.06730#A5.T10); Figure[3](https://arxiv.org/html/2605.06730#A5.F3)\)\. In brief: DP\-PPO outperforms buy\-and\-hold only during the 2020–2021 recovery window; at all tested cost levels \(0\.05%–2\.0%\) DP\-PPO remains below the passive sleeve\.
## 6Discussion
#### Representation vs\. algorithm\.
Sparse LLM semantics are useful as direct portfolio tilts and high\-conviction supervised overlays, but not as naive ridge features\. Inside RL, the evaluation invariant \(Appendix[A\.1](https://arxiv.org/html/2605.06730#A1.SS1)\) attributes the SAC/DP\-PPO gap to algorithm differences, not representation\. Evaluation\-time masking is not causal; RL as diagnostic, not product; signal validation in Appendix[F](https://arxiv.org/html/2605.06730#A6); Definition 1 is domain\-agnostic \(Appendix[D](https://arxiv.org/html/2605.06730#A4)\)\.
#### Limitations\.
\(1\)Composition confound: the 63pp SFP gap fails within\-stratum controls and reverses at≥0\.2%\{\\geq\}0\.2\\%costs\. \(2\)Single\-market: NASDAQ\-100, 2019–2023 only; sector\-neutral and rolling OOS designs are open\. \(3\)SAC budget: 7 seeds vs\. 21 DP\-PPO seeds; Sharpe gap is significant \(p=0\.027p\{=\}0\.027\) but CR is not \(p=0\.604p\{=\}0\.604\); hyperparameter parity not guaranteed\. \(4\)Open encoder gaps: FinBERT/FinGPT as RL state vectors remain untested\. This work is a retrospective benchmark; see Appendix[G](https://arxiv.org/html/2605.06730#A7)for ethics and Appendix[B](https://arxiv.org/html/2605.06730#A2)for reproducibility details\.
## 7Conclusion
We introduced SSAI as a controlled evaluation framework for disentangling text representation from optimisation effects in LLM\-augmented decision systems, and applied it to a US equity portfolio task\. The main findings are cautionary: the apparent 63pp SFP advantage over buy\-and\-hold is a basket\-selection composition artefact \(fails within\-stratum controls, reverses at≥0\.2%\{\\geq\}0\.2\\%costs\); PC1 loadings are near\-equal\-weighted \(0\.257/0\.260/0\.240/0\.242\), and an equal\-weight Softmax\-SFP achieves 439\.8% CR—matching PC1’s 433\.6%—showing the named\-axis interface costs nothing over PC1; the 126pp gap to SSAI\-SFP \(307\.2%\) traces to the ridge weighting rule, not the axes themselves; and RL performance depends critically on algorithm choice, not representation\. The SSAI template, code artifact, and diagnostic protocol are offered as reusable infrastructure for future sparse\-text decision\-making research\.
## References
- Bengio et al\. \[2013\]Yoshua Bengio, Aaron Courville, and Pascal Vincent\.Representation learning: A review and new perspectives\.*IEEE Transactions on Pattern Analysis and Machine Intelligence*, 35\(8\):1798–1828, 2013\.
- Benhenda \[2025\]Mostapha Benhenda\.FinRL\-DeepSeek: LLM\-infused risk\-sensitive reinforcement learning for trading agents\.*arXiv preprint arXiv:2502\.07393*, 2025\.
- Brown et al\. \[2020\]Tom Brown, Benjamin Mann, Nick Ryder, et al\.Language models are few\-shot learners\.*Advances in Neural Information Processing Systems*, 33:1877–1901, 2020\.
- Dong et al\. \[2024\]Zihan Dong, Xinyu Fan, and Zhiyuan Peng\.FNSPID: A comprehensive financial news dataset in time series\.*arXiv preprint arXiv:2402\.06698*, 2024\.Dataset mirror:[https://huggingface\.co/datasets/Zihan1004/FNSPID](https://huggingface.co/datasets/Zihan1004/FNSPID)\.
- Higgins et al\. \[2017\]Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner\.beta\-VAE: Learning basic visual concepts with a constrained variational framework\.In*International Conference on Learning Representations*, 2017\.
- Kritzman and Li \[2010\]Mark Kritzman and Yuanzhen Li\.Skulls, financial turbulence, and risk management\.*Financial Analysts Journal*, 66\(5\):30–41, 2010\.
- Liu et al\. \[2021\]Xiao\-Yang Liu, Hongyang Yang, Qian Chen, Runjia Zhang, Liuqing Yang, Bowen Xiao, and Christina Dan Wang\.FinRL: A deep reinforcement learning library for automated stock trading in quantitative finance\.*arXiv preprint arXiv:2011\.09607*, 2021\.
- Liu et al\. \[2022\]Xiao\-Yang Liu, Jingyang Rui, Jiechao Gao, Liuqing Yang, Hongyang Yang, Zhaoran Wang, Christina Dan Wang, and Jian Guo\.FinRL\-Meta: Market environments and benchmarks for data\-driven financial reinforcement learning\.*arXiv preprint arXiv:2211\.03107*, 2022\.
- Liu et al\. \[2023\]Xiao\-Yang Liu, Guoxuan Wang, and Daochen Zha\.FinGPT: Open\-source financial large language models\.*arXiv preprint arXiv:2306\.06031*, 2023\.
- Lopez\-Lira and Tang \[2023\]Alejandro Lopez\-Lira and Yuehua Tang\.Can ChatGPT forecast stock price movements? Return predictability and large language models\.*arXiv preprint arXiv:2304\.07619*, 2023\.
- Lundberg and Lee \[2017\]Scott M Lundberg and Su\-In Lee\.A unified approach to interpreting model predictions\.In*Advances in Neural Information Processing Systems*, volume 30, 2017\.
- Mnih et al\. \[2015\]Volodymyr Mnih, Koray Kavukcuoglu, David Silver, et al\.Human\-level control through deep reinforcement learning\.*Nature*, 518\(7540\):529–533, 2015\.
- Ray et al\. \[2019\]Alex Ray, Joshua Achiam, and Dario Amodei\.Benchmarking safe exploration in deep reinforcement learning\.*arXiv preprint arXiv:1910\.01708*, 2019\.
- Ribeiro et al\. \[2016\]Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin\."why should I trust you?": Explaining the predictions of any classifier\.In*Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining*, pages 1135–1144, 2016\.
- Schulman et al\. \[2017\]John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov\.Proximal policy optimization algorithms\.*arXiv preprint arXiv:1707\.06347*, 2017\.
- Tamar et al\. \[2015\]Aviv Tamar, Yonatan Glassner, and Shie Mannor\.Optimizing the CVaR via stochastic gradient\.*Proceedings of the AAAI Conference on Artificial Intelligence*, 29, 2015\.
- Wu et al\. \[2023\]Shijie Wu, Ozan Irsoy, Steven Lu, et al\.BloombergGPT: A large language model for finance\.*arXiv preprint arXiv:2303\.17564*, 2023\.
- Yang et al\. \[2020a\]Hongyang Yang, Xiao\-Yang Liu, Shan Zhong, and Anwar Walid\.FinRL: A deep reinforcement learning library for automated stock trading in quantitative finance\.*NeurIPS Workshop on Deep RL*, 2020a\.
- Yang et al\. \[2020b\]Hongyang Yang, Xiao\-Yang Liu, Shan Zhong, and Anwar Walid\.Deep reinforcement learning for automated stock trading: An ensemble strategy\.*Proceedings of the First ACM International Conference on AI in Finance*, 2020b\.
## Appendix AFormal Framework Definition
### A\.1SSAI Evaluation Invariant
> Invariant\.Letϕ\\phibe a fixed SSAI map and\{f1,…,fM\}\\\{f\_\{1\},\\ldots,f\_\{M\}\\\}a set of estimators sharingϕ\\phias their sole text interface\. Then differences in outcomeΔ\(fi,fj\)\\Delta\(f\_\{i\},f\_\{j\}\)are attributable to algorithm differences, not representation differences, sinceϕ\\phiis common\. This holds when \(i\)ϕ\\phiis frozen before any estimator trains, \(ii\) no estimator fine\-tunesϕ\\phi, and \(iii\) the same observation vector𝐬d\\mathbf\{s\}\_\{d\}is used across estimators\. All three conditions are satisfied in this study\.
Definition 1 \(Semantic State Abstraction Interface, SSAI\)\.Let𝒩s,d\\mathcal\{N\}\_\{s,d\}denote the set of text documents for assetsson daydd\. A SSAI is a functionϕ:𝒩s,d→ℝK\\phi:\\mathcal\{N\}\_\{s,d\}\\rightarrow\\mathbb\{R\}^\{K\}mapping raw text toKKinterpretable scalar axes, where each axis is \(i\) defined in natural language, \(ii\) elicited by a fixed prompt from a frozen LLM, and \(iii\) replaced by a neutral defaultϕ0∈ℝK\\phi\_\{0\}\\in\\mathbb\{R\}^\{K\}when𝒩s,d=∅\\mathcal\{N\}\_\{s,d\}=\\emptyset\. HereK=4K\{=\}4\(sentiment, risk, confidence, volatility forecast\),ϕ0=\[3,3,3,3\]\\phi\_\{0\}=\[3,3,3,3\], and the prompt is fixed for all tickers and dates\. The SSAI is*auditable*\(each axis is human\-interpretable\),*sparse\-by\-design*\(neutral default on no\-news days\), and*optimizer\-agnostic*\(the sameϕ\\phifeeds both direct factor rules and RL state vectors\)\.
## Appendix BImplementation Details
### B\.1Pipeline schematic
Financial news→\\rightarrowLLM prompt→\\rightarrow4 semantic factors sentiment / risk / confidence / volatility forecast↓\\downarrowaggregate by ticker\-day \(33\-day window\)FinRL state: cash, prices, holdings, indicators, semantic factors→\\rightarrowRL policy: DP\-PPO or SAC→\\rightarrowportfolio trades
Figure 2:FinRL\-MultiSignal pipeline\. The LLM supplies structured semantic factors; trading decisions are learned by the RL policy from market state plus those factors\.
### B\.2LLM Scoring Prompt
The system prompt used for multi\-signal scoring is:
> You are a quantitative financial analyst\. For each news snippet about a stock you will output exactly four integer scores on a 1\-5 scale, separated by a pipe ‘\|’, in this exact order: sentiment \| risk \| confidence \| volatility\_forecast\.\(Per\-axis rubrics and canonical few\-shot exemplars are lengthy; the verbatim templates are included with the scoring scripts in the artifact package\.\)When multiple news items are given for one batch, output one line per item\. Never add explanation \-\-\- only scores\.
Two few\-shot examples are prepended to each batch request\. Up to 20 articles are batched per API call to amortise latency\.
### B\.3Network Architecture
### B\.4Data Pipeline
Stock price data: Yahoo Finance viayfinance\(adjusted closes, 2013\-01\-02–2023\-12\-29, 30 NASDAQ\-100 tickers\)\. Technical indicators computed viastockstats\. Turbulence index computed followingKritzman and Li \[[2010](https://arxiv.org/html/2605.06730#bib.bib6)\]\. News data: Hugging Face datasetZihan1004/FNSPID, filtered to the 30 target tickers, 40,850 articles collected \(2009–2023\)\.
#### Ticker universe\.
AAPL, ADBE, ADI, AMAT, AMD, AMZN, ASML, AVGO, CDNS, COST, GOOGL, INTC, INTU, KLAC, LRCX, MCHP, META, MRVL, MSFT, MU, NFLX, NVDA, ORCL, PANW, QCOM, REGN, SNPS, TSLA, TXN, ISRG\.
### B\.5Compute
All experiments were run on a single workstation \(Apple Silicon, 16 GB unified memory\)\. LLM scoring:≈\\approx17\.5 hours for 40,850 articles at batch size 20 via a hosted chat\-completions endpoint\. DP\-PPO training:≈\\approx45 minutes per 30\-epoch run \(4 MPI workers, CPU\-only PyTorch\)\.
### B\.6Drawdown shaping interpretation
LetDt=max\(0,\(Wpeak−Wt\)/Wpeak\)D\_\{t\}=\\max\(0,\(W\_\{\\text\{peak\}\}\-W\_\{t\}\)/W\_\{\\text\{peak\}\}\)denote fractional drawdown\. The penalty−αDt2\-\\alpha D\_\{t\}^\{2\}is a smooth downside\-risk surrogate: for small losses it is mild, but its gradient magnitude grows linearly in drawdown, so deeper excursions create increasingly strong pressure to reduce exposure\. This differs from variance regularisation \(symmetric upside/downside penalties\) and from CVaR\-style constraints \(explicit tail optimisation\)\. The intent is pragmatic shaping compatible with standard policy\-gradient code\.
## Appendix CSAC mechanism hypotheses
The comparison suggests an optimisation explanation rather than a state\-representation explanation alone: SAC exhibits lower realised volatility, smaller drawdown, shorter drawdown duration, and higher bear\-market outperformance than DP\-PPO under identical semantic inputs\. A plausible mechanism is that on\-policy PPO re\-uses each transition only once: when only 14\.8% of stock\-days carry non\-neutral semantic coordinates, PPO rollouts are dominated by neutral\-signal steps, and rare informative transitions are weighted equally with uninformative ones\. SAC uses auniform replay buffer\(standard SpinningUp implementation, not prioritised experience replay\): semantic\-rich transitions are revisited in proportion to their density in the buffer, not up\-weighted\. The “sparse\-signal replay” interpretation therefore does not require PER—it simply notes that when the replay buffer contains 50–100K transitions and only∼\\sim15% are semantically non\-neutral, SAC’s multiple gradient steps per environment step will naturally encounter semantic transitions more often than PPO’s single\-pass rollout\. Entropy regularisation further reduces premature policy collapse before semantic evidence accumulates\. This mechanism hypothesis is not verified here \(verifying it would require ablating the replay buffer or matching PPO on data reuse\); it is offered as a testable hypothesis for future work\.
## Appendix DCross\-domain SSAI instantiations
The SSAI framework \(Definition 1\) is not specific to portfolio trading*as a definition*\. Any sequential system that occasionally observes sparse unstructured text*could*use a fixed\-axis LLM interface as an auditable, sparse\-by\-design augmentation\.The present submission does not run experiments in these domains;the following are*illustrative*sketches only: clinical triage \(patient notes→\\rightarrowurgency/confidence/risk\); supply\-chain optimisation \(disruption news→\\rightarrowrisk/confidence/lead\-time\); recommendation \(reviews→\\rightarrowsentiment/engagement/quality\)\. A diagnostic suite analogous to ours—direct factor tests, residualization, masking, optimiser comparison—*could*be instantiated elsewhere; doing so is future work\. Prompts and the released harness are domain\-agnostic at the software level, but all numbers in the main paper refer to equities only\.
## Appendix EAdditional Results
### E\.1Regime Analysis and Sub\-period Breakdown
The 2019–2023 evaluation period spans three distinct regimes:
- •2019–2020 Q1: bull market followed by COVID crash \(S&P 500−34%\-34\\%in 33 days, February–March 2020\)\.
- •2020 Q2–2021: recovery and growth rally\.
- •2022: Federal Reserve rate hike bear market \(NASDAQ−33%\-33\\%in 2022\)\.
The turbulence index exceeded the threshold of 380 on 87 trading days \(6\.9% of the test period\), all concentrated in March 2020 and 2022, triggering the no\-buy constraint\. Table[9](https://arxiv.org/html/2605.06730#A5.T9)reports a more granular sub\-period breakdown for the released checkpoint\. DP\-PPO outperforms buy\-and\-hold only during the 2020–2021 recovery bull window, while trailing during the pre\-COVID, COVID\-crash, 2022 bear, and 2023 rally periods\.
Table 9:Sub\-period breakdown for the released DP\-PPO checkpoint versus equal\-weight buy\-and\-hold\.
### E\.2Transaction\-Cost Sensitivity
Table[10](https://arxiv.org/html/2605.06730#A5.T10)sweeps per\-trade costs from 0\.05% to 2\.0% for the released DP\-PPO checkpoint\. Because buy\-and\-hold incurs only the initial allocation in this simplified sweep, DP\-PPO remains below the passive sleeve at every tested cost level\.
Table 10:Transaction\-cost sensitivity for the released DP\-PPO checkpoint\. Buy\-and\-hold has no ongoing rebalance cost in this sweep; DP\-PPO remains below the passive sleeve across tested cost levels\.Figure 3:Transaction\-cost sensitivity for the released DP\-PPO checkpoint versus equal\-weight buy\-and\-hold\.
### E\.3Paired daily\-return diagnostics \(non\-RL baselines\)
Table[11](https://arxiv.org/html/2605.06730#A5.T11)reports paired daily\-return comparisons between semantic baselines and equal\-weight buy\-and\-hold using 20\-trading\-day block bootstrap confidence intervals and Wilcoxon signed\-rank tests\. These tests are provided as diagnostics rather than definitive multiple\-comparison\-adjusted claims\.
Table 11:Paired daily\-return diagnostics: semantic baselines plus multi\-seed RL vs EW benchmark\. The confidence interval is a 20\-trading\-day block bootstrap over mean paired daily active returns, in basis points per day\. Wilcoxon tests are paired over daily returns and are reported as diagnostics rather than definitive multiple\-comparison\-adjusted claims\.
### E\.4SFP Transaction\-Cost Sensitivity
Table[12](https://arxiv.org/html/2605.06730#A5.T12)sweeps per\-trade costs from 0\.05% to 2\.0% for the four\-factor SFP, mirroring Table 8 for DP\-PPO\. B&H uses the same equal\-weight price\-average index as Table 2 \(243\.6% CR\); SFP uses the same 0\.1%\-cost evaluated portfolio as reported in the main results \(307\.2% CR at 0\.1%\)\.Key finding:SFP’s daily top\-10 rebalancing generates high turnover; at 0\.1% \(the baseline used in the main text\) SFP exactly reproduces the reported 307\.2% CR, but at 0\.2% the gap over buy\-and\-hold reverses \(214\.5% vs\. 243\.6%\), and at 0\.5%\+ SFP collapses\. This means SFP’s main\-text outperformance is conditional on implementation at very low execution costs \(≤\\leq0\.1%\), which practical portfolio managers should treat as a strong caveat\.
Table 12:SFP transaction\-cost sensitivity \(2019–2023\)\. Each row re\-evaluates the daily top\-10 four\-factor SFP at the stated per\-trade cost\. B&H \(equal\-weight price\-average index, same computation as Table 2, 243\.6% CR\) is cost\-free throughout\.SFP is highly turnover\-sensitive: it outperforms B&H only at≤\\leq0\.1% per trade and underperforms at≥\\geq0\.2%, collapsing at 0\.5%\+\. The paper’s main reported result uses 0\.1%; realistic broker costs \(0\.2–0\.5%\) reverse the advantage\. Mirrors Table 8 \(DP\-PPO tx\-cost sensitivity\) for the direct portfolio\.
### E\.5SFP Sub\-period Performance
Table[13](https://arxiv.org/html/2605.06730#A5.T13)breaks down SFP performance by calendar year and regime window, mirroring Table 7 for DP\-PPO\. SFP outperforms buy\-and\-hold in 3 of 5 calendar years \(2020:\+\+11pp; 2021:\+\+4pp; 2023:\+\+26pp\) and underperforms in 2019 \(−\-8pp\) and 2022 \(−\-2pp\)\. The largest single\-year outperformance \(\+\+26pp in 2023\) coincides with the AI\-investment\-cycle rally that disproportionately benefited NVDA, GOOGL, and AVGO—the high\-coverage NASDAQ names that dominate the SFP basket \(see SRF/SCW coincidence discussion in Section[5\.2](https://arxiv.org/html/2605.06730#S5.SS2)\)\.Caveat:it is not possible to rule out that SFP’s 2023 excess return is a disguised factor exposure to the AI\-driven index rally rather than a semantic signal effect; a sector\-neutral SFP design would be needed to isolate these contributions, which we leave as future work\. Unlike DP\-PPO—which outperforms only in the 2020–21 recovery—SFP’s positive excess return is concentrated in the 2020–21 recovery \(\+\+39pp\) and 2023, both periods of high momentum in the NASDAQ\-100 names that dominate the SFP selection basket\.
Table 13:SFP sub\-period performance \(2019–2023\)\. Mirrors Table 7 \(DP\-PPO sub\-periods\)\. Excess CR = SFP minus equal\-weight B&H \(same computation as Table 2\)\. SFP outperforms in 3 of 5 calendar years \(2020, 2021, 2023\)\. The single largest annual outperformance \(\+26pp\) occurs in 2023, coinciding with the AI investment cycle rally that disproportionately benefited the high\-coverage NASDAQ names \(NVDA, GOOGL, AVGO\) that dominate the SFP basket; it is not possible to rule out that SFP’s 2023 excess is a disguised factor exposure rather than a semantic signal effect\. SFP underperforms in 2019 and 2022 \(rate\-hike, drawdown year\)\. Cost = 0\.1% per trade\.
### E\.6High\-Coverage Underperformance Paradox
The highest\-coverage tercile \(NVDA, GOOGL, AVGO, ASML, MU, QCOM, …\) shows the worst SSAI relative performance \(SFP 232\.6% vs\. B&H 352\.1% within\-stratum\)\. Three mechanisms explain this\. \(1\)Market efficiency: high\-coverage names have the most analyst attention; LLM\-scored news headline information is priced rapidly, leaving little marginal alpha\. \(2\)AI\-cycle momentum: daily rebalancing responds to high\-risk or negative\-sentiment signals by rotating*out*of positions, precisely when AI\-momentum was strongest in these names\. \(3\)Factor model misspecification: SFP uses 2013–2018 training weights; the structural shift of the 2022–2023 AI\-investment cycle invalidates these weights for the highest\-growth test\-period names\. Together, these predict SSAI\-based active management underperforms passive exposure in high\-momentum, efficiently\-priced large caps during momentum rallies—consistent with SSAI being a risk\-management and interpretability interface rather than a return\-generation signal in these names\.
### E\.7SFP/SRF/SCW Numerical Coincidence
SFP, SRF, and SCW share identical Sharpe \(1\.067\) and MDD \(−\-35\.6%\) because all three use the same daily top\-10 rule and basket membership is highly stable across weight configurations—residualising sentiment \(SRF\) and applying conviction scaling \(SCW\) reorder weights within nearly the same basket\. Sortino/Calmar differences \(SFP: 1\.430/0\.911 vs\. SCW: 1\.460/0\.924\) reflect modest conviction\-weighting improvement in downside management\. The dominant driver is basket selection, confirmed by Table[3](https://arxiv.org/html/2605.06730#S5.T3)\.
### E\.8Coverage\-Stratified SFP Analysis
Table[3](https://arxiv.org/html/2605.06730#S5.T3)\(presented in Section[5\.2](https://arxiv.org/html/2605.06730#S5.SS2)\) addresses the coverage confound directly\. We split the 30\-ticker universe into terciles by non\-neutral coverage fraction and evaluate SFP within each tercile, comparing against an equal\-weight B&H sleeve of the*same tickers*\.Key result: SFP does not outperform equal\-weight B&H within any individual coverage tercile\.The full\-portfolio outperformance \(307\.2% vs\. 243\.6%\) reflects a cross\-universe*composition*effect—SFP’s top\-10 rule preferentially selects mid\-coverage names \(INTU, KLAC, LRCX, TSLA, ADBE\) that happened to deliver strong returns, not names on which the semantic signal is demonstrably more predictive\. This result substantially qualifies the main SFP finding; full discussion and the table appear in Section[5\.2](https://arxiv.org/html/2605.06730#S5.SS2)\.
Causal chain from SSAI scores to basket membership\.For completeness: \(1\) DeepSeek\-V3 assigns integer scoresσds∈\{1,…,5\}\\sigma\_\{d\}^\{s\}\\in\\\{1,\\ldots,5\\\}to each \(article, ticker\) pair from the FNSPID corpus; \(2\) neutral imputation fills no\-news days at 3; \(3\) ridge regression on the 2013–2018 training panel fits weightsw∈ℝ4w\\in\\mathbb\{R\}^\{4\}minimising 5\-day return prediction error; \(4\) at each OOS datedd, the composite scorey^ds=w⊤σds\\hat\{y\}\_\{d\}^\{s\}=w^\{\\top\}\\sigma\_\{d\}^\{s\}ranks the 30 tickers; \(5\) the top\-10 byy^\\hat\{y\}form the equal\-weight basket for the next trading day\.*If the composition effect explains the full SFP advantage*, the implication is that the ridge weights learned in step \(3\) are correlated with which tickers are frequently scored—they pick mid\-coverage names as a side effect of the training signal, not because their SSAI scores are predictive on those names\. This is precisely what the within\-stratum analysis \(Table[3](https://arxiv.org/html/2605.06730#S5.T3)\) confirms: within any fixed coverage tercile, SFP does not outperform equal\-weight B&H\. The LLM scores are a necessary ingredient \(without them there is no ranking\), but the performance originates from which names are*selected*, not from the direction of the scores on selected names\.
## Appendix FLLM Signal Validation
Table 14:Descriptive statistics of LLM\-generated signal scores \(39,995 articles\)\.Figure 4:LLM semantic signal validation on the full 2013–2023 panel \(N=83,040N\{=\}83\{,\}040stock\-days, 18,578 non\-neutral\)\. \(A\) Score distributions on non\-neutral days show structured directional skew\. \(B\) Inter\-signal Spearman correlations replicate expected semantic structure\. \(C\) Lag\-1 temporal autocorrelation per ticker confirms scores are not white noise\. \(D\) Signal–return IC: confidence shows a significant positive IC \(p=0\.004p\{=\}0\.004\)\.Table 15:Ticker\-level LLM signal coverage on the 2019–2023 trade panel\. Coverage is the percentage of stock\-days whose score differs from the neutral imputation value 3\.0; rows show the five most\-covered and five least\-covered tickers\.Table 16:Signal validity proxy: pooled Spearman information coefficients between LLM signal coordinates and 5\-day forward returns / absolute returns on the 2019–2023 stock\-date panel\. This is not a human annotation audit, but it checks whether scores align with realized market quantities rather than pure noise\.
## Appendix GEthics, Deployment, and Reproducibility
This work is a retrospective research benchmark, not a trading recommendation or deployable advisory system\. Automated trading systems can amplify losses, liquidity stress, and feedback loops if deployed without governance, risk limits, and human oversight\. Our evaluation omits important operational frictions—slippage, partial fills, exchange outages, short\-sale constraints, tax effects, and capital limits—so reported results should not be interpreted as expected live performance\. The LLM component introduces additional reproducibility concerns: hosted model behaviour can drift over time; prompts may elicit different outputs across providers or versions; and financial news coverage is uneven across firms\. Scored article\-level signals should be cached and released with prompt templates, model identifiers, timestamps, and dataset hashes\. The artifact package provides scripts, processed CSV panels, checkpoints, and dry\-run commands for retrained ablations, but full reproducibility still depends on access to the same market data snapshot and LLM\-scored signal cache\.Similar Articles
AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining
This preprint introduces AlphaSchema, a framework that constructs and explores a structured space of trading semantics for LLM-based alpha mining, decoupling semantic exploration from factor implementation. Experiments on the Chinese stock market show that it discovers factor pools with strong predictive and portfolio performance.
Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems
Proposes Semantic Cooperative Games (SCG) and Semantic Shapley Value (SSV) for contribution attribution in LLM-based multi-agent systems, introducing a single-trajectory algorithm SLIC that reduces computation cost by 93.3% while remaining consistent with Monte Carlo Shapley baselines.
Learning Predictive Ambiguity Sets for Decision-Focused Distributionally Robust Optimization
Proposes learned predictive ambiguity sets (LPAS) for distributionally robust optimization, where a deep contextual model outputs a nominal scenario distribution, state-dependent Wasserstein radius, and ground metric, trained with decision loss and calibration. Applied to portfolio optimization on S&P 500 data, the method achieves higher returns and Sharpe ratio with reduced conservatism compared to fixed-radius baselines.
Consistency Analysis of Sentiment Predictions using Syntactic & Semantic Context Assessment Summarization (SSAS)
This paper presents SSAS (Syntactic & Semantic Context Assessment Summarization), a framework designed to improve consistency in LLM-based sentiment prediction by reducing noise and variance through hierarchical classification and iterative summarization. Empirical evaluation on three industry-standard datasets shows up to 30% improvement in data quality and reliability for enterprise decision-making.
InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance
InsightSR is a framework that leverages Large Language Models to refine the search space for symbolic regression, improving accuracy and physical consistency through iterative semantic and structural guidance.