Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment
Summary
This paper tests the assumption that sentiment tools validated against human labels also predict market signals in financial NLP, finding that benchmark agreement does not determine predictive effectiveness and spam influences volume-based analysis.
View Cached Full Text
Cached at: 09/12/26, 08:22 AM
# Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment
Source: [https://arxiv.org/html/2609.11144](https://arxiv.org/html/2609.11144)
Laven SrivastavaAffiliation:PerssonifyHarsh NandwaniAffiliation:Perssonifyresearch@perssonify\.com
###### Abstract
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal\. This assumes the two evaluations measure the same thing\. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions \(2002\-\-2025\) linking 70,500 X111formerly Twittermessages to abnormal stock returns, with a single\-annotator human labelled gold sample\. Running five instruments \(VADER, Loughran–McDonald, FinBERT, Twitter\-RoBERTa, and an LLM annotator\) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation\. Under conventional method\-specific sampling, human agreement aligns more closely with graded same\-day associations than with one\-day leads\. On a fixed\-nnpanel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak\. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings\. In a conversation that is 17\.6% spam, message volume predicts neither market damage nor settlement size\.
## 1Introduction
Sentiment instruments are validated in natural language processing the way classifiers are generally validated: by agreement with human annotation on a labeled sample\. In finance, they are used for a different purpose, to capture a signal that is associated with, and ideally precedes, movements in asset prices\. These two criteria are routinely treated as though the first implied the second\. An instrument that scores well on a sentiment benchmark is assumed to be the better instrument for market signal extraction\. That assumption is rarely tested, because the two evaluations are normally conducted on different corpora, human labeled benchmarks carry no market outcomes and market datasets carry no human labels\.
This paper tests the assumption directly: We construct a corpus in which the same messages carry both\. Securities class actions provide the linkage: court complaints supply event anchors and litigation metadata, the defendant’s price history supplies abnormal returns, and the surrounding social media conversation supplies text that we annotate both automatically, with five instruments, and manually, on a gold sample\. All instruments use the same aggregation and testing procedures\. We compare both method\-specific samples and a shared set of days, since the days retained can also affect the results\.
The domain is deliberately adverse; corrective disclosures and filings generate genuine investor discussion, but also plaintiff firm solicitation, repeated news headlines, and automated promotion\. The solicitation is structural, the PSLRA requires the first filing plaintiff to publish notice inviting other investors to seek lead plaintiff appointment, so every case mechanically produces a burst of law firm press releases\([Choi et al\., 2024](https://arxiv.org/html/2609.11144#bib.bib43)\)\. Headline repetition is a known feature of financial news\([Tetlock, 2011](https://arxiv.org/html/2609.11144#bib.bib42)\), and bots piggyback the cashtags of newsworthy firms to promote unrelated securities\([Cresci et al\., 2019](https://arxiv.org/html/2609.11144#bib.bib41)\)\. Spam is 17\.6% of our corpus, and solicitation and spam together account for 25,166 messages\. Counting is therefore a poor measure by construction: a count absorbs all of these sources equally\. We ask three questions\.
- •RQ1\.Does agreement with human sentiment judgments predict the strength of an instrument’s association with abnormal returns?
- •RQ2\.If agreement and association diverge, at what horizon do they diverge, and what are the instruments responding to?
- •RQ3\.Which conclusions about litigation discourse are invariant to the choice of instrument, and which are not?
Our contributions are as follows\.
1. 1\.A dual validity corpus, released\.We release 66,890 event anchored messages across 845 securities class actions \(2002\-\-2025\), carrying five instrument sentiment labels and multi axis LLM annotation, anchored by a 400 message human gold standard with its codebook and agreement statistics; to our knowledge, the first public financial social media corpus in which the same messages carry human annotation and linked market outcomes, so both validities can be measured through one pipeline\.222[https://bit\.ly/4iSUf6d](https://bit.ly/4iSUf6d)
2. 2\.A sampling dependent relationship between the two validities\.Under conventional method\-specific sampling, human agreement aligns more closely with graded same\-day associations than with one\-day leads\. On the fixed\-nnpanel, the graded ordering is similar at both horizons, whereas the coarse ordering remains weak\. The relationship between construct and predictive validity therefore depends on horizon, representation, and the treatment of zero\-score days\.
3. 3\.Mechanism, and instability of instrument rankings\.Under conventional method\-specific sampling, the instrument with the largest one\-day point estimate is pretrained on financial news; news accounts carry the strongest attributed signal \(ρ=−0\.261\\rho=\-0\.261\), and the local\-projection placebo fails at negative horizons\.
4. 4\.Domain result and informative nulls\.Content is informative where volume is not, despite systematic pollution; message volume predicts neither the depth of the price decline nor dollar settlement size\.
## 2Related Work
#### Content and volume in financial discourse\.
The distinction between what market participants say and how much they post originates in the message board literature:[Tumarkin and Whitelaw \(2001\)](https://arxiv.org/html/2609.11144#bib.bib9)find little return information in posting activity, while[Antweiler and Frank \(2004\)](https://arxiv.org/html/2609.11144#bib.bib8)show that message*content*carries signal where volume mainly predicts volatility, and[Das and Chen \(2007\)](https://arxiv.org/html/2609.11144#bib.bib10)build the first purpose made sentiment classifier for stock talk\. Media pessimism predicts market activity\([Tetlock, 2007](https://arxiv.org/html/2609.11144#bib.bib11)\), firm level language predicts fundamentals\([Tetlock et al\., 2008](https://arxiv.org/html/2609.11144#bib.bib12)\), and the content of crowd\-sourced investment opinions predicts returns and earnings surprises\([Chen et al\., 2014](https://arxiv.org/html/2609.11144#bib.bib13)\)\. Message volume is one member of a family of attention measures that includes search intensity\([Da et al\., 2011](https://arxiv.org/html/2609.11144#bib.bib14)\)and the attention driven trading it induces\([Barber and Odean, 2008](https://arxiv.org/html/2609.11144#bib.bib15)\)\. We treat this content over volume regularity as established and use it as a domain check; the paper’s claim concerns how the instruments that measure content should be evaluated\.
#### Sentiment instruments and their benchmark evaluation\.
The instruments we compare span the standard toolkit: a social media lexicon\([Hutto and Gilbert, 2014](https://arxiv.org/html/2609.11144#bib.bib18)\), a finance dictionary\([Loughran and McDonald, 2011](https://arxiv.org/html/2609.11144#bib.bib16)\), and transformers pretrained on financial news\([Araci, 2019](https://arxiv.org/html/2609.11144#bib.bib19)\)and on tweets\([Barbieri et al\., 2020](https://arxiv.org/html/2609.11144#bib.bib20)\)\. Such instruments are standardly evaluated by agreement with human annotation, with Financial PhraseBank\([Malo et al\., 2014](https://arxiv.org/html/2609.11144#bib.bib21)\)the reference benchmark; recent suites extend the same paradigm to financial large language models\([Xie et al\., 2023](https://arxiv.org/html/2609.11144#bib.bib39);[Xie et al\., 2024](https://arxiv.org/html/2609.11144#bib.bib40)\), making the question of what agreement predicts downstream more consequential, not less\.[Loughran and McDonald \(2016\)](https://arxiv.org/html/2609.11144#bib.bib17)argue that measurement choices dominate downstream conclusions; our results give that argument a sharp form: the choice between two validated instruments changes which temporal conclusions are recoverable\.
#### Intrinsic versus extrinsic evaluation\.
The distinction we draw between agreement with human labels and association with outcomes is the psychometric distinction between construct and predictive validity\([Cronbach and Meehl, 1955](https://arxiv.org/html/2609.11144#bib.bib22)\), and it has an NLP precedent: intrinsic evaluations of word representations fail to predict extrinsic task performance\([Chiu et al\., 2016](https://arxiv.org/html/2609.11144#bib.bib23)\), and measurement\-theoretic critiques argue that NLP systems routinely operationalise constructs without testing what the operationalisation measures\([Jacobs and Wallach, 2021](https://arxiv.org/html/2609.11144#bib.bib24)\)\. To our knowledge the two evaluations have not previously been conducted on the same financial messages with linked market outcomes, which is what permits the horizon localised comparison in Section[7](https://arxiv.org/html/2609.11144#S7)\.
#### LLMs in financial text\.
Large language models extract return relevant signal from headlines\([Lopez\-Lira and Tang, 2023](https://arxiv.org/html/2609.11144#bib.bib25)\)and outperform transformer and dictionary sentiment in trading settings\([Kirtac and Germano, 2024](https://arxiv.org/html/2609.11144#bib.bib26)\); LLM annotation can match or exceed crowd workers\([Gilardi et al\., 2023](https://arxiv.org/html/2609.11144#bib.bib27)\)\. Two caveats shape our design: LLM sentiment can embed look ahead information from the pretraining window\([Glasserman and Lin, 2023](https://arxiv.org/html/2609.11144#bib.bib28)\), motivating the contamination analysis of Section[9](https://arxiv.org/html/2609.11144#S9), with chronologically consistent training the proposed remedy\([He et al\., 2025](https://arxiv.org/html/2609.11144#bib.bib29)\)\.
#### Social media, returns, and securities litigation\.
Relations between social media mood and market movement are established\([Bollen et al\., 2011](https://arxiv.org/html/2609.11144#bib.bib30);[Sprenger et al\., 2014](https://arxiv.org/html/2609.11144#bib.bib31);[Ranco et al\., 2015](https://arxiv.org/html/2609.11144#bib.bib32)\), as is the link from investor disagreement to trading volume\([Cookson and Niessner, 2020](https://arxiv.org/html/2609.11144#bib.bib33)\); class action event studies measure shareholder wealth effects, litigation risk, and reputational penalties without social media text\([Gande and Lewis, 2009](https://arxiv.org/html/2609.11144#bib.bib34);[Kim and Skinner, 2012](https://arxiv.org/html/2609.11144#bib.bib35);[Karpoff et al\., 2008](https://arxiv.org/html/2609.11144#bib.bib36)\)\. Our setting differs in being event anchored by court filings, carrying eventual legal outcomes, and evaluating the measurement instruments themselves rather than any one instrument’s signal\. The press is a documented fraud detection channel\([Miller, 2006](https://arxiv.org/html/2609.11144#bib.bib37);[Dyck et al\., 2010](https://arxiv.org/html/2609.11144#bib.bib38)\), consistent with our attribution result that news accounts carry the strongest price relevant signal and with the news arrival reading of the leading component\.
## 3Corpus and Annotation
### 3\.1Corpus construction
The corpus contains 845 securities class actions in three cohorts \(Table[1](https://arxiv.org/html/2609.11144#S3.T1)\): prospective 2024 and 2025 cohorts and a retrospective archive of resolved cases from 2002–2021\. Court filings supply company identity, class period bounds, corrective disclosure dates, allegation summaries, defendants, and litigation metadata\. Filings are ingested from PDF and converted to structured records\. High value fields \(company name, ticker, class period bounds, defendants, and disclosure dates\) receive an independent extraction pass, and disagreements are manually reconciled\. These dates define the message retrieval and event windows\.
Table 1:Corpus by cohort\. “Social” counts cases with collected messages; “Msgs\.” counts annotated messages\.
### 3\.2Composition of the conversation
The label distributions in this subsection are produced by the LLM instrument described in Section[3\.4](https://arxiv.org/html/2609.11144#S3.SS4)\. The relevance, topic, and polarity taxonomies are closed sets fixed in the annotation prompt, which is released with the code and data\.
Of 70,500 messages, 62\.0% are relevant, 20\.3% tangential, and 17\.6% spam\. The largest topics are news reporting \(13,880\), law firm solicitation \(12,901\), automated spam \(12,265\), legal procedure \(9,658\), equity analysis \(8,055\), and direct fraud allegations \(7,835\)\. Polarity is 41\.0% negative and 7\.2% positive\. Relevance and polarity agree substantially with human annotation \(κ=0\.746\\kappa=0\.746and0\.7410\.741; Section[3\.5](https://arxiv.org/html/2609.11144#S3.SS5)\), while the topic taxonomy is not separately validated and is therefore used descriptively only\.
A count based attention measure aggregates informative reporting, investor reaction, solicitation, and automation into a single number, motivating the comparison in Section[5](https://arxiv.org/html/2609.11144#S5)\.
### 3\.3Event and market linkage
Messages are assigned to a pre class period baseline, the alleged class period, the corrective disclosure window, and post disclosure and post filing windows, retaining text, timestamp, impressions, and likes where available\. For companyiion daytt, letRi,tR\_\{i,t\}denote the realised return andRm,tR\_\{m,t\}the market return\. A market model
Ri,t=αi\+βiRm,t\+ϵi,tR\_\{i,t\}=\\alpha\_\{i\}\+\\beta\_\{i\}R\_\{m,t\}\+\\epsilon\_\{i,t\}\(1\)is estimated over the 120 trading days ending ten days before the class period begins, so the parameters are fitted on pre event data only\. The*abnormal return*is the part of the day’s return the market does not explain,
𝐴𝑅i,t=Ri,t−\(α^i\+β^iRm,t\),\\mathit\{AR\}\_\{i,t\}=R\_\{i,t\}\-\\big\(\\hat\{\\alpha\}\_\{i\}\+\\hat\{\\beta\}\_\{i\}R\_\{m,t\}\\big\),\(2\)and the*cumulative abnormal return*over an event windowWWis its sum,𝐶𝐴𝑅i=∑t∈W𝐴𝑅i,t\\mathit\{CAR\}\_\{i\}=\\sum\_\{t\\in W\}\\mathit\{AR\}\_\{i,t\}\.
### 3\.4Scoring instruments
Five instruments score message sentiment: VADER, a social media lexicon\([Hutto and Gilbert, 2014](https://arxiv.org/html/2609.11144#bib.bib18)\); Loughran–McDonald, a finance dictionary\([Loughran and McDonald, 2011](https://arxiv.org/html/2609.11144#bib.bib16)\); FinBERT and Twitter\-RoBERTa, transformers pretrained on financial news and on tweets\([Araci, 2019](https://arxiv.org/html/2609.11144#bib.bib19);[Barbieri et al\., 2020](https://arxiv.org/html/2609.11144#bib.bib20)\); and Claude Haiku, an LLM annotator\. The instruments produce different outputs, so we derive coarse and graded negativity scores, with higher values indicating more negative sentiment\. The*coarse*score discards magnitude\. Each message is assigned to one of three classes and mapped to\{\+1,0,−1\}\\\{\+1,0,\-1\\\}with\+1\+1negative, using the predicted class for the transformers and the LLM, and the sign of the score for the lexicons\. The*graded*score keeps magnitude, oriented so that larger is more negative: the negated VADER compound score, the Loughran–McDonald net\-negative word ratio,p\(neg\)−p\(pos\)p\(\\text\{neg\}\)\-p\(\\text\{pos\}\)for the transformers, and the negated−2\-2to\+2\+2intensity label for the LLM\.
Both scores are averaged over the messages for each case day\. All instruments use the same aggregation and tests, but the conventional samples differ because zero\-score days are excluded separately for each instrument\. Section[7\.1](https://arxiv.org/html/2609.11144#S7.SS1)repeats the comparison on a shared set of days\.
### 3\.5Construct validity
We assess agreement with human judgment on a gold set of 400 messages, labelled by a single annotator blind to all model outputs and presented in randomised order\. Messages were drawn by stratified random sampling across predicted relevance×\\timespolarity×\\timescohort \(26 strata\); inverse probability weighted accuracy \(0\.861 polarity, 0\.860 relevance\) matches the unweighted estimates, so stratification does not drive the reported agreement\.
Table[2](https://arxiv.org/html/2609.11144#S3.T2)reports polarity agreement for all five instruments\. We report Cohen’sκ\\kappa\([Cohen, 1960](https://arxiv.org/html/2609.11144#bib.bib1)\)alongside accuracy because the corpus is 41\.0% negative and 7\.2% positive, so an instrument that predicts the majority class attains substantial accuracy while carrying no information\. FinBERT is precisely this case: accuracy 0\.538 withκ=0\.079\\kappa=0\.079\. Under conventional benchmarks\([Landis and Koch, 1977](https://arxiv.org/html/2609.11144#bib.bib2)\)only the LLM instrument reaches substantial agreement; Twitter\-RoBERTa is moderate, Loughran–McDonald fair, and FinBERT indistinguishable from chance\. VADER’s accuracy is below the majority class rate\.
The LLM instrument’s polarity errors concentrate at the negative neutral boundary and are directionally one sided: it labels human neutral messages negative in 49 of 56 disagreements, so LLM derived negativity is if anything slightly inflated\. On relevance it achieves 0\.860 accuracy, 0\.797 macro\-F1, andκ=0\.746\\kappa=0\.746\. This evaluation measures agreement with the intended linguistic constructs; it makes no claim about association with returns, which is the subject of the next two sections\.
Table 2:Polarity agreement with human gold labels on the same 400 messages\. Brackets give 95% bootstrap intervals from 1,000 resamples\([Efron and Tibshirani, 1993](https://arxiv.org/html/2609.11144#bib.bib3)\)\. LM denotes Loughran–McDonald; RoBERTa denotes Twitter\-RoBERTa\.
## 4Empirical Setup
We report Spearmanρ\\rhofor monotonic associations, which is robust to the heavy tails of daily returns\. Daily analyses pool case days; cross sectional analyses collapse each case to one observation\. Significance markers are∗∗∗p<0\.001\{\}^\{\*\*\*\}p<0\.001,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01,∗p<0\.05\{\}^\{\*\}p<0\.05,†p<0\.1\{\}^\{\\dagger\}p<0\.1\. New baseline families use Benjamini–Hochberg FDR adjustment\([Benjamini and Hochberg, 1995](https://arxiv.org/html/2609.11144#bib.bib4)\)\. Because the daily panel is large, statistical significance is easily attained; we therefore compare instruments on effect size and treatpp\-values as evidence only that an association is nonzero\.
Timing proceeds from descriptive to conditional tests\. Lead lag correlations compare sentiment ont−ℓt\-\\ellwithARtAR\_\{t\}\. Order two Granger tests ask whether lagged sentiment improves prediction beyond return history\([Granger, 1969](https://arxiv.org/html/2609.11144#bib.bib5)\)\. A distributed lag model,
ARi,t=∑k=03βksi,t−k\+ηi\+ui,t,AR\_\{i,t\}=\\sum\_\{k=0\}^\{3\}\\beta\_\{k\}s\_\{i,t\-k\}\+\\eta\_\{i\}\+u\_\{i,t\},\(3\)uses case fixed effects and case clustered standard errors\([Petersen, 2009](https://arxiv.org/html/2609.11144#bib.bib7)\)\. Local projections\([Jordà, 2005](https://arxiv.org/html/2609.11144#bib.bib6)\)estimateARi,t\+hAR\_\{i,t\+h\}onsi,ts\_\{i,t\}forh∈\[−5,5\]h\\in\[\-5,5\], with negative horizons serving as a placebo\. These tests establish predictive precedence, not structural causation\. Between\-instrument differences inρ\\rhoare not directly tested; comparisons of their magnitudes are therefore descriptive point\-estimate comparisons, and significance for one instrument but not another does not establish a significant difference between them\.
The panel contains 179 cases and 103,542 case days\.333The panel was re fetched because the original price series were unavailable\. Of 273 archive tickers, 74 no longer resolve\. Published values reproduce within\|Δρ\|≤0\.002\|\\Delta\\rho\|\\leq 0\.002, with identical signs and significance ordering\. All cross method comparisons use the refreshed same data panel\.The conventional correlation sample drops days on which a method emits zero, sonndiffers by method; Section[7\.1](https://arxiv.org/html/2609.11144#S7.SS1)reports a common day intersection in whichnnis held fixed\.
## 5Predictive Validity: Content versus Counting
At case level, attention does not track damage\. In the 2025 cohort, total message volume is uncorrelated with the worst single\-day abnormal return \(ρ=−0\.083\\rho=\-0\.083,p=0\.60p=0\.60,n=43n=43\) and with CAR \(ρ=\+0\.003\\rho=\+0\.003,p=0\.99p=0\.99,n=36n=36\)\.
The daily panel establishes the comparison precisely \(Table[3](https://arxiv.org/html/2609.11144#S5.T3), summarised in Figure[1](https://arxiv.org/html/2609.11144#S5.F1)\)\. Tweet volume reachesρ=−0\.0311\\rho=\-0\.0311\. Every sentiment family contains a representation with a larger association, and the strongest instruments exceed volume by a factor of two to three\. Loughran–McDonald coarse \(0\.80×0\.80\\times\) is the single exception, so the claim applies to method families rather than to every threshold choice\. Impressions, the natural reach based count, are not significant at all\.
Two features of this table matter for what follows\. First, the strongest same day instrument is Twitter RoBERTa, not the LLM, so the content over counting result does not depend on the LLM annotation\. Second, thepp\-values are not comparable across rows: RoBERTa graded attainsp=5\.27×10−18p=5\.27\\times 10^\{\-18\}on 12,910 case days while FinBERT coarse attainsp=1\.87×10−6p=1\.87\\times 10^\{\-6\}on 5,056, and the difference is largely sample size\. We compare theρ\\rhocolumn throughout\.
Metricρ\\rhoppnn×\\timesvol\.CountsTweet volume−0\.0311∗∗∗\-0\.0311^\{\*\*\*\}4\.14×10−44\.14\{\\times\}10^\{\-4\}12,9101\.00×1\.00\\timesImpressions−0\.0317†\-0\.0317^\{\\dagger\}0\.06340\.06343,4391\.02×1\.02\\timesNegative count−0\.0380∗\-0\.0380^\{\*\}0\.01350\.01354,2191\.22×1\.22\\timesLexiconsVADER c\.−0\.0454∗∗∗\-0\.0454^\{\*\*\*\}8\.86×10−68\.86\{\\times\}10^\{\-6\}9,5661\.46×1\.46\\timesVADER g\.−0\.0454∗∗∗\-0\.0454^\{\*\*\*\}3\.78×10−63\.78\{\\times\}10^\{\-6\}10,3391\.46×1\.46\\timesLM c\.−0\.0249∗\-0\.0249^\{\*\}0\.04140\.04146,7130\.80×0\.80\\timesLM g\.−0\.0559∗∗∗\-0\.0559^\{\*\*\*\}3\.60×10−63\.60\{\\times\}10^\{\-6\}6,8711\.80×1\.80\\timesTransformersFinBERT c\.−0\.0670∗∗∗\-0\.0670^\{\*\*\*\}1\.87×10−61\.87\{\\times\}10^\{\-6\}5,0562\.16×2\.16\\timesFinBERT g\.−0\.0526∗∗∗\-0\.0526^\{\*\*\*\}2\.29×10−92\.29\{\\times\}10^\{\-9\}12,9101\.69×1\.69\\timesRoBERTa c\.−0\.0967∗∗∗\-0\.0967^\{\*\*\*\}7\.91×10−97\.91\{\\times\}10^\{\-9\}3,5453\.11×3\.11\\timesRoBERTa g\.−0\.0760∗∗∗\-0\.0760^\{\*\*\*\}5\.27×10−185\.27\{\\times\}10^\{\-18\}12,9102\.45×2\.45\\timesLLM instrumentClaude c\.−0\.0822∗∗∗\-0\.0822^\{\*\*\*\}5\.77×10−95\.77\{\\times\}10^\{\-9\}5,0002\.65×2\.65\\timesClaude g\.−0\.0821∗∗∗\-0\.0821^\{\*\*\*\}6\.12×10−96\.12\{\\times\}10^\{\-9\}5,0042\.64×2\.64\\timesTable 3:Same day correlation with the daily abnormal return\. LM denotes Loughran–McDonald; RoBERTa denotes Twitter\-RoBERTa; c\. and g\. denote coarse and graded scores\.×\\timesvol\. is\|ρ\|\|\\rho\|relative to tweet volume\.Figure 1:Content over volume across scoring instruments\. Panel \(a\) shows the strongest same day representation in each method family;nndiffers because the conventional analysis drops method specific zero days\. Panel \(b\) uses the identical 134 cases for event minus baseline sentiment shift versus CAR\. Bars report exactρ\\rhovalues without treating any instrument as the contribution\.
## 6Temporal Structure
Table[4](https://arxiv.org/html/2609.11144#S6.T4)reports lead lag correlations\. Every instrument shows a negative and significant same day association, so the contemporaneous result of Section[5](https://arxiv.org/html/2609.11144#S5)is not instrument specific\. Under the conventional method\-specific nonzero\-day samples, the one\-day results differ across instruments\. FinBERT coarse has the largestℓ=1\\ell=1point estimate \(−0\.0466\-0\.0466,pFDR=0\.003p\_\{\\mathrm\{FDR\}\}=0\.003\); FinBERT graded and Twitter\-RoBERTa graded also survive FDR, whereas VADER and Loughran–McDonald do not\. Atℓ=2\\ell=2, Claude graded and Twitter\-RoBERTa graded are significant at the unadjusted 5% level\. These comparisons describe point estimates and do not establish statistically significant differences between instruments\.
Table 4:Lead lag correlations, all messages; c\. and g\. denote coarse and graded scores;nnper row as in Table[3](https://arxiv.org/html/2609.11144#S5.T3)\.qqvalues are FDR adjusted across the baseline family; the pre existing LLM reference is not included in that correction\.Conditional tests give the same picture\. Four baseline specifications survive FDR on forward Granger tests: Loughran–McDonald coarse/relevant \(F=4\.388F=4\.388,q=0\.034q=0\.034\), Loughran–McDonald graded/all \(F=4\.056F=4\.056,q=0\.040q=0\.040\), FinBERT coarse/relevant \(F=3\.907F=3\.907,q=0\.044q=0\.044\), and FinBERT graded/all \(F=4\.316F=4\.316,q=0\.034q=0\.034\)\. The LLM reference also passes \(F=5\.435F=5\.435,p=0\.0051p=0\.0051coarse/all\)\. Twitter\-RoBERTa never survives, despite having the strongest same day association\. Same day distributed lag coefficients are negative and significant for almost every instrument; their magnitudes are not comparable across instruments because the underlying signals differ in scale\. Full results appear in Appendix[B](https://arxiv.org/html/2609.11144#A2)\.
## 7Where the Two Validities Diverge
Sections[3\.5](https://arxiv.org/html/2609.11144#S3.SS5)and[6](https://arxiv.org/html/2609.11144#S6)measure two different properties of the same five instruments on the same corpus: agreement with human judgment, and association with abnormal returns\. Table[5](https://arxiv.org/html/2609.11144#S7.T5)places them side by side\. This is the paper’s central comparison, and it is not visible from either table in isolation\.
Construct validityPredictive validity \(\|ρ\|\|\\rho\|\)Instrumentκ\\kappaAccuracyMacro\-F1coarseρ0\\rho\_\{0\}coarseρ1\\rho\_\{1\}gradedρ0\\rho\_\{0\}gradedρ1\\rho\_\{1\}Claude Haiku0\.7410\.8600\.8780\.08220\.02270\.08210\.0376Twitter\-RoBERTa0\.4490\.7330\.5940\.09670\.00580\.07600\.0272Loughran–McDonald0\.2100\.5530\.4460\.02490\.00570\.05590\.0263VADER0\.1710\.3850\.3730\.04540\.01360\.04540\.0193FinBERT0\.0790\.5380\.3790\.06700\.04660\.05260\.0292Rank correlation withκ\\kappa\+0\.50\+0\.50−0\.30\-0\.30\+0\.90†\+0\.90^\{\\dagger\}\+0\.40\+0\.40Rank correlation with accuracy\+0\.60\+0\.60−0\.10\-0\.10\+1\.00∗\+1\.00^\{\*\}\+0\.70\+0\.70Table 5:Construct validity from Table[2](https://arxiv.org/html/2609.11144#S3.T2)joined to predictive validity from Tables[3](https://arxiv.org/html/2609.11144#S5.T3)and[4](https://arxiv.org/html/2609.11144#S6.T4), ordered byκ\\kappa\. On conventional method\-specific nonzero\-day samples, human\-agreement rankings align most closely with graded same\-day association and less consistently with the one\-day lead\. Rank correlations are Spearman correlations over five instruments, with exact two\-sidedpp\-values from all 120 permutations\. These comparisons are descriptive, and no cell survives adjustment for the 12 comparisons in this table\. The corresponding fixed\-nncomparison is reported in Table[6](https://arxiv.org/html/2609.11144#S7.T6)and Section[7\.1](https://arxiv.org/html/2609.11144#S7.SS1)\.Conventional samples show contemporaneous alignment\. For the graded representation in Table[5](https://arxiv.org/html/2609.11144#S7.T5), the ordering by human agreement closely matches the ordering by same\-day association: the rank correlation is\+0\.90\+0\.90withκ\\kappaand\+1\.00\+1\.00with accuracy\. Thus, under this representation and the conventional method\-specific sampling rule, benchmark agreement is informative about contemporaneous association\.
The one\-day relationship is weaker and depends on the score representation\. For graded scores, the rank correlations are\+0\.40\+0\.40withκ\\kappaand\+0\.70\+0\.70with accuracy; for coarse scores, they are−0\.30\-0\.30and−0\.10\-0\.10\. Under conventional sampling, FinBERT has the strongest coarse one\-day association \(ρ1=−0\.0466\\rho\_\{1\}=\-0\.0466\) despite its low human agreement; Claude’s corresponding value is−0\.0227\-0\.0227\. Section[7\.1](https://arxiv.org/html/2609.11144#S7.SS1)shows how these rankings change when all instruments are evaluated on the same case\-day set\.
#### Interpretation\.
Table[4](https://arxiv.org/html/2609.11144#S6.T4)shows a horizon contrast under the conventional, method\-specific samples, but Table[6](https://arxiv.org/html/2609.11144#S7.T6)shows that this contrast is not invariant to sample construction or score representation\. On the fixed\-nngraded panel, agreement has the same rank correlation with association atℓ=0\\ell=0andℓ=1\\ell=1; on the coarse panel, both relationships are weak\. We therefore make the narrower claim that benchmark agreement does not, by itself, determine predictive rankings, and that comparisons should be reported with the horizon, representation, and zero\-day inclusion rule explicitly stated\.
### 7\.1Instrument rankings are unstable
The conventional analysis drops case days on which a given method emits zero, sonndiffers across instruments in Tables[3](https://arxiv.org/html/2609.11144#S5.T3)and[4](https://arxiv.org/html/2609.11144#S6.T4)and the comparison is not made on identical data\. Table[6](https://arxiv.org/html/2609.11144#S7.T6)restricts to case days on which every instrument emits a nonzero score, holdingnnfixed\.
Table 6:Common\-case\-day intersection, all messages\. Relevant\-message panels appear in Appendix[C](https://arxiv.org/html/2609.11144#A3)\.Applying the Table[5](https://arxiv.org/html/2609.11144#S7.T5)rank comparison to the fixed\-nnpanels gives correlations withκ\\kappaof \+0\.00 and \+0\.10 for coarse scores atℓ=0\\ell=0and 1, and \+0\.80 at both horizons for graded scores\. The corresponding values for accuracy and macro\-F1 are \+0\.40 and \+0\.20 for coarse scores, and \+0\.90 at both graded horizons\. The exact two\-sided permutationpp\-values are 1\.000, 0\.950, 0\.133, and 0\.133 forκ\\kappa, and 0\.517, 0\.783, 0\.083, and 0\.083 for accuracy and macro\-F1, respectively\. None of these correlations survives BH adjustment within the 12 fixed\-nnconstruct–predictive comparisons\.
Three observations follow\. First, the same day ranking reverses: on the coarse common days FinBERT is strongest \(−0\.1259\-0\.1259\), where the conventional sample placed Twitter\-RoBERTa first\. Second, VADER is indistinguishable from zero on common days, so its apparent advantage in Table[3](https://arxiv.org/html/2609.11144#S5.T3)was substantially a matter of which days it scored\. Third, several sentiment measures remain associated with returns on the shared days, although the strongest instrument changes\. Table[6](https://arxiv.org/html/2609.11144#S7.T6)does not compare sentiment with volume on those same days, so it does not establish that the content\-over\-counting result is unchanged\. Claims about the strongest instrument therefore need to specify the sample used\.
### 7\.2Divergence at the case level
A cross sectional test reproduces the divergence in a different form\. For each case we compute the event minus baseline change in mean negativity and correlate it with the event window CAR \(Table[7](https://arxiv.org/html/2609.11144#S7.T7)\)\. The LLM instrument attainsρ=−0\.3106\\rho=\-0\.3106\(p=0\.0003p=0\.0003,n=134n=134\), several times any daily panel effect\. No baseline survives FDR correction: FinBERT reaches−0\.1961\-0\.1961\(p=0\.0231p=0\.0231\) butq=0\.185q=0\.185, and Twitter\-RoBERTa−0\.1651\-0\.1651\(p=0\.0566p=0\.0566\)\. The case level shift is a contemporaneous contrast between two windows rather than a lead, so theℓ=0\\ell=0pattern of Table[5](https://arxiv.org/html/2609.11144#S7.T5)would predict the best agreeing instrument to perform well here, which it does\. We do not press that reading further, for two reasons\. Withn=134n=134and five instruments the comparison is underpowered, so the ordering among the baselines is not resolvable; and the result is sensitive to how the case level measure is constructed\. Substituting mean window negativity for negative share, on a reconstruction of the panel, raises Loughran–McDonald and Twitter\-RoBERTa to the same range as the LLM instrument\. The robust content of this table is that a case level contrast exists and is several times larger than any daily panel effect; the identity of the strongest instrument at this level is not established\. Restricting to relevant messages reduces the sample ton=50n=50and renders every instrument nonsignificant\.
Table 7:Event\-minus\-baseline change in the share of negative messages versus CAR, using the same 134 cases for all instruments and including all messages\.
## 8What the Leading Signal Responds To
Section[7](https://arxiv.org/html/2609.11144#S7)shows that the relationship between construct validity and the one\-day lead depends on the sampling convention and score representation\. Two pieces of evidence bear on what the observed leading associations reflect\.
#### The largest conventional\-sample lead is produced by an instrument trained on news\.
FinBERT is pretrained on financial news and evaluated on sentence\-level news sentiment\. On tweets about securities fraud, it agrees with human polarity at close to chance, yet its coarse score has the largest one\-day point estimate under conventional method\-specific sampling\. The natural reading is that it detects the vocabulary of adverse financial news rather than the polarity of investor expression, and that this vocabulary arrives shortly before the price adjustment completes\.
#### News accounts carry the strongest signal\.
We classify accounts from topic mix, cross\-case activity, and spam fraction\. News accounts have the strongest association \(ρ=−0\.261\\rho=\-0\.261,p<0\.001p<0\.001\), followed by cross\-case broadcasters \(−0\.161\-0\.161\), other multi\-case accounts \(−0\.115\-0\.115\), and retail accounts \(−0\.083\-0\.083\)\. The associations for bot/spam accounts \(\+0\.064\+0\.064\) and law firm solicitation \(\+0\.040\+0\.040\) are not significant\. The strongest observed signal therefore comes from news accounts, while these two sources of additional volume show no detectable association with returns\.
## 9Robustness
#### Look ahead contamination\.
The archive cohort \(2002–2021\) lies inside the LLM’s pretraining window, so its results could in principle reflect recognition of publicised cases rather than message level reading\([Glasserman and Lin, 2023](https://arxiv.org/html/2609.11144#bib.bib28)\)\. The baselines provide a control: the two dictionaries have no training window at all, and the two transformers were pretrained without these outcome labels\. Contamination would therefore show up as an LLM advantage that is larger in the archive than in the 2024–2025 cohorts\. On the daily panel it is not\. Across 1,000 case level bootstrap resamples, the archive minus recent difference in the LLM’s advantage is\+0\.0091\+0\.0091\(graded, same day; 95% CI\[−0\.027,\+0\.046\]\[\-0\.027,\+0\.046\]\),−0\.0013\-0\.0013\(graded, one day lag; CI\[−0\.047,\+0\.052\]\[\-0\.047,\+0\.052\]\),\+0\.0061\+0\.0061and\+0\.0017\+0\.0017\(coarse\)\. Every interval includes zero\. The LLM’s own association is also similar across cohorts \(−0\.081\-0\.081in the archive and−0\.090\-0\.090in the recent cohorts\)\. These results do not reveal a clear difference between cohorts, but they do not rule out contamination\. The LLM’s own association is stable across cohorts \(−0\.081\-0\.081archive,−0\.090\-0\.090recent\)\. We retain the archive for the daily analyses\. The case level measure of Section[7\.2](https://arxiv.org/html/2609.11144#S7.SS2)is less settled: split by cohort \(80 archive, 47 recent\), the LLM’s advantage is larger in the archive the direction contamination predicts\. The subsample is too small to treat as evidence either way, so the paper rests its claims on the daily panel and reports the case level result as corroborative \(decomposition in Appendix[D](https://arxiv.org/html/2609.11144#A4)\)\.
#### Relevance filtering\.
Restricting to relevant messages shifts same day correlations by at most0\.01640\.0164; the central result is present on all messages\. Since the relevance labels come from the LLM instrument, this removes a circularity concern \(Appendix[A](https://arxiv.org/html/2609.11144#A1)\)\.
#### Attention and legal materiality\.
Among 133 resolved cases with settlements, total message volume is uncorrelated with settlement amount \(ρ=−0\.041\\rho=\-0\.041,p=0\.64p=0\.64\), as are relevant message \(−0\.062\-0\.062\) and negative message \(−0\.093\-0\.093\) volume; median settlements do not differ across volume tertiles \(H=0\.83H=0\.83,p=0\.66p=0\.66\)\. Public attention is orthogonal to legal materiality, as it was to price damage depth\.
## 10Discussion and Conclusion
We introduce a corpus linking 845 securities class actions to 70,500 event\-anchored messages, pairing a 400\-message single\-annotator human reference sample with abnormal\-return series linked to the messages\. Evaluating five instruments through one pipeline, we find that the relationship between human agreement and return association is sensitive to sampling and representation\. On conventional method\-specific samples, human agreement closely orders graded contemporaneous associations but not one\-day leads; on the fixed\-nnpanel, graded rankings are similar at both horizons, while coarse rankings remain weak\.
Agreement with human labels is therefore evidence of semantic validity, not a sufficient basis for selecting an instrument for predictive use\. Construct and predictive validity should be reported together with the horizon, score representation, and zero\-day inclusion rule\.
The account\-level results and the failed placebo are consistent with a leading signal that partly reflects news arrival\. The daily\-panel comparison finds no clear archive\-recent difference in the LLM’s relative advantage, but does not rule out contamination\. The domain result stands: in a conversation that is 17\.6% spam, content informs where counting does not, and volume predicts neither the depth of the price decline nor the size of the eventual settlement\. The contribution is an evaluation design and a measurement result, not a classifier\.
## Limitations
Human validation covers polarity and relevance but not intensity, emotion, severity, or topic, and it reflects agreement with a single careful annotator rather than inter annotator reliability; no human to human agreement figure is available, so theκ\\kappavalues in Table[2](https://arxiv.org/html/2609.11144#S3.T2)have no measured ceiling\.
The rank comparisons are based on five instruments and are therefore descriptive\. Under conventional method\-specific sampling, human agreement aligns more closely with graded same\-day associations than with one\-day leads\. We therefore conclude only that benchmark agreement does not by itself determine predictive rankings and that the sampling convention and score representation must accompany such comparisons\. Widening the instrument battery is a natural next step\.
FinBERT’s low agreement is partly attributable to domain shift: it was tuned on Financial PhraseBank sentences rather than on tweets\. We regard this as consistent with, rather than an alternative to, our reading, since practitioners apply it off the shelf to exactly this kind of text; but it means thatκ=0\.079\\kappa=0\.079should be read as agreement in deployment rather than as an intrinsic property of the model\.
Account types are behavioural rather than profile derived\. Relevance filtered case level windows become sparse \(nnfalls from 134 to 50\)\. Reach weighted variants have been evaluated only for the LLM instrument\. Timing tests establish predictive precedence rather than structural causality or tradeability, and the local projection placebo fails at negative horizons\. Effect sizes throughout are small\.
The corpus is specific to securities\-class\-action discourse, where solicitation, repeated headlines, and cashtag piggybacking are unusually prevalent\. Although the evaluation design is portable, whether the observed relationships generalize to ordinary stock discussion requires cross\-domain validation\.
In addition, 74 of 273 archive tickers no longer resolve through the original price API\. If unresolved symbols disproportionately represent delisted firms, their absence from the refreshed price panel may truncate the adverse\-return tail and attenuate estimated negative associations\. The reported return relationships should therefore be interpreted subject to this potential survivorship bias\.
## Ethics Statement
The corpus consists of public X/Twitter posts collected via the platform API\. No author profile lookups were performed; account categories are derived solely from posting behaviour, and all attribution results are reported at the aggregate class level\. No personally identifying information is included in released artifacts\. Court filings are public records\. In accordance with the platform’s developer terms, the released artifact contains message identifiers and our annotations rather than message text, together with a script for retrieving text from the platform\.
## Data and Code Availability
The provided archive contains the classification prompts and configuration, the 400 message gold annotation file, message identifiers with all model and human labels, per instrument daily sentiment series, the computed abnormal return panel, and analysis code sufficient to reproduce every table\. The abnormal return panel is included directly because 74 of 273 archive tickers no longer resolve through the original price API\.
## Acknowledgements
We extend our gratitude to Perssonify LLC\. for their continued support of our work, and allocating resources that made this work possible\. In particular we thank Stefan Persson for his valuable inputs and guidance\.
## References
- Antweiler and Frank \(2004\)W\. Antweiler and M\. Z\. FrankIs all that talk just noise? The information content of internet stock message boards\.The Journal of Finance59\(3\),pp\. 1259–1294\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Araci \(2019\)D\. AraciFinBERT: financial sentiment analysis with pre\-trained language models\.arXiv preprint arXiv:1908\.10063\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2609.11144#S3.SS4.p1.1)\.
- Barber and Odean \(2008\)B\. M\. Barber and T\. OdeanAll that glitters: the effect of attention and news on the buying behavior of individual and institutional investors\.Review of Financial Studies21\(2\),pp\. 785–818\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Barbieriet al\.\(2020\)F\. Barbieri, J\. Camacho\-Collados, L\. Espinosa Anke, and L\. NevesTweetEval: unified benchmark and comparative evaluation for tweet classification\.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp\. 1644–1650\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2609.11144#S3.SS4.p1.1)\.
- Benjamini and Hochberg \(1995\)Y\. Benjamini and Y\. HochbergControlling the false discovery rate: a practical and powerful approach to multiple testing\.Journal of the Royal Statistical Society: Series B57\(1\),pp\. 289–300\.Cited by:[§4](https://arxiv.org/html/2609.11144#S4.p1.1)\.
- Bollenet al\.\(2011\)J\. Bollen, H\. Mao, and X\. ZengTwitter mood predicts the stock market\.Journal of Computational Science2\(1\),pp\. 1–8\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Chenet al\.\(2014\)H\. Chen, P\. De, Y\. \(\. Hu, and B\. HwangWisdom of crowds: the value of stock opinions transmitted through social media\.Review of Financial Studies27\(5\),pp\. 1367–1403\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Chiuet al\.\(2016\)B\. Chiu, A\. Korhonen, and S\. PyysaloIntrinsic evaluation of word vectors fails to predict extrinsic performance\.InProceedings of the 1st Workshop on Evaluating Vector\-Space Representations for NLP \(RepEval\),pp\. 1–6\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px3.p1.1)\.
- Choiet al\.\(2024\)S\. J\. Choi, J\. Erickson, and A\. C\. PritchardThe business of securities class action lawyering\.Indiana Law Journal99\(3\)\.Cited by:[§1](https://arxiv.org/html/2609.11144#S1.p3.1)\.
- Cohen \(1960\)J\. CohenA coefficient of agreement for nominal scales\.Educational and Psychological Measurement20\(1\),pp\. 37–46\.Cited by:[§3\.5](https://arxiv.org/html/2609.11144#S3.SS5.p2.1)\.
- Cookson and Niessner \(2020\)J\. A\. Cookson and M\. NiessnerWhy don’t we agree? Evidence from a social network of investors\.The Journal of Finance75\(1\),pp\. 173–228\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Cresciet al\.\(2019\)S\. Cresci, F\. Lillo, D\. Regoli, S\. Tardelli, and M\. TesconiCashtag piggybacking: uncovering spam and bot activity in stock microblogs on Twitter\.ACM Transactions on the Web13\(2\),pp\. 1–27\.External Links:[Document](https://dx.doi.org/10.1145/3313184)Cited by:[§1](https://arxiv.org/html/2609.11144#S1.p3.1)\.
- Cronbach and Meehl \(1955\)L\. J\. Cronbach and P\. E\. MeehlConstruct validity in psychological tests\.Psychological Bulletin52\(4\),pp\. 281–302\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px3.p1.1)\.
- Daet al\.\(2011\)Z\. Da, J\. Engelberg, and P\. GaoIn search of attention\.The Journal of Finance66\(5\),pp\. 1461–1499\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Das and Chen \(2007\)S\. R\. Das and M\. Y\. ChenYahoo\! for Amazon: sentiment extraction from small talk on the web\.Management Science53\(9\),pp\. 1375–1388\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Dycket al\.\(2010\)A\. Dyck, A\. Morse, and L\. ZingalesWho blows the whistle on corporate fraud?\.The Journal of Finance65\(6\),pp\. 2213–2253\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Efron and Tibshirani \(1993\)B\. Efron and R\. J\. TibshiraniAn introduction to the bootstrap\.Chapman & Hall\.Cited by:[Table 2](https://arxiv.org/html/2609.11144#S3.T2)\.
- Gande and Lewis \(2009\)A\. Gande and C\. M\. LewisShareholder\-initiated class action lawsuits: shareholder wealth effects and industry spillovers\.Journal of Financial and Quantitative Analysis44\(4\),pp\. 823–850\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Gilardiet al\.\(2023\)F\. Gilardi, M\. Alizadeh, and M\. KubliChatGPT outperforms crowd workers for text\-annotation tasks\.Proceedings of the National Academy of Sciences120\(30\),pp\. e2305016120\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px4.p1.1)\.
- Glasserman and Lin \(2023\)P\. Glasserman and C\. LinAssessing look\-ahead bias in stock return predictions generated by large language models\.arXiv preprint arXiv:2309\.17322\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px4.p1.1),[§9](https://arxiv.org/html/2609.11144#S9.SS0.SSS0.Px1.p1.1)\.
- Granger \(1969\)C\. W\. J\. GrangerInvestigating causal relations by econometric models and cross\-spectral methods\.Econometrica37\(3\),pp\. 424–438\.Cited by:[§4](https://arxiv.org/html/2609.11144#S4.p2.1)\.
- Heet al\.\(2025\)S\. He, L\. Lv, A\. Manela, and J\. WuChronologically consistent large language models\.arXiv preprint arXiv:2502\.21206\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px4.p1.1)\.
- Hutto and Gilbert \(2014\)C\. J\. Hutto and E\. GilbertVADER: a parsimonious rule\-based model for sentiment analysis of social media text\.InProceedings of the Eighth International AAAI Conference on Weblogs and Social Media,pp\. 216–225\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2609.11144#S3.SS4.p1.1)\.
- Jacobs and Wallach \(2021\)A\. Z\. Jacobs and H\. WallachMeasurement and fairness\.InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency,pp\. 375–385\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px3.p1.1)\.
- Jordà \(2005\)Ò\. JordàEstimation and inference of impulse responses by local projections\.American Economic Review95\(1\),pp\. 161–182\.Cited by:[§4](https://arxiv.org/html/2609.11144#S4.p2.2)\.
- Karpoffet al\.\(2008\)J\. M\. Karpoff, D\. S\. Lee, and G\. S\. MartinThe cost to firms of cooking the books\.Journal of Financial and Quantitative Analysis43\(3\),pp\. 581–611\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Kim and Skinner \(2012\)I\. Kim and D\. J\. SkinnerMeasuring securities litigation risk\.Journal of Accounting and Economics53\(1–2\),pp\. 290–310\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Kirtac and Germano \(2024\)K\. Kirtac and G\. GermanoSentiment trading with large language models\.arXiv preprint arXiv:2412\.19245\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px4.p1.1)\.
- Landis and Koch \(1977\)J\. R\. Landis and G\. G\. KochThe measurement of observer agreement for categorical data\.Biometrics33\(1\),pp\. 159–174\.Cited by:[§3\.5](https://arxiv.org/html/2609.11144#S3.SS5.p2.1)\.
- Lopez\-Lira and Tang \(2023\)A\. Lopez\-Lira and Y\. TangCan ChatGPT forecast stock price movements? Return predictability and large language models\.arXiv preprint arXiv:2304\.07619\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px4.p1.1)\.
- Loughran and McDonald \(2011\)T\. Loughran and B\. McDonaldWhen is a liability not a liability? Textual analysis, dictionaries, and 10\-Ks\.The Journal of Finance66\(1\),pp\. 35–65\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2609.11144#S3.SS4.p1.1)\.
- Loughran and McDonald \(2016\)T\. Loughran and B\. McDonaldTextual analysis in accounting and finance: a survey\.Journal of Accounting Research54\(4\),pp\. 1187–1230\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1)\.
- Maloet al\.\(2014\)P\. Malo, A\. Sinha, P\. Korhonen, J\. Wallenius, and P\. TakalaGood debt or bad debt: detecting semantic orientations in economic texts\.Journal of the Association for Information Science and Technology65\(4\),pp\. 782–796\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1)\.
- Miller \(2006\)G\. S\. MillerThe press as a watchdog for accounting fraud\.Journal of Accounting Research44\(5\),pp\. 1001–1033\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Petersen \(2009\)M\. A\. PetersenEstimating standard errors in finance panel data sets: comparing approaches\.Review of Financial Studies22\(1\),pp\. 435–480\.Cited by:[§4](https://arxiv.org/html/2609.11144#S4.p2.2)\.
- Rancoet al\.\(2015\)G\. Ranco, D\. Aleksovski, G\. Caldarelli, M\. Grčar, and I\. MozetičThe effects of Twitter sentiment on stock price returns\.PLOS ONE10\(9\),pp\. e0138441\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Sprengeret al\.\(2014\)T\. O\. Sprenger, A\. Tumasjan, P\. G\. Sandner, and I\. M\. WelpeTweets and trades: the information content of stock microblogs\.European Financial Management20\(5\),pp\. 926–957\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px5.p1.1)\.
- Tetlocket al\.\(2008\)P\. C\. Tetlock, M\. Saar\-Tsechansky, and S\. MacskassyMore than words: quantifying language to measure firms’ fundamentals\.The Journal of Finance63\(3\),pp\. 1437–1467\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Tetlock \(2007\)P\. C\. TetlockGiving content to investor sentiment: the role of media in the stock market\.The Journal of Finance62\(3\),pp\. 1139–1168\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Tetlock \(2011\)P\. C\. TetlockAll the news that’s fit to reprint: do investors react to stale information?\.The Review of Financial Studies24\(5\),pp\. 1481–1512\.External Links:[Document](https://dx.doi.org/10.1093/rfs/hhq141)Cited by:[§1](https://arxiv.org/html/2609.11144#S1.p3.1)\.
- Tumarkin and Whitelaw \(2001\)R\. Tumarkin and R\. F\. WhitelawNews or noise? Internet postings and stock prices\.Financial Analysts Journal57\(3\),pp\. 41–51\.Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px1.p1.1)\.
- Xieet al\.\(2024\)Q\. Xie, W\. Han, Z\. Chen, R\. Xiang, X\. Zhang, S\. Ananiadou, J\. Huang,et al\.FinBen: a holistic financial benchmark for large language models\.InAdvances in Neural Information Processing Systems 37 \(Datasets and Benchmarks Track\),Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1)\.
- Xieet al\.\(2023\)Q\. Xie, W\. Han, X\. Zhang, Y\. Lai, M\. Peng, A\. Lopez\-Lira, and J\. HuangPIXIU: a comprehensive benchmark, instruction dataset and large language model for finance\.InAdvances in Neural Information Processing Systems 36 \(Datasets and Benchmarks Track\),External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/6a386d703b50f1cf1f61ab02a15967bb-Abstract-Datasets_and_Benchmarks.html)Cited by:[§2](https://arxiv.org/html/2609.11144#S2.SS0.SSS0.Px2.p1.1)\.
## Appendix ARelevance Filtering
Table[8](https://arxiv.org/html/2609.11144#A1.T8)repeats the lead–lag analysis of Table[4](https://arxiv.org/html/2609.11144#S6.T4)on messages the LLM instrument labels*relevant*, and Table[9](https://arxiv.org/html/2609.11144#A1.T9)reports the change against the all\-message baseline\. Filtering shifts the same\-day correlation by at most0\.01640\.0164and never reverses a sign, so the central result is not produced by the relevance labels\. This matters because those labels come from one of the instruments under evaluation; if the result depended on them the comparison would be circular\. At case level, filtering reduces the common sample from 134 to 50 because windows containing no relevant message become undefined, and every instrument becomes nonsignificant\.
Table 8:Lead–lag correlations restricted to relevant messages\. LM denotes Loughran–McDonald; RoBERTa denotes Twitter\-RoBERTa\.q1q\_\{1\}is FDR\-adjusted across the baseline family; the LLM reference is not included in that correction \(—\)\. Significance as in Section[4](https://arxiv.org/html/2609.11144#S4)\.Table 9:Effect of relevance filtering on the same\-day correlation\.Δ\\Deltais relevant minus all\.Table 10:Forward\-Granger and distributed\-lag results\. Boldqqmarks the four specifications that survive FDR correction\. All\-message rows usenF=103,184n\_\{F\}=103\{,\}184andnβ=103,005n\_\{\\beta\}=103\{,\}005; relevant\-message rows use99,32899\{,\}328and99,15799\{,\}157\. Coefficients are not comparable across instruments \(see text\)\. Significance as in Section[4](https://arxiv.org/html/2609.11144#S4)\.
## Appendix BForward\-Granger and Distributed\-Lag Tests
Table[10](https://arxiv.org/html/2609.11144#A1.T10)gives the complete battery summarised in Section[6](https://arxiv.org/html/2609.11144#S6)\. Distributed\-lag coefficients are not comparable across instruments because the underlying signals differ in scale; only sign and significance transfer\.
## Appendix CCommon Day Intersection, Relevant Messages
Table[6](https://arxiv.org/html/2609.11144#S7.T6)in the main text holdsnnfixed across instruments using all messages\. Table[11](https://arxiv.org/html/2609.11144#A3.T11)repeats that intersection on relevant messages only\. The reordering described in Section[7\.1](https://arxiv.org/html/2609.11144#S7.SS1)persists: FinBERT is strongest same day on the coarse panel and Twitter\-RoBERTa on the graded panel, while VADER remains indistinguishable from zero throughout\.
Table 11:Common case day intersection on relevant messages\. C and G denote coarse and graded;n=1,328n=1\{,\}328and3,2953\{,\}295respectively\.
## Appendix DLook Ahead Contamination
Section[9](https://arxiv.org/html/2609.11144#S9)reports the headline contamination test\. This appendix gives the decomposition\. Table[12](https://arxiv.org/html/2609.11144#A4.T12)reports the difference in differences against each baseline separately rather than against their mean, so the null is not an artifact of averaging: all sixteen estimates lie between−0\.0400\-0\.0400and\+0\.0281\+0\.0281and every interval spans zero\.
Within the archive cohort, splitting at the median case year gives an LLM advantage of\+0\.0395\+0\.0395for older cases against\+0\.0363\+0\.0363for newer ones on the coarse same day measure, and\+0\.0153\+0\.0153against\+0\.0127\+0\.0127at a one day lag\. The advantage is slightly larger for older cases, but the differences are small\. This comparison alone does not establish contamination\.
We are more cautious about the case level measure\. Splitting that test by cohort leaves 80 archive and 47 recent cases, and on a reconstruction of the panel the LLM’s advantage is larger in the archive than in the recent cohorts, which is the direction contamination predicts\. The subsample is small and we do not treat this as evidence of contamination, but the daily panel null does not extend to it, which is why the paper rests its claims on the daily panel\.
Table 12:Per baseline difference in differences\. Each cell is the LLM’s\|ρ\|\|\\rho\|advantage over that baseline in the archive cohort minus the same advantage in the 2024–2025 cohorts\. No estimate is distinguishable from zero at 95% over 1,000 case level bootstrap resamples\.
## Appendix ERetained Null Results
For completeness: the fear/panic/outrage*share*is insignificant at every lag, while its daily count is not; several event window specifications lose significance once firm and sector controls are added; relevance filtering does not create the daily result; the filtered case level measure becomes sparse and nonsignificant; negative horizon local projections limit causal interpretation and total, relevant, and negative message volume all fail to predict settlement size or case level price damage depth\.
## Appendix FEvent Validation and Local Projections
Mean negative message share rises from 11\.8% in the baseline window to 46\.7% in the disclosure window \(n=184n=184, pairedt=10\.97t=10\.97,p<10−21p<10^\{\-21\}\), confirming that the complaint derived disclosure dates coincide with an information shock rather than background chatter\. With controls for log market capitalisation, class period length, defendant count, and sector, class period negative share predicts maximum drawdown \(n=41n=41,t=−3\.24t=\-3\.24,p=0\.003p=0\.003, adjustedR2=0\.60R^\{2\}=0\.60\)\.
Local projections \(Figure[2](https://arxiv.org/html/2609.11144#A6.F2), bottom\) peak ath=0h=0\(β=−0\.0079\\beta=\-0\.0079,p<0\.001p<0\.001\), remain significant ath=1h=1\(β=−0\.0022\\beta=\-0\.0022,p<0\.05p<0\.05\), and are indistinguishable from zero byh=2h=2\. Coefficients at negative horizons are also significantly negative, and a reverse next\-day correlation is present \(ρ=−0\.046\\rho=\-0\.046,p<0\.01p<0\.01\)\.


Figure 2:Top: the complaint derived disclosure window is a sharp negativity shock\. Bottom: local projection coefficients peak contemporaneously and decay after one day; the significantly negative coefficients at negative horizons are a failed placebo and caution against a causal interpretation\.
## Appendix GRank Correlations Between the Two Validities
Figure 3:Target Corporation over its class period\. Top: daily closing price\. Bottom: messages per day, coloured by the share of negative messages\. Negative\-heavy days concentrate around the major price declines\.Table[13](https://arxiv.org/html/2609.11144#A7.T13)gives every correlation between a construct validity measure and a predictive validity column, completing Table[5](https://arxiv.org/html/2609.11144#S7.T5)\. With five instruments the asymptoticpp\-values reported by standard software are unreliable, so we compute exact two sidedpp\-values by enumerating all5\!=1205\!=120permutations\.
Table 13:Spearman rank correlation between each construct validity measure and each predictive validity column, over the five instruments\. Exact two sidedpp\-values from all 120 permutations in parentheses\. Accuracy and macro\-F1 give identical values because they order the five instruments identically\. Only the graded same day column is nominally significant, and no cell survives adjustment for twelve comparisons\.Across all three agreement measures, the closest alignment is with graded same\-day associations\. At one day, the relationship is weaker for graded scores and weakly negative for coarse scores\. None of the twelve comparisons survives multiple\-testing correction\. With only five instruments, we treat these patterns as descriptive\.
## Appendix HCase Study: Target Corporation
A large cap example shows the mechanism: Figure[3](https://arxiv.org/html/2609.11144#A7.F3)overlays Target Corporation’s stock price with its daily message volume coloured by negative share over the class period\. Days on which more than half the messages are negative cluster around the two largest price declines, while low negativity days dominate the calm stretches and carry little price information\. Message*volume*is similar in negative and quiet periods; it is the classified sentiment that lines up with the moves\. The case illustrates at the single case level the domain result of Section[5](https://arxiv.org/html/2609.11144#S5)\.Similar Articles
FinSMART: Financial Sentiment Analysis for Algorithmic Trading through Market-Aligned Reinforcement Learning
FinSMART introduces a market-aligned reinforcement learning framework for financial sentiment analysis, optimizing sentiment signals with realized market outcomes and achieving a 220% improvement in cumulative trading returns over the strongest baseline.
Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
This paper audits temporal leakage in financial news NLP benchmarks across multiple models, finding that random splits inflate performance metrics and identifies M&A events as a category with a localized positive signal under chronological evaluation.
Enhancing Financial Sentiment Analysis via Retrieval Augmented Large Language Models
This paper introduces a retrieval-augmented LLM framework for financial sentiment analysis, achieving 15-48% improvement in accuracy and F1 score over traditional models and LLMs like ChatGPT and LLaMA.
A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis
This paper presents a unified multi-modal framework integrating reinforcement learning, high-frequency trading, game-theoretic approaches, and cross-modal sentiment analysis for intelligent financial systems, claiming significant improvements over single-domain systems.
Bridging the Gap Between Natural Language and Market Dynamics via High-Dimensional Representation Learning
This paper explores replacing scalar sentiment scores with high-dimensional FinBERT embeddings in a Transformer-based architecture for short-term stock price prediction, showing improved accuracy with Siamese-optimized embeddings.