Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse

arXiv cs.CL Papers

Summary

This paper proposes a diagnostic framework to separate preprocessing pipeline instability from measurement method instability in LLM-based stance analysis of public discourse, finding that cross-method disagreement is larger and more systematic than pipeline effects, and that aggregate metrics can mask these instabilities.

arXiv:2607.10846v1 Announce Type: new Abstract: Computational social science increasingly relies on automated preprocessing pipelines -- speaker diarization, ASR transcript cleaning, sentence segmentation -- to convert raw media into analyzable text. When these pipelines produce different outputs from the same input, two distinct sources of instability can arise: the preprocessing pipeline itself (diarization method, segmentation rules) and the downstream measurement instrument (LLM annotation vs.\ keyword lexicon). Using 256 YouTube interviews across 41 public figures from five domains, we compare two speaker-diarization pipelines and two measurement methods, all targeting the coupling between affective valence and epistemic modality. We find that (1) preprocessing pipeline sensitivity is concentrated in speakers with limited video samples (N $\leq 5$); for the four best-sampled speakers (N $\geq 16$), the mean absolute pipeline-induced change in $r(\text{neg}, \text{emph})$ is only $0.13$; (2) cross-method disagreement is larger and more systematic -- the LLM and keyword-lexicon methods assign opposite coupling directions to several well-sampled speakers, even within the same preprocessing pipeline; and (3) aggregate valence proportions are highly stable ($|\Delta p(\text{neg})| < 6$pp) regardless of pipeline or method, masking both sources of instability. The contribution is a diagnostic framework that separates pipeline effects from measurement effects: researchers studying cross-dimensional relationships in interview data should verify that their conclusions are robust to both sources of variation, with particular attention to measurement method choice.
Original Article
View Cached Full Text

Cached at: 07/14/26, 04:23 AM

# Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse
Source: [https://arxiv.org/html/2607.10846](https://arxiv.org/html/2607.10846)
###### Abstract

Computational social science increasingly relies on automated preprocessing pipelines—speaker diarization, ASR transcript cleaning, sentence segmentation—to convert raw media into analyzable text\. When these pipelines produce different outputs from the same input, two distinct sources of instability can arise: the preprocessing pipeline itself \(diarization method, segmentation rules\) and the downstream measurement instrument \(LLM annotation vs\. keyword lexicon\)\. Using 256 YouTube interviews across 41 public figures from five domains, we compare two speaker\-diarization pipelines and two measurement methods, all targeting the coupling between affective valence and epistemic modality\. We find that \(1\) preprocessing pipeline sensitivity is concentrated in speakers with limited video samples \(N≤5\\leq 5\); for the four best\-sampled speakers \(N≥16\\geq 16\), the mean absolute pipeline\-induced change inr​\(neg,emph\)r\(\\text\{neg\},\\text\{emph\}\)is only0\.130\.13; \(2\) cross\-method disagreement is larger and more systematic—the LLM and keyword\-lexicon methods assign opposite coupling directions to several well\-sampled speakers, even within the same preprocessing pipeline; and \(3\) aggregate valence proportions are highly stable \(\|Δ​p​\(neg\)\|<6\|\\Delta p\(\\text\{neg\}\)\|<6pp\) regardless of pipeline or method, masking both sources of instability\. The contribution is a diagnostic framework that separates pipeline effects from measurement effects: researchers studying cross\-dimensional relationships in interview data should verify that their conclusions are robust to both sources of variation, with particular attention to measurement method choice\.

Keywords:pipeline sensitivity, preprocessing robustness, speaker diarization, LLM annotation, keyword lexicon, affective valence, epistemic modality, computational social science

## 1Introduction

The computational social science \(CSS\) pipeline is increasingly automated\. Raw interview footage passes through automatic speech recognition \(ASR\), speaker diarization, transcript cleaning, and sentence segmentation before any analysis begins\. Each of these preprocessing steps embeds choices—which ASR engine, which diarization method, which sentence boundary detector—and each choice may leave a fingerprint on the resulting data\.

The CSS literature has devoted substantial attention toannotationvalidity: whether LLM\-generated labels match human judgments\[[4](https://arxiv.org/html/2607.10846#bib.bib4),[5](https://arxiv.org/html/2607.10846#bib.bib5)\]\. It has devoted far less attention to two upstream sources of instability:preprocessingvalidity—whether pipeline choices change downstream conclusions—andmeasurementvalidity—whether the choice of annotation method \(LLM vs\. keyword lexicon\) leads to different conclusions even when applied to the same preprocessed text\. These two sources are often conflated, making it difficult to diagnose whether an observed instability originates in the pipeline or in the measurement instrument\.

We proposepipeline sensitivity analysisas a systematic approach to this problem:

> Pipeline sensitivityis the degree to which a derived metric changes under alternative, equally plausible preprocessing configurations, measured per\-unit \(per\-speaker, per\-video\) and aggregated across units to identify conditions under which preprocessing choice is consequential\.

Three features distinguish pipeline sensitivity from classical measurement error\. First, it ismetric\-dependent: aggregate means may be stable while cross\-dimensional correlations reverse sign\. Second, it isnon\-uniform: some speakers or domains may be far more sensitive than others\. Third, it ismethod\-interactive: the same preprocessing divergence may affect LLM\-based and lexicon\-based measurements differently\.

We demonstrate pipeline sensitivity analysis through a case study of speaker diarization\. Using a corpus of 256 YouTube interviews across 41 public figures \(spanning finance, academia, central banking, geopolitics, and media\), we compare:

1. 1\.Two preprocessing pipelines:VTT\-based diarization \(YouTube captions \+ LLM speaker classification\) vs\. AssemblyAI audio\-based diarization\.
2. 2\.Two measurement methods:LLM zero\-shot annotation \(DeepSeek\) vs\. keyword\-lexicon scoring, both applied to the same valence\-modality dimensions\.

We then ask:how much does the preprocessing choice affect the downstream conclusion, and does the answer depend on the measurement method?

Our contributions are:

1. 1\.A two\-source diagnostic frameworkthat separates preprocessing pipeline effects from measurement method effects, with practical recommendations for reporting both sources of instability \(Section[3](https://arxiv.org/html/2607.10846#S3),[6](https://arxiv.org/html/2607.10846#S6)\)\.
2. 2\.Empirical evidence that pipeline sensitivity is bounded and predictable\.For the four best\-sampled speakers \(16–34 videos\), mean\|Δ​r\|=0\.13\|\\Delta r\|=0\.13under LLM annotation; largerΔ​r\\Delta rvalues are concentrated in speakers with≤5\\leq 5videos where correlational estimates are inherently unstable \(Section[5](https://arxiv.org/html/2607.10846#S5)\)\.
3. 3\.Empirical evidence that cross\-method disagreement is the more serious validity threat\.LLM and keyword\-lexicon methods disagree on the sign ofr​\(neg,emph\)r\(\\text\{neg\},\\text\{emph\}\)in49%49\\%of cases, including well\-sampled speakers \(Rogoff, 17 videos\) where both methods are internally stable but give opposite coupling directions, indicating that measurement choice can systematically alter conclusions even when preprocessing is held constant \(Section[5\.3](https://arxiv.org/html/2607.10846#S5.SS3)\)\.

## 2Related Work

### 2\.1Speaker Diarization in CSS

Text\-only speaker diarization uses linguistic content rather than acoustic features to distinguish speakers\. Recent work\[[1](https://arxiv.org/html/2607.10846#bib.bib1),[2](https://arxiv.org/html/2607.10846#bib.bib2),[3](https://arxiv.org/html/2607.10846#bib.bib3)\]has reported 0–4% diarization error rates for LLM\-based methods on short, structured conversations\. However, these validations share a common context: regular turn\-taking, lexically distinct speaker roles, and controlled recording conditions\. Their generalizability to the long\-form, unscripted public discourse typical of CSS research remains untested\.

### 2\.2LLM\-Based Corpus Annotation and Validation

LLMs are increasingly used for zero\-shot annotation in CSS\[[4](https://arxiv.org/html/2607.10846#bib.bib4),[6](https://arxiv.org/html/2607.10846#bib.bib6)\]\. Ziems et al\.\[[5](https://arxiv.org/html/2607.10846#bib.bib5)\]identified prompt sensitivity and provider\-specific biases as key concerns\. Our work extends the validation imperative upstream: we hold annotation quality constant and ask whether preprocessing choices alone can change conclusions\.

### 2\.3Preprocessing Robustness and Measurement Validity

Preprocessing sensitivity has been explored in NLP for specific tasks—machine translation evaluation\[[9](https://arxiv.org/html/2607.10846#bib.bib9)\], reproducibility\[[10](https://arxiv.org/html/2607.10846#bib.bib10)\]—but has not been systematically linked to downstream statistical inference in CSS\. Recent work has begun to examine how LLM\-based annotation is affected by input formatting choices: Zhao et al\.\[[15](https://arxiv.org/html/2607.10846#bib.bib15)\]show that varying context lengths and text segmentation strategies can shift LLM sentiment labels by 5–15 percentage points on the same underlying text, and Liu et al\.\[[16](https://arxiv.org/html/2607.10846#bib.bib16)\]demonstrate that sentence\-level versus paragraph\-level chunking produces systematic differences in fine\-grained emotion classification\. These findings suggest that the same preprocessing pipeline that introduces sentence\-boundary artifacts \(Section[4](https://arxiv.org/html/2607.10846#S4)\) may also amplify measurement\-level biases\. Classical measurement theory\[[7](https://arxiv.org/html/2607.10846#bib.bib7),[8](https://arxiv.org/html/2607.10846#bib.bib8)\]provides tools for understanding noise and bias, but its standard model of random error does not capture the structured, metric\-dependent contamination that preprocessing pipelines can introduce\[[14](https://arxiv.org/html/2607.10846#bib.bib14),[13](https://arxiv.org/html/2607.10846#bib.bib13)\]\.

## 3A Framework for Diagnosing Sources of Instability

### 3\.1Two Sources of Instability

LetDDbe a raw dataset ofNNunits \(speakers, each withkik\_\{i\}interview videos\)\. LetP1P\_\{1\}andP2P\_\{2\}be two preprocessing pipelines—alternative ways of extracting analyzable text fromDD\. Letmmbe a downstream metric \(e\.g\.,r​\(neg,emph\)r\(\\text\{neg\},\\text\{emph\}\)across videos\)\.

We distinguish two sources of instability:

#### Source 1: Preprocessing pipeline divergence\.

For each unitii, the pipeline delta measures how much the metric changes when the preprocessing pipeline changes:

Δipipeline=m​\(P2​\(Di\)\)−m​\(P1​\(Di\)\)\\Delta\_\{i\}^\{\\text\{pipeline\}\}=m\(P\_\{2\}\(D\_\{i\}\)\)\-m\(P\_\{1\}\(D\_\{i\}\)\)\(1\)
We report the per\-unit absolute delta\|Δipipeline\|\|\\Delta\_\{i\}^\{\\text\{pipeline\}\}\|and the proportion of units where the sign ofmmreverses \(m​\(P1\)⋅m​\(P2\)<0m\(P\_\{1\}\)\\cdot m\(P\_\{2\}\)<0\)\. Whenmmis a Pearson correlationrr, we apply Fisher’szz\-transformation \(z=12​ln⁡1\+r1−rz=\\frac\{1\}\{2\}\\ln\\frac\{1\+r\}\{1\-r\}\) for inferential purposes; mean\|Δ​r\|\|\\Delta r\|values are reported inrr\-space for interpretability\.

#### Source 2: Measurement method divergence\.

LetmLLMm^\{\\text\{LLM\}\}andmKWm^\{\\text\{KW\}\}denote the same metric computed by different measurement instruments applied to the same preprocessed textP​\(Di\)P\(D\_\{i\}\)\. The cross\-method divergence is the absolute difference between the methods’ estimates:

Δimethod​\(P\)=\|mLLM​\(P​\(Di\)\)−mKW​\(P​\(Di\)\)\|\\Delta\_\{i\}^\{\\text\{method\}\}\(P\)=\|m^\{\\text\{LLM\}\}\(P\(D\_\{i\}\)\)\-m^\{\\text\{KW\}\}\(P\(D\_\{i\}\)\)\|\(2\)
Values near zero indicate that the two methods give equivalent estimates on the same preprocessed input; large values indicate measurement\-driven disagreement\. Unlike Equation[1](https://arxiv.org/html/2607.10846#S3.E1), which measures sensitivity to preprocessingwithina single method, Equation[2](https://arxiv.org/html/2607.10846#S3.E2)measures sensitivity to measurementwithina single pipeline\. The two deltas can be compared directly: ifΔimethod\>Δipipeline\\Delta\_\{i\}^\{\\text\{method\}\}\>\\Delta\_\{i\}^\{\\text\{pipeline\}\}for a given speaker, measurement choice dominates preprocessing choice as a source of instability for that speaker\.

### 3\.2Decomposing Observed Instability

For any metricmmcomputed from raw dataDD, the total observed variation can arise from either source\. To diagnose which source dominates:

1. 1\.ComputeΔipipeline\\Delta\_\{i\}^\{\\text\{pipeline\}\}separately within each measurement method \(LLM and keyword\)\.
2. 2\.ComputeΔimethod\\Delta\_\{i\}^\{\\text\{method\}\}separately within each preprocessing pipeline \(VTT and AssemblyAI\)\.
3. 3\.Compare the magnitudes: if\|Δipipeline\|≪\|Δimethod\|\|\\Delta\_\{i\}^\{\\text\{pipeline\}\}\|\\ll\|\\Delta\_\{i\}^\{\\text\{method\}\}\|on average, measurement choice is the dominant source of instability\. If the reverse, preprocessing is dominant\.

We apply this diagnostic decomposition in Sections[5](https://arxiv.org/html/2607.10846#S5)–[5\.4](https://arxiv.org/html/2607.10846#S5.SS4)\.

## 4Data and Preprocessing

### 4\.1Corpus

Our empirical testbed consists of 256 YouTube interview videos across 41 public figures from five domains\. Table[1](https://arxiv.org/html/2607.10846#S4.T1)provides the per\-domain, per\-speaker breakdown\. Videos were collected via yt\-dlp; metadata was extracted from YouTube’s JSON output\. A metadata\-based quality filter excluded 9 videos as non\-interview content \(documentaries, third\-person narration, non\-English\)\. Ten additional speakers with fewer than 4 videos were excluded from correlational analysis to avoid small\-sample spurious correlations \(Section[5\.1](https://arxiv.org/html/2607.10846#S5.SS1)\)\. The corpus is used here as a testbed for pipeline sensitivity, not introduced as a dataset contribution\.

Table 1:Corpus composition by domain and speaker \(restricted to speakers with≥4\\geq 4videos\)\.
### 4\.2Preprocessing Pipelines

#### VTT\-based pipeline\.

Raw YouTube VTT subtitle files are parsed to extract timed text segments\. Roll\-up caption artifacts are deduplicated\. Consecutive segments with a gap≥1\.5\\geq 1\.5s form separate utterance turns\. Each turn is submitted to DeepSeek\-V4\-Flash for binary “guest”/“interviewer” classification \(batched in groups of 20\)\. A content\-based pre\-check flags videos with\>\>500 words but zero first\-person speech markers as likely non\-interview content\. Guest text is concatenated into the final per\-video output\.

#### Audio\-based pipeline \(AssemblyAI\)\.

111We selected AssemblyAI \(a commercial cloud API\) over open\-source alternatives such as pyannote\-audio\[[12](https://arxiv.org/html/2607.10846#bib.bib12)\]\+ Whisper for three reasons: \(1\) AssemblyAI provides integrated ASR \+ diarization with speaker labels, avoiding the engineering complexity of stitching separate ASR and diarization outputs; \(2\) preliminary tests on our long\-form interview data showed that AssemblyAI’s diarization quality was substantially more reliable than pyannote’s for multi\-speaker, unscripted content with overlapping speech; and \(3\) the API\-based workflow is reproducible without local GPU resources, making the pipeline accessible to CSS researchers without specialized hardware\. We acknowledge that commercial APIs introduce cost and reproducibility constraints; extending the comparison to open\-source diarizers is discussed under Future Work at the end of the Conclusion\.Audio files \(\.m4a\) are uploaded to AssemblyAI’s API with speaker diarization enabled\. The dominant speaker by word count is identified as the guest \(LLM fallback for ambiguous cases where the second speaker exceeds 25% word share\)\. Guest utterances are concatenated into the final per\-video output\.

Both pipelines produce raw ASR text, differing only in the speaker identification mechanism\. The two pipelines yield systematically different outputs: AssemblyAI detects multi\-speaker structure with guest ratios varying across videos, while the VTT\-based pipeline produces higher and less variable guest ratios, consistent with under\-removal of interviewer speech\. No LLM\-based text cleaning is applied to either output\.

### 4\.3Sentence Segmentation

Guest text from each pipeline is split into sentences using regex boundary detection\. Sentences outside 2–60 words are excluded from downstream annotation\. Table[2](https://arxiv.org/html/2607.10846#S4.T2)reports the resulting sentence corpus statistics\. The two pipelines produce markedly different sentence distributions: VTT\-based output is more fragmented \(mean 10\.2 words/sentence\), reflecting YouTube’s auto\-generated caption segmentation lacking sentence\-final punctuation, while AssemblyAI output forms longer, more coherent sentences \(mean 15\.3 words/sentence\)\. This difference means that the observed pipeline sensitivity is a composite of diarization and segmentation effects; the current design cannot fully isolate them\.

Table 2:Sentence corpus statistics after segmentation and filtering \(2–60 words\), restricted to the 41 speakers with≥4\\geq 4videos\.
### 4\.4LLM Annotation

All sentences from both pipelines are annotated via DeepSeek\-V4\-Flash \(temperature 0\.0\) on valence \(positive/negative/neutral\) and epistemic modality \(emphatic/hedged/neutral\), batched in groups of 50, following the LLM\-as\-annotator paradigm established by\[[4](https://arxiv.org/html/2607.10846#bib.bib4),[5](https://arxiv.org/html/2607.10846#bib.bib5)\]and the dual\-dimension valence\-modality framework of\[[11](https://arxiv.org/html/2607.10846#bib.bib11)\]\. The annotation prompt defines both dimensions with examples and instructs the model to prefer “neutral” when uncertain\. Annotation is per\-video with batch\-level checkpointing for resumption\. All 41 speakers are fully annotated for both pipelines \(108,374 VTT sentences; 82,254 AssemblyAI sentences\)\.

### 4\.5Keyword\-Lexicon Scoring

As an independent measurement method that shares the same video\-level analysis unit as the LLM, we replicate the keyword\-lexicon approach from\[[11](https://arxiv.org/html/2607.10846#bib.bib11)\], which was developed for a parallel study of valence\-modality coupling in financial and political discourse\. The lexicon contains patterns across three categories—Negative Valence \(Fear/Anxiety, Anger/Conflict, Decline/Destruction, Crisis/Urgency, Financial/Political Risk; 97 patterns\), Positive Valence \(Hope/Optimism, Strength/Success, Stability/Order; 33 patterns\), and Modality \(Emphatic/Certain, Hedging/Doubt, Intensity Boosters; 67 patterns\)—with negation handling \(e\.g\., “not strong” is reclassified as Decline/Destruction\)\. Per\-video keyword frequencies are normalized by video character length to a baseline ofNbase=30,000N\_\{\\text\{base\}\}=30\{,\}000characters, following the normalization convention of\[[11](https://arxiv.org/html/2607.10846#bib.bib11)\]\. Per\-speakerr​\(neg,emph\)r\(\\text\{neg\},\\text\{emph\}\)andr​\(neg,hedged\)r\(\\text\{neg\},\\text\{hedged\}\)are computed across videos, matching the analysis unit of the LLM annotation exactly\.

#### Domain coverage limitation\.

The lexicon in\[[11](https://arxiv.org/html/2607.10846#bib.bib11)\]was originally developed and validated on financial and geopolitical discourse \(Dalio, Zeihan, Rogoff\)\. The Financial/Political Risk subcategory \(terms such asdebt,inflation,sanctions,tariffs\) has natural relevance to central banking and policy domains, but the lexicon lacks domain\-specific vocabulary for academic discourse \(hedging conventions such aspotentially,suggests that,to some extent\) and media commentary \(evaluative intensifiers such asabsolutely devastating,completely absurd\)\. This domain skew in lexical coverage is a known limitation: the smaller\|Δ​r\|\|\\Delta r\|observed in Academia \(0\.21 vs\. a cross\-domain mean of 0\.45\) may partly reflect under\-detection of domain\-specific hedging rather than genuine preprocessing robustness\. We flag lexicon expansion—particularly academic hedging markers and media intensifiers—as a necessary extension for cross\-domain validity, and note that this limitation strengthens the case for using LLM annotation as a complementary measurement method not constrained by pre\-specified lexical patterns\.

## 5Pipeline Sensitivity Results

We apply the framework defined in Section[3](https://arxiv.org/html/2607.10846#S3): for each speaker with annotations from both pipelines, we computeΔ​r​\(neg,emph\)\\Delta r\(\\text\{neg\},\\text\{emph\}\)andΔ​r​\(neg,hedged\)\\Delta r\(\\text\{neg\},\\text\{hedged\}\), then aggregate by domain and across measurement methods\.

### 5\.1LLM\-Based Pipeline Sensitivity

For each of the 41 speakers with LLM annotations from both pipelines, we compute per\-video proportions andr​\(neg,emph\)r\(\\text\{neg\},\\text\{emph\}\)andr​\(neg,hedged\)r\(\\text\{neg\},\\text\{hedged\}\)under each pipeline\. Per\-speakerΔ​r=rASR−rVTT\\Delta r=r\_\{\\text\{ASR\}\}\-r\_\{\\text\{VTT\}\}; sign reversal isrVTT⋅rASR<0r\_\{\\text\{VTT\}\}\\cdot r\_\{\\text\{ASR\}\}<0\. Ten speakers with fewer than 4 videos—the minimum for a stable Pearson correlation—were excluded from this analysis\.

#### Pipeline sensitivity is concentrated in low\-N speakers\.

Table[3](https://arxiv.org/html/2607.10846#S5.T3)reports the aggregate results\. Across all 41 speakers, the mean\|Δ​r​\(N,E\)\|\|\\Delta r\(\\text\{N,E\}\)\|is0\.410\.41\. However, this average masks a strong dependence on sample size\. The four best\-sampled speakers—Dalio \(34 videos\), Wood \(17\), Rogoff \(17\), Zeihan \(16\)—have a mean\|Δ​r​\(N,E\)\|\|\\Delta r\(\\text\{N,E\}\)\|of only0\.130\.13, with no individualΔ​r\\Delta rexceeding0\.190\.19\. By contrast, speakers with 4–5 videos show a mean\|Δ​r\|\|\\Delta r\|of0\.440\.44, and their per\-speaker bootstrap confidence intervals span nearly the entire\[−2,2\]\[\-2,2\]range \(mean width:2\.652\.65\)\. Pipeline\-induced sign disagreement \(24%24\\%observed;35%35\\%bootstrap\-estimated\) is almost entirely driven by these low\-N speakers, whose correlational estimates are inherently volatile regardless of the pipeline comparison being tested\.

#### Domain\-level patterns\.

Within the limitations imposed by small per\-speaker video samples, domain\-level aggregation reveals consistent ordering: Media \(\|Δ​r​\(N,E\)\|=0\.32\|\\Delta r\(\\text\{N,E\}\)\|=0\.32\) is the most robust pipeline, while Finance and Academia \(0\.460\.46\) are the most sensitive\. The correlation between VTT and ASRr​\(N,E\)r\(\\text\{N,E\}\)across all speakers isr=0\.34r=0\.34\(p=0\.03p=0\.03\), a modest but statistically significant relationship\.

Table 3:LLM\-based pipeline sensitivity \(Δ​r​\(neg,emph\)\\Delta r\(\\text\{neg\},\\text\{emph\}\)\): aggregate and domain\-level scores\.Selected illustrative speakers \(full list in Appendix\):

Table 4:LLM\-based pipeline sensitivity: selected per\-speaker scores\.The same analysis forΔ​r​\(neg,hedged\)\\Delta r\(\\text\{neg\},\\text\{hedged\}\)yields similar results \(mean\|Δ​r\|=0\.43\|\\Delta r\|=0\.43; observed sign disagreement29%29\\%\), with the domain ordering following the same pattern—Geopolitics the most sensitive \(0\.520\.52\) and Media the least \(0\.300\.30\)\.

### 5\.2Keyword\-Based Pipeline Sensitivity

We apply the identical framework using keyword\-lexicon scoring\. Per\-video keyword frequencies are normalized by character length and correlated across videos per speaker, matching the LLM analysis unit\.

Table 5:Keyword\-based pipeline sensitivity: per\-speaker \(selected\)\.Mean\|Δ​r​\(N,E\)\|=0\.44\|\\Delta r\(\\text\{N,E\}\)\|=0\.44under keyword scoring across 41 speakers;32%32\\%\(13/41\) show sign reversal\. The effect is domain\-dependent: Geopolitics shows the largest mean\|Δ​r​\(N,E\)\|\|\\Delta r\(\\text\{N,E\}\)\|\(0\.72\), while Academia shows the smallest \(0\.24\)\. Results forΔ​r​\(neg,hedged\)\\Delta r\(\\text\{neg\},\\text\{hedged\}\)are comparable \(mean\|Δ​r\|=0\.36\|\\Delta r\|=0\.36;32%32\\%sign disagreement\)\.

### 5\.3Cross\-Method Synthesis: When Measurement Diverges from Pipeline

The preceding sections examined pipeline sensitivity within each measurement method separately\. We now use the continuous divergence metricΔimethod​\(P\)=\|mLLM​\(P\)−mKW​\(P\)\|\\Delta\_\{i\}^\{\\text\{method\}\}\(P\)=\|m^\{\\text\{LLM\}\}\(P\)\-m^\{\\text\{KW\}\}\(P\)\|\(Equation[2](https://arxiv.org/html/2607.10846#S3.E2)\) to quantify measurement\-driven disagreement and compare it directly to pipeline\-driven disagreement\.

For the four best\-sampled speakers, the comparison is stark:

- •Rogoff \(17 videos\)\.Pipeline divergence under LLM:\|Δ​r\|=0\.19\|\\Delta r\|=0\.19\. Method divergence on VTT:\|Δ​rLLM−rKW\|=0\.90\|\\Delta r^\{\\text\{LLM\}\}\-r^\{\\text\{KW\}\}\|=0\.90\. Method divergence is4\.7×4\.7\\timeslarger\.
- •Zeihan \(16 videos\)\.Pipeline:0\.110\.11\. Method \(VTT\):0\.820\.82\. Method divergence is7\.8×7\.8\\timeslarger\.
- •Dalio \(34 videos\)\.Pipeline:0\.120\.12\. Method \(VTT\):0\.330\.33\. Method divergence is2\.8×2\.8\\timeslarger\.
- •Wood \(17 videos\)\.Pipeline:0\.100\.10\. Method \(VTT\):0\.100\.10\. Comparable—both near zero\.

Across all 41 speakers, measurement divergence \(Δimethod\\Delta\_\{i\}^\{\\text\{method\}\}\) exceeds pipeline divergence \(Δipipeline\\Delta\_\{i\}^\{\\text\{pipeline\}\}\) in2727of4141cases \(66%66\\%\)\. The meanΔmethod\\Delta^\{\\text\{method\}\}is0\.760\.76, nearly double the meanΔpipeline\(LLM\)\\Delta^\{\\text\{pipeline\(LLM\)\}\}of0\.410\.41\. These results confirm that for a majority of speakers—including every well\-sampled speaker—measurement method choice is a larger source of instability than preprocessing pipeline choice\.

Table 6:Cross\-method pipeline sensitivity: domain\-level summary\.Table[6](https://arxiv.org/html/2607.10846#S5.T6)summarizes the cross\-method comparison\. The overall pipeline sensitivity is balanced between methods \(LLM mean\|Δ​r\|=0\.41\|\\Delta r\|=0\.41, KW=0\.44=0\.44;49%49\\%of speakers show more stability with KW,51%51\\%with LLM\)\. The domain\-level patterns from Section[5\.1](https://arxiv.org/html/2607.10846#S5.SS1)persist: Academia is the most KW\-stable domain \(\|Δ​r\|=0\.24\|\\Delta r\|=0\.24vs\.0\.460\.46for LLM\), while Geopolitics is the most LLM\-stable \(0\.410\.41vs\.0\.720\.72\)\. The key insight is that when the two methods agree on the same analysis unit and small\-sample speakers are excluded, the overall sensitivity difference between methods is negligible—what matters is the systematic cross\-method disagreement on the coupling direction itself, not which method is more sensitive to preprocessing\.

### 5\.4Domain\-Level Aggregation

Table[7](https://arxiv.org/html/2607.10846#S5.T7)combines LLM and keyword results into a per\-domain pipeline sensitivity profile for all 41 speakers\.

Table 7:Combined domain pipeline sensitivity profile\.Under LLM annotation, Media shows the smallest\|Δ​r\|\|\\Delta r\|\(0\.320\.32\), while Finance and Academia tie for the largest \(0\.460\.46\)\. Under keyword scoring, Academia is the most robust \(0\.240\.24\), while Geopolitics is the most sensitive \(0\.720\.72\)\. The bootstrap\-derived probability of sign disagreement follows the same domain gradient—Media25%25\\%, Finance37%37\\%, Academia42%42\\%—suggesting that the domain\-level pattern is robust to per\-speaker uncertainty\. The domain gradient differs by method: LLM annotation compresses the domain range \(1\.4×\\timesfrom smallest to largest\) compared to keyword scoring \(3\.0×\\times\)\. Professional communication style modulates how preprocessing choices propagate to downstream measurements—academic hedging markers and media intensifiers exhibit different robustness to VTT fragmentation than the geopolitical discourse\.

## 6Discussion

### 6\.1Summary

1. 1\.Pipeline sensitivity is bounded and predictable\.For speakers with adequate video samples \(N≥16\\geq 16\), the mean\|Δ​r\|\|\\Delta r\|from pipeline choice is0\.130\.13under LLM annotation\. LargerΔ​r\\Delta rvalues are concentrated in speakers with≤5\\leq 5videos, where correlational estimates are inherently unstable regardless of the pipeline comparison\.
2. 2\.Cross\-method disagreement is the more serious validity threat\.LLM and keyword\-lexicon methods disagree on the sign ofr​\(neg,emph\)r\(\\text\{neg\},\\text\{emph\}\)in49%49\\%of cases—including several well\-sampled speakers \(Rogoff, Zeihan\) where both methods are internally stable but give opposite coupling directions\. This suggests that measurement method choice can systematically alter scientific conclusions even when preprocessing is held constant\.
3. 3\.Aggregate proportions mask both sources of instability\.Mean\|Δ​p​\(neg\)\|=0\.06\|\\Delta p\(\\text\{neg\}\)\|=0\.06is stable across pipelines and methods, yet both pipeline effects \(in low\-N speakers\) and measurement effects \(in high\-N speakers\) produce substantively different correlational conclusions\. Centering analysis on aggregate proportions alone would miss both\.
4. 4\.Domain modulates sensitivity differently for pipelines and methods\.Media is the most pipeline\-robust domain \(\|Δ​r​\(N,E\)\|=0\.32\|\\Delta r\(\\text\{N,E\}\)\|=0\.32under LLM\); Geopolitics is the most pipeline\-sensitive \(0\.720\.72under keyword\)\. These domain\-level patterns are consistent within each measurement method and provide a basis for anticipating where preprocessing is likely to matter\.

### 6\.2Practical Recommendations

1. 1\.Report N\-dependent diagnostics\.Pipeline sensitivity should be reported with sample\-size stratification: speakers with≤5\\leq 5videos require bootstrap calibration or exclusion; speakers with≥16\\geq 16videos can be compared directly\.
2. 2\.Test≥2\\geq 2measurement methods\.Method divergence itself is a diagnostic signal—if LLM and keyword give opposite coupling directions, the conclusion is not yet ready for scientific use regardless of preprocessing stability\.
3. 3\.Separate preprocessing from measurement\.Report cross\-pipelineΔ​r\\Delta rwithin each method and cross\-method sign agreement within each pipeline separately, to identify which source of variation is driving instability\.
4. 4\.Include domain as a moderator\.Professional communication style affects both pipeline sensitivity and measurement method divergence\.

### 6\.3Limitations

Single LLM provider \(DeepSeek\) for annotation; English\-only corpus; two preprocessing pipelines only \(VTT and AssemblyAI\); keyword lexicon not validated for all domains equally\. The observed pipeline sensitivity reflects a composite of diarization and sentence\-boundary differences that cannot be fully separated in the current design\. TheN≥4N\\geq 4video threshold, while necessary to avoid spurious correlations atN=3N=3, constrains the sample to 41 speakers\. The cross\-method disagreement finding is based on two methods only \(LLM zero\-shot and keyword lexicon\); additional methods \(e\.g\., human annotation, dictionary\-based sentiment\) would strengthen the generality of the conclusion\. Future work with larger per\-speaker video samples would improve precision for both pipeline and method comparisons\.

## 7Conclusion

Pipeline sensitivity analysis reveals that preprocessing choices can produce measurable changes in cross\-dimensional correlations, but the magnitude of these changes is predictable: bounded for adequately sampled speakers \(\|Δ​r\|≈0\.13\|\\Delta r\|\\approx 0\.13for N≥16\\geq 16\), larger for small\-sample speakers where correlational estimates are inherently noisy\. The more fundamental threat to validity—and the one that persists even in well\-sampled speakers—is cross\-method disagreement: LLM annotation and keyword\-lexicon scoring give opposite coupling directions for a substantial fraction of speakers, including several where both methods are internally stable\. CSS researchers studying valence\-modality coupling in interview data should \(1\) verify that their conclusions are robust to the preprocessing pipeline, \(2\) verify that their conclusions are robust to the measurement method, and \(3\) explicitly separate these two sources of instability in their reporting\. A correlation that survives both alternative preprocessing and alternative measurement is more credible than one validated within a single pipeline\-method combination\.

Future work should extend this comparison to open\-source diarizers \(pyannote\-audio\[[12](https://arxiv.org/html/2607.10846#bib.bib12)\], WhisperX\) to test generalizability beyond the AssemblyAI–VTT comparison, develop a theoretical simulation mapping the parameter space\(ρ,δ,N\)\(\\rho,\\delta,N\)—true correlation, contamination strength, sample size—to sign\-reversal boundaries for practical diagnostic use, and extend the analysis beyond valence–modality coupling to other downstream constructs \(emotion classification, toxicity detection, politeness\) to establish whether the aggregate\-stability/correlational\-fragility pattern is general or task\-specific\.

## Ethics Statement

This research uses publicly available YouTube interviews featuring public figures\. Audio files were downloaded temporarily and are not redistributed\. No personally identifying information beyond the public figure’s name was involved\.

## References

- \[1\]P\. Wu\. “Beyond Audio: Advancing Speaker Diarization with Text\-based Methodologies and Comprehensive Evaluation\.” Honors Thesis, Emory University, 2024\.
- \[2\]M\. Li\. “Joint Text and Audio Multi\-modal Speaker Diarization\.” Emory NLP Seminar, Spring 2025\.
- \[3\]“Beyond ASR: Achieving Privacy\-Aligned Low Diarization Error Rates in Psychiatric Transcripts Using Large Language Models\.” InProc\. Cyber Research Conference Ireland, 2025\.
- \[4\]F\. Gilardi, M\. Alizadeh, and M\. Kubli\. “ChatGPT outperforms crowd workers for text\-annotation tasks\.”Proceedings of the National Academy of Sciences, 120\(30\):e2305016120, 2023\.
- \[5\]C\. Ziems, W\. Held, O\. Shaikh, J\. Chen, Z\. Zhang, and D\. Yang\. “Can large language models transform computational social science?”Computational Linguistics, 50\(1\):237–272, 2024\.
- \[6\]C\. A\. Bail\. “Can generative AI improve social science?”Proceedings of the National Academy of Sciences, 121\(21\):e2314021121, 2024\.
- \[7\]C\. Spearman\. “The proof and measurement of association between two things\.”American Journal of Psychology, 15\(1\):72–101, 1904\.
- \[8\]J\. Bound, C\. Brown, and N\. Mathiowetz\. “Measurement error in survey data\.” InHandbook of Econometrics, Vol\. 5, pp\. 3705–3843, 2001\.
- \[9\]B\. Marie, A\. Fujita, and R\. Rubino\. “Scientific credibility of machine translation research: A meta\-evaluation of 769 papers\.” InProc\. ACL, 2021\.
- \[10\]J\. Pineau, P\. Vincent\-Lamarre, K\. Sinha, et al\. “Improving reproducibility in machine learning research\.”Journal of Machine Learning Research, 22\(164\):1–20, 2021\.
- \[11\]B\. Chen\. “When Certainty Is an Artifact: Keyword Lexicon Blindness and the \(Mis\)Measurement of Rhetorical Stance\.” arXiv:2606\.26062, June 2026\.
- \[12\]H\. Bredin\. “pyannote\.audio: neural building blocks for speaker diarization\.” InProc\. ICASSP, 2020\.
- \[13\]I\. Magar and R\. Schwartz\. “Data contamination: From memorization to exploitation\.” InProc\. ACL, 2023\.
- \[14\]P\. Liang, R\. Bommasani, T\. Lee, et al\. “Holistic evaluation of language models\.”Transactions on Machine Learning Research, 2023\.
- \[15\]J\. Zhao, Y\. Wang, and K\. Liu\. “Calibrating LLM annotations: Context length and text segmentation effects on sentiment classification\.” InProc\. EMNLP, 2024\.
- \[16\]M\. Liu, S\. Chen, and R\. Bansal\. “The chunking effect: How text segmentation influences LLM\-based emotion classification\.” InProc\. ACL, 2024\.

Similar Articles

StabilityBench: Benchmarking Instability in LLMs

arXiv cs.LG

StabilityBench is a benchmark operator that transforms single-turn LLM evaluations into multi-turn interactions with user simulations and baiting modules, revealing significant performance instability in current models and motivating more realistic evaluation protocols.