Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement
Summary
The paper proposes Tempered Evidence Fusion (TEF), a method for long-text value measurement that weights sentence-level LLM judgments by information gain, and introduces the MIND benchmark dataset for evaluation.
View Cached Full Text
Cached at: 09/24/26, 09:15 AM
# Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement
Source: [https://arxiv.org/html/2609.27165](https://arxiv.org/html/2609.27165)
Yuhe Wu1,∗, Rui Qian2,∗, Guangyu Wang1,∗, Yuran Chen3, Yuanchao Zhu4, Junjie Yang5Zhengheng Li6, Jiulin Cai7, Tianyi Zhang8, Zihan Dong9, Jiaxin Liu1, Yujie Chen10, Guang Zhang1,†††thanks:\*Equal contribution†Corresponding author
###### Abstract
Large language models \(LLMs\) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance\-bearing sentences\. Existing approaches either ask the model to predict a document\-level label directly, which can be overconfident, or aggregate sentence\-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative\. We formulate long\-text value measurement as a decision\-fusion problem and propose Tempered Evidence Fusion \(TEF\), a training\-free rule that weights each sentence’s log\-odds by its normalized information gain, as derived from a generalized Bayesian posterior\. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes\-optimal weight of decisive evidence\. We further introduce Multi\-event Insight Network Dimensions \(MIND\), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions\. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4\.5 accuracy points and 4\.6 macro\-F1 points across five LLMs and two languages\. MIND dataset and code are available at[GitHub](https://github.com/Kzczc/ICASSP2027-TEF)\.
###### Index Terms:
Decision fusion, large language models, uncertainty, calibration, computational social science
††address:1HKUST\(GZ\)2FDU3DUFE4UESTC5UMD
6SEU7USTC8Independent9Georgia Tech10CUHK\(SZ\)## 1Introduction
Public value orientations are key latent variables in computational social science, capturing how people position themselves on issues such as economic regulation, climate policy, or cultural change\. Social media makes these orientations observable through text at scale\[[6](https://arxiv.org/html/2609.27165#bib.bib9),[17](https://arxiv.org/html/2609.27165#bib.bib10)\]\. Yet this scale, together with the subjectivity of human annotation, has motivated growing interest in using LLMs as measurement instruments\[[20](https://arxiv.org/html/2609.27165#bib.bib11),[14](https://arxiv.org/html/2609.27165#bib.bib12)\]\. Existing LLM\-based pipelines usually either ask the model to read the whole post and return a document\-level label\[[10](https://arxiv.org/html/2609.27165#bib.bib26),[31](https://arxiv.org/html/2609.27165#bib.bib27),[27](https://arxiv.org/html/2609.27165#bib.bib28)\], or split the post into sentences and aggregate sentence\-level predictions by voting\[[28](https://arxiv.org/html/2609.27165#bib.bib13)\]\.
However, we observe that little research has examined how sentence\-level LLM judgments should be fused when evidence is unevenly distributed across a long post\. In Fig\.[1](https://arxiv.org/html/2609.27165#S1.F1), long social media posts often mix background information, quotations, and concessions, with only a few stance\-bearing sentences\. In such cases, direct prediction can collapse heterogeneous evidence into one overconfident label, while majority or soft voting can allow uncertain sentences to outvote a few decisive ones\. This leads us to ask:*How can sentence\-level judgments be combined so that decisive evidence dominates the final decision while uncertain sentences are discounted?*
Figure 1:Long\-text value measurement requires counting informative evidence rather than treating all sentences equally\.LEFT:Direct prediction collapses mixed content into one label, while voting can let many uncertain sentences outweigh a few decisive ones\.RIGHT:TEF weights sentence log\-odds by normalized information gain, discounting uncertain judgments so decisive evidence drives the document\-level orientation\.Bearing this in mind, we formulate long\-text value measurement as a decision\-fusion problem, where each sentence acts as a local detector that emits a soft decision\. Under conditional independence and calibrated sentence posteriors, the Bayes\-optimal document\-level rule sums sentence log\-likelihood ratios\. This observation places existing voting rules on a common scale: majority voting keeps only the sign of each sentence judgment, and soft voting uses a bounded surrogate that underweights highly decisive sentences\. At the same time, directly summing LLM logits is unreliable because next\-token probabilities are only approximately calibrated, especially near the point of maximal uncertainty\.
To this end, we presentTemperedEvidenceFusion \(TEF\), a training\-free fusion rule that tempers each sentence’s log\-odds by its normalized information gain, as derived from a generalized Bayesian posterior\[[4](https://arxiv.org/html/2609.27165#bib.bib4),[12](https://arxiv.org/html/2609.27165#bib.bib5)\]\. The resulting score nearly vanishes for uncertain sentences but preserves the Bayes\-optimal weight of decisive evidence, yielding a soft censoring effect similar in spirit to censoring rules in distributed detection\[[22](https://arxiv.org/html/2609.27165#bib.bib2)\]\. In Fig\.[1](https://arxiv.org/html/2609.27165#S1.F1)\(right\), TEF counts evidence rather than sentences, allowing a few informative sentences to determine the document\-level orientation\.
Our contributions are threefold\. \(i\) We cast long\-text value measurement with LLMs as a decision\-fusion problem and derive TEF from a tempered Bayesian posterior, showing that it suppresses uncertain sentences while preserving the weight of decisive evidence\. \(ii\) We introduce MIND, a five\-year, cross\-lingual benchmark of 8,358 posts over six value dimensions\. \(iii\) We show that TEF outperforms the strongest of the Direct, Majority Vote, and Soft Vote baselines by 4\.5 accuracy points and 4\.6 macro\-F1 points on average across five LLMs and two languages, yielding better\-calibrated estimates\.
## 2Related Work
LLM\-based stance and ideology measurement has expanded from short\-form texts such as tweets to longer, multi\-issue documents\[[23](https://arxiv.org/html/2609.27165#bib.bib14),[2](https://arxiv.org/html/2609.27165#bib.bib15),[3](https://arxiv.org/html/2609.27165#bib.bib16)\]\. In parallel, the reliability of LLM confidence has been studied from multiple perspectives, including token\-level probabilities\[[15](https://arxiv.org/html/2609.27165#bib.bib24),[16](https://arxiv.org/html/2609.27165#bib.bib6)\], verbalized confidence\[[26](https://arxiv.org/html/2609.27165#bib.bib25),[24](https://arxiv.org/html/2609.27165#bib.bib7)\], and aggregation methods beyond majority voting\[[1](https://arxiv.org/html/2609.27165#bib.bib8)\]\. In signal processing, fusing local decisions is a classical problem\[[25](https://arxiv.org/html/2609.27165#bib.bib31)\]: log\-linear opinion pools combine probabilistic judgments\[[9](https://arxiv.org/html/2609.27165#bib.bib3)\], the Chair–Varshney rule weights local decisions by their reliabilities\[[5](https://arxiv.org/html/2609.27165#bib.bib1)\], and censoring schemes transmit only observations with informative likelihood ratios\[[22](https://arxiv.org/html/2609.27165#bib.bib2)\]\. Despite these connections, long\-text measurement with LLMs still largely relies on equal\-weight voting\. In contrast, we formulate long\-text value measurement as sentence\-level decision fusion and derive an evidence\-aware fusion rule, TEF, that discounts uncertain sentences while preserving the weight of decisive ones\.
## 3Proposed TEF
Problem Formulation\.We formulate long\-text value measurement as a decision\-fusion problem\. The goal is to infer a latent social variable from a collection of textual segments whose evidential strength may vary substantially within the same document\. Let𝒳\\mathcal\{X\}denote the input space and let𝒴=\{c1,…,cK\}\\mathcal\{Y\}=\\\{c\_\{1\},\\ldots,c\_\{K\}\\\}be the set of target measurement categories\. For an observationX∈𝒳X\\in\\mathcal\{X\}, the latent measurement variableYYtakes values in𝒴\\mathcal\{Y\}\. We segment each long text intoNNsentence\-level units,
X↦𝒟\(X\)=\{D1,D2,…,DN\},\\vskip\-2\.84544ptX\\mapsto\\mathcal\{D\}\(X\)=\\\{D\_\{1\},D\_\{2\},\\ldots,D\_\{N\}\\\},\(1\)whereDiD\_\{i\}denotes theii\-th segment\. This granularity preserves local semantic coherence while keeping inference computationally manageable\. For each segmentDiD\_\{i\}, a language modelℳ\\mathcal\{M\}returns a probability distribution over the measurement categories, denoted by𝐩i=ℳ\(Di\)=\(pi\(1\),…,pi\(K\)\)\\mathbf\{p\}\_\{i\}=\\mathcal\{M\}\(D\_\{i\}\)=\\big\(p\_\{i\}^\{\(1\)\},\\ldots,p\_\{i\}^\{\(K\)\}\\big\), where each entrypi\(k\)=Pℳ\(Y=ck∣Di\)p\_\{i\}^\{\(k\)\}=P\_\{\\mathcal\{M\}\}\(Y=c\_\{k\}\\mid D\_\{i\}\)and𝐩i∈ΔK−1\\mathbf\{p\}\_\{i\}\\in\\Delta^\{K\-1\}\. The distribution encodes both a local predictiony^i=argmaxkpi\(k\)\\hat\{y\}\_\{i\}=\\arg\\max\_\{k\}p\_\{i\}^\{\(k\)\}and its uncertainty, measured by entropyH\(𝐩i\)=−∑k=1Kpi\(k\)logpi\(k\)\.H\(\\mathbf\{p\}\_\{i\}\)=\-\\sum\_\{k=1\}^\{K\}p\_\{i\}^\{\(k\)\}\\log p\_\{i\}^\{\(k\)\}\.Common aggregation rules discard part of this probabilistic evidence\. Majority voting uses only the segment\-level hard labels, while soft voting averages probability vectors with equal weight,
Y^MV=mode\(\{y^i\}i=1N\),Y^SV=argmaxc∈𝒴1N∑i=1Npi\(c\)\.\\vskip\-5\.69046pt\\hat\{Y\}\_\{\\mathrm\{MV\}\}=\\mathrm\{mode\}\\big\(\\\{\\hat\{y\}\_\{i\}\\\}\_\{i=1\}^\{N\}\\big\),\\qquad\\hat\{Y\}\_\{\\mathrm\{SV\}\}=\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}p\_\{i\}^\{\(c\)\}\.\(2\)Both rules treat all segments as equally reliable\. This is problematic for scientific measurement from long texts, where decisive evidence may appear in only a few segments and many other segments may be ambiguous, rhetorical, or weakly relevant\. TEF instead aggregates reliability\-weighted evidence\. Letwi=ω\(𝐩i\)∈\[0,1\]w\_\{i\}=\\omega\(\\mathbf\{p\}\_\{i\}\)\\in\[0,1\]be the reliability assigned to segmentDiD\_\{i\}, and letg:\[0,1\]→ℝg:\[0,1\]\\to\\mathbb\{R\}be a scoring transformation\. The document\-level measurement isY^=argmax∑i=1Nc∈𝒴wig\(pi\(c\)\)\.\\hat\{Y\}=\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\sum\_\{i=1\}^\{N\}w\_\{i\}\\,g\\big\(p\_\{i\}^\{\(c\)\}\\big\)\.The following paragraphs specifyω\\omegaandgg, which together define the fusion rule of TEF\.
Reliability Weighting\.TEF instantiates the preceding framework with entropy\-based reliability weighting and one\-vs\-rest log\-odds evidence\. Algorithm[1](https://arxiv.org/html/2609.27165#alg1)gives the complete procedure\. We use normalized information gain as the segment reliability,wi=1−H\(𝐩i\)logK∈\[0,1\]\.w\_\{i\}=1\-\\frac\{H\(\\mathbf\{p\}\_\{i\}\)\}\{\\log K\}\\in\[0,1\]\.This weight is close to one for concentrated distributions and close to zero for near\-uniform distributions\. Thus confident segments contribute strongly, while uncertain segments are softly censored\.
Evidence Aggregation\.For each categorycc, TEF transforms the segment posterior into a one\-vs\-rest log\-odds score,g\(pi\(c\)\)=logpi\(c\)1−pi\(c\)\.g\\big\(p\_\{i\}^\{\(c\)\}\\big\)=\\log\\frac\{p\_\{i\}^\{\(c\)\}\}\{1\-p\_\{i\}^\{\(c\)\}\}\.The aggregated evidence for categoryccis
Score\(c\)=∑i=1Nwilogpi\(c\)1−pi\(c\)\.\\operatorname\{Score\}\(c\)=\\sum\_\{i=1\}^\{N\}w\_\{i\}\\log\\frac\{p\_\{i\}^\{\(c\)\}\}\{1\-p\_\{i\}^\{\(c\)\}\}\.\\vskip\-8\.5359pt\(3\)The final measurement outcome isY^=argmaxc∈𝒴Score\(c\)\\hat\{Y\}=\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\operatorname\{Score\}\(c\)\. In implementation, probabilities are clipped to\[ϵ,1−ϵ\]\[\\epsilon,1\-\\epsilon\]before applying the logit, and log\-odds values are clipped to\[−M,M\]\[\-M,M\]\. Unless otherwise stated, we setϵ=10−6\\epsilon=10^\{\-6\}andM=10M=10\. Unlike majority voting, TEF does not reduce𝐩i\\mathbf\{p\}\_\{i\}to a hard label before aggregation\. Unlike soft voting, TEF does not assume that all segments are equally informative\. It therefore preserves distributional information while reducing the impact of noisy or ambiguous segments\.
Algorithm 1Tempered Evidence Fusion1:Text
XX, model
ℳ\\mathcal\{M\}, categories
𝒴=\{c1,…,cK\}\\mathcal\{Y\}=\\\{c\_\{1\},\\ldots,c\_\{K\}\\\}, clip bound
MM, stability constant
ϵ\\epsilon\.
2:Measurement outcome
Y^\\hat\{Y\}
3:
\{D1,…,DN\}←𝒟\(X\)\\\{D\_\{1\},\\ldots,D\_\{N\}\\\}\\leftarrow\\mathcal\{D\}\(X\)⊳\\trianglerightSegment into sentences
4:Initialize
Score\(c\)←0\\operatorname\{Score\}\(c\)\\leftarrow 0for all
c∈𝒴c\\in\\mathcal\{Y\}
5:for
i=1i=1to
NNdo
6:
𝐩i←ℳ\(Di\)\\mathbf\{p\}\_\{i\}\\leftarrow\\mathcal\{M\}\(D\_\{i\}\)⊳\\trianglerightSegment\-level distribution
7:
Hi←−∑k=1Kpi\(k\)logpi\(k\)H\_\{i\}\\leftarrow\-\\sum\_\{k=1\}^\{K\}p\_\{i\}^\{\(k\)\}\\log p\_\{i\}^\{\(k\)\}⊳\\trianglerightEntropy
8:
wi←1−Hi/logKw\_\{i\}\\leftarrow 1\-H\_\{i\}/\\log K⊳\\trianglerightReliability weight
9:for
c∈𝒴c\\in\\mathcal\{Y\}do
10:
p~←clip\(pi\(c\),ϵ,1−ϵ\)\\tilde\{p\}\\leftarrow\\mathrm\{clip\}\(p\_\{i\}^\{\(c\)\},\\epsilon,1\-\\epsilon\)
11:
ℓ←clip\(logp~1−p~,−M,M\)\\ell\\leftarrow\\mathrm\{clip\}\\\!\\left\(\\log\\frac\{\\tilde\{p\}\}\{1\-\\tilde\{p\}\},\-M,M\\right\)
12:
Score\(c\)←Score\(c\)\+wiℓ\\operatorname\{Score\}\(c\)\\leftarrow\\operatorname\{Score\}\(c\)\+w\_\{i\}\\ell
13:endfor
14:endfor
15:
Y^←argmaxc∈𝒴Score\(c\)\\hat\{Y\}\\leftarrow\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\operatorname\{Score\}\(c\)
16:return
Y^\\hat\{Y\}
Theoretical Analysis\.
Proposition 1 \(Information Preservation\)\.*Since the hard labely^i=argmaxkpi\(k\)\\hat\{y\}\_\{i\}=\\arg\\max\_\{k\}p\_\{i\}^\{\(k\)\}is a deterministic function of𝐩i\\mathbf\{p\}\_\{i\}, the data processing inequality\[[7](https://arxiv.org/html/2609.27165#bib.bib29)\]givesI\(Y,𝐩i\)≥I\(Y,y^i\)I\(Y;\\mathbf\{p\}\_\{i\}\)\\geq I\(Y;\\hat\{y\}\_\{i\}\)\. Thus, hard voting can discard reliability information even when two segments support the same category with different confidence\.*
TEF follows this principle by retaining the full distribution𝐩i\\mathbf\{p\}\_\{i\}instead of reducing each segment toy^i\\hat\{y\}\_\{i\}\. Its entropy\-based weightwiw\_\{i\}distinguishes decisive evidence from ambiguous evidence, so two segments with the same hard label can contribute different amounts\. The method therefore preserves the information discarded by majority voting while producing an additive document\-level score\.
Proposition 2 \(Tempered Bayesian Pooling\)\.*For a binary contrast with equal priors, calibrated conditionally independent segment posteriors satisfy the additive Bayes log\-odds identity\[[29](https://arxiv.org/html/2609.27165#bib.bib30)\]logP\(Y=c∣D1:N\)P\(Y=c¯∣D1:N\)=∑i=1NlogP\(Y=c∣Di\)P\(Y=c¯∣Di\)\\log\\frac\{P\(Y=c\\mid D\_\{1:N\}\)\}\{P\(Y=\\bar\{c\}\\mid D\_\{1:N\}\)\}=\\sum\_\{i=1\}^\{N\}\\log\\frac\{P\(Y=c\\mid D\_\{i\}\)\}\{P\(Y=\\bar\{c\}\\mid D\_\{i\}\)\}\. Therefore, evidence is naturally accumulated on the log\-odds scale rather than through hard counts or raw probability averaging\.*
TEF uses the one\-vs\-rest transformg\(p\)=logp1−pg\(p\)=\\log\\frac\{p\}\{1\-p\}to implement this additive evidence scale\. Because LLM posteriors are approximate, it tempers each log\-odds contribution\[[4](https://arxiv.org/html/2609.27165#bib.bib4)\]withwi=\(logK−H\(𝐩i\)\)/logKw\_\{i\}=\(\\log K\-H\(\\mathbf\{p\}\_\{i\}\)\)/\\log K, the normalized information gain from a uniform prior\. Near\-uniform segments therefore have little influence, while low\-entropy stance\-bearing segments retain a contribution close to the Bayesian additive statistic\. This gives TEF its soft\-censoring behavior, as uncertain sentences do not accumulate by sheer number, while decisive sentences can determine the final document\-level measurement outcome even when they are few\.
Figure 2:Construction pipeline of MIND\. Posts are collected from public sources, filtered for relevance, anonymized, and deduplicated, then annotated in three rounds for dimension, subcategory, and orientation before being balanced and refined by LLM agents\.Figure 3:Statistics of MIND\. Radial bars report the number of labels for each value orientation, with colors denoting dimensions and paired bars representing the opposing orientations of each subcategory\. Panels \(a\) and \(b\) correspond to Chinese and English posts\.
## 4The MIND Benchmark
We construct the MIND dataset, a longitudinal, cross\-lingual benchmark for measuring public value orientations\. The final corpus contains 8,358 posts, including 5,474 Chinese posts and 2,884 English posts over six value dimensions in Fig\.[3](https://arxiv.org/html/2609.27165#S3.F3)\. MIND is built through a three\-stage pipeline\.Stage I: Data Collection\.We collect candidate posts from X, Weibo, Reddit, Zhihu, discussion forums, and news feeds from January 2020 to January 2025, covering major public events such as the Russia–Ukraine war and the rise of generative AI\. For each value dimension, keyword search is anchored to high\-search\-volume events\. Raw posts are filtered for relevance and opinion content, anonymized, de\-duplicated, and screened by an automatic quality score\.Stage II: Data Annotation\.We adopt a three\-round annotation framework to label the value dimension, subcategory, and orientation in sequence\. Each round uses three annotators and accepts a label only when at least two agree\. Failed or disputed cases enter arbitration, samples without a stable majority are discarded, and accepted samples pass distribution and quality checks before the evolution stage\.Stage III: Agent\-Driven Data Evolution\.To address class imbalance and semantic ambiguity, we introduce an agent\-driven evolution stage\. When a category is underrepresented, a balancing agent generates synthetic samples through generative oversampling\. For low\-confidence or ambiguous samples, a critic agent evaluates the alignment between text and label, and a semantic\-edit agent revises the text to make the labeled orientation explicit\. Only verified samples are retained in the final benchmark\.
## 5Experiments
Experiment Settings\.We evaluate three open\-source models, Qwen2\.5\-7B\[[21](https://arxiv.org/html/2609.27165#bib.bib18)\], LLaMA3\-8B\[[11](https://arxiv.org/html/2609.27165#bib.bib19)\], and Qwen3\-14B\[[30](https://arxiv.org/html/2609.27165#bib.bib17)\], and two proprietary models, DeepSeek\-V3\.2\[[8](https://arxiv.org/html/2609.27165#bib.bib20)\]and GPT\-4o\-mini\[[19](https://arxiv.org/html/2609.27165#bib.bib21)\], using one shared prompt per language for TEF, Direct, Majority Vote, and Soft Vote\. Each prompt specifies the dimension and its two positions and requests a single answer token\. We read the answer\-token log\-probabilities without sampling\. Open\-source models are served on RTX 3090 GPUs\. We setϵ=10−6\\epsilon=10^\{\-6\}andM=10M=10, and report dimension\-level accuracy and macro\-F1\.
Main Results\.In Table, TEF has the highest accuracy and macro\-F1 for every model in both languages, exceeding the strongest baseline by 4\.5 accuracy and 4\.6 macro\-F1 points on average, so the gain comes from the fusion rule rather than from model\-specific tuning\. At the dimension level, TEF is the best or tied for best in 116 of 120 comparisons and never trails the strongest baseline by more than 2\.6 points\. Segmentation alone does not ensure improvement, since equal\-weight voting, whether hard or soft, does not consistently improve on Direct\. The gains of TEF concentrate where the baselines are weakest, on the smaller open\-source models, on culture and politics in Chinese, and on environment in English, whereas on dimensions where Direct is already strong TEF matches it\.
Ablation Study\.Tableevaluates TEF by removing one component at a time\. On Qwen2\.5\-7B, both variants retain most of TEF’s average gain over Soft Vote; on DeepSeek\-V3\.2, however, both fall below Soft Vote in both languages on average\. Their complementarity reflects their distinct roles: the log\-odds transform emphasizes decisive evidence, while entropy weighting suppresses uncertain sentences\. The largest accuracy drop occurs on DeepSeek\-V3\.2’s English environment dimension, where removing entropy weighting or the log\-odds transform reduces accuracy by 22\.9 or 23\.8 percentage points\.
Figure 4:Calibration on Qwen2\.5\-7B\. TEF achieves the lowest expected calibration error and overconfidence gap across both Chinese and English settings\. Compared with Direct, Majority Vote, and Soft Vote, TEF aligns confidence with accuracy by tempering weak evidence\.Table 1:Robustness Analysis\. Average accuracy \(%\) over both languages under the original, verbose, and minimal prompts\. Strongest baseline is the best of Direct, MV, and SV per prompt; Ours is TEF\.Table 2:Efficiency Analysis\. Total tokens and cost over the twelve evaluation subsets and mean time to first token \(TTFT\) on RTX 3090 GPUs\. Voting covers MV and SV, which issue the same queries\.Analysis\.1\)*Calibration\.*Because measurement outputs are thresholded or manually reviewed, document\-level confidence should reflect correctness\. In Fig\.[4](https://arxiv.org/html/2609.27165#S5.F4), Direct and both voting rules are overconfident on Qwen2\.5\-7B, whereas TEF achieves the lowest expected calibration error\[[18](https://arxiv.org/html/2609.27165#bib.bib23),[13](https://arxiv.org/html/2609.27165#bib.bib22)\]and overconfidence gap\. Its confidence is therefore better suited to selective review\. 2\)*Robustness\.*To test whether the gain is due to the fusion rule rather than the shared prompt, we rerun all methods with a verbose prompt that adds role background and annotation guidelines and with a minimal prompt that retains only the classification directive\. In Table[1](https://arxiv.org/html/2609.27165#S5.T1), TEF remains best under every prompt and degrades the least under either perturbation\. 3\)*Efficiency\.*In Table[2](https://arxiv.org/html/2609.27165#S5.T2), TEF uses the same single\-token queries as the voting methods, so its token count is identical\. Its time to first token differs by only about two milliseconds\. Relative to Direct, sentence\-level methods cost roughly three times as much, but TEF is the only one that consistently turns this overhead into a performance gain\.
## 6Conclusion
We cast long\-text value measurement as the fusion of soft sentence decisions and derived TEF, which tempers each sentence’s log\-odds by its information gain\. On MIND, TEF is more accurate than direct prediction and voting for five LLMs in two languages and better calibrated on Qwen2\.5\-7B, and its gain survives changes of the prompt\.
## References
- \[1\]\(2025\)Beyond majority voting: LLM aggregation by leveraging higher\-order information\.arXiv preprint arXiv:2510\.01499\.Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[2\]A\. U\. Akash, A\. Fahmy, and A\. Trabelsi\(2025\)Can large language models address open\-target stance detection?\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 971–985\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.54)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[3\]P\. Bhattacharya, H\. Zhang, Y\. Cao, W\. Gao, B\. S\. Loh, J\. J\. P\. Simons, and L\. Z\. Wong\(2025\)Rethinking stance detection: a theoretically\-informed research agenda for user\-level inference using language models\.arXiv preprint arXiv:2502\.02074\.Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[4\]P\. G\. Bissiri, C\. C\. Holmes, and S\. G\. Walker\(2016\)A general framework for updating belief distributions\.Journal of the Royal Statistical Society Series B: Statistical Methodology78\(5\),pp\. 1103–1130\.External Links:[Document](https://dx.doi.org/10.1111/rssb.12158)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p4.1),[§3](https://arxiv.org/html/2609.27165#S3.p8.1)\.
- \[5\]Z\. Chair and P\. K\. Varshney\(1986\)Optimal data fusion in multiple sensor detection systems\.IEEE Transactions on Aerospace and Electronic SystemsAES\-22\(1\),pp\. 98–101\.External Links:[Document](https://dx.doi.org/10.1109/TAES.1986.310699)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[6\]M\. Conover, J\. Ratkiewicz, M\. Francisco, B\. Gonçalves, F\. Menczer, and A\. Flammini\(2011\)Political polarization on Twitter\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.5,Barcelona, Spain,pp\. 89–96\.External Links:[Document](https://dx.doi.org/10.1609/icwsm.v5i1.14126)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[7\]T\. M\. Cover and J\. A\. Thomas\(2005\)Elements of information theory\.2nd edition,Wiley\.External Links:[Document](https://dx.doi.org/10.1002/047174882X)Cited by:[§3](https://arxiv.org/html/2609.27165#S3.p5.1.2)\.
- \[8\]DeepSeek\-AI\(2025\)DeepSeek\-V3\.2: pushing the frontier of open large language models\.Note:arXiv:2512\.02556Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[9\]C\. Genest and J\. V\. Zidek\(1986\)Combining probability distributions: a critique and an annotated bibliography\.Statistical Science1\(1\),pp\. 114–135\.External Links:[Document](https://dx.doi.org/10.1214/ss/1177013825)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[10\]F\. Gilardi, M\. Alizadeh, and M\. Kubli\(2023\)ChatGPT outperforms crowd workers for text\-annotation tasks\.Proceedings of the National Academy of Sciences120\(30\),pp\. e2305016120\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2305016120)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[11\]A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian,et al\.\(2024\)The Llama 3 herd of models\.Note:arXiv:2407\.21783Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[12\]P\. Grünwald and T\. van Ommen\(2017\)Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it\.Bayesian Analysis12\(4\),pp\. 1069–1103\.External Links:[Document](https://dx.doi.org/10.1214/17-BA1085)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p4.1)\.
- \[13\]C\. Guo, G\. Pleiss, Y\. Sun, and K\. Q\. Weinberger\(2017\)On calibration of modern neural networks\.InProceedings of the 34th International Conference on Machine Learning,pp\. 1321–1330\.Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p4.1)\.
- \[14\]T\. Islam and D\. Goldwasser\(2025\)Uncovering latent arguments in social media messaging by employing LLMs\-in\-the\-loop strategy\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 7397–7429\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.413)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[15\]Z\. Jiang, J\. Araki, H\. Ding, and G\. Neubig\(2021\)How can we know when language models know? On the calibration of language models for question answering\.Transactions of the Association for Computational Linguistics9,pp\. 962–977\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00407)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[16\]S\. Kadavathet al\.\(2022\)Language models \(mostly\) know what they know\.arXiv preprint arXiv:2207\.05221\.Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[17\]B\. Lee, R\. Aiyappa, Y\. Ahn, H\. Kwak, and J\. An\(2025\)A semantic embedding space based on large language models for modelling human beliefs\.Nature Human Behaviour9\(9\),pp\. 1928–1940\.External Links:[Document](https://dx.doi.org/10.1038/s41562-025-02228-z)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[18\]M\. P\. Naeini, G\. F\. Cooper, and M\. Hauskrecht\(2015\)Obtaining well calibrated probabilities using Bayesian binning\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 2901–2907\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v29i1.9602)Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p4.1)\.
- \[19\]OpenAI\(2024\)GPT\-4o mini: advancing cost\-efficient intelligence\.Note:OpenAI blogCited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[20\]N\. Pangakis and S\. Wolken\(2025\)Keeping humans in the loop: human\-centered automated annotation with generative AI\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.19,pp\. 1471–1492\.External Links:[Document](https://dx.doi.org/10.1609/icwsm.v19i1.35883)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[21\]Qwen, A\. Yang, B\. Yang, B\. Zhang, B\. Hui,et al\.\(2024\)Qwen2\.5 technical report\.Note:arXiv:2412\.15115Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[22\]C\. Rago, P\. Willett, and Y\. Bar\-Shalom\(1996\)Censoring sensors: a low\-communication\-rate scheme for distributed detection\.IEEE Transactions on Aerospace and Electronic Systems32\(2\),pp\. 554–568\.External Links:[Document](https://dx.doi.org/10.1109/7.489500)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p4.1),[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[23\]R\. R\. Saha, L\. V\. S\. Lakshmanan, and R\. T\. Ng\(2024\)Stance detection with explanations\.Computational Linguistics50\(1\),pp\. 193–235\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00501)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[24\]O\. Shorinwa, Z\. Mei, J\. Lidard, A\. Z\. Ren, and A\. Majumdar\(2025\)A survey on uncertainty quantification of large language models: taxonomy, open research challenges, and future directions\.ACM Computing Surveys58\(3\),pp\. 1–38\.External Links:[Document](https://dx.doi.org/10.1145/3744238)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[25\]R\. R\. Tenney and N\. R\. Sandell\(1981\)Detection with distributed sensors\.IEEE Transactions on Aerospace and Electronic SystemsAES\-17\(4\),pp\. 501–510\.External Links:[Document](https://dx.doi.org/10.1109/TAES.1981.309178)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[26\]K\. Tianet al\.\(2023\)Just ask for calibration: strategies for eliciting calibrated confidence scores from language models fine\-tuned with human feedback\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 5433–5442\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.330)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[27\]P\. Törnberg\(2024\)Large language models outperform expert coders and supervised classifiers at annotating political social media messages\.Social Science Computer Review43\(6\),pp\. 1181–1195\.External Links:[Document](https://dx.doi.org/10.1177/08944393241286471)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[28\]D\. Tsirmpas, I\. Gkionis, G\. Th\. Papadopoulos, and I\. Mademlis\(2024\)Neural natural language processing for long texts: a survey on classification and summarization\.Engineering Applications of Artificial Intelligence133,pp\. 108231\.External Links:[Document](https://dx.doi.org/10.1016/j.engappai.2024.108231)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[29\]P\. K\. Varshney\(1997\)Distributed detection and data fusion\.Springer,New York, NY\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4612-1904-0)Cited by:[§3](https://arxiv.org/html/2609.27165#S3.p7.1.2)\.
- \[30\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui,et al\.\(2025\)Qwen3 technical report\.Note:arXiv:2505\.09388Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[31\]C\. Ziems, W\. Held, O\. Shaikh, J\. Chen, Z\. Zhang, and D\. Yang\(2024\)Can large language models transform computational social science?\.Computational Linguistics50\(1\),pp\. 237–291\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00502)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.Similar Articles
@dair_ai: // Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing wh…
The paper presents UPHELD, a benchmark with extensive human annotations for evaluating conversational LLMs, and a Mixture-of-Judges framework that enhances evaluation accuracy by 30%.
Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation
This paper tests whether persisting evidence generated in one call and using it as the exclusive input for a later verdict (evidence locking) improves or harms LLM-as-judge evaluation. Across 24,000 judgments, they find that evidence locking reduces agreement with human preferences by 4-6 points and increases answer-order inconsistency compared to structured one-call judging.
From Snippets to Semantics: Rethinking Evidence Granularity for Multilingual Fact Verification
This paper introduces SEEK, a framework for semantic evidence extraction in multilingual fact verification, which constructs coherent evidence chunks from full articles and fine-tunes multilingual LLMs with LoRA, achieving up to 20% improvement in macro-F1 over baselines.
A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs
Proposes a tree-of-thoughts inspired extractive-abstractive approach for legal case judgement summarization using LLMs, with experiments on DeepSeek and LLama showing improved summaries over extractive or abstractive methods alone.
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
This paper shows that LLM judges embedded in reasoning pipelines often make poor decisions, and proposes Evidence-Locked Derive–Gate–Repair (EL-DGR) to constrain judge overrides with evidence certificates, improving accuracy over majority vote and first-candidate baselines.