计数证据,而非句子:LLM判断的Tempered Evidence Fusion用于长文本价值测量

arXiv cs.CL 论文

摘要

本文提出了Tempered Evidence Fusion (TEF),一种用于长文本价值测量的方法,该方法通过信息增益对句子级别的LLM判断进行加权,并介绍了用于评估的MIND基准数据集。

arXiv:2609.27165v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.
查看原文
查看缓存全文

缓存时间: 2026/09/24 09:15

# Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement
Source: [https://arxiv.org/html/2609.27165](https://arxiv.org/html/2609.27165)
Yuhe Wu1,∗, Rui Qian2,∗, Guangyu Wang1,∗, Yuran Chen3, Yuanchao Zhu4, Junjie Yang5Zhengheng Li6, Jiulin Cai7, Tianyi Zhang8, Zihan Dong9, Jiaxin Liu1, Yujie Chen10, Guang Zhang1,†††thanks:\*Equal contribution†Corresponding author

###### Abstract

Large language models \(LLMs\) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance\-bearing sentences\. Existing approaches either ask the model to predict a document\-level label directly, which can be overconfident, or aggregate sentence\-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative\. We formulate long\-text value measurement as a decision\-fusion problem and propose Tempered Evidence Fusion \(TEF\), a training\-free rule that weights each sentence’s log\-odds by its normalized information gain, as derived from a generalized Bayesian posterior\. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes\-optimal weight of decisive evidence\. We further introduce Multi\-event Insight Network Dimensions \(MIND\), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions\. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4\.5 accuracy points and 4\.6 macro\-F1 points across five LLMs and two languages\. MIND dataset and code are available at[GitHub](https://github.com/Kzczc/ICASSP2027-TEF)\.

###### Index Terms:

Decision fusion, large language models, uncertainty, calibration, computational social science

††address:1HKUST\(GZ\)2FDU3DUFE4UESTC5UMD
6SEU7USTC8Independent9Georgia Tech10CUHK\(SZ\)## 1Introduction

Public value orientations are key latent variables in computational social science, capturing how people position themselves on issues such as economic regulation, climate policy, or cultural change\. Social media makes these orientations observable through text at scale\[[6](https://arxiv.org/html/2609.27165#bib.bib9),[17](https://arxiv.org/html/2609.27165#bib.bib10)\]\. Yet this scale, together with the subjectivity of human annotation, has motivated growing interest in using LLMs as measurement instruments\[[20](https://arxiv.org/html/2609.27165#bib.bib11),[14](https://arxiv.org/html/2609.27165#bib.bib12)\]\. Existing LLM\-based pipelines usually either ask the model to read the whole post and return a document\-level label\[[10](https://arxiv.org/html/2609.27165#bib.bib26),[31](https://arxiv.org/html/2609.27165#bib.bib27),[27](https://arxiv.org/html/2609.27165#bib.bib28)\], or split the post into sentences and aggregate sentence\-level predictions by voting\[[28](https://arxiv.org/html/2609.27165#bib.bib13)\]\.

However, we observe that little research has examined how sentence\-level LLM judgments should be fused when evidence is unevenly distributed across a long post\. In Fig\.[1](https://arxiv.org/html/2609.27165#S1.F1), long social media posts often mix background information, quotations, and concessions, with only a few stance\-bearing sentences\. In such cases, direct prediction can collapse heterogeneous evidence into one overconfident label, while majority or soft voting can allow uncertain sentences to outvote a few decisive ones\. This leads us to ask:*How can sentence\-level judgments be combined so that decisive evidence dominates the final decision while uncertain sentences are discounted?*

![Refer to caption](https://arxiv.org/html/2609.27165v1/teaser.png)Figure 1:Long\-text value measurement requires counting informative evidence rather than treating all sentences equally\.LEFT:Direct prediction collapses mixed content into one label, while voting can let many uncertain sentences outweigh a few decisive ones\.RIGHT:TEF weights sentence log\-odds by normalized information gain, discounting uncertain judgments so decisive evidence drives the document\-level orientation\.Bearing this in mind, we formulate long\-text value measurement as a decision\-fusion problem, where each sentence acts as a local detector that emits a soft decision\. Under conditional independence and calibrated sentence posteriors, the Bayes\-optimal document\-level rule sums sentence log\-likelihood ratios\. This observation places existing voting rules on a common scale: majority voting keeps only the sign of each sentence judgment, and soft voting uses a bounded surrogate that underweights highly decisive sentences\. At the same time, directly summing LLM logits is unreliable because next\-token probabilities are only approximately calibrated, especially near the point of maximal uncertainty\.

To this end, we presentTemperedEvidenceFusion \(TEF\), a training\-free fusion rule that tempers each sentence’s log\-odds by its normalized information gain, as derived from a generalized Bayesian posterior\[[4](https://arxiv.org/html/2609.27165#bib.bib4),[12](https://arxiv.org/html/2609.27165#bib.bib5)\]\. The resulting score nearly vanishes for uncertain sentences but preserves the Bayes\-optimal weight of decisive evidence, yielding a soft censoring effect similar in spirit to censoring rules in distributed detection\[[22](https://arxiv.org/html/2609.27165#bib.bib2)\]\. In Fig\.[1](https://arxiv.org/html/2609.27165#S1.F1)\(right\), TEF counts evidence rather than sentences, allowing a few informative sentences to determine the document\-level orientation\.

Our contributions are threefold\. \(i\) We cast long\-text value measurement with LLMs as a decision\-fusion problem and derive TEF from a tempered Bayesian posterior, showing that it suppresses uncertain sentences while preserving the weight of decisive evidence\. \(ii\) We introduce MIND, a five\-year, cross\-lingual benchmark of 8,358 posts over six value dimensions\. \(iii\) We show that TEF outperforms the strongest of the Direct, Majority Vote, and Soft Vote baselines by 4\.5 accuracy points and 4\.6 macro\-F1 points on average across five LLMs and two languages, yielding better\-calibrated estimates\.

## 2Related Work

LLM\-based stance and ideology measurement has expanded from short\-form texts such as tweets to longer, multi\-issue documents\[[23](https://arxiv.org/html/2609.27165#bib.bib14),[2](https://arxiv.org/html/2609.27165#bib.bib15),[3](https://arxiv.org/html/2609.27165#bib.bib16)\]\. In parallel, the reliability of LLM confidence has been studied from multiple perspectives, including token\-level probabilities\[[15](https://arxiv.org/html/2609.27165#bib.bib24),[16](https://arxiv.org/html/2609.27165#bib.bib6)\], verbalized confidence\[[26](https://arxiv.org/html/2609.27165#bib.bib25),[24](https://arxiv.org/html/2609.27165#bib.bib7)\], and aggregation methods beyond majority voting\[[1](https://arxiv.org/html/2609.27165#bib.bib8)\]\. In signal processing, fusing local decisions is a classical problem\[[25](https://arxiv.org/html/2609.27165#bib.bib31)\]: log\-linear opinion pools combine probabilistic judgments\[[9](https://arxiv.org/html/2609.27165#bib.bib3)\], the Chair–Varshney rule weights local decisions by their reliabilities\[[5](https://arxiv.org/html/2609.27165#bib.bib1)\], and censoring schemes transmit only observations with informative likelihood ratios\[[22](https://arxiv.org/html/2609.27165#bib.bib2)\]\. Despite these connections, long\-text measurement with LLMs still largely relies on equal\-weight voting\. In contrast, we formulate long\-text value measurement as sentence\-level decision fusion and derive an evidence\-aware fusion rule, TEF, that discounts uncertain sentences while preserving the weight of decisive ones\.

## 3Proposed TEF

Problem Formulation\.We formulate long\-text value measurement as a decision\-fusion problem\. The goal is to infer a latent social variable from a collection of textual segments whose evidential strength may vary substantially within the same document\. Let𝒳\\mathcal\{X\}denote the input space and let𝒴=\{c1,…,cK\}\\mathcal\{Y\}=\\\{c\_\{1\},\\ldots,c\_\{K\}\\\}be the set of target measurement categories\. For an observationX∈𝒳X\\in\\mathcal\{X\}, the latent measurement variableYYtakes values in𝒴\\mathcal\{Y\}\. We segment each long text intoNNsentence\-level units,

X↦𝒟⁡\(X\)=\{D1,D2,…,DN\},\\vskip\-2\.84544ptX\\mapsto\\mathcal\{D\}\(X\)=\\\{D\_\{1\},D\_\{2\},\\ldots,D\_\{N\}\\\},\(1\)whereDiD\_\{i\}denotes theii\-th segment\. This granularity preserves local semantic coherence while keeping inference computationally manageable\. For each segmentDiD\_\{i\}, a language modelℳ\\mathcal\{M\}returns a probability distribution over the measurement categories, denoted by𝐩i=ℳ⁡\(Di\)=\(pi\(1\),…,pi\(K\)\)\\mathbf\{p\}\_\{i\}=\\mathcal\{M\}\(D\_\{i\}\)=\\big\(p\_\{i\}^\{\(1\)\},\\ldots,p\_\{i\}^\{\(K\)\}\\big\), where each entrypi\(k\)=Pℳ​\(Y=ck∣Di\)p\_\{i\}^\{\(k\)\}=P\_\{\\mathcal\{M\}\}\(Y=c\_\{k\}\\mid D\_\{i\}\)and𝐩i∈ΔK−1\\mathbf\{p\}\_\{i\}\\in\\Delta^\{K\-1\}\. The distribution encodes both a local predictiony^i=arg⁡maxk⁡pi\(k\)\\hat\{y\}\_\{i\}=\\arg\\max\_\{k\}p\_\{i\}^\{\(k\)\}and its uncertainty, measured by entropyH\(𝐩i\)=−∑k=1Kpi\(k\)logpi\(k\)\.H\(\\mathbf\{p\}\_\{i\}\)=\-\\sum\_\{k=1\}^\{K\}p\_\{i\}^\{\(k\)\}\\log p\_\{i\}^\{\(k\)\}\.Common aggregation rules discard part of this probabilistic evidence\. Majority voting uses only the segment\-level hard labels, while soft voting averages probability vectors with equal weight,

Y^MV=mode⁡\(\{y^i\}i=1N\),Y^SV=arg⁡maxc∈𝒴​1N​∑i=1Npi\(c\)\.\\vskip\-5\.69046pt\\hat\{Y\}\_\{\\mathrm\{MV\}\}=\\mathrm\{mode\}\\big\(\\\{\\hat\{y\}\_\{i\}\\\}\_\{i=1\}^\{N\}\\big\),\\qquad\\hat\{Y\}\_\{\\mathrm\{SV\}\}=\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}p\_\{i\}^\{\(c\)\}\.\(2\)Both rules treat all segments as equally reliable\. This is problematic for scientific measurement from long texts, where decisive evidence may appear in only a few segments and many other segments may be ambiguous, rhetorical, or weakly relevant\. TEF instead aggregates reliability\-weighted evidence\. Letwi=ω⁡\(𝐩i\)∈\[0,1\]w\_\{i\}=\\omega\(\\mathbf\{p\}\_\{i\}\)\\in\[0,1\]be the reliability assigned to segmentDiD\_\{i\}, and letg:\[0,1\]→ℝg:\[0,1\]\\to\\mathbb\{R\}be a scoring transformation\. The document\-level measurement isY^=arg⁡max⁡∑i=1Nc∈𝒴⁡wi​g​\(pi\(c\)\)\.\\hat\{Y\}=\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\sum\_\{i=1\}^\{N\}w\_\{i\}\\,g\\big\(p\_\{i\}^\{\(c\)\}\\big\)\.The following paragraphs specifyω\\omegaandgg, which together define the fusion rule of TEF\.

Reliability Weighting\.TEF instantiates the preceding framework with entropy\-based reliability weighting and one\-vs\-rest log\-odds evidence\. Algorithm[1](https://arxiv.org/html/2609.27165#alg1)gives the complete procedure\. We use normalized information gain as the segment reliability,wi=1−H⁡\(𝐩i\)log⁡K∈\[0,1\]\.w\_\{i\}=1\-\\frac\{H\(\\mathbf\{p\}\_\{i\}\)\}\{\\log K\}\\in\[0,1\]\.This weight is close to one for concentrated distributions and close to zero for near\-uniform distributions\. Thus confident segments contribute strongly, while uncertain segments are softly censored\.

Evidence Aggregation\.For each categorycc, TEF transforms the segment posterior into a one\-vs\-rest log\-odds score,g⁡\(pi\(c\)\)=log⁡pi\(c\)1−pi\(c\)\.g\\big\(p\_\{i\}^\{\(c\)\}\\big\)=\\log\\frac\{p\_\{i\}^\{\(c\)\}\}\{1\-p\_\{i\}^\{\(c\)\}\}\.The aggregated evidence for categoryccis

Score⁡\(c\)=∑i=1Nwi​log⁡pi\(c\)1−pi\(c\)\.\\operatorname\{Score\}\(c\)=\\sum\_\{i=1\}^\{N\}w\_\{i\}\\log\\frac\{p\_\{i\}^\{\(c\)\}\}\{1\-p\_\{i\}^\{\(c\)\}\}\.\\vskip\-8\.5359pt\(3\)The final measurement outcome isY^=arg⁡maxc∈𝒴​Score⁡\(c\)\\hat\{Y\}=\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\operatorname\{Score\}\(c\)\. In implementation, probabilities are clipped to\[ϵ,1−ϵ\]\[\\epsilon,1\-\\epsilon\]before applying the logit, and log\-odds values are clipped to\[−M,M\]\[\-M,M\]\. Unless otherwise stated, we setϵ=10−6\\epsilon=10^\{\-6\}andM=10M=10\. Unlike majority voting, TEF does not reduce𝐩i\\mathbf\{p\}\_\{i\}to a hard label before aggregation\. Unlike soft voting, TEF does not assume that all segments are equally informative\. It therefore preserves distributional information while reducing the impact of noisy or ambiguous segments\.

Algorithm 1Tempered Evidence Fusion1:Text

XX, model

ℳ\\mathcal\{M\}, categories

𝒴=\{c1,…,cK\}\\mathcal\{Y\}=\\\{c\_\{1\},\\ldots,c\_\{K\}\\\}, clip bound

MM, stability constant

ϵ\\epsilon\.

2:Measurement outcome

Y^\\hat\{Y\}
3:

\{D1,…,DN\}←𝒟⁡\(X\)\\\{D\_\{1\},\\ldots,D\_\{N\}\\\}\\leftarrow\\mathcal\{D\}\(X\)⊳\\trianglerightSegment into sentences

4:Initialize

Score⁡\(c\)←0\\operatorname\{Score\}\(c\)\\leftarrow 0for all

c∈𝒴c\\in\\mathcal\{Y\}
5:for

i=1i=1to

NNdo

6:

𝐩i←ℳ⁡\(Di\)\\mathbf\{p\}\_\{i\}\\leftarrow\\mathcal\{M\}\(D\_\{i\}\)⊳\\trianglerightSegment\-level distribution

7:

Hi←−∑k=1Kpi\(k\)logpi\(k\)H\_\{i\}\\leftarrow\-\\sum\_\{k=1\}^\{K\}p\_\{i\}^\{\(k\)\}\\log p\_\{i\}^\{\(k\)\}⊳\\trianglerightEntropy

8:

wi←1−Hi/log⁡Kw\_\{i\}\\leftarrow 1\-H\_\{i\}/\\log K⊳\\trianglerightReliability weight

9:for

c∈𝒴c\\in\\mathcal\{Y\}do

10:

p~←clip⁡\(pi\(c\),ϵ,1−ϵ\)\\tilde\{p\}\\leftarrow\\mathrm\{clip\}\(p\_\{i\}^\{\(c\)\},\\epsilon,1\-\\epsilon\)
11:

ℓ←clip⁡\(log⁡p~1−p~,−M,M\)\\ell\\leftarrow\\mathrm\{clip\}\\\!\\left\(\\log\\frac\{\\tilde\{p\}\}\{1\-\\tilde\{p\}\},\-M,M\\right\)
12:

Score⁡\(c\)←Score⁡\(c\)\+wi​ℓ\\operatorname\{Score\}\(c\)\\leftarrow\\operatorname\{Score\}\(c\)\+w\_\{i\}\\ell
13:endfor

14:endfor

15:

Y^←arg⁡maxc∈𝒴​Score⁡\(c\)\\hat\{Y\}\\leftarrow\\arg\\max\_\{c\\in\\mathcal\{Y\}\}\\operatorname\{Score\}\(c\)
16:return

Y^\\hat\{Y\}

Theoretical Analysis\.

Proposition 1 \(Information Preservation\)\.*Since the hard labely^i=arg⁡maxk⁡pi\(k\)\\hat\{y\}\_\{i\}=\\arg\\max\_\{k\}p\_\{i\}^\{\(k\)\}is a deterministic function of𝐩i\\mathbf\{p\}\_\{i\}, the data processing inequality\[[7](https://arxiv.org/html/2609.27165#bib.bib29)\]givesI⁡\(Y,𝐩i\)≥I⁡\(Y,y^i\)I\(Y;\\mathbf\{p\}\_\{i\}\)\\geq I\(Y;\\hat\{y\}\_\{i\}\)\. Thus, hard voting can discard reliability information even when two segments support the same category with different confidence\.*

TEF follows this principle by retaining the full distribution𝐩i\\mathbf\{p\}\_\{i\}instead of reducing each segment toy^i\\hat\{y\}\_\{i\}\. Its entropy\-based weightwiw\_\{i\}distinguishes decisive evidence from ambiguous evidence, so two segments with the same hard label can contribute different amounts\. The method therefore preserves the information discarded by majority voting while producing an additive document\-level score\.

Proposition 2 \(Tempered Bayesian Pooling\)\.*For a binary contrast with equal priors, calibrated conditionally independent segment posteriors satisfy the additive Bayes log\-odds identity\[[29](https://arxiv.org/html/2609.27165#bib.bib30)\]logP\(Y=c∣D1:N\)P\(Y=c¯∣D1:N\)=∑i=1NlogP⁡\(Y=c∣Di\)P⁡\(Y=c¯∣Di\)\\log\\frac\{P\(Y=c\\mid D\_\{1:N\}\)\}\{P\(Y=\\bar\{c\}\\mid D\_\{1:N\}\)\}=\\sum\_\{i=1\}^\{N\}\\log\\frac\{P\(Y=c\\mid D\_\{i\}\)\}\{P\(Y=\\bar\{c\}\\mid D\_\{i\}\)\}\. Therefore, evidence is naturally accumulated on the log\-odds scale rather than through hard counts or raw probability averaging\.*

TEF uses the one\-vs\-rest transformg⁡\(p\)=log⁡p1−pg\(p\)=\\log\\frac\{p\}\{1\-p\}to implement this additive evidence scale\. Because LLM posteriors are approximate, it tempers each log\-odds contribution\[[4](https://arxiv.org/html/2609.27165#bib.bib4)\]withwi=\(log⁡K−H⁡\(𝐩i\)\)/log⁡Kw\_\{i\}=\(\\log K\-H\(\\mathbf\{p\}\_\{i\}\)\)/\\log K, the normalized information gain from a uniform prior\. Near\-uniform segments therefore have little influence, while low\-entropy stance\-bearing segments retain a contribution close to the Bayesian additive statistic\. This gives TEF its soft\-censoring behavior, as uncertain sentences do not accumulate by sheer number, while decisive sentences can determine the final document\-level measurement outcome even when they are few\.

![Refer to caption](https://arxiv.org/html/2609.27165v1/Fig2_Final.png)Figure 2:Construction pipeline of MIND\. Posts are collected from public sources, filtered for relevance, anonymized, and deduplicated, then annotated in three rounds for dimension, subcategory, and orientation before being balanced and refined by LLM agents\.Figure 3:Statistics of MIND\. Radial bars report the number of labels for each value orientation, with colors denoting dimensions and paired bars representing the opposing orientations of each subcategory\. Panels \(a\) and \(b\) correspond to Chinese and English posts\.
## 4The MIND Benchmark

We construct the MIND dataset, a longitudinal, cross\-lingual benchmark for measuring public value orientations\. The final corpus contains 8,358 posts, including 5,474 Chinese posts and 2,884 English posts over six value dimensions in Fig\.[3](https://arxiv.org/html/2609.27165#S3.F3)\. MIND is built through a three\-stage pipeline\.Stage I: Data Collection\.We collect candidate posts from X, Weibo, Reddit, Zhihu, discussion forums, and news feeds from January 2020 to January 2025, covering major public events such as the Russia–Ukraine war and the rise of generative AI\. For each value dimension, keyword search is anchored to high\-search\-volume events\. Raw posts are filtered for relevance and opinion content, anonymized, de\-duplicated, and screened by an automatic quality score\.Stage II: Data Annotation\.We adopt a three\-round annotation framework to label the value dimension, subcategory, and orientation in sequence\. Each round uses three annotators and accepts a label only when at least two agree\. Failed or disputed cases enter arbitration, samples without a stable majority are discarded, and accepted samples pass distribution and quality checks before the evolution stage\.Stage III: Agent\-Driven Data Evolution\.To address class imbalance and semantic ambiguity, we introduce an agent\-driven evolution stage\. When a category is underrepresented, a balancing agent generates synthetic samples through generative oversampling\. For low\-confidence or ambiguous samples, a critic agent evaluates the alignment between text and label, and a semantic\-edit agent revises the text to make the labeled orientation explicit\. Only verified samples are retained in the final benchmark\.

## 5Experiments

Experiment Settings\.We evaluate three open\-source models, Qwen2\.5\-7B\[[21](https://arxiv.org/html/2609.27165#bib.bib18)\], LLaMA3\-8B\[[11](https://arxiv.org/html/2609.27165#bib.bib19)\], and Qwen3\-14B\[[30](https://arxiv.org/html/2609.27165#bib.bib17)\], and two proprietary models, DeepSeek\-V3\.2\[[8](https://arxiv.org/html/2609.27165#bib.bib20)\]and GPT\-4o\-mini\[[19](https://arxiv.org/html/2609.27165#bib.bib21)\], using one shared prompt per language for TEF, Direct, Majority Vote, and Soft Vote\. Each prompt specifies the dimension and its two positions and requests a single answer token\. We read the answer\-token log\-probabilities without sampling\. Open\-source models are served on RTX 3090 GPUs\. We setϵ=10−6\\epsilon=10^\{\-6\}andM=10M=10, and report dimension\-level accuracy and macro\-F1\.

Main Results\.In Table, TEF has the highest accuracy and macro\-F1 for every model in both languages, exceeding the strongest baseline by 4\.5 accuracy and 4\.6 macro\-F1 points on average, so the gain comes from the fusion rule rather than from model\-specific tuning\. At the dimension level, TEF is the best or tied for best in 116 of 120 comparisons and never trails the strongest baseline by more than 2\.6 points\. Segmentation alone does not ensure improvement, since equal\-weight voting, whether hard or soft, does not consistently improve on Direct\. The gains of TEF concentrate where the baselines are weakest, on the smaller open\-source models, on culture and politics in Chinese, and on environment in English, whereas on dimensions where Direct is already strong TEF matches it\.

Ablation Study\.Tableevaluates TEF by removing one component at a time\. On Qwen2\.5\-7B, both variants retain most of TEF’s average gain over Soft Vote; on DeepSeek\-V3\.2, however, both fall below Soft Vote in both languages on average\. Their complementarity reflects their distinct roles: the log\-odds transform emphasizes decisive evidence, while entropy weighting suppresses uncertain sentences\. The largest accuracy drop occurs on DeepSeek\-V3\.2’s English environment dimension, where removing entropy weighting or the log\-odds transform reduces accuracy by 22\.9 or 23\.8 percentage points\.

Figure 4:Calibration on Qwen2\.5\-7B\. TEF achieves the lowest expected calibration error and overconfidence gap across both Chinese and English settings\. Compared with Direct, Majority Vote, and Soft Vote, TEF aligns confidence with accuracy by tempering weak evidence\.Table 1:Robustness Analysis\. Average accuracy \(%\) over both languages under the original, verbose, and minimal prompts\. Strongest baseline is the best of Direct, MV, and SV per prompt; Ours is TEF\.Table 2:Efficiency Analysis\. Total tokens and cost over the twelve evaluation subsets and mean time to first token \(TTFT\) on RTX 3090 GPUs\. Voting covers MV and SV, which issue the same queries\.Analysis\.1\)*Calibration\.*Because measurement outputs are thresholded or manually reviewed, document\-level confidence should reflect correctness\. In Fig\.[4](https://arxiv.org/html/2609.27165#S5.F4), Direct and both voting rules are overconfident on Qwen2\.5\-7B, whereas TEF achieves the lowest expected calibration error\[[18](https://arxiv.org/html/2609.27165#bib.bib23),[13](https://arxiv.org/html/2609.27165#bib.bib22)\]and overconfidence gap\. Its confidence is therefore better suited to selective review\. 2\)*Robustness\.*To test whether the gain is due to the fusion rule rather than the shared prompt, we rerun all methods with a verbose prompt that adds role background and annotation guidelines and with a minimal prompt that retains only the classification directive\. In Table[1](https://arxiv.org/html/2609.27165#S5.T1), TEF remains best under every prompt and degrades the least under either perturbation\. 3\)*Efficiency\.*In Table[2](https://arxiv.org/html/2609.27165#S5.T2), TEF uses the same single\-token queries as the voting methods, so its token count is identical\. Its time to first token differs by only about two milliseconds\. Relative to Direct, sentence\-level methods cost roughly three times as much, but TEF is the only one that consistently turns this overhead into a performance gain\.

## 6Conclusion

We cast long\-text value measurement as the fusion of soft sentence decisions and derived TEF, which tempers each sentence’s log\-odds by its information gain\. On MIND, TEF is more accurate than direct prediction and voting for five LLMs in two languages and better calibrated on Qwen2\.5\-7B, and its gain survives changes of the prompt\.

## References

- \[1\]\(2025\)Beyond majority voting: LLM aggregation by leveraging higher\-order information\.arXiv preprint arXiv:2510\.01499\.Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[2\]A\. U\. Akash, A\. Fahmy, and A\. Trabelsi\(2025\)Can large language models address open\-target stance detection?\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 971–985\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.54)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[3\]P\. Bhattacharya, H\. Zhang, Y\. Cao, W\. Gao, B\. S\. Loh, J\. J\. P\. Simons, and L\. Z\. Wong\(2025\)Rethinking stance detection: a theoretically\-informed research agenda for user\-level inference using language models\.arXiv preprint arXiv:2502\.02074\.Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[4\]P\. G\. Bissiri, C\. C\. Holmes, and S\. G\. Walker\(2016\)A general framework for updating belief distributions\.Journal of the Royal Statistical Society Series B: Statistical Methodology78\(5\),pp\. 1103–1130\.External Links:[Document](https://dx.doi.org/10.1111/rssb.12158)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p4.1),[§3](https://arxiv.org/html/2609.27165#S3.p8.1)\.
- \[5\]Z\. Chair and P\. K\. Varshney\(1986\)Optimal data fusion in multiple sensor detection systems\.IEEE Transactions on Aerospace and Electronic SystemsAES\-22\(1\),pp\. 98–101\.External Links:[Document](https://dx.doi.org/10.1109/TAES.1986.310699)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[6\]M\. Conover, J\. Ratkiewicz, M\. Francisco, B\. Gonçalves, F\. Menczer, and A\. Flammini\(2011\)Political polarization on Twitter\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.5,Barcelona, Spain,pp\. 89–96\.External Links:[Document](https://dx.doi.org/10.1609/icwsm.v5i1.14126)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[7\]T\. M\. Cover and J\. A\. Thomas\(2005\)Elements of information theory\.2nd edition,Wiley\.External Links:[Document](https://dx.doi.org/10.1002/047174882X)Cited by:[§3](https://arxiv.org/html/2609.27165#S3.p5.1.2)\.
- \[8\]DeepSeek\-AI\(2025\)DeepSeek\-V3\.2: pushing the frontier of open large language models\.Note:arXiv:2512\.02556Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[9\]C\. Genest and J\. V\. Zidek\(1986\)Combining probability distributions: a critique and an annotated bibliography\.Statistical Science1\(1\),pp\. 114–135\.External Links:[Document](https://dx.doi.org/10.1214/ss/1177013825)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[10\]F\. Gilardi, M\. Alizadeh, and M\. Kubli\(2023\)ChatGPT outperforms crowd workers for text\-annotation tasks\.Proceedings of the National Academy of Sciences120\(30\),pp\. e2305016120\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2305016120)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[11\]A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian,et al\.\(2024\)The Llama 3 herd of models\.Note:arXiv:2407\.21783Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[12\]P\. Grünwald and T\. van Ommen\(2017\)Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it\.Bayesian Analysis12\(4\),pp\. 1069–1103\.External Links:[Document](https://dx.doi.org/10.1214/17-BA1085)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p4.1)\.
- \[13\]C\. Guo, G\. Pleiss, Y\. Sun, and K\. Q\. Weinberger\(2017\)On calibration of modern neural networks\.InProceedings of the 34th International Conference on Machine Learning,pp\. 1321–1330\.Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p4.1)\.
- \[14\]T\. Islam and D\. Goldwasser\(2025\)Uncovering latent arguments in social media messaging by employing LLMs\-in\-the\-loop strategy\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 7397–7429\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.413)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[15\]Z\. Jiang, J\. Araki, H\. Ding, and G\. Neubig\(2021\)How can we know when language models know? On the calibration of language models for question answering\.Transactions of the Association for Computational Linguistics9,pp\. 962–977\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00407)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[16\]S\. Kadavathet al\.\(2022\)Language models \(mostly\) know what they know\.arXiv preprint arXiv:2207\.05221\.Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[17\]B\. Lee, R\. Aiyappa, Y\. Ahn, H\. Kwak, and J\. An\(2025\)A semantic embedding space based on large language models for modelling human beliefs\.Nature Human Behaviour9\(9\),pp\. 1928–1940\.External Links:[Document](https://dx.doi.org/10.1038/s41562-025-02228-z)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[18\]M\. P\. Naeini, G\. F\. Cooper, and M\. Hauskrecht\(2015\)Obtaining well calibrated probabilities using Bayesian binning\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 2901–2907\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v29i1.9602)Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p4.1)\.
- \[19\]OpenAI\(2024\)GPT\-4o mini: advancing cost\-efficient intelligence\.Note:OpenAI blogCited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[20\]N\. Pangakis and S\. Wolken\(2025\)Keeping humans in the loop: human\-centered automated annotation with generative AI\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.19,pp\. 1471–1492\.External Links:[Document](https://dx.doi.org/10.1609/icwsm.v19i1.35883)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[21\]Qwen, A\. Yang, B\. Yang, B\. Zhang, B\. Hui,et al\.\(2024\)Qwen2\.5 technical report\.Note:arXiv:2412\.15115Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[22\]C\. Rago, P\. Willett, and Y\. Bar\-Shalom\(1996\)Censoring sensors: a low\-communication\-rate scheme for distributed detection\.IEEE Transactions on Aerospace and Electronic Systems32\(2\),pp\. 554–568\.External Links:[Document](https://dx.doi.org/10.1109/7.489500)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p4.1),[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[23\]R\. R\. Saha, L\. V\. S\. Lakshmanan, and R\. T\. Ng\(2024\)Stance detection with explanations\.Computational Linguistics50\(1\),pp\. 193–235\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00501)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[24\]O\. Shorinwa, Z\. Mei, J\. Lidard, A\. Z\. Ren, and A\. Majumdar\(2025\)A survey on uncertainty quantification of large language models: taxonomy, open research challenges, and future directions\.ACM Computing Surveys58\(3\),pp\. 1–38\.External Links:[Document](https://dx.doi.org/10.1145/3744238)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[25\]R\. R\. Tenney and N\. R\. Sandell\(1981\)Detection with distributed sensors\.IEEE Transactions on Aerospace and Electronic SystemsAES\-17\(4\),pp\. 501–510\.External Links:[Document](https://dx.doi.org/10.1109/TAES.1981.309178)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[26\]K\. Tianet al\.\(2023\)Just ask for calibration: strategies for eliciting calibrated confidence scores from language models fine\-tuned with human feedback\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 5433–5442\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.330)Cited by:[§2](https://arxiv.org/html/2609.27165#S2.p1.1)\.
- \[27\]P\. Törnberg\(2024\)Large language models outperform expert coders and supervised classifiers at annotating political social media messages\.Social Science Computer Review43\(6\),pp\. 1181–1195\.External Links:[Document](https://dx.doi.org/10.1177/08944393241286471)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[28\]D\. Tsirmpas, I\. Gkionis, G\. Th\. Papadopoulos, and I\. Mademlis\(2024\)Neural natural language processing for long texts: a survey on classification and summarization\.Engineering Applications of Artificial Intelligence133,pp\. 108231\.External Links:[Document](https://dx.doi.org/10.1016/j.engappai.2024.108231)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.
- \[29\]P\. K\. Varshney\(1997\)Distributed detection and data fusion\.Springer,New York, NY\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4612-1904-0)Cited by:[§3](https://arxiv.org/html/2609.27165#S3.p7.1.2)\.
- \[30\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui,et al\.\(2025\)Qwen3 technical report\.Note:arXiv:2505\.09388Cited by:[§5](https://arxiv.org/html/2609.27165#S5.p1.1)\.
- \[31\]C\. Ziems, W\. Held, O\. Shaikh, J\. Chen, Z\. Zhang, and D\. Yang\(2024\)Can large language models transform computational social science?\.Computational Linguistics50\(1\),pp\. 237–291\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00502)Cited by:[§1](https://arxiv.org/html/2609.27165#S1.p1.1)\.

相似文章

锁定证据再做评判:冻结接口降低LLM-as-Judge评估质量

arXiv cs.CL

本文测试了在一次调用中生成证据并将其作为后续裁决的唯一输入(证据锁定)是否会改善或损害LLM-as-Judge评估。在24,000次判断中,他们发现,与结构化单次调用评判相比,证据锁定将人类偏好一致性降低了4到6个百分点,并增加了答案顺序不一致性。