Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
摘要
CUE-Bench is a Chinese benchmark for affective stance that models explicit-implicit polarity interaction and provides intent and fine-grained emotion annotations, showing gains in emotion recognition and pragmatic intent detection.
arXiv:2608.10810v1 Announce Type: new
Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotate surface polarity or final emotion categories, while lacking a structured account of how explicit expression, implicit affect, pragmatic intent, and fine grained emotion interact. This limitation makes current evaluations insensitive to cases where affective meaning is concealed, weakened, inverted, or pragmatically reshaped, thereby obscuring model failures in deeper emotion understanding. To address this gap, we introduce CUE Bench, a Chinese Unsaid Emotion benchmark that centers on Affective Stance and covers diverse communicative scenarios. CUE Bench constructs nine human interpretable affective stances from explicit implicit polarity interaction and further provides intent and fine grained emotion annotations for structured affective inference. Experiments show that incorporating Affective Stance improves fine grained emotion recognition by 3.5 percentage points and pragmatic intent detection by 7.8 percentage points over strong baselines.
查看缓存全文
缓存时间: 2026/08/12 08:38
# Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
Source: [https://arxiv.org/html/2608.10810](https://arxiv.org/html/2608.10810)
Zhenyan Zheng11footnotemark:1,Yunyao Zhang,Junxi Sheng,Junqing Yu,Zikai Song Huazhong University of Science and Technology \{ZhenyanZheng, ikostar, skyesong\}@hust\.edu\.cn
###### Abstract
Emotion understanding in discourse requires reasoning beyond surface sentiment, since speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions\. Existing emotion benchmarks mainly annotate surface polarity or final emotion categories, while lacking a structured account of how explicit expression, implicit affect, pragmatic intent, and fine\-grained emotion interact\. This limitation makes current evaluations insensitive to cases where affective meaning is concealed, weakened, inverted, or pragmatically reshaped, thereby obscuring models’ failures in deeper emotion understanding\. To address this gap, we introduceCUE\-Bench, aChineseUnsaidEmotion benchmark that centers onAffective Stanceand covers diverse communicative scenarios\. CUE\-Bench constructs nine human\-interpretable affective stances from Explicit\-Implicit polarity interaction and further provides intent and fine\-grained emotion annotations for structured affective inference\. Experiments show that incorporating Affective Stance improves fine\-grained emotion recognition by3\.5percentage points and pragmatic intent detection by7\.8percentage points over strong baselines\.
Surfacing the Unsaid: CUE\-Bench for Affective Stance in Chinese Discourse
Zhenyan Zheng11footnotemark:1, Yunyao Zhang††thanks:Equal contribution\., Junxi Sheng, Junqing Yu, Zikai Song††thanks:Corresponding authorHuazhong University of Science and Technology\{ZhenyanZheng, ikostar, skyesong\}@hust\.edu\.cn
Figure 1:Overview ofCUE\-Bench\. The benchmark linkswhat is saidandwhat is meantthrough Affective Stance, enabling structured affective inference in Chinese discourse\.## 1Introduction
Understanding emotion in language is a central problem for affective computing and human\-centered NLPPicard \([1997](https://arxiv.org/html/2608.10810#bib.bib31)\); Songet al\.\([2026](https://arxiv.org/html/2608.10810#bib.bib50)\)\. However, real discourse often requires more than recognizing surface sentiment or predicting an isolated emotion categoryZhanget al\.\([2026b](https://arxiv.org/html/2608.10810#bib.bib45),[c](https://arxiv.org/html/2608.10810#bib.bib46)\)\. It requires modeling how explicit expression and implicit affective tendency jointly shape the speaker’s communicative orientation, which we define asAffective StanceDu Bois \([2007](https://arxiv.org/html/2608.10810#bib.bib22)\)\. For example, positive wording may imply criticism, neutral wording may conceal commitment, and negative wording may signal affiliation or support\. Such explicit\-implicit mismatches are common in Chinese discourse, where affect is shaped by politeness, suppression, irony, understatement, and other indirect strategiesBrown and Levinson \([1987](https://arxiv.org/html/2608.10810#bib.bib32)\); Grice \([1975](https://arxiv.org/html/2608.10810#bib.bib33)\); Gross \([1998](https://arxiv.org/html/2608.10810#bib.bib34)\)\. Without modeling this explicit\-implicit relation, existing evaluations may capture either explicit expression or implicit affect in isolation, but miss how their mismatch shapes the speaker’s affective stance\.
Existing emotion benchmarks have advanced affective understanding from three perspectives\. Representative resources include implicit emotion recognition benchmarks such as IESTKlingeret al\.\([2018](https://arxiv.org/html/2608.10810#bib.bib11)\), SMP2020\-EWECTXianweiet al\.\([2021](https://arxiv.org/html/2608.10810#bib.bib14)\), and ResEmoHuet al\.\([2024](https://arxiv.org/html/2608.10810#bib.bib13)\), intent understanding benchmarks such as DailyDialogLiet al\.\([2017](https://arxiv.org/html/2608.10810#bib.bib7)\), MELDPoriaet al\.\([2019](https://arxiv.org/html/2608.10810#bib.bib16)\), and CPEDChenet al\.\([2022](https://arxiv.org/html/2608.10810#bib.bib18)\), and fine\-grained emotion benchmarks such as GoEmotionsDemszkyet al\.\([2020](https://arxiv.org/html/2608.10810#bib.bib9)\)and CMMAZhanget al\.\([2023](https://arxiv.org/html/2608.10810#bib.bib19)\)\. However, most of them either focus on explicit affect or implicit affect separately, and their evaluation settings are usually restricted to a single task or communicative domain\. Consequently, they provide limited diagnostic power for assessing whether models can infer affective stance, pragmatic intent, and fine\-grained emotion when expressed affect and unsaid affect are misaligned\.
To address this gap, we introduceCUE\-Bench, aChineseUnsaidEmotion benchmark for affective understanding\. To operationalize affective stance, we propose theExplicit\-Implicit Stance Matrix, a structured framework that explicitly models the interaction between explicit expression and implicit affective tendency\. By connecting what is expressed with what remains unsaid, this framework provides a unified perspective for analyzing affective stance, pragmatic intent, and fine\-grained emotion\. Built on this framework, CUE\-Bench contains51,823annotated instances with four levels of supervision: explicit and implicit affective layers, nine Affective Stances, eight pragmatic intents, and twenty\-five fine\-grained emotions\. It covers diverse Chinese discourse scenarios beyond dialogue, including open\-domain conversation, social media comments, sarcasm\-oriented text, customer\-service interactions, and question\-answering contentZhanget al\.\([2025](https://arxiv.org/html/2608.10810#bib.bib48)\)\.
CUE\-Bench supportsAffective Stance Recognition,Pragmatic Intent Understanding, andFine\-grained Emotion Classification, spanning explicit\-implicit stance interaction, pragmatic motivation, and final emotion interpretation\. It is built via a hybrid pipeline of dual\-model agreement, human adjudication, and bias\-controlled LLM adjudication, using LLMs as a constrained aid rather than a replacement for human validationGilardiet al\.\([2023](https://arxiv.org/html/2608.10810#bib.bib25)\); Zhenget al\.\([2023](https://arxiv.org/html/2608.10810#bib.bib26)\)\.
In summary, our contributions are:
- ∙\\bulletWe introduceCUE\-Bench, a Chinese discourse benchmark that provides a unified setting for unsaid emotion understanding through three connected tasks:Affective Stance Recognition,Pragmatic Intent Understanding, andFine\-grained Emotion Classification\.
- ∙\\bulletWe propose theExplicit\-Implicit Stance Matrix, a structured framework for modeling affective stance, pragmatic intent, and fine\-grained emotion\.
BenchmarkLang\.ScaleMulti\-Exp\.Imp\.StanceIntentFine\-grainedMulti\-taskdomainaffectaffectemotionDailyDialogLiet al\.\([2017](https://arxiv.org/html/2608.10810#bib.bib7)\)En\.13\.1k✗✓✗✗✓✗✓IESTKlingeret al\.\([2018](https://arxiv.org/html/2608.10810#bib.bib11)\)En\.191\.7k✗✗✓✗✗✗✗MELDPoriaet al\.\([2019](https://arxiv.org/html/2608.10810#bib.bib16)\)En\.13\.0k✗✓✗✗✗✗✗GoEmotionsDemszkyet al\.\([2020](https://arxiv.org/html/2608.10810#bib.bib9)\)En\.58\.0k✗✓✗✗✗✓✗SMP2020\-EWECTXianweiet al\.\([2021](https://arxiv.org/html/2608.10810#bib.bib14)\)Zh\.13\.6k✗✓✗✗✗✗✗CPEDChenet al\.\([2022](https://arxiv.org/html/2608.10810#bib.bib18)\)Zh\.12k✗✓✗✗✓✓✓CMMAZhanget al\.\([2023](https://arxiv.org/html/2608.10810#bib.bib19)\)Zh\.21\.8k✗✓✗✗✓✗✓ResEmoHuet al\.\([2024](https://arxiv.org/html/2608.10810#bib.bib13)\)Zh\.72\.5k✗✓✗✗✗✓✓CUE\-BenchZh\.51\.8k✓✓✓✓✓✓✓
Table 1:Comparison with representative emotion\-related benchmarks\. Scale is reported at the utterance/comment/instance level when available; for CPED and DailyDialog, we report dialogue\-level scale following the original paper\. Exp\. and Imp\. affect denote explicit and implicit affective supervision\. Stance denotes whether the benchmark provides Explicit\-Implicit affective stance labels\. Multi\-domain indicates coverage of diverse communicative scenarios beyond a single dialogue or social\-media domain\.
## 2Related Work
### Emotion Benchmarks
Prior emotion\-related benchmarks cover three related but largely separate lines of work\.\(1\) Implicit emotion recognition\.IESTKlingeret al\.\([2018](https://arxiv.org/html/2608.10810#bib.bib11)\)and Chinese resources such as SMP2020\-EWECTXianweiet al\.\([2021](https://arxiv.org/html/2608.10810#bib.bib14)\)and ResEmoHuet al\.\([2024](https://arxiv.org/html/2608.10810#bib.bib13)\)study emotions that are not directly expressed through emotion words\. Yet they mainly formulate implicitness as emotion prediction, rather than as a structured relation between explicit expression and implicit affective tendency\.\(2\) Pragmatic intent understanding\.DailyDialogLiet al\.\([2017](https://arxiv.org/html/2608.10810#bib.bib7)\), DiplomatLiet al\.\([2023](https://arxiv.org/html/2608.10810#bib.bib23)\), and PUBSravanthiet al\.\([2024](https://arxiv.org/html/2608.10810#bib.bib24)\)provide resources for dialogue acts, situated pragmatic reasoning, or pragmatic capability evaluation\. However, they do not explicitly model how pragmatic intent interacts with explicit and implicit affect to reshape the affective meaning of an utterance\.\(3\) Fine\-grained emotion classification\.GoEmotionsDemszkyet al\.\([2020](https://arxiv.org/html/2608.10810#bib.bib9)\)and CMMAZhanget al\.\([2023](https://arxiv.org/html/2608.10810#bib.bib19)\)offer rich emotion or multi\-affection labels, but still largely treat affective states as independent categories\. Overall, existing benchmarks either evaluate explicit or implicit affect in isolation, or focus on a single task or communicative domain\. They therefore provide limited diagnostic power for jointly assessing affective stance, pragmatic intent, and fine\-grained emotion when what is expressed and what remains unsaid diverge\.
### Affective Modeling Methods
Existing methods also address these three aspects separately\.\(1\) Implicit emotion recognition\.Early methodsChronopoulouet al\.\([2018](https://arxiv.org/html/2608.10810#bib.bib35)\); Devlinet al\.\([2019](https://arxiv.org/html/2608.10810#bib.bib29)\); Cuiet al\.\([2021](https://arxiv.org/html/2608.10810#bib.bib30)\)rely on neural transfer learning, recurrent encoders, attention mechanisms, and pre\-trained language models to infer emotions that are not explicitly expressedLiet al\.\([2026c](https://arxiv.org/html/2608.10810#bib.bib53)\); Chenet al\.\([2026](https://arxiv.org/html/2608.10810#bib.bib59)\); Liet al\.\([2026d](https://arxiv.org/html/2608.10810#bib.bib60)\); however, they usually predict implicit affect directly without modeling its relation to explicit expression\.\(2\) Pragmatic intent understanding\.Intent\-oriented methodsChenet al\.\([2019](https://arxiv.org/html/2608.10810#bib.bib36)\); Caiet al\.\([2022](https://arxiv.org/html/2608.10810#bib.bib37)\); Ghosalet al\.\([2020](https://arxiv.org/html/2608.10810#bib.bib38)\); Zhanget al\.\([2026e](https://arxiv.org/html/2608.10810#bib.bib42),[a](https://arxiv.org/html/2608.10810#bib.bib44)\)model communicative goals through pre\-trained encoders, joint intent\-slot learning, dialogue context, or commonsense\-enhanced conversational reasoningLiet al\.\([2026b](https://arxiv.org/html/2608.10810#bib.bib52),[f](https://arxiv.org/html/2608.10810#bib.bib55),[2025](https://arxiv.org/html/2608.10810#bib.bib56)\), but they often treat intent as an independent target rather than explaining how it reshapes affective meaning\.\(3\) Fine\-grained emotion classification\.Fine\-grained emotion methodsDemszkyet al\.\([2020](https://arxiv.org/html/2608.10810#bib.bib9)\); Zhanget al\.\([2023](https://arxiv.org/html/2608.10810#bib.bib19)\); Devlinet al\.\([2019](https://arxiv.org/html/2608.10810#bib.bib29)\)typically map text into rich emotion taxonomies with supervised classifiers, contextual encoders, or multi\-label Transformer modelsChenet al\.\([2025](https://arxiv.org/html/2608.10810#bib.bib57)\); Liet al\.\([2026e](https://arxiv.org/html/2608.10810#bib.bib66)\), while leaving the intermediate links among explicit affect, implicit affect, and pragmatic intent underexplored\. In contrast, ourExplicit\-Implicit Stance Matrixprovides a structured intermediate framework that connects explicit expression, implicit affective tendency, pragmatic intent, and fine\-grained emotion across the three tasks\.
Figure 2:Overview ofCUE\-Bench\. The benchmark collects context–target utterance pairs from diverse Chinese dialogue scenarios and models deeper affective understanding through theExplicit–Implicit Stance Matrix\. The matrix contrasts the explicit affective signaleie\_\{i\}, i\.e\., what is expressed on the surface, with the implicit affective signalhih\_\{i\}, i\.e\., what remains unsaid\. Guided byMatrix\-Guided CoT, the reasoning pipeline progressively infers affective stance, pragmatic intent, and fine\-grained emotion, moving from surface expression to intended meaning\.
## 3CUE\-Bench
### 3\.1Design Principles
CUE\-Bench is designed to evaluate affective understanding beyond surface emotion classification, following three principles:
1. 1\.Context\-sensitive affective inference\.Each instance includes contextual information, since implicit affect often cannot be inferred from an isolated utterance or text span alone\.
2. 2\.Explicit\-Implicit stance modeling\.The benchmark emphasizes how explicit expression and implicit affective tendency jointly shape the speaker’s affective stance, covering realistic Chinese discourse phenomena such as politeness, suppression, irony, understatement, teasing, and indirect refusal\.
3. 3\.Unified multi\-task evaluation\.The annotation schema supports three connected tasks:Affective Stance Recognition,Pragmatic Intent Understanding, andFine\-grained Emotion Classification, enabling evaluation from stance interaction to intent reasoning and final emotion interpretation\.
### 3\.2Data Collection and Instance Format
We construct CUE\-Bench from diverse Chinese discourse sources, including open\-domain conversations, social media comments, sarcasm\-oriented text, customer\-service interactions, and question\-answering content\. Candidate instances are retained when contextual information plausibly changes affective interpretation, such as when the target text contains explicit affective markers, follows emotionally charged context, involves politeness or irony, or suggests a mismatch between literal wording and implicit affective tendency\.
Each instance is represented as:
xi=\(Ci,ui\),x\_\{i\}=\(C\_\{i\},u\_\{i\}\),whereCiC\_\{i\}denotes the discourse context, anduiu\_\{i\}denotes the target utterance or text span\.
The annotation target is defined as:
yi=\(yiexp,yiimp,yistance,yiintent,yiemotion\),y\_\{i\}=\(y\_\{i\}^\{\\mathrm\{exp\}\},y\_\{i\}^\{\\mathrm\{imp\}\},y\_\{i\}^\{\\mathrm\{stance\}\},y\_\{i\}^\{\\mathrm\{intent\}\},y\_\{i\}^\{\\mathrm\{emotion\}\}\),whereyiexpy\_\{i\}^\{\\mathrm\{exp\}\}andyiimpy\_\{i\}^\{\\mathrm\{imp\}\}denote explicit and implicit affective layers,yistancey\_\{i\}^\{\\mathrm\{stance\}\}denotes the Affective Stance label,yiintenty\_\{i\}^\{\\mathrm\{intent\}\}denotes pragmatic intent, andyiemotiony\_\{i\}^\{\\mathrm\{emotion\}\}denotes fine\-grained emotion\.
Before annotation, we normalize whitespace, remove duplicated instances, and discard samples that require private or external knowledge for interpretation\. Personally identifiable information is masked, and sensitive content is retained only when it is necessary for affective interpretation and can be safely anonymized\.
### 3\.3Annotation Protocols
Our annotation strategy combines model\-assisted candidate generation, human adjudication, and bias\-controlled LLM adjudication\. The goal is to obtain scalable annotations while preserving human verification for ambiguous Explicit\-Implicit affective relations\.
Model\-assisted Candidate Annotation\.We first use two independent models,M1M\_\{1\}andM2M\_\{2\}, to generate candidate annotations for each instance:
y^i\(1\)=M1\(xi\),y^i\(2\)=M2\(xi\)\.\\hat\{y\}\_\{i\}^\{\(1\)\}=M\_\{1\}\(x\_\{i\}\),\\qquad\\hat\{y\}\_\{i\}^\{\(2\)\}=M\_\{2\}\(x\_\{i\}\)\.Consistent outputs are retained as high\-confidence candidates:
𝒟agr=\{xi∣y^i\(1\)=y^i\(2\)\},\\mathcal\{D\}\_\{agr\}=\\\{x\_\{i\}\\mid\\hat\{y\}\_\{i\}^\{\(1\)\}=\\hat\{y\}\_\{i\}^\{\(2\)\}\\\},while inconsistent outputs are treated as ambiguous cases for further adjudication:
𝒟dis=\{xi∣y^i\(1\)≠y^i\(2\)\}\.\\mathcal\{D\}\_\{dis\}=\\\{x\_\{i\}\\mid\\hat\{y\}\_\{i\}^\{\(1\)\}\\neq\\hat\{y\}\_\{i\}^\{\(2\)\}\\\}\.
Human Adjudication and Gold Verification\.For a subset of𝒟dis\\mathcal\{D\}\_\{dis\}, trained annotators verify the candidate annotations by selecting the more appropriate label or revising the annotation when neither candidate is adequate\. This produces a human\-verified subset that serves as gold supervision and calibration data for later quality control\.
Bias\-controlled LLM Adjudication\.For the remaining ambiguous cases, we use LLMs as constrained adjudicators rather than unconstrained annotators\. Each case is adjudicated twice with the candidate order reversed:
bifwd=A\(xi,y^i\(1\),y^i\(2\)\),birev=A\(xi,y^i\(2\),y^i\(1\)\),b\_\{i\}^\{\\mathrm\{fwd\}\}=A\(x\_\{i\},\\hat\{y\}\_\{i\}^\{\(1\)\},\\hat\{y\}\_\{i\}^\{\(2\)\}\),b\_\{i\}^\{\\mathrm\{rev\}\}=A\(x\_\{i\},\\hat\{y\}\_\{i\}^\{\(2\)\},\\hat\{y\}\_\{i\}^\{\(1\)\}\),wherebifwd,birev∈\{1,2\}b\_\{i\}^\{\\mathrm\{fwd\}\},b\_\{i\}^\{\\mathrm\{rev\}\}\\in\\\{1,2\\\}indicate the selected candidate under each order\. We map the reversed decision back to the canonical order asb¯irev=3−birev\\bar\{b\}\_\{i\}^\{\\mathrm\{rev\}\}=3\-b\_\{i\}^\{\\mathrm\{rev\}\}\. The retained annotation is:
y~i=\{y^i\(bifwd\),ifbifwd=b¯irev,∅,otherwise\.\\tilde\{y\}\_\{i\}=\\begin\{cases\}\\hat\{y\}\_\{i\}^\{\(b\_\{i\}^\{\\mathrm\{fwd\}\}\)\},&\\text\{if \}b\_\{i\}^\{\\mathrm\{fwd\}\}=\\bar\{b\}\_\{i\}^\{\\mathrm\{rev\}\},\\\\ \\varnothing,&\\text\{otherwise\}\.\\end\{cases\}Only instances withy~i≠∅\\tilde\{y\}\_\{i\}\\neq\\varnothingare added to the retained LLM\-adjudicated subset\. This consistency check reduces positional bias and filters out unstable adjudication cases\.
MetricAffectiveStancePragmaticIntentFine\-grainedEmotion\# Classes9825Krippendorff’sα\\alpha0\.51970\.33880\.3146Majority Agr\.0\.79330\.61330\.4400Avg\.κ\\kappa0\.54900\.39220\.3369Cond\. Avg\.κ\\kappa–0\.78940\.6689Table 2:Inter\-annotator agreement on 300 expert re\-annotated instances\. Cond\. Avg\.κ\\kappadenotes average Cohen’s kappa computed on instances with consistent Affective Stance\. Detailed construction statistics and pairwise agreement results are reported in Appendix[B\.3](https://arxiv.org/html/2608.10810#A2.SS3)\.
### 3\.4Statistics and Analysis
Dataset Composition\.CUE\-Bench is constructed from 60,000 candidate instances through a hybrid pipeline of model agreement, human verification, and bias\-controlled LLM adjudication\. The final benchmark contains 51,823 retained instances, consisting of high\-confidence model\-agreement annotations, a human\-verified gold subset, and order\-consistent LLM\-adjudicated annotations\. Detailed construction statistics are reported in Appendix[B\.3](https://arxiv.org/html/2608.10810#A2.SS3)\.
Adjudication Validation and Noise Estimate\.We validate LLM adjudication on the human\-verified gold subset, where the adjudicator achieves 89% accuracy when selecting between two model\-generated candidate annotations\. After forward–reverse consistency filtering, we estimate that approximately 1,600 retained instances may remain uncertain, yielding an overall estimated contamination rate of 3\.1% in the final dataset\. This indicates that constrained LLM adjudication introduces limited residual noise while substantially improving dataset coverage\.
Inter\-Annotator Agreement\.To evaluate annotation reliability, we sample 300 instances from the human\-verified gold subset and ask three expert annotators to independently re\-annotate them\. Table[2](https://arxiv.org/html/2608.10810#S3.T2)reports agreement at the three annotation layers\. Affective Stance obtains the strongest agreement, with a Krippendorff’sα\\alphaof 0\.5197, a majority agreement rate of 79\.3%, and an average Cohen’sκ\\kappaof 0\.5490\. Pragmatic Intent and Fine\-grained Emotion show lower raw agreement, withα\\alphascores of 0\.3388 and 0\.3146, respectively, reflecting the higher subjectivity of latent intent and fine\-grained affective inference\. This pattern is consistent with prior findings that fixed IRR thresholds can be overly rigid for subjective annotation tasks, and that fine\-grained emotion annotation often yields moderate or low chance\-corrected agreement scoresWonget al\.\([2021](https://arxiv.org/html/2608.10810#bib.bib27)\); Demszkyet al\.\([2020](https://arxiv.org/html/2608.10810#bib.bib9)\)\. We therefore report both raw and conditional agreement\. When conditioned on consistent Affective Stance, agreement improves substantially: the conditional averageκ\\kappareaches 0\.7894 for Pragmatic Intent and 0\.6689 for Fine\-grained Emotion\. This suggests that Affective Stance provides a useful intermediate structure for localizing annotation ambiguity and reducing downstream uncertainty in intent and emotion labeling\. The annotators’ conflicts are not uniform across categories; the detailed confusion matrix is shown in Figure[3](https://arxiv.org/html/2608.10810#S3.F3)\.
Figure 3:IAA disagreement matrix among three annotators\. Darker cells indicate stronger disagreement between annotator label assignments\.
## 4Explicit\-Implicit Stance Matrix
The core of CUE\-Bench is theExplicit\-Implicit Stance Matrix, a structured framework for modeling how expressed affect and unsaid affect jointly shape the speaker’s affective stance\. Rather than treating explicit and implicit affect as two independent labels, the matrix defines Affective Stance as their compositional relation\.
### 4\.1Explicit and Implicit Affective Signals
For each instancexi=\(Ci,ui\)x\_\{i\}=\(C\_\{i\},u\_\{i\}\), we distinguish two affective signals\. Theexplicit affective signaldescribes the affect directly expressed by the target textuiu\_\{i\}, such as affective words, intensifiers, punctuation, praise, complaint, apology, thanks, or direct evaluation\. Theimplicit affective tendencydescribes the affect inferred from the discourse contextCiC\_\{i\}together with the target textuiu\_\{i\}, including pragmatic force, speaker intention, and contextual implication\.
Letπ\(⋅\)\\pi\(\\cdot\)denote an orientation projection that maps an affective layer onto a three\-way affective orientation space\. Specifically, we define
ei=π\(yiexp\),hi=π\(yiimp\),e\_\{i\}=\\pi\(y\_\{i\}^\{\\mathrm\{exp\}\}\),\\qquad h\_\{i\}=\\pi\(y\_\{i\}^\{\\mathrm\{imp\}\}\),whereei,hi∈𝒪e\_\{i\},h\_\{i\}\\in\\mathcal\{O\}and𝒪=\{\+,0,−\}\\mathcal\{O\}=\\\{\+,0,\-\\\}\. The symbols\+\+,0, and−\-correspond to positive, neutral, and negative affective orientations, respectively\. In this formulation,eie\_\{i\}encodes the explicit affective signal anchored in the surface expression ofuiu\_\{i\}, whilehih\_\{i\}encodes the latent affective tendency inferred from the utterance together with its context\(Ci,ui\)\(C\_\{i\},u\_\{i\}\)\.
### 4\.2Affective Stance
We define Affective Stance as:
si=ϕ\(ei,hi\),si=yistance,s\_\{i\}=\\phi\(e\_\{i\},h\_\{i\}\),\\qquad s\_\{i\}=y\_\{i\}^\{stance\},whereϕ:𝒪×𝒪→𝒮\\phi:\\mathcal\{O\}\\times\\mathcal\{O\}\\rightarrow\\mathcal\{S\}maps each Explicit\-Implicit pair to one of nine stance categories\.
Table[2](https://arxiv.org/html/2608.10810#S2.F2)shows the resulting matrix, where rows correspond to the explicit signaleie\_\{i\}and columns correspond to the implicit tendencyhih\_\{i\}\. This formulation makes Affective Stance a structured intermediate representation rather than an independent free\-form label: if eithereie\_\{i\}orhih\_\{i\}changes, the stance labelsis\_\{i\}changes accordingly\. The matrix therefore provides a transparent mechanism for capturing alignment, neutralization, concealment, and reversal between what is expressed and what remains unsaid\. Detailed definitions and examples of the nine Affective Stances are provided in Appendix[D](https://arxiv.org/html/2608.10810#A4)\.
### 4\.3Matrix\-Guided Chain\-of\-Thought
The Explicit\-Implicit Stance Matrix further provides a structured chain\-of\-thought for affective inference\. Rather than treating the three benchmark tasks as independent predictions, we use the matrix to impose an explicit reasoning order from surface expression to latent meaning and then to downstream affective interpretation\. This turns affective understanding into a progressive inference process: identify what is expressed, infer what remains unsaid, resolve their stance relation, interpret the speaker’s pragmatic motivation, and finally determine the fine\-grained emotion\.
For each instancexix\_\{i\}, the matrix\-guided reasoning path is:
\(e^i,h^i\)\\displaystyle\(\\hat\{e\}\_\{i\},\\hat\{h\}\_\{i\}\)=Fsig\(xi\),\\displaystyle=F\_\{\\mathrm\{sig\}\}\(x\_\{i\}\),s^i\\displaystyle\\hat\{s\}\_\{i\}=ϕ\(e^i,h^i\),\\displaystyle=\\phi\(\\hat\{e\}\_\{i\},\\hat\{h\}\_\{i\}\),y^iintent\\displaystyle\\hat\{y\}\_\{i\}^\{\\mathrm\{intent\}\}=Fprag\(xi,e^i,h^i,s^i\),\\displaystyle=F\_\{\\mathrm\{prag\}\}\(x\_\{i\},\\hat\{e\}\_\{i\},\\hat\{h\}\_\{i\},\\hat\{s\}\_\{i\}\),y^iemotion\\displaystyle\\hat\{y\}\_\{i\}^\{\\mathrm\{emotion\}\}=Femo\(xi,e^i,h^i,s^i,y^iintent\)\.\\displaystyle=F\_\{\\mathrm\{emo\}\}\(x\_\{i\},\\hat\{e\}\_\{i\},\\hat\{h\}\_\{i\},\\hat\{s\}\_\{i\},\\hat\{y\}\_\{i\}^\{\\mathrm\{intent\}\}\)\.Here,e^i\\hat\{e\}\_\{i\}andh^i\\hat\{h\}\_\{i\}denote the predicted explicit and implicit affective orientations,s^i=ϕ\(e^i,h^i\)\\hat\{s\}\_\{i\}=\\phi\(\\hat\{e\}\_\{i\},\\hat\{h\}\_\{i\}\)denotes the predicted Affective Stance,y^iintent\\hat\{y\}\_\{i\}^\{\\mathrm\{intent\}\}denotes the predicted pragmatic intent, andy^iemotion\\hat\{y\}\_\{i\}^\{\\mathrm\{emotion\}\}denotes the predicted fine\-grained emotion\.
This chain gives each prediction a structured dependency: stance is inferred from the Explicit\-Implicit relation, intent is interpreted under the resulting stance, and fine\-grained emotion is decided with both stance and intent as intermediate evidence\. In LLM evaluation, we instantiate this process as a normalized prompting protocol, requiring the model to output intermediate fields in the order of explicit signal, implicit tendency, Affective Stance, pragmatic intent, and fine\-grained emotion\. Compared with direct label prediction, this matrix\-guided CoT exposes the model’s reasoning path and enables error analysis at each level of affective inference\.
ModelMethodAffective Stance↑\\uparrowPragmatic Intent↑\\uparrowFine\-grained Emotion↑\\uparrowAvg\.↑\\uparrowAcc\.F1W\-F1Acc\.F1W\-F1Acc\.F1W\-F1![[Uncaptioned image]](https://arxiv.org/html/2608.10810v1/Figure/icon1-deepseek-color.png)DeepSeek V4\-FlashDirect0\.4920\.4080\.5000\.3250\.2770\.3300\.2470\.1700\.2480\.333Few\-shot0\.4800\.4120\.4980\.3390\.2880\.3440\.2470\.1740\.2460\.336CoT0\.4980\.4250\.5160\.3320\.2770\.3430\.2510\.1820\.2530\.342Ours0\.5100\.4660\.5270\.4590\.3610\.4650\.3130\.1940\.3270\.402Δ\\Delta\+0\.012\+0\.041\+0\.011\+0\.120\+0\.073\+0\.121\+0\.062\+0\.012\+0\.074\+0\.061![[Uncaptioned image]](https://arxiv.org/html/2608.10810v1/Figure/icon2-openai.png)GPT 4o\-miniDirect0\.3170\.2500\.2920\.2810\.2630\.2850\.1150\.0890\.1120\.223Few\-shot0\.3410\.2810\.3360\.3020\.2740\.2980\.2260\.1530\.2290\.271CoT0\.4290\.3430\.4310\.2810\.2620\.2970\.2260\.1500\.2260\.294Ours0\.4580\.3360\.4440\.3900\.2640\.3830\.2750\.1470\.2660\.329Δ\\Delta\+0\.029\-0\.007\+0\.013\+0\.088\-0\.010\+0\.085\+0\.049\-0\.006\+0\.037\+0\.035![[Uncaptioned image]](https://arxiv.org/html/2608.10810v1/Figure/icon3-meta-color.png)LLaMA 4 MaverickDirect0\.3380\.2700\.3260\.2730\.2450\.2730\.2370\.1620\.2320\.262Few\-shot0\.3670\.3150\.3850\.2930\.2710\.3070\.2330\.1690\.2330\.286CoT0\.4540\.3610\.4630\.2480\.2160\.2560\.2530\.1670\.2470\.296Ours0\.5120\.4100\.5110\.4530\.3420\.4520\.3390\.1830\.3310\.393Δ\\Delta\+0\.058\+0\.049\+0\.048\+0\.160\+0\.071\+0\.145\+0\.086\+0\.014\+0\.084\+0\.096![[Uncaptioned image]](https://arxiv.org/html/2608.10810v1/Figure/icon3-meta-color.png)LLaMA 3\.1 8BDirect0\.2030\.1940\.2210\.2320\.2080\.2330\.1420\.0860\.1510\.186Few\-shot0\.2810\.2560\.3070\.2590\.2310\.2770\.1670\.1120\.1660\.228CoT0\.2770\.2560\.3110\.2260\.2160\.2480\.1290\.1010\.1350\.211Ours0\.3620\.2930\.3690\.2850\.2360\.2920\.1900\.0970\.1840\.256Δ\\Delta\+0\.081\+0\.037\+0\.058\+0\.026\+0\.005\+0\.015\+0\.023\-0\.015\+0\.018\+0\.027![[Uncaptioned image]](https://arxiv.org/html/2608.10810v1/Figure/icon4-qwen.png)Qwen 3\-8BDirect0\.1800\.1540\.1720\.1960\.2100\.2420\.2090\.1280\.2250\.191Few\-shot0\.1810\.1830\.1970\.2310\.2140\.2270\.2120\.1210\.2210\.199CoT0\.4220\.3250\.4300\.2540\.2260\.2570\.2120\.1320\.2060\.274Ours0\.4240\.3090\.4120\.3640\.2710\.3670\.2600\.1370\.2620\.312Δ\\Delta\+0\.002\-0\.016\-0\.018\+0\.110\+0\.045\+0\.110\+0\.048\+0\.005\+0\.037\+0\.038Table 3:Main results\.We report Accuracy, macro\-F1, and weighted\-F1 on three CUE\-Bench tasks: Affective Stance, Pragmatic Intent, and Fine\-grained Emotion\.↑\\uparrowindicates that higher values are better\.
## 5Experiments
### 5\.1Settings
#### Models\.
We evaluate a diverse set of large language models covering both proprietary and open\-source families\. \(1\) For proprietary models, we include GPT\-4o\-mini and DeepSeek\-V4\-Flash, which represent widely used lightweight instruction\-following models with strong general reasoning ability\. \(2\) For open\-source models, we evaluate LLaMA\-4\-Maverick, LLaMA\-3\.1\-8B, and Qwen\-series models, including Qwen\-3\-8B\.
#### LLM prompting baselines\.
We compare our matrix\-guided reasoning method with three standard LLM prompting baselines: \(1\)Direct prompting, which asks the model to directly output the predicted label from the context and target utterance; \(2\)Few\-shot prompting, which provides annotated demonstrations before prediction; and \(3\)CoT prompting, which elicits free\-form intermediate reasoning before the final prediction\.
#### Metrics\.
We reportAccuracy \(Acc\.\),macro\-F1 \(F1\), andweighted\-F1 \(W\-F1\)\. Accuracy measures overall correctness, macro\-F1 gives equal weight to each class and reflects performance on minority categories, while weighted\-F1 accounts for label imbalance by weighting class\-wise F1 scores by class frequency\.
### 5\.2Main Results
As shown in Table[3](https://arxiv.org/html/2608.10810#S4.T3), we draw three observations\.
Our matrix\-guided method achieves the best overall performance\.Our method obtains the highest average score across all evaluated models, outperforming the strongest baseline by\+0\.027\+0\.027to\+0\.096\+0\.096\. The gains are especially clear on DeepSeek\-V4\-Flash and LLaMA\-4\-Maverick, showing the effectiveness of jointly modeling surface expression and hidden affective tendency\.
Pragmatic intent shows the most consistent gains\.Our method improves Pragmatic Intent accuracy across all models, with gains from\+0\.026\+0\.026to\+0\.160\+0\.160, and also consistently improves weighted\-F1\. This indicates that communicative intent benefits strongly from the explicit–implicit affective distinction\.
Fine\-grained emotion improves but remains harder\.Our method improves Fine\-grained Emotion accuracy and weighted\-F1 across all models, while macro\-F1 gains are less stable\. This suggests that stance\-guided reasoning helps infer emotional tendency, but distinguishing fine\-grained emotion categories remains challenging\.
ModelTaskAcc\.F1W\-F1DeepSeek\-V4\-FlashI0\.7910\.7030\.790II0\.3920\.2430\.379III0\.4320\.2190\.389LLaMA\-4\-MaverickI0\.7390\.6580\.745II0\.4300\.2510\.436III0\.4890\.2540\.466Table 4:Ablations on CUE\-Bench\. Task I: Affective Stance→\\rightarrowPragmatic Intent; Task II: Affective Stance→\\rightarrowFine\-grained Emotion; Task III: Affective Stance \+ Pragmatic Intent→\\rightarrowFine\-grained Emotion\.
### 5\.3Ablations
To examine whether the proposed reasoning path benefits downstream affective inference, we conduct oracle\-conditioning ablations with DeepSeek\-V4\-Flash and LLaMA\-4\-Maverick\. As shown in Table[4](https://arxiv.org/html/2608.10810#S5.T4), we test whether gold Affective Stance and Pragmatic Intent can serve as intermediate evidence for predicting Pragmatic Intent and Fine\-grained Emotion\. We draw three observations\.
Affective Stance provides strong evidence for pragmatic intent\.With gold Affective Stance, Pragmatic Intent prediction reaches 0\.703 macro\-F1 on DeepSeek\-V4\-Flash and 0\.658 on LLaMA\-4\-Maverick, confirming its value as an intermediate representation\.
Pragmatic Intent adds complementary evidence for emotion inference\.Adding gold Pragmatic Intent on top of gold Affective Stance improves Fine\-grained Emotion accuracy and weighted\-F1 for both models, showing that intent contributes information beyond stance alone\.
Oracle signals help but do not replace category\-level emotion discrimination\.Although gold intermediate labels improve Fine\-grained Emotion performance, the gains are not uniform across all metrics\. This suggests that the reasoning chain provides useful affective evidence, while final emotion prediction still requires direct discrimination among emotion categories\.
Figure 4:Wide light bars show audited support, narrow dark bars show problem\-case counts, and the red line shows the problem\-case rate\.
### 5\.4Analysis
We focus on two research questions that explain the distributional and annotation patterns observed in the CUE\-Bench\.
#### RQ1: Why are negative and indirect cases frequent?
CUE\-Bench contains many negative or negative\-leaning cases because we intentionally retain sources rich in implicit affect, such as sarcastic, hostile, conflictual, and emotionally charged online discourse\. This increases the density of cases where literal wording is insufficient, making the benchmark more diagnostic for unsaid affect\.
Veiled negative cases dominate the benchmark\.Veiled Negativeaccounts for 22\.3% of the data, where neutral surface wording often implies dissatisfaction, reluctance, pressure, or indirect criticism\. These cases are challenging because models must recover hidden negative affect from context rather than explicit lexical cues\.
Sarcastic negative cases reflect deliberate stress\-test design\.Sarcastic Negativealso appears frequently \(10\.9%\) because the corpus includes sarcasm\-oriented and hostile\-comment data\. Thus, the negative skew should be viewed as a benchmark feature rather than a natural base\-rate estimate: CUE\-Bench is designed to evaluate affective inference under pragmatic mismatch\.
#### RQ2: Where does human–AI annotation friction arise?
We analyze a 1,500\-instance audit sample from the model\-disagreement pool, where GPT adjudication was run in both candidate orders and compared with human review\. A case is marked as problematic if the two adjudication orders are inconsistent, or if a consistent GPT decision disagrees with the human\-accepted label\.
Veiled negative cases are the main source of friction\.As shown in Figure[4](https://arxiv.org/html/2608.10810#S5.F4), problematic cases concentrate heavily inVeiled Negative: 458 cases, covering 62% of auditedVeiled Negativeinstances\. This suggests that neutral\-looking utterances with negative implication are especially difficult for AI adjudication\.
Friction appears when surface affect and implied affect diverge\.Sarcastic NegativeandUnderstated Positiveoften require pragmatic reversal or concealment, whileFormulaic Positiverequires distinguishing genuine positivity from scripted politeness\. These patterns show that human review is most valuable for stance categories with strong explicit–implicit mismatch\.
## 6Conclusion
We introduce CUE\-Bench, a Chinese benchmark for unsaid emotion understanding centered on the Explicit\-Implicit Stance Matrix\. By decomposing affective meaning into explicit signal, implicit tendency, Affective Stance, Pragmatic Intent, and Fine\-grained Emotion, CUE\-Bench makes the path from what is said to what is meant directly evaluable\. Experiments show that matrix\-guided prompting improves overall performance, while oracle\-conditioning ablations confirm the value of Affective Stance as an intermediate representation\. CUE\-Bench provides both a benchmark and an analysis framework for studying polite, suppressed, ironic, indirect, and otherwise unsaid affect in Chinese discourse\.
## 7Limitations
CUE\-Bench has three main limitations\.\(1\) Residual annotation noise\.Although our dual\-model annotation pipeline is calibrated and verified with human\-labeled data, the final benchmark may still contain unavoidable residual noise\. In particular, LLM adjudication is used only as a constrained component with consistency filtering, but some uncertain or contaminated instances may remain\.\(2\) Coarse explicit–implicit orientation space\.The three\-way explicit/implicit orientation space makes Affective Stance interpretable and easy to operationalize, but it inevitably abstracts away finer affective distinctions that must be recovered at the Fine\-grained Emotion stage\.\(3\) Long\-tailed and culturally situated labels\.CUE\-Bench is long\-tailed: categories such asReportive Negative,Empathy, and several rare emotions have limited support, so macro\-F1 and weighted\-F1 should be read together\. Moreover, implicit affect and pragmatic intent remain partly subjective and culturally situated, even with detailed guidelines and adjudication\. Future extensions can add richer conversational metadata, multimodal signals, and multilingual comparisons while preserving the explicit–implicit stance structure\.
## References
- P\. Brown and S\. C\. Levinson \(1987\)Politeness: some universals in language usage\.Vol\.4,Cambridge university press\.Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- F\. Cai, W\. Zhou, F\. Mi, and B\. Faltings \(2022\)Slim: explicit slot\-intent mapping with bert for joint multi\-intent detection and slot filling\.pp\. 7607–7611\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- CCAC 2024 Chinese Sarcasm Calculation Organizers \(2024\)Chinese sarcasm calculation evaluation task at CCAC 2024\.Note:[https://github\.com/pjzj220113/chinese\-sarcasm\-calculation](https://github.com/pjzj220113/chinese-sarcasm-calculation)Evaluation task dataset and instructionsCited by:[§B\.2](https://arxiv.org/html/2608.10810#A2.SS2.p1.1)\.
- M\. Chen, R\. Liu, L\. Shen, S\. Yuan, J\. Zhou, Y\. Wu, X\. He, and B\. Zhou \(2020\)The JDDC corpus: a large\-scale multi\-turn Chinese dialogue dataset for E\-commerce customer service\.InProceedings of the Twelfth Language Resources and Evaluation Conference,Marseille, France,pp\. 459–466\.External Links:[Link](https://aclanthology.org/2020.lrec-1.58/)Cited by:[§B\.2](https://arxiv.org/html/2608.10810#A2.SS2.p1.1)\.
- Q\. Chen, Z\. Zhuo, and W\. Wang \(2019\)BERT for joint intent classification and slot filling\.arXiv preprint arXiv:1902\.10909\.External Links:[Link](https://arxiv.org/abs/1902.10909)Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Y\. Chen, W\. Fan, X\. Xing, J\. Pang, M\. Huang, W\. Han, Q\. Tie, and X\. Xu \(2022\)CPED: a large\-scale chinese personalized and emotional dialogue dataset for conversational ai\.External Links:2205\.14727,[Link](https://arxiv.org/abs/2205.14727)Cited by:[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.8.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1)\.
- Z\. Chen, Y\. Hu, Z\. Fu, Z\. Li, J\. Huang, Q\. Huang, and Y\. Wei \(2026\)INTENT: invariance and discrimination\-aware noise mitigation for robust composed image retrieval\.InAAAI,Vol\.40,pp\. 20463–20471\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Z\. Chen, Y\. Hu, Z\. Li, Z\. Fu, X\. Song, and L\. Nie \(2025\)OFFSET: segmentation\-based focus shift revision for composed image retrieval\.InACM MM,pp\. 6113–6122\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- A\. Chronopoulou, A\. Margatina, C\. Baziotis, and A\. Potamianos \(2018\)NTUA\-SLP at IEST 2018: ensemble of neural transfer methods for implicit emotion classification\.InProceedings of the 9th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis,A\. Balahur, S\. M\. Mohammad, V\. Hoste, and R\. Klinger \(Eds\.\),Brussels, Belgium,pp\. 57–64\.External Links:[Link](https://aclanthology.org/W18-6209/),[Document](https://dx.doi.org/10.18653/v1/W18-6209)Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Y\. Cui, W\. Che, T\. Liu, B\. Qin, and Z\. Yang \(2021\)Pre\-training with whole word masking for chinese bert\.IEEE/ACM transactions on audio, speech, and language processing29,pp\. 3504–3514\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- D\. Demszky, D\. Movshovitz\-Attias, J\. Ko, A\. Cowen, G\. Nemade, and S\. Ravi \(2020\)GoEmotions: a dataset of fine\-grained emotions\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 4040–4054\.External Links:[Link](https://aclanthology.org/2020.acl-main.372/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.372)Cited by:[§B\.4](https://arxiv.org/html/2608.10810#A2.SS4.SSS0.Px1.p1.5),[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.6.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1),[§3\.4](https://arxiv.org/html/2608.10810#S3.SS4.p3.4)\.
- J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova \(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 4171–4186\.External Links:[Link](https://aclanthology.org/N19-1423/),[Document](https://dx.doi.org/10.18653/v1/N19-1423)Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Devon018 \(2025\)CN\-SarcasmBench\.Note:[https://huggingface\.co/datasets/Devon018/CN\-SarcasmBench](https://huggingface.co/datasets/Devon018/CN-SarcasmBench)Hugging Face dataset; licensed under CC\-BY\-NC\-4\.0Cited by:[§B\.2](https://arxiv.org/html/2608.10810#A2.SS2.p1.1)\.
- J\. W\. Du Bois \(2007\)The stance triangle\.Stancetaking in discourse: Subjectivity, evaluation, interaction164\(3\),pp\. 139–182\.Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- Q\. Du and V\. Hoste \(2025\)Another approach to agreement measurement and prediction with emotion annotations\.InProceedings of the 19th Linguistic Annotation Workshop \(LAW\-XIX\-2025\),S\. Peng and I\. Rehbein \(Eds\.\),Vienna, Austria,pp\. 87–102\.External Links:[Link](https://aclanthology.org/2025.law-1.7/),[Document](https://dx.doi.org/10.18653/v1/2025.law-1.7),ISBN 979\-8\-89176\-262\-6Cited by:[§B\.4](https://arxiv.org/html/2608.10810#A2.SS4.SSS0.Px1.p1.5)\.
- D\. Ghosal, N\. Majumder, A\. Gelbukh, R\. Mihalcea, and S\. Poria \(2020\)COSMIC: COmmonSense knowledge for eMotion identification in conversations\.InFindings of the Association for Computational Linguistics: EMNLP 2020,T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 2470–2481\.External Links:[Link](https://aclanthology.org/2020.findings-emnlp.224/),[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.224)Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- F\. Gilardi, M\. Alizadeh, and M\. Kubli \(2023\)ChatGPT outperforms crowd workers for text\-annotation tasks\.Proceedings of the National Academy of Sciences120\(30\)\.External Links:ISSN 1091\-6490,[Link](http://dx.doi.org/10.1073/pnas.2305016120),[Document](https://dx.doi.org/10.1073/pnas.2305016120)Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p4.1)\.
- H\. P\. Grice \(1975\)Logic and conversation\.InSpeech acts,pp\. 41–58\.Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- J\. J\. Gross \(1998\)The emerging field of emotion regulation: an integrative review\.Review of general psychology2\(3\),pp\. 271–299\.Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- B\. Hu, M\. Zhang, C\. Xie, Y\. Tian, Y\. Song, and Z\. Mao \(2024\)RESEMO: a benchmark Chinese dataset for studying responsive emotion from social media content\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 16375–16387\.External Links:[Link](https://aclanthology.org/2024.findings-acl.970/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.970)Cited by:[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.10.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1)\.
- R\. Klinger, O\. De Clercq, S\. Mohammad, and A\. Balahur \(2018\)IEST: WASSA\-2018 implicit emotions shared task\.InProceedings of the 9th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis,A\. Balahur, S\. M\. Mohammad, V\. Hoste, and R\. Klinger \(Eds\.\),Brussels, Belgium,pp\. 31–42\.External Links:[Link](https://aclanthology.org/W18-6206/),[Document](https://dx.doi.org/10.18653/v1/W18-6206)Cited by:[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.4.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1)\.
- H\. Li, S\. Zhu, and Z\. Zheng \(2023\)Diplomat: a dialogue dataset for situated pragmatic reasoning\.Advances in Neural Information Processing Systems36,pp\. 46856–46884\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1)\.
- W\. Li, Z\. Song, H\. Zhou, J\. Yu, Y\. Zhang, and W\. Yang \(2026a\)LoRA\-mixer: coordinate modular lora experts through serial attention routing\.InInternational Conference on Learning Representations,C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(Eds\.\),Vol\.2026,pp\. 14694–14716\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/file/18610cbcd1da57854aa05ebdc5cd3168-Paper-Conference.pdf)Cited by:[§G\.1](https://arxiv.org/html/2608.10810#A7.SS1.p1.1)\.
- Y\. Li, H\. Su, X\. Shen, W\. Li, Z\. Cao, and S\. Niu \(2017\)DailyDialog: a manually labelled multi\-turn dialogue dataset\.InProceedings of the Eighth International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),G\. Kondrak and T\. Watanabe \(Eds\.\),Taipei, Taiwan,pp\. 986–995\.External Links:[Link](https://aclanthology.org/I17-1099/)Cited by:[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.3.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1)\.
- Z\. Li, Z\. Chen, H\. Wen, Z\. Fu, Y\. Hu, and W\. Guan \(2025\)Encoder: entity mining and modification relation binding for composed image retrieval\.InAAAI,Vol\.39,pp\. 5101–5109\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Z\. Li, Y\. Hu, Z\. Chen, H\. Wen, X\. Song, and L\. Nie \(2026b\)COMBINER: composed image retrieval guided by attribute\-based neighbor relations\.IEEE TIP\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Z\. Li, Y\. Hu, Z\. Chen, M\. Zhang, Z\. Fu, and L\. Nie \(2026c\)Conesep: cone\-based robust noise\-unlearning compositional network for composed image retrieval\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 16897–16909\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Z\. Li, Y\. Hu, Z\. Chen, S\. Zhang, Q\. Huang, Z\. Fu, and Y\. Wei \(2026d\)HABIT: chrono\-synergia robust progressive learning framework for composed image retrieval\.InAAAI,Vol\.40,pp\. 6762–6770\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Z\. Li, Y\. Hu, Z\. Fu, Z\. Chen, W\. Guan, and L\. Nie \(2026e\)R3: composed video retrieval via reasoning\-guided recalling and re\-ranking\.arXiv preprint arXiv:2606\.01113\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Z\. Li, Y\. Hu, Z\. Fu, Z\. Chen, Y\. Li, and L\. Nie \(2026f\)Tema: anchor the image, follow the text for multi\-modification composed image retrieval\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 24421–24442\.Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- H\. Liu, L\. Ding, and R\. Henao \(2026\)Learning to control summaries with score ranking\.arXiv preprint arXiv:2604\.17197\.Cited by:[§F\.2](https://arxiv.org/html/2608.10810#A6.SS2.p1.1)\.
- H\. Liu and R\. Henao \(2025\)Learning to substitute words with model\-based score ranking\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 11551–11565\.Cited by:[§F\.2](https://arxiv.org/html/2608.10810#A6.SS2.p1.1)\.
- R\. W\. Picard \(1997\)Affective computing\.Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- R\. Plutchik \(2001\)The nature of emotions: human emotions have deep evolutionary roots, a fact that may explain their complexity and provide tools for clinical practice\.American scientist89\(4\),pp\. 344–350\.Cited by:[Appendix A](https://arxiv.org/html/2608.10810#A1.SS0.SSS0.Px2.p1.1)\.
- S\. Poria, D\. Hazarika, N\. Majumder, G\. Naik, E\. Cambria, and R\. Mihalcea \(2019\)MELD: a multimodal multi\-party dataset for emotion recognition in conversations\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 527–536\.External Links:[Link](https://aclanthology.org/P19-1050/),[Document](https://dx.doi.org/10.18653/v1/P19-1050)Cited by:[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.5.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1)\.
- Z\. Song, X\. Li, Y\. Zhang, X\. Zhang, W\. Yang, and J\. Yu \(2026\)Social intelligence modeling: a comprehensive survey from social perception to social simulation\.External Links:[Document](https://dx.doi.org/10.13140/RG.2.2.21157.87528)Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- S\. Sravanthi, M\. Doshi, P\. Tankala, R\. Murthy, R\. Dabre, and P\. Bhattacharyya \(2024\)PUB: a pragmatics understanding benchmark for assessing LLMs’ pragmatics capabilities\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 12075–12097\.External Links:[Link](https://aclanthology.org/2024.findings-acl.719/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.719)Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1)\.
- D\. Wang, Y\. Zhang, J\. Yu, Y\. P\. Chen, C\. Xu, and Z\. Song \(2026\)Seeing further and wider: joint spatio\-temporal enlargement for micro\-video popularity prediction\.External Links:2604\.20311,[Link](https://arxiv.org/abs/2604.20311)Cited by:[§G\.1](https://arxiv.org/html/2608.10810#A7.SS1.p1.1)\.
- Y\. Wang, P\. Ke, Y\. Zheng, K\. Huang, Y\. Jiang, X\. Zhu, and M\. Huang \(2020\)A large\-scale chinese short\-text conversation dataset\.InNLPCC,External Links:[Link](https://arxiv.org/abs/2008.03946)Cited by:[§B\.2](https://arxiv.org/html/2608.10810#A2.SS2.p1.1)\.
- K\. Wong, P\. Paritosh, and L\. Aroyo \(2021\)Cross\-replication reliability \- an empirical approach to interpreting inter\-rater reliability\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 7053–7065\.External Links:[Link](https://aclanthology.org/2021.acl-long.548/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.548)Cited by:[§B\.4](https://arxiv.org/html/2608.10810#A2.SS4.SSS0.Px1.p1.5),[§3\.4](https://arxiv.org/html/2608.10810#S3.SS4.p3.4)\.
- Y\. Wu, Y\. Zhang, L\. Ye, G\. Zeng, J\. Yu, C\. Xu, and Z\. Song \(2026\)HotComment: a benchmark for evaluating popularity of online comments\.arXiv preprint arXiv:2604\.25614\.Cited by:[Appendix A](https://arxiv.org/html/2608.10810#A1.SS0.SSS0.Px1.p1.1)\.
- G\. Xianwei, L\. Hua, X\. Yan, Y\. Zhengtao, and H\. Yuxin \(2021\)Emotion classification of COVID\-19 Chinese microblogs based on the emotion category description\.InProceedings of the 20th Chinese National Conference on Computational Linguistics,S\. Li, M\. Sun, Y\. Liu, H\. Wu, K\. Liu, W\. Che, S\. He, and G\. Rao \(Eds\.\),Huhhot, China,pp\. 916–927\(eng\)\.External Links:[Link](https://aclanthology.org/2021.ccl-1.82/)Cited by:[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.7.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1)\.
- X\. Zhang, Y\. Zhang, Z\. Chen, J\. Yu, W\. Yang, and Z\. Song \(2026a\)Logical phase transitions: understanding collapse in LLM logical reasoning\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 18836–18860\.External Links:[Link](https://aclanthology.org/2026.acl-long.858/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.858),ISBN 979\-8\-89176\-390\-6Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Y\. Zhang, Y\. Yu, Q\. Guo, B\. Wang, D\. Zhao, S\. Uprety, D\. Song, Q\. Li, and J\. Qin \(2023\)CMMA: benchmarking multi\-affection detection in chinese multi\-modal conversations\.Advances in Neural Information Processing Systems36,pp\. 18794–18805\.Cited by:[Table 1](https://arxiv.org/html/2608.10810#S1.T1.1.1.9.1),[§1](https://arxiv.org/html/2608.10810#S1.p2.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx1.p1.1),[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- Y\. Zhang, Y\. Ai, Z\. Ying, Q\. Mi, J\. Yu, W\. Yang, and Z\. Song \(2026b\)Coupling macro dynamics and micro states for long\-horizon social simulation\.arXiv preprint arXiv:2604\.05516\.Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- Y\. Zhang, Z\. Song, H\. Zhou, W\. Ren, Y\. P\. Chen, J\. Yu, and W\. Yang \(2025\)GA−S3GA\-S^\{3\}: Comprehensive social network simulation with group agents\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 8950–8970\.External Links:[Link](https://aclanthology.org/2025.findings-acl.468/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.468),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p3.1)\.
- Y\. Zhang, Z\. Ying, X\. Zhang, J\. Yu, P\. Fang, X\. Chen, W\. Yang, and Z\. Song \(2026c\)IntervenSim: intervention\-aware social network simulation for opinion dynamics\.External Links:2604\.06600,[Link](https://arxiv.org/abs/2604.06600)Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p1.1)\.
- Y\. Zhang, X\. Zhang, Z\. Chen, J\. Yu, and Z\. Song \(2026d\)Semiotic logical hexagon theory for llm logical reasoning\.External Links:2607\.21933,[Link](https://arxiv.org/abs/2607.21933)Cited by:[§G\.1](https://arxiv.org/html/2608.10810#A7.SS1.p1.1)\.
- Y\. Zhang, X\. Zhang, J\. Sheng, W\. Li, J\. Yu, Y\. P\. Chen, W\. Yang, and Z\. Song \(2026e\)Semantic\-aware logical reasoning via a semiotic framework\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 18349–18374\.External Links:[Link](https://aclanthology.org/2026.acl-long.835/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.835),ISBN 979\-8\-89176\-390\-6Cited by:[§2](https://arxiv.org/html/2608.10810#S2.SSx2.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica \(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§1](https://arxiv.org/html/2608.10810#S1.p4.1)\.
## Appendix
This appendix provides implementation details for annotation, examples, prompts, and release documentation\.
## The Usage of LLM
In accordance with \*CL policy, LLMs were used as writing and annotation\-support tools\. For dataset construction, model outputs are treated as candidate labels and rationales\. Final labels are human\-validated\.
## Appendix AAdditional Distributional Analysis
Figure 5:Distribution of the three label layers in CUE\-Bench\. Affective Stance and Pragmatic Intent are shown in full; Fine\-grained Emotion shows the ten most frequent categories for readability\.#### Context\-bound stance categories\.
The rarest Affective Stance isReportive Negative\(0\.6%\), where negative surface wording is used with a largely neutral, reportive force\. This pattern is more natural in news\-style reporting, incident summaries, or analytical long\-form writing than in casual online interaction\. Because CUE\-Bench is dominated by internet dialogue and social\-media discourseWuet al\.\([2026](https://arxiv.org/html/2608.10810#bib.bib49)\),Reportive Negativeremains sparse; Zhihu\-style long\-form text is more likely to contain such cases, but it is not the dominant source\. By contrast,Formulaic Positive\(9\.2%\) is tied to service\-oriented interaction, where thanks, apologies, honorifics, and blessing formulas are often routine rather than deeply positive\.
#### Intent and emotion distributions\.
Authenticityis the most frequent Pragmatic Intent \(34\.7%\), while socially mediated intents such asPoliteness\(16\.8%\),Suppression\(14\.5%\),Irony\(12\.4%\), andResistance\(9\.5%\) remain substantial\. Fine\-grained Emotion shows a stronger long tail\. The right panel of Figure[5](https://arxiv.org/html/2608.10810#A1.F5)shows the high\-frequency head; low\-frequency emotions are omitted from the visualization for readability but remain part of the benchmark and evaluation\. The emotion inventory is organized with reference to Plutchik’s emotion wheel\(Plutchik,[2001](https://arxiv.org/html/2608.10810#bib.bib15)\)\. The prominence ofContempt\(20\.1%\) is consistent with the negative and sarcastic bias of the initial corpus construction, while categories such asSubmissionandHopeshow that the benchmark also captures relational and anticipatory affect\.
## Appendix BCUE\-Bench Details
### B\.1Release Format and Data Split
The released data follow a unified format that supports all three benchmark tasks\. Table[5](https://arxiv.org/html/2608.10810#A2.T5)summarizes the core fields\.
We split the dataset by source instance rather than by individual text span to avoid context leakage across training, development, and test sets\. The split is stratified by Affective Stance where possible, since the nine\-way stance distribution is naturally imbalanced\.
### B\.2Source Data and Label Independence
CUE\-Bench draws raw Chinese text and dialogue context from LCCC\(Wanget al\.,[2020](https://arxiv.org/html/2608.10810#bib.bib3)\), JDDC\(Chenet al\.,[2020](https://arxiv.org/html/2608.10810#bib.bib4)\), the CCAC 2024 Chinese Sarcasm Calculation dataset\(CCAC 2024 Chinese Sarcasm Calculation Organizers,[2024](https://arxiv.org/html/2608.10810#bib.bib5)\), Zhihu QA, and CN\-SarcasmBench\(Devon018,[2025](https://arxiv.org/html/2608.10810#bib.bib6)\)\. These sources are used only as text pools\. We do not reuse their original sentiment labels, sarcasm labels, dialogue labels, or any other source\-provided annotations as CUE\-Bench labels\. All released Affective Stance, Pragmatic Intent, and Fine\-grained Emotion annotations are produced through our own screening, candidate annotation, human verification, and adjudication pipeline\.
### B\.3Annotation Pipeline and Quality Control
We construct CUE\-Bench through a hybrid annotation pipeline that combines model\-assisted pre\-annotation, human verification, and consistency\-based LLM adjudication\. Starting from 60,000 candidate instances, GPT\-4o\-mini and DeepSeek\-V4 independently produce candidate annotations\. The 20,000 instances on which the two models agree are retained as high\-confidence annotations\. For the remaining 40,000 disagreement cases, we manually verify 10,000 instances to construct the human\-verified gold subset\. During this process, annotators select the more appropriate label from the model\-generated candidates rather than assuming that all model annotations are incorrect\.
To further assess annotation reliability, we sample 300 instances from the human\-verified gold subset and ask three expert annotators to independently re\-annotate them\. As reported in Table[2](https://arxiv.org/html/2608.10810#S3.T2), Affective Stance obtains the strongest agreement, while Pragmatic Intent and Fine\-grained Emotion show lower raw agreement due to their greater subjectivity and dependence on latent affective inference\. However, agreement improves substantially when downstream labels are evaluated under consistent Affective Stance\. For Pragmatic Intent, the conditional Krippendorff’sα\\alphaincreases to 0\.7688\. This supports the role of Affective Stance as an intermediate structure for localizing ambiguity and reducing uncertainty in intent and emotion annotation\.
For the remaining 30,000 model\-disagreement instances, we use GPT\-4o\-mini as an adjudicator\. Validation on the human\-verified gold subset shows that the adjudicator achieves 89% accuracy when selecting between two model\-generated candidate annotations\. To reduce positional bias, each instance is adjudicated twice with the order of candidate labels reversed\. Only instances with consistent forward–reverse adjudication decisions are retained, yielding 21,823 additional instances and discarding 8,177 uncertain cases\. The final benchmark contains 51,823 instances\.
FieldDescriptionsample\_idAnonymized instance identifier\.contextContext used for affective interpretation\.target\_textTarget utterance or text span to be labeled\.explicit\_layerSurface affective information\.implicit\_layerContext\-inferred implicit affective information\.affective\_stanceNine\-way Affective Stance label\.pragmatic\_intentPragmatic intent label\.fine\_grained\_emotionFine\-grained emotion label\.Table 5:Core fields in the CUE\-Bench release format\.
### B\.4Detailed Pairwise Agreement
Table[6](https://arxiv.org/html/2608.10810#A2.T6)reports pairwise Cohen’sκ\\kappascores between the gold labels and each expert re\-annotation\. These results complement the aggregate agreement statistics in Table[2](https://arxiv.org/html/2608.10810#S3.T2)by showing how each expert annotation compares with the gold labels\. Conditional scores are computed on instances with consistent Affective Stance\.
LayerExpertRawκ\\kappaCond\.κ\\kappaAffective StanceA0\.5116–B0\.5864–Pragmatic IntentA0\.40570\.8746B0\.37860\.7042Fine\-grained EmotionA0\.33250\.6915B0\.34120\.6462Table 6:Pairwise Cohen’sκ\\kappascores between the gold labels and expert re\-annotations\. Conditional scores are computed on instances with consistent Affective Stance\.#### Interpreting Agreement Scores\.
We interpret agreement scores in the context of subjective affective annotation rather than relying on a single universal threshold\. Prior work has argued that fixed IRR thresholds such asκ\\kappaorα\>0\.6\\alpha\>0\.6can be overly rigid for subjective tasks with genuine ambiguityWonget al\.\([2021](https://arxiv.org/html/2608.10810#bib.bib27)\)\. Fine\-grained emotion annotation also commonly exhibits moderate or low chance\-corrected agreement: GoEmotions reports an average Cohen’sκ\\kappaof approximately 0\.29 across 27 emotion categoriesDemszkyet al\.\([2020](https://arxiv.org/html/2608.10810#bib.bib9)\), and recent work on emotion annotation reports Fleiss’κ\\kappavalues of 0\.19–0\.33 and Krippendorff’sα\\alphavalues of 0\.22–0\.64 for emotion and valence annotation settingsDu and Hoste \([2025](https://arxiv.org/html/2608.10810#bib.bib28)\)\. These findings motivate our use of both raw and conditional agreement scores, where the latter evaluates whether downstream intent and emotion labels become more reliable under a shared Affective Stance interpretation\.
## Appendix CAnnotation Guidelines
### C\.1Explicit Affective Signal
Annotators label the explicit affective signal using only the target utterance\. Context may be shown for orientation, but the decision must be justified by surface evidence in the utterance itself\.
Positive signal\.Usepositivewhen the utterance directly expresses appreciation, happiness, agreement, praise, relief, or encouragement\.
Negative signal\.Usenegativewhen the utterance directly expresses blame, anger, disappointment, sadness, anxiety, refusal, or complaint\.
Neutral signal\.Useneutralwhen the utterance contains no clear surface affective expression, even if the surrounding context is emotional\.
### C\.2Implicit Affective Tendency
Annotators label the implicit affective tendency using the dialogue context and the target utterance together\. The label should capture the affect implied by the speaker’s stance, conversational goal, or pragmatic force, rather than simply repeating the surface wording\.
Positive tendency\.Usepositivewhen the utterance implies acceptance, care, support, relief, or friendly intent\.
Negative tendency\.Usenegativewhen the utterance implies rejection, dissatisfaction, pressure, hostility, disappointment, or sarcasm\.
Neutral tendency\.Useneutralwhen context does not support a clear positive or negative affective inference\.
### C\.3Affective Stance Assignment
Affective Stance is derived from the ordered pair of explicit affective signal and implicit affective tendency\. Annotators therefore do not invent an independent stance label\. They first verify the two base signals, map the pair to the Explicit\-Implicit Stance Matrix, and then check whether the resulting stance matches the intended interpretation\. If the mapped stance feels implausible, annotators must revisit the explicit or implicit signal rather than manually overriding the stance\.
Authentic expression\.UseAuthenticitywhen the utterance directly expresses the speaker’s genuine affective position\.
Polite mitigation\.UsePolitenesswhen positive or softened wording primarily serves social etiquette, deference, apology, or service\-script politeness\.
Affective concealment\.UseSuppressionwhen the speaker hides, weakens, or withholds the implied affect\.
Contrastive meaning\.UseIronywhen the utterance relies on reversal, sarcasm, exaggerated praise, or contrast between literal and intended meaning\.
Oppositional stance\.UseResistancewhen the utterance implies refusal, opposition, complaint, pressure, or dissatisfaction\.
Task\-oriented communication\.UseFunctionalwhen the utterance mainly reports, informs, requests, or describes without strong interpersonal strategy\.
Playful framing\.UseHumorwhen the utterance uses playfulness, teasing, or comic framing as the main pragmatic force\.
Supportive orientation\.UseEmpathywhen the utterance primarily conveys care, comfort, solidarity, or perspective\-taking toward another person\.
### C\.4Fine\-grained Emotion Decision Rules
Annotators label Fine\-grained Emotion as the final affective interpretation, using the target utterance, context, Affective Stance, and Pragmatic Intent together\. The label should describe the speaker’s inferred affective state, not the emotion that the utterance may cause in the reader\.
Specificity preference\.Prefer the most specific emotion supported by contextual evidence; use broad labels such asNeutralonly when no specific affect is recoverable\.
Underlying affect\.Distinguish outward negativity from the underlying state: complaint may indicateOutrage,Disappointment,Contempt, orAnxietydepending on context\.
Relational and future\-oriented affect\.Use relational and anticipatory labels such asSubmission,Hope,Pessimism, orOptimismwhen the utterance encodes expectation, dependence, resignation, or future orientation\.
Ambiguity resolution\.When multiple emotions are plausible, choose the label best supported by the target utterance and its immediate conversational context, and flag genuinely ambiguous cases for review\.
## Appendix DDefinitions and Examples of Affective Stances
Table[7](https://arxiv.org/html/2608.10810#A4.T7)provides detailed definitions and examples for the nine Affective Stances in the Explicit\-Implicit Stance Matrix\.
eie\_\{i\}hih\_\{i\}Affective StanceTypical PhenomenonChinese–English Example\+\+\+\+PositiveDirect praise, joy, gratitude, approval, support, or affection\.Zh:太好了!你终于熬出来了,我真的为你骄傲。
En:That’s wonderful\! You finally made it through\. I’m truly proud of you\.\+\+0Formulaic
PositiveCustomer\-service politeness, routine thanks, greetings, scripted apologies, or socially expected positive wording\.Zh:亲,真的非常抱歉给您带来不便,感谢您的理解,祝您生活愉快。
En:Dear customer, we sincerely apologize for the inconvenience\. Thank you for your understanding, and have a nice day\.\+\+−\-Sarcastic
NegativeSarcasm, irony, backhanded praise, mock compliment, or exaggerated praise used to express criticism\.Zh:你可真是天才,每次都能精准踩雷。
En:What a genius you are, managing to step on the exact landmine every time\.0\+\+Understated
PositiveModesty, humblebragging, restrained pride, indirect approval, or neutral wording that hides satisfaction\.Zh:也就一般吧,省赛第一而已。
En:It was nothing special, just first place in the provincial contest\.00NeutralFactual, procedural, descriptive, or informational utterances without clear affective commitment\.Zh:会议改到下午三点,地点不变。
En:The meeting has been moved to 3 p\.m\.; the location remains unchanged\.0−\-Veiled
NegativeNeutral wording that implies dissatisfaction, reluctance, compromise, pressure, rejection, or concealed criticism\.Zh:行,你继续按这个方案来,结果我就不评价了。
En:Fine, keep following this plan\. I won’t comment on the result\.−\-\+\+Affiliative
PositiveAggressive joking among friends, affectionate blame, self\-deprecation for attention, or criticism mixed with care and expectation\.Zh:你个猪,终于知道好好考一次了。
En:You little pig, you finally learned to take an exam seriously\.−\-0Reportive
NegativeNews reporting, objective analysis, factual description, or neutral discussion of negative events\.Zh:事故造成多人受伤,现场交通一度中断。
En:The accident injured several people, and traffic at the scene was temporarily suspended\.−\-−\-NegativeDirect anger, blame, rejection, disappointment, anxiety, sadness, disgust, or dissatisfaction\.Zh:蠢货,给我滚出这里。
En:You idiot, get out of here\.Table 7:Definitions and examples of the nine Affective Stances\. Each stance is determined by the composition of explicit affective signaleie\_\{i\}and implicit affective tendencyhih\_\{i\}\.
## Appendix EFormal Setup
### E\.1Notation
LetD=\(u1,…,uT\)D=\(u\_\{1\},\\ldots,u\_\{T\}\)be a dialogue\. For a target utteranceutu\_\{t\}, the context isCt=\(ut−k,…,ut−1\)C\_\{t\}=\(u\_\{t\-k\},\\ldots,u\_\{t\-1\}\)for a configurable window sizekk, or the full previous dialogue history when available\. Each benchmark instance is represented asxi=\(Ci,ui\)x\_\{i\}=\(C\_\{i\},u\_\{i\}\)and annotated with explicit affective signal, implicit affective tendency, Affective Stance, Pragmatic Intent, and Fine\-grained Emotion\.
### E\.2Well\-Formedness
An annotation is well formed only when the Affective Stance label matches the ordered pair\(yiexp,yiimp\)\(y\_\{i\}^\{exp\},y\_\{i\}^\{imp\}\)underϕ\\phi\. Invalid stance combinations are rejected during export validation\. Pragmatic Intent and Fine\-grained Emotion are not deterministic functions of the stance, but they must be justified by the same context–target pair and by the intermediate stance interpretation\.
## Appendix FDataset Card
### F\.1Intended Use
The dataset is intended for research on Chinese discourse emotion understanding, Affective Stance Recognition, Pragmatic Intent Understanding, Fine\-grained Emotion Classification, and robust affect\-aware dialogue systems\. It is designed for evaluating how models infer unsaid affect from context rather than for reusing source\-dataset sentiment or sarcasm annotations\.
### F\.2Out\-of\-Scope Use
The dataset should not be used to infer private mental states about real individuals, rankLiu and Henao \([2025](https://arxiv.org/html/2608.10810#bib.bib39)\); Liuet al\.\([2026](https://arxiv.org/html/2608.10810#bib.bib40)\)users by emotional tendency, or make high\-stakes decisions about employment, health, credit, or legal status\.
### F\.3Source and License Notes
The release records source identifiers so users can trace the text pool from which each instance was drawn\. Users should comply with the licenses and terms of the underlying source datasets, especially for sources with non\-commercial restrictions\. CUE\-Bench annotations are newly created and should not be interpreted as inherited labels from the original datasets\.
### F\.4Recommended Reporting
Papers using CUE\-Bench should report the input setting, context window, prompt or model configuration, split version, Accuracy, macro\-F1, weighted\-F1, and whether any low\-confidence or filtered examples are included\. For Task 1, Affective Stance Recognition, reports should include a confusion matrix over the nine stance categories, because this task corresponds to the Explicit\-Implicit Stance Matrix\. For Task 2 and Task 3, reports should include class\-wise F1 or per\-class error analysis when space permits, since Pragmatic Intent and Fine\-grained Emotion are long\-tailed and more subjective\. When external data augmentation or additional annotation is used, papers should distinguish it clearly from the official CUE\-Bench training, development, and test splits\.
## Appendix GAdditional Evaluation Details
### G\.1Input Formatting for LLM Evaluation
Each LLMLiet al\.\([2026a](https://arxiv.org/html/2608.10810#bib.bib51)\); Zhanget al\.\([2026d](https://arxiv.org/html/2608.10810#bib.bib43)\); Wanget al\.\([2026](https://arxiv.org/html/2608.10810#bib.bib47)\)input contains the dialogue context, the target utterance, task\-specific label definitions, and a constrained output schema\. For direct prompting settings, the model is asked to predict the target label directly from the given context and target utterance\. For chain\-based or matrix\-guided prompting settings, the prompt specifies an intermediate reasoning order, but the intermediate fields are inferred by the model itself rather than provided as gold labels\.
For Affective Stance Recognition, matrix\-guided prompting asks the model to first infer the explicit affective signal and implicit affective tendency, and then map the ordered pair to one of the nine stance labels\. For Pragmatic Intent and Fine\-grained Emotion, matrix\-guided prompting similarly requires the model to derive the relevant intermediate fields before predicting the final label\. Gold intermediate labels are provided only in the oracle\-conditioning ablation settings, where they are inserted into the prompt to test the contribution of each intermediate variable\.
### G\.2Prompting Protocol
Zero\-shot prompting includes label definitions and a constrained output format\. Few\-shot prompting adds demonstrations covering aligned affect, neutral\-surface latent affect, and contrastive affect\. Free\-form CoT prompting asks the model to explain its reasoning before giving the final label, while matrix\-guided prompting fixes the reasoning order from explicit signal to implicit tendency, Affective Stance, Pragmatic Intent, and Fine\-grained Emotion\. All prompts are evaluated with deterministic decoding when the API or model interface supports it\.
### G\.3Prediction Normalization
Model outputs are normalized before scoring\. We strip formatting artifacts, map aliases to canonical label names, and reject outputs that cannot be matched to the task label space\. When a response contains both rationale text and a final answer, only the final structured label field is used for metric computation\. This prevents verbose reasoning from being treated as additional labels\.
### G\.4Oracle\-conditioning Settings
For the ablation study, gold intermediate labels are inserted into the prompt while the target label remains hidden\. Task I provides gold Affective Stance and predicts Pragmatic Intent\. Task II provides gold Affective Stance and predicts Fine\-grained Emotion\. Task III provides both gold Affective Stance and gold Pragmatic Intent and predicts Fine\-grained Emotion\.相似文章
VocalAffectBench:评估AI音频模型中的声音情感识别
VocalAffectBench 被引入作为一个公共基准,用于评估AI音频模型中的声音情感识别,展示了当前基线模型的准确性有限,尤其是对于非中性情感。
CAREBench:通过评估认知评价推理来检验LLM的情感理解能力
介绍CAREBench,一个基于评价理论的基准测试,通过认知评价推理评估LLM的情感理解能力,表明当前模型在推理和积极情绪识别方面存在困难,尽管在某些下游任务上与人类表现相当。
文本嵌入中情感线索的跨心理学情绪理论比较研究
本文评估了十二种最新文本编码器在三种心理学情绪理论中编码情感线索的能力,发现指令感知的开源权重编码器在单词级别上达到或超过专有编码器,而任务微调嵌入在句子级别上更优。
EIBench:基于模拟器的基准测试与面向情感管理的回合信用强化学习
EIBench 引入了一个基于模拟器的交互式情感管理基准测试,通过每轮用户状态反馈实现评估与训练。作者提出了 CTC-GRPO,一种强化学习方法,在多个基准测试上提升了情感管理表现。
CallBench:电话助手双目标协调基准测试
CallBench是一个中文基准测试,用于评估电话助手中的双目标协调能力,包含50,000轮跨越六个场景的多轮对话,并采用预设感知的评估协议,覆盖语义理解、安全性和对话节奏。