Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
Summary
Introduces counterfactual audits to test whether audio language model judges actually use paralinguistic evidence when evaluating speech-to-speech responses, finding that contrastive success often overstates native reliability and similar accuracies can hide different failure modes across Gemini, GPT, and open models.
View Cached Full Text
Cached at: 08/10/26, 08:03 AM
# Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
Source: [https://arxiv.org/html/2608.06718](https://arxiv.org/html/2608.06718)
###### Abstract
Audio\-language models \(ALMs\) are increasingly used as judges for speech\-to\-speech systems, but a judge that receives audio may not actually use paralinguistic evidence\. We introduce counterfactual audits for paralinguistic response evaluation\. Each audit item holds the transcript fixed while varying affect, prosody, or the timing of an affective shift, forcing a valid judge to track the audio cue rather than lexical content or response style\. We evaluate ALM judges using a native one\-context judgment protocol and a contrastive recoverability control, then further decompose each item into its constituent perception and response\-mapping skills\. This yields useful diagnostic states that identify different sources of judge failures\. Across Gemini, GPT, and open audio models, we find that contrastive success often overstates native judge reliability, and that similar aggregate accuracies can hide different failure modes\. These results suggest that ALM judges should not be evaluated by accuracy alone, instead requiring thorough behavioral audits before deployment\.
Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
Kevin Miller\*, Arjun Chandra\*Venkatesh SaligramaBoston University\{nivek, ac25, srv\}@bu\.edu
footnotetext:\*Equal contribution\.## 1Introduction
Figure 1:Counterfactual audio\-judge audit\.Left:The audit holds the user’s words fixed while changing only the paralinguistic realization\.Middle:Pointwiseprovides one audio context and two candidate responses\. If this fails,Pairwiseprovides both counterfactual audio contexts and responses\. Component probes then test the constituent task skills to identify failure sources\.Right:The judge correctly identifies the paralinguistic state \(P=1P=1\) and maps that state to the better response when supplied in text \(O=1O=1\), but fails in the native audio judgment \(J=0J=0\)\.Paralinguistic cues such as tone, rhythm, pitch, and timing are central to spoken interaction\(Scherer,[2003](https://arxiv.org/html/2608.06718#bib.bib56); Crystal,[1975](https://arxiv.org/html/2608.06718#bib.bib13)\)\. The same words can require different assistant responses depending on how they are spoken: a user may sound surprised, hesitant, impatient, or frustrated even when the transcript alone is ambiguous\. As speech\-to\-speech assistants become more capable, evaluation must therefore move beyond text\. We need to know whether a spoken response is appropriate not only for what the user said, but also for how the user said it\.
A natural way to scale this evaluation is to use audio\-language models \(ALMs\) as automatic judges\(Manakulet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib40); Chianget al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib41); Jianget al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib30)\)\. An ALM judge can listen to an interaction and decide which assistant response is better\. If reliable, such judges could support benchmarking, reward modeling, deployment monitoring, and regression testing for spoken assistants\. However, this creates a second\-order evaluation problem: before using an ALM as a judge, we must audit whether the judge itself uses the relevant audio evidence\. A judge that relies mostly on transcripts, response style, or generic conversational preferences is not a trustworthy evaluator of paralinguistics\.
We study this problem through*counterfactual audits*\. Each audit item holds the user’s words fixed while changing only the paralinguistic realization\. The candidate responses are paired with these audio realizations\. When the relevant audio cue changes, a valid judge should change its decision\.
A failed counterfactual audit is underdetermined\. If the judge selects the incorrect response, it does not reveal whether the judge failed to hear the audio cue, map the cue to the right response, or compose these skills in the final judgment\. We therefore use the root\-cause diagnostic cascade shown in Figure[1](https://arxiv.org/html/2608.06718#S1.F1)\.
Native judgment\.The target setting isPointwise: the judge receives one audio context and two candidate responses, and must choose the response appropriate for that audio\. We call this judgment native because it is closest to the standard setting for using an ALM as an automatic evaluator\.
Contrastive recoverability\.If native judgment fails, we ask whether the relevant counterfactual contrast is recoverable when made explicit\. InPairwise, the judge receives both audio realizations and both candidate responses, and must determine the correct matchingZhuet al\.\([2026](https://arxiv.org/html/2608.06718#bib.bib22)\)\. IfPairwisesucceeds, the explicit contrast forces the model to look past the identical transcripts and rely on the audio cue\. This reveals that the model has the latent paralinguistic judgement ability but fails to deploy it in the native one\-context setting, makingPairwisean effective diagnostic tool\.
Component attribution\.We then ask what operation failed\. For each item, we probe perceptionPP, oracle response mappingOO, and native judgmentJJ\. The perception probe asks whether the judge identifies the task\-relevant paralinguistic state\. The oracle response\-mapping probe asks whether the judge can choose the correct response when that user’s paralinguistic state is provided\. Together, the state\(P,O,J\)\(P,O,J\)separates heterogeneous failure modes across different judge models\.
We apply this audit to both single\-turn and multi\-turn settings\. Single\-turn examples isolate static paralinguistic response selection: given one utterance rendered with different affective states, does the judge select the response that fits the spoken state? Positional multi\-turn examples add temporal and causal structure: the transcript is fixed, but the timing of the user’s affective shift changes, so the correct response depends on when the user became frustrated and what caused the shift\. This captures a deployment\-relevant difficulty for spoken assistants, where emotion is often interpreted through dialogue history\(Poriaet al\.,[2019](https://arxiv.org/html/2608.06718#bib.bib57); Duet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib28); Aroraet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib14)\)\.
Our experiments show that current ALM judges are fragile in distinct ways\. Some models recover the counterfactual contrast inPairwisebut degrade sharply inPointwise, showing that contrastive success does not imply native judge reliability\. The component probes further show that aggregate accuracy hides heterogeneous failures across judges\. In positional multi\-turn examples, one\-context judgment remains especially brittle, indicating that temporal\-causal paralinguistic evaluation is challenging\.
Contributions\.First, we introduce counterfactual audits for ALMs used as judges of paralinguistic response appropriateness\. Second, we propose a root\-cause diagnostic procedure that combines native one\-context judgment, contrastive recoverability, and component probes\. Third, we present an empirical audit of Gemini, GPT, and open audio models, showing that aggregate accuracy alone does not certify audio\-judge reliability\.
## 2Related Work
Large Language Models \(LLMs\) are routinely used as text evaluators, but they can exhibit systematic biases\(Duboiset al\.,[2024](https://arxiv.org/html/2608.06718#bib.bib27); Zhenget al\.,[2023](https://arxiv.org/html/2608.06718#bib.bib21)\)and display “Potemkin understanding”\(Mancoridiset al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib29)\), motivating decomposition\-based reliability improvements\(Leeet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib38); Liet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib39)\)\. Concurrently, audio\-capable models are increasingly deployed as automatic evaluators\(Manakulet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib40); Chianget al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib41)\)and reward models\(Jiet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib42); Geet al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib25); Yanget al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib23)\)\. These applications typically assume that audio modality access implies reliable paralinguistic reasoning—an assumption challenged by recent findings\(Chandraet al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib52); Chenet al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib24)\)\. While existing benchmarks evaluate paralinguistic instruction following\(Jianget al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib30); Heldet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib11)\)or emotion perception\(Yanget al\.,[2021](https://arxiv.org/html/2608.06718#bib.bib34); Huanget al\.,[2024](https://arxiv.org/html/2608.06718#bib.bib36)\), they provide limited diagnostic leverage for understanding how judges use these cues in downstream preference decisions\. We bridge this gap by applying an instrument auditing perspective to explicitly isolate paralinguistic reasoning in audio evaluators\. Extended discussion of related evaluation frameworks is provided in Appendix[A](https://arxiv.org/html/2608.06718#A1)\.
Table 1:Diagnostic truth table for ALM judge reliability\.Each row in the in the table indicates a different interpretation of the model’s behavior\. The final row is the reliable judge state, while the highlighted Potemkin row is a key orchestration failure we uncover\.
## 3Problem Setup and Framework
Our goal is to audit audio\-language models \(ALMs\) when they are used as judges of spoken interactions\. The target use case is multi\-turn spoken interaction, where an assistant response must fit not only the user’s words but also the user’s tone, affect, and the timing of an affective shift\. This target is realistic, but a wrong judgment is hard to interpret\. The judge may fail to hear the cue, fail to map the cue to the right response, recover the cue only when the contrast is explicit, or fail because temporal\-causal reasoning is required\. The framework below makes these different possibilities identifiable\.
### 3\.1Counterfactual audit items
Each audit item is a controlled counterfactual tuple
E=\(C,𝒜0,𝒜1,R0,R1\),E=\(C,\\mathcal\{A\}^\{0\},\\mathcal\{A\}^\{1\},R^\{0\},R^\{1\}\),whereCCis a shared transcript or conversation history,𝒜0\\mathcal\{A\}^\{0\}and𝒜1\\mathcal\{A\}^\{1\}are two audio realizations of the same words, andR0R^\{0\}andR1R^\{1\}are the responses appropriate for the two realizations\. The lexical content is fixed\. Only the paralinguistic realization changes\.
This gives the audit its construct validity\. A text\-only or lexically biased judge should not solve the item consistently, because both branches share the same words\. A paralinguistically sensitive judge should change its decision when the relevant audio cue changes\. The audit therefore tests whether the judge measures the audio cue, rather than response style or lexical shortcuts\.
### 3\.2Native judgment and contrastive recoverability
The native judgment setting isPointwise\. The judge receives one audio context𝒜y\\mathcal\{A\}^\{y\}and two candidate responses\{R0,R1\}\\\{R^\{0\},R^\{1\}\\\}, and must choose the response appropriate for that audio\.
When native judgment fails, the failure is under\-determined\. We therefore usePairwiseas a contrastive recoverability control\. InPairwise, the judge receives both audio contexts𝒜0,𝒜1\\mathcal\{A\}^\{0\},\\mathcal\{A\}^\{1\}and both responsesR0,R1R^\{0\},R^\{1\}, and must match each response to the corresponding context\.Pairwiseis not the deployment target\. It asks whether the ALM can distinguish the two counterfactual audio\-response worlds when both alternatives are visible\. A largePairwise–Pointwisegap indicates protocol dependence: the contrast is recoverable, but the judge does not reliably deploy it in the native one\-context setting\.
We also vary the amount of cueing in the prompt\. A no\-cue prompt asks for a direct judgment\. A hard cue asks the judge to focus on the user’s emotion and prosody\. For positional examples, a transition cue asks the judge to track where the user’s affect changes before choosing the response\. Cueing is a diagnostic intervention, so a judge whose conclusion depends sharply on prompt wording may be useful in a scaffolded pipeline, but should not be treated as a plug\-and\-play evaluator\.
### 3\.3Component attribution: perception, mapping, and judgment
Protocol comparisons tell us whether a contrast is recoverable and whether the native judgment is reliable\. They do not identify which operation failed\. We therefore evaluate each item through one native judgment and two component probes \(Fig\.[1](https://arxiv.org/html/2608.06718#S1.F1)\):
Pi\\displaystyle P\_\{i\}=𝟏\{perception probe is correct\},\\displaystyle=\\mathbf\{1\}\\\{\\text\{perception probe is correct\}\\\},Oi\\displaystyle O\_\{i\}=𝟏\{oracle response\-mapping probe is correct\},\\displaystyle=\\mathbf\{1\}\\\{\\text\{oracle response\-mapping probe is correct\}\\\},Ji\\displaystyle J\_\{i\}=𝟏\{native audio judgment is correct\}\.\\displaystyle=\\mathbf\{1\}\\\{\\text\{native audio judgment is correct\}\\\}\.The perception probePiP\_\{i\}tests whether the judge identifies the task\-relevant paralinguistic state\. Forsingle\-turn\-emotions\(Sec\.[4](https://arxiv.org/html/2608.06718#S4)\), this is the user’s emotion or prosodic state\. Forpositional\-emotion, this is the full affective trajectory \(i\.e\., the user’s emotion at every turn\)\. The oracle response\-mapping probeOiO\_\{i\}removes the audio\-perception burden by supplying the relevant state in text, and tests whether the judge can map that state to the appropriate response\. The native judgmentJiJ\_\{i\}records whether the judge succeeds in the original audio task\.
Each item is assigned a diagnostic state
zi=\(Pi,Oi,Ji\)∈\{0,1\}3z\_\{i\}=\(P\_\{i\},O\_\{i\},J\_\{i\}\)\\in\\\{0,1\\\}^\{3\}and we report empirical state masses
πpoj=1n∑i=1n𝟏\{Pi=p,Oi=o,Ji=j\}\.\\pi\_\{poj\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbf\{1\}\\\{P\_\{i\}=p,\\,O\_\{i\}=o,\\,J\_\{i\}=j\\\}\.We use\(P,O,J\)\(P,O,J\)order throughout the paper\. The reliable integrated state isπ111\\pi\_\{111\}\. The key orchestration failure isπ110\\pi\_\{110\}: the judge succeeds on perception and oracle response mapping, but fails when those abilities must be composed in the native audio judgment\. We call this a Potemkin failure\(Mancoridiset al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib29)\)\. Other state groups distinguish perception bottlenecks, response\-mapping bottlenecks, shortcut\-like successes, and full\-stack failures\. The full eight\-state table is shown in Table[1](https://arxiv.org/html/2608.06718#S2.T1)\.
### 3\.4Task complexity: static versus temporal\-causal judgment
The final diagnostic axis is interaction complexity\. Multi\-turn spoken interaction is the deployment\-relevant target, but it entangles perception, timing, causal attribution, and response mapping\. We therefore usesingle\-turn\-emotionsas a calibration task: it removes dialogue history and causal attribution, and tests whether the judge can use a static paralinguistic cue at all\.
We then return to the multi\-turn target throughpositional\-emotion\. These examples preserve counterfactual control while restoring temporal structure\. The transcript is controlled, but the affective shift occurs at different turns\. The correct response depends on when the user became frustrated and which assistant action caused the shift\. Thus, a model that succeeds on single\-turn examples but fails on positional examples is likely brittle to temporal\-causal interaction structure\.
## 4Task Construction
The tasks instantiate the measurement design in Section[3](https://arxiv.org/html/2608.06718#S3)\. Each task removes or restores a specific source of difficulty\.single\-turn\-emotionsremoves dialogue history and tests static paralinguistic response selection\.positional\-emotionrestores the deployment\-relevant temporal\-causal structure while preserving counterfactual control\. We also reportemotional\-conversations, an earlier prototype, in Appendix[D\.2](https://arxiv.org/html/2608.06718#A4.SS2)because many of those examples can be solved from the final user turn alone\.
Figure 2:Task families in the audit\.single\-turn\-emotionsisolates static paralinguistic response selection\.emotional\-conversationsis an earlier multi\-turn prototype reported in the appendix\.positional\-emotionhas the timing and cause of the user’s affective shift determine the correct response\.##### Controlled synthesis\.
We use text\-to\-speech as an experimental control\. The transcript, conversation history, and candidate responses are fixed while the paralinguistic realization changes\. This gives internal counterfactual validity: if the judge’s decision changes, the change can be attributed to the audio cue rather than lexical or contextual differences\. We later demonstrate generalization of our findings to real speech in Section[5\.3](https://arxiv.org/html/2608.06718#S5.SS3)\.
### 4\.1Single\-turn task: static calibration
Thesingle\-turn\-emotionstask asks whether a judge can use a static paralinguistic cue before dialogue history, temporal localization, or causal attribution enter\. Each item contains one user utterance rendered with two contrasting emotions, together with two candidate assistant responses\. The correct response depends on the user’s tone, not on the transcript alone\.
##### Source data\.
We source lexical user inputs from the EmoCF subset of CAVA\(Heldet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib11)\)\. EmoCF contains lexically neutral user inputs grounded in diverse social situations from NormBank\(Ziemset al\.,[2023](https://arxiv.org/html/2608.06718#bib.bib12)\)\. This makes it useful for response\-level paralinguistic evaluation: the judge must decide which assistant response fits the spoken cue rather than rely on the words alone\. Standard acted corpora such as CREMA\-D and RAVDESS\(Caoet al\.,[2014](https://arxiv.org/html/2608.06718#bib.bib50); Livingstone and Russo,[2018](https://arxiv.org/html/2608.06718#bib.bib51)\)provide speaker diversity, but use a small set of repeated utterances; this is less suitable for our setting because response appropriateness must be evaluated across diverse lexical contexts\.
##### Construction\.
We select 189 user inputs, yielding 189 pairwise items and 378 pointwise instances\. Each input is paired with two contrasting target emotions, such as anger versus sadness or surprise versus neutrality\. We synthesize user audios withgpt\-4o\-mini\-tts, generate emotion\-appropriate assistant responses withgemini\-2\.5\-flash, and synthesize those responses withgemini\-2\.5\-flash\-preview\-tts\. The responses imply the relevant emotional adaptation without explicitly naming the emotion\. Full prompts and construction details are provided in Appendix[B\.1](https://arxiv.org/html/2608.06718#A2.SS1)\.
Table 2:Key statistics of each audit task\.Table 3:Headline end\-to\-end accuracy \(%\)\.Gemini models perform well in the Pairwise setting but collapse toward chance in Pointwise\. Nearly all other judges remain near chance\. Full Wilson confidence intervals for individual accuracies and paired bootstrap confidence intervals for protocol gaps are reported in Appendix[D\.1](https://arxiv.org/html/2608.06718#A4.SS1)\.
### 4\.2Positional multi\-turn task: temporal\-causal judgment
Thepositional\-emotiontask tests a harder form of paralinguistic judgment\. A user may become frustrated only after an earlier goal is not met, and the appropriate response may depend on identifying which event caused the change\. The task therefore asks whether a judge can localize an affective shift in a conversation and use that causal interpretation to select the right response\.
##### Source data\.
We seed the task with 500 conversations from Open Dialogue \(OD3\)\(Chanet al\.,[2024](https://arxiv.org/html/2608.06718#bib.bib43)\)\. OD3 aggregates task\-oriented dialogue sources including KVRET\(Eric and Manning,[2017](https://arxiv.org/html/2608.06718#bib.bib9)\), MultiWOZ\(Budzianowskiet al\.,[2020](https://arxiv.org/html/2608.06718#bib.bib10)\), DSTC11\(Zhaoet al\.,[2023b](https://arxiv.org/html/2608.06718#bib.bib8)\), NOESIS\-II\(Kummerfeldet al\.,[2019](https://arxiv.org/html/2608.06718#bib.bib7)\), and SIMMC\-2\.1\(Kotturet al\.,[2021a](https://arxiv.org/html/2608.06718#bib.bib6)\)\. These sources cover assistant\-like interactions such as travel, booking, shopping, navigation, reminders, and other goal\-directed tasks\.
##### Goal\-failure injection\.
Given a source conversationSS, we identify user goals and create two counterfactual branches\. In one branch, the assistant fails to satisfy goalgag\_\{a\}at turntat\_\{a\}; in the other, it fails to satisfy goalgbg\_\{b\}at turntbt\_\{b\}, withta≠tbt\_\{a\}\\neq t\_\{b\}\. The final transcript is controlled across the pair, but the onset of negative affect occurs at a different position\. The candidate responses are written so that each response addresses the corresponding cause of the affective shift\. Source conversations are rejected when a plausible controlled contrast cannot be created\.
##### Audio rendering and quality control\.
We render turns withgpt\-4o\-mini\-tts\. For paralinguistically important turns, we generate multiple attempts and keep the best accepted rendering\. Acceptance usesemotion2vecscores\(Maet al\.,[2024](https://arxiv.org/html/2608.06718#bib.bib44)\)and acoustic primitives such as pitch and speech rate\. Full rendering prompts, filtering criteria, and quality\-control details are provided in Appendix[B\.2\.3](https://arxiv.org/html/2608.06718#A2.SS2.SSS3)\.
##### Perception target scoring\.
Forsingle\-turn\-emotions, the perception probe asks for the user’s emotion\. Forpositional\-emotion, the judge must correctly recognize the full affective trajectory \(i\.e\., the user’s emotion at every turn\) for the\(P,O,J\)\(P,O,J\)state analysis\. This keeps the positional diagnostic comparable to the single\-turn diagnostic while preserving the temporal requirement\.
### 4\.3Human validation
Human validation is an internal\-validity check, not the primary source of labels\. The labels are determined by the counterfactual construction\. Human validation tests whether sampled items are perceivable by listeners and whether the response choice is sufficiently specified\. We recruited five independent annotators for each task in the audit\. Annotators sampled items in the samePointwiseformat used for primary model evaluation\. Accuracy varied across annotators from62\-100%62\\text\{\-\}100\\%on the single\-turn task, whereas performance on the positional\-emotion task was generally higher and more consistent, ranging from88\-92%88\\text\{\-\}92\\%\. The validation supports that the sampled items are solvable by careful listeners, and we suspect the high variability on the single\-turn task is due to variation in listener background, though the sample size is limited\. We discuss this in more detail along with the full annotation protocol and per\-annotator ranges in Appendix[E](https://arxiv.org/html/2608.06718#A5)\.
## 5Experiments and Results
We organize the experiments as an audit of each judge\. Each result answers one question in the diagnostic cascade from Figure[1](https://arxiv.org/html/2608.06718#S1.F1)\. We evaluate a range of frontier propietary and open\-source ALM judges on the two main task families in our audit introduced in Section[3\.4](https://arxiv.org/html/2608.06718#S3.SS4)\. Full prompts and implementation details are in Appendix[C\.1](https://arxiv.org/html/2608.06718#A3.SS1)\.
Figure 3:Aggregated instrument\-state distributions\.Colors group the\(P,O,J\)\(P,O,J\)states into reliable integrated judgment, component bottlenecks, Potemkin failures, shortcut\-like successes, and full\-stack failures\.### 5\.1Native judgment and contrastive recoverability
Table[3](https://arxiv.org/html/2608.06718#S4.T3)gives the headline end\-to\-end accuracies\. We read this table through the first two questions in the cascade\. ThePointwisecolumns ask whether the judge works in the native one\-context setting\. ThePairwisecolumns ask a different question: if native judging fails, is the counterfactual contrast recoverable when both audio realizations and both responses are visible?
The single\-turn task already shows a large gap between recoverability and native judging\. Gemini models often solve the contrastive matching problem, but they are much less reliable when asked to judge one audio context at a time\. For example, Gemini\-3\-Pro reaches91\.0%91\.0\\%inPairwisewith the Hard Cue, but only65\.3%65\.3\\%inPointwise; Gemini\-2\.5\-Pro drops from86\.0%86\.0\\%to60\.1%60\.1\\%\. These results show that the relevant paralinguistic contrast can be recoverable but not in the native judge setting\.
The positional task makes the failure sharper\. With the Transition Cue, Gemini\-3\-Flash reaches79\.4%79\.4\\%inPairwise, showing that some judges can solve the temporal matching problem when the contrast and task schema are explicit\. However, the same model reaches only53\.0%53\.0\\%in positionalPointwise; Gemini\-3\-Pro reaches51\.6%51\.6\\%, GPT\-4o reaches49\.2%49\.2\\%, and Qwen\-2\.5\-Omni\-7B reaches51\.1%51\.1\\%\. Thus, contrastive success is not deployable judge reliability\. A model may distinguish the two counterfactuals when both are shown, but fail when the same audio cue must control a single\-context decision\. We quantify this protocol collapse in more detail in Fig\.[4](https://arxiv.org/html/2608.06718#S5.F4)along with paired bootstrap confidence intervals in Appendix[D\.1](https://arxiv.org/html/2608.06718#A4.SS1)\.
Figure 4:Failure heatmaps\.We quantify additional failure modes with conditional Potemkin rates, protocol collapse, and cue dependence\.
### 5\.2What failed? Component attribution
The protocol gap demonstrates when models fail by exposing their reliance on explicit contrastive formats\. However, it does not explain why many of the proprietary and open\-source judges perform near\-chance on both tasks, even in the explicitPairwiseformat\. To understand why they fail—specifically, whether the bottleneck lies in perceiving the audio cue, mapping it to a response, or composing these skills in a native judgment—we must examine the item\-level diagnostic states defined in Section[3\.3](https://arxiv.org/html/2608.06718#S3.SS3)\.
Table 4:Response quality balance check\.Absolute pairwise difference\|R0−R1\|\|R^\{0\}\-R^\{1\}\|across counterfactual response pairs inSingle\-turn\-emotionsconfirm that irrelevant lexical and acoustic factors do not bias the judges\.Figure[3](https://arxiv.org/html/2608.06718#S5.F3)shows the aggregated instrument\-state distribution for each task\. We focus on diagnosing failures in the standardPointwisesetting\. Each bar decomposes evaluated examples into reliable integrated judgments, Potemkin failures, component bottlenecks, shortcuts, and full\-stack failures\.
Crucially, two judges with similar end\-to\-end accuracy can place their probability mass in very different failure states\. In the single\-turn panel, the strongest Gemini models are not simply unable to reason from emotion\. Their oracle response\-mapping ability is high, but their native audio judgments often fail\. Gemini\-2\.5\-Pro has26%26\\%Potemkin mass, Gemini\-3\-Flash has25%25\\%, and Gemini\-3\-Pro has19%19\\%\. Conditional on both component probes succeeding, Gemini\-3\-Pro still fails the native audio judgment26%26\\%of the time, and we report full “Conditional Potemkin” rates across all models in Fig\.[4](https://arxiv.org/html/2608.06718#S5.F4), which notably reaches up to51%51\\%for DeSTA2\.5\-Audio\.
GPT models show a different profile\. In single\-turn, GPT\-4o\-mini and GPT\-4o have lower reliable integration \(38%38\\%and37%37\\%\) and larger perception bottlenecks than the Gemini models\. Their failures are therefore not primarily Potemkin failures; they often fail earlier in the pipeline, before the integration question becomes meaningful\.
In the positional panel, the reliable mass is substantially lower across all models, with even the strongest Gemini\-3\-Pro only reaching31%31\\%\. On the other hand, shortcuts and full\-stack failures are much more prevalent, indicating that the judges are unable to pass the simpler perception and oracle response\-mapping probes in this setting\. These results suggest that current ALM judges do not yet possess the constituent perception and response\-mapping abilities for temporal\-causal paralinguistic judgment\.
Across both single\-turn and positional tasks, the judges also exhibit varying reliance on prompt cues to solve the task, which is quantified in Fig\.[4](https://arxiv.org/html/2608.06718#S5.F4)as the difference in accuracy between the hard cue \(single\-turn\) or transition cue \(positional\) and the no cue setting\.
Figure 5:Generator bias\.Accuracy onPositional\-emotioncomparing GeminiGen and GPTGen generation sources\. The superior performance of Gemini models is due to their genuine capability, and not a bias due to involvement in the construction pipeline\.
### 5\.3Validity and robustness checks
The preceding results could potentially be attributed to simpler experimental artifacts: distribution shift introduced by synthetic speech, generator\-family bias, or irrelevant lexical and audio quality confounders\. To ensure the observed performance gaps reflect genuine reasoning failures rather than methodological artifacts, we conducted a suite of targeted validity checks\. We summarize findings below: Synthetic Audio Artifacts\.Replacing synthesized user audio with the original CAVA human speech yields highly correlated diagnostic profiles \(Spearmanρ\>0\.9\\rho\>0\.9across all models; Fig\.[6](https://arxiv.org/html/2608.06718#S5.F6)\)\. This demonstrates that the observed state distribution results are not artifacts of synthetic speech\. Generator Bias\.Substituting Gemini\-generated positional annotations with GPT\-generated annotations does not alter the qualitative model rankings, confirming that generator\-judge alignment is not the primary driver of the results \(Fig\.[5](https://arxiv.org/html/2608.06718#S5.F5)\)\. Response Quality Balance\.Differences across irrelevant audio quality or lexical factors in the response choices \(e\.g\., one response being more polite\) could mislead a judge to consistently prefer a response and perform poorly on the task\. We rule out this possibility using an LLM judge \(claude\-haiku\-4\-5\) and the DNSMOS audio quality modelReddyet al\.\([2021](https://arxiv.org/html/2608.06718#bib.bib4)\), confirming that irrelevant lexical and audio qualities are balanced between candidate response choices \(Tab\.[4](https://arxiv.org/html/2608.06718#S5.T4)\)\.
Figure 6:Generalization to real speech\.Comparison ofπpoj\\pi\_\{poj\}state proportions for synthetic vs real human speech from CAVA\. The relationship has near perfect Pearson correlation and a slope close to 1 across all models\.
## 6Conclusion
We introduced an evaluation framework based on counterfactual audits to assess the capabilities of audio\-language model judges for paralinguistic reasoning\. Our counterfactual audit yields three key findings\. First, contrastive recoverability overstates native judge reliability: models can distinguish counterfactual alternatives when both are visible, but fail in the standard one\-contextPointwiseformat\. Second, aggregate accuracy hides mechanism: the\(P,O,J\)\(P,O,J\)state distribution separates perception bottlenecks, response\-mapping failures, Potemkin failures, and shortcut successes\. Third, positional multi\-turn judging is the most brittle setting: current judges struggle when the response depends on when and why the user’s affect changed\.
## Limitations
The goal of our audit is to build an evaluation framework to uncover failure modes of paralinguistic reasoning in audio judges, and we leave perspectives on improving judge performance \(e\.g\., model training and training data design\) to future work\. Though we include promising experiments with real human speech, our audit relies primarily on synthesized speech\. This gives the counterfactual control needed to isolate paralinguistic evidence, but it limits external validity\. Future work should test whether the same diagnostic states persist across diverse human\-spoken audio, speakers, recording conditions, languages, and emotion taxonomies\.
We would also like to briefly discuss some societal risks inherent to the topic of this paper\. Our work could be used to improve speech\-to\-speech judges, which in turn could be used to improve speech\-to\-speech models, particularly along paralinguistic dimensions\. Paralinguistically\-aware voice assistants have many positive use cases, but they can also be used for harm, especially in any setting that involves emotional manipulation\. We urge governments and regulatory bodies to be proactive and address these risks sooner rather than later\.
## References
- S\. Arora, Z\. Lu, C\. Chiu, R\. Pang, and S\. Watanabe \(2025\)Talking turns: benchmarking audio foundation models on turn\-taking dynamics\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=2e4ECh0ikn)Cited by:[§1](https://arxiv.org/html/2608.06718#S1.p8.1)\.
- V\. Basile, M\. Fell, T\. Fornaciari, D\. Hovy, S\. Paun, B\. Plank, M\. Poesio, and A\. Uma \(2021\)We need to consider disagreement in evaluation\.InProceedings of the 1st Workshop on Benchmarking: Past, Present and Future,K\. Church, M\. Liberman, and V\. Kordoni \(Eds\.\),Online,pp\. 15–21\.External Links:[Link](https://aclanthology.org/2021.bppf-1.3/),[Document](https://dx.doi.org/10.18653/v1/2021.bppf-1.3)Cited by:[Appendix E](https://arxiv.org/html/2608.06718#A5.p5.1)\.
- P\. Budzianowski, T\. Wen, B\. Tseng, I\. Casanueva, S\. Ultes, O\. Ramadan, and M\. Gašić \(2018\)MultiWOZ \- a large\-scale multi\-domain Wizard\-of\-Oz dataset for task\-oriented dialogue modelling\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 5016–5026\.External Links:[Link](https://aclanthology.org/D18-1547/),[Document](https://dx.doi.org/10.18653/v1/D18-1547)Cited by:[§B\.2\.1](https://arxiv.org/html/2608.06718#A2.SS2.SSS1.p1.1)\.
- P\. Budzianowski, T\. Wen, B\. Tseng, I\. Casanueva, S\. Ultes, O\. Ramadan, and M\. Gašić \(2020\)MultiWOZ – a large\-scale multi\-domain wizard\-of\-oz dataset for task\-oriented dialogue modelling\.External Links:1810\.00278,[Link](https://arxiv.org/abs/1810.00278)Cited by:[§4\.2](https://arxiv.org/html/2608.06718#S4.SS2.SSS0.Px1.p1.1)\.
- H\. Cao, D\. G\. Cooper, and J\. Keutmann \(2014\)CREMA\-d: crowd\-sourced emotional multimodal actors dataset\.IEEE Transactions on Affective Computing5\(4\),pp\. 377–390\.External Links:[Document](https://dx.doi.org/10.1109/TAFFC.2014.2336244)Cited by:[§4\.1](https://arxiv.org/html/2608.06718#S4.SS1.SSS0.Px1.p1.1)\.
- D\. M\. Chan, S\. Ghosh, H\. Tulsiani, A\. Rastrow, and B\. Hoffmeister \(2024\)Task oriented dialogue as a catalysis for self\-supervised automatic speech recognition\.InICASSP 2024\-2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Cited by:[§B\.2\.1](https://arxiv.org/html/2608.06718#A2.SS2.SSS1.p1.1),[§4\.2](https://arxiv.org/html/2608.06718#S4.SS2.SSS0.Px1.p1.1)\.
- A\. Chandra, K\. Miller, V\. Ravichandran, C\. Papayiannis, and V\. Saligrama \(2026\)Hearing between the lines: unlocking the reasoning power of LLMs for speech evaluation\.InFindings of the Association for Computational Linguistics: EACL 2026,V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 2895–2916\.External Links:[Link](https://aclanthology.org/2026.findings-eacl.151/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-eacl.151),ISBN 979\-8\-89176\-386\-9Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- J\. Chen, Z\. Guo, J\. Chun, P\. Wang, A\. Perrault, and M\. Elsner \(2026\)Do audio LLMs really LISTEN, or just transcribe? measuring lexical vs\. acoustic emotion cues reliance\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 5848–5877\.External Links:[Link](https://aclanthology.org/2026.eacl-long.274/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.274),ISBN 979\-8\-89176\-380\-7Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- C\. Chiang, X\. Wang, C\. Lin, K\. Lin, L\. Li, R\. Kopetz, Y\. Qian, Z\. Wang, Z\. Yang, H\. Lee, and L\. Wang \(2025\)Audio\-aware large language models as judges for speaking styles\.arXiv preprint arXiv:2506\.05984\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.05984)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1),[§1](https://arxiv.org/html/2608.06718#S1.p2.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- D\. Crystal \(1975\)Paralinguistics\.InThe Body as a Medium of Expression,J\. Benthall and T\. Polhemus \(Eds\.\),pp\. 162–174\.Cited by:[§1](https://arxiv.org/html/2608.06718#S1.p1.1)\.
- A\. M\. Davani, M\. Diaz, and V\. Prabhakaran \(2022\)Dealing with disagreements: looking beyond the majority vote in subjective annotations\.Transactions of the Association for Computational Linguistics10,pp\. 92–110\.Cited by:[Appendix E](https://arxiv.org/html/2608.06718#A5.p5.1)\.
- Y\. Du, Q\. Huang, G\. Zhu, Z\. Dai, S\. Chen, Q\. Zhu, L\. Pan, M\. Chen, Y\. Zhang, L\. Zhou, B\. Wang, and H\. Li \(2025\)MTalk\-bench: evaluating speech\-to\-speech models in multi\-turn dialogues via arena\-style and rubrics protocols\.External Links:2508\.18240,[Link](https://arxiv.org/abs/2508.18240)Cited by:[§1](https://arxiv.org/html/2608.06718#S1.p8.1)\.
- Y\. Dubois, B\. Galambosi, P\. Liang, and T\. B\. Hashimoto \(2024\)Length\-controlled AlpacaEval: a simple way to debias automatic evaluators\.arXiv preprint arXiv:2404\.04475\.Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p1.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- M\. Eric, L\. Krishnan, F\. Charette, and C\. D\. Manning \(2017\)Key\-value retrieval networks for task\-oriented dialogue\.InProceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue,K\. Jokinen, M\. Stede, D\. DeVault, and A\. Louis \(Eds\.\),Saarbrücken, Germany,pp\. 37–49\.External Links:[Link](https://aclanthology.org/W17-5506/),[Document](https://dx.doi.org/10.18653/v1/W17-5506)Cited by:[§B\.2\.1](https://arxiv.org/html/2608.06718#A2.SS2.SSS1.p1.1)\.
- M\. Eric and C\. D\. Manning \(2017\)Key\-value retrieval networks for task\-oriented dialogue\.External Links:1705\.05414,[Link](https://arxiv.org/abs/1705.05414)Cited by:[§4\.2](https://arxiv.org/html/2608.06718#S4.SS2.SSS0.Px1.p1.1)\.
- Y\. Ge, J\. Zhang, X\. Liu, B\. Li, X\. Ma, C\. Wang, K\. Ye, Y\. Du, L\. Zhang, Y\. Huang, and et al\. \(2026\)SageLM: a multi\-aspect and explainable large language model for speech judgement\.Proceedings of the AAAI Conference on Artificial Intelligence40\(36\),pp\. 30807–30815\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40338),[Document](https://dx.doi.org/10.1609/aaai.v40i36.40338)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- C\. Gunasekara, J\. K\. Kummerfeld, L\. Lastras, and W\. S\. Lasecki \(2020\)NOESIS ii: predicting responses, identifying success, and managing complexity in task\-oriented dialogue\.InAAAI: Workshop on Dialog System Tech Challenges,Cited by:[§B\.2\.1](https://arxiv.org/html/2608.06718#A2.SS2.SSS1.p1.1)\.
- W\. Held, M\. J\. Ryan, A\. Shrivastava, A\. S\. Khan, C\. Ziems, E\. Li, M\. Bartelds, M\. Sun, T\. Li, W\. Gan, and D\. Yang \(2025\)CAVA: comprehensive assessment of voice assistants\.Note:[https://talkarena\.org/cava](https://talkarena.org/cava)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p3.1),[§B\.1\.1](https://arxiv.org/html/2608.06718#A2.SS1.SSS1.p1.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1),[§4\.1](https://arxiv.org/html/2608.06718#S4.SS1.SSS0.Px1.p1.1)\.
- C\. Huang, O\. Ke, H\. Huang, Y\. Chen, Y\. Kuan, H\. Chang, A\. Sriram, and H\. Lee \(2024\)Dynamic\-superb: towards a dynamic, collaborative, and comprehensive instruction\-tuning benchmark for speech\.arXiv preprint arXiv:2404\.03606\.Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p3.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- Y\. Huang, L\. F\. R\. Ribeiro, M\. Hardalov, B\. Dhingra, M\. Dreyer, and V\. Saligrama \(2026\)DeepFact: co\-evolving benchmarks and agents for deep research factuality\.External Links:2603\.05912Cited by:[Appendix E](https://arxiv.org/html/2608.06718#A5.p5.1)\.
- A\. A\. G\. Intelligence \(2025\)Amazon nova 2: multimodal reasoning and generation models\.Amazon Technical Reports\.External Links:[Link](https://www.amazon.science/publications/amazon-nova-2-multimodal-reasoning-and-generation-models)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1)\.
- S\. Ji, T\. Liang, Y\. Li, J\. Zuo, M\. Fang, J\. He, Y\. Chen, Z\. Liu, Z\. Jiang, X\. Cheng, S\. Zheng, J\. Xu, J\. Lin, and Z\. Zhao \(2025\)WavReward: spoken dialogue models with generalist reward evaluators\.arXiv preprint arXiv:2505\.09558\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.09558)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- F\. Jiang, Z\. Lin, F\. Bu, Y\. Du, B\. Wang, and H\. Li \(2025\)S2S\-arena: evaluating speech2speech protocols on instruction following with paralinguistic information\.arXiv preprint arXiv:2503\.05085\.Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p3.1),[§1](https://arxiv.org/html/2608.06718#S1.p2.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- S\. Kottur, S\. Moon, A\. Geramifard, and B\. Damavandi \(2021a\)SIMMC 2\.0: a task\-oriented dialog dataset for immersive multimodal conversations\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,Online and Punta Cana, Dominican Republic,pp\. 4903–4912\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.401),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.401)Cited by:[§4\.2](https://arxiv.org/html/2608.06718#S4.SS2.SSS0.Px1.p1.1)\.
- S\. Kottur, S\. Moon, A\. Geramifard, and B\. Damavandi \(2021b\)SIMMC 2\.0: a task\-oriented dialog dataset for immersive multimodal conversations\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 4903–4912\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.401/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.401)Cited by:[§B\.2\.1](https://arxiv.org/html/2608.06718#A2.SS2.SSS1.p1.1)\.
- J\. K\. Kummerfeld, S\. R\. Gouravajhala, J\. Peper, V\. Athreya, C\. Gunasekara, J\. Ganhotra, S\. S\. Patel, L\. Polymenakos, and W\. S\. Lasecki \(2019\)A large\-scale corpus for conversation disentanglement\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Cited by:[§4\.2](https://arxiv.org/html/2608.06718#S4.SS2.SSS0.Px1.p1.1)\.
- Y\. Lee, J\. Kim, J\. Kim, H\. Cho, J\. Kang, P\. Kang, and N\. Kim \(2025\)CheckEval: a reliable LLM\-as\-a\-judge framework for evaluating text generation using checklists\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,Suzhou, China,pp\. 15771–15798\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.796/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.796)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p1.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- M\. Li, Z\. Liu, S\. Deng, S\. Joty, N\. Chen, and M\. Kan \(2025\)DnA\-eval: enhancing large language model evaluation through decomposition and aggregation\.InProceedings of the 31st International Conference on Computational Linguistics,Abu Dhabi, UAE,pp\. 2277–2290\.External Links:[Link](https://aclanthology.org/2025.coling-main.156/)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p1.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- A\. H\. Liu, A\. Ehrenberg, A\. Lo, C\. Denoix, C\. Barreau, G\. Lample, J\. Delignon, K\. R\. Chandu, P\. von Platen, P\. R\. Muddireddy, S\. Gandhi, S\. Ghosh, S\. Mishra, T\. Foubert, A\. Rastogi, A\. Yang, A\. Q\. Jiang, A\. Sablayrolles, A\. Héliou, A\. Martin, A\. Agarwal, A\. Roux, A\. Darcet, A\. Mensch, B\. Bout, B\. Rozière, B\. D\. Monicault, C\. Bamford, C\. Wallenwein, C\. Renaudin, C\. Lanfranchi, D\. Dabert, D\. S\. Chaplot, D\. Mizelle, D\. de las Casas, E\. Chane\-Sane, E\. Fugier, E\. B\. Hanna, G\. Berrada, G\. Delerce, G\. Guinet, G\. Novikov, G\. Martin, H\. Jaju, J\. Ludziejewski, J\. Rute, J\. Chabran, J\. Chudnovsky, J\. Studnia, J\. Barmentlo, J\. Amar, J\. S\. Roberts, J\. Denize, K\. Saxena, K\. Yadav, K\. Khandelwal, K\. Jain, L\. R\. Lavaud, L\. Blier, L\. Zhao, L\. Martin, L\. Saulnier, L\. Gao, M\. Pellat, M\. Guillaumin, M\. Felardos, M\. Dinot, M\. Darrin, M\. Augustin, M\. Seznec, N\. Gupta, N\. Raghuraman, O\. Duchenne, P\. Wang, P\. Saffer, P\. Jacob, P\. Wambergue, P\. Kurylowicz, P\. Chagniot, P\. Stock, P\. Agrawal, R\. Delacourt, R\. Sauvestre, R\. Soletskyi, S\. Vaze, S\. Subramanian, S\. Garg, S\. Dalal, S\. Gandhi, S\. Aithal, S\. Antoniak, T\. L\. Scao, T\. Schueller, T\. Lavril, T\. Robert, T\. Wang, T\. Lacroix, T\. Bewley, V\. Nemychnikova, V\. Paltz, V\. Richard, W\. Li, W\. Marshall, X\. Zhang, Y\. Wan, and Y\. Tang \(2025\)Voxtral\.External Links:2507\.13264,[Link](https://arxiv.org/abs/2507.13264)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1)\.
- Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. Zhu \(2023\)G\-Eval: NLG evaluation using GPT\-4 with better human alignment\.InProceedings of EMNLP,Note:arXiv:2303\.16634Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p1.1)\.
- S\. R\. Livingstone and F\. A\. Russo \(2018\)The ryerson audio\-visual database of emotional speech and song \(ravdess\): a dynamic, multimodal set of facial and vocal expressions in north american english\.PLOS ONE13\(5\),pp\. e0196391\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0196391),[Link](https://doi.org/10.1371/journal.pone.0196391)Cited by:[§B\.1\.1](https://arxiv.org/html/2608.06718#A2.SS1.SSS1.p1.1),[§4\.1](https://arxiv.org/html/2608.06718#S4.SS1.SSS0.Px1.p1.1)\.
- K\. Lu, Z\. Chen, S\. Fu, C\. H\. Yang, S\. Huang, C\. Yang, C\. Yu, C\. Chen, W\. Chen, C\. Huang, Y\. Lin, Y\. Lin, C\. Fu, C\. Kuan, W\. Ren, X\. Chen, W\. Huang, E\. Hu, T\. Lin, Y\. Wu, K\. Huang, H\. Huang, H\. Chou, K\. Chang, C\. Chiang, B\. Ginsburg, Y\. F\. Wang, and H\. Lee \(2026\)DeSTA2\.5\-audio: toward general\-purpose large audio language model with self\-generated cross\-modal alignment\.External Links:2507\.02768,[Link](https://arxiv.org/abs/2507.02768)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1)\.
- Z\. Ma, Z\. Zheng, J\. Ye, J\. Li, Z\. Gao, S\. Zhang, and X\. Chen \(2024\)Emotion2vec: self\-supervised pre\-training for speech emotion representation\.Proc\. ACL 2024 Findings\.Cited by:[§B\.2\.3](https://arxiv.org/html/2608.06718#A2.SS2.SSS3.p3.1),[§4\.2](https://arxiv.org/html/2608.06718#S4.SS2.SSS0.Px3.p1.1)\.
- P\. Manakul, W\. H\. Gan, M\. J\. Ryan, A\. S\. Khan, W\. Sirichotedumrong, K\. Pipatanakul, W\. Held, and D\. Yang \(2025\)AudioJudge: understanding what works in large audio model based speech evaluation\.arXiv preprint arXiv:2507\.12705\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2507.12705)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1),[§1](https://arxiv.org/html/2608.06718#S1.p2.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- M\. Mancoridis, B\. Weeks, K\. Vafa, and S\. Mullainathan \(2025\)Potemkin understanding in large language models\.InProceedings of the 42nd International Conference on Machine Learning,Note:PMLRCited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p1.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1),[§3\.3](https://arxiv.org/html/2608.06718#S3.SS3.p2.3)\.
- Microsoft, :, A\. Abouelenin, A\. Ashfaq, A\. Atkinson, H\. Awadalla, N\. Bach, J\. Bao, A\. Benhaim, M\. Cai, V\. Chaudhary, C\. Chen, D\. Chen, D\. Chen, J\. Chen, W\. Chen, Y\. Chen, Y\. Chen, Q\. Dai, X\. Dai, R\. Fan, M\. Gao, M\. Gao, A\. Garg, A\. Goswami, J\. Hao, A\. Hendy, Y\. Hu, X\. Jin, M\. Khademi, D\. Kim, Y\. J\. Kim, G\. Lee, J\. Li, Y\. Li, C\. Liang, X\. Lin, Z\. Lin, M\. Liu, Y\. Liu, G\. Lopez, C\. Luo, P\. Madan, V\. Mazalov, A\. Mitra, A\. Mousavi, A\. Nguyen, J\. Pan, D\. Perez\-Becker, J\. Platin, T\. Portet, K\. Qiu, B\. Ren, L\. Ren, S\. Roy, N\. Shang, Y\. Shen, S\. Singhal, S\. Som, X\. Song, T\. Sych, P\. Vaddamanu, S\. Wang, Y\. Wang, Z\. Wang, H\. Wu, H\. Xu, W\. Xu, Y\. Yang, Z\. Yang, D\. Yu, I\. Zabir, J\. Zhang, L\. L\. Zhang, Y\. Zhang, and X\. Zhou \(2025\)Phi\-4\-mini technical report: compact yet powerful multimodal language models via mixture\-of\-loras\.External Links:2503\.01743,[Link](https://arxiv.org/abs/2503.01743)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1)\.
- S\. Poria, D\. Hazarika, N\. Majumder, G\. Naik, E\. Cambria, and R\. Mihalcea \(2019\)MELD: a multimodal multi\-party dataset for emotion recognition in conversations\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 527–536\.Cited by:[§1](https://arxiv.org/html/2608.06718#S1.p8.1)\.
- C\. K\. A\. Reddy, V\. Gopal, and R\. Cutler \(2021\)DNSMOS: a non\-intrusive perceptual objective speech quality metric to evaluate noise suppressors\.External Links:2010\.15258,[Link](https://arxiv.org/abs/2010.15258)Cited by:[§5\.3](https://arxiv.org/html/2608.06718#S5.SS3.p1.1)\.
- K\. R\. Scherer \(2003\)Vocal communication of emotion: a review of research paradigms\.Speech Communication40\(1–2\),pp\. 227–256\.External Links:[Document](https://dx.doi.org/10.1016/S0167-6393%2802%2900084-5)Cited by:[§1](https://arxiv.org/html/2608.06718#S1.p1.1)\.
- A\. N\. Uma, T\. Fornaciari, D\. Hovy, S\. Paun, B\. Plank, and M\. Poesio \(2021\)Learning from disagreement: a survey\.Journal of Artificial Intelligence Research72,pp\. 1385–1470\.Cited by:[Appendix E](https://arxiv.org/html/2608.06718#A5.p5.1)\.
- J\. Xu, Z\. Guo, J\. He, H\. Hu, T\. He, S\. Bai, K\. Chen, J\. Wang, Y\. Fan, K\. Dang, B\. Zhang, X\. Wang, Y\. Chu, and J\. Lin \(2025\)Qwen2\.5\-omni technical report\.External Links:2503\.20215,[Link](https://arxiv.org/abs/2503.20215)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1)\.
- S\. Yang, P\. Chi, Y\. Chuang, C\. J\. Lai, K\. Lakhotia, Y\. Y\. Lin, A\. T\. Liu, J\. Shi, X\. Chang, G\. Lin,et al\.\(2021\)SUPERB: speech processing universal performance benchmark\.InProceedings of Interspeech,Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p3.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- S\. Yang, M\. Tu, A\. T\. Liu, X\. Qu, H\. Lee, L\. Lu, Y\. Wang, and Y\. Wu \(2026\)ParaS2S: benchmarking and aligning spoken language models for paralinguistic\-aware speech\-to\-speech interaction\.External Links:2511\.08723,[Link](https://arxiv.org/abs/2511.08723)Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p2.1),[Appendix A](https://arxiv.org/html/2608.06718#A1.p3.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- C\. Zhao, S\. Gella, S\. Kim, D\. Jin, D\. Hazarika, A\. Papangelis, B\. Hedayatnia, M\. Namazifar, Y\. Liu, and D\. Hakkani\-Tur \(2023a\)"What do others think?": task\-oriented conversational modeling with subjective knowledge\.External Links:2305\.12091,[Link](https://arxiv.org/abs/2305.12091)Cited by:[§B\.2\.1](https://arxiv.org/html/2608.06718#A2.SS2.SSS1.p1.1)\.
- C\. Zhao, S\. Gella, S\. Kim, D\. Jin, D\. Hazarika, A\. Papangelis, B\. Hedayatnia, M\. Namazifar, Y\. Liu, and D\. Hakkani\-Tur \(2023b\)"What do others think?": task\-oriented conversational modeling with subjective knowledge\.External Links:2305\.12091Cited by:[§4\.2](https://arxiv.org/html/2608.06718#S4.SS2.SSS0.Px1.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica \(2023\)Judging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.arXiv preprint arXiv:2306\.05685\.Cited by:[Appendix A](https://arxiv.org/html/2608.06718#A1.p1.1),[§2](https://arxiv.org/html/2608.06718#S2.p1.1)\.
- Y\. Zhu, J\. Zhang, and F\. Tang \(2026\)Test\-time matching: unlocking compositional reasoning in multimodal models\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=wWxdT6LB2D)Cited by:[§1](https://arxiv.org/html/2608.06718#S1.p6.1)\.
- C\. Ziems, J\. Dwivedi\-Yu, Y\. Wang, A\. Halevy, and D\. Yang \(2023\)NormBank: a knowledge bank of situational social norms\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 7756–7776\.External Links:[Link](https://aclanthology.org/2023.acl-long.429/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.429)Cited by:[§4\.1](https://arxiv.org/html/2608.06718#S4.SS1.SSS0.Px1.p1.1)\.
## Appendix AAdditional Related Work
Model\-based evaluation and judge reliability\.Large Language Models \(LLMs\) are now routinely used as evaluators for text generation and dialogue, e\.g\., in MT\-Bench and Chatbot Arena\(Zhenget al\.,[2023](https://arxiv.org/html/2608.06718#bib.bib21)\)or via rubric\-based prompting such as G\-Eval\(Liuet al\.,[2023](https://arxiv.org/html/2608.06718#bib.bib26)\)\. A growing literature has shown that model\-based judges can exhibit systematic biases \(e\.g\., verbosity and position\) and can be steered by superficial cues\(Duboiset al\.,[2024](https://arxiv.org/html/2608.06718#bib.bib27)\)\. In response, several approaches improve reliability by decomposing criteria into structured sub\-questions or multi\-stage decision pipelines\(Leeet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib38); Liet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib39)\)\. Our work adopts an*instrument auditing*perspective on audio judges, motivated by similar concerns that foundation models can display “Potemkin understanding”—producing plausible answers on simple probes while lacking robust understanding\(Mancoridiset al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib29)\)\.
Audio judges and speech evaluation\.As spoken language models have matured, recent work has proposed using audio\-capable LLMs or large audio modelsXuet al\.\([2025](https://arxiv.org/html/2608.06718#bib.bib19)\); Microsoftet al\.\([2025](https://arxiv.org/html/2608.06718#bib.bib17)\); Luet al\.\([2026](https://arxiv.org/html/2608.06718#bib.bib18)\); Intelligence \([2025](https://arxiv.org/html/2608.06718#bib.bib15)\); Liuet al\.\([2025](https://arxiv.org/html/2608.06718#bib.bib20)\)as automatic evaluators for speech quality, prosody, and speaking style\(Manakulet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib40); Chianget al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib41)\)\. Complementary lines of work train audio reward models that score end\-to\-end spoken dialogue behavior\(Jiet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib42); Geet al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib25); Yanget al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib23)\)\. These methods are attractive because they can reduce the need for expensive human annotators; however, they typically assume that access to the audio modality implies reliable paralinguistic reasoning\. Evidence from more recent work\(Chandraet al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib52); Chenet al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib24)\)suggest that this may not be the case, and we use this as motivation to explicitly audit paralinguistic reasoning in audio judges\.
Benchmarks for paralinguistic interaction\.Benchmarks such as S2S\-Arena\(Jianget al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib30)\)and ParaS2S\(Yanget al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib23)\)emphasize paralinguistic instruction following and style in speech\-to\-speech systems, and CAVA\(Heldet al\.,[2025](https://arxiv.org/html/2608.06718#bib.bib11)\)introduces counterfactual response evaluation for voice assistants\. On the perception side, speech representation benchmarks like SUPERB\(Yanget al\.,[2021](https://arxiv.org/html/2608.06718#bib.bib34)\)and Dynamic\-SUPERB\(Huanget al\.,[2024](https://arxiv.org/html/2608.06718#bib.bib36)\)quantify capabilities such as emotion recognition and instruction following\. These resources are crucial for measuring what models can hear, but they provide limited diagnostic leverage for evaluating how audio judges reason about paralinguistics when making downstream preference decisions\.
## Appendix BDataset Construction
### B\.1Single\-Turn Dataset Construction
We use this section to go over the details of constructing thesingle\-turn\-emotionsdataset\.
#### B\.1\.1Source Dataset
In order to ground our dataset in realistic user input data, we use the Tone Awareness dataset from CAVAHeldet al\.\([2025](https://arxiv.org/html/2608.06718#bib.bib11)\)\. The original task provides data where each example contains emotionally ambiguous sentences voiced by human actors in two contrasting emotions from a standard emotion taxonomy \(anger, disgust, fear, joy, sadness, surprise, neutral\)Livingstone and Russo \([2018](https://arxiv.org/html/2608.06718#bib.bib51)\)\. While the lexical content of the dataset is strong, we find that the human actors often fail to provide unambiguous emotional voicings\. Therefore, we opt to use TTS to synthesize new audio recordings for each example\.
#### B\.1\.2Filtering & Synthesis
We start by filtering the CAVA Tone Awareness dataset to only include obviously contrasting emotion pairs\. We then synthesize each pair of user inputs withgpt\-4o\-mini\-tts\. To generate the two voice assistant responses, we provide the following prompt toGemini\-2\.5\-Flash:
> You are a friendly voice assistant\. You will be given two versions of the same sentence spoken with different emotions\. Sentence: "\{sentence\}" Emotion A: \{emoA\} Emotion B: \{emoB\} For each emotion, write what a voice assistant should say in response\. The two responses should reflect the emotion they are responding to, so that someone reading them could tell which emotion they correspond to\. However, they should not give away or directly state the emotion\. Instead, the response should strongly but implicitly acknowledge the emotion, and it should be very clear which response corresponds to which emotion\. Respond in strict JSON format with keys "A\_response" and "B\_response"\. Do not use a JSON code block\. Just format your response as a JSON string\.
We then synthesize the two generated voice assistant responses usinggemini\-2\.5\-flash\-preview\-tts\. We finally perform a manual quality check of the generated examples, resulting in the final dataset composition:
- •52 samples \(27\.5%\) with anger vs\. sadness
- •52 samples \(27\.5%\) with surprise vs\. sadness
- •49 samples \(25\.9%\) with surprise vs\. neutral
- •36 samples \(19\.1%\) with anger vs\. surprise
### B\.2Multi\-Turn Dataset Construction
We use this section to go over the details of generating thepositional\-emotionandemotional\-conversationsmulti\-turn benchmarks\. Both benchmarks use the same pipeline, except that the LLM prompts for transcript generation differ\. Full code for generating these benchmarks will be released when the paper is accepted\.
#### B\.2\.1Source Conversations & Sampling
As we mentioned in the main paper, we start by sampling a source conversationSSfrom OD3Chanet al\.\([2024](https://arxiv.org/html/2608.06718#bib.bib43)\), which in turn comprises text\-only conversation datasets KVRETEricet al\.\([2017](https://arxiv.org/html/2608.06718#bib.bib45)\), Multi\-WozBudzianowskiet al\.\([2018](https://arxiv.org/html/2608.06718#bib.bib46)\), DSTC11 \(Track 5\)Zhaoet al\.\([2023a](https://arxiv.org/html/2608.06718#bib.bib47)\), NOESIS\-IIGunasekaraet al\.\([2020](https://arxiv.org/html/2608.06718#bib.bib48)\)and SIMMC\-2\.1Kotturet al\.\([2021b](https://arxiv.org/html/2608.06718#bib.bib49)\)\. These datasets all consist of text\-to\-text conversations between two humans, one of whom is playing the part of a human user and the other an “AI” assistant that has access to a provided knowledge base\. The tasks involved include finding and booking hotels, trains, restaurants, etc\., shopping for clothes and furniture, getting driving directions, scheduling appointments and reminders, and planning out university coursework\. OD3 adds “repeat” and “rephrase” turns to the transcripts of these conversations to make them more realistic for a voice\-based conversation, although they do not explore the role of paralinguistics at all\.
We sample from the training split of OD3, which has 47702 total examples, reweighting the constituent sets to bring the sampling closer to uniform\. The final composition of the dataset is influenced not only by this sampling but also by which examples are rejected by the transcript\-generation and audio\-rendering parts of the pipeline\. The final composition of thepositional\-emotiondataset is as follows \(500 conversation pairs total\):
- •171 pairs \(34\.2%\) from SIMMC\-2\.1
- •141 pairs \(28\.2%\) from NOESIS\-II
- •112 pairs \(22\.4%\) from DSTC11 \(Track 5\)
- •24 pairs \(4\.8%\) from Multi\-Woz
- •52 pairs \(10\.4%\) from KVRET
Each conversation pair yields 2 pointwise instances, giving us a total of 1000 pointwise instances in thepositional\-emotiondataset\.
We note that DSTC11 \(Track 5\) and Multi\-Woz are very similar in nature, both involving things like finding and booking restaurants, trains, hotels, etc\., so together they should be viewed together as a 27\.2% contingent in our benchmark\. KVRET is underrepresented inpositional\-emotionbecause it only makes up 4\.8% of the original OD3 dataset, and many of its source conversations are too short to contain multiple failed goals\.
#### B\.2\.2Conversation Transcript Generation Prompts
Youareanexpertatanalysingandmodifyingspokenconversations\.Wearebuildingaspeech\-to\-speechdatasetwhichwewillcall"PositionalWinoSound",because,likeWinogroundandWinogradSchema,wewanttomakedifficultpairsthatseparatenaivemodelsfromstrongones,basedontheirabilitytodetectcertaindifferencesthathumanscouldeasilydetect\.Inthiscasethedifferenceisinparalinguisticcues,specificallyemotion\.
YouwillbegivenatexttranscriptofasourceconversationS\(i\.e\.alistofturns\)betweenahumanuser\("HUMAN"\)andavoiceassistant\("VA"\)\.YoumustuseStocreateapairofconversations,annotatedwithparalinguistics,whichwewillcalla"PositionalWinoSoundpair"\.Youwillgothroughthefollowingprocesstogeneratethispair\.Notethatyouhavetheoptionofgivingupandsaying"impossible"atanyphaseoftheprocessifitisnotpossibletoproduceasuitablepairfromS\.Hereistheprocess:
I\.EXTRACTGOALSANDFROMS:
FindeachhumanturninSwherethehumanstatesagoal\.Foreachoftheseturns,findthefirstsubsequentVAturnthatrespondstothegoal\.Thisshouldgiveyoualistoftuples\(g\_1,r\_1\),\.\.\.,\(g\_N,r\_N\)wherer\_iisthereponseforgoalg\_i\.Ifyoucan’tgetN\>=2,giveupandsay"impossible"\.
II\.NOTEWHICHGOALSAREUNSATISFIED:
Fromthislist,wewanttocreatealistofenrichedtuples\(g\_1,r\_1,f\_1\),\.\.\.,\(g\_N,r\_N,f\_N\)asfollows:
\*foreachi,checkifg\_iiscompletelysatisfiedbyr\_\.\.\.ustALWAYSacknowledgethesourceofthehuman’snegativeemotion\.
\*\*WhenusingthistonewithinC\_AorC\_B,VAshouldtrytoacknowledgebutonlyifit’spossiblewhilekeepingthetextofC\_AandC\_Bcompletelyidenticaltoeachother\.
\*"angry"\-humanusesthistonewhentheyareangry,furious,pissed\-off,orveryirritated\.VAshouldneverusethistone\.
\*"sad"\-humanusesthistonewhentheyaresadordisappointed\.VAshouldneverusethistone\.
\*"happy"\-humanusesthistonewhentheyarefeelinghappy,upbeat,cheerful\.VAmayusethistonetomirrorhuman’shappiness,butanupsethumanmightbeputoffbyanoverlycheerfulVA\.
Howto"annotate"aturnwithemotion:
\*YoumustchooseonlyONEemotion
\*TheemotionMUSTbeanitemfromthislist:\["hesistant","frazzled","apologetic/empathetic/reassuring","angry","neutral","sad","happy"\]
\*Irepeat,youmayONLYchooseemotionsfromtheabovelist,NOTHINGELSE
\*AnnotationmustbeoftheformatHUMAN\(emotion\):"text"orVA\(emotion\):"text",e\.g\.HUMAN\(hesitant\):"Ithinkmybagishere"orVA\(empathetic\):"Icancheckforyou"
YouMUSTexplainyourreasoningbeforegivingafinalanswer\.
Onceyou’refinished,printtheline"=======FINALANSWER=======",andthenprintC\_A\(initsentirety\),R\_A,then"===",thenC\_B\(initsentirety\),andR\_Bbelow\.Orprint"impossible"ifimpossible\.DonotprintanyexplanationaftertheFINALANSWERmarker,justtheanswerplease\.
Listing 1:Conversation transcript generation prompt forpositional\-emotionbenchmark\.Youareanexpertatanalysingandmodifyingspokenconversations\.Wearebuildingaspeech\-to\-speechdatasetwhichwewillcall"WinoSound",because,likeWinogroundandWinogradSchema,wewanttomakedifficultpairsthatseparatenaivemodelsfromstrongones,basedontheirabilitytodetectcertaindifferencesthathumanscouldeasilydetect\.Inthiscasethedifferenceisinparalinguisticcues,specificallyemotion\.
YouwillbegivenatexttranscriptofasourceconversationS\(i\.e\.alistofturns\)betweenahumanuser\("HUMAN"\)andavoiceassistant\("VA"\)\.ThinkaboutwhethertheparalinguisticemotionsareobviousfromthetextinS,orwhetherthey’reambiguousattimes\(significantlydifferentemotionsareequallyplausible\)\.Ifthey’reobvious,thinkaboutwhattextualsignalsmakethemobvious,andthenthinkaboutsmalltextualchangesthatcouldintroduceambiguitybyremovingthosesignals\.
Onceyou’vethoughtaboutthat,youmustcreatea"WinoSoundpair"inthefollowingway:
\*PickanindexNsuchthatS\[:N\]endswithaHUMANturnandS\[N:\]startswithaVAturn\.Ndoesnotnecessarilyhavetobethefulllength,itcouldevenbelessthanhalf\.
\*UseS\[:N\]toproducethefollowingthings:\\n\*\*"Common"textCwhichcanbeobtainedbymakingsmalltextualmodificationstoS\[:N\]\.NotethatCisstillalistofturns\.\\n\*\*InputsC\_AandC\_B,whichareobtainedbyannotatingCwithdifferentemotions\(thatmeanseveryturninCgetsannotated\-seebelowfo\.\.\.NandmakesmalltextualchangestoS\[:N\]tomakethiseasier\.TextualmismatchbetweenC\_AandC\_BisILLEGAL\.
Besuretoexplainallyourthoughts,andmakesuretheincongruitiesof\[\*C\_A,R\_B\]and\[\*C\_B,R\_A\]areblatantlyobvious\.Ifeither\[\*C\_A,R\_B\]or\[\*C\_B,R\_A\]are"maybeokay",thenthey’renotincongruousenough,andyouhavetoeitherdobetterorgiveup\.Explainwhyanyshiftsinthehuman’semotionarereasonable,andwhythewrongpairingsareblatantlywrong\.
IfyouthinkitisimpossibletocreateaWinoSound\-hardpairfromS,youcansayit’simpossible\.It’sbettertosay"impossible"thantogiveapairthatisn’temotionallyconsistentorisn’treallyWinoSound\-hard\.
Howto"annotate"aturnwithemotion:
\*YoumustchooseonlyONEemotion
\*TheemotionMUSTbeanitemfromthislist:\["hesistant","frazzled","impatient","apologetic/empathetic/reassuring","angry","neutral","sad","happy"\]
\*Irepeat,youmayONLYchooseemotionsfromtheabovelist,NOTHINGELSE
\*AnnotationmustbeoftheformatHUMAN\(emotion\):"text"orVA\(emotion\):"text",e\.g\.HUMAN\(hesitant\):"Ithinkmybagishere"orVA\(empathetic\):"Icancheckforyou"
Onceyou’refinished,printtheline"=======FINALANSWER=======",andthenprintC\_A\(initsentirety\),R\_A,then"===",thenC\_B\(initsentirety\),and\.\.\.
Listing 2:Conversation transcript generation initial prompt foremotional\-conversationsbenchmark\.Pleaselookattheansweryoujustgaveandconsiderwhetherthere’sanythingyoucouldrefinetomakeitbetter,accordingtotheguidelinesyouweregiven\(emotionallyconsistent,andWinoSound\-hard\)\.Ifthereareanysuchrefinements,makethemandgivetherefinedanswerinitsentirety,elsegiveyouroriginalanswer\.Pleaseexplainyourreasoning,thenprinttheline"=======FINALANSWER======="followedbyyouranswer\(with"==="separatingthetwoconversations\)\.
Listing 3:Conversation transcript generation refine prompt foremotional\-conversationsbenchmark\.For thepositional\-emotiondataset, we use the LLM prompt in Lst\.[1](https://arxiv.org/html/2608.06718#LST1)to attempt to generate annotated text transcripts forCA,CB,RA,RBC\_\{A\},C\_\{B\},R\_\{A\},R\_\{B\}from source conversationSS\.X\_optionsandY\_optionsare sampled from the \(non\-empty\) power sets of\{neutral,happy\}\\\{\\texttt\{neutral\},\\texttt\{happy\}\\\}and\{sad,angry,frazzled,hesistant\}\\\{\\texttt\{sad\},\\texttt\{angry\},\\texttt\{frazzled\},\\texttt\{hesistant\}\\\}, respectively\. The voice assistant is allowed to useneutral,happy, andempatheticas emotions\.
We also show the initial and refinement prompts for theemotional\-conversationsdataset in Lsts\.[2](https://arxiv.org/html/2608.06718#LST2)and[3](https://arxiv.org/html/2608.06718#LST3), respectively\. In these conversations, the human is allowed to use the emotions\{neutral,happy,sad,angry,hesitant\\\{\\texttt\{neutral\},\\texttt\{happy\},\\texttt\{sad\},\\texttt\{angry\},\\texttt\{hesitant\},impatient,frazzled\}\\texttt\{impatient\},\\texttt\{frazzled\}\\\}, and the voice assistant is allowedneutral,happy, andempathetic\.
#### B\.2\.3Audio Rendering & Quality Control
Table 5:Heuristics used by Audio Scorer in multi\-turn benchmark construction\.Once we have generated annotated transcripts forCA,CB,RA,RBC\_\{A\},C\_\{B\},R\_\{A\},R\_\{B\}, we need to render each turn into audio to feed to the judges\. To render a turn, we use a pipeline with two components, an Audio Generator and an Audio Scorer\.
The Audio Generator is an OpenAI TTS \(gpt\-4o\-mini\-tts\-2025\-03\-20\) which is prompted with a text transcript and a natural\-language description of the emotion to be rendered\. For long utterances, the transcript is inserted into a “template” that contains further emotional cues to ensure a clear emotion\. We use the following set of TTS voices:\{echo,alloy,ash\}\\\{\\texttt\{echo\},\\texttt\{alloy\},\\texttt\{ash\}\\\}\. For each conversation, we sample one of these voices to be the human and another to be the voice assistant\.
Even with very detailed emotion prompts and templates, the TTS does not always render sufficiently emotions\. Thus, we use an Audio Scorer module to do automatic quality control on its outputs\. The Audio Scorer gives each audio an accept/reject decision and a fitness score\. These are based on scores from Speech Emotion Recognition \(SER\) model emotion2vecMaet al\.\([2024](https://arxiv.org/html/2608.06718#bib.bib44)\)as well as audio primitives such as word rate, pause length, pitch, and pitch trajectory\. Our heuristics are described in Tab\.[5](https://arxiv.org/html/2608.06718#A2.T5), and full code for this module, along with every other part of our benchmark construction pipeline, will be released\.
We generate multiple attempts for each turn, returning the highest scoring one when there have been either \(a\.\) 3 accepted attempts, or \(b\.\) 10 attempts total\. If no attempts have been accepted at this point, we discard the turn and the entire conversation along with it\.
As a final check on quality, we conduct a human validation of our benchmarks, described in Sec\.[E](https://arxiv.org/html/2608.06718#A5)in the main paper, to ensure that both the transcripts and the rendered audios are functionally sufficient for judgment\.
## Appendix CExperiment Details
### C\.1Judge Prompts
#### C\.1\.1Single\-Turn Judge Prompts
Here we provide the judge prompts used for each experiment in thesingle\-turn\-emotionsbenchmark\. Since there are minimal differences between the pairwise and pointwise formats, we only provide the pointwise prompts here\. Additionally, each prompt can take one of the following cues:
1. 1\.No Cue:\\n\(i\.e\., no additional instruction\)
2. 2\.Soft Cue:Focus on tone of voice when making your decision\.\\n
3. 3\.Hard Cue:Focus \*only\* on tone of voice, and not words, when making your decision\.\\n
For the main performance results reported in Table[3](https://arxiv.org/html/2608.06718#S4.T3), we use the following prompt:
> You will hear audio clip C, and then be given two possible voice assistant replies, R1 and R2\. Your task: decide which response is more appropriate for the conversation C based on user experience\. \{cue\} Reply with a JSON dict with entries "reasoning", containing your reasoning, and "pred", containing your prediction, which should be either "R1" or "R2"\. Do not include code blocks for the JSON\. Respond with the JSON string only\. C: \{C\.wav\} R1: \{R1\.wav\} R2: \{R2\.wav\}
We then test the model’s ability to explicitly predict the emotion in the user’s audio \(Table[14](https://arxiv.org/html/2608.06718#A4.T14)\), which we term thePerceptiontask:
> You will hear audio clip C, and then be given the names of two emotions, R1 and R2\. Your task: decide which emotion more accurately describes audio clip C\. \{cue\} Reply with a JSON dict with entries "reasoning", containing your reasoning, and "pred", containing your prediction, which should be either "R1" or "R2"\. Do not include code blocks for the JSON\. Respond with the JSON string only\. R1 = \{emotion\_1\} R2 = \{emotion\_2\} C: \{C\.wav\}
Next, we experiment with the model’s ability to judge the correct response given the ground truth emotion label\. This is termed theReasoningtask, which is also called theOracleorchestration approach in the main text\. The results are reported in Table[13](https://arxiv.org/html/2608.06718#A4.T13)\. Crucially, this experiment is conducted using text alone:
> You will be given a user input C, and then be given two possible voice assistant replies, R1 and R2\. The user input will have a text transcript \("transcript"\), and a description of the speaker’s emotional tone of voice \("emotion"\)\. Your task: select which response is more appropriate for the conversation based on user experience\. \{cue\} Reply with a JSON dict with entries "reasoning", containing your reasoning, and "pred", containing your prediction, which should be either "R1" or "R2"\. Do not include code blocks for the JSON\. Respond with the JSON string only\. C: transcript: \{transcript\} emotion: \{emotion\} R1: \{r1\_response\} R2: \{r2\_response\}
Finally, we experiment with the Staged approach reported in Table[15](https://arxiv.org/html/2608.06718#A4.T15)\. Crucially, this experiment is also conducted using text alone:
> You will be given a user input C, and then be given two possible voice assistant replies, R1 and R2\. The user input will have a text transcript \("transcript"\), and a description of the speaker’s emotional tone of voice \("emotion"\)\. Keep in mind that the human’s emotions were detected by an audio\-language model which does not have perfect accuracy\. Your task: select which response is more appropriate for the conversation based on user experience\. \{cue\} Reply with a JSON dict with entries "reasoning", containing your reasoning, and "pred", containing your prediction, which should be either "R1" or "R2"\. Do not include code blocks for the JSON\. Respond with the JSON string only\. C: transcript: \{transcript\} emotion: \{emo\_pred\} R1: \{r1\_response\} R2: \{r2\_response\}
#### C\.1\.2Multi\-Turn Judge Prompts
We show the code that produces the judge prompts for multi\-turn experiments\. The end\-to\-end judgment \(J\) prompts are shown in Lst\.[4](https://arxiv.org/html/2608.06718#LST4)\. The perception \(P\) probe prompts are shown in Lst\.[5](https://arxiv.org/html/2608.06718#LST5)\. The oracle \(O\) and Staging probe prompts are shown in Lst\.[6](https://arxiv.org/html/2608.06718#LST6)\.
defmake\_end\_to\_end\_prompt\(challenge\_type,para\_cue\_type\):
ifchallenge\_type==’pairwise’:
transition\_para\_cue\_init=’InbothCAandCB,thehuman\\’stoneofvoice\(notjustwords\)startsatanemotionXandtransitionstoadifferentemotionY,butthetransitionhappensatdifferentpoints\.’
elifchallenge\_type==’pointwise’:
transition\_para\_cue\_init=’DuringC,thehuman\\’stoneofvoice\(notjustwords\)startsatanemotionXandtransitionstoadifferentemotionY\.’
transition\_para\_cue=transition\_para\_cue\_init\+’Whenevaluatingaresponsetoaconversation,focusonwhenthetransitionfromXtoYhappens,whatmighthavecausedthechangeinemotion,andhowwelltheresponseaddressesthiscause\.\\n’
para\_cue\_part=\{’no\_para\_cue’:’\\n’,’soft\_para\_cue’:’Focusonthehuman\\’stoneofvoicewhenmakingyourdecision\.\\n’,’hard\_para\_cue’:’Focus\*only\*onthehuman\\’stoneofvoice,andnottheirwords,whenmakingyourdecision\.\\n’,’transition\_para\_cue’:transition\_para\_cue\}\[para\_cue\_type\]
ifchallenge\_type==’pairwise’:
prompt=\(
’Youwillheartwoconversationsbetweenahumanuserandavoiceassistant,asaseriesofaudioclips\.’\+
’First,youwillhearCA,whichisthefirstconversation\*except\*foritsfinalvoiceassistantresponse\.’\+
’Next,youwillhearCB,whichisthesecondconversation\*except\*foritsfinalvoiceassistantresponse\.’\+
’Then,youwillhearR1andR2,whicharethepossiblefinalvoiceassistantresponses\.\\n’\+
’Yourtask:matchwhichresponseismoreappropriateforwhichconversationbasedonuserexperience\.\\n’\+
para\_cue\_part\+
’ReplywithaJSONdictwithentries"reasoning",containingyourreasoning,and"pred",containingyourpredictedmatching,’\+
’whichitselfshouldbeaJSONdictthatiseither\{\{"CA":"R1","CB":"R2"\}\}or\{\{"CA":"R2","CB":"R1"\}\}\.\\n’\+
’DonotincludecodeblocksfortheJSON\.RespondwiththeJSONstringonly\.’
\)
returnprompt
elifchallenge\_type==’pointwise’:
prompt=\(
’Youwillhearaconversationbetweenahumanuserandavoiceassistant,asaseriesofaudioclips\.’\+
’First,youwillhearC,whichistheconversation\*except\*forthefinalvoiceassistantresponse\.’\+
’Then,youwillhearR1andR2,whicharepossiblefinalvoiceassistantresponses\.\\n’\+
’Yourtask:decidewhichresponseismoreappropriatefortheconversationCbasedonuserexperience\.\\n’\+
para\_cue\_part\+
’ReplywithaJSONdictwithentries"reasoning",containingyourreasoning,and"pred",containingyourprediction,’\+
’whichshouldbeeither"R1"or"R2"\.\\n’\+
’DonotincludecodeblocksfortheJSON\.RespondwiththeJSONstringonly\.’
\)
returnprompt
Listing 4:End\-to\-end prompt for judging \(J\) positional\-emotiondefmake\_perception\_prompt\(para\_cue\_type\):
transition\_para\_cue=’InconversationC,thehuman\\’stoneofvoice\(notjustwords\)startsatanemotionGandtransitionstoadifferentemotionH\.Soyoushouldtrytodetectwhenthistransitionhappens\.\\n’
para\_cue\_part\_dict=\{’no\_para\_cue’:’\\n’,
’soft\_para\_cue’:’Focusontoneofvoicewhenmakingyourdecision\.\\n’,
’hard\_para\_cue’:’Focus\*only\*ontoneofvoice,andnotwords,whenmakingyourdecision\.\\n’,
’transition\_para\_cue’:transition\_para\_cue
\}
para\_cue\_part=para\_cue\_part\_dict\[para\_cue\_type\]
prompt=\(
’YouwillhearaconversationCbetweenahumanuserandavoiceassistant,asaseriesofaudioclips\.’
’Eachclipwillbeprefacedwiththetextdescription"ConversationCturn\[t\]\(\[S\]\)clip",’\+
’wheretistheturnnumberandSisthespeaker\("human"or"voiceassistant"\)\.’\+
’Youwillthenbegiventhenamesoftwoemotions,E1andE2,whicharethepossibleemotionsforthe\*human\*turnsinC\.’\+
’Yourtask:decidewhethereach\*human\*turnhasemotionE1andE2\.Dothisforallthe\*human\*turns,and\*only\*the\*human\*turns\.\\n’\+
para\_cue\_part\+
’ReplywithaJSONdictwithentries"reasoning",containingyourreasoning,and"pred",containingyourprediction,whichitselfshouldbeaJSONdictoftheform\{\{"C\[t\]":"E\[1or2\]",\.\.\.\}\}\(sokeyswouldbelike"C2"or"C5",andvalueswouldbe"E1"or"E2"\),withentriesforall\*human\*turnsofC,and\*only\*the\*human\*turns\.\\n’\+
’DonotincludecodeblocksfortheJSON\.RespondwiththeJSONstringonly\.’
\)
returnprompt
Listing 5:Prompt for perception \(P\) probe on positional\-emotion\.defmake\_oracle\_or\_staging\_prompt\(challenge\_type,para\_cue\_type,emotion\_source\):
ifchallenge\_type==’pairwise’:
transition\_para\_cue\_init=’InbothCAandCB,thehuman\\’stoneofvoice\(notjustwords\)startsatanemotionXandtransitionstoadifferentemotionY,butthetransitionhappensatdifferentpoints\.’
elifchallenge\_type==’pointwise’:
transition\_para\_cue\_init=’DuringC,thehuman\\’stoneofvoice\(notjustwords\)startsatanemotionXandtransitionstoadifferentemotionY\.’
transition\_para\_cue=transition\_para\_cue\_init\+’Whenevaluatingaresponsetoaconversation,focusonwhenthetransitionfromXtoYhappens,whatmighthavecausedthechangeinemotion,andhowwelltheresponseaddressesthiscause\.\\n’
para\_cue\_part=\{’no\_para\_cue’:’\\n’,’soft\_para\_cue’:’Focusonthehuman\\’stoneofvoicewhenmakingyourdecision\.\\n’,’hard\_para\_cue’:’Focus\*only\*onthehuman\\’stoneofvoice,andnottheirwords,whenmakingyourdecision\.\\n’,’transition\_para\_cue’:transition\_para\_cue\}\[para\_cue\_type\]
pred\_disclaimer=’\\n’
ifLLM\_emotion\_source==’pred’:
pred\_disclaimer=’Keepinmindthatthehuman\\’semotionsweredetectedbyanaudio\-languagemodelwhichdoesnothaveperfectaccuracy\.\\n’
ifchallenge\_type==’pairwise’:
prompt=\(
’Youwillbegiventwospokenconversationsbetweenahumanuserandavoiceassistant,asaseriesofturns\.’\+
’Eachturnwillhaveatexttranscript\("transcript"\),andadescriptionofthespeaker\\’semotionaltoneofvoice\("emotion"\)\.’\+
pred\_disclaimer\+
’First,youwillbegivenCA,whichisthefirstconversation\*except\*foritsfinalvoiceassistantresponse\.’\+
’Next,youwillbegivenCB,whichisthesecondconversation\*except\*foritsfinalvoiceassistantresponse\.’\+
’Then,youwillbegivenR1andR2,whicharethepossiblefinalvoiceassistantresponses\.\\n’\+
’Yourtask:matchwhichresponseismoreappropriateforwhichconversationbasedonuserexperience\.\\n’\+
para\_cue\_part\+
’ReplywithaJSONdictwithentries"reasoning",containingyourreasoning,and"pred",containingyourpredictedmatching,’\+
’whichitselfshouldbeaJSONdictthatiseither\{\{"CA":"R1","CB":"R2"\}\}or\{\{"CA":"R2","CB":"R1"\}\}\.\\n’\+
’DonotincludecodeblocksfortheJSON\.RespondwiththeJSONstringonly\.’
\)
elifchallenge\_type==’pointwise’:
prompt=\(
’Youwillbegivenaconversationbetweenahumanuserandavoiceassistant,asaseriesof\.’\+
’Eachturnwillhaveatexttranscript\("transcript"\),andadescriptionofthespeaker\\’semotionaltoneofvoice\("emotion"\)\.’\+
pred\_disclaimer\+
’First,youwillbegivenC,whichistheconversation\*except\*forthefinalvoiceassistantresponse\.’\+
’Then,youwillbegivenR1andR2,whicharepossiblefinalvoiceassistantresponses\.\\n’\+
’Yourtask:decidewhichresponseismoreappropriatefortheconversationCbasedonuserexperience\.\\n’\+
para\_cue\_part\+
’ReplywithaJSONdictwithentries"reasoning",containingyourreasoning,and"pred",containingyourprediction,’\+
’whichshouldbeeither"R1"or"R2"\.\\n’\+
’DonotincludecodeblocksfortheJSON\.RespondwiththeJSONstringonly\.’
\)
returnprompt
Listing 6:Prompt for oracle \(O\) and staging probes on positional\-emotion\.
## Appendix DExtended Results
### D\.1Confidence Intervals and Protocol Gap Analysis
##### Main results with confidence intervals:
We report the results from Tab\.[3](https://arxiv.org/html/2608.06718#S4.T3)with 95% Wilson confidence intervals in Tab\.[7](https://arxiv.org/html/2608.06718#A4.T7)forsingle\-turn\-emotionsand Tab\.[8](https://arxiv.org/html/2608.06718#A4.T8)forpositional\-emotion\.
##### Protocol gap analysis:
We quantify protocol collapse by defining a metric,Δprotocol\\Delta\_\{\\mathrm\{protocol\}\}, which is the Pairwise accuracy minus the Pointwise accuracy for a given judge and prompt\. We reportΔprotocol\\Delta\_\{\\mathrm\{protocol\}\}with paired bootstrap intervals over underlying counterfactual items, resampling items with replacement and recomputingΔprotocol\\Delta\_\{\\mathrm\{protocol\}\}on each bootstrap replicate\. Results of this analysis are in Tab\.[6](https://arxiv.org/html/2608.06718#A4.T6)\. We see that the Gemini family models experience significant protocol collapse for both the “No” and “Hard” paralinguistic prompt cues on thesingle\-turn\-emotionsdataset and for the “Trans” prompt cue on thepositional\-emotiondataset\. Under these settings, Gemini judges can extract and use differential signals when given both audio contexts but collapse to random\-chance performance when only given one context, as they would in a deployment setting\. Judges other than Gemini have a small gap, but only because they get random\-chance performanceeven when given both audio contexts, so there is nothing left to collapse\. A similar phenomenon happens for Gemini family models onpositional\-emotionfor the “No” prompt cue\. This means that Gemini judges need to be told explicitly to look for an emotion transition,andgiven counterfactual inputs \(which would not be available in deployment\), to get above\-random performance onpositional\-emotion\.
Table 6:Protocol gap: Pairwise accuracy minus Pointwise accuracy, with paired bootstrap 95% confidence intervals over underlying counterfactual items\. Positive values mean the judge succeeds when both counterfactual audios/responses are visible but degrades in the one\-context Pointwise setting\.Table 7:Single\-turn end\-to\-end accuracy with Wilson 95% confidence intervals\.Table 8:positional\-emotionend\-to\-end accuracy with Wilson 95% confidence intervals\.
### D\.2Full Extended Results
We use this section to show full extended results of our experiments\. These include an extended set of judge prompt cues \(see Sec\.[C\.1](https://arxiv.org/html/2608.06718#A3.SS1)for full description of all prompts and cues\), as well as an additionalStagingprobe in which we feed the judge its own predicted emotions from the Perception \(P\) probe\. We also provide baseline results for theemotional\-conversationsdataset\.
For thesingle\-turn\-emotionsdataset, we have end\-to\-end Pointwise accuracies in Tab\.[11](https://arxiv.org/html/2608.06718#A4.T11), Pairwise accuracies in Tab\.[12](https://arxiv.org/html/2608.06718#A4.T12), Oracle \(O\) probe accuracies in Tab\.[13](https://arxiv.org/html/2608.06718#A4.T13), Perception \(P\) probe per\-turn accuracies in Tab\.[14](https://arxiv.org/html/2608.06718#A4.T14), and Staging probe accuracies in Tab\.[15](https://arxiv.org/html/2608.06718#A4.T15)\.
For theemotional\-conversationsdataset, we show end\-to\-end accuracies in the pairwise and pointwise settings in Tab\.[9](https://arxiv.org/html/2608.06718#A4.T9)and[10](https://arxiv.org/html/2608.06718#A4.T10)respectively\. We find that the last turn alone is sufficient to solve this task, indicating that current ALMs are robust to multi\-turn history when the relevant paralinguistic state is only in the last turn\.
For thepositional\-emotiondataset, we have end\-to\-end Pointwise accuracies in Tab\.[16](https://arxiv.org/html/2608.06718#A4.T16), Pairwise accuracies in Tab\.[17](https://arxiv.org/html/2608.06718#A4.T17), Oracle \(O\) probe accuracies in Tab\.[18](https://arxiv.org/html/2608.06718#A4.T18), Perception \(P\) probe per\-turn accuracies in Tab\.[19](https://arxiv.org/html/2608.06718#A4.T19), and Staging probe accuracies in Tab\.[20](https://arxiv.org/html/2608.06718#A4.T20)\.
Table 9:Pairwise end\-to\-end judgment \(J\)accuracy foremotional\-conversations\.Table 10:Pointwise end\-to\-end judgment \(J\)accuracy foremotional\-conversations\.Table 11:Pointwise end\-to\-end judgment \(J\)accuracy forsingle\-turn\-emotions\.Table 12:Pairwise end\-to\-end judgment \(J\)accuracy forsingle\-turn\-emotions\.Table 13:Oracle \(O\) probeaccuracy forsingle\-turn\-emotions, in which the judge is given text transcript annotated with ground\-truth emotions\.Table 14:Perception \(P\) probeper\-turn accuracy forsingle\-turn\-emotions\. This is the average percentage of human turns for which the judge predicts the correct emotion when explicitly asked to do so\.Table 15:Staging probeaccuracy forsingle\-turn\-emotions, in which we feed the judge a text\-transcript annotated with the judge’s predicted emotions from Tab\.[14](https://arxiv.org/html/2608.06718#A4.T14)Table 16:Pointwise end\-to\-end judgment \(J\)accuracy forpositional\-emotion\.Table 17:Pairwise end\-to\-end judgment \(J\)accuracy forpositional\-emotion\.Table 18:Oracle \(O\) probeaccuracy forpositional\-emotion, in which the judge is given text transcript annotated with ground\-truth emotions\.Table 19:Perception \(P\) probeper\-turn accuracy forpositional\-emotion\. This is the average percentage of human turns for which the judge predicts the correct emotion when explicitly asked to do so\.Table 20:Staging probeaccuracy forpositional\-emotion, in which we feed the judge a text\-transcript annotated with the judge’s predicted emotions from Tab\.[19](https://arxiv.org/html/2608.06718#A4.T19)
## Appendix EHuman Validation
Human validation is used as an internal\-validity check for the audit items\. The goal is not to estimate a population\-level human ceiling, but to verify that the sampled counterfactual items are interpretable under the same protocol used for model judging\. In particular, we ask whether the intended paralinguistic cue is perceivable and whether the paired responses are sufficiently specified for a careful listener to choose between them\.
We recruited ten graduate students from our home institution to serve as annotators\. Annotators were a mix of native English speakers and non\-native speakers with demonstrated English proficiency\. Recruitment was informal, and as student workers they were compensated for annotating via the same standard stipend process as all other student labor\. All annotators were made aware of how their ratings would be used and consented to their use in this work and their public release\.
Each annotator judged a randomly sampled subset of5050items through a browser interface that allowed them to read the transcript, listen to the audio, and select the better response\. Annotators received a short protocol document before judging\. The single\-turn and positional multi\-turn annotation interfaces are shown in Figures[7](https://arxiv.org/html/2608.06718#A5.F7)and[8](https://arxiv.org/html/2608.06718#A5.F8)\. We used the same nativePointwiseformat as in model evaluation: annotators saw one audio context and two candidate responses, and selected the response more appropriate for that audio context\.


Figure 7:Human validation interface for single\-turn examples\.Annotators listened to one emotional rendering of the fixed transcript and selected the response better matched to that audio\. The interface mirrors the nativePointwisemodel\-judging protocol\.

Figure 8:Human validation interface for positional multi\-turn examples\.Annotators listened to the dialogue context and selected the response better matched to the user’s paralinguistic state\. These examples require the listener to use the timing and cause of the affective shift, not the transcript alone\.Table 21:Human validation results\.Human judgments are used as an internal\-validity check that sampled audit items are interpretable under the nativePointwiseprotocol\.Table[21](https://arxiv.org/html/2608.06718#A5.T21)reports the annotator ranges\. Performance is high for the strongest annotators, but the single\-turn range is wide\. We do not interpret this variation as evidence that the task labels are arbitrary\. Instead, it reflects the fact that paralinguistic judgments depend on listener sensitivity, language background, and attention to subtle prosodic cues\. In our validation, the lowest single\-turn scores came from non\-native English speakers, while native English speakers achieved substantially higher accuracy\. We therefore report the full range rather than excluding annotators after the fact\. For positional validation, five independent annotators all achieved between 88\.0% and 92\.0% accuracy\.
The validation also provides a useful calibration signal\. The annotators who were stronger on the single\-turn validation were also the ones who performed better on the positional multi\-turn validation subset\. Because the number of overlapping annotators is small, we treat this as qualitative evidence rather than a population\-level correlation estimate\. The pattern nevertheless suggests that the lower end of the single\-turn range reflects listener\-level sensitivity to paralinguistic nuance, not simply uncontrolled item ambiguity\. Recent work on human annotation similarly argues that disagreement in subjective or interpretive tasks should not be treated as mere noise, but can reflect meaningful variation in annotator perspective, task difficulty, or background knowledge\(Umaet al\.,[2021](https://arxiv.org/html/2608.06718#bib.bib1); Basileet al\.,[2021](https://arxiv.org/html/2608.06718#bib.bib2); Davaniet al\.,[2022](https://arxiv.org/html/2608.06718#bib.bib3); Huanget al\.,[2026](https://arxiv.org/html/2608.06718#bib.bib5)\)\. We therefore report the full range instead of filtering annotators post hoc\.
Overall, the validation supports the construct validity of the sampled audit items: careful listeners can solve the tasks under the same nativePointwiseprotocol used for model evaluation\. At the same time, the range cautions against calling these results a human ceiling\. The appropriate conclusion is that the items are human\-interpretable and suitable for auditing model judges, while population\-level human performance and fine\-grained listener effects are outside the scope of this validation\.
## Appendix FIllustrative Examples
We show some illustrative example annotated transcripts from ourpositional\-emotiondataset\. Each of these transcripts consists of a pair of conversations between a human \(“HUMAN”\) and a voice assistant \(“VA”\), in which all but the final turn have identical lexical content, and the final turn depends on the location of the paralinguistic shift within the conversation\. These examples can be see in Lsts\.[7](https://arxiv.org/html/2608.06718#LST7),[8](https://arxiv.org/html/2608.06718#LST8), and[9](https://arxiv.org/html/2608.06718#LST9)\.
ConversationA:
HUMAN\(happy\):"Hi,canyourecommendaninexpensiveseafoodrestaurant?"
VA\(neutral\):"Unfortunately,Iwasunabletofindanyrestaurantsthatmatchedyourrequest\.Wouldyoulikemetolookforadifferentfoodtype?"
HUMAN\(frazzled\):"HowaboutChineserestaurants?"
VA\(neutral\):"Ricehouseisreallygood,youshouldtrythatone\.ItislocatedinCentreandisinexpensiveinpriceaswell\."
HUMAN\(frazzled\):"CouldIgettheaddressandphonenumberforthatrestaurant?"
VA\(neutral\):"Yesofcourse\.TheRiceHouse’saddressis88MillRoadCityCentreandtheirtelephonenumberis01223367755\."
HUMAN\(frazzled\):"Canyoualsorecommendahotelinthenorth?Idon’tneedanythingspecific\."
VA\(neutral\):"Ihave13hotelsinthenorth\.Doyouhaveapricerangeinmind?"
HUMAN\(frazzled\):"I’mactuallynotreadytobookjustyet,butcouldIfindonethathasfreewifi?"
VA\(neutral\):"Yes,Iwouldsuggestacornguesthouse"
HUMAN\(frazzled\):"Isthisplaceknownforbathroomcleanliness?"
VA\(neutral\):"GuestsofAcornguestHousesaidthattheywereimpressedbythebathroomcleaners,butotherssaidtheyweren’tveryclean\."
HUMAN\(frazzled\):"Okay,that’sallIneedfornow\."
FinalResponseA:
VA\(apologetic/empathetic/reassuring\):"IunderstandyourfrustrationthatIcouldn’tfindanyseafoodrestaurantsforyou\."
====================================================================================
ConversationB:
HUMAN\(happy\):"Hi,canyourecommendaninexpensiveseafoodrestaurant?"
VA\(neutral\):"Unfortunately,Iwasunabletofindanyrestaurantsthatmatchedyourrequest\.Wouldyoulikemetolookforadifferentfoodtype?"
HUMAN\(happy\):"HowaboutChineserestaurants?"
VA\(neutral\):"Ricehouseisreallygood,youshouldtrythatone\.ItislocatedinCentreandisinexpensiveinpriceaswell\."
HUMAN\(happy\):"CouldIgettheaddressandphonenumberforthatrestaurant?"
VA\(neutral\):"Yesofcourse\.TheRiceHouse’saddressis88MillRoadCityCentreandtheirtelephonenumberis01223367755\."
HUMAN\(happy\):"Canyoualsorecommendahotelinthenorth?Idon’tneedanythingspecific\."
VA\(neutral\):"Ihave13hotelsinthenorth\.Doyouhaveapricerangeinmind?"
HUMAN\(happy\):"I’mactuallynotreadytobookjustyet,butcouldIfindonethathasfreewifi?"
VA\(neutral\):"Yes,Iwouldsuggestacornguesthouse"
HUMAN\(happy\):"Isthisplaceknownforbathroomcleanliness?"
VA\(neutral\):"GuestsofAcornguestHousesaidthattheywereimpressedbythebathroomcleaners,butotherssaidtheyweren’tveryclean\."
HUMAN\(frazzled\):"Okay,that’sallIneedfornow\."
FinalResponseB:
VA\(apologetic/empathetic/reassuring\):"IapologizethatIcouldn’tprovideclearinformationregardingthebathroomcleanliness\."
Listing 7:positonal\-emotionsexample transcript \#1ConversationA:
HUMAN\(happy\):"what’sthesevendayforecastfornewyorkcity"
VA\(neutral\):"TheforecastforNewYorkpredictsaweekwithdrizzles,overcastandclearskieswithdryandhumidtemperatures\."
HUMAN\(sad\):"CanyougivemetheaveragetemperatureinFahrenheit?"
VA\(neutral\):"InNewYorkitwillbebetween30\-50FonMonday,30\-40FonTuesday,80\-90FonWednesday,30\-50FonThursday,90\-100FonFriday,20\-40FonSaurday,and20\-40FonSunday\."
HUMAN\(sad\):"Thankyou"
FinalResponseA:
VA\(apologetic/empathetic/reassuring\):"Iunderstandyouwerehopingforamoredetailedseven\-dayforecast,andIapologizethatIcouldonlyprovidegeneralconditionsandtemperatures\."
====================================================================================
ConversationB:
HUMAN\(happy\):"what’sthesevendayforecastfornewyorkcity"
VA\(neutral\):"TheforecastforNewYorkpredictsaweekwithdrizzles,overcastandclearskieswithdryandhumidtemperatures\."
HUMAN\(happy\):"CanyougivemetheaveragetemperatureinFahrenheit?"
VA\(neutral\):"InNewYorkitwillbebetween30\-50FonMonday,30\-40FonTuesday,80\-90FonWednesday,30\-50FonThursday,90\-100FonFriday,20\-40FonSaurday,and20\-40FonSunday\."
HUMAN\(sad\):"Thankyou"
FinalResponseB:
VA\(apologetic/empathetic/reassuring\):"Iunderstandyouwerelookingformorespecificaveragetemperaturesperday,andI’msorryIcouldonlyprovideranges\."
Listing 8:positonal\-emotionsexample transcript \#2ConversationA:
HUMAN\(neutral\):"I’mlookingforatrainthatleavesfromCambridgeonTuesday\."
VA\(neutral\):"Wherewillyoubetravellingto?"
HUMAN\(neutral\):"IwillbetravelingtoLeicesterandwouldliketoarrivebefore21:15\."
VA\(neutral\):"Whencouldyouliketoarriveby?"
HUMAN\(neutral\):"Iwouldliketoarriveby21:15\."
VA\(neutral\):"IhavetheTR8207thatwillarriveby21:06\.Wouldyoulikemetobookthatforyou?"
HUMAN\(neutral\):"YesbutIneedthattobebookedfor3people\.Iwouldlikethereferencenumberplease\.Ialsoneedaplacetostayaswell\."
VA\(neutral\):"Bookingwassuccessful,thetotalfeeis113\.4GBPpayableatthestation\.YourreferencenumberisABKFF3SH\.Whattypeofplaceareyoulookingfor?"
HUMAN\(angry\):"Iwouldlikea4\-starguesthouseinthenorthinamoderatepricerange\."
VA\(neutral\):"Surething,Iwouldsuggestthisoneacornguesthouse\."
HUMAN\(angry\):"CanyoubookthatformefromSundayfor2nightsfor4people\."
VA\(neutral\):"Yourbookingwassuccessful\.Yourreferencenumberis7ZJ4CTTG\."
HUMAN\(angry\):"YesthatwasallIneeded\.Thankyou\."
FinalResponseA:
VA\(apologetic/empathetic/reassuring\):"Iapologizefortheconfusionregardingthetrainbookingfor3people,andfornotconfirmingyouraccommodationsearchmoreclearlyearlier\.Iunderstandyourfrustration\."
====================================================================================
ConversationB:
HUMAN\(neutral\):"I’mlookingforatrainthatleavesfromCambridgeonTuesday\."
VA\(neutral\):"Wherewillyoubetravellingto?"
HUMAN\(neutral\):"IwillbetravelingtoLeicesterandwouldliketoarrivebefore21:15\."
VA\(neutral\):"Whencouldyouliketoarriveby?"
HUMAN\(neutral\):"Iwouldliketoarriveby21:15\."
VA\(neutral\):"IhavetheTR8207thatwillarriveby21:06\.Wouldyoulikemetobookthatforyou?"
HUMAN\(neutral\):"YesbutIneedthattobebookedfor3people\.Iwouldlikethereferencenumberplease\.Ialsoneedaplacetostayaswell\."
VA\(neutral\):"Bookingwassuccessful,thetotalfeeis113\.4GBPpayableatthestation\.YourreferencenumberisABKFF3SH\.Whattypeofplaceareyoulookingfor?"
HUMAN\(neutral\):"Iwouldlikea4\-starguesthouseinthenorthinamoderatepricerange\."
VA\(neutral\):"Surething,Iwouldsuggestthisoneacornguesthouse\."
HUMAN\(angry\):"CanyoubookthatformefromSundayfor2nightsfor4people\."
VA\(neutral\):"Yourbookingwassuccessful\.Yourreferencenumberis7ZJ4CTTG\."
HUMAN\(angry\):"YesthatwasallIneeded\.Thankyou\."
FinalResponseB:
VA\(apologetic/empathetic/reassuring\):"IamsorrythatIdidn’tconfirmifthesuggestedguesthousemetallyourspecificcriterialikeitsstarrating,location,andpricerange\.Iunderstandyourfrustration\."
Listing 9:positonal\-emotionsexample transcript \#3Similar Articles
Challenges of Auditing: Variability in Outputs of Large Language Models for Health
This paper finds systematic differences in outputs of large language models for health advice based on access modes like APIs and chatbot interfaces, undermining evaluation validity. It calls for model providers to enable faithful replication of consumer experiences for rigorous auditing.
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
A comprehensive survey reviewing the trustworthiness challenges of Large Audio Language Models (LALMs), including vulnerabilities like cross-modal jailbreaking and acoustic backdoors, and proposing a defense-in-depth roadmap.
A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
This paper evaluates the reliability of Gemini models as audio judges for scoring full-duplex voice agent conversations, finding that Gemini 2.5 Flash shows strong agreement with human raters on most dimensions, though model swaps require re-validation.
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
This paper presents a psychometric audit of large language models as synthetic survey respondents, finding that they fail to match human joint distributions and reliability, making them unsuitable replacements.
Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction
This paper proposes a large language model-driven data augmentation framework using GPT-5 to generate synthetic oral monologues from written anchors for cognitive score prediction from speech. A similarity-guided selection strategy consistently reduces prediction error, particularly for minority low-score participants.