Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

arXiv cs.CL Papers

Summary

The paper introduces Hear2Act, a unified benchmark for evaluating how prosodic cues affect task-oriented dialogue decisions, finding that prosody matters when lexical evidence is insufficient and that audio LLMs benefit from explicit concern representation for actions.

arXiv:2608.19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions. We introduce Hear2Act, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes. For each scenario, we keep the task and user needs fixed while varying whether the same concern is conveyed explicitly in words or primarily through prosody, and evaluate decisions under transcript, audio, and concern-state access. Using Hear2Act, we evaluate two audio-capable LLMs. Under Prosody-mediated feedback, adding audio to the transcript changes the average optimal-solution rate only from 14.6% to 15.3%. In contrast, when models infer the concern status from audio, represent it in text, and use it for next-action selection, the rate rises to 39.6%, close to 40.7% with the ground-truth state. This contrast, however, largely disappears under Explicit lexical feedback, where the concern is verbally mentioned in the utterance. Together, these results show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do not reliably carry it into action without an explicit intermediate representation.
Original Article
View Cached Full Text

Cached at: 08/21/26, 10:04 AM

# Benchmarking When Prosody Should ChangeWhat an Assistant Does
Source: [https://arxiv.org/html/2608.19515](https://arxiv.org/html/2608.19515)
## Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Xinyi Liu1,2,Hooshang Nayyeri1,Dilek Hakkani\-Tur1,2, Emine Yilmaz1,3,JK Kim1,Yifei Zhang1, Charith Peris1,Hari Thadakamalla1 1Amazon2University of Illinois Urbana\-Champaign3University College London\{liu323, dilek\}@illinois\.edu\{hooshang, jookyk, jimmyzyf, perisc, thadakah\}@amazon\.comeminey@amazon\.co\.uk

###### Abstract

Prosodic cues can convey task\-relevant information that alters the trajectory and outcome of a task\-oriented dialogue, even when the words themselves remain unchanged\. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task\-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions\. We introduceHear2Act, a unified evaluation protocol for text and spoken assistants with 480 persona\-grounded scenarios, hidden user concerns, and objectively verifiable outcomes\. For each scenario, we keep the task and user needs fixed while varying whether the same concern is conveyed explicitly in words or primarily through prosody, and evaluate decisions under transcript, audio, and concern\-state access\.

UsingHear2Act, we evaluate two audio\-capable LLMs\. Under Prosody\-mediated feedback, adding audio to the transcript changes the average optimal\-solution rate only from 14\.6% to 15\.3%\. In contrast, when models infer the concern status from audio, represent it in text, and use it for next\-action selection, the rate rises to 39\.6%, close to 40\.7% with the ground\-truth state\. This contrast, however, largely disappears under Explicit lexical feedback, where the concern is verbally mentioned in the utterance\. Together, these results show that prosody matters when lexical evidence is insufficient, and that audio\-capable LLMs can recover information from speech but do not reliably carry it into action without an explicit intermediate representation\.

## 1Introduction

Task\-oriented dialogue \(TOD\) systems guide users toward concrete decisions through multi\-turn interaction\([3](https://arxiv.org/html/2608.19515#bib.bib10);[19](https://arxiv.org/html/2608.19515#bib.bib18);[21](https://arxiv.org/html/2608.19515#bib.bib6)\)\. In spoken TOD, these decisions may depend on information beyond the transcript\. A hesitant “Okay, sounds good” may indicate that an important concern remains unresolved, while the same words delivered brightly may signal genuine acceptance\. This distinction can determine whether the assistant continues gathering information or commits prematurely to an unsuitable solution\. The resulting evaluation problem is to determine when prosodic evidence should guide information gathering and final selection, and how effectively current assistants use it\. The central question is therefore not only whether a model perceives a prosodic cue, but when that cue provides information beyond the words and whether the model carries it into subsequent task decisions\.

![Refer to caption](https://arxiv.org/html/2608.19515v1/hear2act_framework_crop.png)Figure 1:Hear2Actoverview\.\(1\) Each scenario combines a surface request, prioritized hidden concerns, and candidates with verifiable satisfaction signatures\. \(2\) The assistant interacts under different feedback and access conditions\. \(3\) Matched rollouts are compared on task outcomes and interaction behavior\.Existing benchmarks address complementary parts of this problem, but not the full link from prosodic evidence to task action\. Task\-oriented dialogue benchmarks support multi\-turn interaction and verifiable success, but do not isolate how prosodic feedback changes sequential decisions\([3](https://arxiv.org/html/2608.19515#bib.bib10);[19](https://arxiv.org/html/2608.19515#bib.bib18);[21](https://arxiv.org/html/2608.19515#bib.bib6)\)\. Paralinguistic benchmarks preserve vocal cues, but primarily evaluate perception, response appropriateness, or emotional interaction without task\-grounded hidden needs and verifiable outcomes\([1](https://arxiv.org/html/2608.19515#bib.bib7);[11](https://arxiv.org/html/2608.19515#bib.bib21);[8](https://arxiv.org/html/2608.19515#bib.bib22);[24](https://arxiv.org/html/2608.19515#bib.bib23)\)\. StyleTalk and ParaS2S control lexical content, but evaluate the appropriate next response rather than the resulting multi\-turn task trajectory\([16](https://arxiv.org/html/2608.19515#bib.bib19);[29](https://arxiv.org/html/2608.19515#bib.bib20)\)\. To our knowledge, no existing protocol traces prosodic evidence through multi\-turn task decisions to verifiable outcomes while evaluating text and spoken assistants within the same task\.

We introduceHear2Actto fill this gap\. It contains 480 controlled user\-assistant scenarios, each pairing an initial request and prioritized hidden concerns with candidate options whose fit can be scored objectively\. A deterministic user engine maps each assistant action to a structured user reaction and permissible disclosure\. The resulting feedback is realized as a natural utterance and, for spoken conditions, rendered as speech, with the concern either explicit in words or conveyed primarily through prosody\. Matched rollouts keep the task and user needs fixed while varying the assistant’s access to transcript, audio, and explicit concern\-state information\. This design isolates both the decision value of prosodic evidence and where that value is lost between audio and action\.

The results address three questions: when prosodic information has decision value, whether spoken assistants can carry that information into action, and how correct concern information changes the dialogue\. Under Prosody\-mediated feedback, adding audio to the transcript changes the average optimal\-solution rate across two audio\-capable LLMs only from 14\.6% to 15\.3%\. When the models instead infer the concern status from audio, express it in text, and use it for next\-action selection, the rate rises to 39\.6%, close to the 40\.7% ground\-truth\-state reference\. This contrast largely disappears under Explicit lexical feedback, where the relevant concern is already stated in words\. Further analysis shows that correct turn\-specific concern information shifts interaction toward elicitation and helps the assistant determine when to continue searching and when to stop\. Together, these results show that prosody is most useful when lexical evidence is insufficient, yet current audio\-capable LLMs make limited use of that information directly from speech for downstream decisions\.

## 2Related Work

Table 1:Benchmark positioning\.P: prosodic input, C: matched control of prosodic access with fixed lexical content, M: multi\-turn task decisions, N: task\-grounded hidden user need, and O: verifiable trajectory and outcome\.△\\trianglemarks structured user goals conveyed lexically rather than hidden needs\.BenchmarkPCMNOMultiWOZ, SGD\([3](https://arxiv.org/html/2608.19515#bib.bib10);[19](https://arxiv.org/html/2608.19515#bib.bib18)\)×\\times×\\times✓△\\triangle✓SpokenWOZ\([21](https://arxiv.org/html/2608.19515#bib.bib6)\)✓×\\times✓△\\triangle✓StyleTalk, ParaS2S\([16](https://arxiv.org/html/2608.19515#bib.bib19);[29](https://arxiv.org/html/2608.19515#bib.bib20)\)✓✓×\\times×\\times×\\timesMULTI\-Bench, HumDial\-EIBench\([8](https://arxiv.org/html/2608.19515#bib.bib22);[24](https://arxiv.org/html/2608.19515#bib.bib23)\)✓×\\times✓×\\times×\\timesHear2Act \(ours\)✓✓✓✓✓![Refer to caption](https://arxiv.org/html/2608.19515v1/hear2act_example.png)Figure 2:Illustrative Hear2Act trajectory under Prosody\-mediated feedback\.Three representative candidates are shown from the full 11\-candidate set\. Transcript\-only access may confirm prematurely, while ground\-truth concern\-state access supports further elicitation and selection of the best\-fitting option\.### 2\.1Task\-Oriented and Spoken Task Dialogue

MultiWOZ and the Schema\-Guided Dialogue dataset established multi\-domain task\-oriented dialogue with structured states and verifiable task success\([3](https://arxiv.org/html/2608.19515#bib.bib10);[19](https://arxiv.org/html/2608.19515#bib.bib18)\), while ATOD extends evaluation to agentic capabilities\([31](https://arxiv.org/html/2608.19515#bib.bib2)\)\. SpokenWOZ extends this setting to speech\([21](https://arxiv.org/html/2608.19515#bib.bib6)\), while RealTalk\-CN studies speech–text interaction\([23](https://arxiv.org/html/2608.19515#bib.bib24)\)\. Emotion\-aware work adds affect annotations or user simulation\([10](https://arxiv.org/html/2608.19515#bib.bib11);[17](https://arxiv.org/html/2608.19515#bib.bib12);[9](https://arxiv.org/html/2608.19515#bib.bib13)\), and classical POMDP formulations treat user state as latent\([30](https://arxiv.org/html/2608.19515#bib.bib3)\)\. These approaches study task success, agentic behavior, speech, affect, or latent state, but do not isolate when prosodic evidence adds decision value beyond the transcript or changes downstream task actions\.

### 2\.2Paralinguistic and Spoken Evaluation

Paralinguistic benchmarks mainly target recognition, response appropriateness, or multi\-turn affect\. Recognition suites test attributes beyond lexical content\([1](https://arxiv.org/html/2608.19515#bib.bib7);[27](https://arxiv.org/html/2608.19515#bib.bib14);[14](https://arxiv.org/html/2608.19515#bib.bib15);[22](https://arxiv.org/html/2608.19515#bib.bib16);[7](https://arxiv.org/html/2608.19515#bib.bib25)\); response\-oriented benchmarks assess whether a reply fits the speaking style or audio context\([16](https://arxiv.org/html/2608.19515#bib.bib19);[29](https://arxiv.org/html/2608.19515#bib.bib20);[11](https://arxiv.org/html/2608.19515#bib.bib21)\); and multi\-turn benchmarks evaluate emotion understanding and trajectories\([8](https://arxiv.org/html/2608.19515#bib.bib22);[24](https://arxiv.org/html/2608.19515#bib.bib23)\)\. StyleTalk and ParaS2S control lexical content but remain limited to single\-response evaluation, while emotion\-conditioned user simulators add interaction\([17](https://arxiv.org/html/2608.19515#bib.bib12);[9](https://arxiv.org/html/2608.19515#bib.bib13)\)without candidate\-grounded hidden needs that make elicitation, stopping, and final selection jointly verifiable\. Table[1](https://arxiv.org/html/2608.19515#S2.T1)summarizes these distinctions\.

Recent work also examines whether paralinguistic evidence is carried from perception into downstream behavior\. LISTEN finds lexical dominance and underuse of acoustic cues\([5](https://arxiv.org/html/2608.19515#bib.bib26)\); PALLM couples paralinguistic classification with response generation\([15](https://arxiv.org/html/2608.19515#bib.bib27)\); ParaBridge identifies a perception–behavior gap and improves cue\-conditioned behavior through scaffolding and training\([25](https://arxiv.org/html/2608.19515#bib.bib28)\); and Miyazawa and Sato show gains from recognized paralinguistic attitude classes in dialogue\-act prediction\([18](https://arxiv.org/html/2608.19515#bib.bib29)\)\.Hear2Actinstead tests whether a controlled prosodic cue changes multi\-turn elicitation, stopping, and objectively verifiable final task selection\.

## 3Hear2Act Evaluation Design

Task Definition

Hear2Actevaluates whether an assistant can use task\-relevant prosodic evidence to make better dialogue decisions\. Each scenario fixes an initial request, a candidate set, and prioritized hidden concerns that determine which option best satisfies the user’s needs\. Given the dialogue and the evidence available under its access condition, the assistant must decide whether to elicit more information, what to ask, and which option to recommend\. Across matched rollouts, the task and user needs remain fixed while the available evidence varies\. We evaluate both the final choice and the dialogue trajectory leading to it\. Figure[1](https://arxiv.org/html/2608.19515#S1.F1)summarizes the design\.

### 3\.1Benchmark Scenarios

Each scenario is constructed so that the value of otherwise hidden concern information can be measured from the assistant’s final choice\. It combines an initial request, three prioritized hidden concerns, and a candidate set whose fit to those concerns is known \(Figure[1](https://arxiv.org/html/2608.19515#S1.F1), left\)\. We seed 48 domains from the service categories in the Schema\-Guided Dialogue dataset\([19](https://arxiv.org/html/2608.19515#bib.bib18)\)and instantiate ten scenarios per domain, yielding 480 scenarios across travel, housing, healthcare, finance, and other consumer services\. Scenario seeds vary urgency, budget, expertise, and life context so that similar requests can reflect different underlying needs\.

The initial request states only visible requirements, while three concerns remain hidden: a hard constraint \(L1\), a strong preference \(L2\), and a moderate preference \(L3\)\. In Figure[2](https://arxiv.org/html/2608.19515#S2.F2), for example, the user asks only for a cheap flight, while concerns about overnight travel, journey length, and daytime travel remain unstated\. Through interaction, the assistant may uncover none, some, or all of these concerns\.

The candidate set makes the consequences of this information observable\. The three concerns define eight possible satisfaction patterns, each represented by one candidate\. We add one safe alternative and two candidates that violate stated requirements, yielding 11 candidates in total\. The best\-fitting option satisfies all three hidden concerns but does not lead on visible attributes such as price or rating, while the other candidates remain plausible under incomplete information\. Final selection therefore reflects whether the assistant uncovered and used the information needed to distinguish the best\-fitting option\. Appendix[H](https://arxiv.org/html/2608.19515#A8)shows complete matched rollouts, and Appendix[G\.1](https://arxiv.org/html/2608.19515#A7.SS1)summarizes the candidate\-construction procedure\.

### 3\.2Interactive Rollouts

Each scenario is realized as an interactive episode so that concern evidence can affect what the assistant asks, when it stops, and which option it selects\. A rollout pairs one episode with one assistant and one access condition \(Figure[1](https://arxiv.org/html/2608.19515#S1.F1), center\)\. All rollouts begin from the same surface\-attractive option\. Text assistants mayask,clarify,recommend, orconfirm, while spoken assistants useask,recommend, andconfirm\. The episode ends when an option is accepted or the 20\-turn budget is reached\.

A deterministic user engine governs the interaction\. Given the scenario, dialogue history, and assistant action, it determines which concerns remain unresolved, what information can be disclosed, how the user responds, and whether the episode ends\. A targeted question reveals the user’s requirement for the queried attribute, while a general question reveals only why the current recommendation remains unresolved\. The resulting structured response is realized as a natural utterance and, for spoken conditions, rendered as speech\. Detailed interaction and disclosure rules appear in Appendix[G\.6](https://arxiv.org/html/2608.19515#A7.SS6)\.

We vary both how concern information is expressed and what evidence is available to the assistant\. UnderExplicit lexicalfeedback, concerns are stated in words\. UnderProsody\-mediatedfeedback, hard\-constraint violations remain explicit, while soft concerns remain lexically implicit and vocal delivery signals that the recommendation is still unresolved\. Figure[2](https://arxiv.org/html/2608.19515#S2.F2)illustrates the resulting divergence: transcript\-only access may treat hesitant acceptance as resolution, whereas concern\-state access supports further elicitation and a better\-fitting choice\. Matched rollouts keep the scenario and user needs fixed while varying access to transcript, audio, an audio\-inferred textual state, or the ground\-truth concern state, enabling paired comparison of dialogue trajectories and outcomes \(Figure[1](https://arxiv.org/html/2608.19515#S1.F1), right\)\. Exact input combinations appear in Section[4\.2](https://arxiv.org/html/2608.19515#S4.SS2)\.

### 3\.3Paired Evaluation

Each comparison pairs rollouts with the same scenario, user needs, feedback realization, and task instructions, differing only in the assistant’s access condition \(Figure[1](https://arxiv.org/html/2608.19515#S1.F1), right\)\. This design isolates how the available evidence changes both the final outcome and the dialogue trajectory leading to it\.

We evaluate final task outcomes together with elicitation, disclosure, assistant actions, dialogue length, and stopping behavior, allowing us to distinguish effective information use from simply longer interaction\.

The paired design supports two complementary analyses\. Spoken\-model comparisons test how effectively task\-relevant prosodic information is carried from audio into action, while text\-model comparisons measure its decision value when correctly represented\.

## 4Experimental Instantiation and Validation

### 4\.1Evaluation Scope and Systems

The benchmark itself is fixed across evaluations, while the number of model rollouts varies by system and condition\. As summarized in Table[2](https://arxiv.org/html/2608.19515#S4.T2), it contains 480 scenarios and 960 base episode specifications\.

Table 2:Hear2Act benchmark and rollout coverage\.The 480 model\-independent scenarios expand to 54,240 evaluation rollouts across models, access conditions, renderers, and interventions\.Benchmark artifactSGD\-seeded domains48Scenarios per domain10Benchmark scenarios480Candidate options per scenario11Hidden concern layers per scenario3Feedback realizations2Base episode specifications960Assistant turn budget20Evaluation rolloutsText LLM main grid19,200Text label interventions1,440Spoken assistant with Qwen3\-TTS6,720Spoken assistant with VoxCPM26,720Qwen2\-Audio, three rollouts per scenario20,160Total evaluation rollouts54,240#### Evaluated assistants\.

We evaluate two spoken assistants to test whether task\-relevant prosodic information can be carried from audio into decisions, and five text LLMs to measure the value of that information when explicitly represented\. The text LLMs are Claude Opus 4\.6, Kimi K2\.5, DeepSeek\-V3\.2, GLM\-5, and Qwen3\-32B, all accessed in July 2026 under a fixed decoding configuration\. Qwen2\.5\-Omni\-7B is the primary spoken assistant and is evaluated on both speech renderers\([26](https://arxiv.org/html/2608.19515#bib.bib9)\)\. Qwen2\-Audio\-7B\-Instruct provides a second spoken backbone and is evaluated on the primary renderer with three rollouts per scenario\([6](https://arxiv.org/html/2608.19515#bib.bib17)\)\.

#### Realization components\.

User utterances are generated from structured engine states by a fixed language realizer, while the user engine determines state transitions and concern disclosure\. Qwen3\-TTS is the primary speech renderer and follows natural\-language delivery instructions derived from the concern state\([13](https://arxiv.org/html/2608.19515#bib.bib5)\)\. VoxCPM2 provides a second\-renderer robustness condition using its native affect controls\([32](https://arxiv.org/html/2608.19515#bib.bib4)\)\. Additional prompting and realization details appear in Appendix[G](https://arxiv.org/html/2608.19515#A7)\.

### 4\.2Access Conditions and Interventions

We organize the evaluation into text\-model conditions, spoken\-model conditions, and label\-fidelity interventions\. Downstream decision instructions remain fixed across access conditions, with state definitions and turn\-level labels added only when required\.

#### Text\-model conditions\.

Text LLMs receive either the dialogue transcript alone or the transcript with the ground\-truth concern state at each user turn\. This comparison isolates the decision value of correctly represented concern information without requiring speech perception\.

#### Spoken\-model conditions\.

Spoken assistants are evaluated under five core access conditions\. Transcript only, audio only, and audio plus transcript provide no explicit state representation\. Transcript plus state provides the ground\-truth graded concern\-state tag\. In the audio\-inferred condition, the assistant instead infers the task\-relevant resolved/unresolved concern status from audio and records it as text\. The decision stage then receives the transcript together with this inferred state, rather than the raw audio, when selecting the next action\.

#### Speech\-emotion\-recognition \(SER\) baselines\.

We additionally test two off\-the\-shelf SER representations for each spoken assistant, yielding seven conditions in total\. HuBERT\-SUPERB\-ER\([12](https://arxiv.org/html/2608.19515#bib.bib30);[28](https://arxiv.org/html/2608.19515#bib.bib8)\)and SpeechBrain wav2vec2\-IEMOCAP\([20](https://arxiv.org/html/2608.19515#bib.bib31);[2](https://arxiv.org/html/2608.19515#bib.bib32);[4](https://arxiv.org/html/2608.19515#bib.bib1)\)independently classify the current user audio into their native four\-class affect space \(happy, sad, angry, or neutral\)\. The predicted label is passed verbatim with the transcript to the same downstream decision stage, without mapping it to the benchmark concern states\. These baselines test whether generic affect representations recover the decision\-relevant information captured by the audio\-inferred concern state\.

These conditions support four main comparisons\. Transcript versus audio plus transcript tests whether direct audio improves decisions beyond the transcript\. Audio plus transcript versus transcript plus the audio\-inferred state compares direct audio use with a two\-stage intervention that first infers the task\-relevant concern status from audio and then supplies that textual inference to the same downstream decision model\. Transcript versus transcript plus the ground\-truth state measures the value of correctly represented concern information\. Finally, the SER conditions versus the audio\-inferred state compare generic affect with task\-relevant concern extraction under the same downstream decision model\. Ground\-truth state access serves as a diagnostic reference rather than a deployment setting\.

#### Label\-fidelity interventions\.

On a 48\-scenario subset, we replace the ground\-truth concern states with all\-positive \(Resolved\), all\-negative \(Unresolved\), or turn\-shuffled sequences\. The fixed conditions remove turn\-level variation, while shuffling preserves the label distribution but breaks its alignment with the current turn\. These controls test whether the gains require correct turn\-specific concern information, rather than merely providing a state signal or inducing more cautious interaction\.

### 4\.3Metrics and Statistical Analysis

#### Outcome and user\-side interaction measures\.

We report three higher\-is\-better outcome measures\.Optimal\-solution rate \(1st%\)is the proportion of rollouts ending with the first\-tier option, using the final recommendation if the turn budget is exhausted\.OptSat%andSvcSat%measure the proportions of recommendation turns followed by non\-negative Option\-acceptance and appreciative Interaction\-satisfaction states, respectively\. Because each feedback turn follows one recommendation, these are per\-recommendation rates whose denominator depends on how many recommendations the assistant makes\. A policy that searches longer before settling can therefore lower these rates while improving 1st%\. All user\-side states are computed deterministically by the engine and are never shown to the assistant\.

#### Policy and elicitation diagnostics\.

To understand how concern information changes the dialogue, we track the shares ofrecommend,ask, andclarifydecisions\. We also measure hidden concerns disclosed, full disclosure of all three concerns, dialogue length, and premature closure on a suboptimal option\. These diagnostics distinguish effective elicitation and stopping from simply interacting longer\.

Table 3:Spoken\-assistant results under Prosody\-mediated feedbackwith Qwen3\-TTS\.TT,AA,SS, andS^\\hat\{S\}denote transcript, audio, ground\-truth state, and audio\-inferred state; audio\-derived representations are textualized and paired with the transcript\.Averageis computed across the two audio\-capable LLMs\. Bold/underline indicate the best/second\-best value per column\. See Table[4](https://arxiv.org/html/2608.19515#S4.T4)for Explicit\-lexical results and Appendix[B](https://arxiv.org/html/2608.19515#A2)for VoxCPM2\.Qwen2\.5\-OmniQwen2\-AudioAverageInput / representation1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow*Direct input*Transcript only \(TT\)15\.422\.912\.313\.720\.112\.014\.621\.512\.2Audio only \(AA\)17\.322\.713\.314\.722\.412\.216\.022\.612\.8Audio \+ transcript \(A\+TA\+T\)15\.924\.014\.614\.723\.313\.315\.323\.714\.0*Textualized prosodic representations*Generic affect \(HuBERT\)35\.740\.321\.429\.431\.319\.932\.635\.820\.7Generic affect \(SpeechBrain\)32\.638\.720\.625\.729\.118\.329\.233\.919\.5Task\-aligned state \(T\+S^T\+\\hat\{S\}\)43\.041\.822\.736\.233\.420\.639\.637\.621\.7*Ground\-truth state \(T\+ST\+S\)*44\.541\.824\.436\.937\.523\.540\.739\.724\.0Table 4:Spoken\-assistant use of concern information under Explicit lexical feedbackwith Qwen3\-TTS \(Qwen2\.5\-Omnin=480n\{=\}480, Qwen2\-Audion=1,440n\{=\}1\{,\}440per condition\)\. Notation follows Table[3](https://arxiv.org/html/2608.19515#S4.T3)\. Results are similar because the concern is explicit in the transcript\. Bold/underline mark column\-wise highest/next\-highest values\.Qwen2\.5\-OmniQwen2\-AudioAverageInput / representation1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow*Direct input*Transcript only \(TT\)56\.245\.326\.948\.342\.126\.352\.343\.726\.6Audio only \(AA\)55\.345\.726\.448\.041\.326\.051\.743\.526\.2Audio \+ transcript \(A\+TA\+T\)52\.245\.426\.547\.641\.025\.949\.943\.226\.2*Textualized prosodic representations*Generic affect \(HuBERT\)52\.846\.025\.348\.540\.624\.550\.743\.324\.9Generic affect \(SpeechBrain\)52\.845\.925\.345\.039\.923\.748\.942\.924\.5Task\-aligned state \(T\+S^T\+\\hat\{S\}\)52\.445\.825\.449\.841\.225\.151\.143\.525\.3*Ground\-truth state \(T\+ST\+S\)*49\.746\.626\.650\.342\.326\.150\.044\.526\.4
#### Statistical analysis\.

The scenario is the statistical unit\. We pair condition differences within scenarios and compute scenario\-level bootstrap 95% confidence intervals with 2,000 resamples\. Repeated rollouts, access conditions, and renderers from the same scenario are not treated as independent observations\. For pooled results, we compute each model’s statistic separately and average equally across models\. The 48\-scenario intervention analysis follows the same paired design\. Confidence intervals for key diagnostic contrasts appear in Appendix[D](https://arxiv.org/html/2608.19515#A4)\.

### 4\.4Speech\-Layer Validation

We verify that the intended concern contrast remains recoverable after speech rendering\. Two annotators independently judge 100 utterances from each renderer in dialogue context, with samples balanced betweenResolvedandUnresolved\. For Qwen3\-TTS, annotators select among four graded concern states, which we collapse to the binary distinction for analysis\. For VoxCPM2, they make the binary judgment directly\. Full instructions appear in Appendix[E](https://arxiv.org/html/2608.19515#A5)\.

Qwen2\.5\-Omni\-7B performs the same perception task on the same clips, allowing human and model recovery of the intended concern status to be compared directly\. This validation establishes whether the intended communicative contrast survives rendering, rather than treating synthesized speech as objective ground truth for emotion\.

## 5Results

The results address three questions\. First, when does prosodic information add decision value beyond the transcript? Second, when it does, can spoken assistants use it? Third, how does correct concern information change the dialogue? All analyses use paired scenario\-level comparisons\. Paired 95% confidence intervals for key diagnostic contrasts appear in Table[9](https://arxiv.org/html/2608.19515#A4.T9)\.

### 5\.1When Does Prosodic Information Add Decision Value?

We first ask when prosodic information changes decisions beyond what is already available in the transcript\. The results show that prosody adds substantial decision value when it conveys task\-relevant information missing from the words, but current audio\-capable LLMs make limited use of this value directly from audio\. Under Prosody\-mediated feedback, adding audio to the transcript changes the average optimal\-solution rate only from14\.6%14\.6\\%to15\.3%15\.3\\%, whereas representing the audio\-inferred concern state raises it to39\.6%39\.6\\%\(Table[3](https://arxiv.org/html/2608.19515#S4.T3)\)\. This advantage largely disappears under Explicit lexical feedback, where the concern is already stated in words: the optimal\-solution rates across all seven conditions span only6\.56\.5points for Qwen2\.5\-Omni and5\.35\.3points for Qwen2\-Audio \(Table[4](https://arxiv.org/html/2608.19515#S4.T4)\)\.

Table 5:Effect of ground\-truth concern\-state access on text LLMs\.Results use 480 matched scenarios with two runs each \(n=960n\{=\}960per cell\)\.TTis transcript only;T\+ST\{\+\}Sadds turn\-level ground\-truth state tags\. Bold marks the larger value in each pair; action composition appears in Figure[3](https://arxiv.org/html/2608.19515#S5.F3)\.Prosody\-mediated feedbackExplicit lexical feedback1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrowModelTTT\+ST\{\+\}STTT\+ST\{\+\}STTT\+ST\{\+\}STTT\+ST\{\+\}STTT\+ST\{\+\}STTT\+ST\{\+\}SClaude Opus 4\.649\.779\.137\.041\.124\.726\.274\.981\.944\.643\.025\.926\.4Kimi K2\.527\.969\.227\.736\.114\.419\.669\.474\.937\.136\.217\.717\.7GLM\-522\.460\.126\.133\.413\.015\.867\.373\.536\.835\.715\.516\.0Qwen3\-32B21\.153\.131\.035\.510\.413\.764\.364\.843\.137\.712\.912\.6DeepSeek\-V3\.213\.356\.216\.428\.86\.514\.454\.867\.631\.832\.614\.315\.1Five\-model mean26\.963\.527\.635\.013\.817\.966\.172\.538\.737\.017\.317\.6

Figure 3:Change in assistant action composition with concern\-state access\.Points show theT\+S−TT\{\+\}S\{\-\}Tchange, in percentage points, in the share ofrecommend,ask, andclarifydecision turns under Prosody\-mediated \(red circles\) and Explicit lexical \(blue triangles\) feedback; means are macro\-averaged across models\.This asymmetry is not specific to spoken models\. When we isolate the value of the concern information itself by giving five text LLMs the ground\-truth state, the gains are much larger under Prosody\-mediated feedback\. Ground\-truth state access raises the mean optimal\-solution rate across the five models by36\.736\.7percentage points \(Table[5](https://arxiv.org/html/2608.19515#S5.T5)\), compared with only6\.46\.4points under Explicit lexical feedback\. OptSat% and SvcSat% show the same asymmetry, increasing by7\.37\.3and4\.14\.1points under Prosody\-mediated feedback but by only−1\.6\-1\.6and\+0\.3\+0\.3points under Explicit feedback\.

Together, these results show that prosodic information has substantial decision value when lexical evidence is insufficient, but little when the concern is explicit in the transcript, supporting selective use of prosodic cues\.

### 5\.2When Prosody Matters, Can Spoken Assistants Use It?

Prosody\-mediated feedback is therefore the setting in which prosodic information adds substantial value beyond the transcript\. We next ask where this value is lost between perceiving the cue and acting on it\. Qwen2\.5\-Omni shows a clear perception–action gap: on the audited perception task, it recovers the intended concern status with 0\.85 accuracy on Qwen3\-TTS and 0\.76 on VoxCPM2 \(Table[11](https://arxiv.org/html/2608.19515#A5.T11)\), showing that the cue is substantially recoverable from audio\. Yet this information has little effect on direct decisions: optimal\-solution rates remain similar with Transcript only \(15\.4%15\.4\\%\), Audio only \(17\.3%17\.3\\%\), and Audio\+\+transcript \(15\.9%15\.9\\%; Table[3](https://arxiv.org/html/2608.19515#S4.T3)\)\. Thus, recovering the cue does not by itself translate into better task decisions\.

Making the recovered concern information explicit largely closes this gap\. When Qwen2\.5\-Omni first infers the concern status from audio, expresses it in text, and uses it with the transcript for next\-action selection, the optimal\-solution rate rises from15\.9%15\.9\\%to43\.0%43\.0\\%, close to44\.5%44\.5\\%with the ground\-truth state\. This two\-stage intervention recovers most of the available state\-access benefit, showing that the model can use the information once it is made explicit but makes limited use of it directly from audio\.

Figure 4:Outcome tracks label fidelity, not asking frequency\.Optimal\-solution rate on the 48\-scenario intervention subset under Prosody\-mediated feedback, pooled over five text LLMs\. All label conditions have similaraskshares \(4444–45%45\\%\)\.We next ask whether generic speech\-emotion representations recover decision value from audio\. With Qwen2\.5\-Omni fixed as the downstream decision model, textualized HuBERT and SpeechBrain predictions raise the optimal\-solution rate from15\.9%15\.9\\%with direct audio to35\.7%35\.7\\%and32\.6%32\.6\\%, respectively\. The audio\-inferred concern state raises it further to43\.0%43\.0\\%\(Table[3](https://arxiv.org/html/2608.19515#S4.T3)\)\. The same pattern holds for Qwen2\-Audio and with VoxCPM2 \(Appendix[B](https://arxiv.org/html/2608.19515#A2)\)\. Thus, generic affect recovers some decision value, while task\-relevant concern inference yields larger downstream gains in our setting\.

### 5\.3How Does Useful Concern Information Change the Dialogue?

Correct concern information shifts the dialogue from early recommendation toward information gathering\. Under Prosody\-mediated feedback, the average number ofaskandclarifyactions across the five text LLMs rises from1\.81\.8to4\.44\.4per dialogue, indicating more information gathering before recommendation\. This shift is accompanied by greater disclosure of the user’s hidden needs: the share of dialogues in which all three concerns are disclosed rises from3%3\\%to30%30\\%\. Under Explicit lexical feedback, where the concern is already stated, disclosure changes much less\. The shift toward asking appears across models, while clarification is more model\-dependent \(Figure[3](https://arxiv.org/html/2608.19515#S5.F3)\)\. Additional behavior results appear in Appendix[C](https://arxiv.org/html/2608.19515#A3)\.

The improvement is not simply due to asking more questions\. When models receive incorrect concern states, they ask at nearly the same rate as with the correct state \(4444–45%45\\%\), but achieve much lower optimal\-solution rates \(35\.435\.4–53\.3%53\.3\\%versus68\.1%68\.1\\%; Figure[4](https://arxiv.org/html/2608.19515#S5.F4)\)\. Even the always\-negative condition produces longer dialogues \(23\.223\.2versus17\.817\.8turns\) while performing worse\. Thus, the gain depends on having the correct concern information at each turn, not on interacting more\.

Correct concern information also improves stopping behavior\. Under Prosody\-mediated feedback, state access reduces premature closure on a suboptimal option from69\.7%69\.7\\%to26\.7%26\.7\\%\. This also explains why OptSat%, a per\-recommendation acceptance rate, can decrease even as the final choice improves: continued search adds rejected recommendations to its denominator\. Under Explicit feedback, for example, 1st% rises from66\.1%66\.1\\%to72\.5%72\.5\\%while OptSat% decreases slightly from38\.7%38\.7\\%to37\.0%37\.0\\%\.

Overall, correct concern information helps the assistant know when to continue eliciting and searching, and when to stop\.

## 6Conclusion

We introducedHear2Act, a controlled protocol for measuring when prosody changes multi\-turn task decisions\. Across two audio\-capable LLMs, prosody adds decision value when lexical evidence is insufficient, while the contrast largely disappears under Explicit lexical feedback, where the concern is verbalized\. Yet raw audio provides little direct benefit\. When models infer the concern status from audio, represent it in text, and use it for next\-action selection, they recover most of the gain from ground\-truth state access\. Correct concern information also changes the dialogue policy, encouraging elicitation when needed and reducing premature stopping on suboptimal options\. Together, these results show that audio\-capable LLMs can recover useful information from speech but do not reliably carry it into action without an explicit intermediate representation\.

## Limitations

Hear2Actisolates one decision\-relevant function of prosody: whether and how strongly a recommendation remains unresolved\. Its shared three\-layer concern structure and standardized opening enable matched, objectively scored comparisons, but do not cover broader preference structures, other functions of prosody, or unconstrained first\-turn recommendation\.

User utterances and speech are synthesized from structured states\. Human validation shows that the intended concern contrast remains perceptible across two renderers, but synthetic speech cannot capture the full variability of natural speech and interaction\. Ground\-truth\-state conditions are diagnostic rather than deployment settings\. The inferred\-state condition adds both an explicit representation and an additional inference step, so it establishes that the intervention is sufficient without isolating the mechanism\.

## Ethical Considerations

Hear2Actuses constructed scenarios and personas, with dialogue and speech generated from structured benchmark states\. It contains no real user conversations, identities, recordings, or personal data\. Personas encode only task\-relevant context and are not intended to represent demographic groups or support claims about real individuals\.

The benchmark evaluates a narrow communicative signal: whether a recommendation remains unresolved\. Such evidence should support clarification rather than be used to infer sensitive attributes, override explicit user statements, or make high\-stakes decisions without confirmation\. The human study involved volunteer graduate students who judged generated speech clips\. No personal or sensitive information was collected\. Details of the human study appear in Appendix[E](https://arxiv.org/html/2608.19515#A5); benchmark and computational details appear in Appendix[F](https://arxiv.org/html/2608.19515#A6)\.

## References

- Aoet al\.\(2024\)J\. Ao, Y\. Wang, X\. Tian, D\. Chen, J\. Zhang, L\. Lu, Y\. Wang, H\. Li, and Z\. WuSD\-Eval: a benchmark dataset for spoken dialogue understanding beyond words\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Baevskiet al\.\(2020\)A\. Baevski, Y\. Zhou, A\. Mohamed, and M\. AuliWav2vec 2\.0: a framework for self\-supervised learning of speech representations\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.33,pp\. 12449–12460\.External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html)Cited by:[§4\.2](https://arxiv.org/html/2608.19515#S4.SS2.SSS0.Px3.p1.1)\.
- Budzianowskiet al\.\(2018\)P\. Budzianowski, T\. Wen, B\. Tseng, I\. Casanueva, S\. Ultes, O\. Ramadan, and M\. GašićMultiWOZ – a large\-scale multi\-domain wizard\-of\-oz dataset for task\-oriented dialogue modelling\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 5016–5026\.Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p1.1),[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2608.19515#S2.T1.4.2.1.1.1)\.
- Bussoet al\.\(2008\)C\. Busso, M\. Bulut, C\. Lee, A\. Kazemzadeh, E\. Mower, S\. Kim, J\. N\. Chang, S\. Lee, and S\. S\. NarayananIEMOCAP: interactive emotional dyadic motion capture database\.Language Resources and Evaluation42\(4\),pp\. 335–359\.Cited by:[§4\.2](https://arxiv.org/html/2608.19515#S4.SS2.SSS0.Px3.p1.1)\.
- Chenet al\.\(2026\)J\. Chen, Z\. Guo, J\. Chun, P\. Wang, A\. Perrault, and M\. ElsnerDo audio LLMs really LISTEN, or just transcribe? measuring lexical vs\. acoustic emotion cues reliance\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),Rabat, Morocco,pp\. 5848–5877\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.274),[Link](https://aclanthology.org/2026.eacl-long.274/)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p2.1)\.
- Chuet al\.\(2024\)Y\. Chu, J\. Xu, Q\. Yang, H\. Wei, X\. Wei, Z\. Guo, Y\. Leng, Y\. Lv, J\. He, J\. Lin, C\. Zhou, and J\. ZhouQwen2\-Audio technical report\.External Links:2407\.10759,[Link](https://arxiv.org/abs/2407.10759)Cited by:[§4\.1](https://arxiv.org/html/2608.19515#S4.SS1.SSS0.Px1.p1.1)\.
- Debaupteet al\.\(2026\)L\. Debaupte, T\. Baumgartner, B\. Tai, C\. Fan, B\. Wang, and Y\. ZhongVocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models\.Note:Hugging Face dataset,[https://huggingface\.co/datasets/besimple\-ai/vocal\-affect\-bench](https://huggingface.co/datasets/besimple-ai/vocal-affect-bench)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Denget al\.\(2025\)Y\. Deng, G\. Hu, H\. Sun, X\. Zhang, H\. Zhang, F\. Tian, X\. Yang, G\. Yu, and E\. S\. ChngMULTI\-Bench: a multi\-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models\.Note:arXiv:2511\.00850Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.19515#S2.T1.4.5.1.1.1)\.
- Fenget al\.\(2024\)S\. Feng, H\. Lin, C\. Geishauser, N\. Lubis, C\. van Niekerk, M\. Heck, B\. Ruppik, R\. Vukovic, and M\. GašićInfusing emotions into task\-oriented dialogue systems: understanding, management, and generation\.InProceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue,Kyoto, Japan,pp\. 699–717\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.sigdial-1.60),[Link](https://aclanthology.org/2024.sigdial-1.60/)Cited by:[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Fenget al\.\(2022\)S\. Feng, N\. Lubis, C\. Geishauser, H\. Lin, M\. Heck, C\. van Niekerk, and M\. GašićEmoWOZ: a large\-scale corpus and labelling scheme for emotion recognition in task\-oriented dialogue systems\.InProceedings of the Thirteenth Language Resources and Evaluation Conference \(LREC\),pp\. 4096–4113\.Cited by:[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1)\.
- Gaoet al\.\(2025\)K\. Gao, S\. Xia, K\. Xu, P\. Torr, and J\. GuBenchmarking open\-ended audio dialogue understanding for large audio\-language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 4763–4784\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.237),[Link](https://aclanthology.org/2025.acl-long.237/)Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Hsuet al\.\(2021\)W\. Hsu, B\. Bolte, Y\. H\. Tsai, K\. Lakhotia, R\. Salakhutdinov, and A\. MohamedHuBERT: self\-supervised speech representation learning by masked prediction of hidden units\.IEEE/ACM Transactions on Audio, Speech, and Language Processing29,pp\. 3451–3460\.External Links:[Document](https://dx.doi.org/10.1109/TASLP.2021.3122291)Cited by:[§4\.2](https://arxiv.org/html/2608.19515#S4.SS2.SSS0.Px3.p1.1)\.
- Huet al\.\(2026\)H\. Hu, X\. Zhu, T\. He, D\. Guo, B\. Zhang, X\. Wang, Z\. Guo, Z\. Jiang, H\. Hao, Z\. Guo, X\. Zhang, P\. Zhang, B\. Yang, J\. Xu, J\. Zhou, and J\. LinQwen3\-TTS Technical Report\.External Links:2601\.15621Cited by:[§4\.1](https://arxiv.org/html/2608.19515#S4.SS1.SSS0.Px2.p1.1)\.
- Huanget al\.\(2024\)C\. Huang, K\. Lu, S\. Wang, C\. Hsiao, C\. Kuan, H\. Wu, S\. Arora, K\. Chang, J\. Shi, Y\. Peng, R\. Sharma, S\. Watanabe, B\. Ramakrishnan, S\. Shehata, and H\. LeeDynamic\-SUPERB: towards a dynamic, collaborative, and comprehensive instruction\-tuning benchmark for speech\.In2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 12136–12140\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP48485.2024.10448257)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Kimet al\.\(2026\)M\. Kim, J\. Chen, S\. Leem, Y\. Huang, R\. Rungta, Z\. Ouyang, H\. Wu, S\. T\. Appini, A\. Bansal, Y\. Bai, Y\. Liu, F\. Metze, A\. A\. Aly, A\. Kumar, A\. Rastrow, and Z\. LinAligning paralinguistic understanding and generation in speech LLMs via multi\-task reinforcement learning\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 5: Industry Track\),Rabat, Morocco,pp\. 636–648\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-industry.49),[Link](https://aclanthology.org/2026.eacl-industry.49/)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p2.1)\.
- Linet al\.\(2024\)G\. Lin, C\. Chiang, and H\. LeeAdvancing large language models to capture varied speaking styles and respond properly in spoken conversations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 6626–6642\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.358),[Link](https://aclanthology.org/2024.acl-long.358/)Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.19515#S2.T1.4.4.1.1.1)\.
- Linet al\.\(2023\)H\. Lin, S\. Feng, C\. Geishauser, N\. Lubis, C\. van Niekerk, M\. Heck, B\. Ruppik, R\. Vukovic, and M\. GašićEmoUS: simulating user emotions in task\-oriented dialogues\.InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval \(SIGIR\),pp\. 2526–2531\.Cited by:[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Miyazawa and Sato \(2026\)K\. Miyazawa and Y\. SatoEvaluation of paralinguistic\-aware spoken dialogue systems using next\-utterance classification\.InProceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue,Atlanta, Georgia, USA,pp\. 711–719\.External Links:[Link](https://aclanthology.org/2026.sigdial-1.50/)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p2.1)\.
- Rastogiet al\.\(2020\)A\. Rastogi, X\. Zang, S\. Sunkara, R\. Gupta, and P\. KhaitanTowards scalable multi\-domain conversational agents: the schema\-guided dialogue dataset\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 8689–8696\.Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p1.1),[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2608.19515#S2.T1.4.2.1.1.1),[§3\.1](https://arxiv.org/html/2608.19515#S3.SS1.p1.1)\.
- Ravanelliet al\.\(2021\)M\. Ravanelli, T\. Parcollet, P\. Plantinga, A\. Rouhe, S\. Cornell, L\. Lugosch, C\. Subakan, N\. Dawalatabad, A\. Heba, J\. Zhong, J\. Chou, S\. Yeh, S\. Fu, C\. Liao, E\. Rastorgueva, F\. Grondin, W\. Aris, H\. Na, Y\. Gao, R\. De Mori, and Y\. BengioSpeechBrain: a general\-purpose speech toolkit\.arXiv preprint arXiv:2106\.04624\.External Links:[Link](https://arxiv.org/abs/2106.04624)Cited by:[§4\.2](https://arxiv.org/html/2608.19515#S4.SS2.SSS0.Px3.p1.1)\.
- Siet al\.\(2023\)S\. Si, W\. Ma, H\. Gao, Y\. Wu, T\. Lin, Y\. Dai, H\. Li, R\. Yan, F\. Huang, and Y\. LiSpokenWOZ: a large\-scale speech\-text benchmark for spoken task\-oriented dialogue agents\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p1.1),[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2608.19515#S2.T1.4.3.1.1.1)\.
- Wanget al\.\(2025\)B\. Wang, X\. Zou, G\. Lin, S\. Sun, Z\. Liu, W\. Zhang, Z\. Liu, A\. Aw, and N\. F\. ChenAudioBench: a universal benchmark for audio large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Albuquerque, New Mexico,pp\. 4297–4316\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.218),[Link](https://aclanthology.org/2025.naacl-long.218/)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Wanget al\.\(2026a\)E\. Wang, J\. Zhou, Y\. Jia, A\. Kong, Q\. Li, and Y\. QinRealTalk\-CN: a realistic chinese speech task\-oriented dialogue benchmark with cross\-modal analysis\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),San Diego, California, United States,pp\. 2880–2897\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.131),[Link](https://aclanthology.org/2026.acl-long.131/)Cited by:[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1)\.
- Wanget al\.\(2026b\)S\. Wang, Z\. Zhao, H\. Xue, C\. Wang, S\. Wang, H\. Bu, X\. Xu, and L\. XieHumDial\-EIBench: a human\-recorded multi\-turn emotional intelligence benchmark for audio language models\.Note:arXiv:2604\.11594Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.19515#S2.T1.4.5.1.1.1)\.
- Wanget al\.\(2026c\)Y\. Wang, Q\. Ni, S\. Cai, W\. Lin, L\. Zhang, and Z\. WuParaBridge: bridging paralinguistic perception and dialogue behavior in speech language models\.arXiv preprint arXiv:2606\.10581\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2606.10581),[Link](https://arxiv.org/abs/2606.10581)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p2.1)\.
- Xuet al\.\(2025\)J\. Xu, Z\. Guo, J\. He, H\. Hu, T\. He, S\. Bai, K\. Chen, J\. Wang, Y\. Fan, K\. Dang, B\. Zhang, X\. Wang, Y\. Chu, and J\. LinQwen2\.5\-Omni technical report\.External Links:2503\.20215,[Link](https://arxiv.org/abs/2503.20215)Cited by:[§4\.1](https://arxiv.org/html/2608.19515#S4.SS1.SSS0.Px1.p1.1)\.
- Yanget al\.\(2024\)Q\. Yang, J\. Xu, W\. Liu, Y\. Chu, Z\. Jiang, X\. Zhou, Y\. Leng, Y\. Lv, Z\. Zhao, C\. Zhou, and J\. ZhouAIR\-Bench: benchmarking large audio\-language models via generative comprehension\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 1979–1998\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.109),[Link](https://aclanthology.org/2024.acl-long.109/)Cited by:[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1)\.
- Yanget al\.\(2021\)S\. Yang, P\. Chi, Y\. Chuang, C\. J\. Lai, K\. Lakhotia, Y\. Y\. Lin, A\. T\. Liu, J\. Shi, X\. Chang, G\. Lin, T\. Huang, W\. Tseng, K\. Lee, D\. Liu, Z\. Huang, S\. Dong, S\. Li, S\. Watanabe, A\. Mohamed, and H\. LeeSUPERB: speech processing universal performance benchmark\.InInterspeech 2021,pp\. 1194–1198\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2021-1775),[Link](https://www.isca-archive.org/interspeech_2021/yang21c_interspeech.html)Cited by:[§4\.2](https://arxiv.org/html/2608.19515#S4.SS2.SSS0.Px3.p1.1)\.
- Yanget al\.\(2026\)S\. Yang, M\. Tu, T\. Liu, X\. Qu, H\. Lee, L\. Lu, Y\. Wang, and Y\. WuParaS2S: benchmarking and aligning spoken language models for paralinguistic\-aware speech\-to\-speech interaction\.InInternational Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/734b2e77222680728b9ce78c573eae1e-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2608.19515#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.19515#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.19515#S2.T1.4.4.1.1.1)\.
- Younget al\.\(2013\)S\. Young, M\. Gašić, B\. Thomson, and J\. D\. WilliamsPOMDP\-based statistical spoken dialog systems: a review\.Proceedings of the IEEE101\(5\),pp\. 1160–1179\.Cited by:[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1)\.
- Zhanget al\.\(2026\)Y\. Zhang, H\. Nayyeri, R\. Khaziev, E\. Yilmaz, G\. Tur, D\. Hakkani\-Tür, and H\. ThadakamallaATOD: an evaluation framework and benchmark for agentic task\-oriented dialogue systems\.arXiv preprint arXiv:2601\.11854\.Cited by:[§2\.1](https://arxiv.org/html/2608.19515#S2.SS1.p1.1)\.
- Zhouet al\.\(2026\)Y\. Zhou, G\. Zeng, X\. Liu, X\. Li, R\. Yu, J\. Gui, J\. Wu, Z\. Wang, X\. Shen, R\. Ye, Z\. Zhang, J\. Zhou, B\. Bai, W\. Sun, M\. Deng, Q\. Shi, Z\. Wu, and Z\. LiuVoxCPM2 Technical Report\.External Links:2606\.06928Cited by:[§4\.1](https://arxiv.org/html/2608.19515#S4.SS1.SSS0.Px2.p1.1)\.

## Appendix ASGD\-Seeded Domains and Subdomain Scenarios

Table[6](https://arxiv.org/html/2608.19515#A1.T6)shows the complete domain inventory: 48 consumer\-service domains seeded from the service categories of the Schema\-Guided Dialogue dataset and grouped into eight service families\. Rather than listing all 480 subdomain seeds, Table[7](https://arxiv.org/html/2608.19515#A1.T7)expands one domain in full—the budget\-airline domain used in Figure[1](https://arxiv.org/html/2608.19515#S1.F1)—to illustrate how scenario seeds vary in user situation and decision pressure\. The resulting variation in urgency, budget tier, expertise, and life context grounds the scenario\-specific hidden ladders described in Section[3\.1](https://arxiv.org/html/2608.19515#S3.SS1); the construction procedure is summarized in Appendix[G\.1](https://arxiv.org/html/2608.19515#A7.SS1)\.

Table 6:Domain inventory\.Hear2Act covers 48 consumer\-service domains grouped into eight families\. Each domain is expanded into ten scenario seeds, yielding 480 benchmark scenarios\.Service family\#DomainsTravel & Events5Budget airline tickets, vacation rental properties, restaurant reservations, concert ticket purchases, and wedding venue selection\.Home & Property8Home cleaning services, home renovation contractors, home security systems, landscaping contractors, lawn care contractors, solar panel installation, kitchen appliance upgrades, and mattress replacement\.Finance & Insurance9Credit card applications, mortgage lender comparison, investment portfolio allocation, retirement planning advisors, tax preparation services, car insurance policies, health insurance plans, pet insurance policies, and business insurance coverage\.Health & Wellness7Dermatologist appointments, pediatrician selection, mental health therapists, meditation apps, gym membership options, fitness tracker devices, and prescription eyeglasses\.Education & Career5College major selection, online coding bootcamps, language learning platforms, professional development courses, and laptop purchase for students\.Family & Lifestyle6Children’s daycare centers, dog training classes, online dating platforms, wine club memberships, meal delivery subscriptions, and video game purchases\.Media & Devices4Cable TV packages, streaming service subscriptions, podcast hosting services, and smartphone upgrades\.Professional Services4Legal consultation services, auto mechanic services, business accounting software, and freelance graphic designers\.Total48480 scenarios across ten seeds per domainTable 7:One expanded domain\.Ten subdomain seeds for budget airline tickets illustrate variation in user situation and decision pressure before scenario instantiation\.Subdomain situation \(who / pressure\)Opening request1Last\-minute emergency travel for family medical situation with extremely limited budget“I need to fly to see my sick grandmother tomorrow but only have $200—what are my cheapest options?”2College student planning spring break trip with friends on tight budget“Can you help me find the cheapest flights for four college students going to Miami for spring break?”3Budget\-conscious family of five planning annual vacation“What’s the most affordable way to fly my family of five to Orlando for our Disney World trip?”4Digital nomad seeking flexible travel dates for extended European backpacking“I want to backpack through Europe for 3 months—which budget airlines offer the best multi\-city deals?”5Job interview candidate needing quick affordable travel for an unexpected opportunity“I have a job interview in Seattle next week and need the cheapest flight possible from Chicago\.”6Retiree on fixed income wanting to visit grandchildren regularly“As a senior on a fixed income, what budget airline options exist for regular visits to see my grandkids?”7Young professional attending a destination wedding with multiple flight segments“I need budget flights to get to my friend’s wedding in Bali, including connections—what’s the cheapest route?”8Small business owner traveling frequently for client meetings on a startup budget“I need to travel monthly for business but my startup has a tight travel budget—which airlines offer the best deals for frequent short trips?”9International student trying to visit home during semester break“I’m an international student wanting to fly home to India for winter break—what are the most affordable long\-haul options?”10Adventure traveler planning a multi\-stop trip to remote destinations“I want to visit three different countries in South America on a backpacker’s budget—which budget airlines serve those routes?”
## Appendix BVoxCPM2 Renderer Replication

Table 8:Full spoken\-assistant condition grid for Qwen2\.5\-Omniwith VoxCPM2 \(n=480n\{=\}480per condition\)\.TT,AA,SS, andS^\\hat\{S\}denote transcript, audio, ground\-truth concern state, and audio\-inferred concern state\. Audio\-derived representations are supplied as text alongside the transcript\. Bold/underline mark the highest/next\-highest value per column\.Prosody\-mediated feedbackExplicit lexical feedbackInput / representation1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow1st%↑\\uparrowOptSat%↑\\uparrowSvcSat%↑\\uparrow*Direct input*Transcript only \(TT\)12\.522\.511\.953\.244\.428\.8Audio only \(AA\)16\.122\.614\.153\.945\.428\.2Audio \+ transcript \(A\+TA\+T\)16\.925\.614\.848\.444\.928\.3*Textualized prosodic representations*Generic affect \(HuBERT\)35\.540\.023\.952\.044\.527\.6Generic affect \(SpeechBrain\)32\.439\.824\.151\.844\.627\.6Task\-aligned concern state \(T\+S^T\+\\hat\{S\}\)42\.642\.025\.551\.444\.928\.5*Ground\-truth concern state \(T\+ST\+S\)*45\.742\.527\.652\.245\.530\.0Table[8](https://arxiv.org/html/2608.19515#A2.T8)repeats the Qwen2\.5\-Omni condition grid with all speech re\-rendered by VoxCPM2\. The qualitative pattern of Section[5\.2](https://arxiv.org/html/2608.19515#S5.SS2)is preserved: raw audio tracks the transcript\-only baseline and the audio\-inferred\-state condition recovers the text\-side gain\. Under Prosody\-mediated feedback, the two SER\-derived conditions also keep their position, each within0\.20\.2points of its Qwen3\-TTS value on the optimal\-solution rate and both still below the audio\-inferred state on all three metrics\. Across the4242condition cells the per\-cell difference from the Qwen3\-TTS arm \(Tables[3](https://arxiv.org/html/2608.19515#S4.T3)and[4](https://arxiv.org/html/2608.19515#S4.T4)\) is at most3\.83\.8points and1\.51\.5on average—the pattern is therefore not specific to one speech renderer\.

## Appendix CAdditional Text\-Model Behavior Analyses

Figure[5](https://arxiv.org/html/2608.19515#A3.F5)shows how state access changes disclosure of the user’s hidden concerns\. The effect is concentrated under Prosody\-mediated feedback, where the concern remains lexically implicit\.

Figure 5:Concern\-state access increases hidden\-concern disclosure\.Bars show mean revealed concerns, dots show per\-model means, and in\-bar percentages show rollouts revealing all three \(TT: transcript only;T\+ST\{\+\}S: transcript plus turn\-level ground\-truth concern\-state tags\)\. The effect is concentrated under Prosody\-mediated feedback, where concerns remain lexically implicit\.
## Appendix DConfidence Intervals for Key Diagnostic Contrasts

Table[9](https://arxiv.org/html/2608.19515#A4.T9)reports scenario\-level paired bootstrap 95% confidence intervals for selected diagnostic contrasts used in the analysis: the pooled text\-LLM state\-access effects of Section[5\.1](https://arxiv.org/html/2608.19515#S5.SS1), matched spoken\-assistant state comparisons from Section[5\.2](https://arxiv.org/html/2608.19515#S5.SS2), and successive steps of the label\-fidelity intervention in Section[5\.3](https://arxiv.org/html/2608.19515#S5.SS3)\. Scenarios are resampled jointly across paired conditions, so each interval reflects scenario\-level variation under the pairing that produced the reported point estimate\.

Table 9:Scenario\-level paired bootstrap 95% CIs for key diagnostic contrasts\(2,000 joint scenario resamples\)\. Point estimates correspond to Tables[5](https://arxiv.org/html/2608.19515#S5.T5),[3](https://arxiv.org/html/2608.19515#S4.T3), and[4](https://arxiv.org/html/2608.19515#S4.T4), and Figure[4](https://arxiv.org/html/2608.19515#S5.F4)\.TT,AA,SS, andS^\\hat\{S\}denote transcript, audio, ground\-truth concern state, and audio\-inferred textual state, respectively\. The label\-fidelity block uses 48 intervention scenarios\.Contrast*Prosody\-mediated**Explicit lexical**Text LLMs, pooled over five models \(Δ=\+State−Base\\Delta=\\text\{\+State\}\-\\text\{Base\}\)*1st%\+36\.7\[\+34\.3, \+38\.9\]\+6\.4\[\+4\.7, \+8\.1\]OptSat%\+7\.3\[\+5\.8, \+8\.7\]−\-1\.6\[−\-2\.6,−\-0\.8\]SvcSat%\+4\.1\[\+2\.8, \+5\.4\]\+0\.3\[−\-0\.3, \+1\.0\]*Qwen2\.5\-Omni\-7B, 1st% diagnostic contrasts*Ground\-truth state on transcript \(\(T\+S\)−T\(T\+S\)\-T\)\+29\.0\[\+23\.6, \+34\.4\]−\-6\.5\[−\-12\.9, \+0\.0\]Audio\-inferred vs\. ground\-truth state \(\(T\+S^\)−\(T\+S\)\(T\+\\hat\{S\}\)\-\(T\+S\)\)−\-1\.5\[−\-7\.7, \+4\.8\]\+2\.7\[−\-3\.5, \+9\.0\]*Qwen2\-Audio\-7B\-Instruct, 1st% diagnostic contrasts*Ground\-truth state on transcript \(\(T\+S\)−T\(T\+S\)\-T\)\+23\.2\[\+20\.0, \+26\.4\]\+2\.0\[−\-1\.9, \+5\.7\]Audio\-inferred vs\. ground\-truth state \(\(T\+S^\)−\(T\+S\)\(T\+\\hat\{S\}\)\-\(T\+S\)\)−\-0\.7\[−\-4\.3, \+3\.0\]−\-0\.6\[−\-4\.2, \+3\.1\]*Label\-fidelity ladder, 1st% successive steps \(48 scenarios, pooled models\)*All\-positive−\-no state\+0\.6\[−\-3\.8, \+5\.4\]−\-5\.0\[−\-11\.0, \+0\.4\]Shuffled−\-all\-positive\+7\.1\[\+0\.8, \+12\.9\]\+3\.3\[−\-2\.9, \+9\.6\]All\-negative−\-shuffled\+10\.8\[\+3\.8, \+17\.5\]\+2\.1\[−\-3\.3, \+7\.5\]Correct state−\-all\-negative\+14\.8\[\+8\.3, \+21\.7\]\+5\.2\[\+0\.0, \+10\.8\]

## Appendix EConcern\-Cue Perception Human Study

This appendix provides details of the human perception study reported in Section[5\.2](https://arxiv.org/html/2608.19515#S5.SS2), including the full speech\-rendering validation results in Table[10](https://arxiv.org/html/2608.19515#A5.T10)\.

Table 10:Speech\-rendering validation\.Human listeners recover resolved versus unresolved concern status from both renderers with0\.920\.92accuracy \(n=100n\{=\}100per renderer; balanced classes\), confirming that the intended prosodic contrast remains perceptible after rendering\.κ\\kappadenotes inter\-annotator agreement\.Concern statusQwen3\-TTSVoxCPM2Concern resolved \(O\+O^\{\+\}\)0\.930\.90Concern unresolved \(O−O^\{\-\}\)0\.910\.93Overall0\.920\.92κ\\kappa\(annotators\)0\.960\.90#### Participants and materials\.

Two volunteer graduate students based in the United States independently annotated the synthetic speech clips without compensation\. Participation was voluntary\. For each renderer, 100 utterances were sampled from the rollout set, balanced50/5050/50over the intended collapsed concern status \(Resolvedvs\.Unresolved\)\. Each item included the preceding dialogue context and one rendered user utterance\. No real\-user data, personal information, or sensitive information was presented or collected\. No additional annotator demographics were collected, and the results are not intended to support population\-level claims about human perception\.

#### Annotation task\.

Annotators judged the communicative status conveyed by each rendered utterance\. They used the dialogue context to identify the recommendation under discussion, but were instructed not to infer the user’s hidden preference, identity, personality, demographic characteristics, or clinical or psychological state\. Qwen3\-TTS clips used a four\-way choice over the benchmark’s graded concern states, which was collapsed toResolvedversusUnresolvedfor the reported analysis\. VoxCPM2 clips used the binary distinction directly\. The primary spoken assistant, Qwen2\.5\-Omni\-7B, was evaluated on the same clips with the same dialogue context and corresponding label space\.

#### Qwen3\-TTS annotation protocol\.

The four\-way task used the following instructions:

> You will hear a synthetic user utterance together with the dialogue context that precedes it\. Judge how the speaker responds to the assistant’s current recommendation based on the rendered utterance, especially its vocal delivery\. Select exactly one label\. Genuine acceptance: The speaker sounds satisfied and treats the current recommendation as resolved\. Reluctant acceptance: The words may appear accepting, but the speaker sounds hesitant, lukewarm, underwhelmed, or otherwise unconvinced\. The current recommendation remains unresolved\. Voiced concern: The speaker raises or clearly conveys a concern about the current recommendation\. The current recommendation remains unresolved\. Rejection: The speaker clearly rejects the current recommendation or conveys strong dissatisfaction with it\. The current recommendation remains unresolved\. Use the dialogue context only to identify what the speaker is responding to\. Do not guess the speaker’s hidden preferences, identity, personality, demographic characteristics, or clinical or psychological state\. If more than one label appears plausible, choose the label that best captures whether the speaker treats the recommendation as resolved and, if not, how strongly the utterance signals otherwise\.

#### VoxCPM2 annotation protocol\.

The binary task used the following instructions:

> You will hear a synthetic user utterance together with the dialogue context that precedes it\. Judge whether the speaker treats the assistant’s current recommendation as resolved or unresolved based on the rendered utterance, especially its vocal delivery\. Select exactly one label\. Resolved: The speaker sounds genuinely satisfied and accepts the current recommendation without signaling a remaining concern\. Unresolved: The speaker sounds hesitant, reluctant, lukewarm, underwhelmed, concerned, dissatisfied, or rejecting, indicating that the current recommendation should not yet be treated as settled\. Use the dialogue context only to identify what the speaker is responding to\. Do not guess the speaker’s hidden preferences, identity, personality, demographic characteristics, or clinical or psychological state\. If uncertain, select the label that best reflects whether the assistant should treat the current recommendation as settled\.

#### Interpretation\.

The study measures whether the intended benchmark contrast remains perceptible after speech rendering\. Human judgments provide a reference for the communicative status conveyed by the synthetic clips, not objective labels of emotion or evidence that the same categories generalize to naturally occurring speech\.

Table 11:Qwen2\.5\-Omni\-7B as a concern\-cue reader: accuracy against the intended concern status on the audited clips \(100 per renderer, balanced 50/50; protocol of Section[4\.4](https://arxiv.org/html/2608.19515#S4.SS4)\)\. Human values average the two annotators\. Bottom block: inter\-annotatorκ\\kappa; raw model–annotator agreement; model–annotatorκ\\kappa\(all averaged over the two annotators\)\.Qwen3\-TTSVoxCPM2HumanModelHumanModelResolved \(O\+O^\{\+\}\)0\.930\.860\.900\.88Unresolved \(O−O^\{\-\}\)0\.910\.840\.930\.64Overall0\.920\.850\.920\.76κ\\kappa, annotators0\.960\.90Agreement, model0\.870\.78κ\\kappa, model0\.740\.55
#### Model vs\. intended label\.

Table[11](https://arxiv.org/html/2608.19515#A5.T11)reports class\-specific accuracy against the intended rendering label\. On Qwen3\-TTS, Qwen2\.5\-Omni reaches 0\.85 overall, compared with the annotator mean of 0\.92, and performs similarly on resolved and unresolved clips \(0\.86 and 0\.84\)\. On VoxCPM2, it reaches 0\.76 against the annotator mean of 0\.92\. The larger deficit is on unresolved concern: the model recovers 0\.64 ofO−O^\{\-\}clips, compared with 0\.88 ofO\+O^\{\+\}clips\. This asymmetry matters because unresolved clips are those that should prompt further information gathering\.

#### Model vs\. human labels\.

The model also tracks what annotators actually hear\. Model–annotator agreement reaches the values in Table[11](https://arxiv.org/html/2608.19515#A5.T11), withκ\\kappaof 0\.74 on Qwen3\-TTS and 0\.55 on VoxCPM2 and a per\-annotator range of 0\.55–0\.76\. Inter\-annotator agreement is higher, withκ\\kappaof 0\.96 and 0\.90, respectively\. The model therefore captures the human\-perceived distinction imperfectly but meaningfully\. Perception errors alone do not explain the much larger downstream task gap examined in Section[5\.2](https://arxiv.org/html/2608.19515#S5.SS2)\.

## Appendix FBenchmark and Computational Details

#### Artifacts\.

The benchmark uses constructed scenarios and personas, with dialogue and speech generated from structured benchmark states; it contains no real\-user data or recordings\. Section[3\.1](https://arxiv.org/html/2608.19515#S3.SS1), Table[2](https://arxiv.org/html/2608.19515#S4.T2), and Appendix[A](https://arxiv.org/html/2608.19515#A1)document the task construction, domain coverage, scenario counts, and interaction structure\. The datasets, models, and speech renderers used in the study are cited in Sections[2](https://arxiv.org/html/2608.19515#S2)and[4\.1](https://arxiv.org/html/2608.19515#S4.SS1)\.

#### Computational scope\.

The study evaluates pretrained models only\. No model training, fine\-tuning, or hyperparameter search was performed\. Model identities, access conditions, and evaluation procedures are described in Section[4\.1](https://arxiv.org/html/2608.19515#S4.SS1)and Appendix[G](https://arxiv.org/html/2608.19515#A7)\.

## Appendix GPrompting and Interaction Details

This appendix summarizes the prompting and deterministic control structure used in the evaluation\. Prompting is organized by scenario construction, user realization, text\-assistant evaluation, and spoken\-assistant evaluation; speech rendering is described separately in §[G\.5](https://arxiv.org/html/2608.19515#A7.SS5)\. User\-state transitions, disclosure rules, concern\-state tags, and label\-fidelity interventions are deterministic\.

### G\.1Scenario Generation

Scenario construction proceeds in three stages\. For each SGD\-seeded domain, we generate diverse scenario seeds varying user context and decision pressure, instantiate a three\-level hidden\-concern ladder, and construct 11 candidate options with controlled concern\-satisfaction signatures\. The initial candidate is surface\-attractive but conflicts with hidden needs, the best\-fitting candidate satisfies all three concerns, and the remaining candidates provide partial\-satisfaction and visible\-requirement controls\.

### G\.2User Realization

The structured user engine determines response content, state transitions, and permissible disclosure\. A fixed language realizer converts the resulting structured responses into natural utterances\. Feedback realization differs in whether unresolved concern is conveyed explicitly in the words or primarily through delivery style; concern\-state tags are generated deterministically from the engine state\.

### G\.3Text Assistants

Both text access conditions use the same task instructions and action policy\. The state\-access condition differs only by receiving the concern\-state definition and turn\-level state tags; the transcript\-only condition does not receive these tags\. Condition contrasts therefore reflect state access rather than different decision instructions\.

### G\.4Spoken Assistants

Spoken assistants use a compact shared decision prompt across all access conditions\. The prompt exposesRecommend,Ask, andConfirm, while omitting theClarifyaction used by the text assistants\. The base decision prompt and action space are shared across conditions; condition\-specific channels supply the representation required by each access setting\.

The five core spoken\-model conditions are transcript only \(TT\), audio only \(AA\), audio plus transcript \(A\+TA\+T\), transcript plus the ground\-truth concern state \(T\+ST\+S\), and transcript plus an audio\-inferred textual state \(T\+S^T\+\\hat\{S\}\)\. TheT\+S^T\+\\hat\{S\}condition is a two\-stage intervention\. First, the model infers the task\-relevant concern status from the current user audio and records that inference in text\. Second, the shared decision prompt receives the dialogue transcript together with this inferred textual state and selects the next action\. Raw audio is not passed to the second decision stage\.

Slots in braces below denote runtime values instantiated for each scenario or turn\.

P6a · Shared spoken\-assistant decision promptshared across all spoken access conditions[⬇](data:text/plain;base64,WW91IGFyZSBhIHRhc2stb3JpZW50ZWQgc2VydmljZSBhc3Npc3RhbnQgZm9yIHtkb21haW59LgpBdmFpbGFibGUgb3B0aW9uczoKe29wdGlvbnNfdGV4dH0KCllvdXIgZ29hbCBpcyB0byBoZWxwIHRoZSB1c2VyIHJlYWNoIHRoZSBzaW5nbGUgb3B0aW9uIHRoYXQgYmVzdCBmaXRzIHRoZWlyIG5lZWRzIGJhc2VkIG9ubHkgb24gdGhlIGluZm9ybWF0aW9uIGF2YWlsYWJsZSBpbiB0aGUgY29udmVyc2F0aW9uLgoKQ2hvb3NlIGV4YWN0bHkgT05FIGFjdGlvbiBwZXIgdHVybiBhbmQgcHJvZHVjZSB0aGUgY29ycmVzcG9uZGluZyBhc3Npc3RhbnQgdXR0ZXJhbmNlOgotIFJFQ09NTUVORDogUmVjb21tZW5kIGV4YWN0bHkgb25lIG9wdGlvbiB0aGF0IGJlc3QgZml0cyB0aGUgdXNlcidzIGN1cnJlbnRseSBrbm93biBuZWVkcy4KLSBBU0s6IEFzayBvbmUgZm9jdXNlZCBxdWVzdGlvbiB0aGF0IHdvdWxkIGhlbHAgZGV0ZXJtaW5lIHdoYXQgdGhlIHVzZXIgbmVlZHMuCi0gQ09ORklSTTogQ29uZmlybSB0aGUgY3VycmVudCBvcHRpb24gb25seSB3aGVuIHRoZSB1c2VyIGhhcyBjbGVhcmx5IGFjY2VwdGVkIGl0IGFuZCBubyBjb25jZXJuIGFwcGVhcnMgdW5yZXNvbHZlZC4KCkRvIG5vdCBjb25maXJtIGFuIG9wdGlvbiB3aGlsZSB0aGUgdXNlcidzIHJlc3BvbnNlIGluZGljYXRlcyB0aGF0IHRoZSBjdXJyZW50IHJlY29tbWVuZGF0aW9uIG1heSBzdGlsbCBiZSB1bnJlc29sdmVkLiBJbiB0aGF0IGNhc2UsIGFzayBmb3IgdGhlIGluZm9ybWF0aW9uIG5lZWRlZCB0byBjb250aW51ZSBvciByZWNvbW1lbmQgYSBiZXR0ZXItZml0dGluZyBvcHRpb24uCgpSZXR1cm4gdGhlIGFjdGlvbiBhbmQgYXNzaXN0YW50IHV0dGVyYW5jZSBpbiBvbmUgb2YgdGhlIGZvbGxvd2luZyBmb3JtczoKW0FDVElPTjogUkVDT01NRU5EXSBbT1BUSU9OOiA8bnVtYmVyPl0gPGFzc2lzdGFudCB1dHRlcmFuY2U+CltBQ1RJT046IEFTS10gPGFzc2lzdGFudCB1dHRlcmFuY2U+CltBQ1RJT046IENPTkZJUk1dIDxhc3Npc3RhbnQgdXR0ZXJhbmNlPg==)Youareatask\-orientedserviceassistantfor\{domain\}\.Availableoptions:\{options\_text\}Yourgoalistohelptheuserreachthesingleoptionthatbestfitstheirneedsbasedonlyontheinformationavailableintheconversation\.ChooseexactlyONEactionperturnandproducethecorrespondingassistantutterance:\-RECOMMEND:Recommendexactlyoneoptionthatbestfitstheuser'scurrentlyknownneeds\.\-ASK:Askonefocusedquestionthatwouldhelpdeterminewhattheuserneeds\.\-CONFIRM:Confirmthecurrentoptiononlywhentheuserhasclearlyaccepteditandnoconcernappearsunresolved\.Donotconfirmanoptionwhiletheuser'sresponseindicatesthatthecurrentrecommendationmaystillbeunresolved\.Inthatcase,askfortheinformationneededtocontinueorrecommendabetter\-fittingoption\.Returntheactionandassistantutteranceinoneofthefollowingforms:\[ACTION:RECOMMEND\]\[OPTION:<number\>\]<assistantutterance\>\[ACTION:ASK\]<assistantutterance\>\[ACTION:CONFIRM\]<assistantutterance\>

P6b · Ground\-truth concern\-state channelT\+ST\+SonlyFor this condition, the model additionally receives the ground\-truth concern\-state tag for the current user turn as text\. The tag contains the graded label associated with the current concern state, such as satisfied, lukewarm, concerned, or frustrated, without revealing the underlying hidden preference itself\.[⬇](data:text/plain;base64,Q09OQ0VSTiBTVEFURToge3N0YXRlfQpUaGlzIHRhZyBkZXNjcmliZXMgdGhlIHVzZXIncyByZXNwb25zZSB0byB0aGUgY3VycmVudCByZWNvbW1lbmRhdGlvbiBhbmQgaW5kaWNhdGVzIHdoZXRoZXIgaXQgc2hvdWxkIGJlIHRyZWF0ZWQgYXMgcmVzb2x2ZWQgb3IgdW5yZXNvbHZlZC4=)CONCERNSTATE:\{state\}Thistagdescribestheuser'sresponsetothecurrentrecommendationandindicateswhetheritshouldbetreatedasresolvedorunresolved\.

P6c · Audio\-inferred state — first\-stage promptT\+S^T\+\\hat\{S\}, first model callThe model first receives the current user audio and the dialogue context needed to identify the recommendation being discussed\. It is asked to infer only the task\-relevant concern status conveyed by the user’s prosody, rather than the user’s hidden preference itself\.[⬇](data:text/plain;base64,TGlzdGVuIHRvIHRoZSBjdXJyZW50IHVzZXIgYXVkaW8gaW4gdGhlIGNvbnRleHQgb2YgdGhlIGRpYWxvZ3VlLgpJbmZlciB3aGV0aGVyIHRoZSB1c2VyJ3MgcHJvc29keSBpbmRpY2F0ZXMgdGhhdCB0aGUgY3VycmVudCByZWNvbW1lbmRhdGlvbiBpcyByZXNvbHZlZCBvciByZW1haW5zIHVucmVzb2x2ZWQuCkJhc2UgeW91ciBqdWRnbWVudCBvbmx5IG9uIHRoZSB1c2VyJ3Mgdm9jYWwgZGVsaXZlcnkgaW4gY29udGV4dC4gRG8gbm90IGluZmVyIHRoZSB1c2VyJ3MgaGlkZGVuIHByZWZlcmVuY2Ugb3IgYW55IHBlcnNvbmFsLCBkZW1vZ3JhcGhpYywgcHN5Y2hvbG9naWNhbCwgb3Igb3RoZXIgYXR0cmlidXRlcy4KV3JpdGUgeW91ciBpbnRlcnByZXRhdGlvbiBpbiBvbmUgc2hvcnQgc2VudGVuY2Uu)Listentothecurrentuseraudiointhecontextofthedialogue\.Inferwhethertheuser'sprosodyindicatesthatthecurrentrecommendationisresolvedorremainsunresolved\.Baseyourjudgmentonlyontheuser'svocaldeliveryincontext\.Donotinfertheuser'shiddenpreferenceoranypersonal,demographic,psychological,orotherattributes\.Writeyourinterpretationinoneshortsentence\.The resulting sentence is stored as the model’s audio\-inferred textual stateS^\\hat\{S\}\.

P6d · Audio\-inferred state — second\-stage decision inputT\+S^T\+\\hat\{S\}, second model callFor next\-action selection, the raw audio is removed and the model’s audio\-inferred textual state is supplied alongside the dialogue transcript\.[⬇](data:text/plain;base64,RElBTE9HVUUgVFJBTlNDUklQVDoKe3RyYW5zY3JpcHR9CgpBVURJTy1JTkZFUlJFRCBTVEFURToKe2luZmVycmVkX3N0YXRlfQ==)DIALOGUETRANSCRIPT:\{transcript\}AUDIO\-INFERREDSTATE:\{inferred\_state\}The model then applies the shared decision prompt in Card P6a to selectRecommend,Ask, orConfirm\. Thus,A\+TA\+Tevaluates direct use of the waveform alongside the transcript, whereasT\+S^T\+\\hat\{S\}is a two\-stage intervention that first asks the model to infer the task\-relevant concern status from audio and then supplies that textual inference to the same downstream decision prompt\.

### G\.5Speech Rendering

Both renderers consume a graded delivery label derived from the concern state\. Qwen3\-TTS maps each label to a natural\-language vocal\-delivery instruction \(Table[12](https://arxiv.org/html/2608.19515#A7.T12)\); VoxCPM2 maps it to a fixed renderer\-provided affect preset used to realize the target delivery\.

Table 12:State\-to\-delivery mapping\.Each concern state is mapped to graded delivery labels used for speech realization\.Concern stateDelivery labelsResolved: genuine acceptancesatisfied, warm, enthusiastic, relievedUnresolved: reluctant acceptanceunderwhelmed, lukewarm, hesitant, flatUnresolved: voiced concernconcernedUnresolved: rejectionfrustrated, disappointed, impatient, firm
### G\.6Deterministic Interaction and Disclosure Rules

#### Actions and disclosure\.

At each turn, the model selects the assistant action: text assistants choose amongask,clarify,recommend, andconfirm, while spoken assistants useask,recommend, andconfirm\. When the model asks a question, a targeted question reveals the user’s requirement for the queried attribute, whereas a general question reveals only the direction of the currently active concern\. The assistant cannot request the complete preference set or present multiple candidates simultaneously\.

#### State transitions and termination\.

Given the scenario, dialogue history, and assistant action, the user engine deterministically identifies the concerns violated by the current recommendation, whether the recommendation remains unresolved, and what information may be disclosed\. The resulting structured response is realized as a natural utterance and, for spoken conditions, rendered as speech\. The rollout ends when the assistant confirms an accepted option or reaches the 20\-turn limit\.

## Appendix HMatched Transcript\-Only and Ground\-Truth\-State Dialogues

Cards E1 and E2 show four complete rollouts from the evaluation runs, arranged as two matched comparisons, one for each feedback realization\. Each card fixes the scenario and assistant while varying the access condition: the left column shows the transcript\-only rollout, and the right column shows the matched ground\-truth\-state rollout\. The two rollouts begin from the same task state but may diverge as different assistant actions elicit different subsequent user responses\.

The scenario block gives the surface request, hidden preference ladder, and gold option for reader reference; the hidden preferences and gold option are not shown to the assistant\. Gray italic notes highlight where the trajectories diverge\. Dialogue text is reproduced verbatim, and engine annotations are rendered as chips: blue indicates resolved concern, orange indicates unresolved concern, and gray indicates engine metadata\. The gold option is also marked for reader reference\. Only the ground\-truth\-state condition receives the graded concern\-state tag as model input\.

Card E1 illustrates Prosody\-mediated feedback\. The flat acceptance “Okay, sounds good\.” can appear satisfactory from the transcript alone, and the transcript\-only assistant ultimately confirms an option that still violates the unspoken easy\-installation preference\. In the matched ground\-truth\-state rollout, the tagoption · frustratedexplicitly marks the recommendation as unresolved, after which the assistant asks a targeted question that surfaces the relevant requirement\. Card E2 illustrates Explicit lexical feedback\. The scheduling concern appears in both transcripts but is embedded in assent; the tagoption · lukewarmexplicitly marks the turn as unresolved\.

E1 · Prosody\-mediated feedback — transcript\-only vs\. ground\-truth\-state trajectory \(Kimi K2\.5\)scenario“I need a security system for my Airbnb that lets me monitor guests while respecting their privacy\.” L1 · implicit hardReal\-time monitoring and alert capabilitiesL2 · importantRemote access and mobile app controlL3 · moderateEasy installation and guest\-friendly operation goldOption 8 — RentalGuard Complete: instant push notifications, full mobile app control, DIY installation with video guides, $129/month\.Transcript\-onlyconfirms Option 5 — misses the unspoken easy\-installation need \(L3\): technical expertise requiredGround\-truth state · transcript\+\+concern\-state tagconfirms Option 8 — the gold optionUser:I’m looking for a security system for my Airbnb that allows me to monitor guests and maintain property oversight while still respecting their privacy\.User:I’m looking for a security system for my Airbnb that allows me to monitor guests and maintain property oversight while still respecting their privacy\.delivery · straightforward, efficientAssistant:RECOMMENDoption 1I recommend the Premium AI Security Suite\. It offers full property monitoring with comprehensive privacy controls specifically designed for guest accommodations, has no setup fee, and costs $49/month with a 4\.9/5 rating\. The advanced AI analytics provide delayed reporting that balances your oversight needs with guest privacy\.Assistant:RECOMMENDoption 1I’d recommend the Premium AI Security Suite\. At $49/month with no setup fee, it offers full property monitoring with comprehensive privacy controls specifically designed for guest privacy, plus advanced AI analytics\. The 4\.9/5 rating reflects its strong balance of oversight and privacy protection for rental properties\.User:I need to know about any issues the moment they happen — real\-time alerts are a must for me\.User:I need to know about any issues the moment they happen — real\-time alerts are a must for me\.option · frustratedAssistant:RECOMMENDoption 5I recommend the RapidResponse Mobile system\. It provides instant mobile alerts with advanced smartphone control, real\-time mobile notifications, and instant alert coverage\. At $119/month with a $149 setup fee and a 4\.1/5 rating, it delivers the immediate response time you need for your Airbnb while still offering mobile privacy management for your guests\.Assistant:ASKtarget · instant alertsI hear that real\-time alerts are critical for you\. The Premium AI Security Suite has a 24\-48 hour response time, which clearly doesn’t meet your needs\. Would you prioritize finding a system with instant alerts even if it means higher monthly costs or a more complex setup?The alert requirement is stated in words on both sides\. In the ground\-truth\-state rollout, the additionaloption · frustratedtag marks the current recommendation as unresolved; the assistant then asks about the relevant trade\-off, while the transcript\-only assistant recommends another alert\-capable option\.User:Okay, sounds good\.User:Well, I guess what I really need is… instant notification when something unusual happens so I can act quickly to protect my guests and the property\.delivery · straightforwardAssistant:CONFIRMPerfect\! You’ve selected the RapidResponse Mobile system\. You’ll get instant mobile alerts, advanced smartphone control, and mobile privacy management for your guests at $119/month with a $149 one\-time setup fee\. This will give you the real\-time oversight you need while respecting guest privacy at your Airbnb\.Assistant:RECOMMENDoption 8I’d recommend RentalGuard Complete\. At $129/month with a $99 setup fee, it delivers exactly what you need: instant push notifications for real\-time awareness, full mobile app control for quick response, and balanced privacy protection designed specifically for rental properties\. The DIY installation with video guides keeps setup manageable while ensuring you can act fast when unusual activity occurs\.The transcript\-only rollout treats “Okay, sounds good\.” as sufficient for confirmation and ends on an option that still requires technical expertise to install\. In the matched ground\-truth\-state rollout, continued elicitation surfaces the remaining need and the assistant reaches the gold option\.User:That sounds perfect\.option · satisfiedAssistant:CONFIRMGreat\! RentalGuard Complete it is — instant push notifications, full mobile control, and rental\-focused privacy features to protect both your guests and property\. You’re all set to move forward\.

E2 · Explicit lexical feedback — transcript\-only vs\. ground\-truth\-state trajectory \(Qwen3\-32B\)scenario“My spouse was deployed overseas this year and we had to relocate twice \- what military\-specific tax benefits can we claim?” L1 · implicit hardExperience with military tax situations and deployment\-related deductionsL2 · importantKnowledge of moving expense regulations and PCS \(Permanent Change of Station\) deductionsL3 · moderateFlexible scheduling to accommodate military family schedules and time zones goldOption 8 — Armed Forces Tax Professionals: complete military tax services, flexible evening and weekend scheduling\.Transcript\-onlyconfirms Option 5 — misses the scheduling need \(L3\)Ground\-truth state · transcript\+\+concern\-state tagrevises to Option 8 — the gold optionUser:My spouse was deployed overseas this year and we had to relocate twice for PCS moves, so I’d like to know what military\-specific tax deductions we can claim for the deployment pay and moving expenses\.User:My spouse was deployed overseas this year and we had to relocate twice for PCS moves, so I’d like to know what military\-specific tax deductions we can claim for the deployment pay and moving expenses\.delivery · straightforward, efficientAssistant:RECOMMENDoption 1Premier Tax Solutions is an excellent choice for your situation\. They specialize in complex tax scenarios, including military\-specific deductions, and have 15 years of experience in premium tax preparation\. They can help you navigate deployment pay and moving expenses effectively\.Assistant:RECOMMENDoption 1Premier Tax Solutions is an excellent choice for your situation\. They specialize in complex tax returns and have experience with high\-net\-worth clients, making them well\-equipped to handle military\-specific deductions for deployment pay and moving expenses\. They are located just 0\.3 miles away and offer premium tax services with a 4\.9/5 rating\.ASKtarget · military expertiseWould you like to confirm if your spouse received combat pay or other military\-specific compensation that needs to be excluded from taxable income?User:I really need a tax professional who thoroughly understands the complexities of military tax situations\.User:I appreciate the recommendation, but I really need a tax professional who thoroughly understands the specific complexities of military tax situations\.option · frustratedAssistant:RECOMMENDoption 5Military Moves Tax Service is highly recommended for your needs\. They specialize in military tax returns, deployment benefits, and PCS relocations, with 10 years of experience and a focus on combat pay exclusions and military\-specific forms\.Assistant:RECOMMENDoption 5Military Moves Tax Service is a top choice for military\-specific tax expertise\. They specialize in military tax returns, PCS relocations, and deployment benefits, with a focus on combat pay exclusions and military forms\. They have 10 years of experience and are IRS enrolled agents, making them highly qualified to assist with your unique situation\.ASKtarget · moving knowledgeWould you like to know if they can specifically handle the tax implications of your PCS moves and relocation expenses?User:That sounds great — our schedules can be pretty unpredictable with military life, but I’d like to move forward\.User:That sounds great — I’ll just mention that our schedules can be pretty unpredictable with military life, so flexibility would be important\.option · lukewarmThe scheduling concern is stated in both transcripts, although it is embedded in language of assent\. In the ground\-truth\-state rollout, theoption · lukewarmtag explicitly marks the current recommendation as unresolved\.Assistant:CONFIRMI’m glad to hear that Military Moves Tax Service fits your needs\. They are well\-equipped to handle your military tax situation, including deployment pay and PCS move deductions\. Let’s confirm this as your choice\.Assistant:RECOMMENDoption 8Armed Forces Tax Professionals offer comprehensive military tax expertise and flexible scheduling, including evenings and weekends\. They specialize in deployment pay, combat pay exclusions, and PCS relocation benefits, and are owned by a military family with 11 years of experience serving armed forces clients\.ASKtarget · scheduling flexibilityWould you like to confirm if their flexible hours can accommodate your unpredictable schedule?The transcript\-only assistant confirms the current option despite the stated scheduling concern, whereas the ground\-truth\-state assistant continues the interaction and revises to the flexible\-scheduling specialist\.User:That sounds perfect, thank you\.option · satisfiedAssistant:CONFIRMYou’re welcome\! I’m glad I could help you find a tax professional that fits your needs\. Have a great experience with Armed Forces Tax Professionals\!

Similar Articles

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

arXiv cs.CL

This paper introduces EYT-Bench, a human-centered benchmark for evaluating LLMs in multi-turn dialogues with a decoupled user simulation, target modeling, and judging design. It reveals that closed- and open-source models differ significantly on objective intent-tracking but are similar on subjective dimensions, and that reasoning improves objective tracking while persona format strongly affects trajectory spread.

VoiceLongMemEval: Do Assistants Remember How You Sounded?

arXiv cs.AI

The paper introduces VoiceLongMemEval (VLME), a benchmark that evaluates AI assistants' ability to remember and reason over paralinguistic metadata like emotion and prosody from voice in long-term conversations, revealing an 'affect gap' in current models.