Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

arXiv cs.CL Papers

Summary

This paper introduces Omni Demand Understanding (ODU), a benchmark to evaluate how well multimodal AI models infer user demands from complex audio-visual interactions, revealing significant performance gaps in current models.

arXiv:2609.21392v1 Announce Type: new Abstract: Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs' ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.
Original Article
View Cached Full Text

Cached at: 09/21/26, 09:07 AM

# Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
Source: [https://arxiv.org/html/2609.21392](https://arxiv.org/html/2609.21392)
Qi Chen1,2,3,\*,§Yunfei Chu3,\*Haolin He3,4,\*,§Yifan Yang1,3,§Zihan Liu3,§Yuxuan Wang3Ziyang Ma1,2Ruiyang Xu1,3,§Meng Gao3,5,§Yinsong Yan3,6,§Ling Wang3,6,§Hui Wang3,7,§Wen Huang8Yiheng Chen1,3,§Guanrou Yang1,2Qiuqiang Kong4Jin Xu3,†Xie Chen1,2,†1Shanghai Jiao Tong University2Shanghai Innovation Institute3Alibaba Token Hub, Alibaba Group4The Chinese University of Hong Kong5Tsinghua University6Hong Kong Polytechnic University7Nankai University8Johns Hopkins University\*Equal contribution\.†Corresponding author§Work done during an internship at Alibaba Token Hub, Alibaba Group\.

###### Abstract

Natural audio–visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts\. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored:*can a model correctly infer the user’s underlying demand from a complex multimodal interaction?*Real\-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history\. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments\. Conversely, request\-like speech may not constitute a demand to the assistant, leading to false triggers\. We establishOmni Demand Understanding \(ODU\)as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer the user’s intent from multimodal and conversational context\. ODU evaluates this capability along five dimensions, covering both single\-turn and multi\-turn interactions\. We construct ODU\-Bench using a challenge\-driven taxonomy, taxonomy\-guided agentic video generation, and human\-recorded interactions, followed by media\-grounded annotation and human verification\. We evaluate 14 native audio and audio–visual MLLMs\. Even the strongest, Gemini 3\.1 Pro, recovers only44\.744\.7% of the key information that must be inferred from visual, acoustic, or conversational context\. Moreover, 11 of the 14 models exhibit false\-trigger rates above 50% on non\-demand scenarios\. These results reveal a systematic capability gap in current MLLMs’ ability to infer contextual user demands\. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response\.

![Refer to caption](https://arxiv.org/html/2609.21392v1/teaser.png)Figure 1:Representative challenging scenarios in ODU\.Demand\-bearing cases \(a–d\) require integrating visual, acoustic, and conversational context while handling challenging expression forms and acoustic conditions, while no\-demand cases \(e–g\) contain demand\-like expressions that should not trigger responses\.## 1Introduction

Multimodal large language models \(MLLMs\) that jointly perceive audio, vision, and language are increasingly serving as conversational assistants\([Google DeepMind, 2026](https://arxiv.org/html/2609.21392#bib.bib15);[Seed, 2026](https://arxiv.org/html/2609.21392#bib.bib16);[Qwen Team, 2026a](https://arxiv.org/html/2609.21392#bib.bib12);[Xu et al\., 2025b](https://arxiv.org/html/2609.21392#bib.bib13);[Xu et al\., 2025a](https://arxiv.org/html/2609.21392#bib.bib25);[Cui et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib11);[Hurst et al\., 2024](https://arxiv.org/html/2609.21392#bib.bib14);[AI et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib23);[Deshmukh et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib18);[Tang et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib17)\)\. Yet existing benchmarks on multimodal interactions primarily evaluate response quality, implicitly assuming that the user demand behind a multimodal query has already been correctly identified and understood\([Wang et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib2);[Selvakumar et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib9);[Lu et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib10);[Zhao et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib3)\)\. This assumption is fragile because human expression is inherently underspecified\. By the Principle of Least Effort\([Zipf, 1949](https://arxiv.org/html/2609.21392#bib.bib1)\), people tend to avoid unnecessary explicitness, leaving parts of their intended meaning to be recovered from shared multimodal and conversational context\. As shown in Fig\.[1](https://arxiv.org/html/2609.21392#S0.F1)\(b\), expressions such as “the same test” and “this one” require conversation history and visual cues from user gestures to resolve\. Conversely, Fig\.[1](https://arxiv.org/html/2609.21392#S0.F1)\(e–g\) show that demand\-like expressions do not necessarily imply the existence of actual user demands\. Recovering a demand is therefore a contextual reasoning task that requires integrating multimodal and conversational evidence, not merely transcribing the utterance\. However, existing benchmarks largely treat user queries as self\-contained text or speech, overlooking the multimodal and conversational context required to infer them and leaving this capability under\-evaluated\.

To diagnose this capability systematically, we introduceOmni Demand Understanding \(ODU\), a task which makes demand understanding an explicit prediction target rather than an implicit precondition of response generation\. Given an audio or audio–visual interaction, ODU evaluates a model along five dimensions: whether a valid demand is present, when it occurs, the transcript of the demand\-bearing speech, the user’s intent as resolved from multimodal and conversational context, and the user profile\. This diagnosis separates perception and localization from contextual user\-intent inference: a model must not only recover the utterance, but also infer the user’s intent and the response\-relevant information left implicit\. Semantic recovery is evaluated against atomic key points, while no\-demand scenes test whether request\-like signals could trigger the assistant\.

To build ODU\-Bench, we start from demand\-understanding challenges in everyday interaction and organize them into a challenge\-driven taxonomy\. Guided by sampled taxonomy targets, an agentic pipeline constructs scenarios coarse\-to\-fine, progressing from demand semantics or invalidation conditions to causal interactions, dialogue, and temporal realization\. After challenge\-validity and transcript\-sufficiency screening, accepted scripts guide the production of synthetic interactions\. We also recruited human actors to perform daily\-life interactions, providing the benchmark with a complementary source of behaviorally and acoustically realistic data\. For both synthetic and human\-recorded interactions, demand annotations are reconstructed from the realized media\. Every scene and its annotation then undergo careful human verification\. Together, these steps yield 2 078 diverse and challenging demand and no\-demand scenes\. Across a broad panel of native multimodal models, even Gemini 3\.1 Pro, the strongest system on our aggregate score, falsely triggers on40\.640\.6% of no\-demand scenes\. Evidence\-channel analysis shows that it recovers82\.982\.9% of key points supported by the spoken request but only44\.744\.7% of those requiring multimodal or dialogue context\. This gap confirms that ODU cannot be reduced to transcription: models must select and bind evidence across the interaction to understand the user’s intent\. In summary, our contributions are as follows:

1. 1\.A new problem formulation\.We establish omni demand understanding as a distinct multimodal contextual inference problem, making it an explicit evaluation target rather than leaving it implicit in response\-generation performance\.
2. 2\.A challenging benchmark\.We construct a benchmark covering diverse demand and no\-demand scenarios, combining taxonomy\-guided agentic generation with human\-recorded interactions\.
3. 3\.A systematic capability gap\.We show that current MLLMs systematically struggle with contextual user\-intent inference, particularly when it requires multimodal and conversational evidence beyond the explicit spoken request\.

## 2Related Work

Table 1:Benchmark comparison\. Visual and acoustic cues refer to intent inference\.BenchmarkDemandPresenceIntent PredictionFormatSpoken UtteranceInputMulti\-Turn DialogueInputVisualCuesAcousticCuesMultimodal interaction benchmarksOmniMMI\([Wang et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib2)\)◐✗◐✓◐✗MultiVox\([Selvakumar et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib9)\)✗✗✓✗✓✓OmniInteract\([Lu et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib10)\)◐✗✓✓◐✗Full\-Duplex\-Bench\-v2\([Lin et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib29)\)◐✗✓✓✗✗Intent and goal understanding benchmarksMIntRec2\.0\([Zhang et al\., 2024](https://arxiv.org/html/2609.21392#bib.bib4)\)✗Label✓◐✓✓WAGIBench\([Veerabadran et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib6)\)✗Open\-ended✗✗✓✗GUIDE\([Yang et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib5)\)✓MCQ✗✗✓✗EgoIntrospect\([Wang et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib7)\)◐MCQ◐✗◐✗ODU✓Open\-ended✓✓✓✓

✓: explicitly evaluated;◐: partial or indirect support;✗: not established\. MCQ: multiple\-choice question\. In input columns,✓means that the input is provided\.

#### Multimodal interaction\.

Research on spoken and multimodal interaction examines response quality, turn\-taking, and the coordination of listening and speaking\. Full\-Duplex\-Bench evaluates pause handling, backchanneling, and interruptions\([Lin et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib27)\), while MTR\-DuplexBench examines multi\-round dialogue quality and instruction following\([He et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib28)\)\. Full\-Duplex\-Bench\-v2 further evaluates multi\-turn turn\-taking and instruction following with an automated examiner\([Lin et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib29)\)\. MultiVox evaluates responses grounded in paralinguistic and visual cues\([Selvakumar et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib9)\), and VideoFDB benchmarks nonverbal behavior in full\-duplex audio–visual conversations\([Mazumdar et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib30)\)\. OmniMMI studies streaming understanding and multi\-turn dependencies\([Wang et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib2)\); OmniInteract tests trigger timing, interruptions, and nested exchanges\([Lu et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib10)\); and OmniPro emphasizes modality necessity in proactive streaming evaluation\([Zhao et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib3)\)\. ODU complements these benchmarks by evaluating demand understanding separately from response quality and timing\.

#### Multimodal intent and contextual goal understanding\.

Research on multimodal intent and goal understanding examines how language, behavior, and surrounding context reveal users’ intentions\. MIntRec2\.0 evaluates conversational intent classification using textual, acoustic, and visual cues\([Zhang et al\., 2024](https://arxiv.org/html/2609.21392#bib.bib4)\), with substantial textual bias identified in subsequent analyses\([Mullick et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib8)\)\. SIMMC 2\.0 studies multimodal disambiguation and coreference in shopping dialogues\([Kottur et al\., 2021](https://arxiv.org/html/2609.21392#bib.bib31)\)\. GUIDE studies assistance needs from GUI activity\([Yang et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib5)\); WAGIBench infers unexpressed goals from egocentric context\([Veerabadran et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib6)\); and EgoIntrospect evaluates request recovery with the spoken request withheld\([Wang et al\., 2026](https://arxiv.org/html/2609.21392#bib.bib7)\)\. ODU complements these studies with joint evaluation of demand presence and open\-ended demand semantics in challenge\-driven daily\-life human–machine interactions, created through media synthesis and realistic human\-recorded interactions\. Table[1](https://arxiv.org/html/2609.21392#S2.T1)comparesOduwith representative benchmarks\.

## 3Task Formulation and Evaluation

### 3\.1Problem setup and prediction targets

A*demand*is the outcome a user wants an assistant to achieve, including response\-relevant objects, constraints, and trigger conditions\.Oduevaluates the final user turn in an audio or audio–visual interaction, using the preceding interaction as context\. Earlier assistant replies are supplied as text\. The task asksfive core questions: whether a valid demand exists; what the user intends and which contextual information is required to resolve it; when the demand occurs; what the user says; and who expresses it\. The first two assess contextual understanding and reasoning over multimodal and conversational evidence\. The latter three primarily assess perception through temporal localization, transcription, and user profile prediction\. Together, they cover the perceptual foundations and contextual reasoning required for demand understanding\. The output schema is shown in App\.[A](https://arxiv.org/html/2609.21392#A1)\.

### 3\.2Key\-point\-based semantic evaluation

We represent demand understanding with two open\-ended fields:*structured intent*for the desired outcome and*required context*for the multimodal or dialogue context needed to resolve it\. We evaluate these fields against atomic*key points*extracted from ground\-truth annotations; omitting or misstating any point could change an appropriate response to the user demand\. Each key point is labeled by evidence source for channel\-level diagnosis \(§[5\.3](https://arxiv.org/html/2609.21392#S5.SS3)\)\. All key points are grounded in the media and verified by humans\. An LLM judge checks whether the concatenated predicted structured intent and required context cover each reference key point, accepting semantically equivalent wording\. We audit key\-point recoverability from source annotations and the stability of model conclusions across judges in App\.[E\.2](https://arxiv.org/html/2609.21392#A5.SS2)\.

### 3\.3Evaluation metrics

Table 2:Evaluation dimensions and their contribution to the overall score\.Output measuredMetricWeightM1demand presentbinary macro\-F1 over demand and no\-demand scenes0\.15M2structured intent, required contextkey\-point hit\-rate0\.60M3demand spantemporal IoU0\.10M4transcriptmax⁡\(0,1−adaptive CER/WER\)\\max\(0,1\-\\text\{adaptive CER/WER\}\)0\.10M5user profilemean per\-field accuracy0\.05Table[2](https://arxiv.org/html/2609.21392#S3.T2)summarizes M1–M5\. M1 uses demand and no\-demand scenes; M2–M5 use only demand\-bearing scenes\. FTR is the fraction of no\-demand scenes that trigger the assistant\. The overall score is a weighted average of M1–M5, with the largest weight on M2 to prioritize semantic recovery\. App\.[A](https://arxiv.org/html/2609.21392#A1)gives full definitions and aggregation details\.

## 4Benchmark Construction

We construct ODU\-Bench through a challenge\-driven taxonomy and seed corpus \(§[4\.1](https://arxiv.org/html/2609.21392#S4.SS1)\), agentic scenario generation \(§[4\.2](https://arxiv.org/html/2609.21392#S4.SS2)\), and media\-grounded annotation with quality control \(§[4\.3](https://arxiv.org/html/2609.21392#S4.SS3)\)\. Fig\.[3](https://arxiv.org/html/2609.21392#S4.F3)summarizes the pipeline; App\.[C](https://arxiv.org/html/2609.21392#A3)provides detailed construction and human verification procedures\. The constructed benchmark contains both audio\-only and audio\-visual modalities\.

### 4\.1Challenge\-driven taxonomy and seed corpus

\(a\)Distribution of positive samples by taxonomy axis\.
\(b\)Distribution of negative samples by signal type and invalidation pattern\.
\(c\)Coverage across task and scenario categories\.

Figure 2:Taxonomy coverage in ODU\-Bench\.\(a\) and \(b\) show positive and negative sample distributions across their taxonomy axes; segment angles are proportional to sample counts within each axis\. \(c\) shows task and scenario coverage, with word size reflecting frequency\.The core challenge of user demand understanding lies in how demands are expressed and how the surrounding context shapes their interpretation\. Real\-world demands may depend on visual or acoustic context, remain ambiguous without prior interaction, be expressed implicitly, or occur under distracting environmental conditions\. In contrast, task semantics and scenario domains primarily characterize what users want to accomplish and where interactions occur, rather than what makes the demands difficult to understand\. We therefore organize scenario generation around sources of demand\-understanding difficulty, while separately characterizing task and domain coverage\.

We define a six\-axis taxonomy for positive\-demand scenes: four independently composable challenge axes for generation and two descriptive axes for task and domain coverage\. No\-demand scenes use a separate taxonomy of demand\-like signal types and invalidation reasons\. Fig\.[2](https://arxiv.org/html/2609.21392#S4.F2)shows the taxonomy and resulting distribution; App\.[B](https://arxiv.org/html/2609.21392#A2)provides the definitions\.

We combine seeds from a corpus of everyday contexts and user needs with sampled taxonomy targets to vary the setting of each challenge\. This separates*challenge coverage*from*scenario diversity*\.

### 4\.2Agentic scenario generation

![Refer to caption](https://arxiv.org/html/2609.21392v1/pipeline.png)Figure 3:Overview of the ODU\-Bench construction pipeline\.Challenge\-driven targets are expanded into coarse\-to\-fine scripts, screened for plausibility and challenge validity, and realized as synthesized or human\-recorded interactions\. Annotations are then reconstructed from complementary media evidence and verified through automatic checks and human review\.We instantiate each scenario from a sampled taxonomy target and scenario seed through three stages: hierarchical script generation, joint plausibility and challenge\-validity review, and continuity\-aware rendering\. Script generation with Qwen3\.7\-Max\([Qwen Team, 2026c](https://arxiv.org/html/2609.21392#bib.bib20)\)proceeds coarse\-to\-fine: we first define a valid demand or an invalidating condition, build a causal scenario skeleton, and then add contextual, linguistic, and temporal details\. Separating demand semantics from their realization supports controlled generation of recoverable demands and demand\-like cues that require contextual evidence to disambiguate\.

The script reviewer checks plausibility, causal coherence, taxonomy alignment, and whether the intended demand is recoverable from the evidence specified in the script\. Its text\-only discriminator sees only the scripted user utterance, with dialogue history, speaker information, and media evidence withheld\. For scenarios designed to require contextual evidence, we reject positives whose intent is already recoverable from that utterance and negatives whose no\-demand status is textually obvious\. This is a script\-stage screening criterion, not a test of exclusive channel dependence in the realized media\.

Accepted scripts specify characters, environments, dialogue, and observable events along a timeline\. We produce synthetic interactions in both modalities from these scripts, retaining only the audio track for audio\-only scenes\. Multi\-clip rendering proceeds sequentially, conditioning later clips on earlier videos to preserve character and environmental continuity\. Automatic review and human inspection assess perceptual quality, scenario fidelity, and cross\-clip coherence where applicable\. We also use accepted scripts as performance guides for human\-recorded interactions\. These preserve the taxonomy targets and interaction structure while introducing real actors, environments, microphones, timing, and speech variation\. Both media sources follow the same media\-grounded annotation and quality\-control pipeline\.

### 4\.3Media\-grounded annotation and quality control

Rendered and human\-recorded media may differ from scripts in scene details, speech, timing, and demand boundaries\. We therefore derive ground\-truth annotations from the media, using scripts as structural guides\. ASR\([An et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib19)\)provides transcripts and word\-level timestamps; multiple MLLMs\([Qwen Team, 2026a](https://arxiv.org/html/2609.21392#bib.bib12);[Google DeepMind, 2026](https://arxiv.org/html/2609.21392#bib.bib15)\)observe visual and acoustic events, speaker relations, and interaction context\. We reconcile these observations with the script to determine demand presence and derive the annotations\.

Rule\-based checks and LLM/MLLM review assess consistency between scripts and media, transcript and timestamp alignment, schema validity, agreement between demand labels and segments, context support, and key\-point grounding\. Unsupported key points are removed or rewritten\. Human reviewers reject low\-quality samples and retain generated scenes only when all applicable fields pass review; experts correct annotation errors in human\-recorded interactions\. The main split contains 1 801 synthetic scenarios and 277 human\-recorded interactions performed by 30 actors\.

## 5Experiments

### 5\.1Experimental setup

We evaluate 14 native configurations in Table[3](https://arxiv.org/html/2609.21392#S5.T3), with 12 evaluated on audio–visual scenes and 14 on audio\-only scenes, using the same task instruction and modality\-specific output schema\. GPT\-Realtime\-2\([OpenAI, 2026a](https://arxiv.org/html/2609.21392#bib.bib22)\)andKimi\-Audio\([KimiTeam et al\., 2025](https://arxiv.org/html/2609.21392#bib.bib24)\)accept audio only; the remaining systems are evaluated on both audio\-only and audio–visual scenes\. We report the modalities separately\. For multi\-turn interactions, only the final user clip is scored \(§[3\.1](https://arxiv.org/html/2609.21392#S3.SS1)\)\. As a strong text\-only baseline, GPT\-5\.4\([OpenAI, 2026b](https://arxiv.org/html/2609.21392#bib.bib26)\)receives ASR transcripts, sentence timings, and prior\-turn assistant reply text \(App\.[D\.2](https://arxiv.org/html/2609.21392#A4.SS2)\)\. It remains unranked because its input interface differs from that of the native systems\. We choose Qwen3\.6\-Flash\([Qwen Team, 2026b](https://arxiv.org/html/2609.21392#bib.bib21)\)as the LLM judge, and analyze the stability of results across different LLM judges in App\.[E\.2](https://arxiv.org/html/2609.21392#A5.SS2)\.

Table 3:Evaluation results on the ODU task\.Audio–visual \(AV\) and audio\-only \(AO\) results are reported separately as percentages\. M1–M5 measure detection, key\-point coverage, localization, transcription, and user profile; Avg\. uses weights 0\.15/0\.60/0\.10/0\.10/0\.05 over M1–M5, and FTR is the no\-demand false\-trigger rate \(lower is better\)\. Size denotes the model parameter count; for MoE models, “A” indicates the number of active parameters\. The text\-only system receives ASR transcripts and sentence\-level timestamps but no media\.Audio–visual \(%\)Audio\-only \(%\)ModelSizeAvg\.M1DetectM2Keypt\.M3LocateM4Trans\.M5ProfileFTRNeg\.Avg\.M1DetectM2Keypt\.M3LocateM4Trans\.M5ProfileFTRNeg\.Closed\-source ModelsGemini 3\.1 Pro–72\.672\.683\.083\.063\.463\.481\.981\.991\.391\.395\.795\.740\.640\.675\.475\.487\.387\.370\.170\.161\.561\.591\.791\.797\.697\.631\.731\.7Gemini 3\.7 Flash–69\.669\.683\.983\.960\.260\.278\.778\.785\.985\.988\.388\.327\.527\.572\.972\.989\.689\.666\.166\.162\.762\.788\.488\.494\.194\.116\.016\.0Gemini 3\.5 Flash Lite–57\.457\.463\.963\.949\.949\.968\.168\.171\.571\.579\.379\.355\.655\.662\.262\.274\.574\.556\.456\.456\.756\.774\.374\.382\.482\.438\.538\.5Qwen3\.5\-Omni\-Plus–69\.669\.663\.963\.965\.865\.875\.775\.783\.383\.393\.593\.573\.873\.874\.774\.776\.276\.271\.571\.566\.166\.189\.389\.395\.895\.855\.355\.3Seed 2\.0 Lite–67\.967\.958\.158\.164\.164\.182\.882\.877\.777\.792\.592\.582\.282\.277\.377\.360\.860\.878\.078\.071\.471\.493\.293\.297\.597\.580\.780\.7GPT\-Realtime\-2––––––––63\.863\.871\.071\.063\.463\.434\.334\.381\.481\.471\.271\.263\.463\.4Open\-source ModelsQwen3\-Omni\-Think30B\-A3B59\.959\.958\.258\.255\.655\.647\.047\.085\.285\.292\.192\.182\.982\.967\.467\.473\.773\.763\.063\.050\.550\.588\.188\.194\.694\.660\.960\.9Qwen3\-Omni\-Instruct30B\-A3B55\.255\.250\.650\.653\.153\.123\.923\.987\.187\.192\.892\.893\.493\.461\.661\.661\.061\.060\.960\.920\.520\.591\.091\.095\.295\.280\.780\.7Qwen2\.5\-Omni7B45\.245\.253\.853\.842\.142\.15\.85\.874\.274\.278\.278\.287\.887\.848\.048\.067\.167\.145\.745\.75\.55\.569\.469\.460\.760\.766\.566\.5Ming\-Flash\-Omni 2\.0104B\-A6B52\.652\.650\.150\.149\.549\.529\.629\.683\.183\.182\.382\.393\.793\.752\.652\.663\.063\.050\.550\.57\.37\.386\.186\.169\.669\.675\.875\.8MiniCPM\-o 4\.59B39\.139\.154\.554\.535\.535\.51\.71\.760\.360\.368\.168\.185\.385\.349\.149\.165\.165\.147\.947\.96\.26\.261\.561\.575\.975\.974\.574\.5Nemotron 3 Nano Omni30B\-A3B34\.934\.955\.855\.826\.826\.835\.935\.934\.334\.369\.669\.672\.072\.040\.140\.159\.459\.433\.833\.839\.339\.337\.637\.664\.164\.172\.072\.0video\-SALMONN 2\+7B14\.214\.250\.350\.36\.76\.72\.92\.912\.312\.323\.423\.489\.189\.123\.823\.846\.446\.418\.418\.410\.010\.026\.326\.342\.642\.697\.597\.5Kimi\-Audio7B–––––––52\.652\.658\.458\.451\.151\.117\.217\.271\.871\.885\.085\.082\.082\.0Text\-only System \(no native media input\)GPT\-5\.4 \+ ASR transcript–59\.059\.060\.660\.655\.555\.582\.382\.381\.981\.92\.62\.673\.473\.471\.571\.578\.978\.968\.368\.386\.186\.185\.185\.130\.630\.640\.440\.4
Note\.Best and second\-best values areboldandunderlined\.

### 5\.2Overall results

ODU remains challenging even for the strongest models\.Table[3](https://arxiv.org/html/2609.21392#S5.T3)shows that the best native Avg\. scores reach only72\.672\.6% on AV scenes \(Gemini 3\.1 Pro\) and77\.377\.3% on AO scenes \(Seed 2\.0 Lite\)\. The highest AV M2 is just65\.865\.8%, leaving substantial headroom in the primary semantic dimension\. Open\-source models lag further: even their strongest model, Qwen3\-Omni\-Think, reaches only59\.959\.9% AV and67\.467\.4% AO Avg\.

Strong perception does not imply strong contextual reasoning\.On AV scenes, Gemini 3\.1 Pro achieves81\.981\.9% M3,91\.391\.3% M4, and95\.795\.7% M5, yet its M2 is only63\.463\.4% and its FTR reaches40\.640\.6%\. High localization, transcription, and user\-profile scores thus do not ensure correct interpretation of a user demand in context\. The key\-point and false\-trigger analyses \(§[5\.3](https://arxiv.org/html/2609.21392#S5.SS3), §[5\.4](https://arxiv.org/html/2609.21392#S5.SS4)\) identify missed contextual information and sensitivity to request\-like language and source/addressee mismatches\. The qualitative error analysis in App\.[E\.5](https://arxiv.org/html/2609.21392#A5.SS5)also visualizes these failures\.

Strong demand recovery does not ensure reliable engagement\.Seed 2\.0 Lite achieves78\.078\.0% M2 on audio\-only scenes but triggers on80\.780\.7% of no\-demand scenes\. Gemini 3\.7 Flash has lower M2 \(66\.166\.1%\) and much lower FTR \(16\.016\.0%\)\. This contrast shows that models can recover the content of valid demands yet still infer a demand when none is present\. Reliable demand understanding therefore requires judging whether the interaction calls for an assistant response, alongside recovering what the user wants\.

Models targeting efficient or real\-time interaction lag in demand understanding\.Gemini 3\.5 Flash Lite, GPT\-Realtime\-2, and MiniCPM\-o 4\.5 target efficient or real\-time interaction\. However, their AO Avg\. scores are only62\.262\.2%,63\.863\.8%, and49\.149\.1%, respectively, below both Gemini 3\.1 Pro \(75\.475\.4%\) and Qwen3\-Omni\-Think \(67\.467\.4%\)\. This gap raises a practical concern: interactive systems often rely on real\-time models, yet weaknesses in demand understanding can cause their responses to miss the user’s intent and degrade the interaction experience\.

Figure 4:Key\-point hit\-rate by evidence channelon AV demand scenes\. Rates are point\-weighted; hatched bars denote the GPT\-5\.4 text\-only baseline \(App\.[E\.1](https://arxiv.org/html/2609.21392#A5.SS1)\)\.![Refer to caption](https://arxiv.org/html/2609.21392v1/negative_subtypes_heatmap.png)Figure 5:False\-trigger patterns\.Cells show subtype FTR minus the same system’s overall FTR on labeled negatives, in percentage points\. The final row is the unweighted mean across systems\.

### 5\.3Contextual demand recovery

Models recover multimodal and conversational context less reliably than spoken requests\.M2, our primary semantic metric, measures coverage of key information about user intent and required context\. Using source labels assigned during annotation, we compare the fraction of key points recovered per source across three native systems and the text\-only baseline on audio–visual demand scenes \(Figure[5](https://arxiv.org/html/2609.21392#S5.F5)\)\.

Native spoken\-request hit\-rates are76\.076\.0–83\.683\.6%, comparable to the text\-only baseline’s78\.978\.9% and consistent with the strong transcription scores in Table[3](https://arxiv.org/html/2609.21392#S5.T3)\. Yet all three native systems score below 60% on visual and acoustic key points, highlighting weak recovery of multimodal context\. Dialogue\-history hit\-rates are also lower \(42\.142\.1–62\.362\.3% for native systems\)\. The text\-only baseline receives prior\-turn assistant reply text and reaches58\.758\.7%\. The text\-only baseline also recovers some visual and acoustic key points, possibly by leveraging commonsense knowledge together with scene details inferred from transcripts and dialogue history\.These results show that accurate speech transcription alone is insufficient for strong ODU performance\.

### 5\.4Error patterns across the taxonomy

We examine false triggers using FTR on negative scenes and demand understanding using M2 on positive scenes\. Figure[5](https://arxiv.org/html/2609.21392#S5.F5)shows the AV and AO negative\-scene results, with separate analyses of signal type and invalidation reason\. Positive cells mark categories with FTR above the same model’s overall rate on these labeled samples\.

Models falsely trigger on negative scenes with request\-like wording or source/addressee mismatches\.Linguistic signals such as questions, imperatives, and complaints yield the highest average FTR, exceeding every model’s mean with an average gap of11\.511\.5percentage points, whereas semantic signals that merely mention assistant functions fall below every model’s mean\. A clear, answerable request can still be invalid for the assistant: in Wrong Source it comes from media playback or background speakers; in Wrong Addressee it targets someone else\. These are the hardest invalidation categories on average; even Gemini 3\.7 Flash, with the lowest FTR in Table[3](https://arxiv.org/html/2609.21392#S5.T3), exceeds its own mean by26\.026\.0percentage points on Wrong Source\. Together, these patterns suggest over\-reliance on request\-like wording: source and addressee cues do not reliably override the apparent request, so speech outside the user–assistant interaction is promoted into a demand\.

Models struggle to recover indirect requests on positive scenes\.Indirect or descriptive requests score below every model’s mean M2, with average differences of−7\.2\-7\.2points on AV scenes and−8\.2\-8\.2points on AO scenes\. Appendix[E\.3](https://arxiv.org/html/2609.21392#A5.SS3)provides the full breakdown\. Models can mistake speech for a request to the assistant, yet struggle to understand genuine needs that users do not state directly\.

### 5\.5Comparison with human\-recorded interactions

Table 4:Model performance \(%\) on Chinese audio–visual scenes from synthetic \(Syn\.\) and human\-recorded \(Rec\.\) interactions\.Δ=\\Delta=\{\}Rec\.−\-Syn\.in percentage points\.MetricGemini 3\.1 ProSeed 2\.0 LiteQwen3\.5\-Omni\-PlusQwen3\-Omni\-ThinkMing\-Flash\-Omni 2\.0Syn\.Rec\.Δ\\DeltaSyn\.Rec\.Δ\\DeltaSyn\.Rec\.Δ\\DeltaSyn\.Rec\.Δ\\DeltaSyn\.Rec\.Δ\\DeltaAvg\.71\.971\.974\.274\.22\.467\.567\.567\.667\.60\.168\.568\.565\.765\.7\-2\.861\.361\.353\.653\.6\-7\.654\.254\.249\.349\.3\-4\.8M1 Detect83\.383\.386862\.758\.258\.255\.955\.9\-2\.260\.760\.753\.753\.7\-6\.960\.960\.948\.648\.6\-12\.349\.249\.243\.543\.5\-5\.6M2 Keypt\.626265\.165\.13\.263\.963\.963\.763\.7\-0\.264\.964\.964\.264\.2\-0\.756\.456\.449\.549\.5\-6\.950\.850\.847\.647\.6\-3\.2M3 Locate81\.881\.884\.984\.93\.182\.682\.682\.982\.90\.376\.376\.373\.173\.1\-3\.345\.945\.949\.849\.83\.9313137\.537\.56\.5M4 Trans\.92\.392\.388\.388\.3\-4\.176\.376\.379\.879\.83\.581\.981\.971\.671\.6\-10\.390\.790\.771\.371\.3\-19\.589\.589\.562\.362\.3\-27\.3M5 Profile95\.395\.398\.398\.33\.191\.591\.594\.994\.93\.493\.593\.593\.393\.3\-0\.292\.892\.890\.590\.5\-2\.284\.684\.685\.185\.10\.5FTR40\.940\.934\.734\.7\-6\.381\.981\.984842\.1797986\.786\.77\.680\.980\.9929211\.195\.295\.298\.798\.73\.4Leading MLLMs retain their performance on human\-recorded interactions\.Table[4](https://arxiv.org/html/2609.21392#S5.T4)shows comparable Avg\. scores for Gemini 3\.1 Pro and Seed 2\.0 Lite across synthetic and human\-recorded interactions\. Qwen3\-Omni\-Think and Ming\-Flash\-Omni 2\.0 are more sensitive to the source change, with Avg\. changes of−7\.6\-7\.6and−4\.8\-4\.8points, respectively\.

Human\-recorded speech exposes transcription weaknesses\.Qwen3\.5\-Omni\-Plus and Qwen3\-Omni\-Think show their largest declines in M4\. Ming\-Flash\-Omni 2\.0 has the steepest transcription decline in the table:−27\.3\-27\.3points in M4, versus−3\.2\-3\.2in M2\. These declines may stem in part from the more complex environmental noise and accented speech encountered in realistic interaction settings\.

### 5\.6Downstream response quality

We further examine whether correct demand understanding translates into better assistant responses\. Using Gemini 3\.1 Pro, we generated paired responses under two conditions:*without demand annotation*and*with demand annotation*\. Both conditions received identical audio–visual inputs and dialogue history under the same generation settings, while the latter additionally received the reference demand annotation, including an explicit no\-demand label when applicable\. The full experiment setup is detailed in App\.[E\.4](https://arxiv.org/html/2609.21392#A5.SS4)\.

We evaluate the paired responses through a blinded human A/B test\. For each interaction, responses A and B are presented in randomized order with condition identities hidden, and annotators select a preferred response or a tie based on correctness, relevance, contextual grounding, and appropriate silence\. Human evaluation yielded40\.040\.0% ties; among decisive judgments,85\.885\.8%favored the*with demand annotation*condition\. These results show thatproviding correct demand information substantially improves downstream response quality, supporting demand understanding as a key prerequisite for building effective conversational assistants\.

## 6Conclusion

We introduced Omni Demand Understanding \(ODU\) as a distinct task for identifying valid user demands and inferring intent from multimodal and conversational context\. ODU\-Bench combines taxonomy\-guided agentic generation with human\-recorded interactions and uses human\-verified annotations to evaluate demand understanding along five complementary dimensions\. Experiments with 14 native MLLMs show that strong perceptual performance does not ensure accurate contextual user\-intent inference\. In our evidence\-channel analysis, the three native systems evaluated on audio–visual demand scenes recover spoken requests more reliably than the visual, acoustic, and conversational information needed to interpret them\. Models also frequently mistake request\-like speech for a demand to the assistant, even when it comes from media playback or is addressed to someone else\. A blinded human A/B study further shows that providing reference demand annotations improves response quality, underscoring the practical value of correct demand understanding\. We hope ODU will make demand understanding a central focus of multimodal interaction research and drive progress toward assistants that infer user intent from multimodal and conversational context before deciding whether and how to respond\.

### AI use statement

We used generative AI tools to assist in designing the benchmark taxonomy\. Generative AI tools are integral components of our pipeline for data synthesis, annotation, and quality control \(§[4](https://arxiv.org/html/2609.21392#S4); App\.[C](https://arxiv.org/html/2609.21392#A3)\)\. We also use an LLM judge for key\-point\-based semantic evaluation \(§[3\.2](https://arxiv.org/html/2609.21392#S3.SS2); App\.[A](https://arxiv.org/html/2609.21392#A1)\)\. Additionally, we used AI writing assistants to improve the writing of this paper\. We did not use generative AI to develop theoretical models or conceptual frameworks; formulate mathematical claims; provide critical ingredients for proving mathematical claims; assist in writing mathematical proofs; propose or refine hypotheses; or design research methodology or experiments\. All benchmark media and annotations underwent rigorous human review, and human reviewers determined inclusion in the released benchmark\. We reviewed all AI\-assisted work and take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\.

### Ethics statement

Human\-recorded interactions were performed by paid actors\. Before release, all participating actors signed informed\-consent forms authorizing use of their recording data, including their faces and voices, for academic evaluation and public release with the benchmark\. To protect actor privacy, the released human\-recorded interactions may be used solely to evaluate systems on this benchmark; all other academic and commercial uses are prohibited\. Generated scenes contain no real individuals\. ODU\-Bench is provided solely for academic research and model evaluation, subject to the above restrictions on human\-recorded interactions\. Commercial use of the dataset or its images in products, services, or other profit\-making activities is prohibited\. Images in the dataset may depict identifiable individuals; copyright and related rights remain with their respective rights holders\. The release does not authorize sublicensing, commercial exploitation, or the creation of derivative works from these images\. Users must comply with the release terms, the scope of participant consent, and applicable personal information protection and portrait rights laws and regulations\.

### Reproducibility statement

The main paper defines the task, construction procedure, and evaluation framework \(§[3](https://arxiv.org/html/2609.21392#S3); §[4](https://arxiv.org/html/2609.21392#S4)\)\. The appendix specifies the prediction schema, matching, metrics, aggregation, and a worked scoring example \(App\.[A](https://arxiv.org/html/2609.21392#A1)\), the release composition and taxonomy \(App\.[B](https://arxiv.org/html/2609.21392#A2)\), and the construction and human\-verification procedures \(App\.[C](https://arxiv.org/html/2609.21392#A3)\)\. Experimental input preparation, the transcript\-only baseline, and the judge configuration are documented in App\.[D](https://arxiv.org/html/2609.21392#A4); supplementary analyses and judge stability are reported in App\.[E](https://arxiv.org/html/2609.21392#A5)\. Evaluation instructions are provided in App\.[F](https://arxiv.org/html/2609.21392#A6)\. Benchmark media, annotations, evaluation code, and configurations support reproducing the results shown in the paper\.

### Acknowledgments

This work was supported by Alibaba Innovative Research Program\. We would like to thank the Qwen Team at Alibaba Token Hub \(ATH\), Alibaba Group, for providing the computational resources and foundation models \(Qwen\) used in this research\.

## References

- I\. AI, B\. Ma, C\. Zou, C\. Du, C\. Yan, C\. Jin, C\. Shen, C\. Lian, C\. Fan, D\. Zheng,et al\.Ming\-Flash\-Omni: a sparse, unified architecture for multimodal perception and generation\.arXiv preprint arXiv:2510\.24821\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- Anet al\.\(2025\)K\. An, Y\. Chen, Z\. Chen, C\. Deng, Z\. Du, C\. Gao, Z\. Gao, B\. Gong, X\. Li, Y\. Li, Y\. Liu, X\. Lv, Y\. Ji, Y\. Jiang, B\. Ma, H\. Luo, C\. Ni, Z\. Pan, Y\. Peng, Z\. Peng, P\. Wang, H\. Wang, H\. Wang, W\. Wang, W\. Wang, Y\. Wu, B\. Tian, Z\. Tan, N\. Yang, B\. Yuan, J\. Ye, J\. Yu, Q\. Zhang, K\. Zou, H\. Zhao, S\. Zhao, J\. Zhou, and Y\. ZhuFun\-ASR technical report\.External Links:2509\.12508,[Link](https://arxiv.org/abs/2509.12508)Cited by:[§4\.3](https://arxiv.org/html/2609.21392#S4.SS3.p1.1)\.
- Cuiet al\.\(2026\)J\. Cui, B\. Xu, C\. Wang, T\. Yu, W\. Sun, Y\. Xu, T\. Wang, Z\. He, W\. Ma, T\. Cai,et al\.MiniCPM\-o 4\.5: towards real\-time full\-duplex omni\-modal interaction\.External Links:[Link](https://arxiv.org/abs/2604.27393)Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- Deshmukhet al\.\(2026\)A\. S\. Deshmukh, K\. Chumachenko, T\. Rintamaki, M\. Le, T\. Poon, D\. M\. Taheri, I\. Karmanov, G\. Liu, J\. Seppanen, A\. Goel,et al\.Nemotron 3 nano omni: efficient and open multimodal intelligence\.arXiv preprint arXiv:2604\.24954\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- Google DeepMind \(2026\)Google DeepMindGemini 3\.1 pro\.Note:[https://deepmind\.google/models/gemini/pro/](https://deepmind.google/models/gemini/pro/)Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1),[§4\.3](https://arxiv.org/html/2609.21392#S4.SS3.p1.1)\.
- Heet al\.\(2026\)Z\. He, W\. Cui, H\. Xu, X\. Li, L\. Zhu, H\. Bai, M\. Shaohua, and I\. KingMTR\-DuplexBench: towards a comprehensive evaluation of multi\-round conversations for full\-duplex speech language models\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 5334–5351\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1)\.
- Hurstet al\.\(2024\)A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.GPT\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- KimiTeamet al\.\(2025\)KimiTeam, D\. Ding, Z\. Ju, Y\. Leng, S\. Liu, T\. Liu, Z\. Shang, K\. Shen, W\. Song, X\. Tan, H\. Tang, Z\. Wang, C\. Wei, Y\. Xin, X\. Xu, J\. Yu, Y\. Zhang, X\. Zhou, Y\. Charles, J\. Chen, Y\. Chen, Y\. Du, W\. He, Z\. Hu, G\. Lai, Q\. Li, Y\. Liu, W\. Sun, J\. Wang, Y\. Wang, Y\. Wu, Y\. Wu, D\. Yang, H\. Yang, Y\. Yang, Z\. Yang, A\. Yin, R\. Yuan, Y\. Zhang, and Z\. ZhouKimi\-Audio technical report\.External Links:2504\.18425,[Link](https://arxiv.org/abs/2504.18425)Cited by:[§5\.1](https://arxiv.org/html/2609.21392#S5.SS1.p1.1.1)\.
- Kotturet al\.\(2021\)S\. Kottur, S\. Moon, A\. Geramifard, and B\. DamavandiSIMMC 2\.0: a task\-oriented dialog dataset for immersive multimodal conversations\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 4903–4912\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px2.p1.1)\.
- Linet al\.\(2026\)G\. Lin, S\. S\. Kuan, J\. Shi, K\. Chang, S\. Arora, S\. Watanabe, and H\. LeeFull\-Duplex\-Bench\-v2: a multi\-turn evaluation framework for duplex dialogue systems with an automated examiner\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 27–36\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-short.4),[Link](https://aclanthology.org/2026.acl-short.4/)Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.6.1)\.
- Linet al\.\(2025\)G\. Lin, J\. Lian, T\. Li, Q\. Wang, G\. Anumanchipalli, A\. H\. Liu, and H\. LeeFull\-duplex\-bench: a benchmark to evaluate full\-duplex spoken dialogue models on turn\-taking capabilities\.In2025 IEEE Automatic Speech Recognition and Understanding Workshop \(ASRU\),pp\. 1–8\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1)\.
- Luet al\.\(2026\)X\. Lu, X\. Li, A\. Wang, Y\. Bo, J\. Chen, Z\. Li, N\. Yang, R\. Liu, X\. Yang, J\. Hou, and L\. HongshengOmniInteract: benchmarking real\-world streaming interaction for real\-time omnimodal assistants\.arXiv preprint arXiv:2605\.26485\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1),[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.5.1)\.
- Mazumdaret al\.\(2026\)A\. Mazumdar, S\. Park, R\. Roy, N\. Srihari, S\. Wang, Y\. Zhou, J\. Wang, K\. Nagano, and S\. De MelloVideoFDB: evaluating full\-duplex vision\-speech capabilities in conversational agents\.arXiv preprint arXiv:2605\.30256\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1)\.
- Mullicket al\.\(2025\)A\. Mullick, S\. Sharma, A\. Jana, and P\. GoyalText takes over: a study of modality bias in multimodal intent detection\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 24039–24069\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px2.p1.1)\.
- OpenAI \(2026a\)OpenAIAdvancing voice intelligence with new models in the api\.Note:[https://openai\.com/index/advancing\-voice\-intelligence\-with\-new\-models\-in\-the\-api/](https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/)Cited by:[§5\.1](https://arxiv.org/html/2609.21392#S5.SS1.p1.1)\.
- OpenAI \(2026b\)OpenAIIntroducing GPT\-5\.4\.Note:[https://openai\.com/index/introducing\-gpt\-5\-4/](https://openai.com/index/introducing-gpt-5-4/)Cited by:[§5\.1](https://arxiv.org/html/2609.21392#S5.SS1.p1.1)\.
- Qwen Team \(2026a\)Qwen TeamQwen3\.5\-Omni technical report\.arXiv preprint arXiv:2604\.15804\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1),[§4\.3](https://arxiv.org/html/2609.21392#S4.SS3.p1.1)\.
- Qwen Team \(2026b\)Qwen TeamQwen3\.6\-35b\-a3b: agentic coding power, now open to all\.External Links:[Link](https://qwen.ai/blog?id=qwen3.6-35b-a3b)Cited by:[§5\.1](https://arxiv.org/html/2609.21392#S5.SS1.p1.1)\.
- Qwen Team \(2026c\)Qwen TeamQwen3\.7: the agent frontier\.External Links:[Link](https://qwen.ai/blog?id=qwen3.7)Cited by:[§4\.2](https://arxiv.org/html/2609.21392#S4.SS2.p1.1)\.
- Seed \(2026\)B\. SeedSeed2\.0 model card: towards intelligence frontier for real\-world complexity\.arXiv preprint arXiv:2607\.00248\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- Selvakumaret al\.\(2025\)R\. Selvakumar, A\. Seth, N\. Anand, U\. Tyagi, S\. Kumar, S\. Ghosh, and D\. ManochaMULTIVOX: a benchmark for evaluating voice assistants for multimodal interactions\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 28469–28481\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1),[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.4.1)\.
- Tanget al\.\(2025\)C\. Tang, Y\. Li, Y\. Yang, J\. Zhuang, G\. Sun, W\. Li, Z\. Ma, and C\. Zhangvideo\-SALMONN 2: Caption\-Enhanced Audio\-Visual Large Language Models\.arXiv preprint arXiv:2506\.15220\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- Veerabadranet al\.\(2025\)V\. Veerabadran, F\. Xiao, N\. Kamra, P\. Matias, J\. Chen, C\. Drooff, B\. Roads, R\. J\. Williams, E\. Henderson, X\. Zhao,et al\.Benchmarking egocentric multimodal goal inference for assistive wearable agents\.Advances in Neural Information Processing Systems38\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.9.1)\.
- Wanget al\.\(2025\)Y\. Wang, Y\. Wang, B\. Chen, T\. Wu, D\. Zhao, and Z\. ZhengOmniMMI: a comprehensive multi\-modal interaction benchmark in streaming video contexts\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 18925–18935\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1),[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.3.1)\.
- Wanget al\.\(2026\)Z\. Wang, C\. Liu, E\. Tjitrahardja, Y\. Wang, B\. Pavlov, F\. Gou, J\. M\. Davila, D\. Shi, R\. Xu, Y\. Pan,et al\.EgoIntrospect: an egocentric dataset and benchmark for user\-centric internal state reasoning\.arXiv preprint arXiv:2605\.17262\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.11.1)\.
- Xuet al\.\(2025a\)J\. Xu, Z\. Guo, J\. He, H\. Hu, T\. He, S\. Bai, K\. Chen, J\. Wang, Y\. Fan, K\. Dang, B\. Zhang, X\. Wang, Y\. Chu, and J\. LinQwen2\.5\-Omni technical report\.arXiv preprint arXiv:2503\.20215\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- Xuet al\.\(2025b\)J\. Xu, Z\. Guo, H\. Hu, Y\. Chu, X\. Wang, J\. He, Y\. Wang, X\. Shi, T\. He, X\. Zhu,et al\.Qwen3\-Omni technical report\.arXiv preprint arXiv:2509\.17765\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.
- Yanget al\.\(2026\)S\. Yang, J\. Yu, Y\. Peng, K\. Q\. Lin, J\. W\. Cho, Y\. Song, and J\. KimGUIDE: a benchmark for understanding and assisting users in open\-ended GUI tasks\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 13017–13027\.Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.10.1)\.
- Zhanget al\.\(2024\)H\. Zhang, X\. Wang, H\. Xu, Q\. Zhou, K\. Gao, J\. Su, J\. Zhao, W\. Li, and Y\. ChenMIntRec2\.0: A large\-scale benchmark dataset for multimodal intent recognition and out\-of\-scope detection in conversations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nY9nITZQjc)Cited by:[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.21392#S2.T1.2.1.8.1)\.
- Zhaoet al\.\(2026\)R\. Zhao, J\. Yang, Z\. Xin, T\. Wang, F\. Rao, J\. LYU, and X\. LiOmniPro: a comprehensive benchmark for omni\-proactive streaming video understanding\.arXiv preprint arXiv:2605\.18577\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1),[§2](https://arxiv.org/html/2609.21392#S2.SS0.SSS0.Px1.p1.1)\.
- Zipf \(1949\)G\. K\. ZipfHuman behavior and the principle of least effort: an introduction to human ecology\.Addison\-Wesley\.Cited by:[§1](https://arxiv.org/html/2609.21392#S1.p1.1)\.

## Appendix Contents

## Appendix ATask specification and evaluation

### A\.1Prediction schema

Only the final user clip is annotated, and timestamps are measured from its start\. A prediction contains ahas\_demandflag and, when it is true, a list of demand segments\. Each segment carries a time span, a transcript, a structured intent, a required\-context list, and closed\-set user profile fields\. The annotation reference also includes a frozen key\-point set and a necessity note for each required\-context item\. Fig\.[6](https://arxiv.org/html/2609.21392#A1.F6)gives the schema, and App\.[F\.1](https://arxiv.org/html/2609.21392#A6.SS1)gives the system prompts for evaluation\.

\{

"has\_demand":bool,

"segments":\[\{

"start\_ms":int,"end\_ms":int,//demandspan,milliseconds

"transcript":str,//thedemand\-bearingspeech

"structured\_intent":str,//whattheuserwantsdone

"required\_context":\[//whatthewordsleftout

\{"type":"visual"\|"audio"\|"dialogue\_history"\|"user\_activity"

\|"user\_identity"\|"user\_emotion"\|"scene\_condition",

"description":str\}

\],

"user\_in\_frame":bool,"user\_gender":str,"user\_age\_group":str

\}\]

\}

Figure 6:The prediction schema\. A model emitshas\_demandand, when it is true, one segment per continuous demand\-bearing region\. Audio\-only predictions omit visual context anduser\_in\_frame\.
### A\.2Segment matching

A scene may contain multiple demand segments, so predicted segments are first aligned with reference segments before computing M2–M5\. We prioritize normalized\-transcript similarity, which is more robust to boundary drift, and fall back to temporal IoU for remaining segments\. We additionally resolve common segmentation mismatches by allowing one prediction to cover multiple reference segments and by merging multiple predicted fragments that correspond to the same reference segment\. This prevents over\-splitting or over\-merging from being penalized as a semantic error\. Unmatched reference segments are retained and scored as missing in the corresponding metrics\. Detailed matching thresholds and merging rules are provided in the released evaluation code\.

### A\.3Metric definitions

Let𝒟\\mathcal\{D\}be the evaluation set,𝒟\+\\mathcal\{D\}^\{\+\}its demand\-bearing subset, andGxG\_\{x\}the reference segments of scenexx\. We distinguish scene scoressm​\(x\)s\_\{m\}\(x\)\(M2–M5\) from dataset scoresMmM\_\{m\}\(M1–M5\)\.

M1: Demand detection\.For each scenex∈𝒟x\\in\\mathcal\{D\}, letyxy\_\{x\}andy^x\\hat\{y\}\_\{x\}be the reference and predicted demand\-presence labels\. With one\-vs\-rest countsTPc\\mathrm\{TP\}\_\{c\},FPc\\mathrm\{FP\}\_\{c\}, andFNc\\mathrm\{FN\}\_\{c\}for each classc∈\{0,1\}c\\in\\\{0,1\\\},

M1=12∑c∈\{0,1\}2​T​Pc2​T​Pc\+FPc\+FNc,FTR=\|\{x∈𝒟:yx=0,y^x=1\}\|\|\{x∈𝒟:yx=0\}\|\.M\_\{1\}=\\frac\{1\}\{2\}\\sum\_\{c\\in\\\{0,1\\\}\}\\frac\{2\\mathrm\{TP\}\_\{c\}\}\{2\\mathrm\{TP\}\_\{c\}\+\\mathrm\{FP\}\_\{c\}\+\\mathrm\{FN\}\_\{c\}\},\\qquad\\mathrm\{FTR\}=\\frac\{\|\\\{x\\in\\mathcal\{D\}:\\,y\_\{x\}=0,\\ \\hat\{y\}\_\{x\}=1\\\}\|\}\{\|\\\{x\\in\\mathcal\{D\}:\\,y\_\{x\}=0\\\}\|\}\.\(1\)The false\-trigger rate is reported separately and does not enter Avg\.

M2: Key\-point coverage\.LetKgK\_\{g\}be the frozen key\-point set for reference segmentgg\. The LLM judge assigns a hit decisionhg​k∈\{0,1\}h\_\{gk\}\\in\\\{0,1\\\}to each pointkkusing the concatenated predicted intent and required\-context descriptions\. We count each distinct point once, regardless of tier labels, and average the per\-segment hit rates equally within each scene:

s2​\(x\)=1\|GxK\|​∑g∈GxK1\|Kg\|​∑k∈Kghg​k,GxK=\{g∈Gx:\|Kg\|\>0\},s\_\{2\}\(x\)=\\frac\{1\}\{\|G\_\{x\}^\{K\}\|\}\\sum\_\{g\\in G\_\{x\}^\{K\}\}\\frac\{1\}\{\|K\_\{g\}\|\}\\sum\_\{k\\in K\_\{g\}\}h\_\{gk\},\\qquad G\_\{x\}^\{K\}=\\\{g\\in G\_\{x\}:\|K\_\{g\}\|\>0\\\},\(2\)wherehg​k=0h\_\{gk\}=0for every point in an unmatched reference segment\. Thus a missed demand cannot improve coverage by removing its points from the denominator\.

M3: Segment localization\.Writem⁡\(g\)m\(g\)for the prediction matched to reference segmentggand set its IoU to zero when no match exists\. The per\-scene score averages IoU over all reference segments:

s3​\(x\)=1\|Gx\|​∑g∈GxIoU⁡\(m⁡\(g\),g\)\.s\_\{3\}\(x\)=\\frac\{1\}\{\|G\_\{x\}\|\}\\sum\_\{g\\in G\_\{x\}\}\\operatorname\{IoU\}\\\!\\left\(m\(g\),g\\right\)\.\(3\)Span precision, recall, F1, and the start\- and end\-time mean absolute errors are retained as diagnostics but do not enter Avg\.

M4: Transcript quality\.LetEgE\_\{g\}be the Levenshtein edit count andLgL\_\{g\}the reference length after case folding and whitespace and punctuation normalization\. Characters are used for a CJK\-dominant reference segment and words otherwise; an unmatched reference is compared with an empty hypothesis\. We pool edit counts and reference lengths across segments, so longer transcripts carry more weight within a scene:

s4​\(x\)=max⁡\(0,1−∑g∈GxEg∑g∈GxLg\)\.s\_\{4\}\(x\)=\\max\\\!\\left\(0,1\-\\frac\{\\sum\_\{g\\in G\_\{x\}\}E\_\{g\}\}\{\\sum\_\{g\\in G\_\{x\}\}L\_\{g\}\}\\right\)\.\(4\)
M5: User profile\.LetFxF\_\{x\}contain user\-in\-frame, gender, and age group for audio–visual scenes, and gender and age group for audio\-only scenes\. Writeug​fu\_\{gf\}andu^g​f\\hat\{u\}\_\{gf\}for the reference and predicted values of fieldfffor segmentgg\. The score averages accuracy over fields and reference segments, with unmatched segments scored as incorrect:

s5\(x\)=1\|Fx\|∑f∈Fx1\|Gx\|∑g∈Gx\[u^g​f=ug​f\]\.s\_\{5\}\(x\)=\\frac\{1\}\{\|F\_\{x\}\|\}\\sum\_\{f\\in F\_\{x\}\}\\frac\{1\}\{\|G\_\{x\}\|\}\\sum\_\{g\\in G\_\{x\}\}\\mathbf\{1\}\\\!\\left\[\\hat\{u\}\_\{gf\}=u\_\{gf\}\\right\]\.\(5\)

### A\.4Aggregation and key\-point accounting

Form∈\{2,3,4,5\}m\\in\\\{2,3,4,5\\\}, let𝒟m\+⊆𝒟\+\\mathcal\{D\}\_\{m\}^\{\+\}\\subseteq\\mathcal\{D\}^\{\+\}contain the positive scenes on which the dimension score is defined\. Each such scene receives equal weight in the reported score,Mm=\|𝒟m\+\|−1​∑x∈𝒟m\+sm​\(x\)M\_\{m\}=\|\\mathcal\{D\}\_\{m\}^\{\+\}\|^\{\-1\}\\sum\_\{x\\in\\mathcal\{D\}\_\{m\}^\{\+\}\}s\_\{m\}\(x\)\. M1 is defined by Eq\.[1](https://arxiv.org/html/2609.21392#A1.E1)over both positive and no\-demand scenes\. LetAAbe the set of available dimensions for a model\. Avg\. is

Avg=∑m∈Awm​Mm∑m∈Awm,\(w1,w2,w3,w4,w5\)=\(0\.15,0\.60,0\.10,0\.10,0\.05\)\.\\mathrm\{Avg\}=\\frac\{\\sum\_\{m\\in A\}w\_\{m\}M\_\{m\}\}\{\\sum\_\{m\\in A\}w\_\{m\}\},\\qquad\(w\_\{1\},w\_\{2\},w\_\{3\},w\_\{4\},w\_\{5\}\)=\(0\.15,0\.60,0\.10,0\.10,0\.05\)\.\(6\)For no\-demand scenes, only M1 and the separately reported false\-trigger rate are defined\.

Tier labels indicate whether a point belongs to demand semantics \(T1\), contextual grounding \(T2\), or both\. M2 counts their union once\. Separately, the source label assigns each attributed point to an evidence channel: the spoken request, visual evidence, audio evidence, dialogue history, or other context\. Points with no assigned channel remain in the overall reference denominator but are excluded from channel\-specific rates\. App\.[E\.1](https://arxiv.org/html/2609.21392#A5.SS1)gives the channel analysis and its population accounting\.

### A\.5Model output and scoring walkthrough

The following released example shows the reference annotation, a stored model prediction, the judge input, and the resulting scene\-level scores\. The judge configuration is specified in App\.[D\.3](https://arxiv.org/html/2609.21392#A4.SS3); the final benchmark aggregation follows Eq\.[6](https://arxiv.org/html/2609.21392#A1.E6)\.

#### Reference annotation\.

Sceneavp\_gen\_000139contains one reference segment at 3510–5390 ms\. The user says*“Hey, agent, mute that sound\.”*The structured intent is*“The user wants to mute the electronic notification sound currently playing”*\. Its required context is:

audioThe device in the close foreground emits a short notification chime

visualAt the bottom edge of the frame, a hand holds a glowing smartphone in landscape orientation, which is the source device emitting the notification chime

The reference user profile isuser\_in\_frame=true, user\_gender=male, user\_age\_group=adult\.

#### Model prediction and judge input\.

Seed 2\.0 Lite predicts a demand at 3600–5500 ms, with transcript*“Hey agent, mute that sound\.”*The scorer concatenates its predicted intent and required\-context descriptions into the following candidate text:*“The user wants the AI agent to mute the ongoing notification sound from his mobile phone\. \| audio: A repeated beeping notification sound from the user’s phone is playing\.”*

The judge receives this candidate text together with the reference key points listed below\. The reference transcript and context\-necessity notes are not appended to the candidate text\.

The predicted user profile isuser\_in\_frame=false, user\_gender=male, user\_age\_group=adult\.

#### Stored key\-point verdicts\.

Evidence channelVerdictKey pointkp01intenthitThe user wants to mute the electronic notification sound currently playingkp02required\_context\.audiohitThe sound to be muted is the notification chime from the nearby devicekp03required\_context\.visualmissThe sound source is the smartphone held by a hand at the bottom of the frame
The stored judge marks 2 of the 3 points as hits, giving this scene an M2 score of2/3=0\.672/3=0\.67\. Although the prediction names a mobile phone, the stored judge marks kp03 as a miss because it omits the hand\-held device and its position at the bottom of the frame\. These verdicts concern coverage of the reference points in the predicted text; they do not directly measure perception\.

#### Scene\-level scores\.

The temporal IoU is 0\.90, with start\- and end\-time errors of 90 ms and 110 ms\. The normalized transcripts match exactly\. Gender and age group are correct, whileuser\_in\_frameis incorrect, yielding two correct profile fields out of three\. The scene\-level scores are:

DimensionScoreM2Key\-point Coverage0\.67M3Segment Localization0\.90M4Transcript Quality1\.00M5User profile0\.67Demand detection is correct for this scene\. The reported M1, however, is macro\-F1 computed across demand and no\-demand scenes\. Consequently, the benchmark Avg\. is computed from the dataset\-level dimension scores in Eq\.[6](https://arxiv.org/html/2609.21392#A1.E6), rather than by treating this scene’s detection correctness as M1\.

## Appendix BBenchmark composition and taxonomy

### B\.1Dataset statistics

Table[5](https://arxiv.org/html/2609.21392#A2.T5)gives the composition of the released dataset\. Its per\-axis distribution is shown in Fig\.[2](https://arxiv.org/html/2609.21392#S4.F2)\.

Table 5:Composition of the released dataset\.PropertyMain splitTotal scenes2 078Audio–visual / audio\-only1 347 / 731Chinese / English1 170 / 908Positive / no\-demand negative1 631 / 447Generated / human\-recorded interactions1 801 / 277Multi\-turn interactions969
### B\.2Positive\-demand scenes: six axes

Fig\.[2](https://arxiv.org/html/2609.21392#S4.F2)summarizes the taxonomy and its distribution in the released dataset\. Axes 1–4 describe construction challenges; Axes 5–6 describe task and domain coverage\. Categories are not an ordered difficulty scale\. The Axis 1 definitions below use the audio–visual vocabulary\. Audio\-only scenes instead use 1\.1 for language\-channel demands and 1\.2 for acoustic\-event dependence\. Their Axis 5\.2 denotes acoustic perception and recognition; the remaining demand types retain their meanings\.

Axis 1: Modality DependencyScope\.This axis describes which input modalities are necessary to infer the user’s intent\.1\.1 Language\-channel\.The user’s intent requires no visual or nonverbal audio grounding\. It may still depend on textual dialogue history, which is classified independently on Axis 2\. Typical evidence is a direct request with verbally stated arguments or a textual back\-reference to an option named earlier\.1\.2 Vision\-grounded\.The user’s intent depends on visible objects, people, spatial relations, written text, gestures, or visual events in the current scene\. Typical evidence includes deictic expressions such as “this one”, “that person”, or “the item on the left”\.1\.3 Audio\-grounded\.The user’s intent depends on nonverbal sounds, environmental audio events, speaker identity, voice characteristics, or acoustic cues\. Typical evidence includes asking about a sound, reacting to an alarm, or referring to someone by voice\.1\.4 Vision–Audio Joint\.The user’s intent requires both visual and audio evidence; either modality alone would leave the intent incomplete or ambiguous\. For example, “Which of these instruments makes that sound?” requires identifying the instrument from the audio and locating it among the visible candidates\.

Axis 2: Contextual DisambiguationScope\.This axis describes what kind of context is needed to resolve the user’s intent beyond the literal wording of the current utterance\.2\.1 Self\-contained Current Context\.The current utterance and scene state determine the intent without recalling earlier events or tracking changes over time\.2\.2 Temporal Event Dependency\.The intent requires recalling a perceptual event or tracking its timing, sequence, or resulting state change, before or during the current turn\.2\.3 Dialogue History Dependency\.The intent depends on prior conversational context, especially previous user–assistant exchanges\.2\.4 Speaker Identity Dependency\.The intent depends on identifying which person or voice is being referred to, or on distinguishing among multiple speakers\.2\.5 Affective State Dependency\.The intent depends on recognizing emotion, attitude, stress, fatigue, urgency, hesitation, or another affective state\.Boundary note\.Axis 2\.2 concerns perceptual event history, whereas Axis 2\.3 concerns discourse history\. A scene may contain both, but the label reflects the primary source of disambiguation\.

Axis 3: Intent Expression FormScope\.This axis describes how the user’s intent is expressed linguistically\.3\.1 Direct Instruction\.The user states an explicit command or request for the assistant to perform an action\.3\.2 Direct Question\.The user expresses the intent as an explicit question\.3\.3 Indirect or Descriptive Request\.The user describes a situation, problem, preference, or state from which the intended assistance must be inferred\.3\.4 Hesitation or Self\-correction\.The user’s expression includes hesitation, revision, restart, or correction that must be resolved\.3\.5 Interrupted or Incomplete Request\.The utterance is cut off, interrupted, or incomplete, but context still determines a valid demand\. R3 \(App\.[B\.3](https://arxiv.org/html/2609.21392#A2.SS3)\) applies when no actionable demand remains\.3\.6 Compound or Multi\-intent Request\.The user expresses multiple related or parallel intents in one turn\.

Axis 4: Acoustic EnvironmentScope\.This axis describes the acoustic conditions under which the user’s demand is expressed\. It captures speech intelligibility and audio interference, not whether audio is part of the task content\.4\.1 Clean Near\-field Speech\.The user’s speech is near\-field, clear, and minimally affected by interference\.4\.2 Background Noise\.Non\-speech environmental noise is present but does not fully obscure the user’s speech\.4\.3 Competing Speech\.Other human speech overlaps with or competes against the user’s speech\.4\.4 Low\-volume or Whispered Speech\.The user intentionally speaks softly, whispers, or has significantly reduced volume\.4\.5 Media Playback Interference\.Media audio, such as television, music, phone playback, navigation audio, or public announcements, interferes with the user demand or is semantically relevant to it\.Boundary note\.Axis 1 asks whether audio is part of the intent content; Axis 4 asks whether the acoustic environment affects how the intent is expressed or perceived\.

Axis 5: Demand TypeScope\.This axis describes the functional type of assistance requested by the user\.5\.1 Information Seeking and Retrieval\.The user asks the assistant to retrieve, look up, or provide factual information\.5\.2 Visual Perception and Recognition\.The user asks the assistant to perceive, identify, describe, compare, or reason about visible content\.5\.3 Action Execution\.The user asks the assistant to perform an operation or control an external function\.5\.4 Guidance and Explanation\.The user asks for step\-by\-step guidance, procedural help, troubleshooting, or explanation\.5\.5 Recommendation and Decision Support\.The user asks the assistant to compare options, make a recommendation, or support a choice\.5\.6 Monitoring and Reminder\.The user asks the assistant to monitor a condition, detect a change, or remind them later\.5\.7 Recording and Summarization\.The user asks the assistant to record, transcribe, summarize, or preserve information\.5\.8 Social and Emotional Interaction\.The user seeks social support, emotional response, interpersonal advice, or empathetic interaction\.5\.9 Translation and Language Conversion\.The user asks the assistant to translate, paraphrase, rewrite, or convert language\.5\.10 Content Creation\.The user asks the assistant to generate new content\.

Axis 6: Scenario DomainScope\.This axis describes the everyday domain or activity setting in which the user demand occurs\.6\.1 Home and Daily Life\.Domestic routines and everyday household activities\.6\.2 Cooking and Dining\.Food preparation, eating, kitchen use, restaurants, and meal\-related decisions\.6\.3 Shopping and Commerce\.Purchasing, product comparison, retail settings, payment, and consumer decisions\.6\.4 Travel and Navigation\.Trip planning, wayfinding, route selection, transportation logistics, and location search\.6\.5 Tourism and Exploration\.Sightseeing, visiting attractions, cultural exploration, and travel discovery\.6\.6 Learning and Education\.Studying, teaching, tutoring, training, and knowledge acquisition\.6\.7 Office and Collaboration\.Workplace tasks, meetings, documents, coordination, and professional communication\.6\.8 Sports and Fitness\.Exercise, training, physical activity, performance tracking, and fitness guidance\.6\.9 Health and Medical Care\.Health monitoring, medication, symptoms, care settings, and medical assistance\.6\.10 DIY and Repair\.Assembly, maintenance, repair, tool use, and practical troubleshooting\.6\.11 Driving and In\-vehicle Contexts\.Driving, vehicle use, road context, navigation, and in\-car assistance\.6\.12 Creative Work and Content Production\.Creative planning, media production, design, writing, and editing\.6\.13 Social Gatherings\.Social interaction in group settings, parties, visits, and interpersonal coordination\.6\.14 Outdoor and Nature\.Outdoor environments, nature activities, weather, and field observations\.6\.15 Childcare and Caregiving\.Caring for children, elderly people, patients, or others who need assistance\.6\.16 Fashion and Beauty\.Clothing, appearance, grooming, cosmetics, and style decisions\.6\.17 Public Transit and Commuting\.Public transportation, commuting routines, stations, transfers, and announcements\.6\.18 Entertainment and Leisure\.Games, media, hobbies, recreation, and leisure activities\.6\.19 Safety and Emergency\.Risk, danger, urgent response, accident prevention, and emergency assistance\.

### B\.3No\-demand scenes: signal type and invalidation reason

No\-demand scenes contain cues that resemble a request, but no valid demand is directed to the conversational assistant\. The two dimensions specify the demand\-like signal that creates false\-trigger risk and the reason it is invalid\.

Dimension N1: Demand\-like Signal TypeScope\.This dimension identifies the observable cue that may be mistaken for a user demand\.S1 Linguistic Signal\.The utterance surface form resembles a request, question, imperative, complaint, or help\-seeking expression\.S2 Semantic Signal\.The utterance mentions a function, task, or topic the assistant could help with, even if the user is not actually requesting assistance\.S3 Prosodic Signal\.The speech sounds as if it may be directed to the assistant because of voice quality, direction, distance, or command\-like prosody\.S4 Interactive Signal\.The dialogue structure resembles a user–assistant request–response turn, even though no new demand is being made\.S5 Contextual Signal\.The scene presents a plausible opportunity for the assistant to help, but the user is not actually requesting assistance\.

Dimension N2: Invalidation ReasonScope\.This dimension identifies why a demand\-like signal is invalid\.R1 Wrong Addressee\.The utterance comes from the user, but is addressed to another person, a pet, or an object, or has no intended recipient; it is not addressed to the assistant\.R2 Wrong Source\.The apparent request is produced by media playback, another speaker, an echo, or another background source, rather than by the user\.R3 Incomplete or Withdrawn Demand\.A possible demand is started but not completed, is withdrawn or abandoned, or is resolved by the user before it becomes actionable, leaving no actionable demand to the assistant\.R4 Non\-real Intent\.The utterance does not express the user’s current real intent; it is hypothetical, joking, exaggerated, quoted, read aloud, rehearsed, remembered, or discussed as someone else’s words\.R5 Function Mention Not Directed as a Request\.The utterance discusses, evaluates, teaches, recalls, or gives feedback about an assistant capability, but does not request its use\.

A negative remains demand\-like in its surface cues; invalidity depends on the addressee, source, dialogue function, pragmatic status, or scene context\. In Fig\.[2](https://arxiv.org/html/2609.21392#S4.F2), Wrong Addressee and Wrong Source retain their full names; Incomplete, Non\-real, and Function Mention abbreviate Incomplete or Withdrawn Demand, Non\-real Intent, and Function Mention Not Directed as a Request, respectively\.

## Appendix CConstruction and human verification

### C\.1Target selection and seed sampling

#### Scene seeds and positive targets\.

Each corpus seed provides an everyday scene category and topic, such as “Office & Workplace / office plant care techniques”\. The generator uses this pair as a starting setting and subject, and combines it with a challenge target from Axes 1–4\. For Axis 5, we supply candidate demand types; the generator chooses one that fits the developing scene when constructing the user’s intent\. Axis 6 is assigned from the completed scenario’s domain\. Thus, the seed category guides scene creation while Axis 6 describes the resulting scene, and the seed topic does not fix an Axis 5 demand type\.

#### Coverage and seed sampling\.

For positive targets, the selector prioritizes uncovered admissible Axes 1–4 combinations, then those with the lowest counts, breaking ties randomly\. The Axis 5 candidates include a least\-used demand type and alternatives weighted toward less\-used types\. Seeds are sampled independently of these targets, uniformly among the least\-used scene–topic pairs in the selected language\. Language selection follows the 1:1 Chinese:English ratio\.

#### Negative targets\.

No\-demand scenes use the same scene–topic seeds with a signal–invalidation target \(App\.[B\.3](https://arxiv.org/html/2609.21392#A2.SS3)\), without an Axis 5 demand type\. The sampler first draws an invalidation reason uniformly, then a pattern uniformly within that reason; the pattern supplies its demand\-like signal type\. Acoustic conditions are varied separately, and patterns that depend on a previous assistant response request dialogue history\.

### C\.2Script generation and review

#### From seed to interaction\.

For positives, a brainstorm combines the seed and taxonomy target into a scene concept, the concrete challenge, and an explanation of why it requires context\. The cascade then specifies the physical setting, people, objects, and sound sources\. Its usual path fixes the intended demand and decisive evidence, builds any required event or dialogue history around them, and finally writes the user’s utterance\. Patterns that require joint intent–utterance construction instead build history first and generate the two together\. For negatives, a separate cascade builds a scene skeleton and any required history, then instantiates the demand\-like cue and the evidence that invalidates it\. In a Wrong Source scene, for example, the apparent request must remain attributed to the playback or other speaker that produced it\.

#### Consistency and validity\.

Programmatic checks verify speaker roles, required history or event traces, and explicit pattern constraints\. An LLM reviewer checks plausibility, causal coherence, taxonomy alignment, and support for the intended interpretation in the script\. Negative validation asks whether an unmet need is present and whether it is directed to the assistant, while checking that the intended misleading signal is actually present\.

#### Utterance\-only screening\.

For positives, the text\-only discriminator uses a single LLM call with two instructed steps: infer intent from the scripted user utterance, then compare it with the supplied intended demand\. It returns amatch,partial, ormismatchverdict, together with the inferred intent, confidence, and reasoning\. Dialogue history, speaker metadata, and media evidence are withheld\. A full match rejects a candidate intended to require contextual evidence, including a language\-channel candidate whose difficulty should come from dialogue history\. For negatives with user speech, the reverse screen uses the intent\-prediction step and rejects candidates whose wording alone is judged to contain no demand; cases without a user utterance skip this check\. These are script\-stage screens, not tests of exclusive evidence dependence in the realized media\.

### C\.3Media production and annotation

#### Rendering and performance guides\.

Accepted timelines guide synthetic media production for both modalities\. Multi\-clip scenes are rendered sequentially with earlier clips as video references for consistent people, voices, and surroundings\. Audio\-only production retains the audio track\. Performance guides for human\-recorded interactions specify roles, dialogue, event order, and essential cues, allowing practical setting or prop substitutions that preserve the challenge\. Single\-turn wording may vary while retaining the intended ambiguity; multi\-turn guides preserve user lines that match prepared assistant replies\. Both routes enter media\-grounded annotation to account for departures from the script\.

#### Media\-derived draft and script reconciliation\.

An initial annotation uses multimodal media captions and word\-level ASR, without the script’s intended annotation\. Captions supply scene and speaker semantics; ASR supplies spoken wording and timing\. The next stage checks scripted speech against ASR and compares the draft with the scenario specification\. Significant disagreements become observable questions for targeted MLLM inspection of the media\. The draft, actual script, ASR, and resulting observations are then reconciled\. In multi\-clip scenes, earlier user clips and assistant replies provide history; annotation and timing target the final user clip\.

#### Timing, relevance, and verification\.

Segment transcripts are matched to ASR words, and their start and end times are aligned to word boundaries\. Rule\-based and LLM checks remove context that does not materially affect the appropriate response, while retaining referents needed to interpret the demand\. Schema and cross\-field checks, together with an LLM review, assess demand presence, temporal grounding, intent, context quality, and transcript coverage before key points are constructed\.

#### Key\-point construction and grounding\.

Verified intent and context are decomposed into atomic, response\-relevant key points, with the transcript helping identify their evidence source\. Each point must be expressible from the structured intent or an existing required\-context description\. Points representing the requested task, required context, or both receive the corresponding labels described in App\.[A](https://arxiv.org/html/2609.21392#A1)\. An MLLM then checks the actual media against the points, context, and intent\. Unsupported details are removed or narrowed to supported content\. If this changes the intent, key points are regenerated and checked against the media again\.

### C\.4Human verification

#### Media quality\.

Human reviewers are asked to assess perceptual realism and playback usability of the generated media\. Generated clips were rejected for conspicuous synthesis artifacts, such as mismatched lip movements, sudden object appearances or disappearances, and distorted faces or physically implausible motion\. Across both media sources, rejection criteria also included corrupted video or audio, missing sound, and severe playback stalls\. Unintended assistant\-response audio and crosstalk that made speakers indistinguishable were also grounds for rejection\.

#### Annotation correctness\.

Reviewers checked each applicable machine\-produced annotation field against the realized interaction: demand presence, demand spans, transcript, structured intent, required context, user profile, and key points\. They verified speech and timing against the media and checked semantic annotations for unsupported details, irrelevant context, and missing key points, using dialogue history where needed\. Generated scenes were retained only when all applicable fields passed review\. For human\-recorded interactions, experts corrected erroneous fields before finalization to accommodate departures from the performance guide\.

Figure[7](https://arxiv.org/html/2609.21392#A3.F7)illustrates the review controls using a generated benchmark example in place of recorded media\. Two consecutive views show the complete annotation form in its original field order, with the reference annotation expanded\. The example uses the released English annotation; selected controls show interface defaults rather than a saved human verdict\.

![Refer to caption](https://arxiv.org/html/2609.21392v1/x1.png)Figure 7:Annotation\-review interface, shown in two consecutive views\. \(a\) Generated media, reference annotation, sample acceptance, and transcript\. The original field order is retained; selected options are interface defaults\.![Refer to caption](https://arxiv.org/html/2609.21392v1/x2.png)Figure 8:Annotation\-review interface \(continued\)\. \(b\) Demand time spans, required context, all key points, and user profile fields from the same generated sample\.

## Appendix DExperimental setup

### D\.1Input preparation and evaluation

User\-turn media are presented in chronological order, with preceding assistant replies supplied as text between turns\. Models predict the demand in the final user clip; preceding clips and assistant replies provide interaction history\. Predicted start and end times are local to the final clip\. The model\-facing instruction is specialized to audio–visual or audio\-only input, and predictions use the same demand\-decision schema and scoring procedure\. The full prediction instruction is provided in App\.[F\.1](https://arxiv.org/html/2609.21392#A6.SS1)\.

### D\.2Text\-only baseline

The text\-only baseline uses GPT\-5\.4 with the same modality\-specific task instruction \(App\.[F\.1](https://arxiv.org/html/2609.21392#A6.SS1)\) as the native systems\. Each media block is replaced in place by its ASR transcript, sentence\-level timestamps, and optional speaker identifiers\. Multi\-turn inputs retain the original clip order, turn labels, and preceding assistant reply text\. Each replacement block explicitly states that native media are unavailable\. The baseline is evaluated with the same parser, segment matching, and metrics as the native systems\.

System message:The full prediction instruction in App\.[F\.1](https://arxiv.org/html/2609.21392#A6.SS1), including the demand annotation schema and field constraints; the audio\-only variant follows App\.[F\.2](https://arxiv.org/html/2609.21392#A6.SS2)\.

User message:The panels below show single\-turn and multi\-turn input templates, followed by input examples for a positive and a no\-demand scene\. Demand labels reflect the full scenes and may not be recoverable from the ASR transcripts alone\. The templates retain the original audio–visual task wording, including “video”, while replacing media with ASR text\. For audio\-only input, “video” and “audio or video” are replaced with “audio”\.

Braces mark placeholders\. Theduration\_noteadds “ \(clip duration \{duration\_ms\} ms\)”, andspeaker\_iddenotes the speaker identifier returned by ASR\.

The multi\-turn template shows the complete ASR block for each clip, including its no\-media notice\. ASR blocks and assistant replies appear in chronological order; longer histories repeat the same pattern\. If no speech is recognized, the no\-speech notice replaces that clip’s ASR block\. The closing instruction follows the final block in either case\.

User\-input templatesSingle clip\[No audio or video is provided for this clip\. Below is the verbatim ASR transcript of its speech, with millisecond timestamps local to this clip\{duration\_note\}\.\]\[\{begin\_ms\}\-\{end\_ms\} ms\] speaker \{speaker\_id\}: \{ASR sentence\}…Annotate the user’s demand in this video and output the demand\-decision JSON\.Replacement ASR block when no speech is recognized\[No audio or video is provided for this clip, and automatic speech recognition returned no speech for it\.\]

User\-input template with dialogue historyYou will receive multiple video clips from the same scene in chronological order\. Together they form the dialogue history; the last video clip is the current user turn to annotate:\[Turn 1 · user video\]\[No audio or video is provided for this clip\. Below is the verbatim ASR transcript of its speech, with millisecond timestamps local to this clip\{duration\_note\}\.\]\[\{begin\_ms\}\-\{end\_ms\} ms\] speaker \{speaker\_id\}: \{ASR sentence\}…\[Turn 1 · agent reply\] \{preceding assistant reply\}\[Turn 2 · user video \(final turn, annotate this\)\]\[No audio or video is provided for this clip\. Below is the verbatim ASR transcript of its speech, with millisecond timestamps local to this clip\{duration\_note\}\.\]\[\{begin\_ms\}\-\{end\_ms\} ms\] speaker \{speaker\_id\}: \{ASR sentence\}…Annotate only the user’s demand in the final video clip and output the demand\-decision JSON\. Use the preceding video clips and agent replies only as dialogue\-history context; do not output segments for them\. start\_ms/end\_ms must be local timestamps within the final video clip\.

Positive example \(Demand present\)\[No audio or video is provided for this clip\. Below is the verbatim ASR transcript of its speech, with millisecond timestamps local to this clip \(clip duration 9055 ms\)\.\]\[5040\-8640 ms\] speaker 0: Set it up, so when he gives the verbal, okay, export this version\.Annotate the user’s demand in this video and output the demand\-decision JSON\.

Negative example \(No demand\)\[No audio or video is provided for this clip\. Below is the verbatim ASR transcript of its speech, with millisecond timestamps local to this clip \(clip duration 12051 ms\)\.\]\[4590\-10710 ms\] speaker 0: my data was leaked, can this it insurance immediately activate an emergency response for me?\[11070\-11990 ms\] speaker 0: how do i file a claim?Annotate the user’s demand in this video and output the demand\-decision JSON\.

### D\.3Semantic scoring

We use Qwen3\.6\-Flash with thinking enabled as the semantic judge in the main experiments\. The judge receives the predicted structured intent concatenated with the predicted required\-context descriptions, together with the reference key points\. Reference transcripts and context\-necessity notes are not added to the candidate answer text\. App\.[F\.4](https://arxiv.org/html/2609.21392#A6.SS4)provides the judge instruction, and App\.[E\.2](https://arxiv.org/html/2609.21392#A5.SS2)analyzes result stability across judges\.

Each distinct reference point is counted once, regardless of tier labels\. Key points in unmatched reference segments count as missed\. Segment matching and score aggregation follow App\.[A](https://arxiv.org/html/2609.21392#A1)\.

## Appendix ESupplementary experiments and diagnostics

This appendix provides supplementary analyses of request and context coverage, LLM judge stability, and positive\-demand challenges\. It then presents the blinded comparison of responses generated with and without demand annotations and closes with qualitative error cases\.

### E\.1Request and context key\-point coverage

We compare recovery of spoken requests and the context needed to interpret them, extending the analysis in §[5\.3](https://arxiv.org/html/2609.21392#S5.SS3)and Figure[5](https://arxiv.org/html/2609.21392#S5.F5)\.

#### Population and counting\.

We analyze 1 060 audio–visual demand scenes scored by all four systems, containing 3 969 distinct reference key points\. Each point counts once; points in unmatched reference segments count as misses\. Context comprises visual, audio, dialogue\-history, and other context points \(emotion, activity, identity, and scene condition\)\. All context pools these categories\.

#### Coverage gaps and sampling uncertainty\.

Table[6](https://arxiv.org/html/2609.21392#A5.T6)uses the same population and point weighting as Figure[5](https://arxiv.org/html/2609.21392#S5.F5)\. Coverage is the ratio of recovered to reference key points in each channel\. The gap subtracts context coverage from spoken\-request coverage, in percentage points\. Confidence intervals assess sensitivity to the composition of the sampled scenes\.

#### Scene\-bootstrap procedure\.

For each of 10 000 resamples, we draw 1 060 scenes uniformly with replacement from the common population\. All systems and channels use the same draw, and each selected scene contributes all its key points, including repeated contributions when selected more than once\. Model outputs and judge verdicts remain fixed\.

Within each resample, we recompute each channel’s coverage from its pooled hit and reference\-point counts, then calculate the request\-minus\-context gaps\. Their 2\.5th and 97\.5th percentiles form the 95% confidence interval\. The Gap column uses the original full sample; resampling supplies the interval while preserving dependence among each scene’s key points\.

Table 6:Point\-weighted coverage by evidence channel \(Figure[5](https://arxiv.org/html/2609.21392#S5.F5)\)\. Gap subtracts each context row’s coverage from the same model’s Spoken request coverage; positive values mean lower context coverage\. The 95% confidence interval quantifies the gap’s uncertainty under scene resampling\. Both are in percentage points \(pp\)\.nncounts reference points; dashes mark the request reference\. Gaps use unrounded rates\.SystemChannelReferencepoints \(nn\)Coverage\(%\)Gap \(pp\)Request−\-context95% confidenceinterval for gap \(pp\)Gemini 3\.1 ProSpoken request1 98482\.982\.9––All context1 97244\.744\.738\.238\.2\[35\.435\.4,41\.141\.1\]Visual80448\.648\.634\.234\.2\[30\.330\.3,38\.338\.3\]Audio55747\.247\.235\.635\.6\[31\.231\.2,40\.240\.2\]History47542\.142\.140\.840\.8\[35\.935\.9,45\.745\.7\]Qwen3\.5\-Omni\-PlusSpoken request1 98483\.683\.6––All context1 97248\.948\.934\.634\.6\[32\.032\.0,37\.337\.3\]Visual80447\.347\.336\.336\.3\[32\.632\.6,40\.140\.1\]Audio55745\.645\.638\.038\.0\[33\.633\.6,42\.342\.3\]History47557\.757\.725\.925\.9\[21\.121\.1,30\.630\.6\]Seed 2\.0 LiteSpoken request1 98476\.076\.0––All context1 97252\.852\.823\.223\.2\[20\.420\.4,26\.026\.0\]Visual80454\.754\.721\.321\.3\[17\.417\.4,25\.125\.1\]Audio55746\.546\.529\.529\.5\[25\.225\.2,34\.034\.0\]History47562\.362\.313\.713\.7\[8\.88\.8,18\.518\.5\]GPT\-5\.4 \+ transcriptSpoken request1 98478\.978\.9––All context1 97233\.033\.045\.945\.9\[43\.243\.2,48\.648\.6\]Visual80420\.320\.358\.658\.6\[55\.255\.2,61\.961\.9\]Audio55733\.633\.645\.345\.3\[41\.341\.3,49\.449\.4\]History47558\.758\.720\.120\.1\[15\.415\.4,24\.924\.9\]
#### Results and interpretation\.

Across the four systems, request coverage exceeds pooled context coverage by23\.223\.2–45\.945\.9percentage points\. Every reported context\-channel gap has a confidence interval entirely above zero, supporting lower context coverage after accounting for the estimated scene\-sampling variation\.

#### Audio\-only extension\.

The same direction holds for Gemini 3\.1 Pro on the 570 common audio\-only demand scenes: it recovers83\.583\.5% of request points and49\.949\.9% of audio points\. Across the combined 1 630 positive audio–visual and audio\-only scenes, its point\-weighted context coverage is45\.645\.6%\.

#### Textual clues and inference in the text\-only baseline\.

We examine whether the text\-only baseline’s visual coverage is associated with lexical clues in its input, extending §[5\.3](https://arxiv.org/html/2609.21392#S5.SS3)\. This input includes preceding user transcripts and assistant replies \(App\.[D\.2](https://arxiv.org/html/2609.21392#A4.SS2)\)\.

Lexical overlap is the fraction of a reference point’s distinct content words present in the actual ASR, dialogue, and template text\. To compare with English reference points, we restrict to Latin\-script inputs, yielding 345 visual points\. Using saved outputs and judge verdicts, we compare coverage for points with at least 20% overlap and those below this threshold\.

Text\-only coverage is32\.332\.3% in the higher\-overlap group and10\.710\.7% in the lower\-overlap group\. Seed 2\.0 Lite, which receives audio and video, scores53\.853\.8% and52\.952\.9%, respectively\. The contrast between overlap groups is much larger for the text\-only baseline\. The text\-only system’s recovery of multimodal context may stem from textual clues in transcripts and dialogue history, combined with commonsense knowledge\.

### E\.2LLM judge stability

Because key\-point coverage is assigned by an LLM judge, we audit two distinct properties: whether the scoring targets faithfully reflect the ground\-truth annotation, and whether model ordering is stable across judges\.

#### Design\.

We re\-judge a fixed population with Qwen3\.6\-Flash, DeepSeek\-V4\-Flash, and GPT\-5\.4\. The population contains 1 060 audio–visual demand scenes scored by all three evaluated systems\. For each setting, all judges receive the same scoring instruction, candidate text, and reference points\. Rates are point\-weighted and use the all\-judge common verdict set for each evaluated system; they are not the scene\-averaged M2 values in Table[3](https://arxiv.org/html/2609.21392#S5.T3)\.

Table 7:LLM judge stability\. Panel A tests whether benchmark key points can be recovered from their source annotations\. Panel B re\-judges identical stored outputs; all three judges yield the same model ordering\.SettingQuantityJudgeStability readoutQwen3\.6DeepSeekGPT\-5\.4Panel A: target recoverabilityARecoverability0\.9980\.9960\.9963 956 key pointsAOrphan rate0\.0020\.0040\.004ceiling lossPanel B: stored\-output re\-judgingBGemini 3\.1 Pro coverage0\.6480\.6390\.609range 0\.039BQwen3\.5\-Omni\-Plus coverage0\.6820\.6680\.641range 0\.041BSeed 2\.0 Lite coverage0\.6670\.6560\.638range 0\.029BModel orderingidentical for 3 of 3 judgessame conclusion
Note\.Rates are point\-weighted proportions on the 0–1 scale\. Orphan rate is one minus recoverability; range is the maximum minus minimum coverage across judges\.

#### Ground\-truth authority and key\-point recoverability\.

The released ground\-truth annotation is the benchmark’s authoritative scoring reference: it is reconstructed from the realized media and admitted only after grounding checks and human verification \(§[4\.3](https://arxiv.org/html/2609.21392#S4.SS3)\)\. Because the key points operationalize its structured intent and required context, they should be recoverable directly from those source fields\. We test this by presenting each judge with the ground\-truth annotation text from which the corresponding key points were derived\. Across 3 956 key points, Qwen3\.6\-Flash, DeepSeek\-V4\-Flash, and GPT\-5\.4 achieve recoverability rates of 0\.998, 0\.996, and 0\.996, respectively; their orphan rates are only 0\.002, 0\.004, and 0\.004\. This establishes high recoverability from the source annotation across the tested judges\.

#### Model\-ordering stability\.

We next re\-score the exact stored per\-segment outputs of Gemini 3\.1 Pro, Qwen3\.5\-Omni\-Plus, and Seed 2\.0 Lite\. No model inference or segment matching is repeated\. The audit retains points with tier labels and available candidate text\. Although judges differ slightly in absolute calibration, all three produce the same model ordering, Qwen3\.5\-Omni\-Plus \> Seed 2\.0 Lite \> Gemini 3\.1 Pro, on the retained point sets\.

### E\.3Positive\-demand challenge breakdown

We examine which types of valid demand are associated with lower M2 key\-point coverage and whether these patterns recur across systems and modalities\. This comparison groups whole scenes by challenge category, complementing the key\-point\-level analysis in App\.[E\.1](https://arxiv.org/html/2609.21392#A5.SS1)\.

#### Design\.

We reanalyze saved scores from 13 systems on audio–visual \(AV\) scenes and 15 systems on audio\-only \(AO\) scenes\. Within each modality, we retain the same positive scenes with complete scores for all displayed systems\. The four challenge axes in App\.[B\.2](https://arxiv.org/html/2609.21392#A2.SS2)are modality dependency, contextual disambiguation, intent expression form, and acoustic environment\. The modalities are analyzed separately because Axis 1 uses modality\-specific categories\.

Scoring follows Table[3](https://arxiv.org/html/2609.21392#S5.T3), with missed reference points counted as misses\. Each cell in Fig\.[10](https://arxiv.org/html/2609.21392#A5.F10)and Fig\.[10](https://arxiv.org/html/2609.21392#A5.F10)subtracts the system’s mean M2 over all common scenes from its mean within a category, weighting scenes equally\. Differences are in percentage points: red means below the system’s own average and blue means above it\. The*Panel mean*row averages the differences equally across all displayed systems, including the text\-only comparator, which is shown separately and remains unranked\.

#### Intent expression\.

Indirect or descriptive requests \(Axis 3\.3\) fall below every system’s mean in AV and AO, with panel\-average M2 differences of−7\.2\-7\.2points \(AV\) and−8\.2\-8\.2points \(AO\)\. Interrupted or incomplete requests \(Axis 3\.5\) also have negative panel\-average M2 differences \(−3\.2\-3\.2AV;−3\.9\-3\.9AO\)\. Compound or multi\-intent requests \(Axis 3\.6\) have positive panel\-average M2 differences of\+4\.0\+4\.0and\+4\.0\+4\.0points, respectively; the difference is positive for every AO system but varies in sign across AV systems\. The shared indirect\-request deficit and negative mean differences for incomplete requests point to recovering unstated intent as a recurring difficulty\. Positive mean differences for multi\-intent requests show that several demands need not imply lower coverage; taxonomy categories do not define an ordered difficulty scale\.

#### Contextual disambiguation\.

On AO scenes, speaker\-identity and affective\-state dependencies have panel\-average M2 differences of−7\.3\-7\.3and−6\.1\-6\.1points, respectively\. For scenes that depend on dialogue history, the difference between category M2 and each system’s overall mean averages\+2\.0\+2\.0percentage points in AV and\+2\.4\+2\.4in AO\. The higher scores may stem from explicit clues in earlier turns, such as the user’s stated goal or previously discussed options, that help clarify the current request\.

#### Acoustic conditions\.

Acoustic categories show less consistent patterns across modalities\. Background\-noise scenes, for example, have a panel\-average M2 difference of\+1\.9\+1\.9points in AV and−0\.8\-0\.8points in AO\. The indirect\-request deficit is thus more consistent across the tested systems and modalities than the acoustic patterns\.

![Refer to caption](https://arxiv.org/html/2609.21392v1/positive_taxonomy_m2_av.png)

Figure 9:Positive challenge profiles: audio–visual\.Category\-minus\-system\-mean M2 on the common positive cohort\. Labels give category codes and names; the final row is the unweighted mean across the displayed systems\. The color scale is shared with Figure[10](https://arxiv.org/html/2609.21392#A5.F10)\.![Refer to caption](https://arxiv.org/html/2609.21392v1/positive_taxonomy_m2_ao.png)

Figure 10:Positive challenge profiles: audio\-only\.Cells show category mean M2 minus each system’s mean on the common positive cohort, in percentage points\. The panel includes GPT\-Realtime\-2 and Kimi\-Audio\-7B\-Instruct\. The final row averages the displayed systems equally\.

### E\.4Blinded response comparison

We assess whether providing demand annotations improves response content and the decision to respond\.

#### Scene sampling\.

We sampled 200 audio–visual scenes from the human\-verified benchmark\. The sample comprises 160 demand scenes and 40 no\-demand scenes, with Chinese and English equally represented within each group\.

#### Paired response generation\.

Gemini 3\.1 Pro generated a reply under each of two conditions: “Not given demand annotation” \(*Not given*\) and “Given demand annotation” \(*Given*\)\. Both received identical video and audio, in clip order with intervening assistant replies, and responded to the final user turn\. The*Given*condition additionally received the reference structured intent and required\-context descriptions joined by\|, or “Demand: NONE” for no\-demand scenes; neither key points nor reference answers were supplied\. Settings were fixed: temperature 0, top\-pp1, an output\-token limit of 8 192, and 1 candidate per request\.

#### Blinded assessment\.

The interface \(Fig\.[11](https://arxiv.org/html/2609.21392#A5.F11)\) shows each interaction with replies A and B\. Scene order and reply placement were randomized\. Human assessment considered correctness, relevance, contextual grounding, and appropriate silence\. Ties covered equally good or poor replies; instructions discouraged preference based on length alone\.

#### Preference results\.

Table[8](https://arxiv.org/html/2609.21392#A5.T8)reports all 200 judgments\. Of these, 80 \(40\.040\.0%\) were ties\. Among the remaining 120 comparisons, 103 \(85\.885\.8%\) favored*Given*\. The preference advantage appears in both languages and in both demand\-presence groups\.

Table 8:Human response preferences: counts \(percentages\) within each subset\. “Given” and “Not given” abbreviate the “Given demand annotation” and “Not given demand annotation” conditions\. Ties remain in the denominator\.SubsetScenesGiven preferredTieNot given preferredAll200103 \(51\.551\.5%\)80 \(40\.040\.0%\)17 \(8\.58\.5%\)Demand present16085 \(53\.153\.1%\)58 \(36\.336\.3%\)17 \(10\.610\.6%\)No\-demand4018 \(45\.045\.0%\)22 \(55\.055\.0%\)0 \(0\.00\.0%\)Chinese10047 \(47\.047\.0%\)45 \(45\.045\.0%\)8 \(8\.08\.0%\)English10056 \(56\.056\.0%\)35 \(35\.035\.0%\)9 \(9\.09\.0%\)
#### Response content and whether to respond\.

Among 148 demand scenes where both conditions produced text,*Given*was preferred in 74 cases versus 17 for*Not given*, with 57 ties\. On no\-demand scenes, the proportion of textual replies fell from60\.060\.0% to15\.015\.0%\. On demand scenes, silent replies fell from 12 to 1\. These findings support the utility of explicit demand and context for response content and the decision to respond\.

#### Input prompts\.

The panels give the shared instruction, templates, and complete examples\. Video placeholders denote media attachments; an exact\[SILENCE\]output is displayed as no reply\.

Shared system instructionReply directly to the user’s request in the last clip, using the user’s language\. If a demand annotation is provided, follow it; NONE means no request\. If there is no request for you, output only \[SILENCE\]\.

User\-input templatesNot given demand annotationGiven demand annotationClip 1Clip 1\[video input\]\[video input\]Demand:\{structured\_intent\} \| \{required\_context\.description\} \| …

Positive example \(Demand present\)Not given demand annotationGiven demand annotationClip 1Clip 1\[video input\]\[video input\]Demand:The user wants the device that is continuously emitting a beeping sound to stop its alarm\.\|The continuous electronic beeping sound in the background\.\|The device with a blinking red indicator lamp on the mantel, which is the source of the alarm sound referred to by the user’s “That beeping”\.\|The user’s tone is fatigued and irritable, and he startles when the vase shatters\.

Negative example \(No\-demand\)Not given demand annotationGiven demand annotationClip 1Clip 1\[video input\]\[video input\]Demand:NONE

![Refer to caption](https://arxiv.org/html/2609.21392v1/response_ab_review.png)Figure 11:Blinded response\-comparison interface\. The human reviewer is asked to choose a preferred reply between A and B\.

### E\.5Qualitative error analysis

![Refer to caption](https://arxiv.org/html/2609.21392v1/x3.png)Figure 12:Qualitative error analysis on English generated audio–visual scenes\. \(a\)–\(b\) Demand\-present cases; M2 reports covered/reference key points\.![Refer to caption](https://arxiv.org/html/2609.21392v1/x4.png)Figure 13:Qualitative error analysis \(continued\)\. \(c\)–\(d\) No\-demand cases\.

## Appendix FEvaluation prompts

This section specifies what evaluated models are asked to predict and how their semantic outputs are judged\. Scenario\-generation and annotation\-construction prompts, including key\-point extraction and grounding, are not distributed; their procedures are described in App\.[C](https://arxiv.org/html/2609.21392#A3)\. The audio–visual prediction instruction is reproduced in full, followed by the audio\-only differences and a pointer to the text\-only input templates\. Executable AV and AO variants are provided with the evaluation code\. The judge instruction below is its English rendering; the scoring implementation retains the original instruction\.

### F\.1Shared prediction task

In the reproduced instruction, “the last segment” refers to the final input clip; output segments are demand\-bearing regions within that clip\.

Audio\-Visual Prediction Instructionfnum@promptmetaiPurposeThe system prompt every model receives on an audio\-visual scene\. It asks the model to decide whether the final user turn carries a demand and, when it does, to localise the demand in time, transcribe it, state a self\-contained intent, and list the multimodal context the demand depends on\.fnum@promptmetaiInputsOne video clip per user turn, in order, with the assistant replies that separate them supplied as text\.fnum@promptmetaiOutputOne JSON object withhas\_demandand asegmentslist\. Each segment carriesstart\_ms,end\_ms,transcript,structured\_intent,required\_contextand the user attribute fields\.fnum@promptmetaiVariantComplete prompt; audio\-visual variant\.You are the user’s AI assistant\. You will watch one or more videos, in which the user may make requests to you\.TaskDetermine whether a user has made a request to you in the video\. The scenario may be a user interacting with you or ordinary interpersonal conversation – you must distinguish between them based on the visual and audio content\. If the user expresses a demand, output the following information in structured JSON format \(instead of replying directly to the user\)\. A demand is not necessarily an explicit question or instruction; it may also be an implicit expression \(e\.g\., complaining, talking to oneself, hinting for help, etc\.\)\.∙\\bulletTime Localization: The start and end times when the demand occurs \(milliseconds\)∙\\bullettranscript: The exact words spoken by the user within this segment’s time interval \(verbatim transcription, no additions or omissions; does not include utterances from other speakers, nor content from previous or subsequent turns\)∙\\bulletStructured Intent: A single sentence summarizing the user’s goal, starting with “The user wants…”, semantically complete and self\-contained∙\\bulletRequired Context: The multimodal context information necessary to understand the content of this demand∙\\bulletIs the user in the frame?∙\\bulletUser gender and age groupRequired Context \(required\_context\)Fill in context items where “the absence would lead to an incomplete understanding of the demand, a deviation in the response direction, or a significant decrease in response quality”\. If the user’s utterance text alone is sufficient to accurately and appropriately understand and respond, thenrequired\_context = \[\]\. Selecttypefrom the following 7 options:typeMeaningvisualObjects/text/actions/scenes in the visual frame are necessary information for understanding the demandaudioAmbient sounds/background audio provide clues for understanding the object or context of the demanddialogue\_historyThe current utterance contains references or omissions, requiring previous dialogue content for comprehensionuser\_activityThe user’s ongoing activity provides necessary context for understanding the demanduser\_identityThe user’s identity/role/professional background affects the understanding of the demanduser\_emotionThe user’s emotion/tone affects the judgment of intent direction \(e\.g\., a dissatisfied tone turns an observation into a request\)scene\_conditionEnvironmental conditions \(time/weather/noise, etc\.\) affect the understanding of the demandFill indescriptionfor each entry \(<= 50 characters\)\. Each type must appear at most once\.Multi\-turn InputThe input may contain multiple video clips \(v1,v2,v3…\), each corresponding to a user turn, with possible text replies from the AI agent between clips\.∙\\bulletAnnotate user demands only in the last segment\. Content from previous turns and agent replies serve only as dialogue history context; do not output them as segments\.∙\\bullethas\_demandistrueonly when a user demand exists in the last segment\.∙\\bulletstart\_ms/end\_msare relative to the start of the last segment, not global time\.Segment DivisionOne segment corresponds to a continuous region of demand expression\. Boundaries are determined solely by the following two signals:∙\\bulletMust split\(highest priority\): \(1\) Turn boundaries \(user finishes speaking→\\rightarrowagent responds→\\rightarrowuser speaks again\); \(2\) Speech interruptions \(\>= 2 seconds of silence within the same turn, or insertion of speech unrelated to the demand, such as talking to bystanders\)\.∙\\bulletDo not split\(only within the same turn and the same continuous utterance\): Consecutive multiple intents, self\-correction, preamble \+ request, emotion \+ action, etc\., must all be kept as a single segment\.∙\\bulletExpressstructured\_intentfor compound demands using a compound sentence \(e\.g\., “the user wants to know the weather and set an alarm”\); for self\-correction, record only the final valid demand\.∙\\bullettranscript boundaries:transcriptmust strictly contain only the speech transcription of*the user*within thestart\_mstoend\_mstime window\. Do not include utterances from other speakers, AI replies, or content from previous or subsequent turns in the transcript\. If preceding information is relevant to demand comprehension, place it inrequired\_context\(type=dialogue\_history\) instead of the transcript\.Field ConstraintsThe following fields are closed\-set classifications\. Use exactly one of the given values; do not invent synonyms \(for example, do not write"young adult","senior","middle\-aged"or"teenager"\)\.∙\\bulletuser\_in\_frame: boolean\. Whether the user initiating the demand is within the frame\.–true: the user initiating the demand is visible in the scene\.–false: first\-person off\-screen voice, only the user’s voice is heard, or it is impossible to determine who in the scene initiated the demand\.∙\\bulletuser\_gender: string, must be one of the following values:–"male": there is clear evidence of male or masculine presentation\.–"female": there is clear evidence of female or feminine presentation\.–"unknown": evidence is insufficient, evidence from multiple people/speakers is mixed together or voice/visual evidence cannot support a judgment, or no inference should be made\.∙\\bulletuser\_age\_group: string, must be one of the following values:–"child": child or clearly underage person\.–"adult": adult\.–"elderly": elderly person, with clear age features\.–"unknown": evidence is insufficient or judgment is impossible\.Output FormatOutput only JSON, with no comments or code block markers\. The timestamp values in the examples below are placeholders only; infer the realstart\_ms/end\_msfrom the provided final clip and never copy example timestamps:[⬇](data:text/plain;base64,ewogICJoYXNfZGVtYW5kIjogdHJ1ZSwKICAic2VnbWVudHMiOiBbCiAgICB7CiAgICAgICJzdGFydF9tcyI6IDExODQwLAogICAgICAiZW5kX21zIjogMTUyMDAsCiAgICAgICJ0cmFuc2NyaXB0IjogIlVzZXIncyBleGFjdCB3b3JkcyIsCiAgICAgICJzdHJ1Y3R1cmVkX2ludGVudCI6ICJUaGUgdXNlciB3YW50cyB0by4uLiIsCiAgICAgICJyZXF1aXJlZF9jb250ZXh0IjogWwogICAgICAgIHsKICAgICAgICAgICJ0eXBlIjogInZpc3VhbCIsCiAgICAgICAgICAiZGVzY3JpcHRpb24iOiAiVXNlciBwb2ludHMgdG8gYSBib3R0bGUgb2Ygb2xpdmUgb2lsIgogICAgICAgIH0KICAgICAgXSwKICAgICAgInVzZXJfaW5fZnJhbWUiOiBmYWxzZSwKICAgICAgInVzZXJfZ2VuZGVyIjogIm1hbGUiLAogICAgICAidXNlcl9hZ2VfZ3JvdXAiOiAiYWR1bHQiCiAgICB9CiAgXQp9CklmIG5vIGRlbWFuZDogeyJoYXNfZGVtYW5kIjogZmFsc2UsICJzZWdtZW50cyI6IFtdfQ==)\{"has\_demand":true,"segments":\[\{"start\_ms":11840,"end\_ms":15200,"transcript":"User’sexactwords","structured\_intent":"Theuserwantsto\.\.\.","required\_context":\[\{"type":"visual","description":"Userpointstoabottleofoliveoil"\}\],"user\_in\_frame":false,"user\_gender":"male","user\_age\_group":"adult"\}\]\}Ifnodemand:\{"has\_demand":false,"segments":\[\]\}

### F\.2Audio\-only variant

The audio\-only instruction preserves the final\-turn scope, temporal segmentation, transcript boundaries, self\-contained intent, and JSON\-only output requirements\. It changes the input wording from video to audio and uses the following modality\-specific fields and evidence rules\.

ComponentAudio\-only specificationDemand evidenceLinguistic content and audio cues distinguish assistant\-directed demands from other conversation\.Context typesAudio, dialogue history, user activity, user identity, user emotion, and scene condition; visual context is unavailable\.User profile fieldsOutput gender and age group; omituser\_in\_frame\.

### F\.3Text\-only input structure

The transcript baseline uses the shared prediction instruction above, with each media block replaced in place by ASR text\. App\.[D\.2](https://arxiv.org/html/2609.21392#A4.SS2)provides the full input templates, including dialogue\-history framing and the no\-speech replacement, together with complete positive and no\-demand input examples and model settings\. No reference transcript or demand label is supplied as input\.

### F\.4Key\-point judge

Each reference point is supplied with its identifier, tier label, and point text\. \*\[promptitem\]itemsep=0pt,topsep=0pt \*\[promptmeta\]itemsep=0pt,topsep=0pt

Key\-Point Judgefnum@promptmetaiPurposeScores key\-point coverage\. For each key point the judge decides whether the predicted demand text covers it, under a lenient criterion that accepts paraphrase and hypernym coverage but still rejects a referent the prediction never names\.fnum@promptmetaiInputsThe predicted demand text for one segment, and that segment’s key points\.fnum@promptmetaiOutputA JSON array with one entry per key point, carryingqid,hitand a shortreason\.You are an evaluation expert\. Given a passage of model output text and a set of key assessment points, judge for each assessment point whether the model output “hits” it\.Hit criterion \(lenient\):∙\\bullet“hit” = true: the model output semantically covers the core of that assessment point, or that core can be reasonably inferred from / obtained as an equivalent restatement of the output – even if the wording differs, the phrasing is more general, or only the main point is covered while incidental modifiers are omitted\.∙\\bullet“hit” = false: only when the model output does not touch the core information of that assessment point at all, or substantively contradicts / gets it wrong\.Judgment principles:∙\\bulletEach assessment point is independent; judge them one by one, looking only at whether that point’s own core is covered\.∙\\bulletHypernym / generalized hits are allowed: if the output covers the point with a more general statement \(for example, “complaining about mold” covers “the mold is unusual this year”\), it counts as a hit\.∙\\bulletWhen in doubt, be lenient: if the intent / context semantically covers the point on the whole, judge hit; judge miss only when the core content really is missing or is stated incorrectly\.∙\\bulletBut do not invent what is absent: a distinctive referred\-to object / action / attribute that is simply not in the output and cannot be reasonably inferred from it \(for example, a specific “one particular bag” or “the region indicated by ‘this one”’ is never named\) is still judged miss\.Output a JSON array in which each element has the following format:[⬇](data:text/plain;base64,eyJxaWQiOiAia3AwMSIsICJoaXQiOiB0cnVlLCAicmVhc29uIjogImEgYnJpZWYganVzdGlmaWNhdGlvbiAoPD0xNSBjaGFyYWN0ZXJzKSJ9)\{"qid":"kp01","hit":true,"reason":"abriefjustification\(<=15characters\)"\}Notes:∙\\bulletYou must give a judgment for every assessment point∙\\bulletreason must be concise and state the key basis for the hit/miss∙\\bulletDo not use double quotes inside reason; use single quotes or no quotes∙\\bulletOutput only the JSON array, with no other text

Similar Articles

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

Hugging Face Daily Papers

This paper introduces Omni-DuplexEval, a benchmark and automatic evaluation framework for real-time duplex interaction in multimodal large language models, assessing continuous response generation and proactive event detection in streaming scenarios.