Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue
Summary
This paper investigates how spoken negation cues in human dialogue are reflected in multimodal nonverbal behavior, using time-series classification models to distinguish negation contexts from control contexts without lexical or acoustic input.
View Cached Full Text
Cached at: 09/16/26, 08:47 AM
# Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue
Source: [https://arxiv.org/html/2609.16396](https://arxiv.org/html/2609.16396)
Patrick SchrottenbacherAlexander MehlerAffiliation:\{hammerla⋅\\cdotschrottenbacher⋅\\cdotmehler\}@em\.uni\-frankfurt\.deAffiliation:Goethe University, Frankfurt am Main, Germany
###### Abstract
Negation is typically modeled through its linguistic realization, although spoken interaction is accompanied by tightly coordinated nonverbal behavior\. We ask whether contexts centered on spoken negation cues contain measurable multimodal behavioral information: whether they can be distinguished from matched control contexts without lexical or acoustic input, where this information occurs in time, which modalities carry it, and whether it extends to the dialogue partner\. We study 27 human\-human interviews conducted in virtual reality, comprising temporally aligned gaze, facial, head, body, hand, and finger behavior and 964 annotated negation cues\. Treating classification as a predictive probe, we compare 20 time\-series models while excluding lexical and acoustic information, and then systematically vary temporal context, interactional source, modality availability, and event timing\. Across grouped 10\-fold cross\-validation, the strongest probes reach up to \.75 mean held\-out AUROC from speaker\-side behavior\. Temporal analyses show that predictive information is concentrated around cue onset but remains detectable over a broader surrounding interval, while dialogue\-partner behavior carries weaker predictive information with a comparatively diffuse temporal profile\. Ablation and timing perturbations further show that facial features produce the largest modality\-ablation effect and that the trained probe is sensitive to the temporal organization of the observed events\.
## 1Introduction
Human communication is inherently multimodal\. Spoken language is accompanied by gaze, facial behavior, head and body motion, and manual gestures, which constitute channels of communication that are closely coordinated in time\([Goldin\-Meadow and Alibali, 2013](https://arxiv.org/html/2609.16396#bib.bib14);[Kelly et al\., 2009](https://arxiv.org/html/2609.16396#bib.bib29)\)\. While computational models increasingly exploit such signals, it remains less clear how specific linguistic phenomena are reflected in nonverbal behavior, and at what temporal scale these effects become observable\([Tsai et al\., 2019](https://arxiv.org/html/2609.16396#bib.bib60)\)\. Negation provides a particularly interesting case\. In computational linguistics, negation has predominantly been studied through its linguistic realization, including the identification of negation cues and their scope\([Morante and Blanco, 2012](https://arxiv.org/html/2609.16396#bib.bib39)\)\. At the same time, negative expressions can be systematically associated with nonverbal behavior, with head movements providing a particularly well\-established example\([Kendon, 2002](https://arxiv.org/html/2609.16396#bib.bib30)\)\. Moreover, dialogue is a joint activity in which listeners continuously produce behavioral responses to speakers\([Clark and Krych, 2004](https://arxiv.org/html/2609.16396#bib.bib7)\)\. Explicit negation cues may therefore be associated with measurable patterns in the surrounding multimodal interaction\. Such patterns may be distributed across modalities, participants, and time rather than being tied to a single gesture or moment\. In this work, we study the temporal multimodal correlates of spoken negation cues in human\-human dialogue\. We use temporally aligned multimodal interaction data and treat spoken words as anchors around which surrounding nonverbal events are observed\. We then ask whether windows centered on negation cues can be distinguished from matched control\-word windows using multimodal behavior without lexical or acoustic input\. We use classification as a predictive probe to quantify how cue\-centered contexts differ from matched controls across temporal context, modality, and interactional source\. This provides a controlled way to investigate*when*,*where*, and*through whom*behavioral information associated with explicit negation cues becomes predictive\. We address the following research questions:
- •Discriminability:Can negation\-cue\-centered contexts be distinguished from matched controls using only surrounding multimodal behavior?
- •Temporal extent:How does this distinguishability change across different temporal contexts?
- •Modality contribution:Which behavioral modalities contribute most strongly to cue\-centered discriminability?
- •Interactional source:How is cue\-associated predictive information distributed between the speaker and the dialogue partner?
- •Temporal sensitivity:Which behavioral event types show the greatest predictive sensitivity to their temporal relation to the spoken word?
By framing prediction as a probe of multimodal interaction, our study examines how explicit negation cues are embedded in coordinated, time\-dependent human behavior rather than treating them only as isolated linguistic labels\.
## 2Related Work
### 2\.1Negation in NLP
Negation has been studied extensively in natural language processing, most commonly through the identification of*negation cues*and the linguistic material falling within their*scope*\([Morante and Sporleder, 2012](https://arxiv.org/html/2609.16396#bib.bib41);[Morante and Blanco, 2012](https://arxiv.org/html/2609.16396#bib.bib39)\)\. Annotated resources such as BioScope\([Szarvas et al\., 2008](https://arxiv.org/html/2609.16396#bib.bib57)\)and the\*SEM2012 shared task\([Morante and Blanco, 2012](https://arxiv.org/html/2609.16396#bib.bib39)\)established widely used formulations of cue, scope, and focus detection\. Early computational approaches included rule\-based systems\([Chapman et al\., 2001](https://arxiv.org/html/2609.16396#bib.bib6);[Peng et al\., 2018](https://arxiv.org/html/2609.16396#bib.bib48)\)and feature\-based statistical models that combined lexical and syntactic information with classifiers such as SVMs and CRFs\([Morante and Daelemans, 2009](https://arxiv.org/html/2609.16396#bib.bib40);[Lapponi et al\., 2012](https://arxiv.org/html/2609.16396#bib.bib34)\)\. Subsequent work increasingly adopted neural architectures\([Fancellu et al\., 2016](https://arxiv.org/html/2609.16396#bib.bib12)\), including transformer\-based systems such asNegBERT\([Khandelwal and Sawant, 2020](https://arxiv.org/html/2609.16396#bib.bib31)\)and syntax\-aware graph attention extensions such asD\-Neg\([Hammerla et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib17)\)\. Despite these advances, negation remains challenging for contemporary language models\. Negated examples are comparatively underrepresented in common NLU benchmarks\([Hossain et al\., 2022](https://arxiv.org/html/2609.16396#bib.bib24)\), pretrained language models can be insensitive to distinctions between affirmative and negated statements\([Kassner and Schütze, 2020](https://arxiv.org/html/2609.16396#bib.bib28)\), and related limitations persist in larger autoregressive language models\([Truong et al\., 2023](https://arxiv.org/html/2609.16396#bib.bib59);[García\-Ferrero et al\., 2023](https://arxiv.org/html/2609.16396#bib.bib13)\)\. These approaches nevertheless treat negation primarily through linguistic input, leaving its manifestation in concurrent nonverbal behavior largely outside the modeling objective\.
### 2\.2Multimodal Negation
Negation is not expressed exclusively through the verbal channel\. Head shakes are among its best documented nonverbal correlates\([Kendon, 2002](https://arxiv.org/html/2609.16396#bib.bib30)\), while manual gestures and combinations of head and hand movements have been shown to systematically accompany spoken negation and align with its node, scope, and focus\([Harrison, 2010](https://arxiv.org/html/2609.16396#bib.bib19);[Harrison, 2014](https://arxiv.org/html/2609.16396#bib.bib20);[Harrison and Larrivée, 2016](https://arxiv.org/html/2609.16396#bib.bib22)\)\. Prosodic and gestural cues can further affect the interpretation of negative utterances\([Prieto et al\., 2013](https://arxiv.org/html/2609.16396#bib.bib49);[Tubau et al\., 2015](https://arxiv.org/html/2609.16396#bib.bib61);[González\-Fuente et al\., 2015](https://arxiv.org/html/2609.16396#bib.bib15)\), and gesture has been related to agreement and refusal\([Guidetti, 2005](https://arxiv.org/html/2609.16396#bib.bib16)\), scope disambiguation\([Brown and Kamiya, 2019](https://arxiv.org/html/2609.16396#bib.bib4)\), and covert negative meaning\([Inbar and Shor, 2019](https://arxiv.org/html/2609.16396#bib.bib25)\); see[Harrison \(2024\)](https://arxiv.org/html/2609.16396#bib.bib21)for a recent overview\. Computational work has increasingly examined whether such distinctions are captured by multimodal models\. Pretrained vision\-language models have been shown to struggle with negation\([Dobreva and Keller, 2021](https://arxiv.org/html/2609.16396#bib.bib11)\), motivating negation\-focused multimodal evaluation\([Parcalabescu et al\., 2022](https://arxiv.org/html/2609.16396#bib.bib44);[Sato et al\., 2023](https://arxiv.org/html/2609.16396#bib.bib54)\)and negation\-aware video retrieval\([Wang et al\., 2022](https://arxiv.org/html/2609.16396#bib.bib63)\)\. Subsequent work has analyzed how negation is represented withinCLIP\([Quantmeyer et al\., 2024](https://arxiv.org/html/2609.16396#bib.bib50)\), while more recent studies have developed dedicated benchmarks and training strategies for improving negation understanding in vision\-language models\([Park et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib45);[Alhamoud et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib2)\)\. Most recently,[AbuSaleh et al\. \(2026\)](https://arxiv.org/html/2609.16396#bib.bib1)study cross\-modal negation in video\-text data through representation analysis and cross\-modal attention fusion\.
### 2\.3Temporal Multimodal Interaction
Nonverbal behavior is temporally coordinated with speech without necessarily being strictly synchronous, as shown for gesture\-speech timing\([Leonard and Cummins, 2011](https://arxiv.org/html/2609.16396#bib.bib35)\)and for gaze coupling between speakers and listeners\([Richardson and Dale, 2005](https://arxiv.org/html/2609.16396#bib.bib51);[Richardson et al\., 2007](https://arxiv.org/html/2609.16396#bib.bib52)\)\. Dialogue is further shaped by behavioral feedback from interlocutors, which speakers continuously monitor during interaction\([Clark and Krych, 2004](https://arxiv.org/html/2609.16396#bib.bib7)\)\. Computational work reflects this temporal and interactional perspective through multimodal corpora combining speech with gaze, gesture, and other behavioral signals\([Carletta et al\., 2006](https://arxiv.org/html/2609.16396#bib.bib5);[Kontogiorgos et al\., 2018](https://arxiv.org/html/2609.16396#bib.bib33);[Kim et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib32)\), as well as models designed to capture temporally unaligned cross\-modal dependencies\([Tsai et al\., 2019](https://arxiv.org/html/2609.16396#bib.bib60)\)\. These findings suggest that both the temporal context of a linguistic event and the source of the surrounding behavior may affect how much multimodal information it carries\.
Taken together, previous work establishes systematic links between spoken negation and nonverbal behavior, shows that multimodal models remain challenged by negation, and demonstrates that verbal and nonverbal behavior are temporally structured within dialogue\. However, the behavioral correlates surrounding spoken negation cues have not been systematically characterized from fine\-grained, temporally aligned event streams spanning multiple modalities\. We address this gap using event logs from human\-human interaction in virtual reality \(VR\), quantifying how cue\-centered discriminability varies across temporal contexts, modalities, and interactional roles\.
## 3Multimodal Dialogue Corpus
Figure 1:Interactional VR interview setting with the interviewer \(left\) navigating the survey questionnaire and the interviewee seated opposite\.Our experiments use a German\-language multimodal dialogue corpus from 27 individual survey interviews conducted in VR\.111Project code and processed datasets are available at[https://github\.com/vrneg/vrneg01](https://github.com/vrneg/vrneg01)and[https://huggingface\.co/VR\-Faces\-Neg](https://huggingface.co/VR-Faces-Neg)\. The dataset paper will be linked here once available\.The recordings involve3030unique participants, with33serving as interviewers and2727as interviewees\. During the study, participants were represented by avatars whose features were controlled through the Meta Quest Pro headsets they wore\. These headsets not only recorded their head and body positions within the room, but also their hands, fingers, gaze, face, and voice\. All of these features were recorded at approximately18\.518\.5Hz throughout interviews lasting27\.727\.7minutes on average\. The survey itself spanned five broad categories and purposefully included questions that were known to cause discomfort due to the sensitive nature of the topics \(e\.g\. income and Machiavellianism\)\. A depiction of the interview environment and avatars can be seen in Figure[1](https://arxiv.org/html/2609.16396#S3.F1)\. We organize the recordings in an event\-basedSurrealDBdatabase\([SurrealDB Ltd\., 2026](https://arxiv.org/html/2609.16396#bib.bib56)\), with separate temporally aligned streams for body, eye, facial, head, hand, and finger events\. The audio recordings are stored in approximately55s chunks, whileCrisperWhisper \(v1\)\([Zusag et al\., 2024](https://arxiv.org/html/2609.16396#bib.bib68)\)is applied to the full recordings to obtain word\-level transcriptions and timestamps for alignment with the multimodal streams\. We annotate the resulting transcriptions for negation usingD\-Neg\([Hammerla et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib17)\), identifying both negation cues and their corresponding scopes \(exact hyperparameter configurations for both models are provided in the Appendix[A](https://arxiv.org/html/2609.16396#A1)\)\. Across57,97157,971spoken words in the corpus, we identify953953negation instances, comprising964964cue words and2,8112,811scope words\. The cue distribution is skewed towardnicht, which accounts for approximately 80% of annotated cue instances\. The shared temporal representation enables extraction of multimodal windows around spoken words and locally matched controls from the surrounding conversational context\.
## 4Task Formulation
We operationalize the study of multimodal correlates surrounding explicit negation cues as a binary cue\-centered classification problem\. Each spoken word serves as a temporal*anchor*\. The positive class comprises words annotated as negation cues, whereas the negative class consists of words outside any annotated negation cue or scope\. Neither the lexical identity of the anchor word nor acoustic information is provided to the classifier; predictions are based exclusively on multimodal behavioral events surrounding the anchor\. We separately assess the potential contribution of visual articulation through a targeted lower\-face ablation in Section[7\.3](https://arxiv.org/html/2609.16396#S7.SS3)\. For an anchor wordwiw\_\{i\}at timetit\_\{i\}, we define a temporal context by left and right window sizeswLw\_\{L\}andwRw\_\{R\}and collect
EiwL,wR=\{ej∣ti−wL≤tj≤ti\+wR\}\.E\_\{i\}^\{w\_\{L\},w\_\{R\}\}=\\\{e\_\{j\}\\mid t\_\{i\}\-w\_\{L\}\\leq t\_\{j\}\\leq t\_\{i\}\+w\_\{R\}\\\}\.The two window sizes can be varied independently, with symmetric windows satisfyingwL=wRw\_\{L\}=w\_\{R\}and asymmetric windows allowingwL≠wRw\_\{L\}\\neq w\_\{R\}\. Furthermore, setting a negative boundary \(wL<0w\_\{L\}<0orwR<0w\_\{R\}<0\) creates an offset sliding window, effectively excluding the target event from the window itself\. Importantly, we do not define the context relative to the full word interval, i\.e\., fromwLw\_\{L\}before word onset towRw\_\{R\}after word offset\. Such a formulation would make the total observation interval dependent on the duration of the anchor word and could indirectly introduce lexical information, allowing the classifier to exploit systematic duration differences of frequent negation cues\. By default, event windows contain only behavior produced by the speaker of the anchor word\. To examine whether cue\-associated predictive information extends beyond the speaker, we additionally consider a*partner\-only*condition containing events from the interlocutor and a*dyadic*condition containing events from both participants\. These conditions allow us to test whether cue\-associated predictive information is carried primarily by the speaker’s own behavior or is also present in the behavior of the dialogue partner\. Together, these choices define two main dimensions of the task: the temporal context around the anchor, and the interactional source of the nonverbal events\.
## 5Experimental Setup
### 5\.1Dataset Construction and Conditions
We instantiate the task along the two main dimensions defined above: \(i\) the temporal context, and \(ii\) the interactional source\. For each of the 964 negation cue words \(positive anchors\), we sample one control word produced by the same speaker within the same55s audio chunk of the same interview, excluding all words annotated as negation cues or scopes\. To avoid selecting temporally near\-identical anchors, control words are required to occur more than11s from the corresponding negation cue\. If no eligible control is available within the same audio chunk, we progressively expand the search to temporally adjacent chunks from the same speaker and interview until an eligible control is found\. This yields a balanced dataset of 1,928 anchors while matching positive and negative examples as closely as possible in speaker identity and local conversational context\. We reuse the same positive anchors and sampled controls across all subsequent conditions\. For \(i\), we evaluate symmetric contexts withwL=wR∈\{0\.1,0\.5,1,2\.5\}w\_\{L\}=w\_\{R\}\\in\\\{0\.1,0\.5,1,2\.5\\\}s, corresponding to total context spans of0\.20\.2,11,22, and55s, respectively\. Subsequent analyses explore asymmetric contexts by varyingwLw\_\{L\}andwRw\_\{R\}independently\. Furthermore, we employ sliding windows to localize cue\-centered discriminability in time\. For \(ii\), each temporal condition is instantiated as*speaker\-only*,*partner\-only*, or*dyadic*, according to the interactional source of the nonverbal events\. The same data partitions are retained across all conditions, enabling paired comparisons across temporal contexts and interactional sources\.
### 5\.2Input Representations
Each example contains the nonverbal events surrounding an anchor word across eight modalities: eye, facial, head, body, left/right hand, and left/right finger behavior\. Lexical identity, audio, absolute timestamps, participant identifiers, and database metadata are excluded\. For fixed\-grid models, irregular event streams are converted into modality\-specific numerical features, including gaze, facial\-expression, head, body, hand, and finger\-tracking measurements, together with derived linear and angular velocity features; a complete inventory of channels and their modality assignments is provided in Appendix[B\.1](https://arxiv.org/html/2609.16396#A2.SS1)\. Continuous features are normalized using training data only and interpolated to 32 equidistant time steps, yielding𝐗∈ℝ698×32\\mathbf\{X\}\\in\\mathbb\{R\}^\{698\\times 32\}\. Missing positions are zero\-filled, with presence indicators retained separately\. The event Transformer instead operates directly on the irregular event sequence using modality, interactional source, and relative timing\. High\-frequency streams are limited to 128 observations per modality and window\.
### 5\.3Comparison Models
We compare a broad set of time\-series classifiers spanning several modeling paradigms \(the exact model configurations are provided in the Appendix[B](https://arxiv.org/html/2609.16396#A2)\)\. Random\-convolution methods compriseMiniRocket\([Dempster et al\., 2021](https://arxiv.org/html/2609.16396#bib.bib9)\), standard and sparseMultiRocket\+Hydra\([Tan et al\., 2022](https://arxiv.org/html/2609.16396#bib.bib58);[Dempster et al\., 2023](https://arxiv.org/html/2609.16396#bib.bib10)\), andSelF\-Rocket\([Lo et al\., 2026](https://arxiv.org/html/2609.16396#bib.bib36)\)\. Feature\-, dictionary\-, shapelet\-, and interval\-based models includeCastor\([Samsten and Lee, 2024](https://arxiv.org/html/2609.16396#bib.bib53)\),Weasel 2\.0\([Schäfer and Leser, 2023](https://arxiv.org/html/2609.16396#bib.bib55)\),MrSQM\([Nguyen and Ifrim, 2023](https://arxiv.org/html/2609.16396#bib.bib42)\), andDrCIF\([Middlehurst et al\., 2021](https://arxiv.org/html/2609.16396#bib.bib38)\), whileHIVE\-COTE 2\.0\([Middlehurst et al\., 2021](https://arxiv.org/html/2609.16396#bib.bib38)\)provides an ensemble of complementary time\-series classifiers\. We further evaluate three combinations of random\-convolution features and pretrained tabular models:RocketPFN\([O’Rourke et al\., 2026](https://arxiv.org/html/2609.16396#bib.bib43)\),MASHT[Cüppers and Vreeken \(2026\)](https://arxiv.org/html/2609.16396#bib.bib8), and adaptiveRocketPFN\. Our neural comparison models compriseInception\-TCN\([Bai et al\., 2018](https://arxiv.org/html/2609.16396#bib.bib3);[Ismail Fawaz et al\., 2020](https://arxiv.org/html/2609.16396#bib.bib26)\), a compact fusion\-TCN, and an irregular\-event Transformer\. We additionally evaluate classification adaptations ofT2M\-GPT\([Zhang et al\., 2023](https://arxiv.org/html/2609.16396#bib.bib66)\), a continuous\-latentT2M\-GPT\-v2ablation of the same architecture,MotionGPT\([Jiang et al\., 2023](https://arxiv.org/html/2609.16396#bib.bib27)\), the continuous diffusion\-head mechanism ofMotionGPT3\([Zhu et al\., 2026](https://arxiv.org/html/2609.16396#bib.bib67)\), and the hierarchical short\- and long\-spanG\-HTTarchitecture\([Wen et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib64)\)\. The comparison assesses whether the observed discriminability is robust across modeling approaches rather than introducing a new classification architecture\.
### 5\.4Cross\-Validation and Evaluation
We use stratified, group\-based 10\-fold cross\-validation with VR experiments as groups, ensuring that no recording session is shared across training, validation, and test data\. In each rotation, one fold is used for testing, one for validation, and the remaining eight for training\. All data\-dependent preprocessing and model selection are performed without access to the test fold\.
Our primary metric is AUROC, which measures class distinguishability independently of a decision threshold; we additionally report macro\-F1\. Results are averaged over ten test folds with identical splits across conditions\. We report bootstrap 95% CIs over fold\-wise scores, or fold\-wise differences for paired contrasts; complete intervals are provided in the figures and tables\.
## 6Results
We first compare classifiers to assess whether cue\-centered discriminability is robust across modeling approaches, and then use one high\-performing probe for the detailed analyses\.
0\.60\.60\.70\.7RocketPFNHIVE\-COTE 2\.0DrCIFMultiRocket\+HydraMASHTSelFRocketMiniRocketAdaptiveRocketPFNWEASEL 2\.0GHTTSparseMultiRocketInception\-TCNCompact FusionTCNT2M\-GPT v2Event TransformerMotionGPTT2M\-GPTMrSQMCastorMotionGPT\-30\.50\.5AUROC→\\rightarrow±\\pm0\.1±\\pm0\.5±\\pm1±\\pm2\.50\.60\.60\.70\.70\.50\.5Macro\-F1F\_\{1\}→\\rightarrowwindow \(s\)chanceFigure 2:Mean Test AUROC \(left\) and macro\-F1F\_\{1\}\(right\) across the four temporal windows\. Markers show window\-specific scores and grey bars denote bootstrap 95% confidence intervals of the window\-averaged score\.### 6\.1Model Comparison
Across the four coarse temporal windows in the speaker\-only condition, the strongest models consistently distinguish negation\-cue\-centered windows from matched control windows \(Table[4](https://arxiv.org/html/2609.16396#A2.T4); Figure[2](https://arxiv.org/html/2609.16396#S6.F2)\)\.RocketPFNachieves the highest mean AUROC \(\.728\), followed closely byHIVE\-COTE 2\.0\(\.723\) andDrCIF\(\.722\)\. Validation and test rankings are closely aligned \(Table[5](https://arxiv.org/html/2609.16396#A2.T5)\)\. Macro\-F1F\_\{1\}yields nearly the same ordering \(Spearmanρ=\.97\\rho=\.97\), withHIVE\-COTE 2\.0\(\.666\) andRocketPFN\(\.665\) effectively tied at the top\. The larger neural sequence models, including the event Transformer and the motion\-language adaptations, perform substantially worse than the strongest time\-series probes and exhibit pronounced train\-validation gaps, indicating overfitting under the available data regime\. Several models achieve comparable performance, indicating that cue\-centered discriminability is not specific to a particular classifier\. We useRocketPFNas a computationally efficient representative probe for the remaining analyses\. Across models, the±0\.1\\pm 0\.1s context performs worst, whereas the three wider windows yield similar AUROC \(\.660\-\.663\)\. This motivates the finer\-grained temporal analysis in Section[6\.2](https://arxiv.org/html/2609.16396#S6.SS2)\.
### 6\.2Temporal Profile of Cue\-Centered Discriminability
−\-5−\-2\.5−\-1\.5−\-0\.5−\-0\.10\.10\.51\.52\.55\.50\.50\.55\.55\.60\.60\.65\.65\.70\.70\.75\.75\.80\.80before onsetafter onsetchancewindow extent relative to cue onset \(s\)AUROCspeakerlistenerbothFigure 3:Cue\-centered discrimination AUROC across pre\- and post\-onset one\-sided windows and interactional sources\. Lines show means over ten folds; shading denotes bootstrap 95% CIs\.\.50\.50\.60\.60\.70\.70speaker\.50\.50\.60\.60\.70\.70listenerAUROC−2\.5\-2\.5−1\.5\-1\.5−0\.5\-0\.50\.50\.51\.51\.52\.52\.5\.50\.50\.60\.60\.70\.70bothtime relative to cue onset \(s\)Figure 4:AUROC in non\-overlapping 500,ms windows around cue onset across interactional sources\. Whiskers denote bootstrap 95% CIs\.We next examine how discriminability changes with temporal context across speaker\-only, partner\-only, and dyadic conditions\. We focus here on the speaker temporal profile and analyze differences between interactional sources in Section[7\.2](https://arxiv.org/html/2609.16396#S7.SS2)\. To characterize where cue\-associated predictive information occurs relative to cue onset, we re\-evaluateRocketPFNon one\-sided windows with extentst∈\{\.1,\.25,\.5,1,1\.5,2,2\.5,5\}t\\in\\\{\.1,\.25,\.5,1,1\.5,2,2\.5,5\\\}s, taken either before cue onset,\[−t,0\]\[\-t,0\], or after it,\[0,t\]\[0,t\]\(Table[6](https://arxiv.org/html/2609.16396#A2.T6); Figure[3](https://arxiv.org/html/2609.16396#S6.F3)\)\. Because larger windows cumulatively add events, the curves reflect the amount of usable evidence available up to each extent rather than the contribution of individual temporal segments\. For speaker behavior, discriminability is already above chance before cue onset: AUROC reaches\.660\.660\(95% CI\[\.619,\.698\]\[\.619,\.698\]\) for\[−0\.1,0\]\[\-0\.1,0\]s and increases to\.742\.742\(95% CI\[\.704,\.777\]\[\.704,\.777\]\) at2\.52\.5s\. Thus, behavior preceding the spoken cue alone is sufficient to distinguish cue\-centered from matched control contexts\. Post\-onset speaker windows show a similar pattern, rising from\.678\.678to a maximum of approximately\.73\.73\. Matched pre\- and post\-onset windows differ only modestly across temporal extents \(pairedΔAUROC\\Delta\\mathrm\{AUROC\}ranging from−\.018\-\.018to\.035\.035\), with intervals including zero except att=5t=5s, where a pre\-onset advantage emerges \(ΔAUROC=\.035\\Delta\\mathrm\{AUROC\}=\.035, 95% CI\[\.010,\.061\]\[\.010,\.061\]\)\. Performance otherwise largely saturates between1\.51\.5and2\.52\.5s\. Neither side improves further when extending the window to55s\. To separate temporal location from the amount of accumulated evidence, we re\-run the probe on ten disjoint0\.50\.5s windows tiling\[−2\.5,2\.5\]\[\-2\.5,2\.5\]s \(Figure[4](https://arxiv.org/html/2609.16396#S6.F4)\)\. Speaker performance peaks in the two windows adjacent to cue onset and declines steadily with temporal distance \(ρ=−\.96\\rho=\-\.96\), from approximately\.710\.710near onset to\.643\.643at±2\.25\\pm 2\.25s\. Nevertheless, the bootstrap 95% confidence interval lies entirely aboveAUROC=\.5\\mathrm\{AUROC\}=\.5in every window, including\[−2\.5,−2\]\[\-2\.5,\-2\]s \(AUROC\.648\.648, 95% CI\[\.604,\.691\]\[\.604,\.691\]\), showing that the pre\-cue discriminability is not solely driven by behavior immediately adjacent to cue onset\. The two onset\-adjacent windows show similar discriminability, withΔAUROC=\.010\\Delta\\mathrm\{AUROC\}=\.010for\[0,\.5\]\[0,\.5\]s relative to\[−\.5,0\]\[\-\.5,0\]s \(95% CI\[−\.007,\.024\]\[\-\.007,\.024\]\), indicating that performance is strongest around cue onset without a clear pre/post asymmetry at this temporal resolution\.
## 7Analysis
Having established cue\-centered discriminability and its temporal profile, we next assess its robustness to stricter control matching and investigate its interactional source, modality contributions, and temporal sensitivity\.
### 7\.1Robustness to Alternative Control Matching
Our primary controls match speaker identity and local conversational context but may differ from negation cues in grammatical category or turn position\. We therefore construct stricter controls additionally matched by part of speech and turn\-relative position, relaxing the same\-chunk requirement and retaining 830 of 964 cues\. On the same anchors, we also evaluate weaker controls sampled randomly outside negation cues and scopes\. At the symmetric±1\\pm 1s context,RocketPFNreaches mean held\-out AUROC of0\.7230\.723,0\.7120\.712, and0\.7220\.722under stricter, default, and weaker controls, respectively\. Relative to the default condition, neither stricter matching \(Δ\\DeltaAUROC=\+0\.012=\+0\.012, 95% CI\[−0\.007,0\.030\]\[\-0\.007,0\.030\]\) nor weaker matching \(\+0\.010\+0\.010,\[−0\.019,0\.038\]\[\-0\.019,0\.038\]\) reliably changes performance, indicating that additionally matching part of speech and turn position does not substantially reduce discriminability\. We next test whether discriminability primarily reflects the broader context of a negated expression\. We sample negative anchors from within the annotated scope of the corresponding cue, requiring at least1\.11\.1s separation from cue onset\. This yields 311 cues\. On exactly these anchors, performance drops from0\.7300\.730with the default controls to0\.6770\.677with within\-scope controls \(Δ\\DeltaAUROC=−0\.053=\-0\.053, 95% CI\[−0\.096,−0\.012\]\[\-0\.096,\-0\.012\]\)\. Thus, matching within the same negation scope reduces but does not eliminate discriminability: cue\-centered behavior remains distinguishable from behavior centered on another word within the same scope\.
### 7\.2Speaker and Dialogue\-Partner Contributions
We first compare speaker\-only, partner\-only, and dyadic contexts to determine how predictive information is distributed across participants \(Figures[3](https://arxiv.org/html/2609.16396#S6.F3)and[4](https://arxiv.org/html/2609.16396#S6.F4)\)\. Across the cumulative windows, speaker behavior is consistently more informative than partner behavior, averaging AUROC\.712\.712versus\.614\.614\(pairedΔAUROC=\.098\\Delta\\mathrm\{AUROC\}=\.098, 95% CI\[\.063,\.139\]\[\.063,\.139\]\), with higher speaker performance in all 16 temporal conditions\. Partner\-only performance nevertheless yields mean AUROC above\.5\.5throughout, with a minimum of\.596\.596\(95% CI\[\.558,\.628\]\[\.558,\.628\]\), indicating that the interlocutor’s behavior alone carries information predictive of whether the speaker’s anchor is a negation cue\. The sliding\-window analysis reveals that the partner temporal profile differs qualitatively from the speaker profile\. While speaker performance peaks around cue onset, partner performance remains nearly constant across\[−2\.5,2\.5\]\[\-2\.5,2\.5\]s \(ρ=−\.07\\rho=\-\.07; AUROC\.591\.591\-\.622\.622\), with mean AUROC exceeding\.5\.5in all ten windows\. Thus, the partner temporal profile is more consistent with broader interactional or contextual information associated with cue occurrence than with a response time\-locked to the cue\. The dyadic condition performs worst in both analyses \(AUROC\.575\.575cumulative;\.546\.546sliding\-window\)\. This likely reflects the fixed\-grid representation, which averages participants and removes actor identity\.
### 7\.3Modality Contributions
00\.025\.025\.05\.05\.075\.075Facial\(131\)↪\\hookrightarrowlower face\(78\)Head\(14\)Eye\(31\)Body\(14\)Right hand\(14\)Left hand\(14\)Right finger\(240\)Left finger\(240\)performance drop when modality removedAUROCmacro\-F1F\_\{1\}Figure 5:Leave\-one\-modality\-out performance decrease atwL=wR=1w\_\{L\}=w\_\{R\}=1,s\. The indented lower\-face row is a nested facial ablation; whiskers denote bootstrap 95% CIs\.For the modality and temporal\-robustness analyses, we fix the context towL=wR=1w\_\{L\}=w\_\{R\}=1s, an intermediate window within the performance plateau observed above\. We estimate modality\-specific predictive contributions using leave\-one\-modality\-out ablation\. For each modality, we retrain and evaluate theRocketPFNprobe with that modality excluded from both training and test inputs, while retaining identical cross\-validation splits across conditions \(Figure[5](https://arxiv.org/html/2609.16396#S7.F5)\)\. We defineΔm=mfull−mablated\\Delta m=m\_\{\\mathrm\{full\}\}\-m\_\{\\mathrm\{ablated\}\}; positive values therefore indicate performance lost when a modality is unavailable, rather than causal necessity or uniquely attributable information\. The full\-input probe reaches AUROC\.734\.734and macro\-F1F\_\{1\}\.667\.667\. Facial behavior is the only modality whose removal reliably degrades the probe \(ΔAUROC=\.061\\Delta\\mathrm\{AUROC\}=\.061, 95% CI\[\.044,\.084\]\[\.044,\.084\];Δ\\Deltamacro\-F1=\.055F\_\{1\}=\.055\)\. To assess whether this effect is driven primarily by visual articulation of the cue word, we additionally remove the mouth\-, lip\-, chin\-, cheek\-, and jaw\-related lower\-face subset defined in Appendix[B\.1](https://arxiv.org/html/2609.16396#A2.SS1)\. This produces a smaller but reliable performance decrease \(ΔAUROC=\.026\\Delta\\mathrm\{AUROC\}=\.026, 95% CI\[\.016,\.037\]\[\.016,\.037\]\)\. Thus, lower\-face behavior contributes predictive information, but the selected articulatory features account for only part of the facial contribution and only part of overall cue\-centered discriminability: substantial discriminability remains when lower\-face features are unavailable, and even when the entire facial modality is removed\. No other modality shifts AUROC by more than\.009\.009, and all remaining intervals overlap zero except left\-finger ablation, which marginally improves performance \(95% CI\[−\.016,−\.001\]\[\-\.016,\-\.001\]\), consistent with the sparse coverage of the finger streams \(6060\-65%65\\%of windows absent\)\.
### 7\.4Temporal Event Robustness
shift \(s\)time scale−\.129\-\.129\+\.129\+\.129×\.8\\times\.8×1\.2\\times 1\.2ΔF1\\Delta F\_\{1\}\|Δp\|\|\\Delta p\|All modalities\.030∗\.034∗\.031∗\.051∗\.061\.076Facial\.022\.025∗\.022∗\.024∗\.020\.043Eye\.003\.003\.004\.002−\-\.004\.017Right hand\.005\.008∗\.008\.014∗\.005\.028Head\.003\.006∗\.003\.005−\-\.001\.015Body\.004∗\.004\.006∗\.008∗\.001\.022Left hand\.003\.000\.002\.005−\-\.000\.019Right finger\.000−\-\.002−\-\.001\.006\.001\.014Left finger−\-\.002−\-\.001−\-\.000\.001−\-\.004\.011AUROC decrease under timing warpFigure 6:Timing\-warp robustness of the frozenRocketPFNprobe\. Cells show meanΔ\\DeltaAUROC over ten folds for joint and modality\-specific temporal perturbations; final columns reportΔ\\Deltamacro\-F1 and mean absolute prediction change\|Δp\|\|\\Delta p\|\.The ablation analysis identifies which modalities contribute predictive information, but not whether the trained probe is sensitive to their temporal organization\. We therefore perturb the frozenRocketPFNcheckpoints at inference time using rigid shifts of±2\\pm 2grid steps \(±\.129\\pm\.129s\) and temporal rescaling by\.8\.8or1\.21\.2around the window centre\. Warps are applied either jointly to all modalities or to one modality in isolation \(Figure[6](https://arxiv.org/html/2609.16396#S7.F6)\)\. Joint warping reduces AUROC by\.030\.030\-\.051\.051from the\.734\.734baseline\. Rescaling is more damaging than rigid shifting \(ΔAUROC=\.009\\Delta\\mathrm\{AUROC\}=\.009, 95% CI\[\.001,\.015\]\[\.001,\.015\]\), and dilation more damaging than compression \(\.020\.020, 95% CI\[\.008,\.033\]\[\.008,\.033\]\), indicating sensitivity to changes in temporal organization\. Single\-modality effects are smaller and less consistent, led by facial \(ΔAUROC=\.023\\Delta\\mathrm\{AUROC\}=\.023\) and right\-hand \(\.009\.009\) timing, and are therefore treated as descriptive\.
## 8Conclusion
We investigated whether and to what extent explicit spoken negation cues are associated with distinguishable patterns of multimodal behavior during human\-human dialogue\. Using lexical and acoustic information only to define temporal anchors, we treated multimodal classification as a predictive probe of gaze, facial, head, body, hand, and finger behavior surrounding negation cues\. For speakers, strong statistical time\-series models reliably distinguish negation\-cue\-centered from matched control contexts using multimodal behavior without lexical or acoustic input, including under stricter POS\- and turn\-position\-matched controls, reaching up to \.750 fold\-mean test AUROC\. Under the harder within\-scope control, performance decreases from \.730 to \.677 on matched anchors \(Δ\\DeltaAUROC=−\.053=\-\.053, 95% CI\[−\.096,−\.012\]\[\-\.096,\-\.012\]\), indicating that broader negation context contributes to, but does not fully account for, cue\-centered discriminability\. Increasing the available context improves discriminability up to approximately2\.52\.5s before and1\.51\.5s after negation cue onset, while the sliding\-window analysis shows that discriminability is highest in the 0\.5 s immediately before and after the cue\. Dialogue\-partner behavior is also predictive, reaching \.634 AUROC, but remains nearly constant across both cumulative temporal extents and disjoint sliding windows, consistent with diffuse interactional or contextual information rather than a response tightly locked to the negation cue\. Facial features produce the largest modality\-ablation effect, but removing only mouth\-, lip\-, chin\-, cheek\-, and jaw\-related lower\-face channels reduces AUROC by just\.026\.026\(from\.734\.734to approximately\.708\.708\), showing that visual articulation contributes to the signal but does not account for it: substantial predictive information remains without these articulatory features\. Finally, perturbing event timing reduces AUROC by \.030\-\.051, showing that the trained probe is sensitive to changes in the temporal organization of its inputs\. Together, these findings show that explicit negation cues in this dialogue setting are embedded in distinguishable and temporally structured patterns of multimodal behavior, characterized by cue\-localized speaker information, weaker and temporally diffuse partner information, and the largest modality\-ablation effect for facial features\.
## Limitations
Our findings should be interpreted within several constraints\. The study is based on 27 human\-human interviews conducted in a single VR survey setting, so it remains unclear to what extent the observed behavioral patterns generalize to face\-to\-face interaction, other conversational tasks, populations, or languages\. In addition, negation cues and their word\-level timestamps are obtained automatically usingD\-NegandCrisperWhisper; residual annotation or alignment errors may therefore affect particularly fine\-grained temporal analyses\. Our formulation focuses on explicitly marked negation cues and does not cover the broader space of implicit negative meaning, disagreement, refusal, or other pragmatically negative constructions\. Finally, the cue inventory is dominated bynicht\(≈80%\\approx 80\\%of cue instances\)\. Our results therefore primarily characterize explicit negation cues under the lexical distribution of this corpus, wherenichtis the predominant realization, and do not establish invariance across different lexical realizations of negation\.
## Ethical Considerations
The underlying VR recordings contain potentially sensitive behavioral data, including facial, gaze, and body\-motion signals\. We therefore do not release the original continuous recordings or metadata that would allow observations to be linked back to individual participants or recording sessions\. The released research data are restricted to the multimodal event windows used for the negation experiments and are stripped of participant identifiers, experiment identifiers, absolute timestamps, and other linkage metadata\. Consequently, individual examples cannot be directly associated with a particular participant, interview, or position within an interview from the released dataset\. This data\-minimization strategy is intended to reduce privacy and re\-identification risks while retaining the information necessary to reproduce the analyses reported in this work\.
## References
- AbuSaleh et al\. \(2026\)Ali AbuSaleh, Leon Hammerla, and Alexander Mehler\. 2026\.[Learning to detect cross\-modal negation: An analysis of latent representations and an attention\-based solution](https://doi.org/10.1109/ICNLP69856.2026.11527861)\.In*2026 8th International Conference on Natural Language Processing \(ICNLP\)*, pages 613–622\.
- Alhamoud et al\. \(2025\)Kumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li, Philip H\.S\. Torr, Yoon Kim, and Marzyeh Ghassemi\. 2025\.Vision\-language models do not understand negation\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, pages 29612–29622\.
- Bai et al\. \(2018\)Shaojie Bai, J\. Zico Kolter, and Vladlen Koltun\. 2018\.[An empirical evaluation of generic convolutional and recurrent networks for sequence modeling](https://arxiv.org/abs/1803.01271)\.*Preprint*, arXiv:1803\.01271\.
- Brown and Kamiya \(2019\)Amanda Brown and Masaaki Kamiya\. 2019\.[Gesture in contexts of scopal ambiguity: Negation and quantification in english](https://doi.org/10.1017/S014271641900016X)\.*Applied Psycholinguistics*, 40\(5\):1141–1172\.
- Carletta et al\. \(2006\)Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner\. 2006\.The ami meeting corpus: A pre\-announcement\.In*Machine Learning for Multimodal Interaction*, pages 28–39, Berlin, Heidelberg\. Springer Berlin Heidelberg\.
- Chapman et al\. \(2001\)Wendy W\. Chapman, Will Bridewell, Paul Hanbury, Gregory F\. Cooper, and Bruce G\. Buchanan\. 2001\.[A simple algorithm for identifying negated findings and diseases in discharge summaries](https://doi.org/10.1006/jbin.2001.1029)\.*Journal of Biomedical Informatics*, 34\(5\):301–310\.
- Clark and Krych \(2004\)Herbert H\. Clark and Meredyth A\. Krych\. 2004\.[Speaking while monitoring addressees for understanding](https://doi.org/10.1016/j.jml.2003.08.004)\.*Journal of Memory and Language*, 50\(1\):62–81\.
- Cüppers and Vreeken \(2026\)Joscha Cüppers and Jilles Vreeken\. 2026\.[In\-context time series classification with random convolutional features](https://arxiv.org/abs/2607.19234)\.*Preprint*, arXiv:2607\.19234\.
- Dempster et al\. \(2021\)Angus Dempster, Daniel F\. Schmidt, and Geoffrey I\. Webb\. 2021\.[Minirocket: A very fast \(almost\) deterministic transform for time series classification](https://doi.org/10.1145/3447548.3467231)\.In*Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining*, KDD ’21, page 248–257, New York, NY, USA\. Association for Computing Machinery\.
- Dempster et al\. \(2023\)Angus Dempster, Daniel F\. Schmidt, and Geoffrey I\. Webb\. 2023\.[Hydra: competing convolutional kernels for fast and accurate time series classification](https://doi.org/10.1007/s10618-023-00939-3)\.*Data Mining and Knowledge Discovery*, 37\(5\):1779–1805\.
- Dobreva and Keller \(2021\)Radina Dobreva and Frank Keller\. 2021\.[Investigating negation in pre\-trained vision\-and\-language models](https://doi.org/10.18653/v1/2021.blackboxnlp-1.27)\.In*Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP*, pages 350–362, Punta Cana, Dominican Republic\. Association for Computational Linguistics\.
- Fancellu et al\. \(2016\)Federico Fancellu, Adam Lopez, and Bonnie Webber\. 2016\.[Neural networks for negation scope detection](https://doi.org/10.18653/v1/P16-1047)\.In*Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 495–504, Berlin, Germany\. Association for Computational Linguistics\.
- García\-Ferrero et al\. \(2023\)Iker García\-Ferrero, Begoña Altuna, Javier Alvez, Itziar Gonzalez\-Dios, and German Rigau\. 2023\.[This is not a dataset: A large negation benchmark to challenge large language models](https://doi.org/10.18653/v1/2023.emnlp-main.531)\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 8596–8615, Singapore\. Association for Computational Linguistics\.
- Goldin\-Meadow and Alibali \(2013\)Susan Goldin\-Meadow and Martha Wagner Alibali\. 2013\.[Gesture’s role in speaking, learning, and creating language](https://doi.org/10.1146/annurev-psych-113011-143802)\.*Annual Review of Psychology*, 64\(1\):257–283\.
- González\-Fuente et al\. \(2015\)Santiago González\-Fuente, Susagna Tubau, Mª Teresa Espinal, and Pilar Prieto\. 2015\.[Is there a universal answering strategy for rejecting negative propositions? typological evidence on the use of prosody and gesture](https://doi.org/10.3389/fpsyg.2015.00899)\.*Frontiers in Psychology*, Volume 6 \- 2015\.
- Guidetti \(2005\)Michéle Guidetti\. 2005\.[Yes or no? how young french children combine gestures and speech to agree and refuse](https://doi.org/10.1017/S0305000905007038)\.*Journal of Child Language*, 32\(4\):911–924\.
- Hammerla et al\. \(2025\)Leon Hammerla, Andy Lücking, Carolin Reinert, and Alexander Mehler\. 2025\.[D\-neg: Syntax\-aware graph reasoning for negation detection](https://doi.org/10.18653/v1/2025.findings-ijcnlp.89)\.In*Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics*, pages 1432–1454, Mumbai, India\. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics\.
- Harris et al\. \(2020\)Charles R\. Harris, K\. Jarrod Millman, Stéfan J\. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J\. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H\. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Río, Mark Wiebe, Pearu Peterson, and 7 others\. 2020\.[Array programming with NumPy](https://doi.org/10.1038/s41586-020-2649-2)\.*Nature*, 585\(7825\):357–362\.
- Harrison \(2010\)Simon Harrison\. 2010\.[Evidence for node and scope of negation in coverbal gesture](https://doi.org/10.1075/gest.10.1.03har)\.*Gesture*, 10\(1\):29–51\.
- Harrison \(2014\)Simon Harrison\. 2014\.[The organisation of kinesic ensembles associated with negation](https://doi.org/10.1075/gest.14.2.01har)\.*Gesture*, 14\(2\):117–140\.
- Harrison \(2024\)Simon Harrison\. 2024\.On grammar\-gesture relations: Gestures associated with negation\.In Alan Cienki, editor,*The Cambridge Handbook of Gesture Studies*, Cambridge Handbooks in Language and Linguistics, pages 446–474\. Cambridge University Press, Cambridge\.
- Harrison and Larrivée \(2016\)Simon Harrison and Pierre Larrivée\. 2016\.[Morphosyntactic correlates of gestures: A gesture associated with negation in French and its organisation with speech](https://doi.org/10.1007/978-3-319-17464-8_4)\.In Pierre Larrivée and Chungmin Lee, editors,*Negation and Polarity: Experimental Perspectives*, pages 75–94\. Springer International Publishing, Cham\.
- Hollmann et al\. \(2025\)Noah Hollmann, Samuel Müller, Lennart Purucker, Arjun Krishnakumar, Max Körfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter\. 2025\.[Accurate predictions on small data with a tabular foundation model](https://doi.org/10.1038/s41586-024-08328-6)\.*Nature*, 637\(8045\):319–326\.
- Hossain et al\. \(2022\)Md Mosharaf Hossain, Dhivya Chinnappa, and Eduardo Blanco\. 2022\.[An analysis of negation in natural language understanding corpora](https://doi.org/10.18653/v1/2022.acl-short.81)\.In*Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 716–723, Dublin, Ireland\. Association for Computational Linguistics\.
- Inbar and Shor \(2019\)Anna Inbar and Leon Shor\. 2019\.[Covert negation in israeli hebrew: Evidence from co\-speech gestures](https://doi.org/10.1016/j.pragma.2019.02.011)\.*Journal of Pragmatics*, 143:85–95\.
- Ismail Fawaz et al\. \(2020\)Hassan Ismail Fawaz, Benjamin Lucas, Germain Forestier, Charlotte Pelletier, Daniel F\. Schmidt, Jonathan Weber, Geoffrey I\. Webb, Lhassane Idoumghar, Pierre\-Alain Muller, and François Petitjean\. 2020\.[Inceptiontime: Finding alexnet for time series classification](https://doi.org/10.1007/s10618-020-00710-y)\.*Data Mining and Knowledge Discovery*, 34\(6\):1936–1962\.
- Jiang et al\. \(2023\)Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen\. 2023\.[Motiongpt: Human motion as a foreign language](https://doi.org/10.52202/075280-0880)\.In*Advances in Neural Information Processing Systems*, volume 36, pages 20067–20079\. Curran Associates, Inc\.
- Kassner and Schütze \(2020\)Nora Kassner and Hinrich Schütze\. 2020\.[Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly](https://doi.org/10.18653/v1/2020.acl-main.698)\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 7811–7818, Online\. Association for Computational Linguistics\.
- Kelly et al\. \(2009\)Spencer D\. Kelly, Aslı Özyürek, and Eric Maris\. 2009\.[Two sides of the same coin: Speech and gesture mutually interact to enhance comprehension](https://doi.org/10.1177/0956797609357327)\.*Psychological Science*, 21\(2\):260–267\.
- Kendon \(2002\)Adam Kendon\. 2002\.[Some uses of the head shake](https://doi.org/10.1075/gest.2.2.03ken)\.*Gesture*, 2\(2\):147–182\.
- Khandelwal and Sawant \(2020\)Aditya Khandelwal and Suraj Sawant\. 2020\.[NegBERT: A transfer learning approach for negation detection and scope resolution](https://aclanthology.org/2020.lrec-1.704/)\.In*Proceedings of the Twelfth Language Resources and Evaluation Conference*, pages 5739–5748, Marseille, France\. European Language Resources Association\.
- Kim et al\. \(2025\)Youngmin Kim, Jiwan Chung, Jisoo Kim, Sunghyun Lee, Sangkyu Lee, Junhyeok Kim, Cheoljong Yang, and Youngjae Yu\. 2025\.[Speaking beyond language: A large\-scale multimodal dataset for learning nonverbal cues from video\-grounded dialogues](https://doi.org/10.18653/v1/2025.acl-long.112)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 2247–2265, Vienna, Austria\. Association for Computational Linguistics\.
- Kontogiorgos et al\. \(2018\)Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, and Joakim Gustafson\. 2018\.[A multimodal corpus for mutual gaze and joint attention in multiparty situated interaction](https://aclanthology.org/L18-1019/)\.In*Proceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC 2018\)*, Miyazaki, Japan\. European Language Resources Association \(ELRA\)\.
- Lapponi et al\. \(2012\)Emanuele Lapponi, Erik Velldal, Lilja Øvrelid, and Jonathon Read\. 2012\.[UiO 2: Sequence\-labeling negation using dependency features](https://aclanthology.org/S12-1042/)\.In*\*SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation \(SemEval 2012\)*, pages 319–327, Montréal, Canada\. Association for Computational Linguistics\.
- Leonard and Cummins \(2011\)Thomas Leonard and Fred Cummins\. 2011\.[The temporal relation between beat gestures and speech](https://doi.org/10.1080/01690965.2010.500218)\.*Language and Cognitive Processes*, 26\(10\):1457–1471\.
- Lo et al\. \(2026\)Mouhamadou Mansour Lo, Gildas Morvan, Mathieu Rossi, Fabrice Morganti, and David Mercier\. 2026\.[Time series classification with random convolution kernels: pooling operators and input representations matter](https://arxiv.org/abs/2409.01115)\.*Preprint*, arXiv:2409\.01115\.
- Middlehurst et al\. \(2024\)Matthew Middlehurst, Ali Ismail\-Fawaz, Antoine Guillaume, Christopher Holder, David Guijo\-Rubio, Guzal Bulatova, Leonidas Tsaprounis, Lukasz Mentel, Martin Walter, Patrick Schäfer, and Anthony Bagnall\. 2024\.[aeon: a python toolkit for learning from time series](http://jmlr.org/papers/v25/23-1444.html)\.*Journal of Machine Learning Research*, 25\(289\):1–10\.
- Middlehurst et al\. \(2021\)Matthew Middlehurst, James Large, Michael Flynn, Jason Lines, Aaron Bostrom, and Anthony Bagnall\. 2021\.[Hive\-cote 2\.0: a new meta ensemble for time series classification](https://doi.org/10.1007/s10994-021-06057-9)\.*Machine Learning*, 110\(11\-12\):3211–3243\.
- Morante and Blanco \(2012\)Roser Morante and Eduardo Blanco\. 2012\.[\*SEM 2012 shared task: Resolving the scope and focus of negation](https://aclanthology.org/S12-1035/)\.In*\*SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation \(SemEval 2012\)*, pages 265–274, Montréal, Canada\. Association for Computational Linguistics\.
- Morante and Daelemans \(2009\)Roser Morante and Walter Daelemans\. 2009\.[A metalearning approach to processing the scope of negation](https://aclanthology.org/W09-1105/)\.In*Proceedings of the Thirteenth Conference on Computational Natural Language Learning \(CoNLL\-2009\)*, pages 21–29, Boulder, Colorado\. Association for Computational Linguistics\.
- Morante and Sporleder \(2012\)Roser Morante and Caroline Sporleder\. 2012\.[Modality and negation: An introduction to the special issue](https://doi.org/10.1162/COLI_a_00095)\.*Computational Linguistics*, 38\(2\):223–260\.
- Nguyen and Ifrim \(2023\)Thach Le Nguyen and Georgiana Ifrim\. 2023\.Fast time series classification with random symbolic subsequences\.In*Advanced Analytics and Learning on Temporal Data*, pages 50–65, Cham\. Springer International Publishing\.
- O’Rourke et al\. \(2026\)Franco Martino O’Rourke, Ana Trisovic, and Dimitris Bertsimas\. 2026\.[Rocketpfn: Accurate time series classification via in\-context learning](https://arxiv.org/abs/2606.21786)\.*Preprint*, arXiv:2606\.21786\.
- Parcalabescu et al\. \(2022\)Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, and Albert Gatt\. 2022\.[VALSE: A task\-independent benchmark for vision and language models centered on linguistic phenomena](https://doi.org/10.18653/v1/2022.acl-long.567)\.In*Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 8253–8280, Dublin, Ireland\. Association for Computational Linguistics\.
- Park et al\. \(2025\)Junsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu, Dahuin Jung, and Sungroh Yoon\. 2025\.Know ”no” better: A data\-driven approach for enhancing negation awareness in clip\.In*Proceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 2825–2835\.
- Paszke et al\. \(2019\)Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, and 2 others\. 2019\.[Pytorch: An imperative style, high\-performance deep learning library](https://proceedings.neurips.cc/paper_files/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf)\.In*Advances in Neural Information Processing Systems*, volume 32\. Curran Associates, Inc\.
- Pedregosa et al\. \(2011\)Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay\. 2011\.[Scikit\-learn: Machine learning in python](http://jmlr.org/papers/v12/pedregosa11a.html)\.*Journal of Machine Learning Research*, 12\(85\):2825–2830\.
- Peng et al\. \(2018\)Yifan Peng, Xiaosong Wang, Le Lu, Mohammadhadi Bagheri, Ronald Summers, and Zhiyong Lu\. 2018\.[Negbio: a high\-performance tool for negation and uncertainty detection in radiology reports](https://pubmed.ncbi.nlm.nih.gov/29888070/)\.*AMIA Joint Summits on Translational Science proceedings\. AMIA Joint Summits on Translational Science*, 2017:188–196\.
- Prieto et al\. \(2013\)Pilar Prieto, Joan Borràs\-Comes, Susagna Tubau, and M\. Teresa Espinal\. 2013\.[Prosody and gesture constrain the interpretation of double negation](https://doi.org/10.1016/j.lingua.2013.02.008)\.*Lingua*, 131:136–150\.
- Quantmeyer et al\. \(2024\)Vincent Quantmeyer, Pablo Mosteiro, and Albert Gatt\. 2024\.[How and where does CLIP process negation?](https://doi.org/10.18653/v1/2024.alvr-1.5)In*Proceedings of the 3rd Workshop on Advances in Language and Vision Research \(ALVR\)*, pages 59–72, Bangkok, Thailand\. Association for Computational Linguistics\.
- Richardson and Dale \(2005\)Daniel C\. Richardson and Rick Dale\. 2005\.[Looking to understand: The coupling between speakers’ and listeners’ eye movements and its relationship to discourse comprehension](https://doi.org/10.1207/s15516709cog0000_29)\.*Cognitive Science*, 29\(6\):1045–1060\.
- Richardson et al\. \(2007\)Daniel C\. Richardson, Rick Dale, and Natasha Z\. Kirkham\. 2007\.[The art of conversation is coordination](https://doi.org/10.1111/j.1467-9280.2007.01914.x)\.*Psychological Science*, 18\(5\):407–413\.PMID: 17576280\.
- Samsten and Lee \(2024\)Isak Samsten and Zed Lee\. 2024\.[Castor: Competing shapelets for fast and accurate time series classification](https://arxiv.org/abs/2403.13176)\.*Preprint*, arXiv:2403\.13176\.
- Sato et al\. \(2023\)Yuri Sato, Koji Mineshima, and Kazuhiro Ueda\. 2023\.[Can negation be depicted? comparing human and machine understanding of visual representations](https://doi.org/10.1111/cogs.13258)\.*Cognitive Science*, 47\(3\):e13258\.
- Schäfer and Leser \(2023\)Patrick Schäfer and Ulf Leser\. 2023\.[Weasel 2\.0: a random dilated dictionary transform for fast, accurate and memory constrained time series classification](https://doi.org/10.1007/s10994-023-06395-w)\.*Machine Learning*, 112\(12\):4763–4788\.
- SurrealDB Ltd\. \(2026\)SurrealDB Ltd\. 2026\.Surrealdb\.[https://github\.com/surrealdb/surrealdb](https://github.com/surrealdb/surrealdb)\.Version 3\.0\.
- Szarvas et al\. \(2008\)György Szarvas, Veronika Vincze, Richárd Farkas, and János Csirik\. 2008\.[The BioScope corpus: annotation for negation, uncertainty and their scope in biomedical texts](https://aclanthology.org/W08-0606/)\.In*Proceedings of the Workshop on Current Trends in Biomedical Natural Language Processing*, pages 38–45, Columbus, Ohio\. Association for Computational Linguistics\.
- Tan et al\. \(2022\)Chang Wei Tan, Angus Dempster, Christoph Bergmeir, and Geoffrey I\. Webb\. 2022\.[Multirocket: multiple pooling operators and transformations for fast and effective time series classification](https://doi.org/10.1007/s10618-022-00844-1)\.*Data Mining and Knowledge Discovery*, 36\(5\):1623–1646\.
- Truong et al\. \(2023\)Thinh Hung Truong, Timothy Baldwin, Karin Verspoor, and Trevor Cohn\. 2023\.[Language models are not naysayers: an analysis of language models on negation benchmarks](https://doi.org/10.18653/v1/2023.starsem-1.10)\.In*Proceedings of the 12th Joint Conference on Lexical and Computational Semantics \(\*SEM 2023\)*, pages 101–114, Toronto, Canada\. Association for Computational Linguistics\.
- Tsai et al\. \(2019\)Yao\-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J\. Zico Kolter, Louis\-Philippe Morency, and Ruslan Salakhutdinov\. 2019\.[Multimodal transformer for unaligned multimodal language sequences](https://doi.org/10.18653/v1/P19-1656)\.In*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics*, pages 6558–6569, Florence, Italy\. Association for Computational Linguistics\.
- Tubau et al\. \(2015\)Susagna Tubau, Santiago González\-Fuente, Pilar Prieto, and Maria Teresa Espinal\. 2015\.[Prosody and gesture in the interpretation of yes\-answers to negative yes/no\-questions](https://doi.org/doi:10.1515/tlr-2014-0016)\.*The Linguistic Review*, 32\(1\):115–142\.
- Virtanen et al\. \(2020\)Pauli Virtanen, Ralf Gommers, Travis E\. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J\. van der Walt, Matthew Brett, Joshua Wilson, K\. Jarrod Millman, Nikolay Mayorov, Andrew R\. J\. Nelson, Eric Jones, Robert Kern, Eric Larson, and 16 others\. 2020\.[SciPy 1\.0: Fundamental Algorithms for Scientific Computing in Python](https://doi.org/10.1038/s41592-019-0686-2)\.*Nature Methods*, 17:261–272\.
- Wang et al\. \(2022\)Ziyue Wang, Aozhu Chen, Fan Hu, and Xirong Li\. 2022\.[Learn to understand negation in video retrieval](https://doi.org/10.1145/3503161.3547968)\.In*Proceedings of the 30th ACM International Conference on Multimedia*, MM ’22, page 434–443, New York, NY, USA\. Association for Computing Machinery\.
- Wen et al\. \(2025\)Yilin Wen, Hao Pan, Takehiko Ohkawa, Lei Yang, Jia Pan, Yoichi Sato, Taku Komura, and Wenping Wang\. 2025\.Generative hierarchical temporal transformer for hand pose and action modeling\.In*Computer Vision – ECCV 2024 Workshops*, pages 49–67, Cham\. Springer Nature Switzerland\.
- Wolf et al\. \(2020\)Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, and 3 others\. 2020\.[Transformers: State\-of\-the\-art natural language processing](https://doi.org/10.18653/v1/2020.emnlp-demos.6)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pages 38–45, Online\. Association for Computational Linguistics\.
- Zhang et al\. \(2023\)Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang, Hongwei Zhao, Hongtao Lu, Xi Shen, and Ying Shan\. 2023\.Generating human motion from textual descriptions with discrete representations\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, pages 14730–14740\.
- Zhu et al\. \(2026\)Bingfan Zhu, Biao Jiang, Sunyi Wang, SHIXIANG TANG, Tao Chen, Linjie Luo, Youyi Zheng, and Xin Chen\. 2026\.[Motiongpt3: Human motion as a second modality](https://proceedings.iclr.cc/paper_files/paper/2026/file/1551c01d7a3d0bf21e2518331e9f7074-Paper-Conference.pdf)\.In*International Conference on Learning Representations*, volume 2026, pages 12631–12657\.
- Zusag et al\. \(2024\)Mario Zusag, Laurin Wagner, and Bernhad Thallinger\. 2024\.[Crisperwhisper: Accurate timestamps on verbatim speech transcriptions](https://doi.org/10.21437/interspeech.2024-731)\.In*Interspeech 2024*, interspeech 2024, page 1265–1269\. ISCA\.
## Appendix AAnnotation Model Configurations
#### CrisperWhisper
We usenyrahealth/CrisperWhisperwith German as the decoding language and beam search with five beams\. The resulting word\-level timestamps are retained for alignment with the multimodal event stream\.
#### D\-Neg
Negation cues and scopes are annotated usingD\-Negv0\.1\.1\([Hammerla et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib17)\)with the German syntax\-aware configuration\. We use theD\-NEG/cue\-gat\-de\-sfuandD\-NEG/scope\-gat\-de\-sfucheckpoints for cue detection and scope resolution, respectively, with a maximum sequence length of 256\. German syntactic annotations are produced with spaCy’sde\_core\_news\_smmodel v3\.8\.0\. All remaining inference parameters use the respective package defaults\.
## Appendix BMain Model Configurations
### B\.1Multimodal Feature Inventory
Table 1:Fixed\-grid channel allocation\. Global indices are zero\-based positions along the first axis of𝐗\\mathbf\{X\}\. Recorded dimensions contain gaze measurements, facial blendshape weights and tracking metadata, rigid\-body poses, and finger\-tracking measurements\. Derived dimensions contain first\-order linear, angular, blendshape, bone, and pinch motion features as applicable\. The eye modality corresponds to the dedicated gaze\-tracking stream; ocular expression blendshapes supplied by the facial tracker remain part of the facial modality\.ModalityLocal channel index or nameFeature mappingEye00;99Left\- and right\-eye gaze confidence, respectively\.Eye11;1010Left\- and right\-eye gaze\-validity indicators\.Eye22–44;1111–1313Left\- and right\-eye gaze positions\(x,y,z\)\(x,y,z\)\.Eye55–88;1414–1717Left\- and right\-eye gaze orientations as quaternions\(x,y,z,w\)\(x,y,z,w\)\.Eye1818–2020;2424–2626Left\- and right\-eye linear velocities\(vx,vy,vz\)\(v\_\{x\},v\_\{y\},v\_\{z\}\)\.Eye2121–2323;2727–2929Left\- and right\-eye angular velocities\(ωx,ωy,ωz\)\(\\omega\_\{x\},\\omega\_\{y\},\\omega\_\{z\}\)\.Facial00–6262The 63 facial expression weights enumerated in Table[3](https://arxiv.org/html/2609.16396#A2.T3)\.Facial6363–6464The two entries of the recordedexpressionWeightConfidencesarray, in stored order\.Facial6565The facial\-trackingstatus\.IsValidindicator\.Facial6666Thestatus\.IsEyeFollowingBlendshapesValidindicator\.Facial6767–129129Blendshape velocities: local facial index67\+i67\+iis the time derivative of expression weightii, fori=0,…,62i=0,\\ldots,62\.Head, body, left hand, right hand00–22Position\(x,y,z\)\(x,y,z\)\.Head, body, left hand, right hand33–66Orientation quaternion\(x,y,z,w\)\(x,y,z,w\)\.Head, body, left hand, right hand77–99Linear velocity\(vx,vy,vz\)\(v\_\{x\},v\_\{y\},v\_\{z\}\)\.Head, body, left hand, right hand1010–1212Angular velocity\(ωx,ωy,ωz\)\(\\omega\_\{x\},\\omega\_\{y\},\\omega\_\{z\}\)\.Left/right finger00;11Hand confidence and hand scale, respectively\.Left/right finger22–88Pointer pose: position\(x,y,z\)\(x,y,z\)followed by quaternion\(x,y,z,w\)\(x,y,z,w\)\.Left/right finger99–1515Root pose: position\(x,y,z\)\(x,y,z\)followed by quaternion\(x,y,z,w\)\(x,y,z,w\)\.Left/right finger16\+4b16\+4b–19\+4b19\+4b,b=0,…,25b=0,\\ldots,25Quaternion\(x,y,z,w\)\(x,y,z,w\)forboneRotations\[b\]; the source\-array order is preserved\.Left/right finger120120–124124The fivefingerConfidencesentries, in stored order\.Left/right finger125125–129129The fivepinchStrengthentries, in stored order\.Left/right finger130130Hand\-scale velocity\.Left/right finger131131–136136Pointer linear velocity\(vx,vy,vz\)\(v\_\{x\},v\_\{y\},v\_\{z\}\)followed by angular velocity\(ωx,ωy,ωz\)\(\\omega\_\{x\},\\omega\_\{y\},\\omega\_\{z\}\)\.Left/right finger137137–142142Root linear velocity followed by angular velocity, in the same order\.Left/right finger143\+3b143\+3b–145\+3b145\+3b,b=0,…,25b=0,\\ldots,25Angular velocity\(ωx,ωy,ωz\)\(\\omega\_\{x\},\\omega\_\{y\},\\omega\_\{z\}\)ofboneRotations\[b\]\.Left/right finger221221–225225Time derivatives of the five pinch strengths\.Left/right fingerflag\_0–flag\_7Bits00–77of the recorded hand\-status bitmask, decoded as separate binary indicators\.Left/right fingerflag\_8–flag\_12Bits00–44of the recorded pinch bitmask, decoded as separate binary indicators\.Every modalitypresentOne inside the temporally observed/interpolated support of that modality and zero elsewhere\.Table 2:Zero\-based local feature mappings within each modality\. Unless a row explicitly names a flag or presence channel, its entries are indicesiiinM\.feature\_i\. For rows containing two index ranges, the first and second ranges refer to the left and right eye, respectively\. Finger bone and finger\-level arrays retain the order in which they are stored in the source events\.Table 3:Mapping of the 63 facial expression\-weight indices\. For every listed indexii, the facial channel with local index67\+i67\+icontains the corresponding blendshape velocity\. A dagger marks expression weights included in the targeted lower\-face ablation\.Table 4:Cue\-centered discrimination from speaker\-side multimodal events: mean test performance over 10 cross\-validation folds, for 20 time\-series classifiers×\\times4 temporal windows \(s around word onset\), with the half\-width of the bootstrap 95% confidence interval below each value\.AUROC\#ModelVal\.TestTest \#1RocketPFN\.728\.72812HIVE\-COTE 2\.723\.72323DrCIF\.722\.72234MultiRocket\.708\.70745MASHT\.695\.69556SelfRocket\.695\.69167MiniRocket\.689\.69078AdaptiveRocketPFN\.688\.68989WEASEL 2\.678\.670910Inception\-TCN\.669\.65612↓\\downarrow211GHTT\.667\.66310↑\\uparrow112T2M\-GPT 2\.663\.63614↓\\downarrow213SparseMultiRocket\-Hydra\.661\.65911↑\\uparrow214EventTransformer\.654\.62815↓\\downarrow115CompactFusion\-TCN\.643\.63713↑\\uparrow216T2M\-GPT\.639\.61417↓\\downarrow117MotionGPT\.635\.62216↑\\uparrow118MrSQM\.604\.6081819CASTOR\.596\.5901920MotionGPT\-3\.506\.48020Table 5:Model ranking under validation and test AUROC, averaged over the four symmetric windows and ordered by validation rank\. The two orderings agree closely \(Spearmanρ=\.985\\rho=\.985, Kendallτ=\.937\\tau=\.937; 6 of 190 pairs discordant\)\. Ranks 1–9 and 18–20 are identical across splits; only the eight models in the middle block \(bold\) change rank, each by at most two positions and only between models separated by≤\.024\\leq\.024AUROC\. Arrows give the rank change from validation to test\.Table 6:Cue\-centered discrimination performance of the ROCKETPFN probe across one\-sided observation windows and interactional sources\. Columns give the window extentttin seconds; each value is the mean over ten test folds, with the half\-width of its bootstrap 95% confidence interval below\. Best window per row in bold\.We report the configurations used for all comparison models\. Unless stated otherwise, the same configuration is used across temporal contexts, modality analyses, and interactional\-source conditions\. No hyperparameter is selected using a test fold\.
### B\.2Shared Input Processing
Fixed\-grid classifiers receive the same modality\-by\-time representation\. For each window, at most 128 observations per modality are retained using uniform temporal subsampling\. Duplicate timestamps are averaged, and each continuous channel is linearly interpolated at 32 equidistant time points between the left and right window boundaries\. Interpolation is restricted to the interval between the first and last observation of a stream, and positions outside this support are set to zero\. A stream containing only one observation is placed at its nearest grid position\. A separate presence channel indicates the temporal support of each modality\. Numeric event features are standardized using the mean and population standard deviation estimated exclusively from the training split of the current fold\. The 13 finger\-status and pinch indicators per finger modality and the modality\-presence channels are not standardized\. The resulting input has shape𝐗∈ℝ698×32\\mathbf\{X\}\\in\\mathbb\{R\}^\{698\\times 32\}\. Channels are concatenated in the modality order shown in Table[1](https://arxiv.org/html/2609.16396#A2.T1)\. Within a modalityMM, continuous channels are namedM\.feature\_i, whereiiis the zero\-based local index defined in Table[2](https://arxiv.org/html/2609.16396#A2.T2)\. Finger bit indicators are namedM\.flag\_i, and the support indicator is namedM\.present\. Positions and linear velocities use\(x,y,z\)\(x,y,z\)component order\. All orientations are normalized quaternions in\(x,y,z,w\)\(x,y,z,w\)order, with the quaternion hemisphere fixed so thatw≥0w\\geq 0\. Angular velocity is computed from the shortest\-path relative quaternion and is represented by its three\-axis rotation vector\. Facial and scalar motion features are ordinary first\-order rates\. When no preceding observation is available, the corresponding derived motion features are zero\. The targeted lower\-face sub\-ablation retrains and evaluates the probe with 39 mouth\-, lip\-, jaw\-, cheek\-, and chin\-related expression weights and their 39 corresponding velocity channels excluded from both training and test inputs\. These expression weights are marked by daggers in Table[3](https://arxiv.org/html/2609.16396#A2.T3)\. Equivalently, the excluded raw blendshape indices are2–9,24–27,30–54,61–62\{2\\text\{\-\-\}9,24\\text\{\-\-\}27,30\\text\{\-\-\}54,61\\text\{\-\-\}62\}, and the excluded local facial velocity indices are69–76,91–94,97–121,128–129\{69\\text\{\-\-\}76,91\\text\{\-\-\}94,97\\text\{\-\-\}121,128\\text\{\-\-\}129\}\. Thus, 78 channels are excluded in total\. The two facial confidence channels, both facial\-validity channels, and the facial presence channel remain available in this ablation\. For the speaker\-only, partner\-only, and dyadic conditions, the corresponding actor filters areanchor,other, andall\. In the dyadic fixed\-grid representation, simultaneous observations from both participants within the same modality are averaged\. The irregular\-event Transformer does not use the fixed\-grid representation\. It uses the same local continuous feature mappings in Table[2](https://arxiv.org/html/2609.16396#A2.T2)and the same limit of 128 observations per modality, but preserves the resulting irregular event sequence and does not add modality\-presence channels\. The actor of each event is encoded as identical to the anchor speaker, different from the anchor speaker, or unknown\. Observations sharing a timestamp and actor are fused into a single temporal token\. Each token is augmented with its time relative to the anchor and its temporal distance to the preceding and following tokens\. Relative times are clipped to\[−10,10\]\[\-10,10\]s\. Learned anchor and classification tokens are added to the sequence\. Absolute timestamps, participant and experiment identifiers, database identifiers, lexical information, and audio are excluded from both representations\.
### B\.3Cross\-Validation, Optimization, and Decisions
We use ten stratified group partitions, with the recording session as the grouping variable\. For outer foldkk, partitionkkis used as the test set, partition\(k\+1\)mod10\(k\+1\)\\bmod 10as the validation set, and the remaining eight partitions as the training set\. Thus, no recording session occurs in more than one split within a fold\. Neural networks use AdamW, binary cross\-entropy for discriminative fine\-tuning, gradient clipping at1\.01\.0, and aReduceLROnPlateaulearning\-rate schedule unless specified otherwise\. Neural checkpoints are selected by minimum validation loss\. AUROC is the primary evaluation metric and is computed from continuous model scores, making it independent of a decision threshold\. For threshold\-dependent metrics, probabilistic classifiers use a threshold of0\.50\.5unless stated otherwise\.MotionGPTinstead selects its decision threshold on the validation split to maximize macro\-F1F\_\{1\}\. No positive\-class weighting is used unless stated otherwise\. The ridge classifiers used byMiniRocket,MultiRocket\+Hydra, andSelF\-Rocketselectα\\alphausing efficient leave\-one\-out ridge cross\-validation with squared\-error scoring\. Their common grid contains 19 half\-decade values,α∈\{10−3,10−2\.5,…,106\}\\alpha\\in\\\{10^\{\-3\},10^\{\-2\.5\},\\ldots,10^\{6\}\\\}\. This selection uses only the outer training split\.
### B\.4Random\-Convolution Models
#### MiniRocket
We use 10,000 nominal kernels, at most 32 dilations per kernel, and four CPU workers\. Transformed features are scaled without mean centering and classified using ridge regression over the commonα\\alphagrid\.
#### MultiRocket\+Hydra
MultiRocketuses 6,250 nominal kernels, at most 32 dilations per kernel, four pooling features per kernel, the native raw and first\-difference representations, and no per\-instance normalization\.Hydrauses eight kernels per group, 64 groups, and at most eight input channels per group\.MultiRocketfeatures are scaled without centering, whileHydracounts use sparse square\-root standardization\. The two branches are concatenated and classified using ridge regression over the commonα\\alphagrid\.
#### SparseMultiRocket
The sparse variant uses the same random transforms as the denseMultiRocket\+Hydramodel\. Within each outer training fold, ANOVAFFscores selectk∈\{5,000,10,000,15,000\}k\\in\\\{5\{,\}000,10\{,\}000,15\{,\}000\\\}features\. Three\-fold stratified inner cross\-validation, scored by balanced accuracy, jointly selectskkand the downstream linear classifier\. Candidate classifiers are ridge regression withα∈\{1,10,100,103,104\}\\alpha\\in\\\{1,10,100,10^\{3\},10^\{4\}\\\},L2L\_\{2\}logistic regression withC∈\{10−3,10−2,10−1,1,10\}C\\in\\\{10^\{\-3\},10^\{\-2\},10^\{\-1\},1,10\\\}, elastic\-net logistic SGD withα∈\{10−3,10−2,10−1\}\\alpha\\in\\\{10^\{\-3\},10^\{\-2\},10^\{\-1\}\\\}andl1l\_\{1\}ratio∈\{0\.2,0\.5,0\.8\}\\in\\\{0\.2,0\.5,0\.8\\\}, or a linear SVM withC∈\{10−3,10−2,10−1,1\}C\\in\\\{10^\{\-3\},10^\{\-2\},10^\{\-1\},1\\\}\. The iteration limit is 2,000 and the convergence tolerance is10−310^\{\-3\}\.
#### SelF\-Rocket
We use 10,000 nominal kernels, at most 32 dilations per kernel, and no per\-instance normalization\. We evaluate all 15 combinations of the base, first\-difference, and mixed representations with PPV, zero\-crossing, MPV, MIPV, and LSPV pooling\. Candidate selection uses two folds repeated ten times, at most 500 cases and 2,500 features per candidate, and ten logarithmically spaced ridge penalties from10−310^\{\-3\}to10310^\{3\}\. The candidate with the highest median validation accuracy is retained if it ranks among the top five in at least 90% of runs\. Otherwise, PPV\-MIX is used for the present sequence length \(32<51232<512\)\. The final ridge classifier uses the commonα\\alphagrid\.
### B\.5Feature\-, Dictionary\-, Shapelet\-, and Interval\-Based Models
#### Castor
We use 128 groups with 16 shapelets per group, shapelet length 9, Euclidean distance, normalization probability0\.50\.5, and lower and upper quantiles of0\.010\.01and0\.200\.20\. Soft minimum and soft threshold are enabled, while soft maximum is disabled\. Class labels are available during shapelet sampling \(ignore\_y=false\)\. Half of the groups operate on the raw series and half on first differences\. Sparse features are square\-root transformed and standardized with zero\-frequency exponent 4\. Classification uses ridge regression withα∈\{0\.01,1,10\}\\alpha\\in\\\{0\.01,1,10\\\}\.
#### Weasel 2\.0
The minimum window length is 4, window normalization is disabled, and word lengths 7 and 8 are used\. Both raw and first\-difference variants are included\. Chi\-squared top\-kkselection retains at most 30,000 features\. The dataset\-size rule yields 100 ensemble configurations per fold\. At most 32 channels are selected using training data only\. Ridge classification uses ten logarithmically spaced penalties from10−110^\{\-1\}to10510^\{5\}\.
#### MrSQM
We use the random\-search strategy with 500 retained and 2,000 preselected features per representation\. The representation uses zero SAX and five SFA representations\. SFA normalization and first differences are enabled\. At most eight channels are selected using the training split\. Classification uses class\-balanced logistic regression withC=1C=1, the Newton–CG solver, and at most 1,000 iterations\.
#### DrCIF
The standaloneDrCIFclassifier uses 200 trees\. Interval counts are\(4,sqrt\-div\)\(4,\\texttt\{sqrt\-div\}\), with minimum interval length 3 and maximum interval length0\.50\.5of the series\. Ten attributes are sampled per interval\.Catch22features are disabled\. No time contract is used, and numerically near\-constant intervals are stabilized\.
#### HIVE\-COTE 2\.0
We use the four standardHIVE\-COTE 2\.0components with fourth\-power CAWPE weighting\. A nominal 360\-minute contract is provided per outer fold\. The implementation\-specific component configurations are given below\.
- •Shapelet Transform Classifier:10,000 shapelet samples are used, with unconstrained maximum shapelet count and length, batch size 100, and a contract cap of 10,000 samples\. Rotation Forest uses 200 trees with a contract cap of 200\.
- •DrCIF component:The internalDrCIFcomponent uses 500 trees, interval counts\(4,sqrt\-div\)\(4,\\texttt\{sqrt\-div\}\), minimum interval length 3, maximum proportional interval length0\.50\.5, and ten attributes per interval\.Catch22is disabled and near\-constant interval stabilization is enabled\. This configuration is distinct from the standaloneDrCIFbaseline above\.
- •Arsenal:The ROCKET transform uses 2,000 kernels per estimator, producing maximum and PPV features\. The ensemble contains 25 estimators with a contract cap of 25 estimators\. The configured maximum\-dilation and features\-per\-kernel arguments are inactive for the aeon ROCKET branch\.
- •Temporal Dictionary Ensemble:We use 250 parameter samples, a maximum ensemble size of 50, a maximum window\-length proportion of1\.01\.0, and a minimum window length of 10\. Fifty parameter settings are selected randomly, bigram configuration is automatic, the dimension threshold is0\.850\.85, and at most 20 dimensions are retained\. The contract cap is 250 parameter samples\.
### B\.6Random\-Convolution Features with Tabular Inference
#### RocketPFN
We use ten independent ROCKET groups with 1,000 kernels each and per\-instance normalization\. Each kernel contributes maximum and PPV features, yielding 2,000 features per group\. One TabPFN v2\.5 classifier with eight internal estimators is fitted per group\. Predicted probabilities are averaged across groups\. TabPFN usesfit\_preprocessorsmode, automatic device selection, memory\-saving and precision settings, one preprocessing worker, and no probability balancing\.
#### MASHT
The nominal feature budget is selected deterministically from the total numberNNof examples\. The budget is 10,000 forN<1,000N<1\{,\}000, 2,000 for1,000≤N<100,0001\{,\}000\\leq N<100\{,\}000, and 200 otherwise\. Only split sizes, rather than held\-out values or labels, enter this rule\. The balanced datasets considered here fall in the second range and therefore use a nominal budget of 2,000 features\. Half of this budget is assigned toMultiRocketand half toHydra\.MultiRocketdivides its budget equally between raw and first\-difference representations, uses four pooling features per kernel, at most 32 dilations, and no per\-instance normalization\.Hydrauses eight kernels per group and at most eight channels\. Only theHydrabranch receives sparse square\-root scaling\. The concatenated representation is classified by TabPFN 3 with eight estimators and automatic estimator scaling\.
#### AdaptiveRocketPFN
Our adaptive variant constructs raw, first\-difference, second\-difference, five\-point smoothed, high\-pass, and local\-normalized views withϵ=10−6\\epsilon=10^\{\-6\}\. Each view is transformed usingMultiRocketbanks with dilation caps 1 and 32, 625 kernels per bank, four pooling features per kernel, and no per\-instance normalization\. The candidate bank additionally containsHydrawith eight kernels per group, 128 groups, and at most eight channels\. It further contains at most 512 random dilated shapelets of lengths 7, 9, and 11, with normalization probability0\.80\.8and similarity threshold0\.50\.5\. Morphology, dynamics, and prototype experts each select 500 features using three\-fold, two\-repeat inner selection\. At most 12 feature families are active per expert, with at least eight features per active family\. The selection\-pool multiplier is 3, the allocation temperature is0\.50\.5, and the redundancy threshold is0\.980\.98, estimated from at most 256 cases\. Three\-fold out\-of\-fold predictions from two\-estimator TabPFN models determine nonnegative ensemble weights by regularized log loss withL2=0\.05L\_\{2\}=0\.05\. Final experts use TabPFN 3 with eight estimators\.
### B\.7Neural Time\-Series and Event Models
#### Inception\-TCN
The model contains eight modality\-specific branches\. Each branch uses 16 projection channels and two inception blocks with eight channels per branch and kernels of size 3, 5, and 9\. These are followed by depthwise\-separable TCN blocks with dilations 1, 2, and 4\. Mask\-aware mean and maximum pooling and the observed fraction are concatenated across modalities\. The classifier hidden size is 64, dropout is0\.300\.30, and normalized relative time is appended\. Training uses batch size 32 and batch size 64 for evaluation\. Models are trained for at most 100 epochs with learning rate3×10−43\\times 10^\{\-4\}and weight decay10−310^\{\-3\}\. Early\-stopping patience is 15, and scheduler patience is 5 with factor0\.50\.5and minimum learning rate10−610^\{\-6\}\. Mixed precision is enabled\.
#### Compact Fusion\-TCN
Per\-modality projection widths are\(8,16,8,8,8,8,24,24\)\(8,16,8,8,8,8,24,24\)for eye, facial, head, body, left/right hand, and left/right finger modalities, respectively\. The fused representation has width 48\. The shared TCN uses kernel size 5, dilations 1, 2, and 4, and two convolutions per block\. Global, pre\-anchor, and post\-anchor regions are pooled using mean, maximum, and standard deviation together with modality\-presence indicators\. The classifier hidden size is 32, dropout is0\.450\.45, and modality dropout is0\.150\.15\. Training uses batch size 32 and batch size 64 for evaluation\. Models are trained for at most 100 epochs with learning rate3×10−43\\times 10^\{\-4\}and weight decay10−210^\{\-2\}\. Gaussian input noise withσ=0\.02\\sigma=0\.02and temporal shifts of at most one grid step are applied during training\. Early\-stopping patience is 20, and scheduler patience is 6 with factor0\.50\.5and minimum learning rate10−610^\{\-6\}\. Mixed precision is enabled\. Results are aggregated over seeds\{17,42,73\}\\\{17,42,73\\\}\.
#### Irregular\-Event Transformer
The model uses one pre\-norm Transformer encoder layer with model width 32, four attention heads, feed\-forward width 128, and dropout0\.350\.35\. Modality encoders and the modality\-fusion MLP have hidden width 64\. The time encoder has width 32, and the classifier hidden layer has width 16\. Complete modality observations are masked with probability0\.200\.20and reconstructed during ten self\-supervised epochs on the outer training split\. Self\-supervised pretraining uses learning rate10−410^\{\-4\}and weight decay10−310^\{\-3\}\. Fine\-tuning uses batch size 8 for at most 60 epochs\. Backbone and classifier learning rates are10−410^\{\-4\}and3×10−43\\times 10^\{\-4\}, respectively, with weight decay10−210^\{\-2\}\. The pretrained backbone is frozen for the first three fine\-tuning epochs\. Early\-stopping patience is 10, and scheduler patience is 3 with factor0\.50\.5and minimum learning rate10−610^\{\-6\}\. Mixed precision is enabled\.
### B\.8Motion\-Language and Hierarchical Motion Models
All five motion\-model adaptations in this subsection use the same normalized fixed\-grid input with 32 temporal positions\.T2M\-GPTandMotionGPTuse the same VQ\-VAE tokenizer\. The tokenizer has 128 codes of dimension 64, hidden width 128, two temporal downsampling layers, two residual blocks, dilation growth 3, ReLU activations, and no dropout\. The EMA codebook uses decay0\.990\.99, codebookϵ=10−5\\epsilon=10^\{\-5\}, and reset threshold 1\. The 32 fixed\-grid positions are thereby compressed to eight motion tokens\. The tokenizer is trained for at most 100 epochs using batch size 32 and batch size 64 for evaluation\. Training uses AdamW with learning rate2×10−42\\times 10^\{\-4\}and zero weight decay\. The reconstruction objective is smooth\-L1L\_\{1\}augmented by velocity and commitment losses with weights0\.50\.5and0\.020\.02, respectively\. The tokenizer checkpoint is selected by validation reconstruction loss with patience 12\. Its learning\-rate scheduler uses patience 6, factor0\.50\.5, and minimum learning rate10−610^\{\-6\}\. Mixed precision is disabled for all motion models\.
#### T2M\-GPT
We use a causal discriminative Transformer with width 64, four attention heads, two layers, feed\-forward width 256, dropout0\.300\.30, a 32\-unit classifier, last\-token pooling, and token\-corruption rate0\.100\.10\. Before supervised fine\-tuning, the backbone undergoes 20 epochs of next\-token pretraining with batch size 32, learning rate3×10−43\\times 10^\{\-4\}, and weight decay10−210^\{\-2\}\. Fine\-tuning uses batch size 32 and batch size 64 for evaluation\. The model is fine\-tuned for at most 60 epochs with learning rate3×10−43\\times 10^\{\-4\}and weight decay10−210^\{\-2\}\. Early\-stopping patience is 10, and scheduler patience is 4 with factor0\.50\.5and minimum learning rate10−610^\{\-6\}\. The fixed decision threshold is0\.50\.5\.
#### T2M\-GPT\-v2
T2M\-GPT\-v2is our continuous\-latent ablation ofT2M\-GPT; it is not a separate published architecture\. The discrete VQ\-VAE is replaced by a variational autoencoder with latent dimension 64, hidden width 128, two temporal downsampling layers, two residual blocks, dilation growth 3, ReLU activations, and no dropout\. The 32 fixed\-grid positions are thereby compressed to eight continuous latent vectors\. The VAE is trained for at most 100 epochs using batch size 32 and batch size 64 for evaluation\. Training uses AdamW with learning rate2×10−42\\times 10^\{\-4\}and zero weight decay\. Its objective combines smooth\-L1L\_\{1\}reconstruction loss, velocity loss with weight0\.50\.5, and KL divergence with weight10−410^\{\-4\}; free bits are disabled\. The VAE checkpoint is selected by validation loss with early\-stopping patience 12 and minimum improvement10−510^\{\-5\}\. Its learning\-rate scheduler uses patience 6, factor0\.50\.5, and minimum learning rate10−610^\{\-6\}\. The causal Transformer has width 64, four attention heads, two layers, feed\-forward width 256, dropout0\.300\.30, a 32\-unit classifier, last\-token pooling, and latent\-corruption rate0\.100\.10\. The backbone undergoes 20 epochs of next\-latent pretraining using smooth\-L1L\_\{1\}regression, batch size 32, learning rate3×10−43\\times 10^\{\-4\}, and weight decay10−210^\{\-2\}\. Fine\-tuning is performed for at most 60 epochs with batch size 32, evaluation batch size 64, learning rate3×10−43\\times 10^\{\-4\}, and weight decay10−210^\{\-2\}\. Early\-stopping patience is 10, and scheduler patience is 4 with factor0\.50\.5and minimum learning rate10−610^\{\-6\}\.
#### MotionGPT
We use a motion\-to\-text encoder–decoder with model width 64, four attention heads, two encoder layers, two decoder layers, feed\-forward width 256, dropout0\.300\.30, tied word embeddings, maximum sequence length 64, and eight sentinel tokens\. The model is pretrained for 25 epochs on denoising, future\-prediction, and in\-between completion objectives\. Pretraining uses batch size 32, learning rate3×10−43\\times 10^\{\-4\}, and weight decay10−210^\{\-2\}\. Span corruption is0\.250\.25with mean span length 2\. Prediction\-context and in\-between\-span fractions are0\.500\.50and0\.250\.25, respectively\. Supervised tuning predicts the label word and otherwise uses the same optimization settings asT2M\-GPT\. The answer\-token log\-odds threshold is selected on the validation split to maximize macro\-F1F\_\{1\}\.
#### MotionGPT3
We evaluate a continuous\-latent classification adaptation ofMotionGPT3\. Because no textual descriptions are available, its language branch and cross\-modal attention are omitted; the portable continuous representation and diffusion\-head mechanism are retained\. The model uses the same continuous VAE and VAE\-training configuration asT2M\-GPT\-v2, producing eight 64\-dimensional latent vectors per window\. A non\-causal Transformer summarizer with width 64, four attention heads, two layers, feed\-forward width 256, and dropout0\.300\.30processes these vectors and produces a mean\-pooled motion representation\. The class\-conditioned diffusion head has hidden width 256, two residual blocks, and dropout0\.100\.10\. It uses 100 diffusion timesteps and a linear noise schedule fromβ1=10−4\\beta\_\{1\}=10^\{\-4\}toβ100=0\.02\\beta\_\{100\}=0\.02\. The denoising objective predicts the added Gaussian noise using mean\-squared error\. Training and classification scores average four independently sampled timestep–noise pairs per motion window\. Classification compares the denoising losses obtained under the positive\- and negative\-class embeddings\. Unlike theT2M\-GPTvariants, the diffusion classifier receives no separate next\-token or next\-latent pretraining\. It is trained for at most 60 epochs using batch size 32, evaluation batch size 64, AdamW with learning rate3×10−43\\times 10^\{\-4\}and weight decay10−210^\{\-2\}, and early\-stopping patience 10\. The learning\-rate scheduler uses patience 4, factor0\.50\.5, and minimum learning rate10−610^\{\-6\}\.
#### G\-HTT
OurG\-HTTadaptation hierarchically models short\- and long\-span temporal structure\. Each 32\-position window is divided into four non\-overlapping clips of length eight\. The short\-span pose block is a Transformer VAE that encodes the first four positions of each clip, reconstructs those positions, and predicts the remaining four\. Its encoder and two decoders use width 64, four attention heads, two layers, feed\-forward width 256, dropout0\.100\.10, and a 32\-dimensional latent bottleneck\. The component\-reconstruction and trajectory\-prediction losses both have weight 1\. Each is a smooth\-L1L\_\{1\}objective augmented by a velocity loss with weight0\.50\.5\. The pose\-block KL term has weight10−510^\{\-5\}, and free bits are disabled\. The pose block is trained for at most 100 epochs with batch size 32, evaluation batch size 64, AdamW learning rate2×10−42\\times 10^\{\-4\}, and zero weight decay\. Its early\-stopping patience is 12, while scheduler patience is 6 with factor0\.50\.5and minimum learning rate10−610^\{\-6\}\. The posterior mean of each clip yields a sequence of four 32\-dimensional mid\-level representations\. The long\-span action block is another Transformer VAE with width 64, four attention heads, two layers, feed\-forward width 256, dropout0\.300\.30, and latent dimension 32\. Its classifier contains one hidden layer of width 32\. The action\-block objective combines smooth\-L1L\_\{1\}reconstruction of the mid\-level sequence with weight 1, binary cross\-entropy with weight0\.10\.1, and KL divergence with weight10−510^\{\-5\}\. The action block is trained for at most 60 epochs using batch size 32, evaluation batch size 64, AdamW learning rate3×10−43\\times 10^\{\-4\}, and weight decay10−210^\{\-2\}\. Early\-stopping patience is 10, and scheduler patience is 4 with factor0\.50\.5and minimum learning rate10−610^\{\-6\}\.
### B\.9Software
Experiments use Python 3\.12\.3, NumPy 2\.3\.5\([Harris et al\., 2020](https://arxiv.org/html/2609.16396#bib.bib18)\), scikit\-learn 1\.8\.0\([Pedregosa et al\., 2011](https://arxiv.org/html/2609.16396#bib.bib47)\), SciPy 1\.17\.1\([Virtanen et al\., 2020](https://arxiv.org/html/2609.16396#bib.bib62)\), aeon 1\.5\.0\([Middlehurst et al\., 2024](https://arxiv.org/html/2609.16396#bib.bib37)\), Wildboar 1\.2\.1\([Samsten and Lee, 2024](https://arxiv.org/html/2609.16396#bib.bib53)\), MrSQM 0\.0\.7\([Nguyen and Ifrim, 2023](https://arxiv.org/html/2609.16396#bib.bib42)\), PyTorch 2\.13\.0\([Paszke et al\., 2019](https://arxiv.org/html/2609.16396#bib.bib46)\), Transformers 4\.51\.3\([Wolf et al\., 2020](https://arxiv.org/html/2609.16396#bib.bib65)\), and TabPFN 8\.1\.0\([Hollmann et al\., 2025](https://arxiv.org/html/2609.16396#bib.bib23)\)\.Similar Articles
Disparities In Negation Understanding Across Languages In Vision-Language Models
MIT researchers release the first multilingual negation benchmark covering seven languages and show VLMs like CLIP struggle with non-Latin scripts, while MultiCLIP and SpaceVLM offer uneven improvements across languages.
Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives
This paper investigates the use of conversational temporal dynamics (turn-pair timing) as a lightweight modality for automatic depression detection from dyadic clinical interviews, showing that a compact 24-dimensional timing module achieves strong performance and complements standard acoustic and semantic features when fused.
Context-Aware Multimodal Claim Verification in Spoken Dialogues
This paper introduces MAD2, a new benchmark for multimodal claim verification in spoken dialogues, and proposes a calibrated fusion of audio and text models that leverages conversational context to improve verification accuracy.
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
This paper analyzes multimodal systems for the CHiME-9 MCoRec cocktail-party scenario, comparing design strategies such as audio-visual target speech separation, improved recognition, and LLM-based conversational grouping, finding that speech overlap alone does not explain performance differences.
Commonsense Knowledge with Negation: A Resource to Enhance Negation Understanding
Researchers introduce a method to automatically augment commonsense knowledge corpora with negation, creating 2M+ triples that improve LLM negation understanding when used for pre-training.