How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

arXiv cs.CL Papers

Summary

This paper evaluates the predictive accuracy, cross-task generalizability, and test-retest reliability of multimodal features for measuring conversational states like cognitive load and power in dyadic remote collaborative tasks. Findings show that linguistic features predict well but generalize poorly, acoustic reliability degrades when controlling for speaker identity, and interaction features provide the most reliable signal.

arXiv:2607.17452v1 Announce Type: new Abstract: Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a three-dimensional evaluation framework assessing predictive accuracy, cross-task generalizability, and test-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video-conferencing platform (AVCAffe dataset; 53 dyads, 9 tasks). Our results show that no single feature family dominates all three dimensions. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross-task evaluation, revealing sensitivity to task-specific vocabulary. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state. Interaction features provide the only genuinely reliable signal, unchanged after speaker normalization. Interestingly, classifying power role remained near chance baseline across all conditions, indicating limitations of task-level aggregated behavior for predicting power role in conversation. Our findings reveal three insights: (1) linguistic features predict best but generalize poorly across task contexts; (2) acoustic reliability collapses to near-zero once speaker identity is controlled, challenging standard evaluation practice; and (3) interaction features provide the only genuinely reliable signal, with floor dominance predicting within-dyad cognitive load asymmetry. These results argue for speaker normalization and multi-dimensional evaluation as prerequisites for context-aware, robust multimodal feature selection in conversational systems.
Original Article
View Cached Full Text

Cached at: 07/21/26, 06:45 AM

# How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks
Source: [https://arxiv.org/html/2607.17452](https://arxiv.org/html/2607.17452)
\(2026\)

###### Abstract\.

Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts\. We present a three\-dimensional evaluation framework assessing predictive accuracy, cross\-task generalizability, and test\-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video\-conferencingplatform\(AVCAffedataset;5353dyads, 9 tasks\)\. Our results show that no single feature family dominates all three dimensions\. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross\-task evaluation, revealing sensitivity to task\-specific vocabulary rather than generalizable load signals\. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state\. Interaction features provide the only genuinely reliable signal \(unchanged after normalization, structurally immune to speaker identity inflation\), and floor dominance predicts within\-dyad mental load asymmetry\. Interestingly,classifying power roleremained near chance baseline across all conditions, indicating limitations of task\-level aggregated behavior for predicting power role in conversation\.Our findings reveal three insights: \(1\) linguistic features predict best but generalize poorly across task contexts; \(2\) acoustic reliability collapses to near\-zero once speaker identity is controlled, challenging standard evaluation practice; and \(3\) interaction features provide the only genuinely reliable signal, with floor dominance predicting within\-dyad cognitive load asymmetry\.These results argue for speaker normalization and multi\-dimensional evaluation as prerequisites for context\-aware, robust multimodal feature selection in conversational systems\.

cognitive load, conversational power, multimodal signals, reliability, dyadic interaction, feature evaluation, remote collaboration\.

††journalyear:2026††copyright:cc††conference:INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION; October 05–09, 2026; Napoli, Italy††booktitle:INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION \(ICMI ’26\), October 05–09, 2026, Napoli, Italy††doi:10\.1145/3776574\.3831235††isbn:979\-8\-4007\-2318\-6/2026/10††ccs:Human\-centered computing Empirical studies in collaborative and social computing††ccs:Computing methodologies Machine learning††ccs:Human\-centered computing## 1\.Introduction

Socially aware conversational systems, designed to support collaboration, negotiation, and group decision\-making, depend on reliable measurement of latent conversational states such as cognitive load and conversational power from observable multimodal behavior\. A growing body of work has explored acoustic prosody, turn\-taking dynamics, and linguistic content as informative signals for this measurement task\(Vinciarelliet al\.,[2010](https://arxiv.org/html/2607.17452#bib.bib26); Schulleret al\.,[2013](https://arxiv.org/html/2607.17452#bib.bib27); Vukovicet al\.,[2021](https://arxiv.org/html/2607.17452#bib.bib9)\), yet two critical questions remain largely unaddressed:*do these signals generalize across task contexts?*and*do they measure what they appear to measure, or do they reflect other stable characteristics of speakers?*

Predictive accuracy, the dominant evaluation criterion in multimodal affective computing, does not answer either question\. A feature may predict cognitive load well within a particular task distribution while completely failing on novel task types\. Its reliability estimate may be inflated by stable speaker characteristics \(vocal tract anatomy, speaking style\) rather than genuine behavioral consistency during the conversation\. These distinctions matter enormously for system design: a feature family that is predictive but not generalizable or reliable cannot serve as the foundation of a real\-time measurement system deployed across diverse contexts\.

In this paper, we examine predictive accuracy, reliability, and generalizability, all three dimensions simultaneously\. We evaluate interactional, acoustic, and linguistic feature families using: \(1\) leave\-one\-dyad\-out \(LODO\) cross\-validation for predictive accuracy; \(2\) leave\-one\-task category\-out \(LOCO\) analysis across different task types for generalizability; and \(3\) intraclass correlation coefficients \(ICC\) before and after within\-speaker normalization for reliability\. We apply this three\-dimensional framework to the AVCAffe dataset of remote dyadic collaboration\(Sarkaret al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib1)\), which provides naturalistic Zoom interactions annotated for cognitive load and conversational power across nine distinct tasks\.Specifically, we train Random Forest models on task\-level behavioral features to predict self\-reported cognitive load scores111Mental and temporal demand are used as measures for cognitive load construct throughout this paper; they are two of the six NASA\-TLX subscales\(Hart and Staveland,[1988](https://arxiv.org/html/2607.17452#bib.bib2)\)selected for their moderate within\-dyad agreement and sensitivity to task\-level conversational behavior in dyadic remote work \(Section[3](https://arxiv.org/html/2607.17452#S3)\)\.\(regression\) and power class \(classification\), evaluated under leave\-one\-dyad\-out cross\-validation using Lin’s concordance correlation coefficient \(CCC\) and macro F1 respectively\.

We organize our evaluation around four research questions:

- •RQ1: Which feature family best predicts cognitive load and conversational power under leave\-one\-dyad\-out cross\-validation?
- •RQ2: Do predictive features generalize across task categories?
- •RQ3: How reliable are features across repeated instances of similar tasks, and does reliability change after speaker normalization?
- •RQ4: Does floor dominance predict within\-dyad cognitive load asymmetry?

Our main contributions are:

- •A three\-dimensional evaluation framework \(prediction, generalizability, reliability\) for multimodal feature selection that can be reused beyond the dataset used in this study\.
- •Evidence that linguistic features predict cognitive load significantly better than interaction or acoustic features, but are strongly task\-dependent\.
- •Evidence that raw acoustic features in dyadic remote conversation reflect speaker identity rather than behavioral state: standard within\-speaker normalization eliminates all apparent reliability, raising concerns about inflated observations in multimodal affective computing applications\.
- •Floor dominance, an interaction feature measuring which one speaker controls the conversational floor, is significantly associated with load asymmetry between speakers within a dyad\.

![Refer to caption](https://arxiv.org/html/2607.17452v1/fig7_framework_concept.png)Figure 1\.Our three\-dimensional evaluation framework\. Each line traces one feature family across prediction accuracy \(mean LODO CCC\), cross\-task generalizability \(mean LOCO CCC\), and test\-retest reliability \(mean ICC\(2,1\)\)\. Crossing lines reveal a rank dissociation that no feature family dominates all three dimensions\. Background shaded zones indicate ICC reliability thresholds based on\(Koo and Li,[2016](https://arxiv.org/html/2607.17452#bib.bib5)\)\.Our three\-dimensional evaluation framework\. Each line traces one feature family across prediction accuracy \(mean LODO CCC\), cross\-task generalizability \(mean LOCO CCC\), and test\-retest reliability \(mean ICC\(2,1\)\)\. Crossing lines reveal a rank dissociation that no feature family dominates all three dimensions\. Background shaded zones indicate ICC reliability thresholds based on \\cite\[citep\]\{\(\\@@bibref\{AuthorsPhrase1Year\}\{koo2016icc\}\{\\@@citephrase\{, \}\}\{\}\)\}\.Our findings draw attention to the need for context awareness in social interaction: features trained on one conversational context must generalize to others if they are to support context\-aware interaction systems\.

## 2\.Related Work

### 2\.1\.Speech Features for Cognitive Load

Acoustic features have been the dominant modality for cognitive load detection from speech\. Early work by Yin et al\.\(Yinet al\.,[2007](https://arxiv.org/html/2607.17452#bib.bib6),[2008](https://arxiv.org/html/2607.17452#bib.bib7)\)demonstrated that pitch, speaking rate, and energy carry systematic load signatures\. The Interspeech 2014 Computational Paralinguistics Challenge\(Schulleret al\.,[2014](https://arxiv.org/html/2607.17452#bib.bib8)\)established speech\-based cognitive and physical load estimation as a benchmark problem, spurring a line of work on acoustic features for cognitive load estimation in operational settings including aviation communication\(Vukovicet al\.,[2021](https://arxiv.org/html/2607.17452#bib.bib9); Yanget al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib10)\)and simulated flight\(Xuet al\.,[2025](https://arxiv.org/html/2607.17452#bib.bib13)\)\. Boyer et al\.\(Boyeret al\.,[2018](https://arxiv.org/html/2607.17452#bib.bib11)\)and Taptiklis et al\.\(Taptikliset al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib12)\)confirmed prosodic vocal biomarkers as reliable mental effort indicators in controlled conditions\. However, this body of work evaluates acoustic features in single\-task, within\-corpus settings; cross\-task generalizability and speaker\-level reliability are not assessed\.To the best of our knowledge, no prior work has evaluated acoustic reliability after within\-speaker normalization in remote collaboration tasks, which we address in this work\.

### 2\.2\.Linguistic and Multimodal Features

Khawaja et al\.\(Khawajaet al\.,[2012](https://arxiv.org/html/2607.17452#bib.bib14)\)showed that linguistic features including hedges, fillers, and pronoun rates reflect cognitive load in collaborative communication, directly motivating linguistic feature family in our study\. Abel and Babel\(Abel and Babel,[2017](https://arxiv.org/html/2607.17452#bib.bib16)\)found that cognitive load reduces linguistic convergence in dyads, suggesting that lexical diversity is a theoretically grounded load marker\. Zhou et al\.\(Zhouet al\.,[2018](https://arxiv.org/html/2607.17452#bib.bib17)\)proposed a multimodal cognitive load measurement model combining behavioral and physiological signals, arguing that multimodal data fusion improves model robustness over any single data modality\. These findings motivate our multimodal feature family comparison\. However, none of these prior works has evaluated cross\-task generalizability or speaker\-normalized reliability for cognitive load prediction from multimodal data, which is what we focus on in this work\.

### 2\.3\.Conversational Dynamics and Interaction Features

Turn\-taking patterns, overlap, and floor control have been studied as indicators of social power and conversational dominance in several prior works\(Vinciarelliet al\.,[2010](https://arxiv.org/html/2607.17452#bib.bib26); Gravano and Hirschberg,[2011](https://arxiv.org/html/2607.17452#bib.bib22)\)\. Additionally, silence and response latency have been linked to cognitive disengagement and processing time\(Heldner and Edlund,[2010](https://arxiv.org/html/2607.17452#bib.bib23); Johnstoneet al\.,[1995](https://arxiv.org/html/2607.17452#bib.bib24)\)indicating unequal participation in interaction as an indicator for cognitive load\. These interaction features motivate our dyad\-level feature family to capture the conversation dynamics, as their reliability and cross\-task stability relative to commonly used acoustic and linguistic families has not previously been evaluated\.

### 2\.4\.Datasets for Remote Collaborative Cognitive Load

Sarkar et al\.\(Sarkaret al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib1)\)introduced AVCAffe, the first audio\-visual dataset combining cognitive load and affect annotations for remote work, which we use in this paper\. Their evaluation uses binary classification on 2\-second video clips with deep learning backbones, reporting weighted F1 \(best: 65–67% for mental and temporal demand on long videos\)\. Our work differs fundamentally from their approach: we use task\-level feature summaries consistent with the inherent nature of NASA\-TLX\(Hart,[2006](https://arxiv.org/html/2607.17452#bib.bib3); Hart and Staveland,[1988](https://arxiv.org/html/2607.17452#bib.bib2)\), employ regression with Concordance Correlation Co\-efficient \(CCC\) as our evaluation metric, and focus on measurement validity in this work rather than model benchmarking based on performance only\.

The recently introduced CoAffinity dataset\(Gunasekaranet al\.,[2025](https://arxiv.org/html/2607.17452#bib.bib18)\)provides a complementary resource for similar tasks: 39 participants performing eight remote\-work tasks with audio, video, and physiological signals \(PPG, GSR\) and cognitive load annotations\. Their results confirm that integrating physiological modalities substantially improves load detection, suggesting that pure behavioral signals face an inherent ceiling, a finding consistent with our moderate CCC results across all three feature families\.

### 2\.5\.Reliability and Measurement Validity in Behavioral Signals

We use the standard Intraclass Correlation Co\-efficient \(ICC\) threshold provided by Koo and Li\(Koo and Li,[2016](https://arxiv.org/html/2607.17452#bib.bib5)\)as a statistical measure for reliability throughout our reliability analysis\. Despite ICC being the standard tool for test\-retest reliability in behavioral measurement, its application to multimodal affective computing is relatively rare\. The systematic inflation of raw acoustic ICC by speaker\-specific vocal characteristics \(e\.g\., vocal tract anatomy, stable speaking style\) is a concern we demonstrate empirically for the first time in a dyadic remote conversation dataset\. In a similar work, Koenecke et al\.\(Koeneckeet al\.,[2020](https://arxiv.org/html/2607.17452#bib.bib29)\)demonstrate unequal ASR error rates across speaker backgrounds, which has direct implications for the reliability of linguistic features derived from automatically generated transcripts for datasets with diverse participant pools such as AVCAffe \(18 countries of origin\) with varying accent, gender, and other speaker conditions\.

Table 1\.Our contributions positioned relative to prior related work\.
### 2\.6\.Positioning Our Contribution

Table[1](https://arxiv.org/html/2607.17452#S2.T1)summarizes key distinctions between our work and the most closely related studies\. Our contribution is threefold: \(1\) we introduce a three\-dimensional evaluation framework \(prediction, generalizability, reliability\) that no other work applies jointly; \(2\) we provide a speaker\-normalized ICC analysis of acoustic features in a remote conversation dataset, revealing systematic inflation by speaker identity; and\(3\) we examine whether linguistic features’ high predictive accuracy generalizes across task contexts — a question prior work has not addressed and has direct implications for the design of future context aware conversational systems\.

## 3\.Data

### 3\.1\.AVCAffe Dataset

We use AVCAffe\(Sarkaret al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib1)\), an audio\-visual dataset of remote dyadic collaboration conducted over Zoom\. The dataset consists of 106 participants \(49% male, 50% female, 1% non\-binary; age 18–57; 18 countries of origin\) organized into 53 dyads\. Each dyad completed a session of nine task instances spanning seven task types: open discussion, lighten the mood \(sharing jokes\), Diapix \(spot\-the\-difference in images\), Montclair map \(spatial coordination; completed twice\), Lost at Sea \(group collaborative decision making\), reading comprehension \(two rounds with roles reversed\), and multi\-task \(email writing with interruptions from partners\)\. Tasks were designed to elicit varying levels of cognitive demand, from low \(social/open tasks\) to high \(time\-pressured coordination\)\. Ethics approval was obtained from the General Research Ethics Board at Queen’s University, Canada\.

The participant pool spans a wide range of professional backgrounds \(engineers, scientists, students, nurses, and lawyers among others\) and age groups \(18–57 years\), with approximately 60% from North America and the remainder from India, Iran, and 14 other countries, making it one of the most geographically diverse remote work datasets available for cognitive load research using audio\-visual recordings\.

![Refer to caption](https://arxiv.org/html/2607.17452v1/x1.png)Figure 2\.NASA\-TLX cognitive load scores per task \(mean±\\pmstd,n=53n=53dyads\)\. Bar colors indicate task category\. Social tasks \(Tasks 1–2, green\) elicit the lowest load on both dimensions, explaining why the Social category is the most challenging held\-out condition in cross\-task evaluation\. Map Matching 1 \(Task 4\) and Lost at Sea \(Task 5\) show the highest temporal demand\. The gray dashed line marks the scale midpoint \(10\.5\)\.NASA\-TLX cognitive load scores per task \(mean $\\pm$ std, $n=53$ dyads\)\. Bar colors indicate task category\. Social tasks \(Tasks 1–2, green\) elicit the lowest load on both dimensions, explaining why the Social category is the most challenging held\-out condition in cross\-task evaluation\. Map Matching 1 \(Task 4\) and Lost at Sea \(Task 5\) show the highest temporal demand\. The gray dashed line marks the scale midpoint \(10\.5\)\.
### 3\.2\.Annotations

NASA\-TLX ratings\(Hart and Staveland,[1988](https://arxiv.org/html/2607.17452#bib.bib2)\)were collected at the end of each task on a 0–21 scale across six dimensions\. We use mental demand and temporal demand as primary regression targets, selected for their moderate within\-dyad agreement \(r=0\.341r=0\.341andr=0\.300r=0\.300respectively\) and theoretical relevance to collaborative task engagement\. We exclude frustration \(r=0\.113r=0\.113\) and physical demand \(r=0\.071r=0\.071\) entirely as their low inter\-rater agreement indicates these dimensions reflect individual responses not systematically captured by shared conversational behavior\. Additionally, Effort showed redundant patterns with mental demand and is omitted from reported results\.

### 3\.3\.Prediction Targets

#### 3\.3\.1\.Cognitive Load\.

For this regression task, we use thedyadic meanof both speakers’ rating as the cognitive load target, to be consistent with the retrospective, shared\-task nature of NASA\-TLX annotation\.

#### 3\.3\.2\.Conversational Power\.

Power annotation was collected as a self\-reported label per participant per task across five categories \(powerful, independent, neutral, dependent, powerless\)\. Due to high class imbalance, we collapse these labels to three levels instead \(high level: powerful \+ independent; medium level: neutral; low level: dependent \+ powerless\) yieldingn=354n=354/502502/9494instances in the dataset\.

### 3\.4\.Transcripts

Speaker\-level transcripts were generated using Distil\-Whisper ASR\(Gandhiet al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib28)\)applied to the per\-speaker, per\-task audio files provided as part of the dataset\. We note that AVCAffe participants represent 18 countries of origin\. As ASR transcription quality is known to vary across speaker accents\(Koeneckeet al\.,[2020](https://arxiv.org/html/2607.17452#bib.bib29); Emara and Shaker,[2024](https://arxiv.org/html/2607.17452#bib.bib34); Ngueajio and Washington,[2022](https://arxiv.org/html/2607.17452#bib.bib33)\), this may affect the reliability and accuracy of linguistic features derived from these transcripts\.

## 4\.Methods

### 4\.1\.Feature Families

We evaluate three theoretically grounded feature families, all computed at the task level \(one value per speaker per task\)\.

Interaction features\(14 features, dyad\-level\) capture conversational dynamics motivated by conversation analysis: floor imbalance and turn count ratio\(Sackset al\.,[1974](https://arxiv.org/html/2607.17452#bib.bib19)\), mean turn duration \(both speakers\), overlap ratio\(Gravano and Hirschberg,[2011](https://arxiv.org/html/2607.17452#bib.bib22)\), interruption imbalance and per\-speaker interruption rates, mean and standard deviation of response latency\(Levinson,[2016](https://arxiv.org/html/2607.17452#bib.bib20)\), long silence duration \(≥1\\geq\\\!1s\), maximum inter\-turn silence, and dyadic dominance indices\. These features are identical for both speakers within a dyad\-task and are aggregated from 30\-second window\-level interaction features\.

Acoustic features\(10 features, per speaker\) capture prosodic and spectral characteristics motivated by load and arousal research\. This feature family includes pitch mean, std, and range; RMS loudness mean and std; voiced segments per minute; mean voiced segment duration; spectral centroid and bandwidth means; and estimated words per minute\.

Linguistic features\(21 features, per speaker\) capture speech content motivated by Linguistic Inquiry and Word Count \(LIWC\)\(Tausczik and Pennebaker,[2010](https://arxiv.org/html/2607.17452#bib.bib15)\)and collaborative communication research\(Khawajaet al\.,[2012](https://arxiv.org/html/2607.17452#bib.bib14)\): lexical diversity \(type\-token ratio and its root\-normalized variant\), average word length, hedge and filler rates\(Khawajaet al\.,[2012](https://arxiv.org/html/2607.17452#bib.bib14)\), question and directive rates, first/second person pronoun rates and their ratio, agreement and disagreement rates, backchannel and proposal rates, modal verb, certainty, deference, negation, and politeness marker rates\.

To summarize: interaction features are aggregated from 30\-second windows to task level; acoustic and linguistic features are computed directly over the full task recording and full task transcript respectively, yielding one value per speaker per task for all feature families\. The resulting feature conditions range from 10 features \(acoustic only\) to 44 features \(when all families combined\) \(see Table[2](https://arxiv.org/html/2607.17452#S4.T2)\)\.

### 4\.2\.Task\-Level Analysis

We analyze features at the task level to be consistent with the annotations provided with this dataset\. NASA\-TLX is a retrospective self\-report collected at the end of each task, making the task the natural unit of analysis\. While prior work has taken a window\-level modeling approach by assigning task\-level labels to individual windows, this requires a Multiple Instance Learning\(Ilseet al\.,[2018](https://arxiv.org/html/2607.17452#bib.bib35)\)assumption that we deliberately avoid here, keeping the measurement framework simple and interpretable\.

### 4\.3\.Evaluation Protocol

Evaluation Metrics\.For cognitive load \(regression task\), we report Lin’s concordance correlation coefficient \(CCC\)\(Lin,[1989](https://arxiv.org/html/2607.17452#bib.bib4)\), which jointly penalizes lack of correlation and systematic prediction bias\. For power \(classification task\), we report macro F1 with class weight to compensate for class imbalance\.

#### 4\.3\.1\.RQ1: Leave One Dyad Out \(LODO\) for prediction accuracy\.

We use leave\-one\-dyad\-out cross\-validation over 53 folds for cross validation\. Features are standardized within each fold to prevent data leakage\. We use Random Forest \(100 trees,random\_state=42\) for regression and classification\(Shwartz\-Ziv and Armon,[2022](https://arxiv.org/html/2607.17452#bib.bib30)\)\.

#### 4\.3\.2\.RQ2: Leave One Category Out \(LOCO\) for cross\-task generalizability\.

We apply leave\-one\-category\-out cross\-validation across four task categories: Social/Open \(tasks 1–2\), Cognitive \(tasks 3, 6, 7\), Problem\-solving \(task 5\), and Time\-pressured \(tasks 4, 8, 9\), testing whether features trained on one task context generalize to another\. The tasks are grouped based on their nature and expected speaker interaction\.

#### 4\.3\.3\.RQ3: Intra\-class Correlation Coefficient \(ICC\) for reliability\.

We compute ICC\(2,1\) — two\-way random effects, single rater, absolute agreement\(Koo and Li,[2016](https://arxiv.org/html/2607.17452#bib.bib5)\)on two test\-retest pairs within dyads: map\-matching \(tasks 4 and 9\) and reading comprehension \(tasks 6 and 7\), which share task type and demand level\. Our reliability thresholds follow Koo and Li\(Koo and Li,[2016](https://arxiv.org/html/2607.17452#bib.bib5)\): ICC\>0\.75\>0\.75= good,0\.500\.50–0\.750\.75= moderate,<0\.50<0\.50= poor\.

Table 2\.Leave\-One\-Dyad\-Out \(LODO\) cross\-validation \(n=53n=53dyads\)\. CCC with 95% bootstrap CI in brackets \(2,000 samples; non\-overlapping CIs indicate statistically supported differences\)\. Power F1 = macro F1, 3\-class; chance baseline = 0\.333\.†\\dagger= best per column\. Effort omitted \(redundant with mental demand\); performance omitted \(negative CCC under cross\-task evaluation\)\.
#### 4\.3\.4\.RQ4: Load Asymmetry\.

Floor dominance and within\-dyad load asymmetry are analyzed using Pearson and Spearman correlation\.

Speaker normalization\.To test whether acoustic \(and linguistic\) ICC is inflated by speaker identity rather than behavioral consistency, we apply within\-speakerzz\-score normalization: for each speaker, each feature is standardized across their nine task observations\. This removes between\-speaker variance \(vocal tract anatomy, stable speaking style\) while preserving within\-speaker task\-to\-task variation\.Note that interaction features are dyadic relative measures \(ratios and differences between speakers\) that by construction cancel between\-speaker absolute levels\.Negative ICC values after normalization indicate that within\-speaker task\-to\-task variation is indistinguishable from residual error and no reliable signal remains once speaker identity is removed\.

![Refer to caption](https://arxiv.org/html/2607.17452v1/x2.png)Figure 3\.Predictive accuracy \(LODO,n=53n=53dyads\)\. Dots show mean CCC; lines show 95% bootstrap CIs \(2,000 samples\)\. Filled circles = single\-family conditions; open squares = combined\. Brackets on panel \(b\) indicate non\-overlapping CIs \(p<\.05p<\.05, bootstrap\)\. CCC = Lin’s concordance correlation coefficient\(Lin,[1989](https://arxiv.org/html/2607.17452#bib.bib4)\)\.Predictive accuracy \(LODO, $n=53$ dyads\)\. Dots show mean CCC; lines show 95% bootstrap CIs \(2,000 samples\)\. Filled circles = single\-family conditions; open squares = combined\. Brackets on panel \(b\) indicate non\-overlapping CIs \($p<\.05$, bootstrap\)\. CCC = Lin’s concordance correlation coefficient \\cite\[citep\]\{\(\\@@bibref\{AuthorsPhrase1Year\}\{lin1989ccc\}\{\\@@citephrase\{, \}\}\{\}\)\}\.![Refer to caption](https://arxiv.org/html/2607.17452v1/x3.png)Figure 4\.Top three features per family for each prediction target \(Random Forest, all\-features condition, LODO\)\. Lexical diversity = root type\-token ratio \(proportion of unique words relative to square root of total words; higher diversity = more varied vocabulary\)\.Top three features per family for each prediction target \(Random Forest, all\-features condition, LODO\)\. Lexical diversity = root type\-token ratio \(proportion of unique words relative to square root of total words; higher diversity = more varied vocabulary\)\.

## 5\.Results

### 5\.1\.Predictive Accuracy \(RQ1\)

Statistical Criterion\.We report 95% bootstrap confidence intervals \(2,000 samples, fold\-level resampling\) to support inferential interpretation of the regression and classification task given the small sample size\. We use non\-overlapping 95% bootstrap CIs as our criterion for statistical support rather than a formal statistical test, given that no standard significance test exists for comparing LODO\-fold CCC distributions with correlated folds\(Lin,[1989](https://arxiv.org/html/2607.17452#bib.bib4)\)\.Table[2](https://arxiv.org/html/2607.17452#S4.T2)and Figure[3](https://arxiv.org/html/2607.17452#S4.F3)show LODO cross\-validation results for all feature conditions\.

Linguistic features significantly outperform interaction and acoustic families for temporal demand\.Linguistic features achieve CCC = 0\.545, compared to interaction \(0\.379\) and acoustic \(0\.205\), with non\-overlapping confidence intervals confirming these differences are not attributable to sampling variability\. Lexical diversity \(root type\-token ratio; importance = 0\.234 in the all\-features condition\) is the dominant predictor for temporal demand, consistent with the finding that vocabulary richness decreases as speakers simplify language when under cognitive time pressure\(Khawajaet al\.,[2012](https://arxiv.org/html/2607.17452#bib.bib14); Abel and Babel,[2017](https://arxiv.org/html/2607.17452#bib.bib16)\)\.

Table 3\.LOCO cross\-task generalizability: mental demand CCC when each task category is held out\. Problem\-solving contains one task \(n=106n=106\)\.Range column indicates difference betweenmax\(CCC\) andmin\(CCC\) across the four held\-out categories per condition\.![Refer to caption](https://arxiv.org/html/2607.17452v1/x4.png)Figure 5\.Cross\-task generalizability results\. Bars show CCC when the column category is held out for testing\. Dashed lines in panel \(b\) show within\-distribution LODO performance as baseline\. Linguistic temporal demand degrades from 0\.545 \(LODO\) to 0\.009 on Social tasks\.Cross\-task generalizability results\. Bars show CCC when the column category is held out for testing\. Dashed lines in panel \(b\) show within\-distribution LODO performance as baseline\. Linguistic temporal demand degrades from 0\.545 \(LODO\) to 0\.009 on Social tasks\.For mental demand, differences between conditions are present but not statistically distinguished\.All conditions produce CCC between 0\.184 and 0\.343, with broadly overlapping confidence intervals, reflecting the limited statistical power of the 53\-dyad sample\. Long silence duration \(interaction importance = 0\.115\) and lexical diversity \(linguistic importance = 0\.123\) are the most informative features for mental demand, consistent with silence as a disengagement marker\(Heldner and Edlund,[2010](https://arxiv.org/html/2607.17452#bib.bib23)\)and with linguistic load sensitivity\(Tausczik and Pennebaker,[2010](https://arxiv.org/html/2607.17452#bib.bib15)\)\.

Combining interaction and linguistic features does not significantly improve over linguistic features alone\.The interaction \+ linguistic condition \(CCC = 0\.519 \[0\.458, 0\.580\]\) is statistically indistinguishable from linguistic only, suggesting interaction features contribute complementary but not additive signal for temporal demand within this sample\.

Adding Acoustic features bring nominal value\.Adding acoustic features to interaction improves temporal demand CCC by only 0\.013, a trivial gain by combining multimodal information\.

Power classification is near chance\.Conversational power is predicted as a 3\-class problem \(high/neutral/low\), evaluated with macro F1 under LODO evaluation protocol using a class\-weighted Random Forest\. The 3\-class chance baseline \(F1 = 0\.333\) falls within the 95% bootstrap CI for every feature condition \(range: 0\.310–0\.359\), confirming that no feature family predicts conversational power above chance at the task level\.We interpret this null result as reflecting the inherent subjectivity of power perception, label sparsity in the low\-power class \(9\.9%\), and the limited temporal resolution in the data due to the task\-level aggregation\.

### 5\.2\.Cross\-Task Generalizability \(RQ2\)

We present the results from cross\-task generalizability experiments in Table[3](https://arxiv.org/html/2607.17452#S5.T3)and Figure[5](https://arxiv.org/html/2607.17452#S5.F5)\.

Linguistic features are strongly task\-dependent\.Temporal demand CCC drops from 0\.545 \(within\-distribution, LODO\) to 0\.009 when tested on Social tasks\(Table[3](https://arxiv.org/html/2607.17452#S5.T3), Linguistic row, Social column; Figure[5](https://arxiv.org/html/2607.17452#S5.F5)b\), while remaining high for Time\-pressured tasks \(0\.406\)\. This performance degradation reveals that linguistic features are learning task\-specific vocabulary patterns rather than a generalizable load signal; high\-demand tasks elicit different speech content than social conversations\. The striking contrast between within\-distribution and cross\-task performance \(0\.545 vs\. 0\.009\) demonstrates that predictive accuracy alone is insufficient as a measurement validity criterion\.

Acoustic features show more consistent cross\-task performance\.For mental demand, acoustic CCC ranges from 0\.121 \(Social\) to 0\.355 \(Problem\-solving\), the most consistent profile across categories\. However, our results in the next section \(Section[5\.3](https://arxiv.org/html/2607.17452#S5.SS3)\) show this consistency reflects speaker identity rather than genuine behavioral measurement\.

Table 4\.Test\-retest ICC\(2,1\) before and after within\-speaker z\-score normalization; averaged over map\-matching \(tasks 4 & 9\) and reading comprehension \(tasks 6 & 7\)\. Normalization removes between\-speaker variance; negative normalized ICC means not reliable beyond speaker identity\(Koo and Li,[2016](https://arxiv.org/html/2607.17452#bib.bib5)\)\. Interaction features not normalized \(dyadic relative measures\)\. Thresholds: ICC\>\>0\.75 = good; 0\.50–0\.75 = moderate;<<0\.50 = poor\.Combining features dampens task\-dependence\.The all\-features condition \(mental demand range: 0\.153–0\.234\) is more balanced than any single family, suggesting that combining families reduces sensitivity to task\-specific patterns\.

We note that the Problem\-solving category contains a single task \(n=106n=106test instances\); estimates for this cell should be interpreted with caution\.

### 5\.3\.Measurement Reliability \(RQ3\)

Table[4](https://arxiv.org/html/2607.17452#S5.T4)and Figure[6](https://arxiv.org/html/2607.17452#S5.F6)present ICC results before and after within\-speaker normalization\.

Acoustic ICC collapses entirely after speaker normalization\.Raw acoustic ICC \(mean: 0\.576, moderate\) drops to−0\.112\-0\.112after within\-speaker z\-scoring, a 119% reduction\. The most dramatic drops occur for spectral features: pitch mean \(0\.966→\\to−0\.057\-0\.057\), spectral centroid \(0\.937→\\to−0\.178\-0\.178\), and spectral bandwidth \(0\.936→\\to−0\.131\-0\.131\)\. These are speaker\-intrinsic properties reflecting vocal tract anatomy rather than task\-sensitive behavioral states\. Negative normalized ICC indicates that within\-speaker variation across tasks is indistinguishable from measurement error, no reliable signal remains once speaker identity is removed\.

![Refer to caption](https://arxiv.org/html/2607.17452v1/x5.png)Figure 6\.Test\-retest reliability\. Filled circles \(•\) = raw ICC; open circles \(◦\) = normalized ICC; lines = ICC drop after normalization\. ICC\(2,1\) = two\-way random effects, single rater, absolute agreement\(Koo and Li,[2016](https://arxiv.org/html/2607.17452#bib.bib5)\)\. Normalization removes between\-speaker variance while preserving within\-speaker variation\. Acoustic features \(red\) show dramatic drops; interaction features \(blue, single dots\) remain unchanged\.Test\-retest reliability\. Filled circles \(•\) = raw ICC; open circles \(◦\) = normalized ICC; lines = ICC drop after normalization\. ICC\(2,1\) = two\-way random effects, single rater, absolute agreement \\cite\[citep\]\{\(\\@@bibref\{AuthorsPhrase1Year\}\{koo2016icc\}\{\\@@citephrase\{, \}\}\{\}\)\}\. Normalization removes between\-speaker variance while preserving within\-speaker variation\. Acoustic features \(red\) show dramatic drops; interaction features \(blue, single dots\) remain unchanged\. Note that only representative features are shown here\.Linguistic ICC is largely artificial\.Raw linguistic ICC \(mean: 0\.090\) drops to−0\.095\-0\.095after normalization\. Notably, lexical diversity \(root type\-token ratio\), the most predictive linguistic feature, has ICC = 0\.037 and 0\.035 across both test\-retest pairs, near\-zero even before normalization\. This confirms that lexical diversity is capturing task\-specific vocabulary rather than a stable individual characteristic\.

Interaction features provide genuine reliability\.Table[4](https://arxiv.org/html/2607.17452#S5.T4)confirms this observation empirically as Interaction ICC is identical before and after normalization \(\+0\.193→\+0\.193\+0\.193\\to\+0\.193\)\. Interaction ICC \(mean: 0\.193, poor\) remains unchanged by normalization, as dyadic relative measures \(floor imbalance, interruption ratio\) are structurally immune to speaker identity inflation\.While the overall level is poor by standard thresholds, it represents genuine cross\-task consistency post\-normalization\.

Power label consistency is fair\.Cohen’sκ\\kappabetween map\-matching tasks 4 and 9 is 0\.326 \(fair\), and 0\.237 between reading comprehension tasks 6 and 7, consistent with the near\-chance classification results\.

### 5\.4\.Load Asymmetry \(RQ4\)

Within\-dyad mental load asymmetry \(\|loadA−loadB\|\|\\text\{load\}\_\{A\}\-\\text\{load\}\_\{B\}\|; mean = 5\.57, SD = 4\.57\) was significantly predicted by floor imbalance \(r=0\.315r=0\.315, 95% CI \[0\.232, 0\.394\],p<0\.001p<0\.001,n=475n=475dyad\-task pairs\)\. Dyads with more unequal talk time distributions show greater divergence in experienced cognitive burden\. No other interaction imbalance feature reached significance with medium or greater effect size \(correlation\|r\|<0\.12\|r\|<0\.12, Spearman\|ρ\|<0\.08\|\\rho\|<0\.08\)\. These findings connect an observable behavioral asymmetry to the distribution of subjective experience within dyads between speakers\.

## 6\.Discussion

The three\-dimensional framework as a diagnostic tool\.Our results reveal a systematic dissociation across three measurement dimensions\. Linguistic features are the best predictors but the worst in terms of both generalizability and \(normalized\) reliability\. They are powerful within\-distribution tools whose apparent reliability is an artifact of between\-speaker vocabulary differences and task\-specific content\. Acoustic features appear moderately reliable by conventional ICC but this is entirely attributable to speaker\-intrinsic spectral characteristics\. Interaction features are genuine but modest in all three dimensions\. This suggests that no single family is suitable as a standalone measurement tool\.

### 6\.1\.Practical Implications

We argue for three practical principles\.

- •First,*speaker normalization is mandatory*before reporting acoustic ICC: raw estimates are misleading and overestimate the deployment value of prosodic features\.
- •Second,*cross\-context evaluation*\(our LOCO protocol\) should complement within\-corpus cross\-task validation to account for contextual differences; features that collapse on held\-out task types cannot support adaptive systems deployed across conversational contexts\.
- •Third,*interaction features*, despite their moderate predictive accuracy and poor ICC, are the most defensible foundation for measurement systems because their reliability estimates are trustworthy, meaning they generalize more consistently than linguistic features, and they connect directly to observable behavioral events \(floor control, silence, interruption\) during a conversation\.

Note that our findings are grounded in Zoom\-mediated dyadic interaction and thus the results may differ across other media \(e\.g\., in\-person, audio\-only\) or multiparty interaction beyond dyadic settings\. In multi\-party settings, power asymmetries and floor dynamics may be more pronounced, and acoustic normalization artifacts may vary with recording equipment and platform\-specific compression applied\.

Implications for deployment\.The floor imbalance finding carries a concrete system design implication: participation\-aware conversational systems that monitor the distribution of speaking time in real time could identify when one speaker is carrying disproportionate cognitive burden within a dyad, enabling targeted support such as prompting the floor\-dominant speaker to pause, or alerting a facilitator to rebalance participation\. Unlike acoustic features, floor imbalance is computable in real time from voice activity detection alone, without speaker identification or normalization, making it a practical candidate for deployment in low\-latency conversational systems\. Acoustic and linguistic features in their current form require offline ASR and spectral processing, and are not suitable for real\-time deployment without further engineering adaptation\.

Power as a null finding\.The failure to predict conversational power above chance from any feature family, even combined, points to a fundamental limitation of task\-level behavioral aggregation for this target\. Three factors likely contribute to this\. First, the dyadic setup structurally constrains observable power asymmetries: research on multi\-party interaction has consistently shown that dominance indices become more behaviorally evident as group size increases, with speaking time asymmetries, interruption patterns, and floor competition more discriminative in group than dyadic settings\(Vinciarelliet al\.,[2010](https://arxiv.org/html/2607.17452#bib.bib26); Mast,[2002](https://arxiv.org/html/2607.17452#bib.bib36); Danescu\-Niculescu\-Mizilet al\.,[2012](https://arxiv.org/html/2607.17452#bib.bib38)\)\. Second, computational work on power has found stronger signal in structured multi\-party settings such as institutional discourse and online forums than in collaborative dyadic conversation\(Danescu\-Niculescu\-Mizilet al\.,[2012](https://arxiv.org/html/2607.17452#bib.bib38)\)\. Third, power perception is subjective, dynamically negotiated within conversations, and likely requires finer temporal resolution and richer annotation schemes than task\-level self\-reported class labels\.

Recommendations\.For researchers working with existing multimodal datasets where speaker normalization was not applied during feature extraction, we recommend the following retrospective check: z\-score each acoustic feature within speaker across available task observations, then recompute ICC\(2,1\) on the normalized values\. A drop exceeding 50% relative to the raw ICC should be treated as evidence that the reported reliability estimate reflects speaker\-intrinsic vocal characteristics rather than behavioral consistency, and any system design claims based on that estimate should be qualified and re\-assessed accordingly\. This check requires no new data collection and can be applied to any dataset where multiple task observations per speaker are available\.

### 6\.2\.Limitations

This study has several limitations\. The sample is 53 dyads drawn from a single platform \(Zoom\) and collected over a single study protocol\. While the participant pool was diverse from different geographic regions and cultural context, the lack of systematic diversity may limit generalizability\.

The tasks used in this study appear in fixed order\(Sarkaret al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib1)\), confounding task type with session time, which can affect the implications of our findings\. Transcripts are generated by Distil\-Whisper ASR\. Due to ASR limitations, transcription errors may be unequal across the 18 participant countries of origin, potentially reducing the reliability of linguistic features for speakers from underrepresented language backgrounds\(Koeneckeet al\.,[2020](https://arxiv.org/html/2607.17452#bib.bib29)\)\.

For modeling purposes, we only evaluated Random Forest; other model families may show different relative performance across feature families, which we left for future work\. LOCO estimates for the single\-task Problem\-solving category \(n=106n=106\) are unreliable due to the limited size of the dataset and should be replicated with larger held\-out sets\.

### 6\.3\.Future Work

In this work we used dyad level mean as the target for cognitive load\. A two\-speaker latent state model tracking coupled individual load trajectories with an explicit cross\-speaker influence term is the natural extension of these findings and is the focus of ongoing work\. Dynamical time series analysis techniques may capture coordination dynamics that is not accessible to feature averages\(Fusaroliet al\.,[2012](https://arxiv.org/html/2607.17452#bib.bib32)\), which we plan to explore as well\. Speaker\-normalized acoustic reliability should be replicated on other multimodal datasets to assess whether the identity inflation we document is specific to Zoom or a more general artifact of video\-conferencing audio pipelines due to audio compression and other processing steps\.Finally, these findings are based on remote dyadic setup, and we plan to extend this to multi\-party setting as future work\.

## 7\.Conclusion

We introduced a three\-dimensional framework for evaluating multimodal conversational features — prediction accuracy, cross\-task generalizability, and speaker\-normalized test\-retest reliability — and applied it to dyadic Zoom\-based collaboration\. Our central finding is that acoustic reliability is systematically inflated by speaker identity: raw ICC values collapse to near\-zero after normalization, revealing that commonly used prosodic features measure who speaks rather than how speakers behave under load\. Linguistic features predict best but generalize poorly\. Interaction features provide the most reliable signal\. These findings argue that speaker normalization and multi\-dimensional evaluation should be standard prerequisites for reliable feature selection and model evaluation in multimodal conversational AI — particularly for systems intended to operate across diverse task contexts, speaker populations, and conversational environments\.

## Responsible Innovation Statement

This work uses the publicly released AVCAffe dataset\(Sarkaret al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib1)\), collected under ethics approval from the General Research Ethics Board at Queen’s University, Canada\. No new human data were collected as part of this study\. All analyses use anonymized participant identifiers\. We acknowledge that ASR\-based transcription may introduce systematic errors for speakers from underrepresented language backgrounds\(Koeneckeet al\.,[2020](https://arxiv.org/html/2607.17452#bib.bib29)\), and we encourage future work to validate linguistic features using manual or speaker\-adapted transcription\. The models and features evaluated here are not intended for deployment without further reliability validation in target contexts\. Generative AI assistance was used for code and editorial review during the preparation of this manuscript\. The authors have meticulously reviewed all results and text and they take responsibility of all content of this paper\.

###### Acknowledgements\.

We thank the AIIM lab, Queen’s University, Canada, for collecting, preparing, and making this dataset publicly available\. We also thank the anonymous reviewers for their helpful comments to improve this work\. A special thanks to the students of CS466 Multimodal Interaction and Learning at Colby College \(Spring 2026\) for the discussion that motivated this work\. This research was supported by the Henry Luce Foundation\.

## References

- J\. Abel and M\. Babel \(2017\)Cognitive load reduces perceived linguistic convergence between dyads\.Language and Speech60\(3\),pp\. 479–502\.External Links:[Document](https://dx.doi.org/10.1177/0023830916665652)Cited by:[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.1.1.1.1),[§2\.2](https://arxiv.org/html/2607.17452#S2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2607.17452#S5.SS1.p2.1)\.
- S\. Boyer, P\.V\. Paubel, R\. Ruiz, and R\. E\. Yagoubi \(2018\)Human voice as a measure of mental load level\.Journal of Speech, Language, and Hearing Research61\(6\),pp\. 1403–1415\.External Links:[Document](https://dx.doi.org/10.1044/2018%5FJSLHR-S-18-0066)Cited by:[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.2.1.3.1.1),[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1)\.
- C\. Danescu\-Niculescu\-Mizil, L\. Lee, B\. Pang, and J\. Kleinberg \(2012\)Echoes of power: language effects and power differences in social interaction\.InProceedings of the 21st International Conference on World Wide Web,WWW ’12,New York, NY, USA,pp\. 699–708\.External Links:ISBN 9781450312295,[Link](https://doi.org/10.1145/2187836.2187931),[Document](https://dx.doi.org/10.1145/2187836.2187931)Cited by:[§6\.1](https://arxiv.org/html/2607.17452#S6.SS1.p5.1.1.1)\.
- I\. F\. Emara and N\. H\. Shaker \(2024\)The impact of non\-native english speakers’ phonological and prosodic features on automatic speech recognition accuracy\.Speech Commun\.157\(C\)\.External Links:ISSN 0167\-6393,[Link](https://doi.org/10.1016/j.specom.2024.103038),[Document](https://dx.doi.org/10.1016/j.specom.2024.103038)Cited by:[§3\.4](https://arxiv.org/html/2607.17452#S3.SS4.p1.1)\.
- R\. Fusaroli, B\. Bahrami, K\. Olsen, A\. Roepstorff, G\. Rees, C\. Frith, and K\. Tylén \(2012\)Coming to terms: quantifying the benefits of linguistic coordination\.Psychological Science23\(8\),pp\. 931–939\.External Links:[Document](https://dx.doi.org/10.1177/0956797612436816)Cited by:[§6\.3](https://arxiv.org/html/2607.17452#S6.SS3.p1.1)\.
- S\. Gandhi, P\. von Platen, and A\. M\. Rush \(2023\)Distil\-whisper: robust knowledge distillation via large\-scale pseudo labelling\.External Links:2311\.00430Cited by:[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7),[§3\.4](https://arxiv.org/html/2607.17452#S3.SS4.p1.1)\.
- A\. Gravano and J\. Hirschberg \(2011\)Turn\-taking cues in task\-oriented dialogue\.Computer Speech & Language25\(3\),pp\. 601–634\.External Links:[Document](https://dx.doi.org/10.1016/j.csl.2010.10.003)Cited by:[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.2.7.4.3.1.1),[§2\.3](https://arxiv.org/html/2607.17452#S2.SS3.p1.1),[§4\.1](https://arxiv.org/html/2607.17452#S4.SS1.p2.1)\.
- T\. S\. Gunasekaran, K\. Gupta, Y\. S\. Pai, H\. Bai,et al\.\(2025\)CoAffinity: a multimodal dataset for cognitive load and affect assessment in remote collaboration\.IEEE Transactions on Affective Computing16\(03\)\.External Links:[Document](https://dx.doi.org/10.1109/TAFFC.2025.3562559)Cited by:[§2\.4](https://arxiv.org/html/2607.17452#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2607.17452#S2.T1.1.5.4.1)\.
- S\. G\. Hart and L\. E\. Staveland \(1988\)Development of NASA\-TLX \(Task Load Index\): results of empirical and theoretical research\.InAdvances in Psychology,Vol\.52,pp\. 139–183\.External Links:[Document](https://dx.doi.org/10.1016/S0166-4115%2808%2962386-9)Cited by:[§2\.4](https://arxiv.org/html/2607.17452#S2.SS4.p1.1),[§3\.2](https://arxiv.org/html/2607.17452#S3.SS2.p1.4),[footnote 1](https://arxiv.org/html/2607.17452#footnote1)\.
- S\. G\. Hart \(2006\)NASA\-Task Load Index \(NASA\-TLX\): 20 years later\.InProceedings of the Human Factors and Ergonomics Society Annual Meeting,Vol\.50,pp\. 904–908\.External Links:[Document](https://dx.doi.org/10.1177/154193120605000909)Cited by:[§2\.4](https://arxiv.org/html/2607.17452#S2.SS4.p1.1)\.
- M\. Heldner and J\. Edlund \(2010\)Pauses, gaps and overlaps in conversations\.Journal of Phonetics38\(4\),pp\. 555–568\.External Links:[Document](https://dx.doi.org/10.1016/j.wocn.2010.08.002)Cited by:[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.2.2.3.1.1),[§2\.3](https://arxiv.org/html/2607.17452#S2.SS3.p1.1),[§5\.1](https://arxiv.org/html/2607.17452#S5.SS1.p3.1)\.
- M\. Ilse, J\. Tomczak, and M\. Welling \(2018\)Attention\-based deep multiple instance learning\.InProceedings of the 35th International Conference on Machine Learning,J\. Dy and A\. Krause \(Eds\.\),Proceedings of Machine Learning Research, Vol\.80,pp\. 2127–2136\.External Links:[Link](https://proceedings.mlr.press/v80/ilse18a.html)Cited by:[§4\.2](https://arxiv.org/html/2607.17452#S4.SS2.p1.1)\.
- A\. Johnstone, U\. Berry, T\. Nguyen, and A\. Asper \(1995\)There was a long pause: influencing turn\-taking behaviour in human\-human and human\-computer spoken dialogues\.International Journal of Human\-Computer Studies42\(4\),pp\. 383–411\.External Links:[Document](https://dx.doi.org/10.1006/ijhc.1995.1018)Cited by:[§2\.3](https://arxiv.org/html/2607.17452#S2.SS3.p1.1)\.
- M\. A\. Khawaja, F\. Chen, and N\. Marcus \(2012\)Analysis of collaborative communication for linguistic cues of cognitive load\.Human Factors54\(4\),pp\. 518–529\.External Links:[Document](https://dx.doi.org/10.1177/0018720811431258)Cited by:[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.8.7.3.1.1),[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.1.1.1.1),[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.12.10.3.1.1),[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.5.3.3.1.1),[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.6.4.3.1.1),[§2\.2](https://arxiv.org/html/2607.17452#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2607.17452#S2.T1.1.2.1.1),[§4\.1](https://arxiv.org/html/2607.17452#S4.SS1.p4.1),[§5\.1](https://arxiv.org/html/2607.17452#S5.SS1.p2.1)\.
- A\. Koenecke, A\. Nam, E\. Lake, J\. Nudell, M\. Quartey, Z\. Mengesha, C\. Toups, J\. R\. Rickford, D\. Jurafsky, and S\. Goel \(2020\)Racial disparities in automated speech recognition\.Proceedings of the National Academy of Sciences117\(14\),pp\. 7684–7689\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1915768117)Cited by:[§2\.5](https://arxiv.org/html/2607.17452#S2.SS5.p1.1),[§3\.4](https://arxiv.org/html/2607.17452#S3.SS4.p1.1),[§6\.2](https://arxiv.org/html/2607.17452#S6.SS2.p2.1),[Responsible Innovation Statement](https://arxiv.org/html/2607.17452#Sx1.p1.1)\.
- T\. K\. Koo and M\. Y\. Li \(2016\)A guideline of selecting and reporting intraclass correlation coefficients for reliability research\.Journal of Chiropractic Medicine15\(2\),pp\. 155–163\.External Links:[Document](https://dx.doi.org/10.1016/j.jcm.2017.10.001)Cited by:[Figure 1](https://arxiv.org/html/2607.17452#S1.F1),[§2\.5](https://arxiv.org/html/2607.17452#S2.SS5.p1.1),[§4\.3\.3](https://arxiv.org/html/2607.17452#S4.SS3.SSS3.p1.4),[Figure 6](https://arxiv.org/html/2607.17452#S5.F6),[Table 4](https://arxiv.org/html/2607.17452#S5.T4)\.
- S\. C\. Levinson \(2016\)Turn\-taking in human communication — origins and implications for language processing\.Trends in Cognitive Sciences20\(1\),pp\. 6–14\.External Links:[Document](https://dx.doi.org/10.1016/j.tics.2015.10.010)Cited by:[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.2.11.8.3.1.1),[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.2.5.2.3.1.1),[§4\.1](https://arxiv.org/html/2607.17452#S4.SS1.p2.1)\.
- L\. I\. Lin \(1989\)A concordance correlation coefficient to evaluate reproducibility\.Biometrics45\(1\),pp\. 255–268\.External Links:[Document](https://dx.doi.org/10.2307/2532051)Cited by:[Figure 3](https://arxiv.org/html/2607.17452#S4.F3),[§4\.3](https://arxiv.org/html/2607.17452#S4.SS3.p1.1.1.1),[§5\.1](https://arxiv.org/html/2607.17452#S5.SS1.p1.1.1.1)\.
- M\. S\. Mast \(2002\)Dominance as expressed and inferred through speaking time: a meta\-analysis\.Human Communication Research28\(3\),pp\. 420–450\.External Links:[Document](https://dx.doi.org/10.1111/j.1468-2958.2002.tb00814.x)Cited by:[§6\.1](https://arxiv.org/html/2607.17452#S6.SS1.p5.1.1.1)\.
- M\. K\. Ngueajio and G\. Washington \(2022\)Hey asr system\! why aren’t you more inclusive? automatic speech recognition systems’ bias and proposed bias mitigation techniques\. a literature review\.InHCI International 2022 – Late Breaking Papers: Interacting with EXtended Reality and Artificial Intelligence: 24th International Conference on Human\-Computer Interaction, HCII 2022, Virtual Event, June 26 – July 1, 2022, Proceedings,Berlin, Heidelberg,pp\. 421–440\.External Links:ISBN 978\-3\-031\-21706\-7,[Link](https://doi.org/10.1007/978-3-031-21707-4_30),[Document](https://dx.doi.org/10.1007/978-3-031-21707-4%5F30)Cited by:[§3\.4](https://arxiv.org/html/2607.17452#S3.SS4.p1.1)\.
- H\. Sacks, E\. A\. Schegloff, and G\. Jefferson \(1974\)A simplest systematics for the organization of turn\-taking for conversation\.Language50\(4\),pp\. 696–735\.External Links:[Document](https://dx.doi.org/10.2307/412243)Cited by:[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.1.1.3.1.1),[§4\.1](https://arxiv.org/html/2607.17452#S4.SS1.p2.1)\.
- P\. Sarkar, A\. Posen, and A\. Etemad \(2023\)AVCAffe: a large scale audio\-visual dataset of cognitive load and affect for remote work\.InProceedings of the AAAI Conference on Artificial Intelligence,AAAI’23/IAAI’23/EAAI’23, Vol\.37,pp\. 76–85\.External Links:ISBN 978\-1\-57735\-880\-0,[Document](https://dx.doi.org/10.1609/aaai.v37i1.25078)Cited by:[§1](https://arxiv.org/html/2607.17452#S1.p3.1),[§2\.4](https://arxiv.org/html/2607.17452#S2.SS4.p1.1),[Table 1](https://arxiv.org/html/2607.17452#S2.T1.1.4.3.1),[§3\.1](https://arxiv.org/html/2607.17452#S3.SS1.p1.1),[§6\.2](https://arxiv.org/html/2607.17452#S6.SS2.p2.1),[Responsible Innovation Statement](https://arxiv.org/html/2607.17452#Sx1.p1.1)\.
- B\. Schuller, S\. Steidl, A\. Batliner, F\. Burkhardt, L\. Devillers, C\. Müller, and S\. Narayanan \(2013\)The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism\.InProceedings of Interspeech,pp\. 148–152\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2013-56)Cited by:[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.4.3.3.1.1),[§1](https://arxiv.org/html/2607.17452#S1.p1.1)\.
- B\. Schuller, S\. Steidl, A\. Batliner, J\. Epps,et al\.\(2014\)The Interspeech 2014 computational paralinguistics challenge: cognitive & physical load, multitasking\.InProceedings of Interspeech,pp\. 427–431\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2014-130)Cited by:[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1)\.
- R\. Shwartz\-Ziv and A\. Armon \(2022\)Tabular data: deep learning is not all you need\.Information Fusion81,pp\. 84–90\.External Links:[Document](https://dx.doi.org/10.1016/j.inffus.2021.11.011)Cited by:[§4\.3\.1](https://arxiv.org/html/2607.17452#S4.SS3.SSS1.p1.1)\.
- N\. Taptiklis, M\. Su, J\.H\. Barnett, and C\. Skirrow \(2023\)Prediction of mental effort derived from an automated vocal biomarker using machine learning in a large\-scale remote sample\.Frontiers in Artificial Intelligence6\.External Links:[Document](https://dx.doi.org/10.3389/frai.2023.1171652)Cited by:[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.11.10.3.1.1),[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1)\.
- Y\. R\. Tausczik and J\. W\. Pennebaker \(2010\)The psychological meaning of words: LIWC and computerized text analysis methods\.Journal of Language and Social Psychology29\(1\),pp\. 24–54\.External Links:[Document](https://dx.doi.org/10.1177/0261927X09351676)Cited by:[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.17.15.3.1.1),[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.3.1.3.1.1),[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.8.6.3.1.1),[§4\.1](https://arxiv.org/html/2607.17452#S4.SS1.p4.1),[§5\.1](https://arxiv.org/html/2607.17452#S5.SS1.p3.1)\.
- A\. Vinciarelli, M\. Pantic, and H\. Bourlard \(2010\)Social signal processing: survey of an emerging field\.Signal Processing90\(5\),pp\. 2130–2150\.External Links:[Document](https://dx.doi.org/10.1016/j.sigpro.2009.11.014)Cited by:[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.2.14.11.3.1.1),[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.2.4.1.3.1.1),[Table 5](https://arxiv.org/html/2607.17452#Ax1.T5.2.8.5.3.1.1),[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.5.4.3.1.1),[Table 7](https://arxiv.org/html/2607.17452#Ax1.T7.1.15.13.3.1.1),[§1](https://arxiv.org/html/2607.17452#S1.p1.1),[§2\.3](https://arxiv.org/html/2607.17452#S2.SS3.p1.1),[§6\.1](https://arxiv.org/html/2607.17452#S6.SS1.p5.1.1.1)\.
- M\. Vukovic, M\. Stolar, and M\. Lech \(2021\)Cognitive load estimation from speech commands to simulated aircraft\.IEEE/ACM Transactions on Audio, Speech, and Language Processing29,pp\. 1011–1022\.External Links:[Document](https://dx.doi.org/10.1109/TASLP.2021.3057492)Cited by:[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.2.1.3.1.1),[§1](https://arxiv.org/html/2607.17452#S1.p1.1),[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2607.17452#S2.T1.1.3.2.1)\.
- H\. Xu, L\. Wang, J\. Zou, J\. Zhang, R\. Li,et al\.\(2025\)Recognising and explaining mental workload using low\-interference method by fusing speech, ECG and eye tracking signals during simulated flight\.Ergonomics\.External Links:[Document](https://dx.doi.org/10.1080/00140139.2025.2511877)Cited by:[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1)\.
- J\. Yang, H\. Yang, Z\. Wu, and X\. Wu \(2023\)Cognitive load assessment of air traffic controller based on SCNN\-TransE network using speech data\.Aerospace10\(7\)\.External Links:[Document](https://dx.doi.org/10.3390/aerospace10070584)Cited by:[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1)\.
- B\. Yin, F\. Chen, and N\. Ruiz \(2008\)Speech\-based cognitive load monitoring system\.InIEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 2041–2044\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP.2008.4518756)Cited by:[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.11.10.3.1.1),[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1)\.
- B\. Yin, N\. Ruiz, F\. Chen, and M\.A\. Khawaja \(2007\)Automatic cognitive load detection from speech features\.InProceedings of the 19th Australasian Conference on Computer\-Human Interaction \(OZCHI\),New York, NY, USA\.External Links:ISBN 9781595938725,[Document](https://dx.doi.org/10.1145/1324892.1324946)Cited by:[Table 6](https://arxiv.org/html/2607.17452#Ax1.T6.3.7.6.3.1.1),[§2\.1](https://arxiv.org/html/2607.17452#S2.SS1.p1.1)\.
- J\. Zhou, K\. Yu, F\. Chen, Y\. Wang, and S\. Z\. Arshad \(2018\)Multimodal behavioral and physiological signals as indicators of cognitive load\.InThe Handbook of Multimodal\-Multisensor Interfaces,Vol\.2\.External Links:[Document](https://dx.doi.org/10.1145/3107990.3108002)Cited by:[§2\.2](https://arxiv.org/html/2607.17452#S2.SS2.p1.1)\.

## Appendix: Feature Definitions

Tables[5](https://arxiv.org/html/2607.17452#Ax1.T5)–[7](https://arxiv.org/html/2607.17452#Ax1.T7)provide definitions for all 44 unique features used in this study, organized by feature family\. All features are computed at the task level \(one value per speaker per task\)\. Interaction features are dyad\-level \(identical for both speakers within a dyad\-task\); acoustic and linguistic features are per\-speaker\.

Table 5\.Interaction features \(14 features, dyad\-level\)\. Aggregated from 30\-second window\-level features across each task\. These features are structurally relative — they capture the asymmetry or joint behavior between the two speakers — making them immune to speaker identity inflation in our analysis\.Table 6\.Acoustic features \(10 features, per speaker, task\-level\)\. Computed from per\-speaker audio recordings using voice activity detection and pitch extraction\.Note:As shown in Table 4 of the main paper, raw ICC values for spectral features \(pitch mean, spectral centroid, spectral bandwidth\) collapse to near\-zero after within\-speaker normalization, indicating they primarily reflect stable vocal tract characteristics rather than task\-sensitive behavioral states\.Table 7\.Linguistic features \(21 features, per speaker, task\-level\)\. Extracted from per\-speaker task\-level transcripts generated by Distil\-Whisper ASR\(Gandhiet al\.,[2023](https://arxiv.org/html/2607.17452#bib.bib28)\)\. Rate features are normalized by total word count or segment count\.

Similar Articles

Multimodal Speaker Identification in Classroom Environments

arXiv cs.CL

This paper evaluates a multimodal framework for speaker identification in K-12 classrooms by combining acoustic embeddings (ECAPA-TDNN) with LLM-derived semantic context from transcripts, improving accuracy from 39% to 50.3% overall and from 64.9% to 76.9% for longer utterances.

Evaluating multimodal emotion recognition in proactive conversational agents: A user study

arXiv cs.AI

This paper presents a multimodal emotion recognition module for proactive conversational agents, using facial recognition and linguistic analysis. A user study with 20 participants reveals a 'poker face' effect where visual cues are unreliable, while linguistic analysis proves more accurate; the study also shows agents can elicit emotions through conversational adaptation.

Token Statistics Reveal Conversational Drift in Multi-turn LLM Interaction

arXiv cs.CL

This paper introduces Bipredictability (P) and the Information Digital Twin (IDT), a lightweight method to monitor conversational consistency in multi-turn LLM interactions using token frequency statistics without embeddings or model internals. The approach achieves 100% sensitivity in detecting contradictions and topic shifts while establishing a practical monitoring framework for extended LLM deployments.