Acoustic and Facial Markers of Perceived Conversational Success in Spontaneous Speech
Summary
This paper analyzes spontaneous dyadic Zoom conversations using multimodal features (acoustic, facial, turn-taking) to identify markers of perceived conversational success, finding that entrainment in speech and facial movements correlates with higher interaction quality.
View Cached Full Text
Cached at: 04/20/26, 08:30 AM
# Acoustic and Facial Markers of Perceived Conversational Success in Spontaneous Speech Source: https://arxiv.org/html/2604.15322 ###### Abstract Individuals often align their speaking patterns with their interlocutors, a phenomenon linked to engagement and rapport. While well documented in task-oriented dialogues, less is known about entrainment in naturalistic, non-task and virtual settings. In this study, we analyze a large corpus of spontaneous dyadic Zoom conversations to examine how conversational dynamics relate to perceived interaction quality. We extract multimodal features encompassing turn-taking, pauses, facial movements, and acoustic measures such as pitch and intensity. Perceived conversational success was quantified via factor analysis of post-conversation ratings. Results demonstrate that entrainment is reliably detected in spontaneous speech and correlates with higher perceived success. These findings identify key interactional markers of conversational quality and highlight opportunities for targeted interventions to foster more effective and engaging communication. Index Terms—entrainment, facial, acoustic, turn, pause ## 1 Introduction Social interaction is vital for communication, bonding, and well-being. Partners influence each other across speech, language, and visual cues[18, 13]. In spontaneous get-to-know-you conversations, turn dynamics may reveal interaction quality: shorter turns may reflect alignment or disengagement, while longer turns can indicate comfort and topic engagement. Notably, longer gaps foster connection among friends but have shown disconnection among strangers[14]. In conversation, interlocutors adapt by mirroring emotions, prosody, and body movements, a process termed alignment, entrainment, or mimicry[16]. High entrainment, measured as time-locked behavioral alignment, enhances interaction quality[9, 12]. Both verbal and non-verbal cues are critical for successful interaction[13], which further depends on shared attentiveness, coordination, and mood. Research shows individuals can learn to entrain through targeted training, for example in children with speech difficulties[15]. Identifying alignment features that predict conversational success is therefore vital for developing effective training methods. Entrainment of acoustic-prosodic features such as pitch, intensity, speaking rate, and loudness has been shown to predict conversation quality in task-oriented settings[9, 7, 3]. Non-verbal features including body movements, eye gaze, facial expressions, and head movements also predict interaction quality[10, 2]. However, a major gap remains in understanding unstructured, naturalistic conversations between adult peers in extended Zoom interactions. Prior research has largely focused on task-oriented contexts, such as computer games[8], collaborative problem solving in education[9], and user-agent interactions[5], primarily through acoustic-prosodic features. Other studies have examined face-to-face dynamics in spontaneous code-switching[4] or short get-to-know-you interactions with autistic children[17], but the focus here is on adult peers conversing in the same language in long virtual settings. Most prior spontaneous conversation studies were in-person, yet many contemporary peer interactions including job interviews now occur remotely. Understanding how entrainment during remote communication aligns with conversational success is therefore critical. Our study investigates whether multimodal entrainment predicts self-reported enjoyment and alignment, with the long-term goal of identifying markers of conversation quality in both neurotypical and mixed-neurotype interactions. The main contributions of this work are: we show that features including facial action units, pitch, intensity, turn count, and pause duration are associated with perceived conversational success. We analyzed a large corpus of naturalistic conversations held via Zoom and show that entrainment occurs and correlates with conversation success. In this work, we are focusing on conversational alignment. First, we initialize our study by observing general conversation dynamics like turn and pause trends. Next, utilizing features such as facial expressions, pitch, and intensity, we fine-tune our analysis into temporal window-based and turn-based analysis to investigate whether speaker entrainment is occurring across time. Finally, we aim to see whether there is an association between the aforementioned features and perceived conversation success. ## 2 Dataset For this study we analyze the CANDOR Corpus (Conversation: A Naturalistic Dataset of Online Recordings)[12], collected by BetterUp Labs in collaboration with researchers at the University of Pennsylvania (2023). CANDOR comprises more than 1500 spontaneous, dyadic, 30-minute video and audio recorded conversations conducted over Zoom between previously unacquainted adults aged 19–66 who represent a broad range of gender, educational, ethnic, and generational identities. Interactions were unstructured and non-task-oriented. After each conversation, both participants independently completed a post-conversation survey rating multiple aspects of perceived interaction quality; these ratings were used to construct a composite success score (section 3.1). Separate audio channels were recorded for each participant, obviating the need for speaker diarization. For all analyses, we restricted the sample to sessions flagged by the dataset as free of background noise or interruptions. ## 3 Method ### 3.1 Perceived conversation success Each participant completed both pre- and post-conversation survey questionnaires designed to assess the quality of the interaction. The full instrument contained 229 items, including demographic and speaker-related details. For the purposes of this study, we focused on a subset of 21 constructs that capture dimensions of affect, enjoyability, friendship, and common ground. To identify relevant dimensions of perceived conversational success (PCS), we conducted a principal component analysis (PCA) on these constructs. An initial exploratory PCA revealed two latent dimensions with loadings exceeding 0.4. Accordingly, we performed a subsequent PCA constrained to two dimensions (PCA₁, PCA₂), retaining constructs with loadings above 0.4. Although both components were interpretable, only PCA₁ was adopted as the basis for the PCS measure. This choice was motivated by its stronger discriminative utility: when analyzed separately, constructs grouped under PCA₁ exhibited significantly greater separation on PCS in terms of feature-level entrainment, whereas PCA₂ did not yield comparable distinctions. Responses to the 11 constructs in PCA₁—Affect, Overall affect, Affect at beginning, Affect at middle, Affect at end, Best affect, How much enjoyable, I like you, You like me, Conversationalist, My friends like you—were originally recorded on heterogeneous scales (1–7, 1–9, or 1–100) and normalized to a common range prior to analysis. Ratings were z-score normalized within each construct and averaged across constructs to yield an individual PCS score bounded between 0 and 1. To reduce label ambiguity and enable a high-contrast assessment of entrainment-related discriminative validity, we focused on conversations at the extreme ends of the PCS distribution, retaining only those more than one standard deviation from the median. Given the overall high enjoyment levels, this corresponded to PCS ≤ 0.6 for Low-Successful Conversations (LSCs) and PCS ≥ 0.9 for High-Successful Conversations (HSCs), resulting in 35 LSCs and 91 HSCs. ### 3.2 Turn-taking analysis While turn-taking is a well-established conversational norm[14], the fine-grained structure of turn exchanges is complex, and varied patterns can nonetheless produce highly successful interactions. To test these hypotheses, we derived turn-level measures from the dataset's transcripts, which exclude backchannel utterances from turn units and thus provide a clearer operational definition of conversational turns. While no single definition of a turn is universally accepted, we follow the dataset's convention (turn is defined as a contiguous stretch of speech by one speaker bounded by a change of floor) for computational consistency[12]. For each conversation (including turns from both speakers), we computed summary statistics of turn duration: minimum, maximum, mean, total (sum across all turns), and overall turn count using the annotated start and end times from the transcripts. We then examined their association with PCS. In parallel, we quantified inter-turn silence, defined as the pause (silences exceeding 0.6 s considered significant[12]) between one speaker's offset and the other speaker's onset. Using the same timing annotations, we calculated the minimum, maximum, mean, and total pause duration per conversation and evaluated their relationship to PCS. ### 3.3 Acoustic features analysis (a) Turn duration statistics (b) Pause duration statistics (c) Turn count Fig. 1: Features vs. PCS. (a) Turn duration (min, max, mean, total). (b) Pause duration (min, max, mean, total). (c) Turn count. Each speaker's audio was recorded on an independent channel. All audio recordings were first converted from dual-channel to mono and downsampled from 44.1 kHz to 16 kHz to standardize the signal representation. Features were segmented into speaker turns based on transcript-provided timing boundaries. To analyze acoustic dynamics, we extracted the pitch (F₀) from each speaker's turn-level audio using the Pitch Estimating Neural Network (PENN)[11], which has demonstrated robust performance, including detection of F₀ during creaky-voice regions. Pitch trajectories were subsequently normalized to reduce gender-related variability. In addition, speech intensity was computed for each turn using Praat. For the statistical analysis, minimum, maximum, and mean of each turn's acoustic feature were computed and their entrainment along the conversation duration was tracked. #### 3.3.1 Turn-level proximity entrainment in acoustic features To quantify acoustic entrainment, we calculate turn-level proximity[8], which is the phenomenon where adjacent turns should lie in close proximity compared to non-adjacent turns. First, for each conversation, we indexed the turns in temporal order i = 1, ..., N. For a given acoustic feature statistic f (we use turn-level summary statistics: minimum, maximum, and mean of F₀ and intensity), we calculated the feature value for the current speaker on turn i, denoted by fc_i, and the corresponding feature value for the partner on the next turn fp_(i+1). Adjacent-turn distance is defined as the absolute difference between those two as in equation 1: fd_a(i) = |fc_i − fp_(i+1)| (1) To obtain a non-adjacent baseline, we randomly choose a partner turn's f that is non-adjacent to turn i (fp_j≠i) and took the absolute difference of that with fc_i. This process is repeated a total of 10 times and the average of those 10 differences is computed as shown in equation 2: fd_na(i) = (Σⱼ₌₁¹⁰ |fc_i − fp_(j≠i)|) / 10 (2) which we refer to as the non-adjacent distance. If entrainment is present, adjacent differences fd_a(i) should be smaller than their non-adjacent counterparts fd_na(i). We compute fd_a(i) and fd_na(i) for every turn in every conversation and for each feature statistic f. The resulting distance distributions were compared using the Mann–Whitney U test to evaluate whether adjacent-turn distances were systematically smaller than non-adjacent baselines. Shapiro–Wilk tests revealed significant deviations from normality across features, particularly in HSCs. For example, minimum pitch showed strong non-normality in both groups (LSC: p = 5.15e−8, HSC: p = 4.16e−17). Given this pattern, the Mann–Whitney U test was applied for all analyses. ### 3.4 Facial expression analysis To study entrainment of facial expressions we extracted Facial Action Units (FAUs) using the OpenFace open-source facial behavior analysis toolkit[1], which processes video recordings of each speaker and outputs a wide range of facial movement indicators. For this study, we used the 17 FAUs (mentioned in Table 2) extracted from the default OpenFace settings. #### 3.4.1 Synchrony in facial action units While proximity reflects the extent to which a given feature is similar in magnitude across interlocutors, synchrony captures the temporal alignment of feature trajectories, even when their absolute values may differ. In this work, we investigate whether the synchrony of FAUs aligns with PCS. Specifically, we computed Pearson correlations between the same FAU across the two speakers within each conversation, using non-overlapping 5-second windows. This continuous, fixed-window approach was chosen in place of turn-based segmentation, as meaningful emotional expressions can also occur during pauses, which would otherwise be excluded in turn-based analysis[3]. When participants mirror each other's facial expressions, the correlation between their facial action units increases. To capture this phenomenon, raw correlation values were first converted to Fisher z-transformed values (z_f). We then computed the mean of z_f across each conversation. This measure was calculated independently for each FAU and subsequently examined in relation to PCS ratings. ## 4 Results ### 4.1 Turns and pauses vs. PCS To assess group differences, we conducted Mann–Whitney U tests for each f with the null hypothesis (H₀) that distributions are similar and alternative hypothesis (H₁) that the distributions differ. The corresponding U, z, p, and q (Benjamini–Hochberg False Discovery Rate corrected value) are reported in Table 1. Figure 1(a) shows that in HSCs, maximum, mean, and total turn durations...
Similar Articles
How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks
This paper evaluates the predictive accuracy, cross-task generalizability, and test-retest reliability of multimodal features for measuring conversational states like cognitive load and power in dyadic remote collaborative tasks. Findings show that linguistic features predict well but generalize poorly, acoustic reliability degrades when controlling for speaker identity, and interaction features provide the most reliable signal.
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
This paper analyzes multimodal systems for the CHiME-9 MCoRec cocktail-party scenario, comparing design strategies such as audio-visual target speech separation, improved recognition, and LLM-based conversational grouping, finding that speech overlap alone does not explain performance differences.
Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating
This paper investigates whether speech signals can complement LLM-based prediction of interpersonal attraction from conversation transcripts, using Japanese speed-dating data. Results show conditional improvements in prediction accuracy when combining speech and transcript-based models.
Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives
This paper investigates the use of conversational temporal dynamics (turn-pair timing) as a lightweight modality for automatic depression detection from dyadic clinical interviews, showing that a compact 24-dimensional timing module achieves strong performance and complements standard acoustic and semantic features when fused.
Evaluating multimodal emotion recognition in proactive conversational agents: A user study
This paper presents a multimodal emotion recognition module for proactive conversational agents, using facial recognition and linguistic analysis. A user study with 20 participants reveals a 'poker face' effect where visual cues are unreliable, while linguistic analysis proves more accurate; the study also shows agents can elicit emotions through conversational adaptation.