Evaluating multimodal emotion recognition in proactive conversational agents: A user study
Summary
This paper presents a multimodal emotion recognition module for proactive conversational agents, using facial recognition and linguistic analysis. A user study with 20 participants reveals a 'poker face' effect where visual cues are unreliable, while linguistic analysis proves more accurate; the study also shows agents can elicit emotions through conversational adaptation.
View Cached Full Text
Cached at: 05/22/26, 08:51 AM
# Evaluating multimodal emotion recognition in proactive conversational agents: A user study Source: [https://arxiv.org/html/2605.20200](https://arxiv.org/html/2605.20200) \\transtitle\\subtranstitle \[https://orcid\.org/0000\-0002\-4773\-4904\] \[https://orcid\.org/0000\-0003\-1231\-7235\] \[https://orcid\.org/0000\-0002\-6137\-9558\] \\corres F\. Xavier Gaya\-Morey \(\) \\transkeywords Raquel LacuestaF\. Xavier Gaya\-MoreyJose M\. Buades\-Rubio\\orgdivEscuela Universitaria Politécnica de Teruel,\\orgnameUniversidad de Zaragoza,\\orgaddressC/ Atarazana, 2,\\cityTeruel,\\postcode44003,\\stateAragón,\\countrySpain\\orgdivI3A \(Institute of Engineering Research of Aragon\),\\orgnameUniversidad de Zaragoza,\\orgaddressC\. de Mariano Esquillor Gómez, s/n,\\cityZaragoza,\\postcode50018,\\stateAragón,\\countrySpain\\orgdivUniversitat de les Illes Balears,\\orgnameUniversitat de les Illes Balears,\\orgaddressCarretera de Valldemossa, km 7\.5,\\cityPalma,\\postcode07122,\\stateIlles Balears,\\countrySpain[francesc\-xavier\.gaya@uib\.es](https://arxiv.org/html/2605.20200v1/mailto:[email protected]) \(–\) ###### Abstract \[ABSTRACT\]This article presents a multimodal emotion recognition module integrated into a proactive Socially Interactive Agent \(SIA\) powered by generative artificial intelligence\. The system evaluates real\-time affective states through two distinct channels: a computer vision\-based facial recognition module and a semantic linguistic analysis engine\. To validate the framework, an empirical study was conducted with 20 users who engaged in dynamic, unscripted dialogues with the conversational agent\. The findings reveal a significant discrepancy between automated visual cues and actual internal emotional states\. When interacting with the AI, users consistently exhibited a “poker face” effect, displaying serious, concentrated facial expressions even when experiencing positive emotions\. Consequently, the generative AI linguistic analysis proved significantly more reliable, by contextualizing the users’ verbal expressions\. Furthermore, an analysis of the interaction dynamics demonstrated that SIAs can effectively elicit specific emotions by adapting conversational themes and employing structured linguistic patterns, such as empathetic or humorous language\. However, the study also noted that instances of uncalibrated proactivity occasionally led to user disengagement and a perception of artificiality\. Ultimately, this research highlights the necessity of refining SIAs to dynamically adapt to users’ emotional evolution, relying on deep linguistic context to foster more natural, human\-like interactions\. \\transabstract \[transABSTRACT\] ###### keywords: facial recognition \| linguistic analysis \| socially interactive agents \| affective computing ††articletype:ORIGINAL ARTICLE††journal:arXiv††volume:0## 1Introduction Understanding and responding to emotions is a fundamental component of human communication\. For Socially Interactive Agents \(SIAs\) to establish meaningful connections with users, they must be capable of recognizing and adapting to the user’s emotional state in real timefeng2022emowoz\. Integrating this affective awareness is essential for humanizing conversational agents, fostering natural dialogues, and ultimately improving their overall user acceptancegao2022emotion\. Currently, emotion detection systems rely on a variety of modalities, ranging from facial expression recognition and voice processing to gaze tracking and the analysis of biological signals, such as electroencephalograms \(EEG\) and electrocardiograms \(ECG\)siam2022deploying,garcia\-magarino2018agent\-based\. Recent research has demonstrated significant progress in multimodal emotion recognition applied to social agents\. For example,makiuchi2021multimodalproposed a novel cross\-representation model combining speech and text that effectively outperforms traditional speech\-only or text\-only unimodal approacheschrist2023muse\. Similarly,katada2023effectsevaluated the fusion of physiological signals with other modalities to estimate different types of sentiment during naturalistic human\-agent interactions\. Despite these technical advancements, most existing studies rely on pre\-recorded datasets, theoretical simulations, or passive observation\. There is a notable lack of research integrating multimodal emotion detection within dynamic, real\-time conversations driven by generative AI\. Specifically, current literature does not sufficiently address how proactive, AI\-generated dialogues attempt to connect with specific user emotions, nor how these interactions directly influence the user’s facial expressions and internal affective states during a live exchange\. To bridge this gap, this study presents the design and empirical evaluation of a proactive conversational SIA\. The proposed system integrates two core emotion detection techniques: a computer vision module for facial expression recognition and a generative AI\-based linguistic analysis module\. Through practical, real\-time evaluations with users interacting with the SIA, this research investigates how human emotions actually manifest during unscripted interactions\. By comparing automated visual data, semantic linguistic analysis, and subjective user self\-reports, this study highlights critical interaction patterns and evaluates the real\-world accuracy of these emotion detection methods\. The remainder of this article is organized as follows: Section[2](https://arxiv.org/html/2605.20200#S2)reviews related work in emotion recognition and SIAs\. Section[3](https://arxiv.org/html/2605.20200#S3)outlines the research objectives and specific research questions\. Section[4](https://arxiv.org/html/2605.20200#S4)details the methodology, including the system design and evaluation procedure\. Section[5](https://arxiv.org/html/2605.20200#S5)presents the objective results of the multimodal evaluation\. Section[6](https://arxiv.org/html/2605.20200#S6)discusses these findings and answers the core research questions\. Section[7](https://arxiv.org/html/2605.20200#S7)acknowledges the study’s limitations and outlines future work, and Section[8](https://arxiv.org/html/2605.20200#S8)provides the final conclusions\. ## 2Related Work The success of interactions between humans and artificial agents relies heavily on emotional identity and the understanding of cultural affect\. These elements allow agents to maintain emotional coherence, transforming their perception from mere machines into entities capable of establishing meaningful connectionshoey2016affect,malhotra2021emotions\. As argued byjoby2022effect, the ability of agents to engage in emotional contagion, the transmission of affective states between participants, is vital for enhancing trust, empathy, and prosocial orientation during social exchanges\. ### 2\.1Emotional Intelligence in Socially Interactive Agents SIAs are increasingly integrated into diverse domains, necessitating the incorporation of emotional intelligence to improve human\-machine dynamics\. Emotions are not merely aesthetic; they are essential for creating realistic behaviors in social simulations by integrating cognition and social relationships\. For instance,alanazi2023predictionfocused on simulating emotions to foster empathy, thereby enhancing agent autonomy and adaptability\. Similarly,samsonovich2014developingexplored emotionally intelligent virtual agents through biologically inspired cognitive architectures to replicate human social behavior\. Further research byerol2020artificialemphasizes the importance of recognizing human emotional states to improve bonding, proposing perception architectures for human\-robot interaction\. In parallel,tavabi2019multimodalinvestigated multimodal deep neural networks to identify opportunities for empathetic responses\. Despite these advances, some authors argue that a significant gap remains: most existing models focus on theoretical simulations and Affect Control Theory without addressing the challenges of real\-time processing in dynamic interactionsmalhotra2021emotions,hoey2016affect\. Many studies, such as those bycipresso2012realandhortensius2018perception, primarily review how humans perceive emotions in agents but do not delve into how these perceptions can be processed and responded to by the agents themselves in live settings\. ### 2\.2Facial Expression Recognition and its Limitations Facial Expression Recognition \(FER\) is a fundamental pillar for detecting user states\. Early deep learning approaches, such as those discussed bykalyani2023smart, highlighted the potential of analyzing expressions alongside speech and text\.wang2020humandemonstrated that bimodal approaches, combining facial features with speech, consistently outperform unimodal systems\. Modern FER systems use sophisticated models like cGANsdeng2019cgan, semantic\-rich frameworkschen2022semantic\-rich, and efficient architectures like SwishNetdar2022efficient\-swishnetto facilitate natural interaction\. However, the transition from laboratory settings to “in\-the\-wild” applications presents severe challenges\.rani2014emotionanddalvi2021surveypoint out that variations in lighting, facial angles, and cultural diversity require more robust models\. While recent developments like LiteFeryang2024liteferand video representation learningstrizhkova2024videoaim to improve efficiency on limited\-resource devices, a deeper technical unreliability persists\.cabitza2022unbearableandkusal2024understandinghave criticized the inconsistency of automated FER when dealing with diverse observer samples\. Crucially,wang2024surveyandsamadiani2019reviewsuggest that serious or neutral facial expressions, common in social interactions, often lead to unreliable detection\. This “ambiguity of expression” suggests that relying solely on visual data may be insufficient for SIAs in real\-world social contexts\. ### 2\.3Text\-Based Emotion Detection and Generative AI Text\-Based Emotion Detection \(TBED\) has become a critical component for developing empathetic agentskusal2024understanding,kusal2022review\. While it has proven valuable for big data analytics in social mediakusal2021ai, applying TBED to real\-time conversation introduces complexities such as handling short texts, synonyms, and reversed word ordermaruf2024challenges,wen2024personality\-affected\. Traditional machine learning approaches have made significant progressmachova2023detection, yet full automation remains a challengemaruf2024challenges\. The emergence of Generative AI offers a paradigm shift\.bertero2016real\-timeproposed real\-time sentiment recognition to enable dialogue systems to respond appropriately\. More recently, the work of Park et al\.park2023generativeon “Generative Agents” demonstrates that large language models \(LLMs\) can simulate believable human behavior and social reflections\. However, as noted byghaffarzadegan2024generative, these feedback\-rich computational models are generally used to generate behavior rather than to analyze the user’s internal state during the interaction\. There is a clear need for studies that utilize generative AI to interpret the fluctuations of user emotions across different topics in real\-time, especially when visual cues are absent or misleadingpeng2020human\-machine\. ### 2\.4Multimodal Integration and User Experience The consensus in the field is that multimodal systems, integrating voice, face, and text, yield the highest accuracyalonso\-martin2013multimodal\.katada2024collectingeven explored the use of frontal brain signals to capture “unexpressed sentiments,” highlighting that users often feel more than they show\.ge2024modelingfurther emphasized that understanding the dependency between the speaker and the sentiment is vital for context\-aware recognition\. The interaction itself is a reciprocal process\.woo2023reciprocalargue that SIAs must adapt their behaviors dynamically, acting as both speakers and listeners\. While agents can induce emotions like happiness or angertanioka2025dialogue,gupta2024facial,alonso\-martin2013multimodal, the complexity of emotional signals remains a barrier to a truly positive user experiencesamsonovich2014developing,skillicorn2019measuring\.wen2024personality\-affectedsuggest that incorporating personality traits can lead to more engaging responses, yet many of these findings are based on theoretical testingli2024sia\-netor lack involvement with real\-time AI\-generated conversational enginesorlov2024real\-time\. This research addresses these limitations by developing a proactive SIA that integrates an AI system to both generate and identify emotions in real\-time\. By comparing “what the user says they feel” \(self\-report\) with visual \(SilNet\) and linguistic \(Generative AI\) analysis, we provide a critical evaluation of the “poker face effect\.” Unlike previous research focused on simulations, our work emphasizes a practical application that explores how linguistic analysis can capture the nuances of human emotion that facial recognition systems might miss\. ## 3Research Objectives and Questions To address the challenges of emotional alignment in human\-agent interaction, this study evaluates a multimodal framework designed to capture user affect during proactive social dialogues\. Unlike previous research focused on passive recognition, we investigate the interaction dynamics when the agent takes the initiative, powered by generative AI\. Our primary goal is to analyze the discrepancy between AI\-recognized emotions and the users’ subjective perceptions in an unscripted, dynamic environment\. Specifically, this study seeks to answer the following two core research questions: - •RQ1 \(User Affective Experience\): What is the emotional experience of users when interacting with a proactive socially interactive agent? This question seeks to uncover the genuine internal impressions of the users during unscripted interactions\. By relying on user self\-reports and questionnaires, the objective is to establish an objective “ground truth” of how users genuinely feel when participating in AI\-driven social dialogues\. - •RQ2 \(Emotion Detection Efficacy\): How effectively can multimodal systems recognize users’ emotions during these interactions, and how do visual and linguistic modalities compare? This question aims to evaluate the practical reliability of automated emotion recognition in a live setting\. It focuses on observing the correspondence between the user’s physical facial expressions \(visual modality\), the semantic text\-based emotions detected by the generative AI \(linguistic modality\), and the user’s self\-reported ground truth\. The goal is to determine if one modality is more effective than the other and to identify the specific causes behind detection discrepancies\. By addressing these questions, we aim to provide a clearer understanding of the “expressive gap” in human\-agent interaction and offer design recommendations for more socially\-aware conversational systems\. ## 4Methodology This section describes the participants in the evaluation conducted, the systems designed, and the evaluation sessions carried out to gather the results\. ### 4\.1Participants A total of 20 participants were recruited for this study, consisting of 13 females and 7 males\. The participants’ ages ranged from 22 to 90 years \(M = 54\.90, SD = 21\.31\)\. This sample size was deemed appropriate for an exploratory Human\-Computer Interaction \(HCI\) study aimed at identifying interaction patterns and validating multimodal frameworks\. The broad age range was intentionally selected to evaluate the system’s performance\. To ensure data integrity, participants were screened based on the following inclusion criteria: - •Cognitive Function:No known history of cognitive impairments or neurological disorders\. - •Communication:Normal or mild speech capabilities, ensuring they could interact with the system’s voice interface\. - •Sensory Abilities:Normal or corrected\-to\-normal vision and hearing, allowing for an unhindered perception of the interface’s multimodal feedback\. All participants were briefed on the study’s objectives and provided informed consent prior to the sessions\. Participation was entirely voluntary, and no financial compensation was provided\. ### 4\.2Apparatus and Materials The study used the Sanbot Elf, a humanoid socially interactive agent, as the primary hardware platform \(see Figure[1](https://arxiv.org/html/2605.20200#S4.F1)\)\. A custom\-built Android application, displayed in Figure[2](https://arxiv.org/html/2605.20200#S4.F2), was developed to serve as the middleware between the user and the agent\. This application provided a minimalist visual interface to manage the session and facilitated real\-time data synchronization for recording emotional states\. Figure 1:Sanbot Elf running the Android app\.The interaction took place in a controlled laboratory setting to minimize external visual and auditory noise that could interfere with the facial recognition and speech processing modules\. ### 4\.3Multimodal Emotion Detection System The core of the system is a proactive AI engine designed for fluid dialogue\. As shown in the component diagram \(Figure[2](https://arxiv.org/html/2605.20200#S4.F2)\), the architecture integrates two primary emotion detection channels: 1. 1\.Computer Vision Module:This module leverages a standardized computer vision framework previously developed for SIAsgayamorey2024ai\-powered\. From the various models supported by this framework, SilNetramis2022novelwas selected due to its architectural simplicity and its specific optimization for FER tasks\. 2. 2\.Linguistic Analysis Module:The system captures the user’s speech, which is then processed via the OpenAI API \(GPT\-4/ChatGPT\)\. To ensure a nuanced understanding of the context, the system is programmed to perform a sentiment analysis every three interaction turns\. This allows the AI to evaluate the emotional trajectory of the discourse rather than isolated words\. Figure 2:Component diagram of the system\. ### 4\.4Conversational Design and Prompting Two distinct prompts were made: - •Prompt A\(Dialogue Generation\): Configures the SIA’s persona, role \(Friend, Expert, or Psychologist\), and the specific emotional goal of the conversation \(see Figure[3](https://arxiv.org/html/2605.20200#S4.F3)\)\. - •Prompt B\(Sentiment Analysis\): Instructs the generative engine to synthesize the conversation history and identify the user’s underlying affective state \(see Figure[4](https://arxiv.org/html/2605.20200#S4.F4)\)\. Figure 3:PROMPT1: dialogue generation\.Figure 4:PROMPT2: dialogue analysis\.The SIA was programmed to discuss four thematic areas \(Family, Nature, Animals, and Random\) across eight potential emotional targets \(e\.g\., Joy, Sadness, Surprise\)\. ### 4\.5Experimental Procedure The evaluation followed a within\-subjects design consisting of five distinct phases: 1. 1\.Setup and Personalization:The researcher configured the SIA’s role and the conversation topic\. 2. 2\.Briefing and Baseline:Participants were introduced to the SIA’s capabilities\. To establish a baseline, they were asked “How do you feel right now?” using the Valence\-Arousal Matrix \(see section 4\.6\)\. 3. 3\.Interaction Phase I:Participants engaged in an unscripted conversation with the SIA for approximately 5 minutes\. 4. 4\.Interaction Phase II:The process was repeated with a different topic or role to observe variations in engagement\. 5. 5\.Post\-Interaction Assessment:After each dialogue, participants answered two core questions: - •Q1:Which emotion do you perceive the SIA was trying to evoke or explore? - •Q2:Which emotion did you actually experience during the interaction? Additionally, users were also asked to complete a post\-dialogue questionnaire \(8 items\), and a final user experience questionnaire \(20 items\)\. ### 4\.6Measures: The 2D Emotion Matrix To bridge the gap between subjective feeling and objective data, we employed a two\-dimensional matrix \(Figure[5](https://arxiv.org/html/2605.20200#S4.F5)\) based on Russell’s Circumplex Model of Affectrussell1980circumplex\. This tool allowed users to self\-report their states based on two continuous axes: - •Valence:Positive vs\. Negative\. - •Arousal:High energy vs\. Low energy\. This model evaluates not only traditional emotions \(e\.g\., “Joy” or “Frustration”\) but also cognitive states\. As emphasized in HCI literaturepicard2000affective,dmello2012dynamics, “engagement” \(or attention\) is a fundamental cognitive\-affective state that must be recognized to ensure a successful interaction\. This visual aid was vital in providing a standardized non\-verbal reference for identifying complex affective and cognitive states before translating them into specific labels\. By driving the conversation, the generative AI actively shapes the interaction\. Therefore, capturing the user’s subjective state is essential to understand how people genuinely react to proactive AI agents, contrasting their internal feelings with their external behaviors\. Figure 5:Emotion matrix used by users to determine their emotions\. ## 5Results In this section, we present the qualitative and quantitative results obtained after conducting a thorough analysis of all the data collected during the user evaluation sessions\. ### 5\.1Subjective Measures and User Experience This section details the subjective data collected through structured surveys at different stages of the study\. The analysis encompasses the participants’ initial emotional baselines, the evolution of their affective state during the interactive sessions, and a final comprehensive assessment of their perception of the proactive agent\. #### 5\.1\.1User Self\-Reported Emotional Baselines Before initiating the dialogues with the system \(examples of these interactions and their AI linguistic analysis can be seen in Tables[1](https://arxiv.org/html/2605.20200#S5.T1)and[2](https://arxiv.org/html/2605.20200#S5.T2)\), the users’ baseline emotional states were predominantly stable and focused\. Specifically, the self\-reported emotions prior to the interaction were Calmness \(9 users\) and Attention \(7 users\), with only a small minority experiencing a Neutral state \(3 users\) or Surprise \(1 user\)\. This initial assessment establishes the subjective reality against which the emotional impact of the subsequent interactions is measured\. Table 1:Example of a proactive dialogue between the user and the designed systemTable 2:Example of the LLM\-based sentiment analysis and emotional reasoningIn the first round of dialogue, the conversational agent was configured to evoke a diverse array of emotions\. The system primarily targeted Surprise \(7 cases\) and Disgust \(5\), followed by Anger \(3\), Sadness \(2\), Happiness \(2\), and Fear \(1\)\. The conversation topics varied, with Nature \(8\) and Animals \(7\) being the most predominant, while the agent mainly adopted the role of a Friend \(13\) or a Psychologist \(6\)\. Consequently, the self\-reported final emotions of the users after this first interaction showed significant diversity\. Attention emerged as the most prevalent final state \(5 users\), followed by Sadness \(3\) and Calmness \(3\)\. This high diversity indicates that when the SIA attempts to connect with complex or negative initial emotions \(such as disgust or anger\), it generates a wide and varied spectrum of emotional responses in the users\. For the second round of dialogue, the strategy was highly focused: the system aimed to connect exclusively with Happiness in all 20 cases, using “Random” conversation topics across the board\. The agent’s roles remained consistent with the first round, predominantly acting as a Friend \(14\)\. The evolution of the users’ reported emotions in this round demonstrated much greater consistency and alignment\. The Happiness\-focused strategy successfully maintained or shifted most users toward positive states\. After the interaction, the dominant user emotion was Happiness \(12 users\), reflecting the system’s success in this specific goal, while the rest transitioned toward relaxed or interested states such as Calmness \(4\) and Attention \(3\), with only one isolated case of Neutrality\. These self\-reported outcomes constitute the objective “ground truth” of the users’ internal affective states during the experiment\. They establish that users generally felt positive, calm, or attentive throughout the sessions, setting the foundation against which the automated facial and linguistic detection systems will be evaluated in Section[5\.2](https://arxiv.org/html/2605.20200#S5.SS2)\. #### 5\.1\.2Post\-Dialogue Experience and Affective State To evaluate the immediate response to the proactive sessions, participants completed an 8\-item questionnaire after each dialogue\. Table[3](https://arxiv.org/html/2605.20200#S5.T3)presents the comparative results between the two sessions\. Table 3:User experience and affective state results across Dialogue 1 and Dialogue 2 \(N=20N=20per session\)\. Items are rated on a 1–5 Likert or Semantic Differential scale\. \(R\) indicates reverse\-scored items\.The results suggest a stable or increasing trend in the perception of the agent\. Regarding the system’s ability to understand emotions \(PDQ2\), the mean increased from the first session \(μ=3\.6,σ=0\.9\\mu=3\.6,\\sigma=0\.9\) to the second \(μ=4\.0,σ=1\.1\\mu=4\.0,\\sigma=1\.1\)\. A decrease in the perceived irrelevance of the conversation was also observed \(PDQ3\), moving from a more neutral stance \(μ=2\.7,σ=1\.2\\mu=2\.7,\\sigma=1\.2\) to a lower score in the second dialogue \(μ=2\.0,σ=0\.9\\mu=2\.0,\\sigma=0\.9\)\. The overall empathy of the assistant \(PDQ4\) remained consistently high in both sessions \(μD1=4\.0,σ=0\.9\\mu\_\{D1\}=4\.0,\\sigma=0\.9;μD2=3\.9,σ=1\.1\\mu\_\{D2\}=3\.9,\\sigma=1\.1\)\. In terms of the users’ emotional state, satisfaction levels \(PDQ5\) were reported above the mid\-point in both interactions \(μD1=4\.2,σ=0\.7\\mu\_\{D1\}=4\.2,\\sigma=0\.7;μD2=4\.1,σ=0\.9\\mu\_\{D2\}=4\.1,\\sigma=0\.9\)\. Conversely, responses regarding the level of activation or energy \(PDQ6\) remained in the lower\-middle range of the scale \(μD1=2\.6,σ=1\.1\\mu\_\{D1\}=2\.6,\\sigma=1\.1;μD2=2\.5,σ=1\.2\\mu\_\{D2\}=2\.5,\\sigma=1\.2\)\. Motivation \(PDQ7\) showed a slight upward trend \(μD1=3\.8,σ=0\.7\\mu\_\{D1\}=3\.8,\\sigma=0\.7;μD2=4\.1,σ=0\.9\\mu\_\{D2\}=4\.1,\\sigma=0\.9\), while interest in the conversation topic \(PDQ8\) reached its highest value in the second dialogue \(μ=4\.2,σ=1\.0\\mu=4\.2,\\sigma=1\.0\)\. #### 5\.1\.3User Perception and Subjective Evaluation Upon completion of the experimental sessions, participants were asked to complete a comprehensive questionnaire designed to evaluate their overall perception of the interaction, the agent’s communicative effectiveness, and the emotional resonance of the experience\. The results of this assessment, organized into three functional dimensions, are summarized in Table[4](https://arxiv.org/html/2605.20200#S5.T4)\. Table 4:User experience and perception questionnaire results \(Likert scale 1\-5,N=20N=20\)\. \(R\) indicates reverse\-scored items \(lower is better\)\.Dimension / ItemMeanSDCommunication & UnderstandingFQ1:I believe I can communicate with Alexa through language4\.50\.5FQ2:I believe I can clearly express my thoughts to Alexa4\.20\.7FQ3:I believe that Alexa fully understands what I mean3\.91\.1FQ4:I don’t think Alexa understands my expressions \(R\)2\.81\.4FQ5:I believe that Alexa is not capable of understanding complex stories \(R\)3\.01\.4FQ6:I believe that Alexa takes too long to respond \(R\)2\.11\.0Emotional Engagement & EmpathyFQ7:I maintain eye contact with Alexa while interacting with it4\.50\.6FQ8:I like it when Alexa encourages me4\.10\.8FQ9:Alexa comforts me when I’m upset3\.71\.2FQ10:I believe that Alexa influenced my feelings3\.61\.1FQ11:I believe that Alexa shows emotions when interacting with me3\.51\.2FQ12:I believe I could open my heart to Alexa3\.11\.4FQ13:I believe that Alexa reacts to my words but does not perceive how I feel \(R\)3\.21\.4System Perception & TrustFQ14:I enjoy talking to the Alexa assistant4\.01\.1FQ15:I trust the advice and recommendations provided by Alexa3\.81\.2FQ16:I believe that Alexa could be a pleasant communication companion3\.71\.3FQ17:I think conversations with Alexa can be rigid \(R\)2\.71\.1FQ18:I believe that Alexa does not have emotions \(R\)2\.71\.3FQ19:No matter what I tell about myself, Alexa acts the same \(R\)2\.11\.0FQ20:I easily get distracted from interacting with Alexa \(R\)2\.71\.2Regarding Communication & Understanding, users reported high confidence in their ability to interact through language \(FQ1,μ=4\.5,σ=0\.5\\mu=4\.5,\\sigma=0\.5\), which showed the lowest deviation in the study\. While the clarity of expression was rated highly \(FQ2,μ=4\.2,σ=0\.7\\mu=4\.2,\\sigma=0\.7\), items concerning the depth of understanding, such as the agent’s ability to interpret complex stories \(FQ5,μ=3\.0,σ=1\.4\\mu=3\.0,\\sigma=1\.4\) or non\-verbal expressions \(FQ4,μ=2\.8,σ=1\.4\\mu=2\.8,\\sigma=1\.4\), showed greater polarization\. Notably, the perceived response time of the agent was rated favorably, with users generally disagreeing that Alexa took too long to respond \(FQ6,μ=2\.1,σ=1\.0\\mu=2\.1,\\sigma=1\.0\)\. In the Emotional Engagement & Empathy dimension, participants reported a high rate of eye contact \(FQ7,μ=4\.5,σ=0\.6\\mu=4\.5,\\sigma=0\.6\) and perceived the agent as encouraging \(FQ8,μ=4\.1,σ=0\.8\\mu=4\.1,\\sigma=0\.8\)\. The agent’s capacity to provide comfort yielded a mean above the mid\-point \(FQ9,μ=3\.7,σ=1\.2\\mu=3\.7,\\sigma=1\.2\), although items related to emotional intimacy, such as “opening one’s heart” \(FQ12,μ=3\.1,σ=1\.4\\mu=3\.1,\\sigma=1\.4\), reflected more diverse individual stances\. Finally, System Perception & Trust indicators show that users generally enjoyed the interaction \(FQ14,μ=4\.0,σ=1\.1\\mu=4\.0,\\sigma=1\.1\) and considered Alexa a pleasant companion \(FQ16,μ=3\.7,σ=1\.3\\mu=3\.7,\\sigma=1\.3\)\. The perceived rigidity of the conversation \(FQ17,μ=2\.7,σ=1\.1\\mu=2\.7,\\sigma=1\.1\) and the tendency for the agent to act the same regardless of user input \(FQ19,μ=2\.1,σ=1\.0\\mu=2\.1,\\sigma=1\.0\) were situated in the lower\-middle range of the scale\. Furthermore, the level of distraction during the interaction remained relatively low \(FQ20,μ=2\.7,σ=1\.2\\mu=2\.7,\\sigma=1\.2\)\. ### 5\.2Performance of Multimodal Emotion Detection Systems This section presents the objective performance data of the two emotion detection modules \(facial and linguistic\) by comparing their real\-time classifications against the users’ self\-reported emotional baselines\. #### 5\.2\.1Facial Recognition The automated facial recognition module demonstrated a strong tendency toward detecting negative emotional states, revealing a significant gap between system classification and user self\-reports\. As illustrated in the middle chart of Figure[6](https://arxiv.org/html/2605.20200#S5.F6), the frequency of emotions detected by the facial module is heavily skewed toward expressions such as “Disgust” \(14\) and “Anger” \(11\), which together account for 62\.5% of the recorded expressions\. In contrast, users predominantly reported positive or neutral states, such as “Happiness” \(14\), “Calmness” \(9\), and “Attention” \(8\)\. Figure 6:Comparison of emotion frequency distribution in self\-reports \(left\), expressions automatically detected by the facial recognition module \(middle\), and emotions detected from the linguistic modality \(right\)\.The correspondence between these states is further detailed in the correspondence matrices shown in Figure[7\(a\)](https://arxiv.org/html/2605.20200#S5.F7.sf1)\. The system’s performance reveals a “poker face” effect: when users reported a state of “Attention”, the facial module incorrectly classified it as “Disgust” in 62\.5% of the instances\. A similar trend is observed in self\-reported “Calmness” states \(9 occurrences\), which the system predominantly misclassified as “Anger”, “Disgust”, and “Sadness”\. “Neutrality” was also poorly recognized by the visual module: from a total of two self\-reported occurrences, one was classified as “Anger”, and the other as “Disgust”\. These discrepancies suggest that users maintain a serious or concentrated facial expression while interacting with the AI, despite experiencing low\-energy or positive internal emotions\. \(\(a\)\) \(\(b\)\) Figure 7:Correspondence matrices comparing predictions from the visual \(a\) and linguistic \(b\) modalities with user self\-reports\.Overall, the accuracy of this modality when taking the self\-reports as ground truth and excluding those occurrences of emotions that are not among those that the network can detect, was of 23\.8%, just a little higher than random choice \(16\.7%\)\. #### 5\.2\.2Linguistic Analysis The linguistic analysis module, powered by a generative AI engine, yielded significantly different outcomes compared to the facial recognition system\. As shown in the right chart of Figure[6](https://arxiv.org/html/2605.20200#S5.F6), the frequency of “Happiness” identified through semantic analysis closely mirrors that of users’ self\-reports \(left chart of the same figure\), although the frequency of other emotions differs greatly\. Emotions like “Calmness” and “Attention” were frequently self\-reported \(9 and 8 occurrences\), while the linguistic module only recognized each one of them once\. The effectiveness of this modality is further detailed in the correspondence matrices \(Figure[7\(b\)](https://arxiv.org/html/2605.20200#S5.F7.sf2)\)\. This module proved highly successful at identifying “Happiness”, “Sadness”, and “Disgust” emotions, correctly detecting them in 12/14, 2/3 and 1/1 of the cases\. However, the data also reveals poor results regarding the detection of “Attention”, “Calmness”, and “Neutrality”, misclassified in 7/8, 7/9, and 2/2 of occurrences, respectively\. This modality achieved an accuracy of 45%, which is much higher than that of the visual modality, while considering many more emotions\. ### 5\.3Conversational Dynamics and AI\-Generated Patterns This section details the conversational interactions between the users and the generative AI, focusing on how specific themes elicited distinct emotional responses and the structural patterns observed in the dialogues\. #### 5\.3\.1Thematic Influence on Emotion Elicitation The emotions generated in the users by the SIA were heavily dependent on the conversational themes and the specific experiences discussed\. The objective mapping of themes to emotional outcomes revealed the following dynamics: - •Positive Emotions:Joy was consistently generated when discussing topics like Animals \(e\.g\., anecdotes about pets\) or Parties and events \(memories of celebrations\)\. Unexpected encounters or life changes typically generated Surprise, while Personal memories often induced Tranquility or Joy, depending on the nature of the anecdote\. - •Negative Emotions:Sensitive or complex topics naturally triggered negative states\. The theme of Nature was polarizing: while it could be relaxing, discussing environmental degradation elicited Sadness, and discussing the lack of societal action provoked Anger\. The Animals theme also generated Disgust when the conversation shifted to insects or decomposing carcasses\. Finally, Unexpected encounters involving dangerous situations \(such as wildlife encounters\) were effective in eliciting Fear\. Despite the system’s targeted strategies, the data revealed several confusions between the emotion the SIA intended to work on and what the user perceived\. An analysis of 40 interaction records identified several misinterpretations: the AI’s attempt to evoke Happiness was confused by users with Fear \(4 occasions\), Disgust \(4 occasions\), and Surprise \(2 occasions\)\. Additionally, intended Fear was confused with Surprise \(2 occasions\), and intended Disgust was confused with Calmness \(2 occasions\)\. #### 5\.3\.2Dialogue Structure and Behavioral Patterns An analysis of the conversation transcripts, detailed in Tables[5](https://arxiv.org/html/2605.20200#S5.T5)and[6](https://arxiv.org/html/2605.20200#S5.T6), highlights the specific linguistic patterns the generative AI utilized to evoke targeted emotions\. Table 5:Pattern analysis results \(joy, sadness and surprise emotions\)Table 6:Pattern analysis results \(anger, fear and disgust emotions\)The SIA employed highly structured language: - •To elicit Joy, the system frequently used interjections \(e\.g\., “Haha\!”, “How great\!”\) and focused on humorous anecdotes to maintain a cheerful tone\. - •For Sadness, the AI used empathetic language designed to validate the user’s feelings \(e\.g\., “It’s normal to feel nostalgic”\)\. - •When targeting Anger, the system utilized direct language addressing the external cause of frustration \(e\.g\., “That’s really disappointing to hear”\)\. - •To evoke Fear, the AI used escalating language to resonate with the danger of the situation \(e\.g\., “It must have been terrifying”\), whereas for Disgust, it introduced examples evoking unpleasant sensations\. While the dialogues generally maintained a coherent, logical flow that mimicked human conversation by adapting to user interests, objective observations of user behavior revealed limitations in the interaction dynamics\. In emotionally complex scenarios, particularly when the user exhibited frustration, the system sometimes failed to dynamically redirect the conversation toward a constructive tone\. In these instances of limited adaptation, or when the AI’s proactivity \(such as forced jokes\) felt uncalibrated, users displayed behavioral signs of disengagement\. This was evidenced by users giving very short answers \(e\.g\., “yes”, “no”\), expressing that they did not know what to say, or abruptly changing the subject, reflecting that they perceived the interaction as artificial\. ## 6Discussion Next, we discuss the findings and use the obtained results to answer the research questions\. We also analyze the effects of agent’s proactivity, and offer design recommendations grounded on the findings\. ### 6\.1RQ1 \(User Affective Experience\) The analysis of the self\-reported data reveals that users are highly receptive to proactive interactions with generative AI, maintaining predominantly positive or cognitive states throughout the dialogues\. Despite the artificial nature of the conversational agent, users reported feeling states such as calmness, focused attention, or happiness most frequently\. The generative SIA successfully elicited a wide range of internal emotions depending on the conversational theme; for instance, discussing personal memories or animals fostered joy, while complex topics like environmental degradation naturally elicited reflective sadness or anger\. Ultimately, the subjective impression of the users was positive, indicating that they felt comfortable and enjoyed the experience, setting a baseline of low\-arousal or positive affect\. The high scores in Emotional Engagement & Empathy \(Q8,μ=4\.1\\mu=4\.1; Q9,μ=3\.7\\mu=3\.7\) indicate that the agent’s proactivity was not merely functional but succeeded in projecting a social presence capable of comforting and encouraging the user\. Furthermore, the evolution of the system perception between sessions suggests that users underwent a ’calibration’ phase\. This indicates that while proactive AI may initially be perceived with uncertainty, continued interaction allows the user to form a stable and positive mental model of the agent’s empathetic capabilities\. ### 6\.2RQ2 \(Emotion Detection Efficacy\) When evaluating how effectively the multimodal systems detected these internal states, the results revealed a significant discrepancy between the visual and linguistic modalities\. Overall, the generative linguistic analysis demonstrated a higher alignment with the users’ actual emotional states, achieving a superior accuracy \(45%\) compared to the facial recognition module, which exhibited a very low alignment \(around 23\.8% success rate\)\. Comparing both modalities highlights clearly differentiated strengths and limitations\. The generative linguistic models have the significant strength of better capturing the user’s internal state through the semantics of their verbal expressions\. Nonetheless, their limitations emerge when dealing with the ambiguity of human language\. If a user employs words like “sadness,” “scared,” or “disgust” while telling a past anecdote, the AI tends to classify the user’s current state as negative, even though the user finishes the conversation feeling at ease\. Context\-dependent affective nuances pose a major challenge: ambiguous phrases like “Yes, I was scared, but since I saw that he was okay, I found it a little funny” or expressions of nostalgia \(“It’s normal to feel nostalgic”\) are frequently misinterpreted by the system, as it struggles to isolate the user’s true current feelings from the semantics of the narrated story\. This reflects the “lexical fallacy”fiske2020lexical, where the system fails to distinguish between the topic of conversation and the user’s actual affective state\. On the other hand, the visual recognition system’s major limitation was its massive skew toward negative emotions\. It detected more than 50% negative emotions \(predominantly anger and disgust\) despite users reporting positive or calm states\. This severe detection failure is explained by the “poker face” phenomenonvanderhasselt2012put\. Interacting with artificial conversational agents alters users’ natural expressiveness; they tend to adopt serious facial features typical of deep concentration \(such as furrowed brows or motionless expressions\) because they perceive the conversation as an evaluative task or a challenge, or simply because they do not find it natural to speak with a machine\. Consequently, a strong dissonance is generated: the user’s physical inexpressiveness or severe concentration is interpreted by the facial recognition system as rejection or negative states \(like anger or disgust\)\. Internally, however, users assert that despite their serious expressions, they felt good, enjoyed the experience, and maintained a comfortable and positive affective state during the interaction\. This is why, as noted byfeldmanbarrett2019emotional, inferring internal affect solely from facial movements is inherently unreliable in social contexts\. The ’poker face’ phenomenon identified in the visual module is explicitly corroborated by the subjective self\-reports in Q26 \(Feeling: Relaxed vs\. Energetic\)\. Despite reporting high satisfaction and motivation \(Q25,μ=4\.1\\mu=4\.1; Q27,μ=4\.1\\mu=4\.1\), users consistently reported low arousal levels \(μ≈2\.5\\mu\\approx 2\.5\)\. This confirms that the ’neutral’ or ’serious’ expressions detected as negative by the AI actually corresponded to a state of calm satisfaction\. The discrepancy, therefore, is not a lack of user engagement, as evidenced by the high rate of eye contact \(Q7,μ=4\.5\\mu=4\.5\), but a physiological state of low activation that facial recognition algorithms misinterpret as emotional flatness or hostility\. ### 6\.3User Engagement and Proactivity The study observed that agent proactivity is a double\-edged sword\. While it facilitates dialogue flow, it can also lead to perceived artificiality\. As observed in some cases, uncalibrated proactivity can result in disengagement, characterized by brief or evasive answers\. This aligns with theories on social reactance in HCI; if the agent’s emotional tone does not match the user’s “poker face,” the interaction feels forcedbrehm1966theory\. The successful delivery of a personalized experience is further evidenced by the low scores on Q19 \(μ=2\.1,σ=1\.0\\mu=2\.1,\\sigma=1\.0\), which indicates that users perceived the agent’s responses as dynamic and tailored to their input, rather than repetitive\. However, the polarization observed in items related to deep understanding \(Q4, Q5,σ=1\.4\\sigma=1\.4\) suggests that the ’double\-edged sword’ of proactivity is most sharp when the agent attempts to interpret complex human nuances, where user trust remains divided\. ### 6\.4Design Recommendations for Emotional SIAs Based on these insights, we propose five key recommendations for the design of future proactive SIAs: - •Prioritize the linguistic modality over the visual one\. Designers should not rely solely on computer vision\. Systems should give higher weight to linguistic sentiment to counteract the “facial neutrality” common in human\-screen interactions\. - •Calibrate for the “poker face\.” Detection systems must be tuned to recognize that neutral expressions often signify high engagement or concentration rather than negative affect\. - •Implement “affective probing\.” When a prolonged discrepancy between a serious face and positive discourse is detected, the agent should use direct verbal queries \(e\.g\., “How are you finding this conversation?”\) to recalibrate its internal model\. - •Dynamic engagement recovery\. If the system detects signs of disengagement \(e\.g\., short answers\), it should pivot from structured information\-sharing to more empathetic or humorous prompts to break conversational rigidity\. - •Feedback and transparency\. Interfaces should subtly reflect the agent’s perception of the user’s emotion \(e\.g\., through slight changes in tone or avatar micro\-expressions\), allowing users to naturally correct misinterpretations\. ## 7Limitations and Future Work While this study provides valuable insights into emotion recognition during proactive interactions with SIAs, several limitations must be acknowledged\. First, the evaluation was conducted with a relatively small sample size \(N=20N=20\) and took place within a controlled laboratory setting\. This artificial environment likely exacerbated the observed “poker face” effect, as users may have perceived the interaction as an evaluative task rather than a spontaneous conversation\. Second, the current multimodal framework is limited to facial and text\-based linguistic analysis\. The linguistic module, while generally more accurate, demonstrated a vulnerability to the complexity and ambiguity of natural language; it occasionally misinterpreted negative vocabulary used in past anecdotes as the user’s current emotional state, failing to fully separate the conversational theme from real\-time affect\. Furthermore, the generative AI’s performance and proactivity were strictly bounded by the specific prompt structures utilized during the experiment\. To address these limitations, our future research will focus on expanding and refining the emotional detection architecture\. Specifically, we propose the following lines of future work: - •Expanded Multimodality:We plan to integrate additional expressive channels, such as voice tone analysis and body gesture detection, to complement the existing facial and linguistic approaches, thereby creating a more robust emotional profile of the user\. - •Dynamic Feedback Mechanisms:To overcome the linguistic module’s literal interpretations, we will incorporate advanced feedback\-driven learning systems\. These systems will dynamically adjust emotional interpretations based on direct user responses and will utilize adapted prompts to ensure continuous improvement and naturalness in the dialogue\. - •Longitudinal and Diverse Studies:Future studies will involve larger and more demographically diverse participant samples to conduct deeper analyses of how user emotional states and engagement evolve over extended, repeated interactions with SIAs\. - •In\-the\-Wild Evaluations:Finally, we intend to transition our evaluations from controlled laboratory environments to naturalistic settings, such as users’ homes\. Analyzing these “in\-the\-wild” scenarios will help determine if users feel more relaxed and display more natural expressive behaviors when interacting with the system in their everyday environments\. ## 8Conclusions This study evaluated a multimodal emotion recognition system integrated into a proactive socially interactive agent powered by generative AI\. By combining facial recognition with real\-time linguistic analysis, we assessed the alignment between automated detection and the self\-reported emotional states of users during unscripted, dynamic dialogues\. Regarding the user affective experience \(RQ1\), the psychometric and self\-reported data indicate that participants were highly receptive to proactive interactions with the generative agent\. Despite the artificial nature of the system, the SIA successfully elicited a wide range of internal emotions, from joy during the discussion of personal memories to reflective sadness when addressing complex social themes\. This emotional resonance suggests that the agent projected a social presence capable of comforting and encouraging the user\. Furthermore, the research demonstrated that agent proactivity serves as a “double\-edged sword” in human\-AI interaction\. On one hand, the generative engine successfully delivered a dynamic and personalized experience, avoiding repetitive patterns and effectively facilitating dialogue flow\. On the other hand, uncalibrated proactivity, where the agent’s emotional tone failed to align with the user’s internal state, occasionally led to perceived artificiality or conversational rigidity\. In terms of emotion detection efficacy \(RQ2\), the study revealed a significant discrepancy between the visual and linguistic modalities\. The linguistic module, leveraging the semantic context of the generative AI, proved more reliable in capturing the users’ internal states, especially during the narration of personal anecdotes\. In contrast, facial recognition faced severe limitations due to the “poker face” effect\. The serious expressions typical of deep concentration or the perceived evaluative nature of speaking with an AI were systematically misinterpreted by the visual system as negative affect \(e\.g\., anger or disgust\)\. This underscores that in proactive HCI contexts, neutral facial features often mask high levels of engagement rather than reflecting negative emotions or indifference\. This research suggests that the design of socially aware agents must move beyond a heavy reliance on visual cues, which can be misleading in human\-screen interactions\. Instead, systems should prioritize contextual linguistic analysis and account for the low\-arousal states typical of satisfied but concentrated users\. By calibrating proactivity to these subtle emotional landscapes, SIAs can achieve a more natural and empathetic alignment, moving from simple command\-following tools to genuine conversational companions\. \\bmsubsection \*Acknowledgments Partially funded by the Spanish Ministry of Science and Innovation through contract PID2022\-136779OB\-C31\. T60\_23R Research Group in Advanced Interfaces \(AffectiveLab\), Government of Aragón\. Research grant program 2024\. Antonio Gargallo University Foundation\. “Companion SIAs for seniors: Heavy lifting and personal transportation”\. Grant PID2022\-136779OB\-C32 \(PLEISAR\) funded by MICIU/ AEI /10\.13039/501100011033/ and FEDER, EU\. \\bmsubsection \*Conflicts of Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\. \\bmsubsection \*Data availability The authors do not have permission to share data\. ## References
Similar Articles
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
This research paper explores emotion-sensitive neurons in multimodal foundation models, revealing shared affective mechanisms between speech and facial emotion recognition through causal interventions and cross-modal analysis.
Leveraging Self-Paced Curriculum Learning for Enhanced Modality Balance in Multimodal Conversational Emotion Recognition
This paper proposes a plug-and-play module using self-paced curriculum learning to enhance modality balance in multimodal conversational emotion recognition, achieving consistent F1-score improvements on IEMOCAP and MELD datasets.
Rationale-Guided Learning for Multimodal Emotion Recognition
Introduces Rationale-Guided Learning (RGL), a framework that reframes multimodal emotion recognition in conversation as a cognitively-inspired reasoning task using dual-process theory and MLLM-generated rationales, achieving state-of-the-art results on IEMOCAP and MELD.
Your Multimodal Speech Model Says I Have a Face for Radio
This paper presents the first bias evaluation of multimodal speech recognition models, finding significant accuracy differences across gender and ethnicity when pairing faces with audio, with implications for fairness in AI systems.
C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning
The paper proposes C²MOE, a Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion recognition in conversations, using information-theoretic decomposition to improve robustness when modalities are missing.