MoCA: Implicit Social Context Analysis

arXiv cs.CL Papers

Summary

Introduces MoCA (Implicit Social Context Analysis), a new task and benchmark for modeling implicit social scenarios across affection, intent, and stance, along with a Conflict-Driven Abductive Reasoning (CoDAR) framework. Experiments show state-of-the-art multimodal LLMs struggle on this task, while CoDAR improves performance but still lags behind human reasoning.

arXiv:2608.05825v1 Announce Type: new Abstract: Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through indirect, socially and culturally grounded signals rather than explicit statements. Such implicit social contexts are pervasive in real-world interactions, yet there remains a lack of a formal and systematic framework for studying them. In this paper, we introduce Implicit Social Context Analysis (MoCA), a novel task that systematically models implicit social scenarios along three key dimensions: affection, intent, and stance. We construct a high-quality benchmark containing 3,108 multimodal instances collected from real-world sources, with fine-grained cognitive annotations revealing who expresses what toward whom, as well as how and why it is conveyed. Using the MoCA dataset, we show that state-of-the-art multimodal large language models struggle significantly with this task because of their reliance on explicit cues and limited ability to reason over latent social contexts. To address this challenge, we propose Conflict-Driven Abductive Reasoning (CoDAR), a novel framework that models the discrepancy between observed expressions and expected truthful behavior as cognitive conflict, thereby enabling the inference of hidden mental states. Extensive experiments demonstrate that CoDAR substantially improves model performance. Nevertheless, a large gap from human reasoning remains, highlighting the fundamental difficulty of implicit social understanding.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:52 AM

# MoCA: Implicit Social Context Analysis
Source: [https://arxiv.org/html/2608.05825](https://arxiv.org/html/2608.05825)
\\equalcont

Equal contributors\.

\\equalcont

Equal contributors\. \[2\]\\fnmHao\\surFei

1\]\\orgnameNational University of Singapore,\\orgaddress\\citySingapore,\\countrySingapore

2\]\\orgnameUniversity of Oxford,\\orgaddress\\cityOxford,\\countryUnited Kingdom

3\]\\orgnameWuhan University,\\orgaddress\\cityWuhan,\\countryChina

###### Abstract

Human social communication \(e\.g\., affection, intent\) is often conveyed in highly implicit ways, where underlying meanings are expressed through indirect, socially and culturally grounded signals rather than explicit statements\. Such implicit social contexts are pervasive in real\-world interactions, yet unfortunately there is still a lack of a formal and systematic framework to study them\. In this paper, we introduceImplicit Social Context Analysis\(MoCA\), a novel task that systematically models implicit social scenarios along three key dimensions: affection, intent, and stance\. We construct a high\-quality benchmark of 3,108 multimodal instances collected from real\-world sources, with fine\-grained cognitive annotations to reveal who expresses what toward whom, and how and why it is conveyed\. With MoCA data, we show that state\-of\-the\-art multimodal large language models struggle significantly on MoCA, due to their reliance on explicit cues and limited capability in reasoning over latent social context\. To address this challenge, we propose a novelConflict\-Driven Abductive Reasoning\(CoDAR\) framework, which models the discrepancy between observed expressions and expected truthful behavior as cognitive conflict, enabling the inference of hidden mental states\. Extensive experiments demonstrate that CoDAR substantially improves model performance, while a large gap to human reasoning remains, highlighting the fundamental difficulty of implicit social understanding\.

## 1Introduction

Human social interaction has evolved into a highly complex system, where communication and expression are deeply shaped by cultural background, social roles, and contextual dependencies\[grice1975logic,goffman2023presentation\]\. Constrained by power structures, politeness norms, and self\-protection considerations, people in everyday interactions often avoid expressing their true emotions, thoughts, or stances in a straightforward manner\[brown1987politeness\]\. Instead, they tend to strategically rely on more indirect and implicit forms of expression, cognitively choosing ways to convey underlying meanings without stating them explicitly\[searle1975indirect\]\. In both offline interactions and online social platforms, such expressions are frequently realized through subtle pragmatic cues, facial micro\-expressions, vocal tone variations, fine\-grained body movements, and creative multimodal artifacts such as memes or short videos, often involving irony, sarcasm, metaphor, or indirect insinuation to express complex emotions and intentions\[van2020hierarchy,van2016social,cheshin2020impact\]\. Despite its prevalence and importance, this phenomenon remains largely underexplored, and there has been scarce systematic effort to study it in a unified manner\. Existing work either focuses on implicit affect and intent analysis in purely textual settings\[fei2023reasoning,villarroel2017unveiling,alswaidan2020survey\], whereas real\-world implicit social expression is inherently multimodal; or considers multimodal tasks such as sarcasm detection\[saha2025mustreason\], which are typically limited to shallow label classification without requiring deeper cognitive\-level interpretation of underlying implications; or relies on highly abstract theory\-of\-mind frameworks\[kanske2018social\]that are often detached from concrete implicit social scenarios\. To bridge this gap, this paper is dedicated to a systematic investigation of this form of implicit social computation\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x1.png)Figure 1:Examples of real\-world implicit social expressions, where the underlying real affection, intent, or stance must be analyzed implicitly from the given multimodal cues\.In this work, we formalize these scenarios asImplicit Social Context Analysis\(MoCA\)\. Recent benchmarks motivate a structured treatment of latent social meaning rather than a single generic label\. Multimodal affective benchmarks combine linguistic, visual, and acoustic evidence for emotion reasoning\[cheng2024emotionllama\], while visually grounded dialogue benchmarks examine affective reasoning in conversational context\[haydarov2024affectivevisualdialog\]\. Multimodal Theory\-of\-Mind and intent benchmarks infer goals, beliefs, plans, and communicative purposes from situated behavior\[jin2024mmtom,zhang2024mintrec2\]\. Multimodal stance detection further treats position as target\-dependent and jointly grounded in textual and visual evidence\[liang2024multimodalstance\]\.

MoCA therefore distinguishes three related but non\-equivalent inference targets: an underlying affective state, an intended outcome, and a target\-relative position:

- •Implicit Affectioninfers a concealed or indirectly communicated affective state from ambiguous or conflicting verbal content, prosody, facial behavior, and visual context\. It targets the underlying emotion rather than the overt affective display\.
- •Implicit Intentinfers the outcome sought through an observed expression or action\. Although the behavior may be explicit, its motivating goal must be inferred from the addressee, interaction history, and plausible social consequences\.
- •Implicit Stanceinfers the subject’s evaluative or relational position toward a person, group, topic, or proposition\. It covers support, opposition, distancing, alignment, and power\-sensitive positioning grounded in multimodal context\. Because stance is relational, its target is inferred rather than treated as metadata\.

These dimensions may co\-occur but remain distinct: affection represents an affective state, intent an intended outcome, and stance a target\-relative relation\. Each instance is assigned to a primary scenario that constrains its label and mechanism spaces; Figure[1](https://arxiv.org/html/2608.05825#S1.F1)provides representative examples\.

To facilitate research along this direction, we construct a large\-scale benchmark dataset for MoCA\. We first collect data from a wide range of real\-world scenarios, including memes, online discussions, public debates, and situational comedies, all of which contain rich multimodal signals such as text, images, and videos, from which we categorize instances into affection, intent, and stance scenarios\. For each instance, we further annotate a comprehensive set of labels from a human cognitive perspective, including who is expressing \(subject\), toward whom the expression is directed \(target\), what the underlying meaning is \(label\), and how&why the meaning is conveyed \(mechanism\)\. We design a human annotation pipeline with rigorous verification procedures to ensure high\-quality control, resulting in a MoCA dataset of 3,108 instances\.

The MoCA task is inherently non\-trivial\. Although state\-of\-the\-art \(SoTA\) multimodal large language models \(MLLMs\)\[gpt5systemcard2025,gemini2025\]have shown strong performance across many cross\-modal tasks, they consistently struggle when applied to MoCA scenarios\. The key reason is that many implications in implicit social contexts are not explicitly present in textual or visual cues, whereas MLLMs are primarily designed to interpret and reason over observable signals\. We argue that understanding such implicit meanings requires access to latent social context, such as interpersonal relationships, cultural norms, historical interactions, and power structures, as well as human\-like cognitive reasoning built upon these factors to infer the subject’s hidden mental states\. To address this challenge, we propose a novelConflict\-Driven Abductive Reasoning\(CoDAR\) framework\. Grounded in cognitive\-driven insights, we observe that in implicit social scenarios, individuals often adopt expressions that deviate from what would be expected if their true internal states were directly revealed, due to underlying goals, constraints, and perceived risks\. This suggests that the core of MoCA reasoning lies in identifying and interpreting thecognitive “conflict”embedded in the context, and CoDAR explicitly models the discrepancy between the subject’s observed expression and their expected expression under truthful conditions, leveraging abductive reasoning to recover the underlying affection, intent, and stance\.

We conduct extensive experiments on the MoCA dataset with both general\-purpose multimodal models and domain\-adapted variants, covering both open\-source and proprietary SoTA MLLMs, and observe a substantial performance gap between these models and human\-level reasoning in implicit social scenarios, highlighting the inherent challenge of the task\. Notably, equipping these models with our proposed CoDAR framework leads to significant performance improvements\. We further provide in\-depth analyses of different models and approaches on MoCA, examining their failure modes and exploring potential directions for improvement\. Overall, our main contributions can be summarized as follows\.

- •We introduce a novel task of Implicit Social Context Analysis, providing the first systematic study of how AI can perform human\-level cognitive reasoning in implicit social scenarios\.
- •We construct a high\-quality, real\-world multimodal dataset of 3,108 MoCA instances spanning affection, intent, and stance, and extensively evaluate across SoTA models\.
- •We propose a cognition\-aware conflict\-driven reasoning framework, CoDAR, as a benchmark method for implicit social inference, which effectively enhances the performance of current multimodal large models\.

## 2Related Work

### 2\.1Implicit Social Computing and Social Cognition Analysis

Social computing seeks to endow machines with socially aware capabilities through AI algorithms, enabling computational systems to perceive, model, and respond to human social signals\. It has emerged as an important research direction with a wide range of real\-world applications\[wang2021survey,wu2021modeling\]\. Within this broad agenda, analyzing implicit social scenarios is particularly valuable because it more closely reflects how social interactions unfold in practice\[zhou2021implicit\]\. Unlike explicit settings, where the relevant social state is directly stated or readily observable, implicit analysis requires recovering unstated affect, intention, stance, and communicative goals from contextual evidence\. This requirement shifts the problem from surface\-level recognition toward structured inference over latent social variables and the observable cues through which they are indirectly expressed\.

In the NLP domain, several lines of work have explored individual aspects of implicit social understanding\. For example, implicit sentiment analysis\[wei2020bilstm\]seeks to recognize emotions and affective attributes that are not directly expressed\. Instead, such states may be conveyed through sarcasm, passive aggressiveness, or multimodal incongruity\[wu2021modeling\]\. Similarly, intent understanding\[louvan2020recent\]aims to infer user intentions in task\-oriented settings\[zhang2016joint\]\. A representative example is identifying destination or time preferences in flight\-booking scenarios\[jbene2025intent\]\. These formulations provide useful supervision for recovering specific latent variables, but they commonly prioritize predefined labels, slots, or task\-completion goals\. Consequently, they do not necessarily model why an utterance adopts an indirect form, toward whom it is directed, or which social mechanism produces its implied meaning\.

However, these existing studies are largely confined to unimodal language settings\. In contrast, real\-world implicit social interactions are inherently multimodal\. Large\-scale resources such as CMU\-MOSEI and MELD demonstrate that linguistic, visual, and acoustic streams provide complementary evidence for affective understanding in naturalistic interactions\[zadeh2018cmumosei,poria2019meld\]\. Individuals frequently combine linguistic signals with visual cues, including facial expressions and meme imagery, to convey conflicting or indirect meanings and achieve subtle social goals\[farabi2024survey\]\. Such signals are not always redundant: agreement can reinforce an interpretation, whereas cross\-modal conflict can reverse the literal meaning or expose a concealed communicative intention\. Multimodal sarcasm detection\[castro2019towards,cai2019multimodal,qin2023mmsd2,guo2025multi\]is perhaps the most closely related line of work\. Nevertheless, its task formulation remains relatively shallow, typically focusing on label classification without requiring deeper reasoning about the underlying implications\. In particular, predicting a sarcasm label alone does not require a model to identify the social actors, recover their relational roles, explain the operative mechanism, or verify the resulting interpretation\.

Another related line of research focuses on abstract social cognition modeling\. Social\-IQ frames artificial social intelligence as question answering over in\-the\-wild videos, requiring models to reason about interactions rather than isolated affective labels\[zadeh2019socialiq\]\. Recent work on multimodal Theory of Mind\[jin2024mmtom,villa2025moments,shi2025muma\]and explanatory reasoning has moved beyond explicit grounding toward richer forms of socially informed inference\. Representative efforts includeMuMA\-ToM,MoMentS,Multimodal UNcommonsense, andDixitWorld\[shi2025muma,villa2025moments,son2026multimodal,mo2025dixitworld\]\. Collectively, these studies examine mental\-state inference, abductive hypothesis generation, explanatory reasoning, and the interpretation of events that cannot be resolved through direct perception alone\. Complementary work on abductive inference and commonsense knowledge generation recovers plausible latent causes, intentions, and surrounding events from partial observations\[bhagavatula2020abductive,park2020visualcomet,bosselut2019comet\]\. Nevertheless, these settings often simplify the social space into low\-dimensional mental variables, constrained game\-like environments, or physically unusual events\. These abstractions are useful for isolating particular cognitive capabilities, but they do not fully capture the deeply masked and strategically constructed expressions that characterize natural implicit interactions\. In such interactions, literal content, multimodal evidence, interpersonal relationships, and socio\-cultural expectations may jointly determine the intended meaning\. To address these limitations, this work presents, to our knowledge, the first systematic study and benchmark of implicit social context analysis in a multimodal setting\.

### 2\.2Multimodal Foundation Models for Social Reasoning

In the era of small\-scale neural networks, AI systems were largely incapable of handling or reasoning about such subtle and complex social phenomena\[sap2019social,vinciarelli2009social,forbes2020social,mathur2024advancing\]\. These systems were generally optimized for narrow prediction objectives and local correlations, with limited access to the broad knowledge needed for context\-sensitive social inference\. Recent advances in multimodal large language models \(MLLMs\) have demonstrated strong potential and substantially improved reasoning over multimodal inputs\. This progress is especially evident when models use multimodal Chain\-of\-Thought and compositional reasoning techniques\[gao2025interleaved,mitra2024compositional,wei2022chain,zhang2023multimodal,fei2023reasoning,fei2024videoofthought,zhang2025improve,li2025vegas\]\. These techniques can decompose a complex query into intermediate subproblems, organize heterogeneous evidence, and make parts of the inference process more explicit\.

Despite this progress, most existing approaches remain limited to relatively shallow forms of social reasoning that rely on explicit cues in the input\[mathur2025social\]\. Their evaluations also frequently emphasize answer or label correctness, providing limited evidence that the inferred social explanation is faithful to the multimodal observations\. Recent probing likewise shows that MLLMs can overlook implicit inconsistencies even when the relevant perceptual and reasoning capabilities are available\[yan2025hidden\]\. In complex implicit social scenarios, these models often fail because successful reasoning requires cognitive\-level understanding rather than direct cue matching\. Required knowledge includes social relationships, cultural norms, and background context, together with the ability to uncover deeper sources of conflict\[ziems2023normbank\]\. The model must explain why a particular expression is chosen, what social motivations underlie it, and how apparently inconsistent signals support a coherent interpretation\. The difficulty therefore extends beyond combining modalities: a reasoner must identify the expected social pattern, locate deviations from that expectation, and infer a plausible latent mechanism\. It must also jointly recover the subject, target, mechanism, and latent social state, while checking whether these components remain mutually consistent\. Therefore, this work proposes an effective solution by introducing a Conflict\-Driven Abductive Reasoning framework\. The framework treats conflict as an informative reasoning signal and progressively connects explicit perception, social context, expectation modeling, abductive inference, and consistency verification\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x2.png)Figure 2:The mechanism spaces of implicit social contexts\.

## 3Approaching Implicit Social Context Analysis

### 3\.1Task Definition

Here we try to lay a formal definition of the MoCA task\. MoCA recovers a structured implicit social interpretation, rather than a single label, from the multimodal input shown in Figure[1](https://arxiv.org/html/2608.05825#S1.F1)\. Letxtxtx\_\{\\text\{txt\}\},ximgx\_\{\\text\{img\}\},xvidx\_\{\\text\{vid\}\}, andxaudx\_\{\\text\{aud\}\}denote the textual, image, video, and audio inputs\. For scenarios∈\{Affection,Intent,Stance\}s\\in\\\{\\text\{Affection\},\\text\{Intent\},\\text\{Stance\}\\\}, the task is

\(xtxt,ximg,xvid,xaud\)↦\(ysubj,ytgt,ymechs,ylabels\),\(x\_\{\\text\{txt\}\},x\_\{\\text\{img\}\},x\_\{\\text\{vid\}\},x\_\{\\text\{aud\}\}\)\\mapsto\(y\_\{\\text\{subj\}\},y\_\{\\text\{tgt\}\},y^\{s\}\_\{\\text\{mech\}\},y^\{s\}\_\{\\text\{label\}\}\)\\,,\(1\)whereysubjy\_\{\\text\{subj\}\}andytgty\_\{\\text\{tgt\}\}identify the expressing subject and intended target, andymechsy^\{s\}\_\{\\text\{mech\}\}andylabelsy^\{s\}\_\{\\text\{label\}\}specify the realizing mechanism and latent label\. The superscriptssindicates that mechanism and label are selected from scenario\-specific spaces\.

The three scenarios specify different inference targets\.Affectioncovers the broader domain of latent affective states, not only interpersonal fondness, whose experience and observable expression may diverge\[scherer2005emotion,gross1998emotionregulation\]\.Intentcovers goal\-directed states that organize action through beliefs, desires, and anticipated behavior\[ajzen1991plannedbehavior,malle1997intentionality\]\.Stancecovers evaluative and relational positioning toward a target\[dubois2007stance\]\. These domains may co\-occur;ssmarks the primary inference target and does not imply psychological independence\.

### 3\.2Mechanism Space

MoCA does not assume that a latent label can be recovered from one modality or simple cue aggregation\. It also predicts themechanismthrough which the observed expression constructs or conceals that label\. Each scenario contains four mechanisms:

- •In Affection, mechanisms range from cross\-modal semantic conflict \(Multimodal Incongruity\) to deliberate emotional masking \(Affective Deception\);
- •In Intent, mechanisms range from overt but strategically framed hostility \(Expressive Aggression\) to covert behavioral manipulation \(Malicious Manipulation\);
- •In Stance, mechanisms range from power\-asymmetric disengagement \(Dominant Detachment\) to strategically compliant positioning \(Submissive Alignment\)\.

Figure[2](https://arxiv.org/html/2608.05825#S2.F2)presents the complete mechanism spaces\. A valid mechanism must explain the predicted label given the observed expression and recovered subject–target relation\.

### 3\.3Evaluation

Evaluation is scenario\-wise: models and humans are scored on the sameNsN\_\{s\}instances\. Each response is parsed into\(Subject,Target,Mechanism,Label\)\(\\mathrm\{Subject\},\\mathrm\{Target\},\\mathrm\{Mechanism\},\\mathrm\{Label\}\)and compared with its annotation by normalized exact match\. We lowercase strings, map hyphens and underscores to spaces, and collapse repeated separators; missing, unparsable, or scenario\-invalid fields are incorrect, without dropping the sample\. For𝒦=\{subj,tgt,mech,label\}\\mathcal\{K\}=\\\{\\mathrm\{subj\},\\mathrm\{tgt\},\\mathrm\{mech\},\\mathrm\{label\}\\\}, letci,kc\_\{i,k\}indicate a normalized match on fieldkk\.Atomic Component AccuracyreportsAcck=Ns−1​∑ici,k\\mathrm\{Acc\}\_\{k\}=N\_\{s\}^\{\-1\}\\sum\_\{i\}c\_\{i,k\}for each field;Dyadic AlignmentisDyad=Ns−1​∑ici,subj​ci,tgt\\mathrm\{Dyad\}=N\_\{s\}^\{\-1\}\\sum\_\{i\}c\_\{i,\\mathrm\{subj\}\}c\_\{i,\\mathrm\{tgt\}\};Mechanism FaithfulnessisFaith=∑ici,label​ci,mech/∑ici,label\\mathrm\{Faith\}=\\sum\_\{i\}c\_\{i,\\mathrm\{label\}\}c\_\{i,\\mathrm\{mech\}\}/\\sum\_\{i\}c\_\{i,\\mathrm\{label\}\}; andHolistic AlignmentisHolis=Ns−1​∑i∏k∈𝒦ci,k\\mathrm\{Holis\}=N\_\{s\}^\{\-1\}\\sum\_\{i\}\\prod\_\{k\\in\\mathcal\{K\}\}c\_\{i,k\}\. These metrics test component recovery, directed relations, output\-level mechanism–label consistency, and complete structured prediction, respectively\. Faithfulness does not claim access to the model’s internal reasoning\. Scores are reported as percentages\.

## 4Benchmarking Implicit Social Context Analysis

### 4\.1MoCA Dataset Construction

Unlike conventional datasets focusing on sentiment polarity or semantic consistency, MoCA captures implicit social semantics via structured relational tuples\. Our construction pipeline entails four stages as follows\.

Data Sources and Preprocessing\. To cover diverse implicit social expressions, we merge candidate samples from two complementary sources into a unified pool for downstream filtering and annotation\. The first comprises existing multimodal datasets: MUStARD\[castro2019towards\]and MIntRec\[zhang2022mintrec\]for video, and RedCaps\[desai2021redcaps\], MMSD2\.0\[qin2023mmsd2\], Hateful Memes\[kiela2020hatefulmemes\], MultiMET\[zhang2021multimet\], MemeCap\[hwang2023memecap\], and additional Instagram captions for image\-text data\. The second consists of supplementary collected data, such as meme\-text pairs and online debates, yielding rich indirect expressions shaped by semantic contrast, socio\-cultural contexts, and pragmatics\.

Unimodal Consistency Filtering\. To eliminate modality\-specific biases and ensure genuine cross\-modal dependency, we implement a consistency filtering step\. For each candidate, we evaluate predictions under three settings:visual\-only \(Rimg,RvidR\_\{\\text\{img\}\},R\_\{\\text\{vid\}\}\), text\-only \(RtxtR\_\{\\text\{txt\}\}\), and multimodal \(RmultiR\_\{\\text\{multi\}\}\)\. A sample is retained only if the multimodal interpretation diverges from both unimodal counterparts:

Rmulti∉\{Rimg,Rvid,Rtxt\}R\_\{\\text\{multi\}\}\\notin\\\{R\_\{\\text\{img\}\},R\_\{\\text\{vid\}\},R\_\{\\text\{txt\}\}\\\}\(2\)This strategy filters out instances inferable from a single modality, ensuring the dataset necessitates holistic multimodal reasoning\.

Cross\-Modal Structured Annotation\. Filtered samples are annotated via a 7\-tuple structured schema: \(Scenario, Domain, Culture, Subject, Target, Mechanism, Label\)\. Instead of isolated categorical labels, this schema deconstructs implicit social meaning into relational units\. Specifically,Mechanismencodes how implicit meaning is realized \(e\.g\., cross\-modal conflict, figurative expression, or socio\-cultural dependency\), facilitating evaluation of both predictive outcomes and reasoning structures\.SubjectandTargettranscend visually explicit entities to include implicit or socially presupposed individuals and groups\.

Quality Control and Final Screening\. Because implicit social interpretation is inherently subjective and context\-dependent, we employ a multi\-stage quality control protocol to ensure annotation consistency and logical validity\. The process comprises independent annotation, cross\-checking, and arbitration\. Samples are first annotated under unified guidelines, followed by consistency verification to identify conflicts, and finally adjudicated by experienced annotators to determine whether disputed cases should be revised, retained, or discarded\. All retained samples undergo final validation to ensure coherent scenario assignment, valid Subject–Target relations, pragmatically plausible Mechanisms, and interpretations supported by multimodal evidence\.

### 4\.2Dataset Statistics and Highlights

MoCA spans diverse social interaction contexts, from public discourse to interpersonal communication, covering both image\-text and video\-text modalities across varied socio\-cultural settings\. After multi\-stage filtering and manual verification, the dataset contains 3,108 high\-quality multimodal samples\. Dataset statistics are summarized in Table[1](https://arxiv.org/html/2608.05825#S4.T1)and Figure[3](https://arxiv.org/html/2608.05825#S4.F3)\.

Table 1:General statistics of MoCA dataset\.![Refer to caption](https://arxiv.org/html/2608.05825v1/x3.png)Figure 3:Distributions of labels and scenarios \(left\) and implicit\-social mechanisms \(right\)\.Implicit Social Scenarios and Labels\. Figure[3](https://arxiv.org/html/2608.05825#S4.F3)shows the distribution of three implicit social scenarios: Affection \(30\.12%\), Intent \(37\.93%\) and Stance \(31\.95%\)\. The inner ring denotes scenario categories, and the outer ring the fine\-grained label composition within each scenario\. Affection contains emotion\-related labels \(happy, bad, angry, disgusted, sad, fearful\)\. Intent contains social intention labels \(mock, mitigate, alienate, provoke, dominate, intimidate, condemn, denounce\)\. Stance contains stance\-related labels \(contemptuous, dismissive, disapproving, hostile, concerned, indifferent, skeptical, supportive, sympathetic, neutral, appreciative\)\. This hierarchical structure supports modeling at both scenario and fine\-grained social semantic levels\.

Diverse Domains in Social Interaction Contexts\. Table[1](https://arxiv.org/html/2608.05825#S4.T1)shows coverage across multiple social domains, including Networked, Occupational, Civic, Intimate, Familial, Peer, and Educational contexts\. The benchmark spans both public and interpersonal environments, enabling evaluation of implicit reasoning across heterogeneous communicative settings and social relations\.

Cultural Diversity and Contextual Variation\. Table[1](https://arxiv.org/html/2608.05825#S4.T1)reports the distribution of socio\-cultural contexts\. Most samples belong to General Context; the remainder span Middle Eastern, North American, South Asian, and East Asian contexts\. Models therefore require background knowledge to interpret this variation\.

Complementary Multimodal Coverage\. MoCA includes image\-text and video\-text inputs\. Image\-text samples provide compact symbolic context, often through cross\-modal semantic contrast, while video samples add temporal cues such as tone variation, actions, and interaction dynamics\. Together, these static and dynamic inputs broaden the evidence available for multimodal implicit reasoning\.

Implicit Mechanism Characteristics\. MoCA explicitly annotates Mechanism in addition to Label to characterize how implicit meaning is constructed\. Across the dataset, implicit meaning is realized through mechanisms including cross\-modal semantic conflict, figurative expression, socio\-cultural dependency, affective masking, and socially strategic expression\. Explicit modeling of Mechanism makes the reasoning process analyzable and enables evaluation of whether models capture the formation of implicit meaning\.

## 5CoDAR: Conflict\-Driven Abductive Reasoning

### 5\.1Theoretical Motivation

The MoCA task requires models to jointly recover structured, latent social information from multimodal inputs\. Unlike tasks that map observable cues directly to labels, MoCA must explain why an observed expression takes a specific form in a given situation\. We ground this requirement in three complementary frameworks from cognitive and social psychology\.

Impression ManagementImpression management theory argues that overt social behavior is shaped by situational pressures, anticipated audiences, and interactional goals\[sezer2022impression\]\. Consequently, observable expressions do not always provide transparent access to a person’s underlying affection, intent, or stance\. Individuals may instead suppress, reframe, or strategically mask internal states to protect relationships, preserve status, avoid sanctions, or pursue context\-dependent social outcomes\. In MoCA, explicit expressionxxis therefore socially produced evidence, not a direct readout of the latent state\.

Expectancy Violations TheoryExpectancy Violations Theory proposes that people continuously form context\-sensitive expectations about how others should behave within a given situation\[burgoon2015expectancy,buidze2025expectation\]\. We denote this normative expectation byee, which is conditioned on social roles, cultural norms, event structure, and relational dynamics\. When the observed expressionxxdeparts fromee, the deviation becomes salient because it violates the behavior predicted by the context\. Importantly, this conflict is not any superficial mismatch between modalities; it is a socially meaningful discrepancy between observed conduct and contextual expectation\. We denote this discrepancy by𝒞\\mathcal\{C\}, thereby converting implicit understanding into the problem of identifying and explaining a context\-grounded violation\.

Theory of MindTheory of Mind describes the human capacity to attribute beliefs, motives, intentions, and other latent mental states to social actors\[lake2017building,byom2013theory\]\. After a conflict is detected, an observer must infer which hidden state or strategic objective would make the otherwise unexpected expression rational\. This inference is abductive because the underlying cause is not directly observed and must be recovered as the explanation that best accounts for available evidence\. For MoCA, this perspective connects conflict resolution to the structured recovery of the subject, target, communicative mechanism, and latent social label\.

Together, these theories define a coherent account of implicit social reasoning\. Impression management explains why overt behavior can mask internal states; expectancy violations identify where that masking produces a salient discrepancy; and Theory of Mind explains how interpreters recover its cause\. Letxxdenote the observed explicit expression andeethe normative expectation under the current social context\. We formalize the semantic discrepancy between them as a conflict𝒞\\mathcal\{C\}\. Consequently, recovering latent social meaning is fundamentally equivalent to constructing an abductive explanation of the root cause of𝒞\\mathcal\{C\}\.

### 5\.2Reasoning Architecture

Building on the preceding theoretical motivation, we introduce Conflict\-Driven Abductive Reasoning \(CoDAR\), a five\-step framework for structured implicit social understanding\. CoDAR treats the discrepancy between observable evidence and contextual expectation as the central object of reasoning\. As illustrated in Figure[4](https://arxiv.org/html/2608.05825#S5.F4), the framework supports both image\-text and video\-text inputs\. Its five reasoning steps are organized into three functional stages that connect multimodal evidence, social context, expectation conflict, and structured interpretation\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x4.png)Figure 4:Overview of CoDAR, a conflict\-driven abductive reasoning chain for recovering structured implicit social meaning from multimodal input\.Stage I: Multimodal Evidence Grounding\.The first stage corresponds to the Explicit Perception Module \(Step 1\)\. It extracts directly observable cues from the multimodal input while avoiding premature interpretation of latent social meaning\. In the image\-text setting, the module records textual content together with static visual cues\. In the video\-text setting, it additionally captures temporally extended evidence\. Such evidence includes scene evolution, bodily behavior, interaction dynamics, and other event\-level cues unfolding over time\. Speech contained in a video is converted into textual transcripts and processed jointly with the visual stream\. These observations are consolidated into an explicit perception representation, denoted byxx, which provides the evidentiary basis for subsequent reasoning\.

Stage II: Contextual Expectation and Conflict Construction\.The second stage combines the Social Context Construction Module \(Step 2\) with the Expectation and Conflict Modeling Module \(Step 3\)\. Givenxx, Step 2 first identifies the background knowledge required to interpret the observed expression\. It then formulates focused questions concerning relevant facts, event connections, and social norms\. Each query can be assigned to a model\-prior, external web\-search, or hybrid retrieval route\. The resulting evidence is synthesized into three complementary fields:fact,connection, andsocial norm\. Together, these fields form a social context representationccthat situates the observation within its broader social environment\. This routing follows the broader principle of retrieval\-augmented and reasoning\-action systems, in which external evidence supplements parametric knowledge when the observation alone is insufficient\[lewis2020rag,yao2023react\]\.

Based oncc, Step 3 constructs a context\-conditioned expectationee\. This expectation describes what would normally occur under the given situation according to the relevant norms, event structure, and relational dynamics\. CoDAR then compares the explicit perceptionxxwitheeand extracts their socially meaningful discrepancy:

𝒞=Δ​\(x,e\)\.\\mathcal\{C\}=\\Delta\(x,e\)\.\(3\)Here,𝒞\\mathcal\{C\}identifies which observable cue departs from which contextual expectation\. The module also formulates an abductive question asking why the explicit reality is presented despite that expectation\. This transformation shifts the task from directly assigning a latent label to explaining why the context\-grounded deviation arises\.

Stage III: Staged Abductive Inference and Verification\.The final stage integrates the Abductive Reasoning Module \(Step 4\) with the Consistency Verification Module \(Step 5\)\. Given𝒞\\mathcal\{C\}, Step 4 performs a prompt\-guided sequence of counterfactual analysis, strategic interpretation, participant recovery, and taxonomy\-guided output selection\. It produces a structured interpretation

𝒜=\{ysubj,ytgt,ymechs,ylabels\},\\mathcal\{A\}=\\\{y\_\{\\text\{subj\}\},y\_\{\\text\{tgt\}\},y^\{s\}\_\{\\text\{mech\}\},y^\{s\}\_\{\\text\{label\}\}\\\},\(4\)where the subject and target are selected from instance\-specific candidates\. The mechanism and label are selected from the output spaces defined for scenarioss\.

Section[5\.3](https://arxiv.org/html/2608.05825#S5.SS3)describes this staged inference process in detail\. Step 5 evaluates the recovered interpretation through a Boolean verification function:

Φ​\(𝒜,x,c,𝒞\)∈\{0,1\}\.\\Phi\(\\mathcal\{A\},x,c,\\mathcal\{C\}\)\\in\\\{0,1\\\}\.\(5\)The verifier checks whether𝒜\\mathcal\{A\}is supported by the explicit evidence, compatible with the social context, and coherent with the identified conflict\. When all checks pass,Φ=1\\Phi=1, and the structured interpretation is accepted\. Otherwise,Φ=0\\Phi=0, and the verifier returns a rejection status together with a description of the detected breakpoint\. This feedback initiates a revision round that reconstructs the social context, conflict, and abductive interpretation before verification is performed again\. The cycle terminates when the interpretation is accepted or the predefined revision budget is exhausted\.

### 5\.3Conflict\-Constrained Staged Abductive Inference

![Refer to caption](https://arxiv.org/html/2608.05825v1/x5.png)Figure 5:Illustration of the Conflict\-Constrained Staged Abductive Inference mechanism\. Building on the explicit perceptionxx, social contextcc, and context\-conditioned expectationeeestablished in the preceding stages, CoDAR constructs the conflictC=Δ​\(x,e\)C=\\Delta\(x,e\)as the primary reasoning signal\. CoDAR first infers the counterfactual costrir\_\{i\}of direct, expectation\-aligned expression, and then derives the strategic functiongig\_\{i\}through which the observed conflict avoids this cost while preserving the subject’s social goal\. These inferences guide participant recovery and scenario\-specific mechanism–label assignment, yieldingAi=\{ysubj,ytgt,ymechs,ylabels\}∈Osubj\(i\)×Otgt\(i\)×Ms×LsA\_\{i\}=\\\{y\_\{\\mathrm\{\{subj\}\}\},y\_\{\\mathrm\{\{tgt\}\}\},y^\{s\}\_\{\\mathrm\{\{mech\}\}\},y^\{s\}\_\{\\mathrm\{\{label\}\}\}\\\}\\in O^\{\(i\)\}\_\{\\mathrm\{\{subj\}\}\}\\times O\_\{\\mathrm\{tgt\}\}^\{\(i\)\}\\times M\_\{s\}\\times L\_\{s\}Because conflict𝒞\\mathcal\{C\}may admit several explanations, Step 4 restricts inference to MoCA’s output schema and scenario\-specific taxonomies\. Figure[5](https://arxiv.org/html/2608.05825#S5.F5)summarizes the admissible tuple construction, staged abductive inference, and verification\-guided revision\. For instanceiiin scenarioss, let𝒪subj\(i\)\\mathcal\{O\}^\{\(i\)\}\_\{\\text\{subj\}\}and𝒪tgt\(i\)\\mathcal\{O\}^\{\(i\)\}\_\{\\text\{tgt\}\}denote the provided subject and target candidate sets\. Letℳs\\mathcal\{M\}\_\{s\}andℒs\\mathcal\{L\}\_\{s\}denote the scenario\-specific mechanism and label sets\. The admissible output space is

Ωi,s=𝒪subj\(i\)×𝒪tgt\(i\)×ℳs×ℒs\.\\Omega\_\{i,s\}=\\mathcal\{O\}^\{\(i\)\}\_\{\\text\{subj\}\}\\times\\mathcal\{O\}^\{\(i\)\}\_\{\\text\{tgt\}\}\\times\\mathcal\{M\}\_\{s\}\\times\\mathcal\{L\}\_\{s\}\.\(6\)CoDAR returns one structured tuple:

𝒜i=\{ysubj,ytgt,ymechs,ylabels\}∈Ωi,s\.\\mathcal\{A\}\_\{i\}=\\\{y\_\{\\text\{subj\}\},y\_\{\\text\{tgt\}\},y^\{s\}\_\{\\text\{mech\}\},y^\{s\}\_\{\\text\{label\}\}\\\}\\in\\Omega\_\{i,s\}\.\(7\)This constraint defines the form of𝒜i\\mathcal\{A\}\_\{i\}; it does not enumerate or rankΩi,s\\Omega\_\{i,s\}, learn a scorer, or introduce a separate decoding objective\.

Counterfactual Cost\.Given expectationee, CoDAR considers a direct, expectation\-aligned expression and infers its likely social, relational, strategic, or power\-related cost\. This cost explains why the subject may prefer an implicit expression; no mechanism or label is assigned at this stage\.

Strategic Function of Conflict\.CoDAR next asks how the observed deviation𝒞\\mathcal\{C\}avoids this cost while preserving the subject’s social goal\. The inferred strategy must explain the observed multimodal deviation;𝒞\\mathcal\{C\}itself is not treated as the final label\.

Participant Recovery\.Usingxx,cc, and𝒞\\mathcal\{C\}, CoDAR selects one subject from𝒪subj\(i\)\\mathcal\{O\}^\{\(i\)\}\_\{\\text\{subj\}\}and one target from𝒪tgt\(i\)\\mathcal\{O\}^\{\(i\)\}\_\{\\text\{tgt\}\}\. Selection follows communicative roles, not visual salience\.

Mechanism–Label Assignment\.CoDAR finally selects one mechanism fromℳs\\mathcal\{M\}\_\{s\}and one label fromℒs\\mathcal\{L\}\_\{s\}\. The mechanism specifies how the meaning is constructed or concealed, whereas the label identifies the latent social state\. Their assignment must agree with the inferred cost, strategic function, participant relation, and observed conflict\.

The complete staged process can be summarized as

\(x,c,e,𝒞\)\\displaystyle\(x,c,e,\\mathcal\{C\}\)⟶Counterfactual Cost\\displaystyle\\longrightarrow\\text\{Counterfactual Cost\}\(8\)⟶Strategic Function of Conflict\\displaystyle\\longrightarrow\\text\{Strategic Function of Conflict\}⟶\(ysubj,ytgt\)\\displaystyle\\longrightarrow\(y\_\{\\text\{subj\}\},y\_\{\\text\{tgt\}\}\)⟶\(ymechs,ylabels\)\.\\displaystyle\\longrightarrow\(y^\{s\}\_\{\\text\{mech\}\},y^\{s\}\_\{\\text\{label\}\}\)\.The final mapping uses the scenario taxonomy within prompt\-guided inference; it is not a separately optimized mechanism\-to\-label decoder\. The mechanism remains an explicit intermediate field linking𝒞\\mathcal\{C\}to the predicted label\.

Consistency Verification\.The verification functionΦ\\Phichecks three conditions:

- •Explicit evidence consistency:𝒜i\\mathcal\{A\}\_\{i\}is supported by observable cues inxx;
- •Social context consistency:𝒜i\\mathcal\{A\}\_\{i\}is compatible with the facts, connections, and norms incc;
- •Mechanism consistency:the mechanism explains𝒞\\mathcal\{C\}and its inferred strategic function\.

If all checks pass,𝒜i\\mathcal\{A\}\_\{i\}is accepted\. Otherwise, the reported breakpoint triggers revision over Steps 2–5\.

Table[5](https://arxiv.org/html/2608.05825#S6.T5)evaluates Steps 3 and 4 jointly\. Removing them yields the largest overall degradation:Holistic Alignmentfalls by 5\.64 and 3\.82 percentage points in Affection and Stance, respectively, whileMechanism Faithfulnessfalls by 16\.55 points in Intent\. Because both steps are removed together, this ablation supports the combined unit but cannot attribute the gains to either step or to an individual inference stage\.

## 6Experiments

### 6\.1Settings

Models\.We evaluate three groups of MLLMs under the same MoCA task: open\-weight, affect\-specialized, and proprietary models\. Together, they support comparisons across model scale, reasoning variants, and domain specialization\.

Open\-weight MLLMs\.The open\-weight baselines are LLaVA\-OneVision\-7B and LLaVA\-OneVision\-70B\[li2024llavaonevisioneasyvisualtask\], InternVL3\.5\-38B\[wang2025internvl35advancingopensourcemultimodal\], MiniCPM\-V\-4\.5\[yu2025minicpmv45cookingefficient\], and GLM\-4\.6V\-Flash\-9B\[hong2025glm\]\. The two LLaVA\-OneVision models enable a within\-family comparison across scales\. We additionally evaluate Qwen3\-VL\-32B\-Instruct and Qwen3\-VL\-32B\-Thinking\[bai2025qwen3vltechnicalreport\]\. Their shared family and nominal scale provide a direct comparison between instruction\-oriented and thinking\-oriented variants, although the observed differences cannot be attributed solely to inference style\.

Affect\-specialized MLLMs\.To test whether explicit affect modeling transfers to implicit social reasoning, we evaluate AffectGPT\[lian2025affectgpt\]and Emotion\-Qwen\[huang2025emotionqwenunifiedframeworkemotion\]under the same structured prediction objective\.

Proprietary MLLMs\.We use Gemini\-3\.1\-Pro\[gemini2025\], Claude\-Sonnet\-4\.6\[claudeSonnet46SystemCard2026\], ChatGPT\-5\.4\[gpt5systemcard2025\], and ChatGPT\-4o\[gpt4o2024\]as proprietary capability references\. Their undisclosed architectures and training data preclude controlled architectural attribution\.

Input Processing\.Each image\-text instance is presented as the original caption and its paired image, preserving the evidence required to resolve cross\-modal agreement or conflict\. Each video is uniformly sampled into a temporally ordered frame sequence, while the original textual input remains available to the model\. The reported video results therefore measure reasoning over sampled temporal evidence rather than uninterrupted raw\-video perception\. Speech is transcribed into captions that retain linguistic content and speaking style\. These transcripts are combined with the corresponding visual evidence\. Audio results therefore measure transcript\-conditioned rather than direct acoustic reasoning\.

Evaluation Protocol\.Every model predicts the same four fields: Subject, Target, Mechanism, and Label\. The normalization, invalid\-output rule, and four metrics defined in Section[3\.3](https://arxiv.org/html/2608.05825#S3.SS3)are applied unchanged to model and human responses\.

### 6\.2Main Results and Observations

Tables[2](https://arxiv.org/html/2608.05825#S6.T2)–[4](https://arxiv.org/html/2608.05825#S6.T4)report all scenario results\.

Table 2:Performance \(accuracy\) in the Affection Scenario\. Lab \(Label\), Subj \(Subject\), Tgt \(Target\), Mech \(Mechanism\), Dyad \(Dyadic Alignment\), Faith \(Mechanism Faithfulness\), and Holis \(Holistic Alignment\)\. The best results within each category are ingreen\.Light bluehighlights human performance, whilelight purpleindicates CoDAR\-augmented models\.\+and\-indicate CoDAR’s performance improvement and degradation over the backbones, respectively\.Performance Disparity Across Dimensions\. A consistent pattern is observed across all scenarios: model performance remains relatively strong on dyadic perception but decreases substantially on reasoning\-related dimensions\. Comparing the best\-performing models including CoDAR to human performance, the gap on Dyadic Alignment remains limited \(3\.19%–7\.46%\), indicating that current MLLMs can reliably identify interaction participants\. The gap increases on Label accuracy \(31\.04%–34\.35%\), which requires recovering implicit social states, and further enlarges on Holistic Alignment \(37\.38%–41\.89%\), where all components must be correct simultaneously\. This decline indicates that MoCA’s main difficulty is recovering the communicative configuration, not identifying participants\.

Model Performance Comparison\. Proprietary models consistently outperform open\-weight models, although the magnitude varies across evaluation dimensions\. On Subject and Target, strong open\-weight models such as Qwen3\-VL\-32B\-Instruct achieve performance comparable to proprietary systems\. But on Holistic Alignment, the gap widens: the best proprietary baseline \(24\.69%\) exceeds the best open\-weight baseline \(15\.71%\) by nearly 10 percentage points in Stance, yet both remain far below human performance \(67\.64%\)\. Affect\-specialized models do not compensate for this gap\. AffectGPT achieves 43\.48% Mechanism accuracy in Affection but only 3\.95% Dyadic Alignment and 0\.64% Holistic Alignment, indicating that explicit emotion recognition does not translate to improved implicit relational understanding\. Qwen3\-VL\-32B\-Thinking also performs below its Instruct variant in Affection \(5\.24% versus 9\.08%\), suggesting that explicit reasoning traces do not necessarily improve consistency between predicted labels and mechanisms\.

Effect of CoDAR\. CoDAR yields consistent improvements across model families, with the largest gains in Mechanism Faithfulness and Holistic Alignment\. We apply CoDAR to representative models from each category; Intent additionally includes LLaVA\-OneVision\-70B because of its strong baseline, yielding three evaluated models\. In Affection, Mechanism Faithfulness increases from 42\.48% to 69\.16% for ChatGPT\-5\.4 and from 41\.62% to 59\.68% for Qwen3\-VL\-32B\-Instruct\.

Table[2](https://arxiv.org/html/2608.05825#S6.T2)further localizes the Affection bottleneck\. Relative to human performance, the best machine results are only 2\.39 and 4\.72 percentage points lower on Subject and Target, respectively, but the gaps widen to 31\.11 points on Label and 41\.89 points on Holistic Alignment\.

The same separation appears within individual baselines\. Gemini\-3\.1\-Pro reaches 80\.28% Subject and 82\.96% Target accuracy, but only 14\.58% Holistic Alignment\. AffectGPT attains 43\.48% Mechanism accuracy while obtaining 3\.95% Dyadic Alignment and 0\.64% Holistic Alignment\. These values show that recovering one participant or mechanism does not by itself satisfy the joint four\-output requirement\. CoDAR’s changes reinforce this distinction\. For Qwen3\-VL\-32B\-Instruct, it improves Mechanism Faithfulness by 18\.06 points while raising Target and Holistic Alignment by 2\.46 points each; for ChatGPT\-5\.4, the corresponding gains are 26\.68, 0\.64, and 5\.13 points\. Thus, the principal improvement lies in making predicted labels more consistent with their mechanisms, rather than uniformly increasing every component\. Nevertheless, the best machine Holistic Alignment remains 17\.84%, compared with 59\.73% for human annotators\.

Table 3:Intent\-scenario accuracy; notation and metric abbreviations follow Table[2](https://arxiv.org/html/2608.05825#S6.T2)\.In Intent, Holistic Alignment improves by 7\.80% for ChatGPT\-5\.4 and 9\.33% for Qwen3\-VL\-32B\-Instruct\. Table[3](https://arxiv.org/html/2608.05825#S6.T3)also shows why component\-level scores alone are insufficient: GLM\-4\.6V\-Flash\-9B obtains 70\.23% Mechanism accuracy and 82\.38% Mechanism Faithfulness, yet only 10\.34% Holistic Alignment; ChatGPT\-4o reaches 72\.53%, 84\.54%, and 15\.36% on the same metrics\. Across all five reported CoDAR pairs, Holistic Alignment increases by 6\.43–9\.33 points, despite substantial variation in the accompanying atomic\-component gains\. The consistent improvement in the joint metric therefore reflects better coordination of the four required outputs, rather than strength on one isolated component\.

These gains indicate that organizing reasoning around conflict identification and abductive explanation helps bridge the gap between surface perception and mechanism\-level understanding\. Subject and Target performance remains stable, confirming that improvement originates from deeper reasoning rather than perceptual enhancement\.

Table 4:Stance\-scenario accuracy; notation and metric abbreviations follow Table[2](https://arxiv.org/html/2608.05825#S6.T2)\.Table[4](https://arxiv.org/html/2608.05825#S6.T4)exhibits the same separation between component recovery and complete structured prediction\. Gemini\-3\.1\-Pro is the strongest baseline on Holistic Alignment at 24\.69%, but remains 42\.95 points below human performance\. CoDAR raises Holistic Alignment by 6\.65 points for InternVL3\.5\-38B, 4\.93 points for Qwen3\-VL\-32B\-Instruct, and 4\.64 points for ChatGPT\-5\.4\. For InternVL3\.5\-38B, this improvement accompanies a 17\.02\-point increase in Dyadic Alignment but only a 0\.30\-point change in Mechanism accuracy\. The contrast further shows that recovering a coherent Subject–Target relation and integrating it with the predicted mechanism and label are distinct requirements of the task\.

### 6\.3Ablation Analysis

To verify the contribution of each step in CoDAR, we design three ablation settings:1\) w/o Step 2:removing the social context construction step;2\) w/o Step 3&4:removing both the expectation\-and\-conflict modeling step and the abductive reasoning step, which together constitute the core reasoning unit of CoDAR;3\) w/o Step 5:removing the consistency verification step\. We evaluate all variants across the three scenarios, with the results reported in Table[5](https://arxiv.org/html/2608.05825#S6.T5)\.

Table 5:Ablation study results across three scenarios based on Qwen3\-VL\-32B\-Instruct\. We report the average of the four atomic metrics \(Lab,Subj,Tgt,Mech\) asAtomic Avg\., together withDyad,Faith, andHolis\. Green indicates drops relative to full CoDAR, while red indicates gains\.Overall, the full CoDAR model achieves the most stable performance, and all ablated variants perform worse onHolistic Alignment\. This confirms that CoDAR’s gains arise from the complete reasoning loop rather than any single step in isolation\. Among all ablations, w/o Step 3&4 causes the most substantial and consistent degradation across all metrics, withHolistic Alignmentdropping by 5\.64% in Affection and 3\.82% in Stance, confirming that contextual expectation construction, conflict modeling, and abductive explanation form the central source of CoDAR’s advantage\. An interesting pattern emerges in Intent and Stance, where all ablated variants show higherDyadic Alignmentthan the full model\. However, this improvement co\-occurs with sharp drops inMechanism Faithfulness\(up to 16\.55% in Intent\) andHolistic Alignment, suggesting that without the full reasoning chain, models fall back on surface\-level relational cues to matchSubjectandTarget\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x6.png)Figure 6:Baseline radar comparison across three scenarios with a holistic metric summary\. Solid lines indicate open\-weight models, while dashed lines indicate proprietary models\. Colors are consistent for each model across subplots\.

## 7Analysis and Discussion

### 7\.1Performance across Evaluation Metrics

RQ\-1: Where Does the Performance Gap between Open\-Weight and Proprietary MLLMs Emerge?Figure[6](https://arxiv.org/html/2608.05825#S6.F6)a–c and Tables[2](https://arxiv.org/html/2608.05825#S6.T2)–[4](https://arxiv.org/html/2608.05825#S6.T4)show that the gap is uneven across the seven metrics\. Within the six plotted baselines, the strongest open\-weight and proprietaryMechanismscores remain close in all scenarios\. Their gaps are only 0\.45, 1\.47, and 0\.51 percentage points in Affection, Intent, and Stance, respectively\.Mechanism Faithfulnessis similarly close in Intent, at 81\.04% versus 81\.50%, and Stance, at 77\.38% versus 81\.46%\. In Affection, InternVL3\.5\-38B instead exceeds the best plotted proprietary baseline on this metric, at 46\.53% versus 43\.97%\. The correspondingLabelgaps are 5\.81, 7\.30, and 7\.60 points, while theTargetgaps are 5\.04, 5\.77, and 5\.63 points\. TheSubjectgap increases from 3\.42 points in Affection to 10\.81 and 11\.95 points in Intent and Stance\.

The separation is larger and consistent on the joint metrics\. The strongest plotted proprietary baseline exceeds the strongest open\-weight baseline onDyadic Alignmentby 13\.17, 17\.33, and 16\.22 points across the three scenarios\. The correspondingHolistic Alignmentgaps are 4\.54, 5\.64, and 8\.98 points, with the largest separation in Stance \(Figure[6](https://arxiv.org/html/2608.05825#S6.F6)d\)\. However, even the strongest proprietary baseline remains 45\.15, 45\.18, and 42\.95 points below humanHolistic Alignment, respectively\. These comparisons locate the model\-family gap primarily in joint relational and four\-component consistency, rather than uniform superiority on every atomic component\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x7.png)Figure 7:Comparison of AffectGPT, Emotion\-Qwen, and Gemini\-3\.1\-Pro on seven metrics in the Affection scenario\. Horizontally hatched bars denote AffectGPT, vertically hatched bars denote Emotion\-Qwen, and lightly shaded bars denote Gemini\-3\.1\-Pro\.RQ\-2: Can Affect\-Specialized Models Excel in the Affection Scenario?Figure[7](https://arxiv.org/html/2608.05825#S7.F7)and Table[2](https://arxiv.org/html/2608.05825#S6.T2)show that neither affect\-specialized baseline matches Gemini\-3\.1\-Pro on relational or holistic recovery\. AffectGPT reaches 43\.48%Mechanismaccuracy, compared with 42\.12% for Gemini\-3\.1\-Pro, but records only 3\.95%Dyadic Alignmentand 0\.64%Holistic Alignment\. Emotion\-Qwen records 15\.71%Dyadic Alignment, yet itsHolistic Alignmentremains 0\.75%\. By comparison, Gemini\-3\.1\-Pro obtains 74\.92% and 14\.58% on these two metrics, respectively\.

Table[2](https://arxiv.org/html/2608.05825#S6.T2)further shows that CoDAR raisesHolistic Alignmentfrom 0\.64% to 2\.88% for AffectGPT and from 0\.75% to 1\.82% for Emotion\-Qwen\. These low absolute scores show that component\-level strength does not ensure directed Subject–Target recovery or four\-output alignment\. This finding applies only to these two models and does not rule out broader benefits from affect specialization\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x8.png)Figure 8:Scatter plot showing correlation between Atomic Component Accuracy for Label and Mechanism Faithfulness\. Points with red borders denote CoDAR\-augmented models, partitioned into quadrants by blue dashed lines\.RQ\-3: Does the Thinking Variant Improve Implicit Social Reasoning?Across Tables[2](https://arxiv.org/html/2608.05825#S6.T2)–[4](https://arxiv.org/html/2608.05825#S6.T4), the Thinking variant records lowerMechanismaccuracy than the matched Qwen3\-VL\-32B\-Instruct model in all scenarios\. The differences are 15\.92, 28\.76, and 6\.04 percentage points in Affection, Intent, and Stance\.Holistic Alignmentis also lower by 3\.84, 5\.17, and 3\.63 points\. Target accuracy changes by\+0\.53\+0\.53,\+3\.82\+3\.82, and\+0\.10\+0\.10points, whereas Subject accuracy changes by−11\.33\-11\.33,−9\.58\-9\.58, and−11\.17\-11\.17points\.Mechanism Faithfulnesslikewise changes by−16\.32\-16\.32,−22\.41\-22\.41, and−6\.69\-6\.69points\.

The largest differences therefore concentrate in Subject and mechanism\-related metrics, whereas Target changes comparatively little\. The reported results do not identify why this pattern occurs\. They establish only that the Thinking variant does not improveHolistic Alignmentfor this matched model pair\. This conclusion should not be generalized to all thinking models or reasoning paradigms\.

RQ\-4: To What Extent Do Models Achieve Label Accuracy through Mechanism\-Level Understanding?A model may predict the correct label without identifying the corresponding mechanism\. To quantify this discrepancy, we define a two\-dimensional space using Label accuracyx=PLx=P\_\{L\}and Mechanism Faithfulnessy=PM\|Ly=P\_\{M\|L\}, as shown in Figure[8](https://arxiv.org/html/2608.05825#S7.F8)\. The conditional formulation measures the probability that a model predicts the correct label while failing to recover the corresponding mechanism:

PL∩¬M=PL−PL×PM\|L=PL​\[1−PM\|L\]\.P\_\{L\\cap\\neg M\}=P\_\{L\}\-P\_\{L\}\\times P\_\{M\|L\}=P\_\{L\}\\left\[1\-P\_\{M\|L\}\\right\]\.\(9\)Higher values indicate weaker consistency between Label accuracy and mechanism recovery\. Models in the upper\-right region, such as ChatGPT\-5\.4 in Stance, achieve 38\.57% onPLP\_\{L\}and 81\.46% onPM\|LP\_\{M\|L\}, showing consistent performance across both dimensions\. In contrast, models in the lower\-right region, such as Gemini\-3\.1\-Pro in Affection, achieve comparable Label accuracy of 45\.34% but lower Mechanism Faithfulness of 43\.97%\. After applying CoDAR, multiple models shift toward the upper\-right region, indicating improved alignment between Label accuracy and mechanism recovery\. For example, CoDAR increases Mechanism Faithfulness of ChatGPT\-5\.4 in Affection from 42\.48% to 69\.16%, corresponding to an improvement of 26\.68%\. These shifts indicate that CoDAR better aligns label and mechanism predictions\.

Table 6:Within\-family scaling from LLaVA\-OneVision\-7B to LLaVA\-OneVision\-70B\.Δ\\DeltaDyad andΔ\\DeltaHolis denote percentage\-point changes inDyadic AlignmentandHolistic Alignment\.Csem∣dyadC\_\{\\mathrm\{sem\}\\mid\\mathrm\{dyad\}\}is the percentage of dyad\-correct predictions that are also holistically correct\.RQ\-5: Does Within\-Family Scaling Resolve Structured Implicit Reasoning?Table[6](https://arxiv.org/html/2608.05825#S7.T6)provides a within\-family comparison between LLaVA\-OneVision\-7B and LLaVA\-OneVision\-70B\. Increasing model scale raisesDyadic Alignmentby 37\.53, 35\.24, and 39\.12 points in Affection, Intent, and Stance, respectively\. The corresponding gains inHolistic Alignmentare substantially smaller at 4\.82, 13\.11, and 7\.53 points\. Thus, the larger model recovers interaction participants much more reliably, but only part of this improvement transfers to the complete four\-field output\.

This separation remains visible after normalizing for dyad recovery\. The conditional completion rate rises from 7\.16% to 10\.81% in Affection, from 5\.55% to 24\.73% in Intent, and from 9\.74% to 15\.96% in Stance\. These values remain far below the corresponding human rates of 76\.05%, 76\.41%, and 79\.80%\. The 70B checkpoint therefore alleviates, but does not remove, the compositional bottleneck\. Because the two checkpoints may differ in factors beyond parameter count, this result is interpreted as a within\-family contrast rather than a general scaling law\.

Table 7:Semantic completion after correct Subject–Target recovery\. Baseline and CoDAR values are macro\-averaged over InternVL3\.5\-38B, Qwen3\-VL\-32B\-Instruct, ChatGPT\-5\.4, and Gemini\-3\.1\-Pro\. Values are percentages;Δ\\DeltaCoDAR is the percentage\-point gain\.RQ\-6: What Remains Unresolved after Correct Subject–Target Recovery?Since every holistically correct prediction is necessarily dyad\-correct, we define the conditional semantic completion rate as

Csem∣dyad=Holistic AlignmentDyadic Alignment×100\.C\_\{\\mathrm\{sem\}\\mid\\mathrm\{dyad\}\}=\\frac\{\\textit\{Holistic Alignment\}\}\{\\textit\{Dyadic Alignment\}\}\\times 100\.\(10\)Table[7](https://arxiv.org/html/2608.05825#S7.T7)reports this diagnostic, which estimates how often a model also recovers the correct Mechanism and Label once both interaction participants are correct\. Across the four matched backbones, the baseline macro\-average is 17\.78%, 24\.69%, and 27\.66% for Affection, Intent, and Stance, respectively\. CoDAR increases these rates to 21\.97%, 31\.73%, and 31\.31%, corresponding to gains of 4\.18, 7\.05, and 3\.65 points\.

The remaining gap is nevertheless large\. Even with CoDAR, fewer than one third of dyad\-correct predictions become complete structured outputs in any scenario, whereas human completion ranges from 76\.05% to 79\.80%\. The principal residual error therefore lies after participant identification: models still struggle to bind the recovered relation to a compatible mechanism and latent social label\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x9.png)Figure 9:Scenario\-wise localization of CoDAR gains\. Bars report mean percentage\-point improvements over InternVL3\.5\-38B, Qwen3\-VL\-32B\-Instruct, ChatGPT\-5\.4, and Gemini\-3\.1\-Pro\. Horizontally hatched, solid, and vertically hatched bars denoteΔ\\DeltaDyad,Δ\\DeltaFaith, andΔ\\DeltaHolis, respectively; numbers give exact values, and the shaded Overall group gives the macro\-average across scenarios\.RQ\-7: Does CoDAR Improve the Same Capability in Every Scenario?We localize CoDAR’s gains through 12 matched comparisons comprising four backbones evaluated before and after CoDAR in all three scenarios\. As shown in Figure[9](https://arxiv.org/html/2608.05825#S7.F9), the gain profile is scenario\-dependent\. In Affection, the dominant change is inMechanism Faithfulness\(\+18\.97\+18\.97points on average\), whereasDyadic AlignmentandHolistic Alignmentrise by 3\.58 and 3\.56 points\. In Intent, the largest gain shifts toDyadic Alignment\(\+11\.42\+11\.42\), accompanied by the strongest average improvement inHolistic Alignment\(\+7\.82\+7\.82\)\. Stance exhibits a more balanced profile, with gains of 7\.23, 3\.76, and 4\.36 points on the three metrics\.

All 12 model–scenario pairs improve simultaneously onDyadic Alignment,Mechanism Faithfulness, andHolistic Alignment; their overall mean gains are 7\.41, 8\.62, and 5\.25 points, respectively\. CoDAR therefore does not obtain its holistic gains by trading relation recovery against mechanism consistency\. Instead, the relative source of improvement changes with the scenario, consistent with the different conflict structures that Affection, Intent, and Stance impose\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x10.png)Figure 10:Visual comparison between Qwen3\-VL\-32B\-Instruct w/ vs\. w/o our CoDAR framework, where the baseline remains limited to surface\-level prediction, while CoDAR proceeds through conflict modeling and abductive reasoning\.
### 7\.2Qualitative Case Study

To illustrate the performance gap between baselines and CoDAR, we present a representative test case, as shown in Figure[10](https://arxiv.org/html/2608.05825#S7.F10)\. The baseline Qwen3\-VL\-32B\-Instruct detects polarity contrast between a calm visual cue and the negative textual cue Monday, and interprets the relation as generalizedSurface Mismatching, leading to incorrect prediction ofMultimodal Incongruityand the label Happy\. In contrast, CoDAR constructs an expectation\-driven reasoning process by establishing contextual expectation that Monday implies work and productivity, which conflicts with the sleeping cat\. Through abductive reasoning, CoDAR interprets this conflict as figurative projection of the subject internal unwillingness to begin the week rather than simple image text inconsistency\. By identifying theFigurative Semanticsmechanism and the latent Sad emotion, this example demonstrates CoDAR ability to move beyond shallow visual text matching toward implicit social reasoning grounded in relational structure\.

## 8Conclusion

The paper introduces a novel task,Implicit Social Context Analysis\(MoCA\), that formalizes implicit social context analysis as a structured multimodal reasoning\. By requiring models to jointly recover the subject, target, mechanism, and latent social state across Affection, Intent, and Stance scenarios, MoCA provides the first systematic evaluation of whether multimodal large models can perform human\-level cognitive reasoning in implicit social interactions\. We construct a high\-quality benchmark with fine\-grained cognitive annotations and show that current multimodal large language models struggle significantly on this task due to their reliance on explicit cues\. To address this limitation, we propose theConflict\-Driven Abductive Reasoning\(CoDAR\) framework, which models cognitive conflict to infer hidden mental states\. Experiments reveal that current models exhibit a substantial and widening gap from human performance\. Further results demonstrate that CoDAR substantially improves performance, while a notable gap to human reasoning remains\.

## 9Ethics Statement

The MoCA benchmark was constructed exclusively from publicly available multimodal sources, including existing datasets and supplementary real\-world materials such as meme–text pairs and online discussions\. The development of MoCA was supported by a structured human\-annotation process\. All annotators were graduate students at the master’s level or above, including master’s and doctoral students, and received formal training in the annotation protocol to ensure consistent and reliable annotation\. The annotation process included independent annotation, cross\-checking, disagreement arbitration, and final validation of the scenario, domain, culture, Subject, Target, Mechanism, and Label fields\.

## References

## Appendix Overview

- •Appendix[A](https://arxiv.org/html/2608.05825#A1)defines the tasks and annotation spaces\.
- •Appendix[B](https://arxiv.org/html/2608.05825#A2)details benchmark construction and statistics\.
- •Appendix[C](https://arxiv.org/html/2608.05825#A3)describes the CoDAR implementation and inference process\.
- •Appendix[D](https://arxiv.org/html/2608.05825#A4)presents representative case studies\.

## Appendix ATask Definition Supplement and Concept Boundary

### A\.1Detailed Definitions and Boundaries of Three Implicit Social Scenarios

Affection\.Implicit Affection concerns the subject’s genuine emotional state when that state is not expressed directly\. The state may instead be conveyed through cross\-modal conflict, figurative language, culture\-dependent meaning, or deliberate emotional masking\. The inference target is what the subject feels, rather than the surface polarity of an isolated modality\.

Intent\.Implicit Intent concerns the interactional outcome that the subject seeks to produce\. The relevant goal may be concealed by humor, praise, aggression, avoidance, or another strategically framed action\. The inference target is what the subject aims to achieve, rather than the emotion accompanying the behavior\.

Stance\.Implicit Stance concerns the subject’s evaluative and relational position toward a target\. It includes support or opposition, relational distance, and power posture\. The inference target is how the subject positions themself with respect to the target, rather than the subject’s private emotion or desired outcome\.

Boundary rule\.One observation may support interpretations in all three domains\. Annotation follows the central latent variable required to explain the expression: Affection asks what is felt, Intent asks what outcome is pursued, and Stance asks what relational position is adopted\. Surface emotion words, rhetorical form, and visual salience do not determine the scenario by themselves\.

### A\.2Data Field Definitions

Each instance is represented as

\(Scenario,Domain,Culture,Subject,Target,Mechanism,Label\)\.\(\\mathrm\{Scenario\},\\mathrm\{Domain\},\\mathrm\{Culture\},\\mathrm\{Subject\},\\mathrm\{Target\},\\mathrm\{Mechanism\},\\mathrm\{Label\}\)\.The first three fields delimit the interpretation context, Subject and Target recover the directed social relation, and Mechanism and Label distinguish the construction of implicit meaning from its latent outcome\.

Table 8:Operational definitions of the seven annotation fields\.
### A\.3Mechanism Space Overview

Mechanism is treated as an explanatory bridge between observable evidence and the final Label\. The output spaces are scenario\-specific so that a plausible verbal explanation cannot drift into a mechanism or label belonging to a different scenario\.

Table 9:Mechanism inventory retained by the current MoCA formulation\.
### A\.4Label Space Overview

The label inventory encodes the scenario\-specific latent social outcome and is interpreted jointly with the corresponding Mechanism rather than as a free\-form description\. Table[10](https://arxiv.org/html/2608.05825#A1.T10)lists the admissible labels used for structured evaluation\.

Table 10:Valid label sets under each scenario\.

## Appendix BMoCA Construction and Annotation Details

### B\.1Construction Pipeline

The source pool combines established multimodal benchmarks with publicly accessible real\-world material\. Video\-text candidates are drawn from MUStARD\[castro2019towards\]and MIntRec\[zhang2022mintrec\]\. Image\-text candidates include RedCaps\[desai2021redcaps\], MMSD2\.0\[qin2023mmsd2\], Hateful Memes\[kiela2020hatefulmemes\], MultiMET\[zhang2021multimet\], MemeCap\[hwang2023memecap\], Instagram captions, meme\-text pairs, and online debates\. These sources provide complementary evidence from conversational video, social\-media imagery, figurative expression, semantic incongruity, and public interaction\.

Preprocessing removes invalid links, markup, system artifacts, corrupted media, and duplicate candidates while preserving semantically relevant emojis and hashtags\. Video frames remain temporally ordered, and speech is represented by transcripts that preserve linguistic content and speaking style\. Consequently, the reported audio\-related setting evaluates transcript\-conditioned multimodal reasoning rather than direct acoustic perception\.

#### B\.1\.1Unimodal Consistency Filtering

Candidate instances are evaluated under visual\-only, text\-only, and multimodal conditions\. LetRvisR\_\{\\mathrm\{vis\}\}denote the image\-only or video\-only interpretation,RtxtR\_\{\\mathrm\{txt\}\}the text\-only interpretation, andRmultiR\_\{\\mathrm\{multi\}\}the interpretation from the complete input\. A candidate proceeds to structured annotation only when

Rmulti≠RvisandRmulti≠Rtxt\.R\_\{\\mathrm\{multi\}\}\\neq R\_\{\\mathrm\{vis\}\}\\quad\\text\{and\}\\quad R\_\{\\mathrm\{multi\}\}\\neq R\_\{\\mathrm\{txt\}\}\.\(B\.1\)The criterion rejects cases for which a single modality already determines the intended interpretation\. It is a screening rule for cross\-modal dependence, not a claim that every retained instance contains a literal polarity reversal\.

### B\.2Annotation

All annotators were graduate students at the master’s level or above, including master’s and doctoral students\. They received formal training on scenario boundaries, field definitions, mechanism criteria, and common sources of multimodal ambiguity\. The operational workflow comprised four stages:

1. 1\.Independent annotation\.An annotator reviewed the complete multimodal input and assigned all seven fields under a shared codebook\.
2. 2\.Field\-level cross\-checking\.A separate reviewer examined scenario assignment, Subject–Target direction, Mechanism–Label compatibility, and evidential support\.
3. 3\.Disagreement arbitration\.Disputed fields were revised, retained, or rejected after joint inspection of the multimodal evidence and annotation rationale\.
4. 4\.Final validation\.Every retained sample was checked for a valid scenario, a coherent directed relation, a scenario\-valid Mechanism and Label, and sufficient multimodal support\.

### B\.3Quality Control

This process treats disagreement as structural rather than label\-only\. For example, matching labels do not constitute agreement when annotators assign different Subjects, reverse the relation direction, or select a Mechanism that does not explain the Label\. Numerical agreement coefficients are not introduced here because the current released materials do not include the per\-annotator records required to reproduce them\.

### B\.4Detailed Statistics

The final benchmark contains 3,108 instances\. Its 936 Affection, 1,179 Intent, and 993 Stance samples correspond to 30\.12%, 37\.93%, and 31\.95% of the benchmark\. The modality split comprises 2,230 image\-text and 878 video\-text instances\. All seven structured fields are populated, yielding 21,756 final field values\.

Table 11:Structural and modality composition of the final MoCA benchmark\.

## Appendix CCoDAR Details

### C\.1Evaluation and Input Interface

Models and humans receive the same scenario\-specific instances and return the same four evaluated fields:\(Subject,Target,Mechanism,Label\)\(\\mathrm\{Subject\},\\mathrm\{Target\},\\mathrm\{Mechanism\},\\mathrm\{Label\}\)\. Predictions and annotations are lowercased, hyphens and underscores are mapped to spaces, and repeated separators are collapsed before exact matching\. A missing, unparsable, or scenario\-invalid field is counted as incorrect; the corresponding sample is not removed\. This implementation supports the four metrics defined in the main paper without selective filtering\.

Image\-text inputs contain the original text and paired image\. Video\-text inputs contain the original text and a temporally ordered sequence of sampled frames\. Speech is supplied as a transcript and processed jointly with the visual stream\. No result in the current manuscript should therefore be interpreted as evaluation of uninterrupted raw\-video or direct acoustic perception\.

### C\.2Details of Architecture

CoDAR is implemented as a train\-free sequence of instruction\-guided stages\. The stages pass structured intermediate representations rather than optimize a learned scoring function or a separate decoding objective\. Table[12](https://arxiv.org/html/2608.05825#A3.T12)records the minimum contract used to keep this process reproducible\.

Table 12:Stage\-level input, inference objective, and output contract of CoDAR\.
### C\.3Conflict\-Constrained Staged Inference

For instanceiiin scenarioss, letOsubj\(i\)O\_\{\\mathrm\{subj\}\}^\{\(i\)\}andOtgt\(i\)O\_\{\\mathrm\{tgt\}\}^\{\(i\)\}denote its Subject and Target candidate sets\. Letℳs\\mathcal\{M\}\_\{s\}andℒs\\mathcal\{L\}\_\{s\}denote the scenario\-valid Mechanism and Label sets\. The admissible output space is

Ωi,s=Osubj\(i\)×Otgt\(i\)×ℳs×ℒs\.\\Omega\_\{i,s\}=O\_\{\\mathrm\{subj\}\}^\{\(i\)\}\\times O\_\{\\mathrm\{tgt\}\}^\{\(i\)\}\\times\\mathcal\{M\}\_\{s\}\\times\\mathcal\{L\}\_\{s\}\.\(C\.1\)The model must return exactly one element from each component set:

𝒜i=\{ysubj,ytgt,ymechs,ylabels\}∈Ωi,s\.\\mathcal\{A\}\_\{i\}=\\\{y\_\{\\mathrm\{subj\}\},y\_\{\\mathrm\{tgt\}\},y\_\{\\mathrm\{mech\}\}^\{s\},y\_\{\\mathrm\{label\}\}^\{s\}\\\}\\in\\Omega\_\{i,s\}\.\(C\.2\)
Inference proceeds in a fixed order\. Counterfactual cost asks what social cost would arise if the subject expressed the latent state directly\. Conflict utility asks how the observed deviation reduces that cost or advances an interactional goal\. Participant recovery identifies who deploys the strategy and toward whom it is directed\. Mechanism–Label assignment then selects the scenario\-valid explanatory bridge and latent outcome jointly\. This order constrains open\-ended verbal reasoning without claiming an explicit combinatorial search\.

The final stage evaluates

Φ​\(𝒜i,xi,ci,𝒞i\)∈\{0,1\}\.\\Phi\(\\mathcal\{A\}\_\{i\},x\_\{i\},c\_\{i\},\\mathcal\{C\}\_\{i\}\)\\in\\\{0,1\\\}\.\(C\.3\)A value of zero triggers a targeted return to the stage responsible for the failed evidence, context, relation, or mechanism–label link\. A value of one indicates output\-level consistency only; it does not establish that the model’s internal reasoning is causally faithful\.

Operationally, a failed check is localized through the stage dependency chain\. Unsupported observable evidence returns to Stage 1; an insufficient social premise returns to Stage 2; an unresolved expectation–observation conflict returns to Stage 3; and a scenario\-invalid or internally incompatible tuple returns to Stage 4\. Components that remain supported are retained, so revision targets the failed dependency rather than restarting the complete inference chain\.

![Refer to caption](https://arxiv.org/html/2608.05825v1/x11.png)Figure 11:A case in the Affection scenario\.![Refer to caption](https://arxiv.org/html/2608.05825v1/x12.png)Figure 12:A case in the Stance scenario\.

## Appendix DCase Study

The main paper provides one direct\-versus\-CoDAR comparison\. The two cases below illustrate complementary failure modes; they are descriptive and do not replace the aggregate evaluation\.

Affection case\.In Figure[11](https://arxiv.org/html/2608.05825#A3.F11), correct participant recovery is insufficient for a correct structured interpretation\. The baseline recognizes the participants but assigns a mechanism that does not explain the latent affect\. CoDAR instead anchors the conflict to the Affection\-specific output space\.

Stance case\.Figure[12](https://arxiv.org/html/2608.05825#A3.F12)separates positive surface presentation from relational positioning\. The visible smile and affirmative wording do not establish affiliation by themselves\. CoDAR instead interprets the expression as protective withdrawal under a social\-role constraint\.

Cross\-case interpretation\.Across both examples, the baseline identifies surface cues but maps them directly to an output\. CoDAR instead requires the selected Mechanism to explain both the observed conflict and the recovered Subject–Target relation\. Accordingly,Ωi,s\\Omega\_\{i,s\}removes scenario\-invalid components, whereasΦ\\Phichecks whether the remaining tuple is jointly supported\. The cases illustrate output\-level revision rather than internal causal faithfulness\.

Similar Articles

Context-Aware RL for Agentic and Multimodal LLMs

Hugging Face Daily Papers

Introduces ContextRL, a reinforcement learning approach that teaches LLMs to identify which context supports an answer, achieving gains on agentic and multimodal benchmarks.