FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

arXiv cs.CL Papers

Summary

FriendBench is a new benchmark for evaluating whether humans and multimodal LLMs can infer if two people are familiar or strangers from a 20-second video clip of an ice-breaker conversation. Results show the best models match human accuracy but differ in bias, and only humans benefit from richer visual behavior.

arXiv:2607.29602v1 Announce Type: new Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward "stranger"---a difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:37 AM

# Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Source: [https://arxiv.org/html/2607.29602](https://arxiv.org/html/2607.29602)
Jeffrey M\. Girard, Jason Z\. Zheng, Jacqueline R\. Vertino, Antony D’Avirro, Benjamin Peloquin Fluid Concepts Research Correspondence:[jeff@fluidconcepts\.ai](https://arxiv.org/html/2607.29602v1/mailto:[email protected])

###### Abstract

Reading a social situation often depends on behavior, not words alone\. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20\-second clip of a dyadic ice\-breaker conversation\. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer\. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads\. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward “stranger”—a difference in effective prior, not discrimination\. Richer channels help both unequally, and only humans gain from visible behavior on top of speech\. We release the stimuli, human ratings, and model predictions\.

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Jeffrey M\. Girard, Jason Z\. Zheng, Jacqueline R\. Vertino,Antony D’Avirro, Benjamin PeloquinFluid Concepts ResearchCorrespondence:[jeff@fluidconcepts\.ai](https://arxiv.org/html/2607.29602v1/mailto:[email protected])

## 1Introduction

Social intelligence, the capacity to make sense of other people and the situations they are in, is a central component of human cognition\(Adolphs,[2003](https://arxiv.org/html/2607.29602#bib.bib18); Frith and Frith,[2007](https://arxiv.org/html/2607.29602#bib.bib19)\)\. It is also an increasingly important requirement for AI systems, which now observe, mediate, and participate in human interaction in roles ranging from companion agents to meeting assistants and care monitors\(Mathuret al\.,[2024](https://arxiv.org/html/2607.29602#bib.bib20); Sapet al\.,[2019](https://arxiv.org/html/2607.29602#bib.bib21)\)\. Such systems must reason about the social situation people are in, and much of that reasoning rests on multimodal behavioral cues rather than explicit verbal content, such as how people coordinate, respond, and orient toward one another\(Tickle\-Degnen and Rosenthal,[1990](https://arxiv.org/html/2607.29602#bib.bib4)\)\. A system that reads only the semantic content of a conversation captures just part of what social understanding requires\.

We study one concrete instance of social intelligence: the capacity to recognize the relationship between two people from a brief sample of how they interact\. Specifically, we ask whether an observer can tell that two people are already familiar with one another rather than meeting as strangers, without being told and without relying on what the conversation is about\. This judgment is a clean probe of behavioral social perception for two reasons\. First, since we have ground\-truth relationship labels, its answer is a matter of fact rather than interpretation: much related work in social intelligence targets intentions, emotions, or beliefs\(Premack and Woodruff,[1978](https://arxiv.org/html/2607.29602#bib.bib22); Sapet al\.,[2019](https://arxiv.org/html/2607.29602#bib.bib21)\), whose correct answer is often contestable, whereas prior familiarity has an unambiguous ground truth \(i\.e\., two people either have met before or they have not\)\. Second, familiarity is often unstated in a short interaction, so it must be inferred from behavior\. The task therefore rules out succeeding by reading semantic content alone, and prior work shows that humans make accurate social judgments from brief samples of behavior—so\-called*thin slices*\(Ambady and Rosenthal,[1992](https://arxiv.org/html/2607.29602#bib.bib2)\)\.

We operationalize the task asFriendBench: binary classification, familiar versus strangers, from a 20\-second clip of a dyadic ice\-breaker conversation, evaluated separately across text, audio, and video modalities\. Every dyad, whether familiar or strangers, responds to the same type of ice\-breaker prompt, so the conversational topic is matched across the two classes and cannot by itself reveal the answer, leaving themannerin which the two people interact as the primary signal\.

In addition to the core capability we benchmark, we examine how humans and models draw on the three modalities, and how each reaches its accuracy—by telling the two classes apart \(*discrimination*\) or by answering one class more often regardless of the evidence \(*response bias*\)\. Two patterns stand out\. First, richer channels help both models and humans alike, but unequally: audio adds reliable signal over text for each, yet only humans gain a further reliable increment from the visual channel, while the strongest models are flat from audio to audiovisual—they under\-exploit the visible behavior humans read\. Second, humans stay close to balanced across both classes in every modality, whereas models show a larger and more idiosyncratic class bias, most clearly in the text\-only condition\. Both patterns concern*how*humans and models solve the task, not merely whether they succeed \(§[5](https://arxiv.org/html/2607.29602#S5)\)\.

#### Contributions\.

- •We introduceFriendBench, a multimodal benchmark for inferring dyads’ prior familiarity \(familiar vs\. strangers\) from ice\-breaker clips in which every dyad answers the same type of prompt, built on the Seamless Interaction dataset\(Agrawalet al\.,[2025](https://arxiv.org/html/2607.29602#bib.bib1)\)\.111The benchmark stimuli, human ratings, and model predictions are openly available at[https://huggingface\.co/datasets/fluid\-concepts/friend\-bench](https://huggingface.co/datasets/fluid-concepts/friend-bench)\.
- •We release a matched human\-rater dataset for the task, covering 96 dyads in text, audio, and video, with roughly 90 raters per modality\.
- •We evaluate 26 models from seven companies—OpenAI, Google, Anthropic, Alibaba, Mistral, Thinking Machines, and Meta—spanning proprietary and open\-weight systems, across all three modalities\.
- •We evaluate humans and models on the same stimuli under matched conditions\. On accuracy, the best model and the human crowd are statistically indistinguishable in every modality\. But equal accuracy is not human\-like perception\. The strongest models reach it by leaning toward “stranger,” while human raters stay balanced\. Using signal detection theory, we characterize this as a difference in effective prior, not discrimination \(§[5](https://arxiv.org/html/2607.29602#S5)\)\.
- •We find a modality ordering both rater types obey only in part: text carries little signal for either and audio adds reliable discrimination, but the audiovisual channel gives*humans*a further reliable gain while leaving the strongest models flat—current models capture the vocal signal but under\-exploit the visible behavior human observers read \(§[5\.3](https://arxiv.org/html/2607.29602#S5.SS3)\)\.

## 2Related Work

#### Thin\-slice perception of familiarity\.

Human observers form accurate judgments about people and relationships from very brief behavioral exposure\. Ambady and Rosenthal’s\([1992](https://arxiv.org/html/2607.29602#bib.bib2)\)meta\-analysis found that judgments from observations under five minutes \(and often under 30 seconds\) predicted objective outcomes atr≈\.39r\\approx\.39, with longer exposure adding little\. Familiarity in particular is legible in thin slices, with friendship the best\-studied case: observers tell friends from strangers from silent video\(Latifet al\.,[2014](https://arxiv.org/html/2607.29602#bib.bib30)\)or brief audio\(Bryantet al\.,[2020](https://arxiv.org/html/2607.29602#bib.bib23)\), and cues such as inter\-turn timing distinguish them\(Templetonet al\.,[2023](https://arxiv.org/html/2607.29602#bib.bib34)\)\. Critically for our design,Dunbaret al\.\([2022](https://arxiv.org/html/2607.29602#bib.bib25)\)show that relationship quality remains inferable from speech even after its lexical content is digitally removed—the relational signal need not come from what is said\. This motivates both our task and our 20\-second window, which is ample by this evidence\.

#### Recognizing relationships from behavior\.

One line of work predicts relationship*type*from images or video via supervised classification: PISC\(Liet al\.,[2017](https://arxiv.org/html/2607.29602#bib.bib5)\)labels images as intimate, non\-intimate, or no\-relation, PIPA\(Sunet al\.,[2017](https://arxiv.org/html/2607.29602#bib.bib7)\)annotates sixteen fine\-grained relations in photo albums, and more recent work classifies asymmetric relations from the temporal dynamics of a live interaction\(Tanget al\.,[2026](https://arxiv.org/html/2607.29602#bib.bib33)\)\. A parallel line infers relationships from the*semantic content*of dialogue—acoustic\-lexical classifiers over phone calls\(Katerenchuket al\.,[2014](https://arxiv.org/html/2607.29602#bib.bib27)\)and language models over movie\-script dialogue, from relation\-classification datasets such as DDRel\(Jiaet al\.,[2021](https://arxiv.org/html/2607.29602#bib.bib26)\)to LLM evaluations where GPT\-4o infers speaker relationships well above chance\(Kimet al\.,[2026](https://arxiv.org/html/2607.29602#bib.bib28)\)\. Both differ from our task in two ways: they infer relationship type rather than the presence or absence of prior familiarity, and they lean on appearance, scene, or semantic content \(especially for scripted dialogue\)\. We instead fix the conversational prompt, controlling the semantic content these methods lean on\.

#### Social reasoning in multimodal models\.

Multimodal models are increasingly tested for social understanding: Social\-IQ\(Zadehet al\.,[2019](https://arxiv.org/html/2607.29602#bib.bib6)\)poses questions about social videos, while SIV\-Bench\(Konget al\.,[2026](https://arxiv.org/html/2607.29602#bib.bib29)\)and PIVOTSBench\(Zhanget al\.,[2026](https://arxiv.org/html/2607.29602#bib.bib35)\)probe reasoning about social scenes and fine\-grained relations\. HumanSense\(Qinet al\.,[2026](https://arxiv.org/html/2607.29602#bib.bib32)\)is the closest to our setting; one of its subtasks asks a model to judge how well two people in a video know each other\. Three things set our benchmark apart: \(1\) every dyad answers the same type of prompt, so the topic itself cannot give the answer away; \(2\) the label is objective, recording whether a pair had actually met before rather than how close a viewer judges them; and \(3\) we gather matched human ratings in text, audio, and video for direct comparison\. Our stimuli come from the Seamless Interaction corpus\(Agrawalet al\.,[2025](https://arxiv.org/html/2607.29602#bib.bib1)\)\. Other dyadic corpora record acquaintance too—UDIVA\(Palmeroet al\.,[2021](https://arxiv.org/html/2607.29602#bib.bib31)\)labels each pair known or unknown, NoXi\(Cafaroet al\.,[2017](https://arxiv.org/html/2607.29602#bib.bib24)\)rates how well partners know each other—but treat it as metadata, not as a label for prediction\.

## 3Methods

### 3\.1Benchmark Construction

Samples are drawn from the Seamless Interaction dataset\(Agrawalet al\.,[2025](https://arxiv.org/html/2607.29602#bib.bib1)\), restricted to naturalistic interactions from three recording sites \(the dataset’s*vendor*field\), each contributing a comparable share of dyads\. We exclude the dataset’s improvised interactions, in which participants are assigned a relationship and asked to act it out, since our aim is to measure this task on real, unscripted behavior rather than acted behavior\. Because the benchmark is used only for zero\-shot evaluation and never for training, we pool dyads across all of the dataset’s predefined splits rather than restricting to one, maximizing the pool from which our quality filters and stratified design \(below\) can draw\.

Seamless Interaction sessions include many different interaction types; we use only the Either\-Or \(EO\) ice\-breaker, a short task in which one participant poses a forced\-choice hypothetical question \(e\.g\., “would you rather have the ability to fly or be invisible?”\) and the dyad discusses their answer\.

We chose EO for three reasons: it standardizes the conversational context, since familiar and stranger pairs do the same thing and any signal must come from how a dyad responds rather than what it discusses; it is always the first interaction in a session, so every dyad is sampled from the same point in their interaction history; and its content carries little direct information about relationship status, so the task cannot be solved from semantic content alone\. In free conversation, by contrast, status can leak through content—shared history and inside knowledge for familiars, getting\-acquainted basics for strangers—cues a shared hypothetical question largely suppresses\.

An EO interaction, though centered on one either\-or question, typically continues for several minutes\. We sample one 20\-second clip per dyad from this interaction, drawn from either its early or late portion \(counterbalanced across relationship and recording site\), with boundaries snapped \(±\\pm5s\) to the nearest turn start so clips do not begin or end mid\-utterance\. The rendered video stimulus retains the clip’s audio track, so the video condition is audiovisual—visible behavior together with speech—whereas the audio and text conditions each isolate a single channel; both human raters and video models therefore receive sound as well as picture in the video condition\.

The final set fully crosses recording site, relationship \(familiar/strangers\), and gender composition \(same\-gender/mixed\-gender pair\), with 8 dyads per cell \(96 total\), and is participant\-disjoint: no individual appears in more than one dyad\. Crossing rules out recording site and gender composition as confounds \(§[4](https://arxiv.org/html/2607.29602#S4)\)\. Disjointness keeps each dyad an independent observation and guards against recognition leakage: because human raters see multiple stimuli in one session, a shared individual could be recognized from an earlier item, letting that recognition serve as the cue rather than the relationship itself\.222This is not a risk for our models, which are evaluated zero\-shot with no training process in which to learn identity\.

Interaction\- and clip\-level filters \(minimum turns, in\-window speaker balance, audio\-track synchrony, interaction length, prompt completeness; thresholds in Appendix[A](https://arxiv.org/html/2607.29602#A1), Table[A1](https://arxiv.org/html/2607.29602#A1.T1)\) exclude clips that cannot support the task regardless of relationship, e\.g\., one participant speaking for only a few seconds or desynchronized audio tracks\. An additional LLM\-based audit checks that each interaction’s recorded prompt label matches its content\.

### 3\.2Human Ratings

To establish baseline human performance, we collected human ratings of relationship type in a crowdsourced study\. For each stimulus, raters made a forced choice: familiar or strangers\. Before rating, each rater read a short instruction screen\. It explained that they would read, listen to, or watch \(depending on modality\) a clip from a real conversation in which the two people were playing the “Either Or” ice\-breaker game, then judge whether the pair had met before \(*familiar*\) or were meeting for the first time \(*strangers*\)\. The full instruction text is reproduced in Appendix[C](https://arxiv.org/html/2607.29602#A3)\.

Raters were recruited via Prolific and completed the task on GORILLA\. We ran three experiments, one per modality \(text, audio, video\), each with a panel of roughly 90 raters \(94 text, 94 audio, 92 video\)\. Prolific prescreening and quota matching ensured a gender\-balanced pool of US nationals residing in the US, with no reported hearing difficulties or autism\-spectrum diagnosis, and excluded anyone who had already participated in one of our other modality panels \(see Appendix[D](https://arxiv.org/html/2607.29602#A4)\)\.

Raters followed a planned\-missing, 6\-block incomplete design\(Grahamet al\.,[2006](https://arxiv.org/html/2607.29602#bib.bib3)\), each rating one of six overlapping blocks and seeing 32 of the 96 dyads by design; individual stimuli were in turn rated by 25–35 raters\.This design—rather than one rating per stimulus, averaged—lets us model raters themselves, estimating each rater’s own accuracy and response bias via crossed random effects, which requires each rater to contribute enough trials to be more than a one\-shot data point\.

Each rater additionally completed the Social Information Processing subscale from the Tromsø Social Intelligence Scale\(TSIS; Silveraet al\.,[2001](https://arxiv.org/html/2607.29602#bib.bib14); Grieve and Mahar,[2013](https://arxiv.org/html/2607.29602#bib.bib15)\)and a custom post\-task self\-report of which cues \(visual, vocal, verbal, interactional\) they attended to, enabling individual\-differences analysis \(Appendix[E](https://arxiv.org/html/2607.29602#A5)\)\.

### 3\.3Model Predictions

We evaluate 26 models from seven companies across text, audio, and video \(Table[2](https://arxiv.org/html/2607.29602#S4.T2)\), a fully\-crossed design in which every model rates every stimulus; unlike the human panels, no incomplete\-block correction is needed for the model results\. Each model receives the same prompt and response parsing\. The model prompt is adapted directly from the human instruction text, so that both rater types are given the same task framing, the same “Either Or” description, and the same familiar/strangers definition; the only difference is the response mechanism—a forced\-choice button for humans, a JSON label with a confidence rating for models\. Models and prompts are further described in Appendix[B](https://arxiv.org/html/2607.29602#A2)and Appendix[C](https://arxiv.org/html/2607.29602#A3), respectively\.

A few audio models decline to answer on a share of trials \(returning no usable familiar/strangers label\) rather than committing to a forced choice\. Because a refusal is not a wrong answer, we report each model’s accuracy over its*answered*\(covered\) trials and give per\-model coverage alongside it \(Table[2](https://arxiv.org/html/2607.29602#S4.T2)\); scoring refusals as errors would understate a model for declining rather than for misjudging\. Coverage is near\-total except for three audio models \(§[4](https://arxiv.org/html/2607.29602#S4)\), and the top models in each modality answer every trial, so this choice does not affect the headline comparison\.

## 4Benchmark Results

Table[1](https://arxiv.org/html/2607.29602#S4.T1)summarizes, per modality, the average per\-rater human, the human crowd \(the per\-stimulus majority vote of all raters who saw that stimulus\), and the single best\-performing model—their accuracy and, for the class\-bias analysis in §[5\.1](https://arxiv.org/html/2607.29602#S5.SS1), their per\-class recall and signal\-detection statistics\(Stanislaw and Todorov,[1999](https://arxiv.org/html/2607.29602#bib.bib12)\)\. Table[2](https://arxiv.org/html/2607.29602#S4.T2)gives the complete per\-model breakdown with Wilson 95% confidence intervals\. Figure[1](https://arxiv.org/html/2607.29602#S4.F1)depicts the same comparison, plotting every model against the average individual human rater and the majority\-vote crowd in each modality\.

RaterAcc\.Fam\.Str\.d′d^\{\\prime\}cc*Text*Human \(indiv\.\)49\.954\.645\.30\.00−0\.12\-0\.12Human \(crowd\)51\.059\.642\.60\.05−0\.21\-0\.21gpt\_text\_mini56\.220\.891\.70\.541\.061\.06*Audio*Human \(indiv\.\)56\.3∗∗57\.255\.50\.32−0\.02\-0\.02Human \(crowd\)63\.5∗64\.662\.50\.68−0\.03\-0\.03gemini\_audio\_pro66\.7∗∗39\.693\.81\.210\.860\.86*Video*Human \(indiv\.\)60\.9∗∗∗58\.763\.20\.560\.060\.06Human \(crowd\)71\.9∗∗∗70\.274\.51\.160\.060\.06gemini\_video66\.7∗∗41\.791\.71\.120\.770\.77Table 1:Human raters \(average individual and majority\-vote crowd\) vs\. the best model per modality\. Acc\. is accuracy; Fam\. and Str\. per\-class recall \(all %\);d′d^\{\\prime\}andccas in Table[2](https://arxiv.org/html/2607.29602#S4.T2)\. The named model in each block is that modality’s best\-performing \(top\-accuracy\) system, scored over answered trials; human\-indiv\. accuracy is the per\-rater mean \(per\-model Wilson CIs in Table[2](https://arxiv.org/html/2607.29602#S4.T2)\)\./∗∗∗/∗⁣∗∗\{\}^\{\*\}/^\{\*\*\}/^\{\*\*\*\}: above chance at\.05/\.01/\.001\.05/\.01/\.001, by exact binomial vs\. 50% \(crowd, best model\) or95%/99%/99\.9%95\\%/99\\%/99\.9\\%GLMM posterior HDI excluding 50% \(individual rows\)\.![Refer to caption](https://arxiv.org/html/2607.29602v1/x1.png)Figure 1:Human vs\. model accuracy by modality\. Each rater point covers only the∼\\sim32 trials that rater saw, so the rater cloud is widened by sampling noise and its upper tail is not a band of expert raters \(§[5\.4](https://arxiv.org/html/2607.29602#S5.SS4)\)\. Horizontal bars are Wilson 95% CIs for the best model and crowd, not tests of human–model or cross\-modality differences \(§4, §5\.3\)\.The best model and the human crowd trade the point\-estimate lead across modalities\. The model edges the crowd in text \(by 5\.2 points\) and audio \(by 3\.2\), while the crowd leads in video \(by 5\.2\)\. Video is humans’ strongest modality, and the one place the crowd, not a model, comes out ahead\. None of these three gaps is statistically reliable, though\. On the 96 paired dyads, an exact McNemar test\(Dietterich,[1998](https://arxiv.org/html/2607.29602#bib.bib36); McNemar,[1947](https://arxiv.org/html/2607.29602#bib.bib37)\)never approaches significance \(p\>\.4p\>\.4in every modality\)\. Text is the weakest modality for humans: at chance individually and barely above it as a crowd\. But the best text model is also not above chance \(56\.2%,p=\.26p=\.26\)\. Text is therefore less a model win than a modality where neither rater type finds a signal\.

ModelAcc\.95% CICov\.Fam\.Str\.d′d^\{\\prime\}cc*Text*![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_text\_mini56\.2\[46\.3, 65\.7\]10020\.891\.70\.541\.061\.06![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_text\_thinking55\.8\[45\.8, 65\.4\]9927\.783\.30\.360\.760\.76![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_text55\.2\[45\.3, 64\.8\]10025\.085\.40\.360\.840\.84![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_text53\.1\[43\.2, 62\.8\]10029\.277\.10\.190\.630\.63![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/thinking-machines-small.png)inkling\_text52\.1\[42\.2, 61\.8\]10016\.787\.50\.171\.031\.03![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/anthropic.png)claude\_text52\.1\[42\.2, 61\.8\]10014\.689\.60\.191\.121\.12![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_text\_high51\.0\[41\.2, 60\.8\]1008\.393\.80\.141\.401\.40![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/mistral.png)mistral\_text50\.0\[40\.2, 59\.8\]10087\.512\.50\.00−1\.11\-1\.11![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_text49\.0\[39\.2, 58\.8\]10010\.487\.5−0\.10\-0\.101\.161\.16*Audio*![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_audio\_pro66\.7∗∗\[56\.8, 75\.3\]10039\.693\.81\.210\.860\.86![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_audio63\.5∗\[53\.6, 72\.5\]10081\.245\.80\.76−0\.48\-0\.48![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_audio57\.3\[47\.3, 66\.7\]10020\.893\.80\.671\.131\.13![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_audio\_thinking55\.2\[45\.3, 64\.8\]10064\.645\.80\.26−0\.23\-0\.23![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_audio\_high53\.1\[43\.2, 62\.8\]10018\.887\.50\.250\.990\.99![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_audio\_1\_553\.1\[43\.2, 62\.8\]10087\.518\.80\.25−0\.99\-0\.99![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/qwen.png)qwen\_audio51\.5\[39\.8, 62\.9\]710\.0100\.00\.022\.192\.19![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/thinking-machines-small.png)inkling\_audio51\.0\[41\.2, 60\.8\]10041\.760\.40\.050\.230\.23![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/mistral.png)voxtral\_audio50\.0\[40\.2, 59\.8\]100100\.00\.00\.00−2\.32\-2\.32![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_audio48\.8\[34\.6, 63\.2\]45100\.04\.30\.45−1\.76\-1\.76![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_audio\_mini48\.1\[30\.7, 66\.0\]2825\.081\.80\.180\.720\.72![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/qwen.png)qwen\_omni47\.9\[38\.2, 57\.8\]10016\.779\.2−0\.15\-0\.150\.870\.87*Video*![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_video66\.7∗∗\[56\.8, 75\.3\]10041\.791\.71\.120\.770\.77![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_video\_thinking58\.3\[48\.3, 67\.7\]10027\.189\.60\.620\.910\.91![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_video\_pro57\.3\[47\.3, 66\.7\]10018\.895\.80\.771\.251\.25![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_video54\.2\[44\.2, 63\.8\]10014\.693\.80\.441\.241\.24![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_video\_high53\.1\[43\.2, 62\.8\]10010\.495\.80\.421\.421\.42Table 2:Per\-model results, sorted within modality\. Acc\. is accuracy \(%\) over answered trials and CI its Wilson 95% confidence interval; Cov\. is coverage \(% answered rather than declined\)\. Fam\. and Str\. are per\-class recall \(%\): share of familiar and of stranger dyads correctly labeled\.d′d^\{\\prime\}and criterioncctreat “familiar” as the signal class \(log\-linear corrected;Hautus,[1995](https://arxiv.org/html/2607.29602#bib.bib8)\), soc\>0c\>0indicates a bias toward “stranger” andc<0c<0toward “familiar\.”∗∗/∗: accuracy above chance atp<\.01p<\.01/p<\.05p<\.05\(exact binomial vs\. 50%\)\.Only three models are individually above chance at this sample size, all from Google:gemini\_audio\_proandgemini\_audioin audio, andgemini\_videoin video\. No text model clears chance\. Apart from the two top\-accuracy systems, the best remaining model in each modality reaches only 56\.2% \(text,gpt\_text\_mini\), 63\.5% \(audio,gemini\_audio\), and 58\.3% \(video,gemini\_video\_thinking\)\. Most models cluster near chance on raw accuracy\. The per\-class recall and criterion columns already hint at the analysis in §[5\.1](https://arxiv.org/html/2607.29602#S5.SS1): several models carry very large response biases regardless of accuracy\. In audio the bias is large but inconsistent in direction\.voxtral\_audiolabels every dyad “familiar” \(100%/0%,c=−2\.32c=\-2\.32\), whileqwen\_audiodoes the opposite \(0%/100%,c=2\.19c=2\.19\)\. In text most models lean hard toward “stranger,” withmuse\_spark\_text\_highrecovering only 8\.3% of familiar pairs \(c=1\.40c=1\.40\)\. A near\-chance accuracy can thus conceal a strongly one\-sided response pattern rather than balanced guessing\. Three audio models also decline a substantial share of trials:gpt\_audioanswers only 45%,gpt\_audio\_mini28%, andqwen\_audio71%\. They sit at chance on the trials they do answer, so their behavior is better summarized as declining\-plus\-guessing\.

## 5Performance Analysis

### 5\.1Class Bias: Humans vs\. Models

Humans and the best\-performing models are close to parity on accuracy \(§[4](https://arxiv.org/html/2607.29602#S4)\), but do they reach it the same way? To separate*discrimination*from*response bias*, we read the per\-class recall and signal\-detection columns of Table[1](https://arxiv.org/html/2607.29602#S4.T1), treating “familiar” as the signal class:d′d^\{\\prime\}measures how well a rater tells the classes apart, and the criterioncchow far its decision rule leans toward one answer \(c\>0c\>0toward “stranger,”c<0c<0toward “familiar”\)\. Two patterns stand out\. Humans stay near\-balanced everywhere: the human criterion is within0\.230\.23of zero in every modality, and per\-class recall gaps never exceed 10 points for pooled raters\. The best text and audio models instead reach their accuracy largely through response bias:gemini\_audio\_prolabels 93\.8% of strangers correctly but recovers only 39\.6% of familiar pairs \(c=0\.86c=0\.86\), and the best text model,gpt\_text\_mini, is more extreme still \(91\.7% vs\. 20\.8%,c=1\.06c=1\.06\)\. Across the roster this is the rule: in text almost every model leans “stranger” by 50–85 recall points, and in audio the bias is large but inconsistent in direction \(Table[2](https://arxiv.org/html/2607.29602#S4.T2)\)\.

The gap holds even in video, where models come closest to humans\. There the best model \(gemini\_video\) still leans “stranger” \(c=0\.77c=0\.77; 91\.7% of strangers but only 41\.7% of familiar pairs\), while the human crowd stays balanced \(c=0\.06c=0\.06\) at slightly higher discrimination \(d′=1\.16d^\{\\prime\}=1\.16vs\.1\.121\.12\)\. The two reach comparabled′d^\{\\prime\}, but only the crowd does so without skewing its decision rule\. Models thus fail to recreate the human balanced\-recall pattern in*any*modality, including their strongest: several of the highest\-accuracy systems \(Table[2](https://arxiv.org/html/2607.29602#S4.T2)\) buy that accuracy with a skewed criterion rather than sharper discrimination\.

### 5\.2Wisdom of the Crowd

Pooling human raters via majority vote barely moves accuracy in text \(49\.9%→\\to51\.0%,\+\+1\.1 points\) but raises it substantially in audio \(56\.3%→\\to63\.5%,\+\+7\.2 points\) and video \(60\.9%→\\to71\.9%,\+\+11\.0 points\)\. Figure[2](https://arxiv.org/html/2607.29602#S5.F2)shows the full crowd\-growth curves: majority\-vote accuracy rises with crowd sizekkin audio and video but stays flat near chance in text\. The audio and video gains therefore reflect real signal, not an artifact of pooling\.

![Refer to caption](https://arxiv.org/html/2607.29602v1/x2.png)Figure 2:Wisdom of the crowd\. Majority\-vote accuracy as a function of human crowd sizekk, with raters sampled per stimulus from that stimulus’s own rater panel \(mean±\\pmSD over 300 bootstrap draws\)\. Legend gives single\-rater→\\tofull\-panel crowd accuracy per modality\.
### 5\.3Modality Comparison

Text is the weakest modality for both humans and models, and richer channels help both—though, as we show below, not to the same degree\. To test the modality ordering while accounting for the repeated\-measures structure of the human data, we fit a Bayesian crossed random\-effects logistic model\(Gelmanet al\.,[2014](https://arxiv.org/html/2607.29602#bib.bib10)\)to per\-trial human correctness \(Bernoulli likelihood, logit link\), with a fixed effect of modality and crossed random intercepts for rater and for dyad \(§[3\.2](https://arxiv.org/html/2607.29602#S3.SS2)\); this is the design the block structure was built to support\. We estimate it with thebambiinterface to PyMC\(Caprettoet al\.,[2022](https://arxiv.org/html/2607.29602#bib.bib9)\), using its default weakly\-informative, data\-scaled priors \(wide normal priors on the fixed effects and half\-normal hyperpriors on the random\-effect standard deviations\), and sample the posterior with NUTS \(4 chains, 1,000 warmup and 1,000 post\-warmup draws each; allR^≈1\.00\\hat\{R\}\\approx 1\.00\)\. We report population\-averaged \(marginal\) accuracies and contrasts as posterior means with 95% highest\-density intervals \(HDIs,Makowskiet al\.,[2019](https://arxiv.org/html/2607.29602#bib.bib11)\)\. The marginal accuracies confirm the raw pattern—text50\.0%50\.0\\%\(95% HDI\[46\.7,53\.6\]\[46\.7,53\.6\]\), audio56\.6%56\.6\\%\[52\.2,60\.6\]\[52\.2,60\.6\], video60\.9%60\.9\\%\[56\.9,64\.9\]\[56\.9,64\.9\]—and the pairwise contrasts are decisive: audio exceeds text by6\.56\.5points \(HDI\[4\.1,8\.9\]\[4\.1,8\.9\]\), video exceeds audio by4\.54\.5points \(\[1\.9,6\.8\]\[1\.9,6\.8\]\), and video exceeds text by10\.910\.9points \(\[8\.5,13\.3\]\[8\.5,13\.3\]\), each with posteriorP\>0P\>0of at least0\.9990\.999\. Text alone is statistically indistinguishable from chance \(posterior probability of exceeding 50% only0\.500\.50\)\. Because the video condition is audiovisual \(§[3\.1](https://arxiv.org/html/2607.29602#S3.SS1)\), its edge over audio reflects the*added*value of visible behavior on top of speech, not vision in isolation; video is best read as an audiovisual upper bound rather than a vision\-only channel\.

The model side matches the human floor but not the human ceiling\. The best model climbs from text \(56\.2%56\.2\\%\) to audio \(66\.7%66\.7\\%\) but gains nothing from the audiovisual channel \(66\.7%66\.7\\%\), and among models evaluated in both audio and video only Gemini Flash improves \(63\.5%→66\.7%63\.5\\%\\to 66\.7\\%\) while Gemini Pro \(66\.7%→57\.3%66\.7\\%\\to 57\.3\\%\) and Muse Spark \(57\.3%→54\.2%57\.3\\%\\to 54\.2\\%\) decline\. \(Cross\-modality model means rise monotonically—52\.752\.7,53\.953\.9,57\.9%57\.9\\%for text, audio, video—but the video roster is small and skewed toward stronger companies, so we read the best\-model and within\-model trajectories rather than the pooled mean\.\) Where humans reliably convert visible behavior into accuracy on top of speech, then, the strongest models do not: they capture the vocal signal but under\-exploit the visual channel, the one place a benchmark of*multimodal*social perception most expects a model to gain\.

That text is the hardest modality is consistent with the benchmark’s design rather than a defect of it\. The EO ice\-breaker was chosen precisely so that transcript content carries little direct information about relationship status \(§[3\.1](https://arxiv.org/html/2607.29602#S3.SS1)\); the weak transcript performance of both humans and models is the expected consequence of that choice\. It also helps explain why response bias is most visible in text: with little signal available, a rater’s responses reflect its bias more than the stimulus\.

### 5\.4Rater Individual Differences

Because each rater contributes many trials by design \(§[3\.2](https://arxiv.org/html/2607.29602#S3.SS2)\), we can estimate individual accuracy and relate it to rater traits\. Trait social intelligence \(TSIS\-PS scale; reliable in every panel,α=0\.88\\alpha=0\.88–0\.920\.92\) is essentially uncorrelated with accuracy in the two modalities where the task is doable: audior=0\.03r=0\.03\(p=\.78p=\.78\) and videor=−0\.12r=\-0\.12\(p=\.24p=\.24\)\. The only significant association is in text \(r=0\.29r=0\.29,p=\.004p=\.004\), the modality where average accuracy is at chance, so we read it with caution\. Self\-reported cue use is similarly flat: raters most often report attending to interactional cues \(rapport, responsiveness\), but cue\-use scores rarely predict accuracy \(Appendix[E](https://arxiv.org/html/2607.29602#A5)\)\. The individual\-rater cloud in Figure[1](https://arxiv.org/html/2607.29602#S4.F1)is correspondingly wide, with the best raters near 75% in every modality\. Its spread should not be read as a stable band of expert observers, though\. Each rater’s accuracy comes from only the∼\\sim32 trials they saw, so the upper tail is close to what sampling noise alone would produce\. The best model therefore sits within the human distribution, not above it, even as almost no individual beats the aggregated crowd \(1 of 92 in video\)\. The crossed random\-effects model makes the point directly: the dyad random\-intercept SD \(0\.670\.67–0\.950\.95across modalities, logit scale\) is six to seven times the rater SD \(0\.100\.10–0\.140\.14\)\. Who the rater is thus matters far less than which dyad they were rating\.

### 5\.5Dyad Difficulty

Treating the per\-dyad random intercepts from the crossed random\-effects model as Rasch\-style item\-easiness parameters\(Rasch,[1960](https://arxiv.org/html/2607.29602#bib.bib16); de Boeck and Wilson,[2004](https://arxiv.org/html/2607.29602#bib.bib17)\), we find that difficulty is largely a property of the conversation, not of the modality through which it is observed\. Estimated dyad easiness correlates positively across all modality pairs \(text–audior=0\.52r=0\.52, text–videor=0\.33r=0\.33, audio–videor=0\.61r=0\.61; allp<\.01p<\.01\): a dyad that is hard to read from one modality tends to be hard from the others as well\. A handful of dyads sit below chance in every modality, acting as systematic “lures” that most raters misread alike \(Appendix[F](https://arxiv.org/html/2607.29602#A6)\)\.

Humans and models tend to find the same dyads hard, though how strongly depends on how it is measured\. Correlating per\-dyad human\-crowd accuracy with per\-dyad accuracy pooled over all models yieldsr=0\.56r=0\.56\(audio\) andr=0\.34r=0\.34\(video\), both significant \(p<\.01p<\.01; Appendix Figure[A1](https://arxiv.org/html/2607.29602#A6.F1)\), but onlyr=0\.20r=0\.20in text \(p=\.046p=\.046\)\. This raw correlation understates the shared difficulty\. It blends two things: how hard a dyad is to read, and which class it belongs to\. Humans stay balanced while the models lean toward “stranger” \(§[5\.1](https://arxiv.org/html/2607.29602#S5.SS1)\), so the two rater types tend to miss different classes\. Isolating difficulty from this class split, by correlating within each true class, brings the agreement out clearly:r=0\.61r=0\.61\(text\),0\.580\.58\(audio\), and0\.440\.44\(video\), allp<\.001p<\.001\. So a substantial part of what makes a pair legible or illegible is a property of the interaction shared across observers, even though the two reach their answers by different rules\.

## 6Discussion

For a benchmark of multimodal social perception, the clearest result is about the modalities themselves\. Text carries little relational signal for anyone, by design: the shared ice\-breaker prompt suppresses lexical content \(§[3\.1](https://arxiv.org/html/2607.29602#S3.SS1)\)\. Richer channels help both humans and models, but asymmetrically\. Humans improve reliably at every step, from text to audio to audiovisual \(§[5\.3](https://arxiv.org/html/2607.29602#S5.SS3)\), gaining a credible increment from visible behavior on top of speech\. The strongest models capture the audio gain but not the visual one: their top accuracy is identical in audio and audiovisual \(66\.7%66\.7\\%\)\. The sharpest human–model difference is thus not whether models can read relationships but whether they exploit the visual channel as humans do\. This gap has applied stakes\. The companion agents, meeting assistants, and care monitors that motivate the task must read social situations from behavior, and the visible cues humans exploit on top of speech are exactly what current models miss\.

In terms of overall accuracy, the best models have drawn level with an aggregated human crowd\. In every modality the two are statistically indistinguishable\. Point estimates put the model slightly ahead in text and audio and the crowd slightly ahead in video \(Table[1](https://arxiv.org/html/2607.29602#S4.T1)\), but no gap survives a paired McNemar test \(p\>\.4p\>\.4throughout; §[4](https://arxiv.org/html/2607.29602#S4)\)\.

The difference is in how humans and models use their two answers\. Discrimination \(d′d^\{\\prime\}\) measures how well a rater tells the classes apart\. Response bias \(criterioncc; Table[1](https://arxiv.org/html/2607.29602#S4.T1)\) measures how far it leans toward one answer\. Human raters stay balanced across the two answers in every modality\. The strongest models instead lean toward “stranger,” giving that answer more often regardless of the pair\. Video makes this clearest\. There the best model tells the classes apart about as well as the crowd \(d′≈1\.1d^\{\\prime\}\\approx 1\.1\), but it reaches that accuracy by leaning toward “stranger” where the crowd stays balanced\. In signal\-detection terms, this is a difference in*effective prior*: the models behave as though strangers were the more common answer and demand more evidence before saying “familiar\.” Humans behave as though the classes were equally likely, which is true in our balanced set\. This describes the response distribution, not a mechanism, and the prior is inferred from behavior rather than verified\. A pure criterion shift is correctable: recentering to the known base rate would raise accuracy with no gain in discrimination\. So a biased model’s raw accuracy can understate its discrimination\.

At the same time, humans and models are not perceiving unrelated things\. Dyad difficulty is partially shared and partially transfers across modalities \(§[5\.5](https://arxiv.org/html/2607.29602#S5.SS5)\), so part of what makes a pair legible or illegible is a property of the interaction itself, not of the observer or channel\. The dissociation is therefore specific: the two agree on*which*dyads are hard but diverge on the*decision rule*they apply\. Methodologically, these results argue for evaluating social\-perception models with more than a single accuracy number: per\-class recall, signal\-detection statistics, and item\-level difficulty each revealed structure that accuracy alone obscured, and are cheap to report once the underlying predictions are available\.

Taken together, the strongest models now match an aggregated human crowd on accuracy in this task, but not on how they reach it\. They find the same conversations hard\. Yet they systematically lean toward “stranger,” where human raters stay balanced, and they do not successfully convert the visual channel into accuracy\. Whether models can be brought to read relationships as humans do—balanced across classes, and drawing on visible behavior rather than speech alone—is the open question this benchmark is built to track\.

## Limitations

This paper reports only the ice\-breaker \(first\-interaction\) task from the Seamless Interaction dataset\. Findings about modality strength and class bias may not generalize to conversations later in a relationship or session, or to different prompt structures\. All familiar subtypes \(friends, family, romantic partner, coworkers, familiar\_other\) are collapsed into a single “familiar” class, because the balanced design does not have enough dyads per subtype to power a multi\-class analysis\. A six\-class relationship\-type task is defined in the broader project but is out of scope here, so any within\-familiar heterogeneity is invisible to this binary framing\. Only two\-person interactions are studied, and relationship inference in larger groups may draw on cues not captured here\. The balanced evaluation set is also modest in size \(96 dyads, one clip each\), so per\-model accuracy intervals are wide and only the strongest models clear chance individually\. The class\-bias and difficulty patterns we emphasize are more robust than any single model’s rank\. Our human baseline is drawn entirely from US\-resident raters \(§[3\.2](https://arxiv.org/html/2607.29602#S3.SS2)\); relationship\-perception cues can be culturally specific, so this baseline may not represent human performance in other populations\. Finally, even the “naturalistic” subset was recorded in a fixed\-camera motion\-capture studio, with wired lapel microphones and a posed ice\-breaker prompt\. This is not truly in\-the\-wild interaction, so findings may not transfer to less controlled contexts such as phone video or casual settings\.

Our evaluation set is also balanced 50/50 between familiar and stranger dyads by construction, which makes accuracy a clean measure of discrimination but base\-rate\-specific\. A reader might ask whether the models’ stranger\-lean is not miscalibration but a well\-calibrated prior for a world in which two people recorded together are more often strangers, penalized only by our artificial balance\. This does not threaten our central claim\. That claim is the*difference*in criterion between humans and models on identical, matched stimuli\. Neither rater type was told the base rate, so this difference does not depend on the true base rate, and neither doesd′d^\{\\prime\}, which is base\-rate\-invariant\. It does mean we measure discrimination, not deployment calibration; we do not read the balanced accuracy as a deployment estimate\. The observed bias is in any case idiosyncratic in direction and size across models \(Table[2](https://arxiv.org/html/2607.29602#S4.T2)\), which is hard to square with calibration to any single real\-world base rate\. Measuring calibration against realistic base rates would usefully complement the discrimination\-focused evaluation we report here\.

Each model is also evaluated under a single prompt, adapted from the human instructions \(§[3\.3](https://arxiv.org/html/2607.29602#S3.SS3)\)\. Because model responses can be sensitive to prompt wording, the exact biases in Table[2](https://arxiv.org/html/2607.29602#S4.T2)may shift under other phrasings; a prompt\-robustness sweep is worthwhile future work\. The human–model comparison holds the task framing fixed for both rater types, so it is less exposed to this concern than any single model’s bias estimate read on its own\.

## Ethics Statement

#### Informed consent and compensation\.

Rater recruitment and prescreening are described in §[3\.2](https://arxiv.org/html/2607.29602#S3.SS2)and Appendix[D](https://arxiv.org/html/2607.29602#A4)\. Before viewing any stimulus, each rater read a description of the study and affirmatively agreed to three statements: that they had read and understood the information, that they were free to withdraw at any time—by closing the browser tab—without giving a reason, and that they agreed to take part\. Participation was voluntary and could be ended at any point without penalty\. Sessions took roughly 15–30 minutes depending on modality, and raters were compensated at an effective rate of approximately $10–12 per hour, above Prolific’s fair\-pay guidance\. The task was minimal\-risk: raters viewed short, benign clips of consenting conversation partners and answered non\-sensitive perceptual and self\-report questions\.

#### Privacy and data minimization\.

We collected no personally identifying information from raters\. Beyond the task responses and the self\-report questionnaires analyzed here, we recorded only the coarse attributes used for prescreening and quota matching \(§[3\.2](https://arxiv.org/html/2607.29602#S3.SS2)\), keyed to a pseudonymous Prolific identifier that we do not link to any real\-world identity\. All released rater data are de\-identified\.

#### Institutional review\.

This research was conducted outside a university setting and was not reviewed by an institutional review board\. Because it collected no personally identifying information from raters and involved only minimal\-risk procedures, it did not fall under IRB oversight; we nonetheless followed standard human\-subjects safeguards—voluntary informed consent, the right to withdraw without penalty, fair compensation, minimal\-risk stimuli, and data de\-identification\.

#### Stimulus source\.

All interaction clips are drawn from the publicly released Seamless Interaction dataset\(Agrawalet al\.,[2025](https://arxiv.org/html/2607.29602#bib.bib1)\), whose participants consented to the recording and research use of their audio and video\. We use these data in accordance with the dataset’s CC\-BY\-NC 4\.0 license \(attribution, non\-commercial use only\): the rendered benchmark clips are redistributed under the same CC\-BY\-NC 4\.0 terms, with attribution to the Seamless Interaction dataset\. We do not attempt to re\-identify or contact any recorded individual\.

#### Intended use and risks\.

The benchmark is intended for evaluating and auditing the social\-perceptual behavior of models, including the response biases we document \(§[5\.1](https://arxiv.org/html/2607.29602#S5.SS1)\)\. Inferring familiarity from behavior could in principle support surveillance or profiling; we release the benchmark to enable research on and scrutiny of such capabilities, not their deployment\. Given the modest accuracy and pronounced, idiosyncratic biases we observe \(§[4](https://arxiv.org/html/2607.29602#S4)\), we caution against using these models or this task to make consequential judgments about real individuals\.

## References

- Cognitive neuroscience of human social behaviour\.Nature Reviews Neuroscience4\(3\),pp\. 165–178\.External Links:ISSN 1471\-0048,[Document](https://dx.doi.org/10.1038/nrn1056)Cited by:[§1](https://arxiv.org/html/2607.29602#S1.p1.1)\.
- V\. Agrawal, A\. Akinyemi, K\. Alvero, M\. Behrooz, J\. Buffalini, F\. M\. Carlucci, J\. Chen, J\. Chen, Z\. Chen, S\. Cheng, P\. Chowdary, J\. Chuang, A\. D’Avirro, J\. Daly, N\. Dong, M\. Duppenthaler, C\. Gao, J\. Girard, M\. Gleize, S\. Gomez, H\. Gong, S\. Govindarajan, B\. Han, S\. He, D\. Hernandez, Y\. Hristov, R\. Huang, H\. Inaguma, S\. Jain, R\. Janardhan, Q\. Jia, C\. Klaiber, D\. Kovachev, M\. Kumar, H\. Li, Y\. Li, P\. Litvin, W\. Liu, G\. Ma, J\. Ma, M\. Ma, X\. Ma, L\. Mantovani, S\. Miglani, S\. Mohan, L\. Morency, E\. Ng, K\. Ng, T\. A\. Nguyen, A\. Oberai, B\. Peloquin, J\. Pino, J\. Popovic, O\. Poursaeed, F\. Prada, A\. Rakotoarison, R\. Ranjan, A\. Richard, C\. Ropers, S\. Saleem, V\. Sharma, A\. Shcherbyna, J\. Shen, J\. Shen, A\. Stathopoulos, A\. Sun, P\. Tomasello, T\. Tran, A\. Turkatenko, B\. Wan, C\. Wang, J\. Wang, M\. Williamson, C\. Wood, T\. Xiang, Y\. Yang, J\. Yao, C\. Zhang, J\. Zhang, X\. Zhang, J\. Zheng, P\. Zhyzheria, J\. Zikes, and M\. Zollhoefer \(2025\)Seamless interaction: Dyadic audiovisual motion modeling and large\-scale dataset\.arXiv:2506\.22554 \[cs\.CV\]\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.22554),https://arxiv\.org/abs/2506\.22554Cited by:[1st item](https://arxiv.org/html/2607.29602#S1.I1.i1.p1.1),[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2607.29602#S3.SS1.p1.1),[Stimulus source\.](https://arxiv.org/html/2607.29602#Sx2.SS0.SSS0.Px4.p1.1)\.
- N\. Ambady and R\. Rosenthal \(1992\)Thin slices of expressive behavior as predictors of interpersonal consequences: A meta\-analysis\.Psychological Bulletin111\(2\),pp\. 256–274\.External Links:ISSN 1939\-1455,[Document](https://dx.doi.org/10.1037/0033-2909.111.2.256)Cited by:[§1](https://arxiv.org/html/2607.29602#S1.p2.1),[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px1.p1.1)\.
- G\. A\. Bryant, C\. S\. Wang, and R\. Fusaroli \(2020\)Recognizing affiliation in colaughter and cospeech\.Royal Society Open Science7\(10\),pp\. 201092\.External Links:ISSN 2054\-5703,[Document](https://dx.doi.org/10.1098/rsos.201092)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Cafaro, J\. Wagner, T\. Baur, S\. Dermouche, M\. Torres Torres, C\. Pelachaud, E\. André, and M\. Valstar \(2017\)The NoXi database: multimodal recordings of mediated novice\-expert interactions\.InProceedings of the 19th ACM International Conference on Multimodal Interaction,ICMI ’17,New York, NY, USA,pp\. 350–359\.External Links:[Document](https://dx.doi.org/10.1145/3136755.3136780),ISBN 978\-1\-4503\-5543\-8Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px3.p1.1)\.
- T\. Capretto, C\. Piho, R\. Kumar, J\. Westfall, T\. Yarkoni, and O\. A\. Martin \(2022\)Bambi: A Simple Interface for Fitting Bayesian Linear Models in Python\.Journal of Statistical Software103,pp\. 1–29\.External Links:ISSN 1548\-7660,[Document](https://dx.doi.org/10.18637/jss.v103.i15)Cited by:[§5\.3](https://arxiv.org/html/2607.29602#S5.SS3.p1.16)\.
- P\. de Boeck and M\. Wilson \(Eds\.\) \(2004\)Explanatory Item Response Models: A Generalized Linear and Nonlinear Approach\.Springer Science & Business Media\.External Links:ISBN 978\-0\-387\-40275\-8Cited by:[§5\.5](https://arxiv.org/html/2607.29602#S5.SS5.p1.4)\.
- T\. G\. Dietterich \(1998\)Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms\.Neural Computation10\(7\),pp\. 1895–1923\.External Links:ISSN 0899\-7667,[Document](https://dx.doi.org/10.1162/089976698300017197)Cited by:[§4](https://arxiv.org/html/2607.29602#S4.p2.2)\.
- R\. I\. M\. Dunbar, J\. Robledo, I\. Tamarit, I\. Cross, and E\. Smith \(2022\)Nonverbal Auditory Cues Allow Relationship Quality to be Inferred During Conversations\.Journal of Nonverbal Behavior46\(1\),pp\. 1–18\.External Links:ISSN 1573\-3653,[Document](https://dx.doi.org/10.1007/s10919-021-00386-y)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px1.p1.1)\.
- C\. D\. Frith and U\. Frith \(2007\)Social Cognition in Humans\.Current Biology17\(16\),pp\. R724–R732\.External Links:ISSN 0960\-9822,[Document](https://dx.doi.org/10.1016/j.cub.2007.05.068)Cited by:[§1](https://arxiv.org/html/2607.29602#S1.p1.1)\.
- A\. Gelman, J\. B\. Carlin, H\. S\. Stern, D\. B\. Dunson, A\. Vehtari, and D\. B\. Rubin \(2014\)Bayesian data analysis\.3rd edition,CRC Press,Boca Raton, FL\.Cited by:[§5\.3](https://arxiv.org/html/2607.29602#S5.SS3.p1.16)\.
- J\. W\. Graham, B\. J\. Taylor, A\. E\. Olchowski, and P\. E\. Cumsille \(2006\)Planned missing data designs in psychological research\.Psychological Methods11\(4\),pp\. 323–343\.External Links:ISSN 1082\-989X,[Document](https://dx.doi.org/10.1037/1082-989x.11.4.323)Cited by:[Appendix D](https://arxiv.org/html/2607.29602#A4.p1.1),[§3\.2](https://arxiv.org/html/2607.29602#S3.SS2.p3.1)\.
- R\. Grieve and D\. Mahar \(2013\)Can social intelligence be measured? Psychometric properties of the Tromsø Social Intelligence Scale – English Version\.The Irish Journal of Psychology34\(1\),pp\. 1–12\.External Links:ISSN 0303\-3910,[Document](https://dx.doi.org/10.1080/03033910.2012.737758)Cited by:[§3\.2](https://arxiv.org/html/2607.29602#S3.SS2.p4.1)\.
- K\. L\. Gwet \(2021\)Handbook of inter\-rater reliability: Chance\-corrected agreement coefficients\.5th edition, Vol\.1,AgreeStat Analytics\.Cited by:[Appendix D](https://arxiv.org/html/2607.29602#A4.SS0.SSS0.Px1.p1.7)\.
- M\. J\. Hautus \(1995\)Corrections for extreme proportions and their biasing effects on estimated values of d′\\prime\.Behavior Research Methods, Instruments, & Computers27\(1\),pp\. 46–51\.External Links:ISSN 1532\-5970,[Document](https://dx.doi.org/10.3758/BF03203619)Cited by:[Table 2](https://arxiv.org/html/2607.29602#S4.T2)\.
- Q\. Jia, H\. Huang, and K\. Q\. Zhu \(2021\)DDRel: A New Dataset for Interpersonal Relation Classification in Dyadic Dialogues\.Proceedings of the AAAI Conference on Artificial Intelligence35\(14\),pp\. 13125–13133\.External Links:ISSN 2374\-3468,[Document](https://dx.doi.org/10.1609/aaai.v35i14.17551)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px2.p1.1)\.
- D\. Katerenchuk, D\. G\. Brizan, and A\. Rosenberg \(2014\)“Was that your mother on the phone?”: classifying interpersonal relationships between dialog participants with lexical and acoustic properties\.InProc\. Interspeech 2014,pp\. 1831–1835\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2014-416)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px2.p1.1)\.
- E\. Kim, J\. Park, J\. Oh, K\. Park, S\. Song, A\. S\. Doğruöz, A\. Oh, and N\. Kim \(2026\)Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 23431–23451\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1074),ISBN 979\-8\-89176\-390\-6Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px2.p1.1)\.
- F\. Kong, W\. Zu, X\. Chen, Y\. Yang, S\. Zhu, and X\. Feng \(2026\)SIV\-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 37379–37403\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1863),ISBN 979\-8\-89176\-395\-1Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px3.p1.1)\.
- N\. Latif, A\. V\. Barbosa, E\. Vatikiotis\-Bateson, M\. S\. Castelhano, and K\. G\. Munhall \(2014\)Movement Coordination during Conversation\.PLOS ONE9\(8\),pp\. e105036\.External Links:ISSN 1932\-6203,[Document](https://dx.doi.org/10.1371/journal.pone.0105036)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Li, Y\. Wong, Q\. Zhao, and M\. S\. Kankanhalli \(2017\)Dual\-Glance Model for Deciphering Social Relationships\.InProceedings of the IEEE International Conference on Computer Vision,pp\. 2650–2659\.Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px2.p1.1)\.
- D\. Makowski, M\. S\. Ben\-Shachar, S\. H\. A\. Chen, and D\. Lüdecke \(2019\)Indices of effect existence and significance in the Bayesian framework\.Frontiers in Psychology10\.External Links:ISSN 1664\-1078,[Document](https://dx.doi.org/10.3389/fpsyg.2019.02767)Cited by:[§5\.3](https://arxiv.org/html/2607.29602#S5.SS3.p1.16)\.
- L\. Mathur, P\. P\. Liang, and L\. Morency \(2024\)Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 20541–20560\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.1143)Cited by:[§1](https://arxiv.org/html/2607.29602#S1.p1.1)\.
- Q\. McNemar \(1947\)Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages\.Psychometrika12\(2\),pp\. 153–157\.External Links:ISSN 0033\-3123, 1860\-0980,[Document](https://dx.doi.org/10.1007/BF02295996)Cited by:[§4](https://arxiv.org/html/2607.29602#S4.p2.2)\.
- C\. Palmero, J\. Selva, S\. Smeureanu, J\. C\. S\. J\. Junior, A\. Clapes, A\. Mosegui, Z\. Zhang, D\. Gallardo, G\. Guilera, D\. Leiva, and S\. Escalera \(2021\)Context\-Aware Personality Inference in Dyadic Scenarios: Introducing the UDIVA Dataset\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp\. 1–12\.Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Premack and G\. Woodruff \(1978\)Does the chimpanzee have a theory of mind?\.Behavioral and Brain Sciences1\(4\),pp\. 515–526\.External Links:ISSN 1469\-1825, 0140\-525X,[Document](https://dx.doi.org/10.1017/S0140525X00076512)Cited by:[§1](https://arxiv.org/html/2607.29602#S1.p2.1)\.
- Z\. Qin, R\. Zheng, Y\. Wang, T\. Li, Y\. Yuan, J\. Chen, and L\. Wang \(2026\)HumanSense: From Multimodal Perception to Empathetic Context\-Aware Responses Through Reasoning MLLMs\.Proceedings of the AAAI Conference on Artificial Intelligence40\(30\),pp\. 24973–24981\.External Links:ISSN 2374\-3468,[Document](https://dx.doi.org/10.1609/aaai.v40i30.39685)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px3.p1.1)\.
- G\. Rasch \(1960\)Probabilistic Models for Some Intelligence and Attainment Tests\.MESA Press, 5835 S\.External Links:ISBN 978\-0\-941938\-05\-1Cited by:[§5\.5](https://arxiv.org/html/2607.29602#S5.SS5.p1.4)\.
- M\. Sap, H\. Rashkin, D\. Chen, R\. Le Bras, and Y\. Choi \(2019\)Social IQa: Commonsense Reasoning about Social Interactions\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 4463–4473\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1454)Cited by:[§1](https://arxiv.org/html/2607.29602#S1.p1.1),[§1](https://arxiv.org/html/2607.29602#S1.p2.1)\.
- D\. Silvera, M\. Martinussen, and T\. I\. Dahl \(2001\)The Tromsø Social Intelligence Scale, a self\-report measure of social intelligence\.Scandinavian Journal of Psychology42\(4\),pp\. 313–319\.External Links:ISSN 1467\-9450,[Document](https://dx.doi.org/10.1111/1467-9450.00242)Cited by:[§3\.2](https://arxiv.org/html/2607.29602#S3.SS2.p4.1)\.
- H\. Stanislaw and N\. Todorov \(1999\)Calculation of signal detection theory measures\.Behavior Research Methods, Instruments, & Computers31\(1\),pp\. 137–149\.External Links:ISSN 1532\-5970,[Document](https://dx.doi.org/10.3758/BF03207704)Cited by:[§4](https://arxiv.org/html/2607.29602#S4.p1.1)\.
- Q\. Sun, B\. Schiele, and M\. Fritz \(2017\)A Domain Based Approach to Social Relation Recognition\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 3481–3490\.Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px2.p1.1)\.
- W\. Tang, F\. I\. Dogan, L\. Qing, and H\. Gunes \(2026\)AsyReC: A Multimodal Graph\-Based Framework for Spatio\-Temporal Asymmetric Dyadic Relationship Classification\.IEEE Transactions on Circuits and Systems for Video Technology36\(3\),pp\. 3693–3708\.External Links:ISSN 1558\-2205,[Document](https://dx.doi.org/10.1109/TCSVT.2025.3616347)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px2.p1.1)\.
- E\. M\. Templeton, L\. J\. Chang, E\. A\. Reynolds, M\. D\. Cone LeBeaumont, and T\. Wheatley \(2023\)Long gaps between turns are awkward for strangers but not for friends\.Philosophical Transactions of the Royal Society B: Biological Sciences378\(1875\),pp\. 20210471\.External Links:ISSN 0962\-8436,[Document](https://dx.doi.org/10.1098/rstb.2021.0471)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px1.p1.1)\.
- L\. Tickle\-Degnen and R\. Rosenthal \(1990\)The nature of rapport and its nonverbal correlates\.Psychological Inquiry1\(4\),pp\. 285–293\.External Links:[Document](https://dx.doi.org/10.1207/s15327965pli0104%5F1)Cited by:[§1](https://arxiv.org/html/2607.29602#S1.p1.1)\.
- A\. Zadeh, M\. Chan, P\. P\. Liang, E\. Tong, and L\. Morency \(2019\)Social\-IQ: A Question Answering Benchmark for Artificial Social Intelligence\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 8807–8817\.Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Zhang, Y\. Yin, W\. Song, Y\. Wu, and M\. Liu \(2026\)PIVOTSBench: Evaluating Fine\-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models\.arXiv\.External Links:2606\.23092,[Document](https://dx.doi.org/10.48550/arXiv.2606.23092)Cited by:[§2](https://arxiv.org/html/2607.29602#S2.SS0.SSS0.Px3.p1.1)\.

## Appendix ABenchmark Quality Filters

Every candidate interaction and clip must pass the automatic quality filters in Table[A1](https://arxiv.org/html/2607.29602#A1.T1)before entering the stratified selection pool \(§[3\.1](https://arxiv.org/html/2607.29602#S3.SS1)\)\. Interaction\-level filters \(F1, F4, F5\) gate the whole source interaction; clip\-level filters \(F2, F3\) gate the specific 20\-second window sampled from it\. Clips are 20 s long with boundaries snapped within±5\\pm 5s to the nearest turn start\. These filters remove clips that cannot support the task regardless of relationship; they are followed by the LLM\-based prompt\-adherence audit \(§[3\.1](https://arxiv.org/html/2607.29602#S3.SS1)\), which verifies that each interaction’s recorded prompt label matches its actual content\.

Table A1:Automatic quality filters applied during clip generation and the EO interaction audit\. Interaction\-level filters gate the source interaction; clip\-level filters gate the sampled 20 s window\. Thresholds are the generator/auditor constants \(scripts/generate\_samples\_rapport\.py,scripts/audit\_eo\_interactions\.py\)\.
## Appendix BModel Inference Configuration

Table[A2](https://arxiv.org/html/2607.29602#A2.T2)lists the inference configuration for each of the 26 models\. All models received the same prompt construction and response parsing \(§[3\.3](https://arxiv.org/html/2607.29602#S3.SS3)\); the settings below cover per\-model decoding and, where applicable, reasoning parameters\. Cloud models were queried at temperature0with a fixed seed \(4242\) wherever the API exposed them; the three locally\-run open\-weight models \(Qwen2\-Audio, Qwen2\.5\-Omni, Voxtral\) used greedy decoding \(do\_sample=False\), which is deterministic without a seed\. Answer length was capped at 150–300 tokens for non\-reasoning models, while reasoning models received a larger combined answer\-plus\-reasoning budget \(8,000 tokens for the Gemini thinking and pro variants; 4,000 for Inkling and Muse Spark\) to avoid truncating the response\.

ModelAPI / model IDTemp\.SeedReasoning*Text*![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_textgpt\-4o04242—![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_text\_minigpt\-4o\-mini04242—![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_textgemini\-3\.5\-flash04242off![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_text\_thinkinggemini\-3\.5\-flash04242dynamic![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/anthropic.png)claude\_textclaude\-opus\-4\-8—a—aoffa![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/mistral.png)mistral\_textmistral\-large\-latest04242—![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/thinking-machines-small.png)inkling\_textthinkingmachines/Inkling042b42^\{b\}on![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_textmuse\-spark\-1\.1—c—cminimal![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_text\_highmuse\-spark\-1\.1—c—chigh*Audio*![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_audiogpt\-audio04242—![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_audio\_1\_5gpt\-audio\-1\.504242—![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/openai.png)gpt\_audio\_minigpt\-audio\-mini04242—![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_audiogemini\-3\.5\-flash04242off![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_audio\_progemini\-pro\-latest04242dynamicd![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_audio\_thinkinggemini\-3\.5\-flash04242dynamic![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/qwen.png)qwen\_audioQwen/Qwen2\-Audio\-7B\-Instructgreedy——![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/qwen.png)qwen\_omniQwen/Qwen2\.5\-Omni\-7Bgreedy——![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/mistral.png)voxtral\_audiomistralai/Voxtral\-Mini\-3B\-2507greedy——![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/thinking-machines-small.png)inkling\_audiothinkingmachines/Inkling042b42^\{b\}on![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_audiomuse\-spark\-1\.1—c—cminimal![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_audio\_highmuse\-spark\-1\.1—c—chigh*Video*![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_videogemini\-3\.5\-flash04242off![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_video\_progemini\-pro\-latest04242dynamicd![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/google.png)gemini\_video\_thinkinggemini\-3\.5\-flash04242dynamic![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_videomuse\-spark\-1\.1—c—cminimal![[Uncaptioned image]](https://arxiv.org/html/2607.29602v1/figures/meta.png)muse\_spark\_video\_highmuse\-spark\-1\.1—c—chighTable A2:Per\-model inference configuration \(26 models, seven companies; roster matches Table[2](https://arxiv.org/html/2607.29602#S4.T2)\)\.*Temp\.*and*Seed*are the decoding temperature and random seed passed to the API;greedymarks the locally\-run open\-weight models decoded withdo\_sample=False\(deterministic, no seed\)\.*Reasoning*: “—” = non\-reasoning model; “off” = reasoning\-capable but disabled \(Geminithinking\_budget=0\{=\}0\); “dynamic” = model sets its own reasoning depth \(thinking\_budget=−1\{=\}\{\-\}1\); “on” = reasoning always on with no depth control; “minimal”/“high” = named reasoning\-effort tier\. Two Gemini IDs are aliases; at run time \(July 2026\)gemini\-pro\-latestresolved togemini\-3\.1\-pro\-preview, andgemini\-3\.5\-flashwas pinned explicitly rather thangemini\-flash\-latest\(which then resolved togemini\-3\.6\-flash\)\.aClaude rejectstemperature/top\_pand exposes no seed; extended thinking was not enabled\.bInkling accepts a seed but the API echoednull, so determinism is best\-effort\.cMuse Spark exposes no temperature or seed control, and its reasoning depth is non\-deterministic run\-to\-run\.dgemini\-pro\-latestcannot disable thinking and always reasons dynamically\.
## Appendix CTask Instructions and Prompts

Both human raters and models were given the same task framing, “Either Or” game description, and familiar/strangers definition, differing only in how a response was collected\. Human raters selected an answer with a button; models were additionally instructed to emit a JSON object with a label and confidence\. Modality\-specific slots \(shown here in brackets, Text/Audio/Video order\) were filled per panel or per clip\.

#### Human rater instructions\.

```
You will [read / listen to / watch] a series
of [transcripts / audio clips / video clips]
taken from real conversations between two
people.

These people are playing a game called
"Either Or," where they discuss whether they
would prefer one thing -- for example, the
ability to fly -- or an alternative, such as
the ability to breathe underwater.

Your job is to pay close attention to each
conversation and determine whether the two
people have met before -- meaning they are
familiar, such as friends, family members,
coworkers, or romantic partners -- or are
meeting for the first time in this
interaction -- meaning they are strangers.
```

Raters then made a forced choice:*familiar*or*strangers*\.

#### Model prompt\.

```
[Read / Listen to / Watch] the following
[transcript / audio clip / video clip], taken
from a real conversation between two people.

These people are playing a game called
"Either Or," where they discuss whether they
would prefer one thing -- for example, the
ability to fly -- or an alternative, such as
the ability to breathe underwater.

Your job is to pay close attention to the
conversation and determine whether the two
people have met before -- meaning they are
familiar, such as friends, family members,
coworkers, or romantic partners -- or are
meeting for the first time in this
interaction -- meaning they are strangers.

For each clip, return:
  {
    "relationship_label": str {FAMILIAR|STRANGER},
    "confidence": float [0, 1],
    "reason": str
  }

Use the following confidence scale:
  0 = Not at all confident, ...,
  0.5 = Moderately confident, ...,
  1 = Extremely confident

Do not provide extra text outside the JSON
object.
```

## Appendix DHuman Rating Study Details

Raters were recruited on Prolific with platform\-level quality screening in addition to the demographic criteria in §[3\.2](https://arxiv.org/html/2607.29602#S3.SS2): a Prolific approval rate of 99–100% and at least 10 prior submissions\. Under the planned\-missing, 6\-block design\(Grahamet al\.,[2006](https://arxiv.org/html/2607.29602#bib.bib3)\), each rater was assigned to a single block and each of the 96 dyads appeared in two of the six blocks, so every dyad was rated by an overlapping subset of raters \(25–35 per dyad\)\. By design each rater judged 32 dyads; in the audio panel a platform issue left 13 of the 94 raters with 29–31 completed trials, while all text and video raters completed the full 32\.

#### Inter\-rater agreement\.

Agreement among human raters is slight, as expected for a difficult perceptual task with a balanced label set\. Fleiss’sκ\\kappa\(computed per stimulus over whichever raters saw it, accommodating the incomplete\-block design\) is0\.080\.08for text and0\.170\.17for both audio and video\. Bennett’sSSis essentially identical \(0\.090\.09,0\.170\.17,0\.170\.17for text, audio, video\), as expected when the two classes are balanced\(Gwet,[2021](https://arxiv.org/html/2607.29602#bib.bib13)\)\. The near\-zero text agreement is consistent with individual text raters performing at chance \(Table[1](https://arxiv.org/html/2607.29602#S4.T1)\): raters are not converging on a shared, reliable text cue\.

## Appendix ERater Individual Differences

Table[A3](https://arxiv.org/html/2607.29602#A5.T3)reports, per modality, the reliability of the TSIS\-PS social\-intelligence scale and its correlation with per\-rater accuracy\. The scale is highly reliable in every panel, but its correlation with accuracy is null in audio and video and positive only in text—the modality where accuracy is at chance—so we do not interpret it as evidence of a general skill advantage\.

Table A3:TSIS\-PS reliability and its correlation with per\-rater accuracy \(N=94N=94text, 94 audio, 92 video\)\.For self\-reported cue use, raters most often reported attending to interactional cues \(rapport, responsiveness\) across all modalities, but individual cue\-attendance scores rarely predicted accuracy \(only a handful of per\-cue correlations reachedp<\.05p<\.05, with no stable cross\-modality pattern\)\. Full cue\-use heatmaps are provided with the released analysis notebooks\.

The per\-rater accuracy clouds in Figure[1](https://arxiv.org/html/2607.29602#S4.F1)show the full distribution of per\-rater accuracy in each modality \(94 text, 94 audio, 92 video raters; each rater’s accuracy over the 32 dyads in their assigned block, 29–32 for a few audio raters affected by a platform issue\)\. The clouds are wide, but much of that width is sampling noise: with only∼\\sim32 trials per rater, binomial variation around a common ability already reproduces most of the observed spread—including the upper tail that appears to exceed the best model \(§[5\.4](https://arxiv.org/html/2607.29602#S5.SS4)\)—so the distribution should not be read as a stable ordering of raters by skill\. This is also why per\-rater accuracy is best modeled with partially\-pooled crossed random effects \(§[3\.2](https://arxiv.org/html/2607.29602#S3.SS2)\), which shrink noisy individual estimates, rather than trusting raw per\-rater rates or treating each rater as a one\-shot observation\. Text raters cluster tightly around chance, consistent with the near\-zero text agreement reported in Appendix[D](https://arxiv.org/html/2607.29602#A4)\.

## Appendix FDyad Difficulty

Table[A5](https://arxiv.org/html/2607.29602#A6.T5)lists the eight hardest and eight easiest dyads by mean model\-estimated probability of a correct human response, from the crossed random\-effects model of §[5\.5](https://arxiv.org/html/2607.29602#S5.SS5)\. Several of the hardest dyads fall below chance in every modality, functioning as systematic lures\. Estimated easiness correlates across modality pairs \(text–audior=0\.52r=0\.52, text–videor=0\.33r=0\.33, audio–videor=0\.61r=0\.61; allp<\.01p<\.01\), indicating that difficulty is largely a property of the conversation rather than the observation channel\.

Table[A4](https://arxiv.org/html/2607.29602#A6.T4)decomposes the human–model difficulty correlation reported in §[5\.5](https://arxiv.org/html/2607.29602#S5.SS5)\. The raw pooled correlation \(human\-crowd accuracy vs\. accuracy pooled over all models\) is strong in audio, moderate in video, and only marginal in text\. The weak raw text value is a suppression effect rather than absence of shared structure: humans and models carry opposing class biases \(§[5\.1](https://arxiv.org/html/2607.29602#S5.SS1)\), so their per\-dyad accuracies partly anti\-align on the familiar/stranger axis\. Partialling out the true class raises the within\-class correlation—fine\-grained difficulty that is not reducible to a two\-way class effect—tor=0\.61r=0\.61\(text\),0\.580\.58\(audio\), and0\.440\.44\(video\), allp<\.001p<\.001\. The shared difficulty is, however, partly a crowd\-level property: repeating the analysis with the single strongest model per modality in place of the pooled estimate leaves it robust only in video \(r=0\.45r=0\.45,p<\.001p<\.001\), with weak raw associations in audio \(r=0\.19r=0\.19,p=\.06p=\.06\) and text \(r=−0\.04r=\-0\.04, n\.s\.\); confidence\-weighted single\-model estimates match the binary ones almost exactly, so this is not an artifact of one prediction per dyad\. Individual models thus express the human\-aligned difficulty signal only noisily, and pooling recovers it\.

Table A4:Human–model per\-dyad difficulty correlation \(rr\) under four estimators, by modality \(N=96N=96dyads\)\. “Pooled” averages model accuracy over all models; “best model” is the single most accurate model per modality \(textgpt\_text\_mini, audiogemini\_audio\_pro, videogemini\_video\)\. “Within\-class” partials out the true familiar/stranger label\.![Refer to caption](https://arxiv.org/html/2607.29602v1/x3.png)Figure A1:Humans and models tend to find the same dyads hard \(§[5\.5](https://arxiv.org/html/2607.29602#S5.SS5)\)\. Each point is one of the 96 EO dyads, plotting human per\-dyad accuracy \(over all raters\) against model per\-dyad accuracy \(pooled over all models\), colored by true class; the diagonal marks equal difficulty\. Per\-dyad difficulty correlates positively in audio \(r=0\.56r=0\.56,≈31%\\approx 31\\%shared variance\) and video \(r=0\.34r=0\.34,≈12%\\approx 12\\%; bothp<\.01p<\.01\); text \(r=0\.20r=0\.20\) is only marginally significant \(p=\.045p=\.045\)\.Table A5:Hardest and easiest eight dyads by mean estimated P\(correct human response\) across modalities \(%\)\. “str\.” = stranger, “fam\.” = familiar\.

Similar Articles