familiarity-inference

Tag

Cards List
#familiarity-inference

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

arXiv cs.CL · 2026-08-03 Cached

FriendBench is a new benchmark for evaluating whether humans and multimodal LLMs can infer if two people are familiar or strangers from a 20-second video clip of an ice-breaker conversation. Results show the best models match human accuracy but differ in bias, and only humans benefit from richer visual behavior.

0 favorites 0 likes
← Back to home

Submit Feedback