Tag
FriendBench is a new benchmark for evaluating whether humans and multimodal LLMs can infer if two people are familiar or strangers from a 20-second video clip of an ice-breaker conversation. Results show the best models match human accuracy but differ in bias, and only humans benefit from richer visual behavior.
This paper evaluates the predictive accuracy, cross-task generalizability, and test-retest reliability of multimodal features for measuring conversational states like cognitive load and power in dyadic remote collaborative tasks. Findings show that linguistic features predict well but generalize poorly, acoustic reliability degrades when controlling for speaker identity, and interaction features provide the most reliable signal.
This paper introduces DEPOOL, a controlled benchmark evaluating six temporal aggregation architectures across six frozen speech backbones for depression detection in dyadic interactions, finding that many configurations collapse into single-class predictions and that robustness should be a key criterion.