social-reasoning

Tag

Cards List
#social-reasoning

From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL

arXiv cs.AI · 2026-08-17 Cached

The paper introduces SocialRL, a reinforcement learning approach to enhance social reasoning in small language models, enabling them to negotiate effectively and match or exceed the performance of larger models like GPT-5 in various interaction domains.

0 favorites 0 likes
#social-reasoning

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

arXiv cs.CL · 2026-07-22 Cached

Introduces MeetingToM, a benchmark for evaluating multimodal LLMs on theory-of-mind reasoning in multi-party meetings, with tasks at subject, dyadic, and group levels including pseudo-consensus detection.

0 favorites 0 likes
#social-reasoning

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

arXiv cs.CL · 2026-07-01 Cached

This paper introduces Non-Conversational Planning Theory of Mind (NCP-ToM) and a novel evaluation framework, NCP-ExploreToM, to assess whether LLMs can induce specific belief states in other agents through actions rather than conversation. Testing on frontier models and humans across 600 tasks, GPT-5 achieved ~80% success, outperforming humans, though all models struggled more with false belief states.

0 favorites 0 likes
#social-reasoning

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

arXiv cs.AI · 2026-05-27 Cached

OmniToM introduces a benchmark that evaluates large language models' theory of mind by requiring explicit belief structure extraction and labeling, revealing a bottleneck in tracking actor-specific beliefs despite strong performance on endpoint QA tasks.

0 favorites 0 likes
#social-reasoning

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions

Hugging Face Daily Papers · 2026-05-15 Cached

GRASP is a large-scale dataset for social reasoning in multi-person videos, connecting high-level social questions with fine-grained gaze and gesture events, and introduces Social Grounding Reward to improve multimodal model understanding.

0 favorites 0 likes
#social-reasoning

RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity

arXiv cs.CL · 2026-04-20 Cached

RoleConflictBench is a novel benchmark containing over 13,000 scenarios across 65 roles designed to evaluate how well LLMs handle contextual sensitivity in role conflict situations where multiple social expectations clash. Analysis of 10 LLMs reveals that models predominantly rely on learned role preferences rather than dynamic contextual cues when making decisions.

0 favorites 0 likes
← Back to home

Submit Feedback