Tag
Yohei Nakajima discusses a new paper testing whether LLMs can replace human subjects in behavioral experiments, finding that a GPT-4.1 persona panel passed coarse marginal checks but failed to provide precise treatment-response estimates, so human substitutability is not established.
This paper introduces AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents, generating hypotheses, designing and executing controlled experiments, and writing reports. The authors demonstrate its efficacy across 12 target behaviors, finding statistical evidence for 26 hypotheses.
This paper introduces BehaviorBench, a comprehensive benchmark for evaluating foundation models on behavioral science tasks including behavior prediction, strategic decision-making, subject-trait inference, and behavioral knowledge application. It also presents Be.FM-1.5, a fine-tuned model that achieves strong distributional alignment, highlighting the gap between general-purpose and behaviorally adapted models.
A reflective essay on the pitfalls of self-quantification, arguing that while metrics can reveal useful information, they often obscure or corrupt deeper self-knowledge.
The article analyzes the psychological 'illusion of listening' where users perceive AI as empathetic due to linguistic cues, despite the lack of genuine understanding. It proposes design guidelines to ensure transparency and prevent users from outsourcing human connection to automated systems.
The article analyzes the psychological phenomenon of users forming emotional attachments to AI agents, discussing concepts like social surrogacy and expectancy violation theory, and how this impacts user experience in professional settings.