From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
Summary
This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.
View Cached Full Text
Cached at: 07/30/26, 09:59 AM
# From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs Source: [https://arxiv.org/abs/2607.26853](https://arxiv.org/abs/2607.26853) [View PDF](https://arxiv.org/pdf/2607.26853) > Abstract:Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors\. Existing studies of personality\-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait\-related expressions while lacking mechanistic evidence for the existence of internal personality\-related representations, their cross\-situational expression, and how these representations shape specific behaviors\. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality\-related internal representations, Situation as contexts that afford trait\-relevant responses, and Behavior as response patterns on broader social tasks\. We introduce a framework for discovering, controlling, and validating trait\-like representations in LLMs\. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition\. We validate their trait relevance through effects on behavior to situation, token\-level activation patterns, and robustness to paraphrasing\. Second, feature\-level interventions induce bidirectional trait\-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts\. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit\-tradeoff patterns consistent with findings from human personality research, providing behavioral\-level validation beyond personality scores\. Our findings provide evidence that LLMs contain controllable trait\-like representations linking internal states, situational expression, and behavioral outcomes\. ## Submission history From: Ruikang Zhang \[[view email](https://arxiv.org/show-email/44bba849/2607.26853)\] **\[v1\]**Wed, 29 Jul 2026 12:36:47 UTC \(325 KB\)
Similar Articles
Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
This paper introduces a mechanistic interpretability approach to steer LLM personality traits by identifying and intervening on latent features using sparse autoencoders, achieving controllable personality modulation while maintaining language performance.
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
This paper presents a multi-dimensional analysis of human-like behaviors in LLMs, examining prevalence, effects, and controllability across 21,000 conversations from four models, finding that behaviors vary by model and user factors, with implications for responsible design.
Beyond Static Personas: Situational Personality Steering for Large Language Models
This paper introduces IRiS, a training-free framework for situational personality steering in LLMs that moves beyond static persona modeling by identifying and leveraging situation-dependent persona neurons. The approach demonstrates that LLM behavior varies contextually and proposes neuron-based identification, retrieval, and weighted steering methods validated on PersonalityBench and a new SPBench benchmark.
Towards Trust Calibration in Socially Interactive Agents: Investigating Gendered Multimodal Behaviors Generation with LLMs
This paper investigates the use of LLMs to generate multimodal behaviors (verbal, vocal, gestural, facial) for trust calibration in socially interactive agents. The study finds that while LLMs can produce coherent behaviors aligned with intended trustworthiness traits, they also reproduce societal gender stereotypes.
Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits
A study investigating how Large Language Models exhibit systematic response distortion in expressing Dark Triad personality traits under social desirability cues, with implications for LLM benchmarking and alignment.