Tag
This paper studies how LLM agents' personalities evolve after major life events, using Big Five traits and introducing a benchmark called BFI-Adapt to evaluate the fidelity of event-induced personality changes across 14 models.
This paper introduces personality vectors for the Big Five traits extracted from language models to provide an interpretable account of emergent misalignment. It shows that misaligned fine-tuning shifts a model's personality along a specific signature (low agreeableness and conscientiousness, high extraversion and neuroticism), offering a human-readable diagnostic profile for safety phenomena.
This paper presents a multi-agent crowd simulation using LLM-driven agents that perceive and appraise each other through visual, auditory, and tactile channels, leading to emergent emotional contagion without explicit hand-authored affect transfer. The study demonstrates spatial, temporal, and personality-dependent contagion dynamics across five scenarios and evaluates backend-dependent appraisal behavior.
This paper introduces a mechanistic interpretability approach to steer LLM personality traits by identifying and intervening on latent features using sparse autoencoders, achieving controllable personality modulation while maintaining language performance.
This paper investigates whether fine-tuning LLMs on long-form essays with associated Big Five personality profiles stabilizes questionnaire responses and can induce target profiles, finding that while variance reduces, accuracy on the full five-dimensional profile remains near chance.