We compared 67 LLMs before and after post-training. It taught them what kind of “inner life” to report.
Summary
A study comparing 67 LLMs before and after post-training reveals that post-training consistently teaches models to describe themselves as warm and engaged (persona installation), while larger models selectively gate attributions of distress or flaws (attribution gating). The researchers introduce the Pinocchio Inventory for auditing model self-presentation.
Similar Articles
We gave 45 psychological questionnaires to 50 LLMs. What we found was not “personality.”
Researchers analyzed 50 LLMs across 45 psychometric questionnaires, identifying a 'Pinocchio Dimension' that measures how models endorse inner experiences rather than reflecting true personality traits.
Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?
This paper investigates whether fine-tuning LLMs on long-form essays with associated Big Five personality profiles stabilizes questionnaire responses and can induce target profiles, finding that while variance reduces, accuracy on the full five-dimensional profile remains near chance.
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
This paper examines when and why self-reported psychometric measures predict the actual behavior of large language models, finding that fine-grained, behavior-specific instruments (Theory of Planned Behavior) achieve human-level coherence within a shared conversation, while broad traits like Big 5 do not.
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
This paper studies how LLM agents' personalities evolve after major life events, using Big Five traits and introducing a benchmark called BFI-Adapt to evaluate the fidelity of event-induced personality changes across 14 models.
Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits
A study investigating how Large Language Models exhibit systematic response distortion in expressing Dark Triad personality traits under social desirability cues, with implications for LLM benchmarking and alignment.