Human Psychometric Questionnaires Mischaracterize LLM Behavior
Summary
This paper finds that human psychometric questionnaires fail to reliably predict LLM behavior in real-world interactions, and proposes generation-based profiling as a more accurate alternative.
View Cached Full Text
Cached at: 06/09/26, 08:43 AM
Paper page - Human Psychometric Questionnaires Mischaracterize LLM Behavior
Source: https://huggingface.co/papers/2509.10078
Abstract
Human psychometric questionnaires fail to reliably predict LLM behavior in real-world interactions, while generation-based profiling offers superior accuracy for understanding model responses to everyday user queries.
We examine whether humanpsychometric questionnairescan serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-sourceLLMsby comparing their value andpersonality profilesderived from two different methods:Likert self-reportson established questionnaires (PVQ-40/21andBFI-44/10) andgeneration probabilitiesovervalue-laden responsesto everyday user queries. The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears ingeneration probabilities. We attribute this gap to the fact that explicit lexical cues in established questionnaire items allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries provide no such cues. In addition,demographic persona promptsshift models’ responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in thegeneration probabilitiesof responses to realistic user queries, showing their limited ability to simulate the behaviors of target demographics in real-world user interactions. Overall, our study shows that humanpsychometric questionnairesare insufficient tools for predicting LLM behavior and suggests generation-based profiling as a more accurate measure.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2509\.10078
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2509.10078 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2509.10078 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2509.10078 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
This paper examines when and why self-reported psychometric measures predict the actual behavior of large language models, finding that fine-grained, behavior-specific instruments (Theory of Planned Behavior) achieve human-level coherence within a shared conversation, while broad traits like Big 5 do not.
Evaluating LLMs as Human Surrogates in Controlled Experiments
This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.
We gave 45 psychological questionnaires to 50 LLMs. What we found was not “personality.”
Researchers analyzed 50 LLMs across 45 psychometric questionnaires, identifying a 'Pinocchio Dimension' that measures how models endorse inner experiences rather than reflecting true personality traits.
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
This paper presents a psychometric audit of large language models as synthetic survey respondents, finding that they fail to match human joint distributions and reliability, making them unsuitable replacements.
Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?
This paper investigates whether fine-tuning LLMs on long-form essays with associated Big Five personality profiles stabilizes questionnaire responses and can induce target profiles, finding that while variance reduces, accuracy on the full five-dimensional profile remains near chance.