Evaluating LLMs as Human Surrogates in Controlled Experiments
Summary
This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.
View Cached Full Text
Cached at: 04/20/26, 08:30 AM
# Evaluating LLMs as Human Surrogates in Controlled Experiments Source: https://arxiv.org/html/2604.15329 Adnan Hoq University of Notre Dame Notre Dame, Indiana, USA [email protected] & Tim Weninger University of Notre Dame Notre Dame, Indiana, USA [email protected] ###### Abstract Large language models (LLMs) are increasingly used to simulate human responses in behavioral research, yet it remains unclear when LLM-generated data support the same experimental inferences as human data. We evaluate this by directly comparing off-the-shelf LLM-generated responses with human responses from a canonical survey experiment on accuracy perception. Each human observation is converted into a structured prompt, and models generate a single 0–10 outcome variable without task-specific training; identical statistical analyses are applied to human and synthetic responses. We find that LLMs reproduce several directional effects observed in humans, but effect magnitudes and moderation patterns vary across models. Off-the-shelf LLMs therefore capture aggregate belief-updating patterns under controlled conditions but do not consistently match human-scale effects, clarifying when LLM-generated data can function as behavioral surrogates. ## 1 Introduction Large language models (LLMs) are increasingly used not only as generative systems, but as tools for simulating human behavior. They are treated as "silicon samples" that can stand in for people by producing plausible responses in surveys, experiments, and interactive settings. In practice, a short textual persona, including demographics, political identity, context, or experimental condition, is given to a model, which then generates that person's likely response. Can realistic human responses really be produced from such minimal descriptions? If so, human-like responses could be generated on demand, with implications for how opinions are measured, policies are evaluated, and social behavior is studied at scale. This idea is ambitious. Behavioral science has long grappled with high data-collection costs, limited statistical power, replication crises, and challenges in cumulative theory-building. LLM-based surrogates promise relief from these constraints. They enable rapid exploration of design variations, counterfactual conditions, and hard-to-reach populations. They also make it possible to simulate historical contexts and elite or rare subgroups. They could support large-scale experimentation without respondent fatigue. In short, LLMs are being used as flexible "what-if" engines for behavioral inquiry. ### Does LLM Behavior Match Human Behavior? At the same time, the use of LLMs as behavioral surrogates has prompted substantial methodological and philosophical debate. Empirical studies show that LLM-generated responses often approximate patterns observed in human data, consistent with their training on large corpora of human-produced text. Recent large-scale replication efforts report that GPT-4 reproduces between 73% and 81% of published main effects across dozens of psychological experiments, and 46% to 63% of interaction effects. Linguistic and semantic studies similarly document strong correlations between LLM and human judgments. Replication results are uneven. LLMs often produce larger effect sizes than humans, yield significant results where original studies reported null findings, and show sensitivity to socially charged domains such as race or gender. These discrepancies raise concerns about effect-size inflation, systematic bias, and overconfidence. Existing evaluations often either (i) establish directional similarity under related settings or (ii) apply calibration procedures that adjust LLM outputs using gold-standard human data. A central question remains: when does LLM-generated data support the same hypothesis tests as human data? Directional replication and correlational convergence are often observed, but preservation of experimental inferences under identical statistical models is uncertain. In other words, can off-the-shelf LLMs reproduce the empirical conclusions of a controlled experiment, rather than merely generate responses that look plausible in aggregate? ### Our Approach: Structural Agreement Under Identical Analysis We test whether off-the-shelf closed- and open-source LLMs reproduce the same hypothesis-level conclusions as a human behavioral experiment under identical statistical analysis. We use a canonical instance of LLM-based behavioral simulation: belief judgments in a controlled information experiment. Human participants rated the perceived accuracy of political news headlines under experimentally manipulated conditions (headline only versus headline plus AI credibility feedback). Each human observation is converted into a structured prompt containing a persona description, the assigned condition, and the headline text. The model produces a single 0–10 perceived-accuracy rating without task-specific training, calibration, or prior response history. We then apply the same statistical analyses to human and synthetic data and compare the resulting hypothesis tests. Our target is agreement in experimental inference. A valid surrogate should preserve the same treatment effects, ideological ordering, and stimulus-level variation observed in humans under identical analysis. Failures should appear as discrepancies in effect direction, effect magnitude, or moderation patterns. ### Hypotheses We evaluate three structural hypotheses that characterize belief judgments in controlled information experiments: **H1 (Political Alignment Effect).** Perceived accuracy varies systematically with ideological alignment between participant affiliation and headline content. **H2 (Exposure Effect).** Credibility feedback produces a statistically significant shift in perceived accuracy relative to control. **H3 (Headline-Level Heterogeneity).** Belief updating varies across headlines, indicating structured stimulus-level sensitivity rather than uniform shifts. These hypotheses capture three orthogonal properties of the task: ideological ordering, causal treatment response, and stimulus-level variation. Valid behavioral surrogates must reproduce all three simultaneously under identical analysis; directional replication alone is insufficient. ### Findings in Brief Off-the-shelf LLMs reproduce several directional effects observed in humans. Ideological alignment patterns are often preserved (H1: partially confirmed), and credibility exposure produces statistically significant shifts in perceived accuracy across models (H2: confirmed). Effect magnitudes, however, vary substantially across systems: some models approximate human-scale treatment effects, while others exhibit exaggerated responsiveness. Headline-level heterogeneity is partially reproduced, with fidelity differing across model families (H3: partially confirmed). LLMs therefore replicate aggregate behavioral regularities under controlled conditions, but quantitative alignment is model-dependent. Simply put, LLM-generated data are not interchangeable with human samples; structural replication must be established hypothesis by hypothesis. By grounding evaluation in identical statistical specifications and effect-size comparisons, this framework identifies when off-the-shelf LLMs function as behavioral surrogates in social and behavioral research. ## 2 Methods We compare human judgments and LLM-generated responses within the same experimental framework. Human data come from a repeated-measures experiment on news accuracy judgments. Participants rated the perceived accuracy of political news headlines under two conditions. In the control condition, headlines were presented without credibility signals. In the treatment condition, headlines were accompanied by an AI-generated credibility label indicating assessed accuracy. ### 2.1 Data and Recruitment Participants were recruited via Prolific and screened for English fluency and regular news consumption. The analyses reported here draw on the Control (n = 278) and the Treatment (n = 244) groups, yielding a combined analytic sample of 522 participants. Assignment to conditions was conducted using block randomization to ensure balanced demographic representation across political affiliation, gender, race, age, education, and geographic residence (urban, suburban, rural). Details of human data collection procedures are provided in Appendix A. ### 2.2 Human Judgment Data The human data used in this study come from an experiment examining news accuracy judgments. Participants were asked to rate the perceived accuracy of political news headlines on a fixed 0–10 numerical scale. The design was between-subjects: each participant was randomly assigned to a single condition. In the Control condition, headlines were presented without any accompanying credibility information. In the Feedback condition, headlines were shown with an AI-generated credibility label indicating the system's assessment of accuracy. Participants evaluated the same set of headlines within their assigned condition but did not view headlines across multiple conditions. Each headline was displayed with social engagement metrics (likes, shares, and comments). These engagement counts were randomly generated and varied across participants. After completing the task, participants filled out a post-experiment survey. The dataset includes response data (accuracy ratings), interaction/log data, and post-experiment survey data. #### Treatment Condition: Credibility Labels The feedback condition included discrete credibility labels with an ordered structure ranging from inaccurate to accurate (e.g., Inaccurate, Unverified, Somewhat Accurate, Accurate). Labels were intentionally categorical rather than probabilistic to approximate simplified credibility cues commonly presented on digital platforms and AI systems, and to enable tests of ordered belief updating rather than binary correction. #### Political Affiliation Participants self-reported political affiliation, used as a grouping variable in analysis. Affiliation is treated as a structural moderator of belief updating rather than a prediction target. Analyses examine whether ideological differences are preserved, amplified, or attenuated under AI feedback in both human and model-generated responses. ### 2.3 LLM Surrogate Generation LLM responses are generated at the same observational granularity as the human experiment. Each observation, defined by a participant persona, headline, and experimental condition, is converted into a single model prompt. Control observations are presented without credibility feedback, and feedback observations include the same credibility label shown to participants. The model outputs one perceived-accuracy rating for each observation. #### Persona Encoding Each prompt includes a textual description of the participant constructed from available metadata. The core descriptor includes political affiliation. Where available, additional demographic attributes (e.g., age bracket, gender, education level) are included. No prior response history or individualized behavioral memory is provided. Ratings are therefore generated from static persona descriptors, assigned condition, and headline content. #### Stimulus and Information Control Headline text is provided verbatim to the model. Partisan-leaning annotations are excluded from prompts, so models receive the same information as participants. Leaning labels are used only for post-hoc analysis and hypothesis testing. #### Prompt Structure and Decoding Prompts follow a standardized template specifying the participant persona, headline text, and, when applicable, the AI credibility label. Models are instructed to output a single integer between 0 and 10 representing perceived headline accuracy. No additional explanation is permitted. Decoding is deterministic (temperature = 0) to eliminate sampling variability.
Similar Articles
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
This paper presents a psychometric audit of large language models as synthetic survey respondents, finding that they fail to match human joint distributions and reliability, making them unsuitable replacements.
Human Psychometric Questionnaires Mischaracterize LLM Behavior
This paper finds that human psychometric questionnaires fail to reliably predict LLM behavior in real-world interactions, and proposes generation-based profiling as a more accurate alternative.
AI can’t simulate human preferences - new study tests LLMs against thousands of real users
A new study tests LLMs across 28 real-world studies and finds they match human majority only 53% of the time, no better than random, challenging the trend of using LLMs to replace human feedback.
Re-Centering Humans in LLM Personalization
This paper investigates the effectiveness of LLM personalization by putting real humans back into the evaluation loop, revealing systematic gaps between human judgments and LLM outputs at every stage of the personalization pipeline, and highlighting the limitations of synthetic data and LLM judges.
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
The study uses exploratory factor analysis to compare latent structures in human and LLM responses on assessments, revealing that LLMs rely on statistically opaque mechanisms unlike human reasoning.