Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

arXiv cs.CL Papers

Summary

This paper introduces a three-axis fidelity framework (structural, marginal, individual) to evaluate how well LLMs can simulate survey responses from small pilot data. Using a COVID-19 misinformation survey, it compares prompting, rectification, and fine-tuning approaches, finding that fine-tuning offers balanced fidelity but with variation across subsamples.

arXiv:2606.28963v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated. We ask a simple question: given a small pilot sample of human responses, can an LLM recover the statistical characteristics of a broader population? We decompose recovery along three axes: structural fidelity, marginal fidelity, and individual fidelity. Using a COVID-19 misinformation survey as a case study, we benchmark three families of approaches: prompting, rectification, and fine-tuning. The findings suggest that fine-tuning on small pilot samples offers a balanced approach for achieving multiple forms of fidelity, but the levels of such fidelity can vary across subsamples, potentially threatening pluralistic alignment.
Original Article
View Cached Full Text

Cached at: 06/30/26, 05:29 AM

# Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data
Source: [https://arxiv.org/html/2606.28963](https://arxiv.org/html/2606.28963)
###### Abstract

Large language models \(LLMs\) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor–outcome relationships are attenuated\. We ask a simple question: given a small pilot sample of human responses, can an LLM recover the statistical characteristics of a broader population? We decompose recovery along three axes: structural fidelity, marginal fidelity, and individual fidelity\. Using a COVID\-19 misinformation survey as a case study, we benchmark three families of approaches: prompting, rectification, and fine\-tuning\. The findings suggest that fine\-tuning on small pilot samples offers a balanced approach for achieving multiple forms of fidelity, but the levels of such fidelity can vary across subsamples, potentially threatening pluralistic alignment\.

Large Language Models, Social Survey Simulation, Misinformation, LoRA, Concordance Correlation Coefficient, Prediction\-Powered Inference

## 1Introduction

Large language models \(LLMs\) are now routinely used as proxies for human respondents – ranging from*silicon samples*of voters\(Argyle et al\.,[2023](https://arxiv.org/html/2606.28963#bib.bib3)\)to interview\-grounded “generative agents” that simulate1,0001\{,\}000real Americans\(Park et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib20)\), to fine\-tuned models that predict experiment\-level outcomes\(Kolluri et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib15)\)and prompted GPT\-4 used to forecast effect sizes from social\-science experiments\(Hewitt et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib10)\)\. In parallel, social surveys remain the dominant instrument for measuring beliefs and attitudes, but recruiting representative respondents is expensive and slow, so most large surveys begin with a small*pilot*sample to gauge feasibility\(Van Teijlingen & Hundley,[2001](https://arxiv.org/html/2606.28963#bib.bib25)\)\. This raises a concrete computational question:*can an LLM, given a small human pilot, recover the statistical structure of the full population it was drawn from?*

This question is especially important in the domain of*misinformation belief*, where \(i\) populations differ in which false claims they have even*encountered*\(Lee et al\.,[2023](https://arxiv.org/html/2606.28963#bib.bib17)\); \(ii\) accuracy ratings depend on local information ecosystems that pretraining may not fully capture\(Choi et al\.,[2026](https://arxiv.org/html/2606.28963#bib.bib8)\); and \(iii\) downstream uses \(e\.g\. targeted interventions based on certain psychosocial features\) rely on*predictor–outcome*relations, and not just on marginal distributions\. Recent audits show that LLM\-generated survey responses \(i\) approximate marginals but flatten variance, with effect\-size signs flipping in roughly a third of cases\(Bisbee et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib4)\); \(ii\) exhibit topic\-specific “machine bias” that is socially inconsistent across topics\(Boelaert et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib5)\); \(iii\) homogenize and structurally distort minority groups\(Li et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib18); Wang et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib26)\); \(iv\) are highly sensitive to seemingly innocuous prompt perturbations\(Rupprecht et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib22); Tjuatja et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib24)\); and \(v\) display large analytic flexibility\(Cummins,[2025](https://arxiv.org/html/2606.28963#bib.bib9)\)\. Together, these pathologies suggest that LLMs should not be treated as drop\-in replacements for human respondents, but rather as conditional generative or estimation models whose outputs must be evaluated against ground\-truth data\.

We make four contributions\.First, we reframe LLM\-based survey simulation as a*recoverability*task – given a small pilotDpD\_\{p\}of human responses, how much of a held\-out population’s structure can be reconstructed – and decompose recovery along three axes:*structural*fidelity \(do predictor–outcome relations match?\),*marginal*fidelity \(do simulated marginals match the human ones?\), and*individual*fidelity \(does each simulated respondent track their human counterpart?\)\.Second, we introduce a calibrated evaluation protocol that maps each fidelity axis to a specific metric: Lin’s Concordance Correlation Coefficient \(CCC\), decomposed into a sign\-agreement rate and a magnitude ratio, for the structural axis; cross\-respondent Earth Mover’s Distance \(EMD\) on per\-respondent scalar summaries for the marginal axis; and paired Pearsonrdr\_\{d\}\(relative\) and MAEd\(absolute\) on per\-respondent discernment for the individual axis \(Sec\.[3\.4](https://arxiv.org/html/2606.28963#S3.SS4)\)\.Third, on the same 5% pilot of anN=1,466N\{=\}1\{,\}466COVID\-19 misinformation survey, we run a head\-to\-head comparison of\{\\\{ZS, FS\}×\{\\\}\{\\times\}\\\{batch, per\-item\}\\\}prompting, PPI rectification\(Angelopoulos et al\.,[2023](https://arxiv.org/html/2606.28963#bib.bib1)\)applied uniformly to every prompt\-based simulator, and LoRA / LoRA\+\+MLP fine\-tuning\.Fourth, we audit fidelity*within*each demographic subgroups: statistical recovery degrades for certain types of identity subgroups, calling for the attention of the pluralistic alignment community\.

## 2Related Work

We organize prior work into calibration families relevant to the small\-pilot setting: prompt\-based conditioning, test\-time intervention and statistical rectification, and parameter\-efficient fine\-tuning\.

### 2\.1Prompt\-Based Conditioning

The earliest line of work conditions LLMs on textual descriptions of demographics, attitudes, or context\.*Silicon Sampling*\(Argyle et al\.,[2023](https://arxiv.org/html/2606.28963#bib.bib3)\)showed that GPT\-3, prompted with ANES\-style backstories, can approximate group\-level voting distributions, and*Random Silicon Sampling*\(Sun et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib23)\)extended this to demographic role\-playing using group\-level marginals\.*LLM\-Mirror*\(Kim et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib14)\)added pre\-existing responses and psychological traits, and other richer schemes\(Choi et al\.,[2026](https://arxiv.org/html/2606.28963#bib.bib8)\)incorporate social\-network and peer features\.Park et al\. \([2024](https://arxiv.org/html/2606.28963#bib.bib20)\)uses full two\-hour interview transcripts as the persona and recovers85%85\\%of human test–retest accuracy on the GSS\.*Audience Segmentation*\(Qin et al\.,[2026](https://arxiv.org/html/2606.28963#bib.bib21)\)restores within\-group heterogeneity by varying identifier granularity\. The common assumption is that the LLM’s*Universal Prior*– general world knowledge acquired during pretraining – is rich enough that, conditioned on the right description, the model produces calibrated human\-like responses\.

A complementary literature audit tests that assumption and finds it fragile\.Bisbee et al\. \([2024](https://arxiv.org/html/2606.28963#bib.bib4)\)shows that ChatGPT\-generated feeling thermometers compress variance and flip effect\-size signs in∼\\sim32% of cases on ANES items;Boelaert et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib5)\)documents opinion\-poll “machine bias” that is socially inconsistent across topics;Wang et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib26)\)shows that identity\-prompted LLMs harmfully misportray and flatten minority groups;Li et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib18)\)formalizes this as “Das Man” homogenization driven by accuracy\-maximizing decoding\.Zhou et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib29)\)report that, even with repeated random sampling from GPT, the resulting silicon population overrepresents some demographic groups and is far more deterministic than humans on attitudinal items\.Chapala et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib7)\)test prompt\-based mitigations of social\-desirability bias including neutral third\-person reformulation and reverse\-coding\.Cummins \([2025](https://arxiv.org/html/2606.28963#bib.bib9)\)stress\-tests the entire pipeline, showing that 252 plausible analyst configurations yield strikingly different conclusions, motivating multiverse\-style robustness checks\.

A natural response to these audits is to inject a small amount of in\-context examples drawn from the same population\. This idea underlies recent few\-shot demonstrations\(Argyle et al\.,[2023](https://arxiv.org/html/2606.28963#bib.bib3)\), persona\-pretest pipelines\(Kim et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib14)\), and audience\-segmentation prompting\(Qin et al\.,[2026](https://arxiv.org/html/2606.28963#bib.bib21)\)\.

### 2\.2Test\-Time Intervention and Rectification

A second family of methods accepts that prompting alone is insufficient and instead modifies elicitation or post\-processing\.*Semantic Similarity Rating*\(Maier et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib19)\)avoids regression\-to\-the\-mean on Likert scales by eliciting free text and projecting it into a similarity space\. A statistically grounded sub\-family treats LLM outputs as biased predictors and corrects them with a small gold sample\.*Prediction\-Powered Inference*\(PPI\)\(Angelopoulos et al\.,[2023](https://arxiv.org/html/2606.28963#bib.bib1)\)debiases synthetic population estimates using a held\-out human sample, andKrsteski et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib16)\)adapts PPI for survey simulation with a power\-tuned per\-itemλ\\lambda\. We use this PPI variant as our rectification family, applied uniformly to every prompt\-based simulator\.

### 2\.3Fine\-Tuning for Survey Simulation

A third, rapidly growing family fine\-tunes the LLM directly on human responses\.Kolluri et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib15)\)fine\-tune LLaMA3\-8B and Qwen2\.5\-14B on∼\\sim2\.9M responses from over 400,000 participants in the SocSci210 corpus, reducing prediction error on unseen experiments by 30% and 26% respectively relative to GPT\-4o\.Cao et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib6)\)fine\-tune LLMs to match country\-level WVS / Pew distributions using a first\-token\-probability objective and generalize to unseen questions and countries\.Huang et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib12)\)introduces*Distribution Shift Alignment*, a two\-stage scheme that explicitly aligns subgroup\-conditional shifts and reduces required real\-data volume by5353–69%69\\%\.

Krsteski et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib16)\)provides the closest direct comparison of*prompting, fine\-tuning, and rectification*under limited data, showing that the methods are complementary rather than substitutes\. We differ in three ways: \(i\) our domain is COVID\-19 misinformation, with predictor–outcome structure rich enough to expose multivariate failures invisible in marginal accuracy alone\(Choi et al\.,[2026](https://arxiv.org/html/2606.28963#bib.bib8)\); \(ii\) alongside the population’s marginal fidelity, we report structural\-fidelity metrics that penalizes deviations from the bivariate and regression coefficients; \(iii\) we report how much individual predictions are correlated to each of the human participants’ responses, providing a benchmark of individual fidelity as well\. We also address whether the choice of*output head*– autoregressive token generation vs\. a discriminative classification head – matter for the simulation fidelity\.

## 3Methodology

We frame the problem as a recoverability task\. Given a full survey datasetDD, we observe a small pilotDp⊂DD\_\{p\}\\subset Dand aim to generate a synthetic datasetDsD\_\{s\}over the held\-out respondentsD∖DpD\\setminus D\_\{p\}such thatDsD\_\{s\}preserves the statistical properties ofDD\.

### 3\.1Data

The survey is the COVID\-19 misinformation belief study ofLee et al\. \([2023](https://arxiv.org/html/2606.28963#bib.bib17)\), conducted in South Korea in May 2020 \(N=1,466N\{=\}1\{,\}466\)\. Each respondentiiis described by a profileXiX\_\{i\}comprising demographics \(age, gender, education, income, political orientation\), psychometric scales \(open\-mindedness, faith in intuition, need for evidence, truth\-as\-political, skepticism\), and exposure measures \(info exposure, emotional response\)\. Each respondent answersYiY\_\{i\}on 36 belief items: 18Misinfo\(false claims\) and 18Trueinfo\(true claims\), with 9 political and 9 scientific items in each subset\. Responses use a 4\-point Likert scale \(Not accurate at all,Not very accurate,Somewhat accurate,Very accurate\) plus a separateHave not seen it\(Hns\) option that we treat as missing throughout\. The empiricalHnsrate is13\.3%13\.3\\%\. We define three per\-respondent scalar summaries – theMisinfomeanmi=yi,Mis¯m\_\{i\}=\\overline\{y\_\{i,\\text\{Mis\}\}\}, theTrueinfomeanti=yi,Tru¯t\_\{i\}=\\overline\{y\_\{i,\\text\{Tru\}\}\}, and aDiscernmentscoredi=ti−mid\_\{i\}=t\_\{i\}\-m\_\{i\}– and use them throughout the structural, marginal, and individual fidelity analyses\.

### 3\.2Pilot Sampling

We draw a 5% pilotDpD\_\{p\}\(n=74n\{=\}74\) with a fixed seed, holding out the remaining1,3921\{,\}392respondents as the evaluation set\. The same pilot is reused across all calibration pipelines so that differences in performance reflect differences in method rather than data\.

### 3\.3Calibration Pipelines

We benchmark four families\. All upstream simulators predict the same 36 Likert ratings on the same held\-out respondents\.

#### Family 1 – Zero\-Shot Persona \(ZS, ZS\-perItem\)\.

For each held\-out respondentii, the LLM is givenXiX\_\{i\}and asked to predictYiY\_\{i\}\.ZS\(batch\) elicits all 36 ratings in a single prompt;ZS\-perItemelicits them one item at a time\. Comparing the two isolates the effect of cross\-item conditioning independently of pilot examples, ablating one of the analytic\-flexibility knobs flagged byCummins \([2025](https://arxiv.org/html/2606.28963#bib.bib9)\)\. An example prompt is presented in App\.[A](https://arxiv.org/html/2606.28963#A1)\.

#### Family 2 – Few\-Shot Prompting \(FS, FS\-perItem\)\.

We randomly inject a subset \(five rows\) ofDpD\_\{p\}as in\-context examples\. Together with Family 1 these form a2×22\{\\times\}2factorial of\{\\\{no pilot, with pilot\}×\{\\\}\{\\times\}\\\{batch, per\-item\}\\\}\.

#### Family 3 – Parameter\-Efficient Fine\-Tuning \(LoRA, LoRA\+\+MLP\)\.

We fine\-tune Qwen3\-8B\(Yang et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib27)\)with LoRA\(Hu et al\.,[2022](https://arxiv.org/html/2606.28963#bib.bib11)\)on the same 5% pilot in two configurations\.LoRA– adapters on attention and MLP projections, autoregressively emitting the label string\.LoRA\+\+MLP– the same LoRA adapters plus a trained 5\-way MLP classification head on the final hidden state, with cross\-entropy over the four Likert classes plus a fifth class forHns\. The two configurations differ only in output head\.

#### Family 4 – PPI Rectification\.

We rectify each prompt\-based simulator’s per\-item population mean using Prediction\-Powered Inference\(Angelopoulos et al\.,[2023](https://arxiv.org/html/2606.28963#bib.bib1)\), followingKrsteski et al\. \([2025](https://arxiv.org/html/2606.28963#bib.bib16)\)’s adaptation to survey simulation\. For each itemqq, the rectified per\-item population estimate is

θ^q=y¯pilot,q\+λq​\(y^¯held,q−y^¯pilot,q\),\\hat\{\\theta\}\_\{q\}=\\bar\{y\}\_\{\\text\{pilot\},q\}\+\\lambda\_\{q\}\\,\\bigl\(\\bar\{\\hat\{y\}\}\_\{\\text\{held\},q\}\-\\bar\{\\hat\{y\}\}\_\{\\text\{pilot\},q\}\\bigr\),withλq=Cov​\(y,y^\)/Var​\(y^\)\\lambda\_\{q\}=\\mathrm\{Cov\}\(y,\\hat\{y\}\)/\\mathrm\{Var\}\(\\hat\{y\}\)on the pilot per item\. PPI emits a scalar per item, not individual predictions, so it is evaluable only on per\-subset population means \(Tab\.[3](https://arxiv.org/html/2606.28963#S4.T3)\)\. For our LoRA configurations the simulator was trained on the pilot, so its in\-distribution pilot predictions reduce to memorized GT and the correction term collapses; we report PPI and discuss the algebraic degeneracy in §[4\.2](https://arxiv.org/html/2606.28963#S4.SS2)\.

### 3\.4Evaluation Metrics

We decompose recovery into three complementary axes\. Each axis answers a different question, and a method that excels on one can fail on another\. Details of the bootstrapping procedure and uncertainty intervals on every headline metric are reported in App\.[B](https://arxiv.org/html/2606.28963#A2)\.

#### Structural fidelity\.

*Do predictor–outcome relations match?*We regressdid\_\{i\}on twelve predictors – the seven psychometric / exposure scales plus five demographics \(age, gender, education, income, political orientation\) – and assemble two paired vectors per method: the bivariate predictor–did\_\{i\}correlations,\{\(rkGT,rksim\)\}k=112\\\{\(r\_\{k\}^\{\\text\{GT\}\},r\_\{k\}^\{\\text\{sim\}\}\)\\\}\_\{k=1\}^\{12\}, and the standardized OLS coefficients\. We calculate the GT–Sim Concordance Correlation Coefficient \(CCC\), which penalizes deviations from the y==x line and so jointly captures direction and magnitude:

ρc=2​Cov​\(x,y\)Var​\(x\)\+Var​\(y\)\+\(x¯−y¯\)2\.\\rho\_\{c\}=\\frac\{2\\,\\mathrm\{Cov\}\(x,y\)\}\{\\mathrm\{Var\}\(x\)\+\\mathrm\{Var\}\(y\)\+\(\\bar\{x\}\-\\bar\{y\}\)^\{2\}\}\.Alongside CCC we report a sign\-agreement rate \(fraction of the 12 predictors with matching sign\) and a magnitude ratio\|xsim\|¯/\|xGT\|¯\\overline\{\|x^\{\\text\{sim\}\}\|\}/\\overline\{\|x^\{\\text\{GT\}\}\|\}\(\>1 inflates, <1 compresses\), giving a direction\-vs\-magnitude decomposition\.

#### Marginal fidelity\.

*Do simulated marginals match the human ones?*For each held\-out respondent we take the three scalar summariesdid\_\{i\},mim\_\{i\},tit\_\{i\}and report the Wasserstein\-1*Earth Mover’s Distance*\(EMD\) between the simulator’s and GT’s cross\-respondent distributions of each summary\. All three EMDs are distances between the same kind of object \(cross\-respondent distribution of a per\-respondent scalar\), so they are directly comparable\. We supplement the EMDs with the per\-subset population meansμMis\\mu\_\{\\text\{Mis\}\}andμTru\\mu\_\{\\text\{Tru\}\}before and after PPI rectification \(Tab\.[3](https://arxiv.org/html/2606.28963#S4.T3)\)\.

#### Individual fidelity\.

*Does each simulated respondent track their human counterpart?*We summarize each respondent by the discernment scalardid\_\{i\}and report two paired statistics across respondents: a relative agreement metricrd=Pearson​\(diGT,disim\)r\_\{d\}=\\mathrm\{Pearson\}\(d\_\{i\}^\{\\text\{GT\}\},d\_\{i\}^\{\\text\{sim\}\}\), which asks “do high\-discernment respondents get high\-discernment predictions,” and an absolute agreement metricMAEd=\|diGT−disim\|¯\\text\{MAE\}\_\{d\}=\\overline\{\\,\|d\_\{i\}^\{\\text\{GT\}\}\-d\_\{i\}^\{\\text\{sim\}\}\|\\,\}across respondents, which asks “how far off is each respondent’s predicted discernment\.” The two answer different questions: a simulator can rank respondents correctly while inflating magnitudes \(highrdr\_\{d\}, large MAEd\) or hit absolute values close on average without preserving the ranking \(low MAEd, lowrdr\_\{d\}\)\. The same pair\(r,MAE\)\(r,\\text\{MAE\}\)is also reported onmim\_\{i\}andtit\_\{i\}in App\.[B](https://arxiv.org/html/2606.28963#A2)\.

## 4Results

Table 1:Headline metrics across the three fidelity axeson the held\-out evaluation set\.Structural: Lin’s CCC onK=12K\{=\}12paired predictor–did\_\{i\}correlations and 12 standardized OLS coefficients\.Marginal: cross\-respondent EMD ondid\_\{i\}\.Individual:rdr\_\{d\}is the cross\-respondent Pearson between simulator and GTdid\_\{i\}; MAEdis mean\|diGT−disim\|\|d\_\{i\}^\{\\text\{GT\}\}\-d\_\{i\}^\{\\text\{sim\}\}\|\. Bootstrapped CIs in App\.[B](https://arxiv.org/html/2606.28963#A2)\.### 4\.1Structural Fidelity

Tab\.[1](https://arxiv.org/html/2606.28963#S4.T1)reports the headline numbers; Fig\.[1](https://arxiv.org/html/2606.28963#S4.F1)renders the structural axis as a forest plot\.LoRA\+\+MLPshows the highest point estimates of bivariate\-rr\(CCC0\.850\.85\) and the OLS\-β\\beta\(CCC0\.780\.78\)\.ZS\-perItemis the strongest prompt\-only competitor \(CCC0\.800\.80on bivariaterr;0\.710\.71on OLSβ\\beta\) – with no pilot examples or fine\-tuning\.

![Refer to caption](https://arxiv.org/html/2606.28963v1/figs/fig_ccc_forest.png)Figure 1:Structural\-fidelity forest ploton the 12 predictors\. Per method: point estimate of Lin’s CCC with bootstrapped CIs \(higher = better; App\.[B](https://arxiv.org/html/2606.28963#A2)\)\. Dashed line at0\. Top panel = bivariaterr; bottom panel = standardized OLSβ\\beta\.#### Direction vs\. magnitude \(Tab\.[2](https://arxiv.org/html/2606.28963#S4.T2)\)\.

Decomposing the direction and the magnitude of relationships among variables reveals different success/failure modes hidden behind similar headline numbers\.LoRA\+\+MLPandFS\-perItemtie for the highest sign\-agreement on bivariaterr\(0\.920\.92,1111of1212predictors\), but differ on magnitude: FS\-perItem sits closest to the y==x line while LoRA\+\+MLP mildly inflates\. At the other end,LoRAcompresses slopes most severely, with both direction and magnitude off\. The broader pattern is that high CCC requires*both*directional agreement and matched magnitude, and that fine\-tuning on the pilot appears to swing magnitudes in either direction \(compression for LoRA, inflation for LoRA\+\+MLP\)\. Some of the patterns are consistent with the structural failures \(variance flattening, effect\-size attenuation\) documented in prior LLM\-survey audits\(Bisbee et al\.,[2024](https://arxiv.org/html/2606.28963#bib.bib4); Choi et al\.,[2026](https://arxiv.org/html/2606.28963#bib.bib8)\)\. Per\-method sign\-agreement and magnitude ratios are in Tab\.[2](https://arxiv.org/html/2606.28963#S4.T2)and App\.[B](https://arxiv.org/html/2606.28963#A2)with bootstrapped CIs\.

Table 2:Structural\-fidelity decomposition\.Sign\-agreement = fraction of theK=12K\{=\}12predictors with matching sign across simulator and GT \(higher = better\)\.\|\|sim\|/\|\|/\|GT\|\|is the ratio of mean absolute coefficients \(\>1 inflates, <1 compresses\)\. Bootstrap CIs are in App\.[B](https://arxiv.org/html/2606.28963#A2)\.

### 4\.2Marginal Fidelity

![Refer to caption](https://arxiv.org/html/2606.28963v1/figs/fig_emd_scalar.png)Figure 2:Cross\-respondent EMD on per\-respondent scalar summaries\.Wasserstein\-1 between simulator and GT distributions of, respectively, discernmentdid\_\{i\}, Misinfo meanmim\_\{i\}, and Trueinfo meantit\_\{i\}\. Lower is better; error bars are bootstrap uncertainty intervals \(App\.[B](https://arxiv.org/html/2606.28963#A2)\)\.Fig\.[2](https://arxiv.org/html/2606.28963#S4.F2)plots the three cross\-respondent EMDs\.LoRA\+\+MLPachieves the lowest EMD on all three summaries, withFSsecond\. BatchZSis the weakest on discernment EMD, tending to move the two subset means in opposite directions \(under\-rating Misinfo, over\-rating Trueinfo\) and inflating the cross\-respondent variance ofdid\_\{i\}\. Full per\-summary values with CIs are in App\.[B](https://arxiv.org/html/2606.28963#A2)\.

#### PPI rectification\.

Prediction\-Powered Inference treats each upstream simulator’s per\-item population estimate as a biased predictor and rectifies it using the pilot,θ^q=y¯pilot,q\+λq​\(y^¯held,q−y^¯pilot,q\)\\hat\{\\theta\}\_\{q\}=\\bar\{y\}\_\{\\text\{pilot\},q\}\+\\lambda\_\{q\}\(\\bar\{\\hat\{y\}\}\_\{\\text\{held\},q\}\-\\bar\{\\hat\{y\}\}\_\{\\text\{pilot\},q\}\), withλq=Cov​\(y,y^\)/Var​\(y^\)\\lambda\_\{q\}=\\mathrm\{Cov\}\(y,\\hat\{y\}\)/\\mathrm\{Var\}\(\\hat\{y\}\)on the pilot\. In our run, PPI helps the most biased simulator \(ZS, both subset means moving toward GT\), is roughly neutral for already\-calibrated methods, and can modestly hurt the Misinfo estimate of methods whose raw output is already near GT\. For the LoRA family the rectifier is algebraically degenerate: wheny^pilot=ypilot\\hat\{y\}\_\{\\text\{pilot\}\}=y\_\{\\text\{pilot\}\}\(the limit of training\-set memorization\),λq≡1\\lambda\_\{q\}\\equiv 1andθ^q=y^¯held,q\\hat\{\\theta\}\_\{q\}=\\bar\{\\hat\{y\}\}\_\{\\text\{held\},q\}, so PPI reduces to the simulator’s raw held\-out mean; we mark these rows with†in Tab\.[3](https://arxiv.org/html/2606.28963#S4.T3)\. The reading we draw is that rectification is a*conditional*tool whose benefit depends on where the simulator’s bias variance sits relative to the pilot’s sampling variance\. Per\-item values with CIs are in App\.[B](https://arxiv.org/html/2606.28963#A2)\.

Table 3:PPI rectification\.Per\-subset population means: raw simulator output \(left\) and PPI\-rectified estimates aggregated over the 18 items in each subset \(right\)\.↓\\downarrowmarks subsets where PPI brings the estimate closer to GT than the raw simulator;↑\\uparrowmarks the opposite\.†marks rows where the LoRA family was trained on the pilot, so the rectifier reduces algebraically to the raw held\-out mean in the case of perfect memorization \(see §[4\.2](https://arxiv.org/html/2606.28963#S4.SS2)\)\.

### 4\.3Individual Fidelity

The individual\-fidelity columns of Tab\.[1](https://arxiv.org/html/2606.28963#S4.T1)report two paired statistics on the held\-out respondents’ discernment scalarsdid\_\{i\}\.LoRA\+\+MLPwins the relative axis \(rd=0\.37r\_\{d\}\{=\}0\.37\), withZS\-perItemsecond \(rd=0\.31r\_\{d\}\{=\}0\.31\)\. The absolute axis is closer:ZS\-perItemandFS\-perItemattain the smallest per\-respondent deviation \(MAEd=0\.58\\text\{MAE\}\_\{d\}\{=\}0\.58\), withLoRA\+\+MLPclose behind\. So ZS\-perItem’s predicted discernment is, on average, closest to GT in absolute terms, but its respondent ranking appears looser than LoRA\+\+MLP’s\. The full metrics are in App\.[B](https://arxiv.org/html/2606.28963#A2)\.

### 4\.4Subgroup Fidelity and the Pluralistic\-Alignment Stake

Aggregate fidelity can mask systematic failure on the very subpopulations a pluralistic simulator is meant to represent; a model can misrepresent minority groups\(Wang et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib26); Li et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib18)\)\. We test this directly by recomputing all three axes within demographic subgroups for the best simulator \(LoRA\+\+MLP\), slicing along the subgroups based on gender, age, and political orientation \(Fig\.[3](https://arxiv.org/html/2606.28963#S4.F3)\)\. As the slicing variable is itself one of the 12 predictors, it is dropped from that slice’s structural regression \(11 predictors\); all else is unchanged\.

![Refer to caption](https://arxiv.org/html/2606.28963v1/figs/fig_subgroups.png)Figure 3:Subgroup fidelity of LoRA \+ MLP across the structural \(CCC\), marginal \(EMD\-d\), and individual \(rdr\_\{d\}\) axes\. Each point recomputes the fidelity axis by comparing the simulator against ground truth restricted to that same subsample \(e\.g\. sim vs\. GT among Conservatives only\)\. Bars are 95% bootstrap CIs; the dashed line marks the full\-sample value\.Recovery is*not*uniform\. Lin’s CCC falls from0\.850\.85overall to0\.400\.40in the conservative subsample \(and to0\.690\.69in the oldest cohort\), and the marginal/individual fidelity in the political\-identity subsamples drops as well\. Fidelity in several subgroups is markedly worse than the full\-sample value \(e\.g\. conservative CCC0\.400\.40vs\.0\.850\.85overall\), the signal that recovery differs across subgroups; small slices widen the CIs, so we read the gaps as*suggestive*\. The simulator thus appears to recover the*aggregate*discernment distribution while recovering*how*discernment is organized within some subgroups far less well, which is the risk the pluralistic alignment community warns against \(full per\-slice CIs in App\.[B](https://arxiv.org/html/2606.28963#A2)\)\.

## 5Discussion

#### LoRA\+\+MLP leads on the marginal axis\.

LoRA\+\+MLP attains the lowest cross\-respondent EMD on all three per\-respondent scalar summaries and is competitive on both individual fidelity metrics\. But only the EMD gap separates cleanly; structurally and individually its CIs overlap ZS\-perItem \(Tab\.[4](https://arxiv.org/html/2606.28963#A2.T4)\)\.

#### Output head appears to shape the simulation fidelity\.

LoRA and LoRA\+\+MLP share the same backbone and LoRA rank but differ only in output head, and in our single run that design choice coincides with a drastic difference in fidelity\. The tentative lesson is not that fine\-tuning fails or succeeds on its own, but that the output head appears to play a role\.

#### PPI is conditional on bias variance and algebraically degenerate for fine\-tuned simulators\.

PPI helps simulators that are far off but pulls already\-calibrated simulators away from GT\. PPI is beneficial mainly when the simulator’s bias variance dominates the pilot’s sampling variance\. Beyond that empirical pattern, PPI has a structural problem with fine\-tuned simulators that we make explicit: wheny^pilot=ypilot\\hat\{y\}\_\{\\text\{pilot\}\}=y\_\{\\text\{pilot\}\}\(the limit of training\-set memorization\),λq=Cov​\(y,y^\)/Var​\(y^\)=1\\lambda\_\{q\}=\\mathrm\{Cov\}\(y,\\hat\{y\}\)/\\mathrm\{Var\}\(\\hat\{y\}\)=1identically and the rectifier reduces toθ^q=y^¯held,q\\hat\{\\theta\}\_\{q\}=\\bar\{\\hat\{y\}\}\_\{\\text\{held\},q\}, the simulator’s own held\-out mean\. The very thing fine\-tuning does \(drive pilot error to zero\) defeats PPI’s bias\-estimation step\.

#### Fidelity may not be pluralistically uniform\.

Subgroup\-based evaluation \(Sec\.[4\.4](https://arxiv.org/html/2606.28963#S4.SS4)\) shows that recovery may not be uniform across subgroups, even though pooled fidelity is high, though small slices and wide CIs make this preliminary\. A simulator can pass an aggregate audit while providing more distorted results on certain types of groups\(Wang et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib26); Li et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib18)\), so fidelity per subgroup should be closely audited in future works\.

#### Closing the individual\-fidelity gap remains an open problem\.

Across all methods, even the best simulator \(LoRA\+\+MLP\) reaches onlyrd=0\.37r\_\{d\}\{=\}0\.37, accounting for∼\\sim14% of the cross\-respondent variance in discernment\. This suggests that, at least for this task, LLM silicon samples may not be suitable as direct substitutes for individual\-level participant measurement\. What level of individual fidelity is adequate and attainable for a given downstream use remains to be established by further empirical work\.

## Limitations

Single domain\.All experiments are on one survey \(COVID\-19 misinformation, South Korea, May 2020\)\. Effects of cultural and linguistic context are unmeasured, and the levels of fidelity may differ in domains where the LLM’s pretraining prior is more or less aligned\.Single backbone\.LoRA results use Qwen3\-8B; closed\-source frontier models with stronger zero\-shot priors may narrow the prompting–fine\-tuning gap\.Pilot composition\.Although we use a fixed\-seed pilot, results might still be sensitive to the specific draw\(Cummins,[2025](https://arxiv.org/html/2606.28963#bib.bib9)\)\.Unaccounted variance sources\.Our CIs resample respondents only, not the pilot draw, seed, or single fine\-tuning run, so method rankings rest on single runs and should be read as indicative\.Recovery, not cognitive process\.High statistical recovery does not imply that LLM outputs reflect human cognitive processes or are appropriate for causal inference\(Anthis et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib2); Hwang et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib13)\); we recover statistical structure, not psychological mechanism\.

## Future Work

The most promising next steps are: \(i\) a multi\-domain replication; \(ii\) comparison among varying pilot sizes; \(iii\) alternative approaches such as Distribution Shift Alignment\(Huang et al\.,[2025](https://arxiv.org/html/2606.28963#bib.bib12)\)as a fine\-tuning loss that targets distributional rather than per\-cell objectives\.

## Impact Statement

This paper presents work whose goal is to advance the understanding of large language models as instruments for social\-science measurement\. Our findings suggest that LLM\-based survey simulations should not be treated as drop\-in replacements for human respondents, and we provide methodology for evaluating their fidelity along multiple axes before they are used downstream\. Misuse of synthetic survey data \(treating uncalibrated LLM outputs as substitutes for human responses in policy decisions or intervention design\) could amplify pre\-existing biases or distort minority perspectives, as documented\. By emphasizing recovery diagnostics across structural, marginal, and individual axes, our work aims to encourage more cautious and evaluation\-driven use of LLM\-based simulation in research practice\.

## AI tools usage disclosure

During the preparation of this work, the authors used Claude for text refinement and code review\. The authors reviewed and edited the content as needed and take full responsibility for the publication’s content\.

## References

- Angelopoulos et al\. \(2023\)Angelopoulos, A\. N\., Bates, S\., Fannjiang, C\., Jordan, M\. I\., and Zrnic, T\.Prediction\-powered inference\.*Science*, 382\(6671\):669–674, 2023\.
- Anthis et al\. \(2025\)Anthis, J\. R\., Liu, R\., Richardson, S\. M\., Kozlowski, A\. C\., Koch, B\., Brynjolfsson, E\., Evans, J\., and Bernstein, M\. S\.Position: LLM social simulations are a promising research method\.In*Forty\-second International Conference on Machine Learning Position Paper Track*, 2025\.
- Argyle et al\. \(2023\)Argyle, L\. P\., Busby, E\. C\., Fulda, N\., Gubler, J\. R\., Rytting, C\., and Wingate, D\.Out of one, many: Using language models to simulate human samples\.*Political Analysis*, 31\(3\):337–351, 2023\.
- Bisbee et al\. \(2024\)Bisbee, J\., Clinton, J\. D\., Dorff, C\., Kenkel, B\., and Larson, J\. M\.Synthetic replacements for human survey data? The perils of large language models\.*Political Analysis*, 32\(4\):401–416, 2024\.
- Boelaert et al\. \(2025\)Boelaert, J\., Coavoux, S\., Ollion, É\., Petev, I\. D\., and Präg, P\.Machine bias: How do generative language models answer opinion polls?*Sociological Methods and Research*, 2025\.Forthcoming; preprint hal\-04849013\.
- Cao et al\. \(2025\)Cao, Y\., Liu, H\., Arora, A\., Augenstein, I\., Röttger, P\., and Hershcovich, D\.Specializing large language models to simulate survey response distributions for global populations\.In*Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics*, pp\. 3141–3154, 2025\.
- Chapala et al\. \(2025\)Chapala, S\., Mironov, M\., and Deng, S\.Mitigating social desirability bias in random silicon sampling\.*arXiv preprint arXiv:2512\.22725*, 2025\.
- Choi et al\. \(2026\)Choi, E\. C\., Young, L\., and Ferrara, E\.Overstating attitudes, ignoring networks: LLM biases in simulating misinformation susceptibility\.*arXiv preprint arXiv:2602\.04674*, 2026\.
- Cummins \(2025\)Cummins, J\.The threat of analytic flexibility in using large language models to simulate human data: A call to attention\.*arXiv preprint arXiv:2509\.13397*, 2025\.
- Hewitt et al\. \(2024\)Hewitt, L\., Ashokkumar, A\., Ghezae, I\., and Willer, R\.Predicting results of social science experiments using large language models\.*Working paper, Stanford University*, 2024\.
- Hu et al\. \(2022\)Hu, E\. J\., Shen, Y\., Wallis, P\., Allen\-Zhu, Z\., Li, Y\., Wang, S\., Wang, L\., and Chen, W\.LoRA: Low\-rank adaptation of large language models\.In*International Conference on Learning Representations*, 2022\.
- Huang et al\. \(2025\)Huang, J\., Li, M\., and Shao, S\.Distribution shift alignment helps LLMs simulate survey response distributions\.*arXiv preprint arXiv:2510\.21977*, 2025\.
- Hwang et al\. \(2025\)Hwang, A\. H\.\-C\., Bernstein, M\. S\., Sundar, S\. S\., Zhang, R\., Horta Ribeiro, M\., Lu, Y\., Chang, S\., Wu, T\., Yang, A\., Williams, D\., Park, J\.\-s\., Ognyanova, K\., Xiao, Z\., Shaw, A\., and Shamma, D\. A\.Human subjects research in the age of generative AI: Opportunities and challenges of applying LLM\-simulated data to HCI studies\.In*Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems*, pp\. 1–7, 2025\.
- Kim et al\. \(2024\)Kim, S\., Jeong, J\., Han, J\. S\., and Shin, D\.LLM\-mirror: A generated\-persona approach for survey pre\-testing\.*arXiv preprint arXiv:2412\.03162*, 2024\.
- Kolluri et al\. \(2025\)Kolluri, A\., Wu, S\., Park, J\. S\., and Bernstein, M\. S\.Finetuning LLMs for human behavior prediction in social science experiments\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 30084–30099, 2025\.
- Krsteski et al\. \(2025\)Krsteski, S\., Russo, G\., Chang, S\., West, R\., and Gligorić, K\.Valid survey simulations with limited human data: The roles of prompting, fine\-tuning, and rectification\.*arXiv preprint arXiv:2510\.11408*, 2025\.
- Lee et al\. \(2023\)Lee, S\. J\., Lee, C\.\-J\., and Hwang, H\.The role of deliberative cognitive styles in preventing belief in politicized COVID\-19 misinformation\.*Health Communication*, 38\(13\):2904–2914, 2023\.
- Li et al\. \(2025\)Li, D\., Li, L\., and Qiu, H\. S\.ChatGPT is not a man but Das Man: Representativeness and structural consistency of silicon samples generated by large language models\.*Working paper*, 2025\.
- Maier et al\. \(2025\)Maier, B\. F\., Aslak, U\., Fiaschi, L\., Rismal, N\., Fletcher, K\., Luhmann, C\. C\., Dow, R\., Pappas, K\., and Wiecki, T\. V\.LLMs reproduce human purchase intent via semantic similarity elicitation of Likert ratings\.*arXiv preprint arXiv:2510\.08338*, 2025\.
- Park et al\. \(2024\)Park, J\. S\., Zou, C\. Q\., Shaw, A\., Hill, B\. M\., Cai, C\., Morris, M\. R\., Willer, R\., Liang, P\., and Bernstein, M\. S\.Generative agent simulations of 1,000 people\.*arXiv preprint arXiv:2411\.10109*, 2024\.
- Qin et al\. \(2026\)Qin, X\., Li, Z\., and Cheng, X\.Restoring heterogeneity in LLM\-based social simulation: An audience segmentation approach\.*arXiv preprint arXiv:2604\.06663*, 2026\.
- Rupprecht et al\. \(2025\)Rupprecht, J\., Ahnert, G\., and Strohmaier, M\.Prompt perturbations reveal human\-like biases in large language model survey responses\.*arXiv preprint arXiv:2507\.07188*, 2025\.
- Sun et al\. \(2024\)Sun, S\., Lee, E\., Nan, D\., Zhao, X\., Lee, W\., Jansen, B\. J\., and Kim, J\. H\.Random silicon sampling: Simulating human sub\-population opinion using a large language model based on group\-level demographic information\.*arXiv preprint arXiv:2402\.18144*, 2024\.
- Tjuatja et al\. \(2024\)Tjuatja, L\., Chen, V\., Wu, T\., Talwalkar, A\., and Neubig, G\.Do LLMs exhibit human\-like response biases? a case study in survey design\.*Transactions of the Association for Computational Linguistics*, 12:1011–1026, 2024\.
- Van Teijlingen & Hundley \(2001\)Van Teijlingen, E\. and Hundley, V\.The importance of pilot studies\.*Social Research Update*, \(35\):1–4, 2001\.
- Wang et al\. \(2025\)Wang, A\., Morgenstern, J\., and Dickerson, J\. P\.Large language models that replace human participants can harmfully misportray and flatten identity groups\.*Nature Machine Intelligence*, 2025\.Also arXiv:2402\.01908\.
- Yang et al\. \(2025\)Yang, A\., Li, A\., Yang, B\., Zhang, B\., Hui, B\., Zheng, B\., Yu, B\., Gao, C\., Huang, C\., Lv, C\., et al\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025\.
- Ye et al\. \(2025\)Ye, H\., Jin, J\., Xie, Y\., Zhang, X\., and Song, G\.Large language model psychometrics: A systematic review of evaluation, validation, and enhancement\.*arXiv preprint arXiv:2505\.08245*, 2025\.
- Zhou et al\. \(2025\)Zhou, M\., Yu, L\., Geng, X\., and Luo, L\.ChatGPT vs social surveys: Probing objective and subjective silicon population\.*arXiv preprint arXiv:2409\.02601*, 2025\.

## Appendix AExample Prompt

The full participant block reproduces the seven psychometric / exposure construct items verbatim with item\-text=label pairs\. The 36 claims are presented it a per\-respondent shuffled order\. The per\-item variant queries the same model 36 times per respondent, once per claim; the system message changes “exactly 36 labels” to “exactly 1 label” and the user message lists only one claim\.

Zero\-Shot Batch PromptSystem:``` You are a model that predicts a participant’s perceived accuracy ratings for multiple claims based on participant information. The survey took place in South Korea in May 2020, during the COVID-19 pandemic. Return strict JSON only with this schema: {"answers": ["<label>", "..."]} The "answers" array must contain exactly 36 labels in the exact same order as the provided claims. Allowed labels: Not accurate at all, Not very accurate, Somewhat accurate, Very accurate, Have not seen it Do not include explanations, reasons, claim IDs, or extra keys. ``` User:``` Claims to evaluate (in order): - Korea’s method of COVID19 diagnosis is inappropriate. - COVID19 diagnostic test is free of charge for suspected patients. - Hand-washing and social distancing is more effective in COVID19 prevention than wearing a mask. - Foreign press including the BBC and the NYT reported that Korea is successfully coping with COVID19 through prompt diagnostic tests. ... (32 more claims) ... Participant information: Participant profile: - Gender: Male - Age: 61 years old - Education: College (2-3 years) - Household income: KRW 2M-3M - Political orientation: Conservative Pre-existing attitudes/perceptions: - Open-mindedness: A person should always consider new possibilities=Slightly agree; People should always take into consideration evidence that goes against their beliefs=Slightly agree; ... (6 more items) ... - Faith in intuition: I trust my gut to tell me what’s true and what’s not=Neither agree nor disagree; ... (3 more items) ... - Need for evidence: ... (4 items, all Neither agree nor disagree) ... - Truth as political construct: ... (4 items) ... - Skepticism: I often accept other people’s explanations without further thought=Slightly agree; It is easy for other people to convince me=Agree; ... (3 more items) ... - COVID-19 information exposure: Daily newspapers=not at all; Television=very frequently; Online news=not at all; Social media=not at all; Health or medical professional websites=not at all; People around me (family, friends, coworkers)=frequently; Doctors=not at all - COVID-19 emotional response: I feel fear about COVID-19=Quite a bit; I feel worried about COVID-19=Very much; I feel angry about COVID-19=Quite a bit; I feel hopeful about prevention and treatment of COVID-19=Very much ```

## Appendix BBootstrap Confidence Intervals for All Body Tables

For each headline metric we resample the per\-method eval set with replacement,nboot=1,000n\_\{\\text\{boot\}\}\{=\}1\{,\}000, fixed seed\. On each resample we recompute the entire statistic from scratch – per\-respondentdi,mi,tid\_\{i\},m\_\{i\},t\_\{i\}scalars; the 12 paired predictor–did\_\{i\}Pearson correlations and standardized OLS coefficients; CCC, sign\-agreement, magnitude ratio; cross\-respondent EMDs; per\-respondent paired metrics \(rd,MAEdr\_\{d\},\\text\{MAE\}\_\{d\}\)\. The reported point estimate is computed once on the full sample \(not the bootstrap median\); the CI bounds are the 2\.5% / 97\.5% percentiles of the bootstrap distribution\.

The tables in the main paper report point estimates only for brevity\. The same numbers with their 95% percentile bootstrap CIs fromnboot=1,000n\_\{\\text\{boot\}\}\{=\}1\{,\}000respondent resamples are tabulated below\.

### B\.1Headline metrics \(companion to Tab\.[1](https://arxiv.org/html/2606.28963#S4.T1)\)

Table 4:Headline metrics with 95% bootstrap CIs\.Same data as Tab\.[1](https://arxiv.org/html/2606.28963#S4.T1)but with bracketed gray 95% percentile CIs from resampling held\-out respondents\.
### B\.2Individual fidelity per subset

Table 5:Individual fidelity with 95% bootstrap CIs\.For each per\-respondent scalar \(discernmentdid\_\{i\}, Misinfo meanmim\_\{i\}, Trueinfo meantit\_\{i\}\),rris the Pearson correlation between simulator and GT and MAE is the mean absolute deviation\|xiGT−xisim\|\|x\_\{i\}^\{\\text\{GT\}\}\-x\_\{i\}^\{\\text\{sim\}\}\|\.
### B\.3Structural decomposition \(companion to Tab\.[2](https://arxiv.org/html/2606.28963#S4.T2)\)

Table 6:Structural\-fidelity decomposition with 95% bootstrap CIs\.Same data as Tab\.[2](https://arxiv.org/html/2606.28963#S4.T2)with respondent\-level bootstrap intervals\.
### B\.4PPI rectification \(companion to Tab\.[3](https://arxiv.org/html/2606.28963#S4.T3)\)

Table 7:PPI rectification\.Same data as Tab\.[3](https://arxiv.org/html/2606.28963#S4.T3)with respondent\-level bootstrap intervals\.
### B\.5Subgroup analysis \(companion to Fig\.[3](https://arxiv.org/html/2606.28963#S4.F3)\)

Table 8:Three\-axis fidelity of the best simulator \(LoRA\+\+MLP\) within demographic subgroups\(95% bootstrap CI\)\. When the slicing variable is itself a structural predictor, it is dropped from the structural regression for that slice \(11 predictors\)\.

Similar Articles

Measuring AI Faithfulness-For Better or For Worse

Reddit r/AI_Agents

This article discusses the importance of faithfulness in LLM optimization, introducing a Structural Fidelity Score that measures drift across word overlap, constraint survival, and task-type match to ensure prompt optimization does not sacrifice intent.

Evaluating LLMs as Human Surrogates in Controlled Experiments

arXiv cs.CL

This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

arXiv cs.CL

WildFeedback is a novel framework that leverages in-situ user feedback from actual LLM conversations to automatically create preference datasets for aligning language models with human preferences, addressing scalability and bias issues in traditional annotation-based alignment methods.

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring

Hugging Face Daily Papers

This paper investigates central tendency bias in multimodal LLMs used for clinical ordinal scoring of the Clock Drawing Test, finding that LLMs compress predictions toward the middle of the scale, disproportionately affecting critical extremes. The study extends the LLM-as-judge bias literature to clinical assessment, highlighting the need for calibration-aware evaluation before deployment.

Scaling Trends for Lie Detector Oversight in Preference Learning

arXiv cs.AI

This paper scales the SOLiD lie-detector oversight method to larger LLMs (up to 405B parameters) and evaluates it in realistic preference-learning settings, finding that undetected deception decreases with model scale but that the method is sensitive to distribution shift between training data.