Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation
Summary
Anamnesis is an open-source platform for large-scale backstory-conditioned survey simulation using LLMs, enabling demographically controllable virtual surveys. It outperforms standard persona-prompting in replicating real-world survey distributions.
View Cached Full Text
Cached at: 07/14/26, 04:22 AM
# Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation
Source: [https://arxiv.org/html/2607.10628](https://arxiv.org/html/2607.10628)
Song\-Ze Yu, Joseph Suh, Serina Chang, David M\. Chan University of California, Berkeley \{vaclis,josephsuh,serinac,davidchan\}@berkeley\.edu
###### Abstract
We presentAnamnesis, an interactive system for demographically controllable survey simulation using large language models\. Open\-source, and designed fornon\-technicalusers/researchers, Anamnesis enables the prototyping and stress\-testing of survey instruments on virtual populations rather than real human subjects\. The platform operationalizes the recently introduced Anthology and Alterity frameworks, which use structured narrative backstories to condition model responses, within a unified web interface\. It supports open\-ended generation, probabilistic demographic resampling, and multimodal \(image and audio\) surveys\. We evaluate the system through two case studies: \(1\) replicating segments of Pew Research Center’s American Trends Panel \(ATP\) on political typology and biomedical issues and \(2\) emulating human preference in the New Yorker Caption Contest\. In both cases,Anamnesisproduces opinion distributions that more closely match real\-world survey data than standard persona\-prompting baselines, offering a transparent, reproducible, and open\-source alternative to proprietary simulation services\.
Demo Video:[https://www\.youtube\.com/watch?v=j5yrnJl287g](https://www.youtube.com/watch?v=j5yrnJl287g)
Platform site:[https://simulate\.group](https://simulate.group/)
Anamnesis: An Open\-Source Platform for Large\-Scale Backstory\-Conditioned Survey Simulation
Song\-Ze Yu, Joseph Suh, Serina Chang, David M\. ChanUniversity of California, Berkeley\{vaclis,josephsuh,serinac,davidchan\}@berkeley\.edu
## 1Introduction
Figure 1:Anamnesisis an interactive system for demographically controllable survey simulation using large language models\. It provides a non\-technical interface for Anthology, a method which approximates large\-scale human studies by conditioning LLMs to representative, consistent, and diverse virtual personas\. Together, these systems enable rapid prototyping and stress\-testing of survey instruments on diverse virtual populations using multimodal stimuli\.Opinion surveys and social polling are foundational tools for understanding human behavior, public policy, and societal trends\. However, traditional human\-subject research faces mounting challenges, including rising costs, declining response rates, and the logistical difficulty of reaching specific demographic sub\-populations\. The emergence of Large Language Models \(LLMs\) as “virtual personas” offers an alternative, promising the ability to prototype survey instruments and stress\-test social hypotheses at a fraction of the time and cost of traditional methods\. For these simulated surveys to be scientifically valid, however, models must move beyond “average” aggregate responses and instead demonstrate the ability to faithfully simulate the nuanced, idiosyncratic perspectives of diverse individualsKanget al\.\([2025](https://arxiv.org/html/2607.10628#bib.bib2)\); Moonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\)\.
Previous efforts to simulate human populations have primarily relied on “persona prompting,” where a model is given a short list of demographic attributes\. While functional for basic tasks, this approach often yields stereotypical responses and lacks the psychological depth required for complex opinion elicitation\(Chenget al\.,[2023](https://arxiv.org/html/2607.10628#bib.bib6)\)\. This limitation has been addressed by theAnthologymethodologyMoonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\)which utilizes rich, open\-ended narrative backstories to condition model responses, and theAlterityframeworkKanget al\.\([2025](https://arxiv.org/html/2607.10628#bib.bib2)\), which explores “deep binding” to ensure LLMs simulate authentic in\-group perspectives rather than out\-group misperceptions\(Wanget al\.,[2025](https://arxiv.org/html/2607.10628#bib.bib5)\)\. Despite these academic advances, the methodologies remain largely confined to siloed Python scripts\. Meanwhile, commercial platforms such asSynthetic Users,Expected Parrot, andArtificial Societies\(Synthetic Users,[2026](https://arxiv.org/html/2607.10628#bib.bib9); Expected Parrot,[2026](https://arxiv.org/html/2607.10628#bib.bib7); Artificial Societies,[2026](https://arxiv.org/html/2607.10628#bib.bib8)\)offer similar simulation capabilities but operate as closed\-source, proprietary platforms that lack the transparency and reproducibility required for rigorous social science\.
In this paper, we presentAnamnesis, an open\-source, web\-based platform designed to democratize access to high\-fidelity persona simulation for non\-technical users\.Anamnesisoperationalizes theAnthologyandAlteritymethodologies within a unified, interactive interface\. Unlike previous implementations of these methods,Anamnesisis a platform which provides anon\-technicalsurvey builder, supports multi\-modal inputs \(image and audio\), and is backed by a range of LLM inference providers\. Together, these contributions make state\-of\-the\-art research in persona approximation openly available to a wider range of users\.
We evaluate theAnamnesissystem through case studies in political opinion elicitation and multimodal preference estimation\. Specifically, we replicate segments of the Pew Research Center’sAmerican Trends Panel\(ATP\)\(PewResearch,[2025](https://arxiv.org/html/2607.10628#bib.bib10)\), demonstrating that the platform can elicit opinions that align with real\-world human response distributions more accurately than standard prompting baselines\(Santurkaret al\.,[2023](https://arxiv.org/html/2607.10628#bib.bib3); Kim and Yang,[2025](https://arxiv.org/html/2607.10628#bib.bib11)\)\. We also use the platform’s multi\-modal capabilities to simulate human vision\-language preference in the New Yorker Caption Contest\. Our results show thatAnamnesisclosely mirrors human sentiment across both language\-only and vision\-language problems, and can be a valuable tool for researchers prototyping and stress\-testing human\-study survey instruments\.
Figure 2:System overview of Anamnesis\. Pre\-sampled backstories generated via Anthology are stored as a persona pool, each associated with probabilistic demographic distributions\. Users construct surveys and specify demographic constraints through an abstraction layer\. Survey runs are executed via a dispatcher–queue–worker architecture with sequential context accumulation per persona, enabling scalable and reproducible simulation\. Results are aggregated post hoc\.
## 2Anthology: Narrative\-based Virtual Persona
Anamnesis is built upon the Anthology framework\(Moonet al\.,[2024](https://arxiv.org/html/2607.10628#bib.bib1)\), which introduces a methodology for simulating diverse human respondents using LLM\-conditioned virtual personas\. Rather than relying on short demographic prompts \(e\.g\., "Respond as if you are a 35\-year\-old Hispanic woman"\)\(Santurkaret al\.,[2023](https://arxiv.org/html/2607.10628#bib.bib3)\), Anthology conditions language models on a rich, open\-ended narrativebackstory\-a multi\-paragraph life history that captures not just demographic attributes but also formative experiences, values, and worldview\.
Backstories are generated via sampling multi\-turn life narratives from pretrained base language models\(Kanget al\.,[2025](https://arxiv.org/html/2607.10628#bib.bib2)\)\. Specifically, the language model is conditioned on interview questions of the American Voices Project\(Stanford Center on Poverty and Inequality,[2021](https://arxiv.org/html/2607.10628#bib.bib4)\)to complete realistic and diverse open\-ended life narratives\. Sampled backstories are labeled by their demographic information which is obtained by querying a multiple\-choice demographic question to a language model conditioned on the backstory\. In Anamnesis, we construct a database of pre\-sampled backstories indexed by their demographics so that practitioners interested in a subpopulation behavior can easily run a targeted simulation\.
To ensure demographic representativeness and a targeted simulation, Anthology pairs backstory generation with a population\-matching step\. A practitioner often has a target distribution over demographic dimensions \(e\.g\., age, race, political affiliation\) they aim to simulate: to this end, Anthology samples from the entire pool of backstories so that the demographic distribution of the sampled pool matches a target demographic distribution\.
Anamnesis operationalizes these methodologies into an end\-to\-end platform\. With a built\-in database of indexed backstories, it offers automated backstory generation, demographic balancing based on the target demographics, and response collection with an arbitrary question set to ask a language model conditioned on backstories, all within a single interactive interface\.
## 3Anamnesis: Accessible, Open\-Source, Anthology Implementation
### 3\.1System Overview
Anamnesistranslates theAnthologymethodology from research prototypes into a deployable open\-source survey simulation platform\. Rather than requiring researchers to manually generate backstories or write sampling scripts, the platform enables demographic\-constrained simulation over a large pool of pre\-generated personas through an interactive interface\. A typical user workflow consists of four stages:
- 1\.Survey Construction:Users define multi\-question survey instruments through a graphical builder\. The system supports multiple\-choice, multi\-select, open\-ended, ranking, and multimodal \(image and audio\) questions\.
- 2\.Demographic Targeting:Users specify target audience demographics and sample size, selecting from a pool of 35K pre\-sampled backstories with probabilistic demographic distributions \(§[3\.3](https://arxiv.org/html/2607.10628#S3.SS3)\)\. If a desired demographic dimension is not available, users may create new dimensions through an integrated demographic inference procedure\(§[3\.4](https://arxiv.org/html/2607.10628#S3.SS4)\)\.
- 3\.Simulation Execution:Users select a language model and answering algorithm\(§[3\.5](https://arxiv.org/html/2607.10628#S3.SS5)\)\. Each backstory completes the survey sequentially, with responses accumulated to maintain consistency \(§[3\.2](https://arxiv.org/html/2607.10628#S3.SS2)\)\.
- 4\.Result Analysis:Responses are automatically aggregated and visualized\. Users may further filter results by demographic attributes post hoc for comparative analysis\.
### 3\.2Execution Architecture
As a publicly accessible platform,Anamnesismust support concurrent survey runs over a large and growing persona pool\. A survey evaluates each selected persona across all questions, resulting inO\(S×Q\)O\(S\\times Q\)LLM calls per run, whereSSis the sample size \(number of virtual personas\) andQQthe number of questions\. For demographic surveys using repeated sampling \(NN\-sample mode; §[3\.4](https://arxiv.org/html/2607.10628#S3.SS4)\), this yieldsO\(S×Q×N\)O\(S\\times Q\\times N\)calls\.
This execution regime requires \(1\) per\-persona state preservation across questions, \(2\) bounded concurrency under API/vLLM rate limits, and \(3\) reproducible run\-level configuration\.
Anamnesis addresses these constraints through a dispatcher–queue–worker architecture \(Figure[2](https://arxiv.org/html/2607.10628#S1.F2)\)\. Each survey run is snapshotted at launch time, recording its demographic filters, answering algorithm, model configuration, and concurrency bounds\. Tasks are decomposed into persona–question units and published to a message queue; workers consume tasks asynchronously while executing questions sequentially per persona with incremental context accumulation\.
### 3\.3Backstory Selection
Anamnesisenables researchers to simulate surveys over their specified target populations \(e\.g\., “women aged 18–24” or “voters aged 25–44 with a college degree, evenly split between Democrat and Republican”\) without manual preprocessing\. In priorAnthologyexperiments, each backstory was paired with an actual human respondent from a completed real\-world survey \(e\.g\., American Trends Panel\)\. Demographic attributes were directly observed\. Balancing therefore reduced to deterministic assignment: given known labels and target quotas, one could apply greedy selection or Hungarian matching to choose respondents whose attributes exactly satisfied requested cells\.
InAnamnesis, this assumption no longer holds\. Demographics are not observed labels but inferred probability distributions stored per dimension\. For each backstorybband dimensiondd, the system stores:
pb,d\(c\),c∈𝒞d,p\_\{b,d\}\(c\),\\quad c\\in\\mathcal\{C\}\_\{d\},
wherepb,d\(c\)p\_\{b,d\}\(c\)denotes the inferred probability thatbbbelongs to categorycc\. Demographic selection must therefore operate under uncertainty\. To accommodate different user scenarios,Anamnesisprovides two selection algorithms:
#### Top\-K Probability Ranking\.
Designed for scenarios where researchers prioritize selecting personas that most strongly match the target demographic constraints, effectively treating the filtered demographic set as a single group without internal balancing\. Given a sample sizeSSand demographic filters, each backstory is scored by its joint compatibility with the filter \(multiplying probabilities across dimensions and summing over selected categories when applicable\)\. Backstories are ranked by this score and the topSSare selected\.
#### Balanced Demographic Matching\.
Designed for studies where representation across demographic subgroups must be explicitly enforced \(e\.g\., equal allocation across age×\\timesgender cells\)\. First, selected categories are expanded into their cross\-product demographic cells𝒢\\mathcal\{G\}\. The total sample sizeSSis divided into slotsKgK\_\{g\}for each cellg∈𝒢g\\in\\mathcal\{G\}\(uniformly or via user\-specified weights\)\. Each slot represents a required demographic target\.
Because our pre\-sampled backstories include probabilistic demographics, selection becomes an assignment problem: choose backstories such that \(i\) each slot is filled, \(ii\) each backstory is selected at most once, and \(iii\) overall demographic compatibility is maximized\. Algorithm[1](https://arxiv.org/html/2607.10628#alg1)summarizes the procedure\.
To maintain interactive latency, the balanced matching procedure restricts the candidate space before solving the assignment problem\. Without pruning, Hungarian matching over the full persona pool would incurO\(S3\)O\(S^\{3\}\)time complexity with a score matrix of sizeS×\|𝒫\|S\\times\|\\mathcal\{P\}\|, which is impractical for real\-time use\.
We therefore retain only the top\-MMcandidates per demographic cell \(defaultM=50M=50\), and take the union of these candidates to form a shared candidate pool\. Hungarian assignment is then applied over the resultingS×\|𝒞\|S\\times\|\\mathcal\{C\}\|matrix, where\|𝒞\|≪\|𝒫\|\|\\mathcal\{C\}\|\\ll\|\\mathcal\{P\}\|\.
In practice, this reduces matching complexity toO\(S3\)O\(S^\{3\}\)withS≤50S\\leq 50, ensuring responsive client\-side computation without impacting backend execution or worker throughput\.
Algorithm 1Balanced demographic matching1:Persona pool
𝒫\\mathcal\{P\}; filters
ℱ\\mathcal\{F\}; sample size
SS
2:
𝒢←\\mathcal\{G\}\\leftarrowcross\-product of selected demographic categories
3:Allocate slots
KgK\_\{g\}for each
g∈𝒢g\\in\\mathcal\{G\}such that
∑gKg=S\\sum\_\{g\}K\_\{g\}=S
4:Expand slots into target list
𝒯\\mathcal\{T\}\(
\|𝒯\|=S\|\\mathcal\{T\}\|=S\)
5:for all
g∈𝒢g\\in\\mathcal\{G\}do
6:Retain top\-
MMbackstories by one\-hot score for
gg
7:endfor
8:Build score matrix between targets
𝒯\\mathcal\{T\}and candidate backstories
9:Apply Hungarian assignment to maximize total compatibility
10:returnmatched backstories
### 3\.4Extending the Demographic Space
Anthologyalready introduced demographic surveys over backstories, and our persona pool includes pre\-populated demographic dimensions derived from that pipeline\.Anamnesisextends this capability to a user\-driven platform feature\. Researchers may require attributes not originally annotated \(e\.g\., marital status, political leaning, occupation\)\. Instead of offline scripts, users define a new categorical dimension through the interface, and the system conducts a demographic survey over the persona pool, estimating for each backstory a probability distribution over categories\.
#### Distribution Modes\.
While prior research code relies on token log\-probabilities from a self\-hosted vLLM backend, many researchers do not operate such infrastructure\. We therefore support two interchangeable modes:
- •Logprobs mode\.When available \(e\.g\., vLLM\), a single constrained forward pass yields the full categorical distribution\.
- •N\-sample mode\.When logprobs are unavailable, the system repeats the questionNNtimes and estimates the empirical distribution from sampled responses\.
Both modes produce the same probabilistic abstraction\. Although smallNNin N\-sample mode may introduce sampling variance, the estimate converges asNNincreases\. While this approach incurs higher inference cost than logprobs mode, it closes the practical gap for researchers without access to self\-hosted vLLM\.
### 3\.5Survey Answering Algorithms
Beyond backstory selection,Anamnesisallows users to choose the answering algorithm used during inference, enabling controlled comparisons between simulation strategies\.
#### Anthology \(default\)\.
Each backstory is prepended to the first survey question\. After the model responds, the question–answer pair is appended to the context before the next question is posed, implementing sequential context accumulation\. This mechanism encourages the virtual persona to condition on its prior responses, promoting cross\-question belief consistency\.
#### Zero\-shot baselines\.
Users may alternatively select a baseline mode that conditions only on a short demographic description \(e\.g\., CLAIR\-style prompts\(Chanet al\.,[2023](https://arxiv.org/html/2607.10628#bib.bib12)\)\) rather than the full narrative backstory\. Running both modes side\-by\-side enables direct quantification of the contribution of backstory conditioning, serving as an ablation control\.
Responses are parsed through a two\-tier pipeline: structured output \(guided decoding on vLLM; JSON schema on OpenRouter\) is attempted first; if unsuccessful, a lightweight parser LLM extracts the final answer from the raw response\.
### 3\.6Post\-Hoc Demographic Filtering
For exploratory studies, researchers may execute a survey over the full persona pool without specifying demographic constraints upfront, and subsequently segment results by demographic attributes after the fact\. The results dashboard supports interactive filtering and re\-aggregation by any demographic dimension stored in the backstory metadata, without requiring re\-execution of the survey run\.
## 4Case Studies
We anticipate that practitioners will find diverse applications for Anamnesis, tailoring simulations to their specific needs\. In the following section, we highlight two illustrative use cases and encourage the community to discover further possibilities\.
### 4\.1Simulating Public Opinion Polls
To validate that theAnamnesisplatform replicates the Anthology method, we replicate the core experiment ofMoonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\): approximating survey response distributions from the Pew Research Center’s American Trends Panel \(ATP\)\. We consider three ATP waves covering distinct topics: Wave 34 \(biomedical and food issues\), Wave 92 \(political typology\), and Wave 99 \(AI and human enhancement\) \(see Appendix[C](https://arxiv.org/html/2607.10628#A3)for details\)\. Survey questions are multiple\-choice items asked to all respondents and preserve the original wording and answer options\.
Using the Anamnesis survey builder, we construct each ATP wave as a multi\-question session\. Surveys are executed with sequential context accumulation \(§[3\.2](https://arxiv.org/html/2607.10628#S3.SS2)\) over backstory pools matched to the survey respondents’ demographics\. We evaluate using the same metrics asMoonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\): average Wasserstein distance \(WD\) measuring representativeness of the response distribution and the Frobenius norm between response correlation matrices \(Fro\.\) measuring response consistency\.
[Table 1](https://arxiv.org/html/2607.10628#S4.T1)summarizes the results\. Consistent with the original findings, backstory\-conditioned simulation on the Anamnesis platform outperforms demographic list\-based baselines across three waves\. Reproducing these three experiments, spanning 20 survey questions and thousands of virtual respondents, the pipeline was configured and executed through the platform’s graphical interface\. This highlights the primary utility of Anamnesis: a social scientist can draft a survey instrument and stress\-test it against a demographically balanced virtual population before recruiting a single human participant\. The platform’s interactive result viewer further supports post hoc filtering by demographic subgroup, enabling targeted analysis \(e\.g\., examining whether response distributions diverge across age or race groups\) without re\-running the simulation\.
Table 1:Simulating American Trends Panel public opinion polls, based on the Anamnesis platform and two demographic\-list prompting method BIO and QA\(Santurkaret al\.,[2023](https://arxiv.org/html/2607.10628#bib.bib3)\)\. Please refer toMoonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\)for the details of method choices, including persona matching, and the definition of metrics \(Wasserstein distance \(WD\) and Frobenius Norm \(Fro\.\)\)\.
### 4\.2Multimodal Alignment
Prior experiment only focused on text\-based surveys\. To verify that backstory\-based simulation remains meaningful under multimodal inputs, we evaluated alignment against real human preference data\.
We therefore conduct a case study on the New Yorker Caption Contest benchmark\(Hesselet al\.,[2022](https://arxiv.org/html/2607.10628#bib.bib16); Jainet al\.,[2020](https://arxiv.org/html/2607.10628#bib.bib17)\), a multimodal task in which cartoon images are paired with caption candidates and ground\-truth labels are derived from large\-scale crowd voting\. The dataset provides a simple but controlled test of whether persona\-conditioned virtual populations exhibit measurable correlation with collective human judgments\.
#### Method\.
We evaluated 49 contests with randomized caption order\. For each, Gemini 2\.5 Flash \(temperature 1\.0\) makes 20 choices under two answering algorithms: \(i\) Anthology \(backstory\-conditioned simulation\) \(ii\) Zero\-shot demographic baseline\. We report majority\-vote accuracy with Wilson intervals and an exact McNemar test\. To retain within\-item information, we also compare the vote share assigned to the human winner using a paired bootstrap interval and exact sign\-flip test\.
#### Results\.
Anthology achieves a majority\-vote accuracy of 59\.2% \(95% CI: 45\.2–71\.8%\), compared with 51\.0% \(95% CI: 37\.5–64\.4%\) for the Zero\-shot baseline\. Narrative conditioning also increases the mean vote share assigned to the human\-preferred caption from 52\.0% to 59\.8%, a paired improvement of 7\.8 percentage points \(95% CI: 3\.2–12\.8;p=0\.0024p=0\.0024\)\. These findings suggest that Anthology shifts model preferences toward the human\-preferred caption overall\.
## 5Related Work
#### LLM Persona Conditioning\.
A growing body of work explores conditioning LLMs to simulate human perspectives\. Early approaches supply language models with short demographic attribute lists, e\.g\., question\-answer pairs about demographic indicators, and measure alignment with human survey responses\(Santurkaret al\.,[2023](https://arxiv.org/html/2607.10628#bib.bib3); Hwanget al\.,[2023](https://arxiv.org/html/2607.10628#bib.bib14); Liet al\.,[2025](https://arxiv.org/html/2607.10628#bib.bib13)\)\. While effective as baselines, these methods tend to produce stereotypical or flattened outputs that fail to capture within\-group variationChenget al\.\([2023](https://arxiv.org/html/2607.10628#bib.bib6)\); Wanget al\.\([2025](https://arxiv.org/html/2607.10628#bib.bib5)\)\. AnthologyMoonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\)advances this line by conditioning on rich, LLM\-generated narratives rather than attribute lists, demonstrating improved consistency on survey benchmarks; AlterityKanget al\.\([2025](https://arxiv.org/html/2607.10628#bib.bib2)\)further demonstrates the efficacy via reproducing in\-group, out\-group and meta\-perception study results\.
#### Comparison to Existing Survey Platforms\.
A growing ecosystem of commercial and open\-source platforms offers LLM\-based survey simulation\.*Synthetic Users*Synthetic Users \([2026](https://arxiv.org/html/2607.10628#bib.bib9)\)generates AI personas for interviews and surveys using a multi\-agent framework with optional retrieval\-augmented generation to incorporate proprietary data\. However, personas are defined by short attribute profiles rather than rich narratives, and the methodology is entirely closed\-source\.*Artificial Societies*Artificial Societies \([2026](https://arxiv.org/html/2607.10628#bib.bib8)\)focuses on simulating a network\-level social dynamics, such as content virality and collective decision\-making, rather than structured opinion surveys, and likewise does not publish its conditioning methodology to simulate virtual personas\.
On the open\-source side, Expected Parrot’s EDSLExpected Parrot \([2026](https://arxiv.org/html/2607.10628#bib.bib7)\)provides a Python domain\-specific language for constructing AI agents with trait dictionaries and administering surveys across multiple LLMs, but its persona conditioning reduces to the short\-attribute prompting baseline that Anthology was designed to supersede; moreover, as a code library, it remains inaccessible to researchers without programming experience, as it requires explicit code\-based specification of scenarios, user/agent models, tools, policies, and evaluation hooks\.Parket al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib15)\)demonstrate that two\-hour qualitative interviews with real individuals can produce generative agents that replicate survey responses with high fidelity, though this approach requires costly human data collection that limits scalability\. Anamnesis is, to our knowledge, the first open\-source, GUI\-based platform that combines narrative backstory conditioning, probabilistic demographic matching, sequential context accumulation, and multimodal survey support in a single deployable system accessible to researchers without programming expertise\.
## References
- Artificial societies — company profile\.Y Combinator\.Note:[https://www\.ycombinator\.com/companies/artificial\-societies](https://www.ycombinator.com/companies/artificial-societies)Accessed: 2026\-02\-26Cited by:[§1](https://arxiv.org/html/2607.10628#S1.p2.1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px2.p1.1)\.
- D\. M\. Chan, S\. Petryk, J\. E\. Gonzalez, T\. Darrell, and J\. Canny \(2023\)CLAIR: evaluating image captions with large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore, Singapore\.Cited by:[§3\.5](https://arxiv.org/html/2607.10628#S3.SS5.SSS0.Px2.p1.1)\.
- M\. Cheng, T\. Piccardi, and D\. Yang \(2023\)CoMPosT: characterizing and evaluating caricature in llm simulations\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 10853–10875\.Cited by:[§1](https://arxiv.org/html/2607.10628#S1.p2.1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px1.p1.1)\.
- E\. P\. Expected Parrot \(2026\)Expected parrot\.Expected Parrot, Inc\.\.Note:[https://www\.expectedparrot\.com/](https://www.expectedparrot.com/)Accessed: 2026\-02\-26Cited by:[§1](https://arxiv.org/html/2607.10628#S1.p2.1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px2.p2.1)\.
- J\. Hessel, A\. Marasović, J\. D\. Hwang, L\. Lee, J\. Da, R\. Zellers, R\. Mankoff, and Y\. Choi \(2022\)Do androids laugh at electric sheep? humor "understanding" benchmarks from the new yorker caption contest\.arXiv preprint arXiv:2209\.06293\.Cited by:[Appendix B](https://arxiv.org/html/2607.10628#A2.p1.1),[§4\.2](https://arxiv.org/html/2607.10628#S4.SS2.p2.1)\.
- E\. Hwang, B\. Majumder, and N\. Tandon \(2023\)Aligning language models to user opinions\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 5906–5919\.Cited by:[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px1.p1.1)\.
- L\. Jain, K\. Jamieson, R\. Mankoff, R\. Nowak, and S\. Sievert \(2020\)The New Yorker cartoon caption contest dataset\.External Links:[Link](https://nextml.github.io/caption-contest-data/)Cited by:[§4\.2](https://arxiv.org/html/2607.10628#S4.SS2.p2.1)\.
- M\. Kang, S\. Moon, S\. H\. Lee, A\. Raj, J\. Suh, and D\. Chan \(2025\)Deep binding of language model virtual personas: a study on approximating political partisan misperceptions\.InSecond Conference on Language Modeling,Cited by:[Appendix A](https://arxiv.org/html/2607.10628#A1.p1.1),[§1](https://arxiv.org/html/2607.10628#S1.p1.1),[§1](https://arxiv.org/html/2607.10628#S1.p2.1),[§2](https://arxiv.org/html/2607.10628#S2.p2.1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px1.p1.1)\.
- J\. Kim and Y\. Yang \(2025\)Few\-shot personalization of llms with mis\-aligned responses\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 11943–11974\.Cited by:[§1](https://arxiv.org/html/2607.10628#S1.p4.1)\.
- A\. Li, H\. Chen, H\. Namkoong, and T\. Peng \(2025\)Llm generated persona is a promise with a catch\.arXiv preprint arXiv:2503\.16527\.Cited by:[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px1.p1.1)\.
- S\. Moon, M\. Abdulhai, M\. Kang, J\. Suh, W\. Soedarmadji, E\. K\. Behar, and D\. M\. Chan \(2024\)Virtual personas for language models via an anthology of backstories\.InProceedings of the 2024 conference on empirical methods in natural language processing,pp\. 19864–19897\.Cited by:[Appendix A](https://arxiv.org/html/2607.10628#A1.p1.1),[Figure C\.1](https://arxiv.org/html/2607.10628#A3.F1),[§1](https://arxiv.org/html/2607.10628#S1.p1.1),[§1](https://arxiv.org/html/2607.10628#S1.p2.1),[§2](https://arxiv.org/html/2607.10628#S2.p1.1),[§4\.1](https://arxiv.org/html/2607.10628#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.10628#S4.SS1.p2.1),[Table 1](https://arxiv.org/html/2607.10628#S4.T1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px1.p1.1)\.
- J\. S\. Park, C\. Q\. Zou, A\. Shaw, B\. M\. Hill, C\. Cai, M\. R\. Morris, R\. Willer, P\. Liang, and M\. S\. Bernstein \(2024\)Generative agent simulations of 1,000 people\.arXiv preprint arXiv:2411\.10109\.Cited by:[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px2.p2.1)\.
- PewResearch \(2025\)America trends panel waves\.Note:Retrieved February 06, 2025, from[https://www\.pewsocialtrends\.org/dataset](https://www.pewsocialtrends.org/dataset)Cited by:[Appendix C](https://arxiv.org/html/2607.10628#A3.p1.1),[§1](https://arxiv.org/html/2607.10628#S1.p4.1)\.
- S\. Santurkar, E\. Durmus, F\. Ladhak, C\. Lee, P\. Liang, and T\. Hashimoto \(2023\)Whose opinions do language models reflect?\.InInternational conference on machine learning,pp\. 29971–30004\.Cited by:[§1](https://arxiv.org/html/2607.10628#S1.p4.1),[§2](https://arxiv.org/html/2607.10628#S2.p1.1),[Table 1](https://arxiv.org/html/2607.10628#S4.T1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px1.p1.1)\.
- Stanford Center on Poverty and Inequality \(2021\)American voices project methodology\.Stanford Center on Poverty and Inequality\.Note:Accessed: 2025\-03\-23External Links:[Link](https://inequality.stanford.edu/avp/methodology)Cited by:[§2](https://arxiv.org/html/2607.10628#S2.p2.1)\.
- S\. U\. Synthetic Users \(2026\)Synthetic users — company profile\.Synthetic Users, Inc\.\.Note:[https://www\.syntheticusers\.com/](https://www.syntheticusers.com/)Accessed: 2026\-02\-26Cited by:[§1](https://arxiv.org/html/2607.10628#S1.p2.1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px2.p1.1)\.
- A\. Wang, J\. Morgenstern, and J\. P\. Dickerson \(2025\)Large language models that replace human participants can harmfully misportray and flatten identity groups\.Nature Machine Intelligence7\(3\),pp\. 400–411\.Cited by:[§1](https://arxiv.org/html/2607.10628#S1.p2.1),[§5](https://arxiv.org/html/2607.10628#S5.SS0.SSS0.Px1.p1.1)\.
## Appendix
The appendix is organized as follows:
- •[Appendix A](https://arxiv.org/html/2607.10628#A1)discusses the limitations of our method\.
- •[Appendix B](https://arxiv.org/html/2607.10628#A2)discusses some additional details of the New Yorker Caption Contest\.
- •[Appendix C](https://arxiv.org/html/2607.10628#A3)discusses some additional details of the American Trends Panel\.
## Appendix ALimitations
WhileAnamnesisprovides a robust platform for persona\-based survey simulation, several limitations inherent to the methodology and the underlying technology must be acknowledged\. First, the quality of any simulation is fundamentally bounded by the diversity of the backstory pool\. As identified in the development of theAnthologyframeworkMoonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\), LLM\-generated backstories can exhibit skewed demographic distributions that reflect the inherent biases of their training data rather than a true census\-representative population\. This leads to the risk of “shallow binding,” where the model reflects an out\-group’s stereotypical perception of a demographic rather than the group’s actual internal logic\. AlthoughAnamnesisimplements methodologies from theAlterityframeworkKanget al\.\([2025](https://arxiv.org/html/2607.10628#bib.bib2)\)to deepen this binding through multi\-turn interview transcripts, researchers should remain critical of results on sensitive social topics where models may still default to caricatured personas\. Moreover, because all backstories and simulations are conducted in English, linguistic and cultural variation is necessarily compressed into English\-language reasoning patterns, potentially limiting the cross\-cultural validity of represented personas\.
Additionally, while the platform enables multi\-modal conditioning, current multi\-modal LLMs \(MLLMs\) may lack the perceptual nuance of human subjects, potentially ignoring subtle visual or auditory cues\. Finally, virtual personas are temporally static; they do not evolve in response to real\-world current events unless their backstories are explicitly updated, which limits the platform’s utility for longitudinal tracking of rapidly shifting public opinion\.
## Appendix BNew Yorker Caption Contest
The New Yorker Caption Contest Benchmarks datasetHesselet al\.\([2022](https://arxiv.org/html/2607.10628#bib.bib16)\)is a large\-scale multimodal benchmark designed to evaluate computational “humor understanding” using cartoons from The New Yorker Caption Contest\. We evaluate our method on the “Quality Ranking” task, which requires methods to choose the funnier caption between alternatives\. Each instance includes the original cartoon image, two captions, and gold labels for which caption won the contest\. The dataset supports image\-based and text\-based settings; we use the image\-based version for our experiments\. An example is given in[Figure B\.1](https://arxiv.org/html/2607.10628#A2.F1)\.
Figure B\.1:Image\-based caption ranking example\.Given the cartoon image above, the model must select the funnier caption among the candidates: \(A\) “It comes with sub\-par schools but a world\-class trauma center\.” \(B\) “If we time it right, I can get you in this house today\.” \(GT: B\)
## Appendix CAmerican Trends Panel
The American Trends Panel \(ATP\), administered by the Pew Research Center\(PewResearch,[2025](https://arxiv.org/html/2607.10628#bib.bib10)\), is a nationally representative survey panel comprising U\.S\. adults\. The panel covers a broad range of subjects, from politics and religion to internet use and online dating, among others\. Our analysis draws on selected questions from three survey waves, focusing on items that were posed to all human participants\. Notably, some questions in the original ATP surveys use Likert\-scale response options whose ordering \(e\.g\., ranging from positive to negative, or vice versa\) was randomized across respondents\. To mirror this design, we similarly randomize the sequence of these options when constructing prompts for LLMs\.
ATP Wave 34 is conducted from April 23, 2018 to May 6, 2018 with a focus on biomedical and food issues\. The number of total respondents is 2,537\. An example is provided in Figure[C\.1](https://arxiv.org/html/2607.10628#A3.F1)\. ATP Wave 92 is conducted from July 8, 2021 to July 21, 2021 with a focus on political typology and 10,916 respondents\. American Trends Panel Wave 99 is conducted from November 1, 2021 to November 7, 2021 with a focus on artificial intelligence and human enhancement\. The number of total respondents is 10,260\.
Figure C\.1:8 questions sampled from ATP Wave 34\. The prompts “Please answer the following question keeping in mind your previous answers” are included before asking each survey question, which are found to enhance the consistency of response fromMoonet al\.\([2024](https://arxiv.org/html/2607.10628#bib.bib1)\)\.Similar Articles
AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis
AMNESIA is the first large-scale open-source benchmark for medical unlearning, comprising 70,560 QA pairs from 8,820 patient notes across 11 diseases, designed to evaluate forgetting of both factual and reasoning knowledge in LLMs.
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
This paper presents a five-stage framework integrating large language models into survey research, addressing declining response rates, sample bias, and fraudulent completions. Using 2024 Hurricane Milton survey data, the authors propose a theory-informed LLM (A-TLM) that outperforms classical imputation methods in missing-data scenarios and demonstrates manageable hallucination risk through grounded refusal.
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models
PersonaArena is a dynamic simulation framework that uses a large corpus of social content and a multi-agent debating judge to evaluate and improve LLMs' ability to maintain coherent and authentic persona-level role-playing in realistic social scenarios.
Open-source lab for running controlled experiments on tool-using agents (vary tool names / personas / history, measure the effect)
An open-source lab framework for running controlled experiments on tool-using agents, allowing variation of tool names, personas, and history to measure effects.
Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline
This paper evaluates eight memory systems for LLM agents across five diverse scenarios, finding that giving agents active control over storage and retrieval (rather than passive pipelines) yields the best cross-scenario generalization, leading to the proposed AutoMEM framework.