Tag
This study examines how defendant statements affect LLM-simulated jurors, focusing on persuasion, ideological bias, and background affinity, and introduces the JuryBench benchmark for controversial criminal cases.
GPS-Bench is an evidence-grounded benchmark for governance policy simulation that uses legislative records and public evidence to model actors and outcomes, enabling controlled comparisons of LLM-based methods for policy analysis.
This paper presents INSIDE, a framework that fine-tunes LLMs to generate internal dialogue grounded in Bloom's Taxonomy, enabling student simulators to model both latent reasoning and observable actions. Evaluations show improved action fidelity and reasoning alignment compared to prompting baselines.
This paper introduces CoevolveSim, a framework for studying belief diffusion among networked generalist and specialist LLM agents, showing that specialist models and network structure affect consensus and influence dynamics.
This paper investigates formal mechanisms, such as Mediation, to maintain market stability among self-interested LLM agents (DeepSeek-V3) in a simulated marketplace, finding that Mediation enables recovery even under sustained adversarial attacks.
A new study tests LLMs across 28 real-world studies and finds they match human majority only 53% of the time, no better than random, challenging the trend of using LLMs to replace human feedback.
This paper studies hate speech cascades on Bluesky and uses multi-LLM agents to simulate them, finding that such simulations reproduce key patterns like stance monoculture and toxicity-delta direction, and that amplifier targeting on dense networks yields 7.5–12.9% reduction in hateful content with low benign collateral.
This paper systematically evaluates assumptions about LLM persona prompting and identifies 'persona manifold collapse,' where richer persona descriptions reduce behavioral diversity and simulation fidelity. The findings show that simple age-gender personas often outperform more detailed profiles.
This paper from Google DeepMind and Carnegie Mellon argues that LLM-simulated experiments are actually observational studies due to user drift, where the simulated population shifts with interventions. The authors propose using negative control outcomes to diagnose confounding and show that eliciting setting-relevant confounders can reduce bias.