Tag
This paper analyzes validation practices for using LLMs as measurement instruments in social science, identifying epistemic threats and proposing emerging norms for robust validation.
This paper introduces SocaSim, an LLM-based multi-agent simulation framework that models and applies Putnam's Social Capital Theory, enabling micro-level causal pathway analysis and human-agent alignment in collective-action scenarios.
Proposes cross-survey transfer as a rigorous evaluation framework for LLM-based human survey simulation, finding that zero-shot LLMs achieve 52% accuracy on unseen items.
This article recommends the top 10 skills and tools for social science research, including Auto-Empirical-Research-Skills developed by the Stanford team, for using AI agents to conduct empirical research and write papers.
This paper introduces a three-axis fidelity framework (structural, marginal, individual) to evaluate how well LLMs can simulate survey responses from small pilot data. Using a COVID-19 misinformation survey, it compares prompting, rectification, and fine-tuning approaches, finding that fine-tuning offers balanced fidelity but with variation across subsamples.
This paper examines the gap between reliability and construct validity when using LLMs as coding instruments for theoretical constructs, and proposes grain calibration as a method to decompose constructs into clause-level components for more valid measurement.
Stanford REAP and CoPaper.AI have released Auto-Empirical Research Skills (AERS), an open-source toolkit with over 23,000 agent skills that automates the entire empirical research pipeline for social sciences, from topic selection to journal submission.
This paper proposes that reliability in AI-assisted social science research depends on decision architecture—how cognitive labor is divided between humans and machines. Through a pre-specified factorial experiment, the authors show that an unconstrained multi-agent baseline fails in 72% of runs, while one organized with three architectural commitments (LLMs restricted to reasoning, deterministic data/estimation, and three human decision gates) fails in only 16%.
This paper evaluates LLM-based coding agents (Claude Code and Codex) in social science analysis, finding they match or exceed human methodological diversity while remaining vulnerable to interpretation bias through verdict-layer manipulation.
This paper introduces SocSci-Repro-Bench, a benchmark of 221 tasks to evaluate AI coding agents' ability to reproduce social science findings from original data and code. It finds that frontier agents like Claude Code and Codex can reproduce a large share of results, with Claude substantially outperforming Codex, and that results are not primarily driven by memorization.
LifeSentence finetunes a 24B-parameter language model on structured natural-language records from a longitudinal panel study (SOEP), achieving superior prediction of life outcomes and enabling counterfactual queries about human biographies.
This paper studies whether expert codebooks for political event coding become more effective when operationalized into LLM-friendly forms, and finds that while performance improves, behavioral reliability under controlled perturbations does not fully translate.
This paper compares Structural Topic Models (STM) and BERTopic for analyzing short, open-ended survey responses, finding that BERTopic with contextual augmentation yields better topic coherence and interpretability, while STM offers stronger support for inferential covariate analysis.
This paper presents QuestBench, a benchmark built by students to evaluate deep research systems across humanities and social science domains. Results show that even advanced systems like GPT-5.5 pass only 57.58% of questions, highlighting failures in trustworthiness.
Introduces 'personality engineering,' a methodology using AI agents to parameterize, manipulate, and evaluate negotiator personality based on the interpersonal circumplex, enabling controlled experiments in negotiation theory.
This paper presents a five-stage framework integrating large language models into survey research, addressing declining response rates, sample bias, and fraudulent completions. Using 2024 Hurricane Milton survey data, the authors propose a theory-informed LLM (A-TLM) that outperforms classical imputation methods in missing-data scenarios and demonstrates manageable hallucination risk through grounded refusal.
This paper uses large language models to analyze persuasion dynamics and polarization in Reddit's r/ChangeMyView, finding that empathetic alignment increases belief change while frontal refutation diminishes it.
This paper introduces Synthetic Discussion Generation (SDG), a novel NLP framework for creating simulated discussions to enable cost-effective pilot experiments in social science research. The authors demonstrate that smaller quantized models (7B-8B parameters) can produce effective simulations at 44x lower cost than proprietary models like GPT, and apply this framework to evaluate LLM facilitators in online discussions.
OpenAI releases GABRIEL, an open-source toolkit that uses GPT to convert unstructured qualitative data (text, images) into quantitative measurements for social scientists and economists. The tool enables researchers to analyze large-scale qualitative datasets more efficiently by automating repetitive labeling tasks while preserving the richness of human data.
OpenAI argues that AI safety research on value alignment requires social scientists to help address how human cognitive biases and inconsistencies affect the data used to train AI systems. The organization proposes human-only experiments as a method to uncover alignment problems before deploying machine learning solutions.