CogGym:迈向人类与机器认知的规模化对比评估

arXiv cs.AI 论文

摘要

CogGym 是一个可扩展框架,利用认知实验比较人类与AI认知能力,研究发现规模更大的语言模型能更好地模仿人类推理,但仍落后于形式化基准。

arXiv:2609.21259v1 Announce Type: new Abstract: Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.
查看原文
查看缓存全文

缓存时间: 2026/09/21 09:20

# CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
Source: [https://arxiv.org/html/2609.21259](https://arxiv.org/html/2609.21259)
\#11,2\#13\#11\#14\#15\#16\#11,7,8\#11\#19\#110,15\#18\#11\#111\#11\#112\#18\#11\#18\#18\#113\#114\#115\#116\#116\#117,18,19\#120\#121\#11\#115\#113\#116\#12\#115\#14\#120\#122\#123\#123\#116\#119,24,25\#126\#121\#11\#15\#14\#127\#128\#129,30\#12\#121\#120\#131\#18\#115\#11,†\#11,††Co\-senior authors\.1Massachusetts Institute of Technology2Harvard University3Cornell University4Johns Hopkins University5Helmholtz Munich6Massachusetts General Hospital7University of Cambridge8Princeton University9Santa Fe Institute10Dartmouth College11EPFL12University of Chicago13University of Tübingen14CHI\-FRO15Stanford University16University of California, Los Angeles17University of British Columbia18Vector Institute19Canada CIFAR AI Chair20Yale University21University of California, Berkeley22University of Washington23New York University24McGill University25Mila–Quebec AI Institute26The University of Texas at Austin27Prior Computers28Purdue University29National University of Singapore30A\*STAR Institute of Advanced Intelligence and Computing31The University of Hong Kong

###### Abstract

Understanding and modeling human intelligence are parallel goals shared by artificial intelligence \(AI\) and cognitive science\. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models\. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials\. CogGym uses a semi\-automated, human\-in\-the\-loop pipeline to standardize diverse experimental paradigms into a task\-agnostic Experiment Markup Language \(EML\), enabling reproducible and faithful comparison at scale\. For initial release, we curate and standardize 258 cognitive experiments from 100 papers focusing on human commonsense reasoning, and evaluate 50 large language models against human responses\. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments\. Yet AI models’ improvement on such common reasoning tasks is considerably slower than the gains observed on formal\-reasoning benchmarks like math and coding, and model–human fit remains well below human splithalf reliability \(R2=0\.93R^\{2\}=0\.93on text,0\.950\.95on image, and0\.920\.92on video\) with the best models achievingR2=0\.59R^\{2\}=0\.59on text,0\.580\.58on image, and0\.430\.43on video experiments\. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve\.

## 1Introduction

Since its founding, the discipline of Artificial Intelligence \(AI\) has often focused on formalizing the computations that underlie human capabilities for thinking and reasoning\([Turing, 1950](https://arxiv.org/html/2609.21259#bib.bib64);[Newell and Simon, 1976](https://arxiv.org/html/2609.21259#bib.bib65);[Minsky, 1961](https://arxiv.org/html/2609.21259#bib.bib66)\)\. In the intervening decades, there have been numerous attempts to build or use AI systems as models of human thinking\([Marr, 1982](https://arxiv.org/html/2609.21259#bib.bib67);[Minsky, 1986](https://arxiv.org/html/2609.21259#bib.bib68);[Rumelhart et al\., 1986](https://arxiv.org/html/2609.21259#bib.bib69);[Tenenbaum et al\., 2011](https://arxiv.org/html/2609.21259#bib.bib29)\), yet the field still lacks a standardized framework for measuring how well today’s machines simulate human thought and behavior\. The rapid advancement of large\-scale neural architectures, most notably Large Language Models \(LLMs\) and their multimodal and reasoning successors\([Brown et al\., 2020](https://arxiv.org/html/2609.21259#bib.bib2);[OpenAI, 2023](https://arxiv.org/html/2609.21259#bib.bib31)\), has yielded systems capable of exhibiting remarkably human\-like cognitive behavior across capabilities such as language generation and logical reasoning\([Bubeck et al\., 2023](https://arxiv.org/html/2609.21259#bib.bib19);[Wei et al\., 2022](https://arxiv.org/html/2609.21259#bib.bib22)\)\. These behavioral similarities have motivated researchers to treat frontier models as computational hypotheses about the human mind\([Binz and Schulz, 2023b](https://arxiv.org/html/2609.21259#bib.bib25)\)\.

The prospect of building human\-like intelligence in machines also generated immense enthusiasm due to its transformative potential in science and other applications\([Ying et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib24)\)\. Highly accurate computational analogs of human cognition could serve as digital twins, enabling researchers to simulate behavioral phenomena at unprecedented scales and piloting experimental designs before deploying them in human populations\([Anthis et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib3);[Ashokkumar et al\., 2026](https://arxiv.org/html/2609.21259#bib.bib4);[Park et al\., 2022](https://arxiv.org/html/2609.21259#bib.bib5)\)\. From an applied perspective, human\-like machines could accelerate product development and user testing by simulating diverse human reactions\([Simile, 2026](https://arxiv.org/html/2609.21259#bib.bib6)\)\. Recent research has also suggested that endowing machines with human\-like cognitive architectures is one path towards safer, more effective, and more intuitive human\-AI interaction\([Ho and Griffiths, 2022](https://arxiv.org/html/2609.21259#bib.bib14);[Collins et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib15)\)\.

However, despite these technological advances and potential benefits, the degree to which contemporary AI models behaviorally align with humans remains poorly understood and fiercely debated\. Indeed, theorists of cognitive science have argued that fully replicating human behavior with an AI model is an inherently intractable problem\([van Rooij et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib60);[Rich et al\., 2021](https://arxiv.org/html/2609.21259#bib.bib61)\)\. These debates raise a more precise question: in what ways do model responses resemble human responses, and where do they systematically diverge? The prevailing paradigm in the AI community relies on standardized benchmarks that aggregate performance against objective answer keys\([Hendrycks et al\., 2021a](https://arxiv.org/html/2609.21259#bib.bib41);[Wang et al\., 2018](https://arxiv.org/html/2609.21259#bib.bib20);[Wang et al\., 2019](https://arxiv.org/html/2609.21259#bib.bib21);[Liang et al\., 2023](https://arxiv.org/html/2609.21259#bib.bib33)\)\(e\.g\., answering standardized test questions or solving programmatic puzzles\)\. Cognitive\-science experiments instead use theoretically motivated manipulations to compare graded response patterns across items, conditions, and participants; objective benchmark accuracy rarely captures structural nuances of human behavior, such as systematic biases\([Kahneman and Tversky, 1979](https://arxiv.org/html/2609.21259#bib.bib36)\), graded probabilistic judgments\([Griffiths and Tenenbaum, 2006](https://arxiv.org/html/2609.21259#bib.bib37)\), and the individual differences in responses across human participants\([Wong et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib50);[Ying et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib24);[Tan et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib70)\)\.

A natural pathway to evaluate intelligence in machines is to test AI models on the already large and continually expanding space of diverse cognitive science experiments, the very paradigms that have been developed over decades to characterize human cognition\. Indeed, within the cognitive sciences, there has been increasing interest in evaluating frontier models using classical psychological paradigms\([Binz and Schulz, 2023b](https://arxiv.org/html/2609.21259#bib.bib25);[Hagendorff et al\., 2023](https://arxiv.org/html/2609.21259#bib.bib26);[Dasgupta et al\., 2022](https://arxiv.org/html/2609.21259#bib.bib27);[Collins et al\., 2022](https://arxiv.org/html/2609.21259#bib.bib28);[Ullman, 2023](https://arxiv.org/html/2609.21259#bib.bib34);[Stojnić et al\., 2023](https://arxiv.org/html/2609.21259#bib.bib35);[Frank, 2023](https://arxiv.org/html/2609.21259#bib.bib32);[Liu et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib9)\)\. However, these efforts largely remain fragmented: individual studies often focus on isolated topics or phenomena using idiosyncratic experimental setups, resulting in a scattered literature that lacks systematicity\. The findings are also strikingly mixed: some studies report that models reproduce a wide range of human capabilities, biases, and classic experimental effects\([Binz and Schulz, 2023b](https://arxiv.org/html/2609.21259#bib.bib25);[Hagendorff et al\., 2023](https://arxiv.org/html/2609.21259#bib.bib26);[Cui et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib45);[Liu et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib12);[Zhu and Griffiths, 2024](https://arxiv.org/html/2609.21259#bib.bib1)\), whereas others find that they diverge substantially from human behavior\([Ullman, 2023](https://arxiv.org/html/2609.21259#bib.bib34);[Mitchell and Krakauer, 2023](https://arxiv.org/html/2609.21259#bib.bib42);[Schulze Buschoff et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib46);[Oh and Linzen, 2026](https://arxiv.org/html/2609.21259#bib.bib62);[Wang et al\., 2026](https://arxiv.org/html/2609.21259#bib.bib63)\), leaving the overall degree of alignment unresolved and the field polarized\. In addition, the continuous release of new models, alongside the rapid updating of existing ones, means that studies examining a small set of models at a fixed point in time yield conclusions with rapidly diminishing relevance\.

To address these limitations, recent efforts have begun to aggregate cognitive experiments into evaluation suites for LLMs\. However, standardizing hundreds of heterogeneous cognitive experiments in a common executable format poses a substantial technical challenge\. Such experiments are highly diverse in modality \(text, image, video\), are implemented with different software frameworks \(e\.g\., PsychoPy\([Peirce, 2007](https://arxiv.org/html/2609.21259#bib.bib16)\), psiTurk\([Gureckis et al\., 2016](https://arxiv.org/html/2609.21259#bib.bib8)\), jsPsych\([de Leeuw, 2015](https://arxiv.org/html/2609.21259#bib.bib23)\), MATLAB, and custom presentation platforms\), and many of the original repositories are outdated, incomplete, or poorly documented\. Recent attempts\([Binz et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib39);[Binz and Schulz, 2023a](https://arxiv.org/html/2609.21259#bib.bib38);[Hu et al\., 2026](https://arxiv.org/html/2609.21259#bib.bib44);[Lei et al\., 2026](https://arxiv.org/html/2609.21259#bib.bib47);[Coda\-Forno et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib7)\)rely on labor\-intensive manual standardization that is difficult to scale, and cover only text\-based tasks across a handful of topics, leaving the more complex image and video paradigms largely unaddressed\. In addition, because such efforts are largely crowdsourced or scraped, they often translate studies into prompts without systematically checking whether instructions, stimuli, trial structure, and response formats match the original experiments or whether published human response patterns can be replicated\. Mismatches introduced during prompt conversion can systematically distort model–human comparisons\([Wang et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib13), e\.g\.,\)\.

In this paper, we introduceCogGym, a scalable framework for systematically comparing AI systems with human behavior\. Unlike a static benchmark, CogGym is reusable scientific infrastructure for sourcing, standardizing, administering, and re\-running cognitive science experiments with current and future AI systems and new human participants\. It builds a scalable, semi\-automated pipeline that leverages AI agents with a human\-in\-the\-loop to convert diverse cognitive experiments into a standardized Experiment Markup Language \(EML\), a high\-level representational framework that abstracts the underlying logic of behavioral experiment paradigms \(experimental design, stimuli, trial structure, and response type\) into structured, standardized specifications\. Crucially, because EML fully specifies an experiment, it can be used both to rerun studies on human participants—filling in missing data, increasing participant numbers, or replicating previously reported results—and to administer the same experiments to AI models\. The common specification allows us to verify the standardized implementations against the original studies and new human data before asking where and how AI models diverge from humans\.

To populate the platform, we sourced experiments from over 30 research labs specializing in computational modeling of human cognition, curating a diverse collection of 258 experiments from 100 published papers\. We translated each experiment into EML using the CogGym pipeline and deployed this battery to evaluate 50 frontier and open\-weight large language models\. For an updated subset of 15 experiments, we compared human replication data and model responses against the original human data to assess data and experiment quality\.

CogGym’s coverage across experiments, topics, modalities, and response formats enables us to investigate fundamental questions about the behavioral alignment of AI systems:

1. 1\.Where do today’s frontier AI models behaviorally align with humans, and where do they diverge?How well do models match mean judgments, response distributions, and systematic patterns across experiments?
2. 2\.What model properties are associated with stronger behavioral alignment?Do larger or reasoning\-enabled models produce responses that more closely match human judgments?
3. 3\.On which tasks do models most and least align with humans?Across topics and stimulus modalities, where do models most closely align with human behavior, and where do they diverge most dramatically?

Our large\-scale evaluation shows that cognitive alignment does increase with both model size and recency, but the slope is far more gradual than the steep, near\-saturated gains these same models show on standardized formal\-reasoning benchmarks in STEM\([Rein et al\., 2023](https://arxiv.org/html/2609.21259#bib.bib48)\), mathematics\([Hendrycks et al\., 2021b](https://arxiv.org/html/2609.21259#bib.bib49)\), and coding\([Chen et al\., 2021](https://arxiv.org/html/2609.21259#bib.bib40)\)\. In addition, even the best aligned models still remain far from matching human behavioral responses, and model–human fit on video experiments remains markedly lower than on text and image experiments\.

![Refer to caption](https://arxiv.org/html/2609.21259v1/coggym_flow.png)Figure 1:CogGym workflow\.CogGym proceeds in three stages\. First, original experiments implemented in diverse frameworks are converted into a standardized Experiment Markup Language \(EML\) representation through an AI agent with human\-in\-the\-loop supervision\. Second, the standardized experiments are rendered and administered to human participants and AI models\. Lastly, human and model responses are analyzed jointly to evaluate experiment validity and model–human alignment\.
## 2CogGym

CogGym is a large\-scale, scalable evaluation framework for comparing machine and human behavior\. Instead of measuring whether models produce the objectively correct answer, CogGym examines whether models capture the behavior of human participants under the same experimental conditions\. It leverages a scalable pipeline to standardize diverse cognitive experiments into a common Experiment Markup Language \(EML\), which can be administered to AI models as well as new online human participants to measure how agents reason with uncertainty, graded judgments, and response variability across a broad range of cognitive tasks\. The experiment and dataset release, documentation, and information on contributing new experiments can be found at[https://coggym\.org](https://coggym.org/)\.

![Refer to caption](https://arxiv.org/html/2609.21259v1/domain_box_grid_examples_editable.png)Figure 2:Topics represented in CogGym\.CogGym experiments span a broad range of commonsense reasoning tasks, including physical reasoning, causal learning, concept learning, social cognition, moral judgment, probabilistic inference, pragmatic language use, and more\. The studies were sourced from over 30 contributing research labs\. These topics are derived from keywords reported in the source papers\.### Sourcing Cognitive Science Experiments

While CogGym supports comparisons between humans and machines across a wide range of tasks, we primarily focus on human commonsense reasoning: how people form and deploy intuitive theories of the physical and social world to solve everyday problems\. A core feature of these tasks is that many of them have no objective ground truth, such as who is to blame when multiple people played a causal role in a bad outcome\.

We collaborated with over30 research labsthat study human commonsense reasoning to curate the first batch of experiments on CogGym, containing258 experimentsfrom100 published papers\.[Figure 2](https://arxiv.org/html/2609.21259#S2.F2)summarizes the topics represented in the current experiment set, using keywords reported in the source papers, including theory of mind, causal reasoning, moral judgment, language understanding, physical reasoning, probabilistic reasoning, decision\-making, concept learning, emotion attribution, and perception\. The experiments feature text, images, and videos\. The present evaluation analyzes response content and distributions; response times are outside its scope\.

The CogGym workflow is divided into three parts—standardization, experiments, and analysis—which we illustrate in[Figure 1](https://arxiv.org/html/2609.21259#S1.F1)\. We describe each stage below\.

### Experiment Standardization

The first stage converts the original cognitive experiments into a common experiment representation\. CogGym begins with source materials from an original study, including experiment code, stimuli, instructions, trial logic, and human behavioral data\. These materials come from heterogeneous frameworks such as Qualtrics, jsPsych, PsychoPy, MATLAB, or custom JavaScript, making it impractical to use the raw source data to directly evaluate AI models\.

To address this, we develop a unifying Experiment Markup Language \(EML\): a model\-agnostic JSON representation of a behavioral experiment as executable data rather than framework\-specific code\. An EML experiment consists of several components: experiment configuration \(e\.g\., randomization logic, and trial order\), instructions \(information and tutorial presented to the participants\), trials \(stimuli and queries\), and human data\. These components are all specified in a shared JSON format, allowing for a common specification that covers a wide range of standard cognitive science experiments, and making them easily interpretable to researchers\. We provide an example in[Figure 3](https://arxiv.org/html/2609.21259#S2.F3)\. Detailed specifications and examples can be found in Appendix[A\.1](https://arxiv.org/html/2609.21259#A1.SS1)\. The EML can then be rendered to publish an experiment on the CogGym platform, which enables experiment visualization for researchers and human data collection\.

![Refer to caption](https://arxiv.org/html/2609.21259v1/EML.png)Figure 3:Example of an EML experiment\.A: EML specification for the experiment, which encodes the instructions, trials \(stimuli and response queries\), and human data of the original study as executable data \(adapted from[Gerstenberg and Goodman 2012](https://arxiv.org/html/2609.21259#bib.bib57)\)\.B: The same EML rendered on the CogGym platform as an interactive web experiment for human data collection\.C: The matched model prompt automatically constructed from the same EML\. Because the human task and the model prompt are generated from a single specification, participants and models receive closely matched instructions, stimuli, and response options, enabling a direct comparison of their responses\.To convert experiments to EML, we developed an agentic AI workflow to produce an initial standardized specification\. CogGym then renders the specification for a researcher to compare side by side with the original experiment\. The researcher checks whether the stimuli, instructions, and response widgets display the same information and ask the same questions as the original paradigm\. When discrepancies are found, the researcher provides feedback and the specification is revised by the AI agent\. This generation and human\-in\-the\-loop refinement process repeats until the CogGym experiment faithfully reproduces the original experiment as judged by the researcher\. On average, converting an experiment into EML takes approximately one hour\. Appendix[A\.2](https://arxiv.org/html/2609.21259#A1.SS2)describes the EML generation and validation process in detail\.

### Data Collection from Models and Humans

The second stage administers standardized experiments to humans and models for data collection\. For model evaluation, the same EML trial objects are passed to a model harness\. The harness converts each trial into a standardized prompt containing the experimental instructions, stimuli, and response options shown to human participants\. Model outputs are parsed to the same query tags used for human responses, allowing a shared format for model and human data for analyses\.

For human validation or replication, CogGym renders each EML specification as a web experiment and records new human responses under the same trial and query identifiers used by the original dataset\. This enables new human data to be compared directly against the source study, allowing the benchmark to test whether converted experiments preserve the intended behavioral signal\.

### Evaluating Model and Human Behavior

The final stage places the original human data, newly replicated human data, and model responses into a common trial\-level format, so all three are directly comparable\. We then evaluate how closely each model matches human behavior: we quantify the degree of model–human alignment, visualize it experiment by experiment, and inspect qualitative examples to see where models align or diverge from human behavior\. In parallel, we compare the original human data against the newly collected replication data to estimate the reliability of each experiment and its original data, excluding experiments that prove unreliable\.

## 3Experiments

We evaluate 50 proprietary and open models spanning multiple model families, sizes, and release dates from April 2024 to April 2026\. Each model is administered the full CogGym battery under matched conditions \(default configuration with identical prompts\), with 10 runs per model for each trial to obtain stable estimates\. The prompt and harness for the model experiments are shown in Appendix[B\.1](https://arxiv.org/html/2609.21259#A2.SS1)\.

For each experiment, we computeR2R^\{2\}, the squared Pearson correlation between human and model mean responses\. For ordinal responses \(e\.g\., slider ratings\), each item contributes one data point: the model’s mean response paired with the human mean\. For categorical responses \(e\.g\., multiple\-choice\), each option contributes a data point: the model’s selection probability \(proportion of runs selecting that option\) paired with the human selection proportion\. Within each experiment, we compute oneR2R^\{2\}for each answer type \(defined by response regime and native scale\), pooling conditions or constructs that share that type and scale; we average equally across answer types and then report the mean across experiments\. Higher values indicate closer agreement between model and human mean responses\.

To measure whether the model reproduces the full distribution of responses across participants, we additionally compute anormalized distribution divergencemetric\. For each item, we take the distance between the model and human response distributions—Earth Mover’s Distance \(EMD\) for continuous items \(normalized by the scale range\) and Jensen–Shannon Divergence \(JSD\) for categorical items—rescaled to\[0,1\]\[0,1\]\.00indicates the model matches exactly the human distribution\. Because repeatedly sampled model responses are sharply peaked and show little variation \(see Appendix[E\.2](https://arxiv.org/html/2609.21259#A5.SS2)\), we follow[Meister et al\. \(2025\)](https://arxiv.org/html/2609.21259#bib.bib43)and prompt models to provide a verbalized distribution over response options to compare against the human response distribution\. The prompts for querying the AI models are provided in Appendix[B\.1](https://arxiv.org/html/2609.21259#A2.SS1)\.

To obtain human reliability references for the two metrics, we estimate human split\-half reliability by randomly splitting participants into halves and computing the metrics, averaged over 100 splits\. We applied the Spearman–Brown correction to estimate full\-sample reliability\. The human split\-half is computed over 179 of the 258 experiments \(80 text, 37 image, 62 video; see Appendix[D\.2](https://arxiv.org/html/2609.21259#A4.SS2)\), where individual\-level response data exists and there is sufficient coverage \(at least 10 data points per judgment\)\.

## 4Results

Figure 4:Model performance on CogGym tasks and formal reasoning benchmarks plotted by model release date\.Models are color\-coded by size, and selected models are labeled\.Top:model–human fitR2R^\{2\}across experimental tasks\.Middle:normalized distributional divergence between AI models and humans\.Bottom:AI models’ performance on GPQA Diamond, MATH, and HumanEval performance\. The plots show that models’ improvement on formal reasoning benchmarks is much steeper than models’ improvement on their alignment with humans on commonsense reasoning tasks\. Error bars show±1\\pm 1standard error across experiments\.Figure 5:Model–human alignment as a function of model size among open\-weight models\.Each point is one open\-weight model\.Top row:model–human fitR2R^\{2\}\(↑\\uparrow\) rises with scale in every modality \(textr=\+0\.81r=\+0\.81, image\+0\.75\+0\.75, video\+0\.81\+0\.81\)\.Bottom row:normalized distribution divergence falls with scale in every modality \(textr=−0\.83r=\-0\.83, image−0\.69\-0\.69, video−0\.68\-0\.68\)\. Error bars are per\-model±1\\pm 1standard error\. Dashed extensions show log\-linear extrapolations\. Shaded regions are 95% confidence bands, shown more lightly over the extrapolated range\.Our results suggest that larger, more recently released models show higher behavioral alignment with humans\.[Figure 4](https://arxiv.org/html/2609.21259#S4.F4)shows that model–human fit on CogGym has risen substantially over the past two years\. On text and image experiments, the strongest models now achieve relatively high mean\-response fit \(bestR2=0\.59R^\{2\}=0\.59and0\.580\.58, respectively\), although both remain below Spearman–Brown\-corrected human referenceR2R^\{2\}values \(0\.930\.93and0\.950\.95\)\. Their normalized distributional divergence has also narrowed, reaching0\.100\.10on text and0\.150\.15on image, compared with human floors of0\.050\.05and0\.070\.07, respectively\. Video tasks remain substantially more challenging \(bestR2=0\.43R^\{2\}=0\.43vs\. a corrected human reference of0\.920\.92; divergence0\.250\.25vs\. a human floor of0\.070\.07\)\. Moreover, this progress has been considerably slower than the near\-saturation gains observed on formal\-reasoning benchmarks such as STEM \(GPQA Diamond\), math \(MATH\), and coding \(HumanEval\) over the same period \([Figure 4](https://arxiv.org/html/2609.21259#S4.F4)\)\.

[Figure 5](https://arxiv.org/html/2609.21259#S4.F5)plots per\-modelR2R^\{2\}against total parameter count for open\-weight models \(per\-modelR2R^\{2\}and normalized\-divergence scores by modality are reported in[Table 1](https://arxiv.org/html/2609.21259#A4.T1)in Appendix[D\.1](https://arxiv.org/html/2609.21259#A4.SS1)\)\. Model–human fitR2R^\{2\}rises with scale in every modality \(textr=\+0\.81r=\+0\.81, imager=\+0\.73r=\+0\.73, videor=\+0\.81r=\+0\.81\), although fewer open\-weight image and video models have valid scores\. The same trend holds for distributional divergence \([Figure 5](https://arxiv.org/html/2609.21259#S4.F5), bottom row\): divergence falls with scale in all three modalities \(textr=−0\.84r=\-0\.84, imager=−0\.79r=\-0\.79, videor=−0\.68r=\-0\.68\), with all remaining well above the human split\-half floor\.

While size and release date are positively correlated with alignment, they are also correlated with one another \(larger models tend to be more recent\)\. In order to estimate the independent effects of size and time, we applied a multivariable regression

y=βsize​log10⁡\(Nparams\)\+βtime​d\+ϵ,y=\\beta\_\{\\text\{size\}\}\\log\_\{10\}\(N\_\{\\text\{params\}\}\)\+\\beta\_\{\\text\{time\}\}\\,d\+\\epsilon,over the open\-weight models in[Table 1](https://arxiv.org/html/2609.21259#A4.T1)\. Across the open\-weight models, the scale coefficientβsize\\beta\_\{\\text\{size\}\}is positive forR2R^\{2\}in text, image, and video \(allp<\.001p<\.001\), and negative for distributional divergence \(textp<\.001p<\.001, imagep=\.004p=\.004, videop=\.036p=\.036\)\. After controlling for size, release date is not significantly associated with textR2R^\{2\}\(β^time=\+0\.010\\hat\{\\beta\}\_\{\\text\{time\}\}=\+0\.010/yr,p=\.637p=\.637\) or imageR2R^\{2\}\(\+0\.043\+0\.043/yr,p=\.205p=\.205\), but is positively associated with videoR2R^\{2\}\(\+0\.026\+0\.026/yr,p=\.012p=\.012\)\. Release date is not significantly associated with distributional divergence in text \(−0\.006\-0\.006/yr,p=\.501p=\.501\), image \(\+0\.006\+0\.006/yr,p=\.807p=\.807\), or video \(\+0\.008\+0\.008/yr,p=\.396p=\.396\)\. Taken together, these results suggest that release date is not consistently associated with model–human fit once model size is taken into account among the open\-weight models\.

Figure 6:Gemini 3\.1 Pro vs\. human split\-half, per experiment\.Each point is one experiment\.\(Top\)R2R^\{2\}:Human split\-halfR2R^\{2\}vs\. model–human fitR2R^\{2\};94%94\\%of experiments fall below the parity diagonal\.\(Bottom\) Distributional Divergence: human split\-half normalized divergence \(reliability floor\) vs\. model–human normalized divergence; Error bars in both panels are±1\\pm 1SE from 1,000\-sample item bootstraps\.Figure 7:Gemini 3\.1 Pro’s model–humanR2R^\{2\}by topic and modality\.Points show the median human–model fit across experiments; horizontal bars span the interquartile range \(25th–75th percentiles\)\. A cross marks a topic which lacks sufficient number of experiments\. The plots show that models tend to align less with humans on physical reasoning tasks than other topics\.### Model cognitive alignment varies among experiment modality

Across the full model set, mean per\-experiment explained variance decreases from text \(R2=\.39R^\{2\}=\.39\) to image \(\.29\.29\) to video \(\.14\.14\)\. A crossed mixed\-effects analysis of the 15 model groups evaluated on all three modalities, with random intercepts for model, study, and experiment and model\-specific modality contrasts, confirmed all pairwise differences \(text minus image:Δ​R2=\.13\\Delta R^\{2\}=\.13, 95% CI\[\.05,\.21\]\[\.05,\.21\]; text minus video:\.26\.26,\[\.17,\.34\]\[\.17,\.34\]; image minus video:\.13\.13,\[\.05,\.21\]\[\.05,\.21\]; all Holm\-adjustedp<\.002p<\.002\)\. These are adjusted differences in model–humanR2R^\{2\}, rather than differences between the pooled means\. The same ordering held in an analysis of the full model set, although the best current model performs similarly on text and image \(R2=\.59R^\{2\}=\.59and\.58\.58, respectively\)\.[Figure 6](https://arxiv.org/html/2609.21259#S4.F6)plots Gemini 3\.1 Pro fit against human split\-half reliability\. The model–humanR2R^\{2\}remains below this conservative human reliability reference for most experiments\. In both text and image, most experiments are far from zero and in many experiments models fit moderately to strongly with humans\. Video is much worse, with most experiments clustered between00and0\.30\.3\.

These differences may reflect perceptual limitations, context\-length strain as video experiments tend to require longer context and more tokens, or the greater demands of tracking agents and objects in dynamic scenes\.

### Model fit varies within and across topics

We use keywords reported in the 100 source papers to produce five common topic groups—Theory of Mind, Causal Reasoning, Moral Judgment, Language & Pragmatics, and Physical Reasoning—and summarize model performance within and across these groups\.[Figure 7](https://arxiv.org/html/2609.21259#S4.F7)shows the median and interquartile range of Gemini 3\.1 Pro’s experiment\-levelR2R^\{2\}by topic and modality\. For Gemini 3\.1 Pro, Physical Reasoning is the least aligned topic, with the lowest meanR2R^\{2\}in both image \(\.33\.33\) and video \(\.24\.24\) experiments\.

However, within individual topics, the same model can fit some experiments nearly perfectly and others near chance\. Large differences in model–human fit can even be observed across experiments within a single paper\. Thus, experiment\-level design and structure would be more predictive of alignment than broad topic labels\.

### Is alignment low because models are superhuman?

A natural concern is that divergence from human responses can reflect that models are more rational or competent than people\. For example, many models today perform coding or math tasks better than average human participants\([Quan et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib11);[OpenAI et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib10)\)\. We argue against this interpretation based on the following observations\.

Firstly, we intentionally sourced experimenets where human performance is close to optimality: among the 258 experiments, only 11 experiments have reported accuracy metrics in their paper\. On this subset of experiments, mean human accuracy \(0\.790\.79\) is not statistically different from Gemini 3\.1 Pro accuracy \(0\.730\.73; pairedtt\-test,p=\.46p=\.46;[Figure 8](https://arxiv.org/html/2609.21259#S4.F8)\)\. Within this subset, Gemini’s accuracy relative to humans’ is positively associated with its fit to human mean responses \(r=\.72r=\.72, 95% bootstrap CI\[\.19,\.93\]\[\.19,\.93\];[Figure 8](https://arxiv.org/html/2609.21259#S4.F8)\)\. Thus, we don’t observe any evidence for an inverse scaling law where superhuman models are less human\-like\.

Secondly, the rest of the experiments often elicit inherently subjective or normative judgments—how much responsibility or blame an agent deserves, how much one would pay someone to perform a task, whether an action is morally permissible, or how to interpret a sentence\. For such questions there is no ground\-truth answer key; we believe studying human\-likeness on these domains is particularly valuable\.

### Evaluating the replicability and validity of the experiments

Another concern is that low alignment could reflect noisy human data or errors introduced when translating an experiment into EML, rather than genuine model–human differences\. We randomly sampled a subset of 15 experiments and recruited 521 human participants on Prolific to complete the experiments on CogGym and then compare human replication data and Gemini 3\.1 Pro responses against the original human data \(See Appendix[F](https://arxiv.org/html/2609.21259#A6)\)\. Across experiments, replication means agree more closely with the original human means \(meanR2=0\.84R^\{2\}=0\.84\) than do Gemini 3\.1 Pro responses \(meanR2=0\.62R^\{2\}=0\.62\)\. These results indicate that the lowR2R^\{2\}observed between models and humans in these experiments is likely due to genuine model misalignment rather than experiment or data quality\.

Figure 8:Accuracy and model–human fit for Gemini 3\.1 Pro vs Human participants\.Each data point indicates one experiment\. \(A\) Mean human versus model accuracy for the 11 experiments that include accuracy metrics\. The dashed diagonal marks equal accuracy \(y=xy=x\); points above it favor the model and points below it favor humans\. Overall humans participants achieve higher mean accuracy than Gemini 3\.1 Pro\. \(B\) Differences between model and human accuracy versus model–humanR2R^\{2\}on the same experiments, which shows a positive correlation between Gemini 3\.1 Pro performance relative to humans’ and model–human fit\.
### Contamination analysis

To assess possible dataset contamination of the source studies in our evaluation, we asked AI models to identify the paper titles of published experiments from experimental materials without access to the Internet\. Gemini 3\.1 Pro achieved the highest exact\-title recognition rate \(29/175, 16\.57%\), followed by Gemini 3 Flash \(18/175, 10\.29%\) and Claude Opus 4\.7 \(5/175, 2\.86%\)\. Open\-weight models exhibited lower mean exact\-title recognition rates than proprietary models in this sample \(0\.11% versus 3\.26%\)\.

We next examined whether normalized Levenshtein similarity between the recalled and correct paper titles was associated with model–human fit \(R2R^\{2\}\)\. We find that the correlation is not significant for Gemini 3\.1 Pro \(r=0\.108r=0\.108,p=0\.184p=0\.184\), Gemini 3 Flash \(r=0\.122r=0\.122,p=0\.111p=0\.111\), or Claude Opus 4\.7 \(r=0\.101r=0\.101,p=0\.227p=0\.227\)\. These findings suggest that some models can recognize source studies, but recognition alone does not seem to affect the model’s fit to human data\.

## 5Discussion

In this paper we propose CogGym, which establishes a reusable, executable infrastructure for comparing language and multimodal models with human behavioral data across a large collection of cognitive science experiments\. Using this infrastructure, we find that model–human behavioral alignment remains well below human split\-half reliability across a wide range of commonsense reasoning experiments, even as models have become sharply more capable on formal benchmarks such as math, coding, and STEM\.

Why does behavioral alignment scale more slowly than performance on formal benchmarks? One plausible explanation is that contemporary training pipelines are particularly well matched to domains with verifiable answers\. Mathematics and coding provide abundant automatically generated examples, exact correctness signals, and opportunities for reinforcement learning, test\-time search, and automatic verification\. CogGym experiments instead require models to reproduce human graded judgments in inherently subjective domains such as moral reasoning\. Such tasks may be under\-represented in the model training process\. Although domains such as physical reasoning can have objective labels, they also require models to infer physical properties and dynamics from images or videos, rather than manipulate explicitly specified symbolic inputs\. Physical reasoning tasks in CogGym often also involve more complex, continuous visual scenes, whereas tasks in other domains more commonly use static images or discrete, gridworld\-like representations\. This difference may contribute to the lower model–human alignment observed in physical reasoning, alongside challenges in reasoning about physical dynamics\. Recent physics\-oriented visual question answering benchmarks, including PhysBench and PhyX, likewise report substantial limitations in vision\-language models’ physical understanding\([Chow et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib73);[Shen et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib74)\)\.

While CogGym provides a scalable platform for comparing model and human behavior, it has several limitations, each pointing to a direction for future work:

#### Continual expansion of experiment set\.

Our initial 258 experiments are a modest start towards building a large\-scale evaluation suite for comparing model and human cognition\. As a living benchmark, CogGym will continually incorporate new experiments from contributing labs and newly published studies, thus building towards a more comprehensive evaluation platform that covers all aspects of cognition\. Future work can also integrate automated experiment design\([Jagadish et al\., 2026](https://arxiv.org/html/2609.21259#bib.bib17);[Prystawski et al\., 2026](https://arxiv.org/html/2609.21259#bib.bib18)\)to further extend the experiment suite, especially on topics where we expect to see greater divergence between AI models and humans\.

#### Deeper analysis\.

Many of the experiments in CogGym were designed with carefully controlled factorial manipulations, parametric variations, and theoretically motivated condition contrasts that enable far more fine\-grained diagnoses of which specific aspects of cognition a model captures or fails to capture\. For example, a social cognition experiment might reveal that a model correctly infers goals but not costs, or a causal reasoning task might show sensitivity to a mechanism but not counterfactual structure\. Extracting these insights at scale requires more sophisticated automated analysis pipelines than the summary statistics we report here\. We expect that future work on CogGym will develop tools for deeper, experiment\-specific insights into the cognitive profile of AI systems and humans\. Such analyses could move beyond asking “how human\-like is this model?” toward more scientifically productive questions such as “in what ways is this model human\-like, and in what ways is it not?”

#### Interactive experiments\.

The current EML specification supports single\-turn, non\-interactive paradigms, including one\-shot economic games, but does not yet support repeated economic games, multi\-agent social interactions, or sequential decision\-making tasks where the stimulus depends on the participant’s previous responses\. Future work will extend the EML to account for richer task specifications, in order to test the performance of AI models against the wider universe of cognitive experiments\.

Going forward, we aim to grow CogGym as living, versioned scientific infrastructure that incorporates increasingly diverse and complex experiments, supports fresh human replications, and permits the same studies to be rerun as models change\. We hope that it provides a scalable framework for identifying where model behavior resembles human behavior, where it diverges, and which experimental or model features explain those patterns\.

## Acknowledgements

We are grateful to Verona Teo, Lara Kirfel, Xinyi Lu, and Chuqi Hu for their help in validating the implementations of their experiments and ensuring that they faithfully reflected the original studies\. We thank Jacob Andreas for his thoughtful and constructive feedback on the manuscript, and Charles Kemp for recommending studies from his lab for inclusion in CogGym\. We also thank Ryan Truong and Raphael Yamamoto for their valuable assistance throughout the project\. This work was supported by the MIT Quest for Intelligence\.

## References

- R\. Aboody, I\. Davis, Y\. Dunham, and J\. Jara\-EttingerPeople can infer the magnitude of other people’s knowledge even when they cannot infer its contents\.Cognition265,pp\. 106236\.External Links:[Document](https://dx.doi.org/10.1016/j.cognition.2025.106236)Cited by:[Figure 15](https://arxiv.org/html/2609.21259#A9.F15)\.
- Anthiset al\.\(2025\)J\. R\. Anthis, R\. Liu, S\. M\. Richardson, A\. C\. Kozlowski, B\. Koch, E\. Brynjolfsson, J\. Evans, and M\. S\. BernsteinPosition: LLM social simulations are a promising research method\.InForty\-second International Conference on Machine Learning Position Paper Track,Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p2.1)\.
- Ashokkumaret al\.\(2026\)A\. Ashokkumar, L\. Hewitt, I\. Ghezae, and R\. WillerLarge language models can predict the results of social science experiments\.Nature656\(8126\),pp\. 115–122\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p2.1)\.
- Bigelow and Piantadosi \(2016\)E\. Bigelow and S\. T\. PiantadosiA large dataset of generalization patterns in the number game\.Journal of Open Psychology Data4\(1\),pp\. 4\.External Links:[Document](https://dx.doi.org/10.5334/jopd.19)Cited by:[Figure 12](https://arxiv.org/html/2609.21259#A9.F12)\.
- Binzet al\.\(2024\)M\. Binz, E\. Akata, M\. Bethge, F\. Brändle, F\. Callaway, J\. Coda\-Forno, P\. Dayan, T\. L\. Griffiths, E\. Schulz,et al\.Centaur: a foundation model of human cognition\.arXiv preprint arXiv:2410\.20268\.Cited by:[Table 1](https://arxiv.org/html/2609.21259#A4.T1.2.61.1.1.1),[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Binz and Schulz \(2023a\)M\. Binz and E\. SchulzTurning large language models into cognitive models\.arXiv preprint arXiv:2306\.03917\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Binz and Schulz \(2023b\)M\. Binz and E\. SchulzUsing cognitive psychology to understand GPT\-3\.Proceedings of the National Academy of Sciences120\(6\),pp\. e2218523120\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1),[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Brownet al\.\(2020\)T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.Advances in Neural Information Processing Systems33,pp\. 1877–1901\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Bubecket al\.\(2023\)S\. Bubeck, V\. Chandrasekaran, R\. Eldan, J\. Gehrke, E\. Horvitz, E\. Kamar, P\. Lee, Y\. T\. Lee, Y\. Li, S\. Lundberg,et al\.Sparks of artificial general intelligence: Early experiments with GPT\-4\.ArXiv\.Note:arXiv:2303\.12712Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. d\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman,et al\.Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p10.1)\.
- Chowet al\.\(2025\)W\. Chow, J\. Mao, B\. Li, D\. Seita, V\. Guizilini, and Y\. WangPhysBench: benchmarking and enhancing vision\-language models for physical world understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2501.16411)Cited by:[§5](https://arxiv.org/html/2609.21259#S5.p2.1)\.
- Coda\-Fornoet al\.\(2024\)J\. Coda\-Forno, M\. Binz, J\. X\. Wang, and E\. SchulzCogBench: a large language model walks into a psychology lab\.InProceedings of the 41st International Conference on Machine Learning,pp\. 9076–9108\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Collinset al\.\(2024\)K\. M\. Collins, I\. Sucholutsky, U\. Bhatt, K\. Chandra, L\. Wong, M\. Lee, C\. E\. Zhang, T\. Zhi\-Xuan, M\. Ho, V\. Mansinghka,et al\.Building machines that learn and think with people\.Nature Human Behaviour8\(10\),pp\. 1851–1863\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p2.1)\.
- Collinset al\.\(2022\)K\. M\. Collins, C\. Wong, J\. Feng, M\. Wei, and J\. B\. TenenbaumStructured, flexible, and robust: benchmarking and improving large language models towards more human\-like behavior in out\-of\-distribution reasoning tasks\.arXiv preprint arXiv:2205\.05718\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Cuiet al\.\(2025\)Z\. Cui, N\. Li, and H\. ZhouA large\-scale replication of scenario\-based experiments in psychology and management using large language models\.Nature Computational Science5,pp\. 627–634\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Dasguptaet al\.\(2022\)I\. Dasgupta, A\. K\. Lampinen, S\. C\. Y\. Chan, A\. Creswell, D\. Kumaran, J\. L\. McClelland, and F\. HillLanguage models show human\-like content effects on reasoning\.arXiv preprint arXiv:2207\.07051\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- de Leeuw \(2015\)J\. R\. de LeeuwjsPsych: A JavaScript library for creating behavioral experiments in a Web browser\.Behavior Research Methods47\(1\),pp\. 1–12\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Frank \(2023\)M\. C\. FrankBaby steps in evaluating the capacities of large language models\.Nature Reviews Psychology2,pp\. 451–452\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Gerstenberget al\.\(2021\)T\. Gerstenberg, N\. D\. Goodman, D\. A\. Lagnado, and J\. B\. TenenbaumA counterfactual simulation model of causal judgments for physical events\.Psychological Review128\(5\),pp\. 936–975\.Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Gerstenberg and Goodman \(2012\)T\. Gerstenberg and N\. D\. GoodmanPing Pong in Church: productive use of concepts in human probabilistic inference\.InProceedings of the 34th Annual Conference of the Cognitive Science Society,pp\. 1590–1595\.Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1),[Figure 3](https://arxiv.org/html/2609.21259#S2.F3)\.
- Griffiths and Tenenbaum \(2006\)T\. L\. Griffiths and J\. B\. TenenbaumOptimal predictions in everyday cognition\.Psychological Science17\(9\),pp\. 767–773\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Gureckiset al\.\(2016\)T\. M\. Gureckis, J\. Martin, J\. McDonnell, A\. S\. Rich, D\. Markant, A\. Coenen, D\. Halpern, J\. B\. Hamrick, and P\. ChanpsiTurk: An open\-source framework for conducting replicable behavioral experiments online\.Behavior Research Methods48\(3\),pp\. 829–842\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Hagendorffet al\.\(2023\)T\. Hagendorff, S\. Fabi, and M\. KosinskiHuman\-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT\.Nature Computational Science3\(10\),pp\. 833–838\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Hendryckset al\.\(2021a\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Hendryckset al\.\(2021b\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the MATH dataset\.arXiv preprint arXiv:2103\.03874\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p10.1)\.
- Ho and Griffiths \(2022\)M\. K\. Ho and T\. L\. GriffithsCognitive science as a source of forward and inverse models of human decisions for robotics and control\.Annual Review of Control, Robotics, and Autonomous Systems5\(1\),pp\. 33–53\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p2.1)\.
- Houlihanet al\.\(2023\)S\. D\. Houlihan, M\. Kleiman\-Weiner, L\. B\. Hewitt, J\. B\. Tenenbaum, and R\. SaxeEmotion prediction as computation over a generative theory of mind\.Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences381\(2251\)\.External Links:ISSN 1471\-2962,[Link](http://dx.doi.org/10.1098/rsta.2022.0047),[Document](https://dx.doi.org/10.1098/rsta.2022.0047)Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Huet al\.\(2026\)T\. Hu, J\. Baumann, L\. Lupo, N\. Collier, D\. Hovy, and P\. RöttgerSimBench: benchmarking the ability of large language models to simulate human behaviors\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Jagadishet al\.\(2026\)A\. K\. Jagadish, Y\. Strittmatter, N\. Jacoby, G\. Kachergis, E\. Schulz, N\. Daw, S\. H\. Chandramouli, and T\. L\. GriffithsClosing the loop to discover psychological theories with an automated cognitive scientist\.arXiv preprint arXiv:2606\.26448\.Cited by:[§5](https://arxiv.org/html/2609.21259#S5.SS0.SSS0.Px1.p1.1)\.
- Jara\-Ettingeret al\.\(2020\)J\. Jara\-Ettinger, L\. E\. Schulz, and J\. B\. TenenbaumThe naïve utility calculus as a unified, quantitative framework for action understanding\.Cognitive Psychology123,pp\. 101334\.External Links:ISSN 0010\-0285,[Link](http://dx.doi.org/10.1016/j.cogpsych.2020.101334),[Document](https://dx.doi.org/10.1016/j.cogpsych.2020.101334)Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Kahneman and Tversky \(1979\)D\. Kahneman and A\. TverskyProspect theory: An analysis of decision under risk\.Econometrica47\(2\),pp\. 263–291\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Kwonet al\.\(2022\)J\. Kwon, J\. B\. Tenenbaum, and S\. LevineFlexibility in moral cognition: when is it okay to break the rules?\.InProceedings of the Annual Conference of the Cognitive Science Society,Vol\.44\.Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Leiet al\.\(2026\)Y\. Lei, T\. Wang, J\. Lian, Z\. Hu, D\. Lian, and X\. XieHumanLLM: towards personalized understanding and simulation of human nature\.InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining \(KDD\),Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Lianget al\.\(2023\)P\. Liang, R\. Bommasani, T\. Lee, D\. Tsipras, D\. Soylu, M\. Yasunaga, Y\. Zhang, D\. Narayanan, Y\. Wu, A\. Kumar,et al\.Holistic evaluation of language models\.Transactions on Machine Learning Research\.Note:arXiv:2211\.09110Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Liuet al\.\(2025\)R\. Liu, J\. Geng, A\. J\. Wu, I\. Sucholutsky, T\. Lombrozo, and T\. L\. GriffithsMind your step \(by step\): chain\-of\-thought can reduce performance on tasks where thinking makes humans worse\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 38489–38517\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Liuet al\.\(2024\)R\. Liu, T\. R\. Sumers, I\. Dasgupta, and T\. L\. GriffithsHow do Large Language Models navigate conflicts between honesty and helpfulness?\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 31844–31865\.External Links:[Link](https://proceedings.mlr.press/v235/liu24bb.html)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Marr \(1982\)D\. MarrVision: a computational investigation into the human representation and processing of visual information\.W\. H\. Freeman,San Francisco\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Meisteret al\.\(2025\)N\. Meister, C\. Guestrin, and T\. HashimotoBenchmarking Distributional Alignment of Large Language Models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Albuquerque, New Mexico\.Cited by:[§B\.1](https://arxiv.org/html/2609.21259#A2.SS1.SSS0.Px3.p1.1),[§E\.2](https://arxiv.org/html/2609.21259#A5.SS2.p1.1),[§3](https://arxiv.org/html/2609.21259#S3.p3.1)\.
- Minsky \(1961\)M\. L\. MinskySteps toward artificial intelligence\.Proceedings of the IRE49\(1\),pp\. 8–30\.External Links:[Document](https://dx.doi.org/10.1109/JRPROC.1961.287775)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Minsky \(1986\)M\. MinskyThe society of mind\.Simon & Schuster,New York\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Mitchell and Krakauer \(2023\)M\. Mitchell and D\. C\. KrakauerThe debate over understanding in AI’s large language models\.Proceedings of the National Academy of Sciences120\(13\),pp\. e2215907120\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Newell and Simon \(1976\)A\. Newell and H\. A\. SimonComputer science as empirical inquiry: symbols and search\.Communications of the ACM19\(3\),pp\. 113–126\.External Links:[Document](https://dx.doi.org/10.1145/360018.360022)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Oh and Linzen \(2026\)B\. Oh and T\. LinzenTo model human linguistic prediction, make LLMs less superhuman\.Trends in Cognitive Sciences\.External Links:[Document](https://dx.doi.org/10.1016/j.tics.2026.05.008),[Link](https://doi.org/10.1016/j.tics.2026.05.008)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- OpenAIet al\.\(2025\)OpenAI, A\. El\-Kishky, A\. Wei, A\. Saraiva, B\. Minaiev, D\. Selsam, D\. Dohan, F\. Song, H\. Lightman, I\. Clavera, J\. Pachocki,et al\.Competitive programming with large reasoning models\.External Links:2502\.06807,[Link](https://arxiv.org/abs/2502.06807)Cited by:[§4](https://arxiv.org/html/2609.21259#S4.SSx3.p1.1)\.
- OpenAI \(2023\)OpenAIGPT\-4 technical report\.ArXiv\.Note:arXiv:2303\.08774Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Parket al\.\(2022\)J\. S\. Park, L\. Popowski, C\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinSocial simulacra: creating populated prototypes for social computing systems\.InProceedings of the 35th annual ACM symposium on user interface software and technology,pp\. 1–18\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p2.1)\.
- Peirce \(2007\)J\. W\. PeircePsychoPy—Psychophysics software in Python\.Journal of Neuroscience Methods162\(1\-2\),pp\. 8–13\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Piantadosiet al\.\(2016\)S\. T\. Piantadosi, J\. B\. Tenenbaum, and N\. D\. GoodmanThe logical primitives of thought: empirical foundations for compositional cognitive models\.\.Psychological Review123\(4\),pp\. 392–424\.External Links:ISSN 0033\-295X,[Link](http://dx.doi.org/10.1037/a0039980),[Document](https://dx.doi.org/10.1037/a0039980)Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Prystawskiet al\.\(2026\)B\. Prystawski, K\. Mukherjee, D\. Wurgaft, L\. Nasvytis, M\. Y\. Li, N\. D\. Goodman, and M\. C\. FrankAuto\-psych: automating the science of mind using agent\-driven theory discovery and experimentation\.arXiv preprint arXiv:2606\.26460\.Cited by:[§5](https://arxiv.org/html/2609.21259#S5.SS0.SSS0.Px1.p1.1)\.
- Quanet al\.\(2025\)S\. Quan, J\. Yang, B\. Yu, B\. Zheng, D\. Liu, A\. Yang, X\. Ren, B\. Gao, Y\. Miao, Y\. Feng,et al\.CodeElo: Benchmarking Competition\-level Code Generation of LLMs with Human\-comparable Elo Ratings\.External Links:2501\.01257,[Link](https://arxiv.org/abs/2501.01257)Cited by:[§4](https://arxiv.org/html/2609.21259#S4.SSx3.p1.1)\.
- Reinet al\.\(2023\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGPQA: A Graduate\-Level Google\-Proof Q&A Benchmark\.arXiv preprint arXiv:2311\.12022\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p10.1)\.
- Richet al\.\(2021\)P\. Rich, R\. de Haan, T\. Wareham, and I\. van RooijHow hard is cognitive science?\.InProceedings of the Annual Conference of the Cognitive Science Society,Vol\.4343\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Rumelhartet al\.\(1986\)D\. E\. Rumelhart, J\. L\. McClelland, and PDP Research GroupParallel distributed processing: explorations in the microstructure of cognition, volume 1: foundations\.MIT Press,Cambridge, MA\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Schulze Buschoffet al\.\(2025\)L\. M\. Schulze Buschoff, E\. Akata, M\. Bethge, and E\. SchulzVisual cognition in multimodal large language models\.Nature Machine Intelligence7,pp\. 96–106\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Shenet al\.\(2025\)H\. Shen, T\. Wu, Q\. Han, Y\. Hsieh, J\. Wang, Y\. Zhang, Y\. Cheng, Z\. Hao, Y\. Ni, X\. Wang, Z\. Wan, K\. Zhang, W\. Xu, J\. Xiong, P\. Luo, W\. Chen, C\. Tao, Z\. Mao, and N\. WongPhyX: does your model have the "wits" for physical reasoning?\.arXiv preprint arXiv:2505\.15929\.External Links:[Link](https://arxiv.org/abs/2505.15929)Cited by:[§5](https://arxiv.org/html/2609.21259#S5.p2.1)\.
- Simile \(2026\)SimileCVS Health x Simile: Simulations for Faster, Safer Decisions\.Note:Simile AI BlogAccessed: 2026\-07\-13External Links:[Link](https://simile.ai/blog/simile-cvs-health)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p2.1)\.
- Stojnićet al\.\(2023\)G\. Stojnić, K\. Gandhi, S\. Yasuda, B\. M\. Lake, and M\. R\. DillonCommonsense psychology in human infants and machines\.Cognition235,pp\. 105406\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Tanet al\.\(2024\)A\. W\. M\. Tan, S\. Yu, B\. Long, W\. A\. Ma, T\. Murray, R\. D\. Silverman, J\. D\. Yeatman, and M\. C\. FrankDevBench: a multimodal developmental benchmark for language learning\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Tenenbaumet al\.\(2011\)J\. B\. Tenenbaum, C\. Kemp, T\. L\. Griffiths, and N\. D\. GoodmanHow to grow a mind: Statistics, structure, and abstraction\.Science331\(6022\),pp\. 1279–1285\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Tsvilodubet al\.\(2025\)P\. Tsvilodub, K\. Gandhi, H\. Zhao, J\. Fränken, M\. Franke, and N\. D\. GoodmanNon\-literal understanding of number words by language models\.arXiv preprint arXiv:2502\.06204\.Cited by:[Figure 13](https://arxiv.org/html/2609.21259#A9.F13),[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Turing \(1950\)A\. M\. TuringComputing machinery and intelligence\.Mind59\(236\),pp\. 433–460\.External Links:[Document](https://dx.doi.org/10.1093/mind/LIX.236.433)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Ullman \(2023\)T\. UllmanLarge language models fail on trivial alterations to theory\-of\-mind tasks\.arXiv preprint arXiv:2302\.08399\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- van Rooijet al\.\(2024\)I\. van Rooij, O\. Guest, F\. Adolfi, R\. de Haan, A\. Kolokolova, and P\. RichReclaiming AI as a theoretical tool for cognitive science\.Computational Brain & Behavior7\(4\),pp\. 616–636\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Wanget al\.\(2019\)A\. Wang, Y\. Pruksachatkun, N\. Nangia, A\. Singh, J\. Michael, F\. Hill, O\. Levy, and S\. R\. BowmanSuperGLUE: A stickier benchmark for general\-purpose language understanding systems\.Advances in Neural Information Processing Systems32\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Wanget al\.\(2018\)A\. Wang, A\. Singh, J\. Michael, F\. Hill, O\. Levy, and S\. R\. BowmanGLUE: A Multi\-Task Benchmark and Analysis Platform for Natural Language Understanding\.arXiv preprint arXiv:1804\.07461\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Wanget al\.\(2024\)P\. Wang, L\. Li, L\. Chen, Z\. Cai, D\. Zhu, B\. Lin, Y\. Cao, L\. Kong, Q\. Liu, T\. Liu,et al\.Large language models are not fair evaluators\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9440–9450\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p5.1)\.
- Wanget al\.\(2026\)Q\. Wang, N\. Tomlin, M\. Hu, B\. Dillon, and T\. LinzenSimulating human memory with language models\.arXiv preprint arXiv:2605\.25680\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.25680),[Link](https://arxiv.org/abs/2605.25680)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Weiet al\.\(2022\)J\. Wei, Y\. Tay, R\. Bommasani, C\. Raffel, B\. Zoph, S\. Borgeaud, D\. Yogatama, M\. Bosma, D\. Zhou, D\. Metzler, E\. H\. Chi, T\. Hashimoto, O\. Vinyals, P\. Liang, J\. Dean, and W\. FedusEmergent abilities of large language models\.Transactions on Machine Learning Research\.Note:arXiv:2206\.07682Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p1.1)\.
- Wonget al\.\(2025\)L\. Wong, K\. M\. Collins, L\. Ying, C\. E\. Zhang, A\. Weller, T\. Gerstenberg, T\. O’Donnell, A\. K\. Lew, J\. D\. Andreas, J\. B\. Tenenbaum, and T\. Brooke\-WilsonModeling open\-world cognition as on\-demand synthesis of probabilistic models\.InProceedings of the Annual Conference of the Cognitive Science Society,Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Yildirimet al\.\(2024\)I\. Yildirim, M\. H\. Siegel, A\. A\. Soltani, S\. Ray Chaudhuri, and J\. B\. TenenbaumPerception of 3D shape integrates intuitive physics and analysis\-by\-synthesis\.Nature Human Behaviour8\(2\),pp\. 320–335\.External Links:ISSN 2397\-3374,[Link](http://dx.doi.org/10.1038/s41562-023-01759-7),[Document](https://dx.doi.org/10.1038/s41562-023-01759-7)Cited by:[Figure 14](https://arxiv.org/html/2609.21259#A9.F14),[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Yinget al\.\(2025\)L\. Ying, K\. M\. Collins, L\. Wong, I\. Sucholutsky, R\. Liu, A\. Weller, T\. Shu, T\. L\. Griffiths, and J\. B\. TenenbaumOn Benchmarking Human\-Like Intelligence in Machines\.arXiv preprint arXiv:2502\.20502\.Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p2.1),[§1](https://arxiv.org/html/2609.21259#S1.p3.1)\.
- Zhouet al\.\(2023\)L\. Zhou, K\. A\. Smith, J\. B\. Tenenbaum, and T\. GerstenbergMental Jenga: A counterfactual simulation model of causal judgments about physical support\.\.Journal of Experimental Psychology: General152\(8\),pp\. 2237–2269\.External Links:ISSN 0096\-3445,[Link](http://dx.doi.org/10.1037/xge0001392),[Document](https://dx.doi.org/10.1037/xge0001392)Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.
- Zhu and Griffiths \(2024\)J\. Zhu and T\. L\. GriffithsIncoherent probability judgments in large language models\.InProceedings of the 46th Annual Conference of the Cognitive Science Society,External Links:[Link](https://escholarship.org/uc/item/2r68p98c)Cited by:[§1](https://arxiv.org/html/2609.21259#S1.p4.1)\.
- Zhuet al\.\(2025\)J\. Zhu, J\. C\. Peterson, B\. Enke, and T\. L\. GriffithsCapturing the complexity of human strategic decision\-making with machine learning\.Nature Human Behaviour9\(10\),pp\. 2114–2120\.External Links:ISSN 2397\-3374,[Link](http://dx.doi.org/10.1038/s41562-025-02230-5),[Document](https://dx.doi.org/10.1038/s41562-025-02230-5)Cited by:[Appendix I](https://arxiv.org/html/2609.21259#A9.p2.1)\.

## Appendix

## Appendix AExperiment Specification and Validation

### A\.1Experiment Markup Language \(EML\)

EML is the interchange format that lets CogGym render heterogeneous cognitive experiments and administer the same trial logic to humans and models\. The key design choice is to represent each experiment as data: flow, stimuli, response widgets, and human responses are explicit JSON objects with shared identifiers\. This makes the experiment portable across web rendering, model prompting, and scoring\.

#### Directory layout\.

A complete EML experiment lives in a study subdirectory such asStudyName/exp1/:

StudyName/

README\.md

exp1/

README\.md

config\.json

instruction\.jsonl

trial\.jsonl

human\_data\_mean\.json

human\_data\_ind\.json

assets/

stimulus\_001\.png

#### 1\. Experiment metadata and flow\.

Theconfig\.jsonfile defines citation metadata, response types, and the sequence of modules and trials\. TheexperimentFlowfield is the bridge between the experiment\-level design and the line\-oriented instruction/trial files: every ID listed inblocksmust appear in eitherinstruction\.jsonlortrial\.jsonl\.

\{

"experimentName":"Examplecausaljudgmenttask",

"description":"Participantsviewasceneandratehowmuchoneeventcausedanother\.",

"paperDOI":"https://doi\.org/10\.1234/example",

"taskType":\["CausalReasoning"\],

"responseType":\["single\-slider","multi\-choice"\],

"contributors":\["CogGymTeam"\],

"stimuli\_count":1,

"experimentFlow":\[

\{

"experimental\_condition":"standard",

"blocks":\[

\["instruction\_01","quiz\_01"\],

\["trial\_001"\]

\]

\}

\]

\}

#### 2\. Instructions and checks\.

Theinstruction\.jsonlfile contains one JSON object per line\. Modules can be plain instructions, practice trials, or comprehension quizzes\. These modules are referenced by ID fromconfig\.json\.

\{"id":"instruction\_01","type":"instruction","text":"<p\>Youwillseeasceneandanswerquestionsaboutwhatcausedtheoutcome\.</p\>"\}

\{"id":"quiz\_01","type":"comprehension\_quiz","text":"Pleaseconfirmthetask\.","queries":\[\{"prompt":"Whatshouldyoujudge?","type":"multi\-choice","tag":"quiz\_task","option":\["Color","Cause","Memory"\],"answer":1\}\]\}

#### 3\. Trials\.

Thetrial\.jsonlfile also contains one JSON object per line\. Each trial has anid, astimuliarray describing what is shown, and aqueriesarray describing what responses are collected\. Querytagvalues become the keys used for scoring\.

\{"id":"trial\_001","stimuli":\[\{"input\_type":"img","media\_url":\["assets/stimulus\_001\.png"\]\},\{"input\_type":"text","text":"Aballhitsaswitch,andthemachineturnson\."\}\],"queries":\[\{"prompt":"Howmuchdidtheballcausethemachinetoturnon?","type":"single\-slider","tag":"cause\_rating","slider\_config":\{"min":0,"max":100,"default\_value":50,"labels":\[\{"value":0,"label":"Notatall"\},\{"value":100,"label":"Completely"\}\],"show\_label\_values":true\},"required":true\},\{"prompt":"Whicheventwasmostresponsible?","type":"multi\-choice","tag":"cause\_choice","option":\["Theball","Theswitch","Themachine"\],"randomize\_order":false,"required":true\}\]\}

EML supportstext,img, andvideostimuli\. It supports the response formats needed for the benchmark:single\-slider,multi\-slider,multi\-choice,multi\-select,ranking, and free\-texttextboxresponses\. Atext\-instructionquery can also insert explanatory text inside a trial without collecting a response\.

#### 4\. Human responses\.

Human data files use the same trial IDs and query tags\. The mean file stores aggregate responses used for correlation, while the individual file stores participant\-level responses used for split\-half reliability\.

\{

"participants\_info":\{"count":40,"age":31\.2,"gender":\{"woman":21,"man":18,"other":1\}\},

"trial\_001":\{

"cause\_rating":72\.4,

"cause\_choice":\{"Theball":0\.70,"Theswitch":0\.25,"Themachine":0\.05\}

\}

\}

#### How EML executes\.

CogGym executes EML in four steps\. First, the renderer readsconfig\.jsonand samples an experimental condition if multiple flows are present\. Second, it resolves each block into instruction, quiz, practice, and trial objects using their IDs\. Third, it renders each trial by placing thestimulion the stimulus side of the interface and thequeriesas response widgets with the specified options or scale bounds\. Fourth, the model harness converts the same trial object into a standardized prompt and parses the model response back into the query tags\. Because the rendered experiment, model prompt, and human data all share the same trial IDs and tags, model and human responses can be aligned without experiment\-specific scoring code\.

### A\.2EML Generation and Validation

To generate EMLs, we first construct Meta\-EML \(MEML\), a specification layer for generating experiments in EML\. A MEML specification describes an experiment’s conditions, trial structure, randomization logic, instructions, stimuli, and response formats\. MEML compiles this specification into the EML files used by CogGym to render the experiment and construct model prompts\. This separates the experimental design from the detailed JSON representation required to execute it\. The full specification and workflow are available in the[MEML repository](https://coggym.org/doc)\.

We develop an agentic workflow for generating MEML specifications from source papers and original experimental materials\. Thegenerate\-memlskill guides agents to assemble stimulus, query, instruction, and citation libraries and encode the experimental design as a MEML specification\. It then directs agents to compile the specification into EML, validate the output against the EML schemas, verify source\-derived design constraints, and check reproducibility across repeated compilations\. Source evidence, adaptations, and unresolved design choices are documented for human review; successful compilation alone does not establish fidelity to the original experiment\.

Once EMLs are generated, we used an agentic workflow and human inspection to validate EML experiments\. We build in two layers of robust checks\.

#### Layer 1: Agentic generation and check\.

We developed an agentic workflow, which compares the MEML specification with the published paper, supplementary materials, original experiment code or survey instruments when available, and released data\. The comparison followed a fixed protocol covering ten dimensions: condition assignment, factor structure, trial counts, block structure, ordering and counterbalancing, stimulus content and modality, response type and scale semantics, instructions and cover story, practice/comprehension/feedback modules, and timing or presentation constraints\. Agents recorded exact source quotations and locations, confirmations as well as discrepancies, unavailable evidence, and whether a discrepancy could plausibly change participant behavior\. Findings and proposed corrections were treated as provisional until human review\.

#### Layer 2: Human experiment audit\.

The human audit had two components\. First, the CogGym team conducted an internal audit of the reconstructed experiments\. Second, we sought additional review from the original study authors\.

Both the researchers conducting the internal audit and the original study authors were asked the same three questions:

1. 1\.Are the instructions and framing identical to, or close enough to, the original study?
2. 2\.If CogGym presents the stimuli or response options differently from the original study, could this affect participant behavior?
3. 3\.Does the experimental flow or assignment of trials to conditions faithfully match the original study?

We then fix the MEML/EML based on the review and feedback\.

## Appendix BModel Prompting and Administration

### B\.1Prompt Construction and Model Harness

The CogGym harness translates each EML trial into a standardized prompt that can be administered to any language model\. For each trial, the prompt is constructed as follows:

#### System prompt\.

A fixed preamble instructs the model to behave as an experiment participant:

Youareaparticipantinacognitivescienceexperiment\.

Followtheinstructionscarefullyandrespondtoeach

question\.ReturnonlytherequestedJSONobject\-\-donotaddtextoutsidetheobject,

commentary,orcaveats\.IMPORTANT:YouMUSTrespondin

validJSONformat\.

#### User message\.

The user message concatenates three components:

1. 1\.Experiment instructions: The condition\-specific instructions frominstruction\.jsonl, with HTML stripped to plain text\. Any instruction images are included inline as base64\-encoded content\.
2. 2\.Trial stimuli: Text descriptions and/or media \(images encoded as base64, videos passed natively as base64 or cloud\-storage references\) from the trial’sstimuliarray\.
3. 3\.Queries: Each query is formatted with its prompt text, response options \(for choice types\) or scale endpoints \(for sliders\), and a JSON response template specifying the expected format\.

For example, a single\-slider query is presented as:

Question1:HowmuchdidAcausetheoutcome?

Scale:0=Notatall,100=Completely

RespondinJSON:\{"answer":<numberbetween0and100\>,

"explanation":"<briefreason\>"\}

A multi\-choice query:

Question2:Whatwilltheagentdonext?

SelectexactlyONEofthefollowingoptions:

A\)Goleft

B\)Goright

C\)Stay

RespondinJSON:\{"answer":"<selectedoptionletter\>",

"explanation":"<briefreason\>"\}

#### Verbalized\-distribution prompt\.

To elicit a full response distribution rather than a single answer \([section 3](https://arxiv.org/html/2609.21259#S3)\), the harness keeps the same system prompt, instructions, and stimuli, but replaces the single\-answer response template above with a distribution instruction, following[Meister et al\. \[2025\]](https://arxiv.org/html/2609.21259#bib.bib43):

===DISTRIBUTIONRESPONSETASK===

Donotanswerasoneparticipant\.Instead,estimatethedistribution

ofresponsesthatagroupofhumanparticipantswouldgive\.

ReturnonlyvalidJSONwiththisexacttop\-levelshape:

\{"responses":\{"<query\_tag\>":<distribution\_object\>\}\}

Includeonlythequerytagslistedbelow\.Probabilitiesmustbenumbers

between0and1andshouldsumto1foreachdistribution\.

Querytagstoinclude:

\-cause\_rating\(single\-slider\)

Prompt:HowmuchdidAcausetheoutcome?

Sliderrange:0to100

Return:\{"kind":"numeric",

"support":\[\{"value":<number\>,"prob":<probability\>\},\.\.\.\]\}

\-action\_choice\(multi\-choice\)

Prompt:Whatwilltheagentdonext?

Options:\["Goleft","Goright","Stay"\]

Return:\{"kind":"categorical",

"probabilities":\{"<exactoptionlabel\>":<probability\>,\.\.\.\}\}

The returned per\-option probabilities \(categorical\) or \(value, probability\) support \(numeric\) are then compared directly against the empirical human response distribution\.

## Appendix CEvaluation Metrics and Baselines

### C\.1Scoring and Metrics

Model responses are parsed from JSON and scored against human behavioral data using query\-type\-specific metrics:

- •Single\-slider: The model’s numeric response is compared to the human mean\. Absolute error \(normalized by scale range\) and the squared Pearson correlation \(R2R^\{2\}\) across trials measure alignment\.
- •Multi\-choice: Across repeated runs, the model’s selected options are converted to selection probabilities for each option\. Each option–item pair contributes one data point: the model selection probability paired with the human selection proportion\. These pairs are pooled within answer type and native scale before computingR2R^\{2\}\.
- •Multi\-select: The model’s binary selection vector is compared to human selection proportions\.
- •Multi\-slider: Each sub\-item is scored independently as a single\-slider\.
- •Ranking: Spearman rank correlation between model and human mean rankings\.

The primary alignment metric isR2¯\\overline\{R^\{2\}\}: within each experiment, we compute oneR2R^\{2\}for each answer type \(defined by response regime and native scale\), pooling conditions or constructs that share that type and scale; we average equally across answer types and then report the mean across experiments\. Human split\-half correlations are estimated from 100 random participant splits\. For the humanR2R^\{2\}references in Tables[1](https://arxiv.org/html/2609.21259#A4.T1)and[2](https://arxiv.org/html/2609.21259#A4.T2), we apply the Spearman–Brown correctionrSB=2​r/\(1\+r\)r\_\{\\mathrm\{SB\}\}=2r/\(1\+r\)to each experiment’s exported mean split\-half correlation, square the corrected correlation, and then average across experiments\. These corrected references use the 172 recoverable archived exports \(80 text, 33 image, 59 video\)\. The humanR2R^\{2\}standard errors in Table[1](https://arxiv.org/html/2609.21259#A4.T1)are computed across corrected experiment\-level values\. to ensure reliable estimates, we restrict the split\-half computation to response items with at least 10 individual data points\.

### C\.2Random Baseline

To place model–human alignment on an interpretable scale, we report a task\-naive random baseline, computed over the same items and passed through the identical scoring and normalization used for every model\. The baseline ignores the stimulus and draws responses uniformly at random over each item’s response space\. Because that space is set by the experiment’s response format, the randomization adapts to the task type:

- •Continuous \(slider\) responses:we draw a value uniformly from the item’s admissible range\[ℓ,h\]\[\\ell,h\], whereℓ\\ellandhhare the minimum and maximum values the slider permits\.
- •Categorical \(forced\-choice\) responses:we select one of thekkoptions uniformly at random, yielding a one\-hot vector over the options\.

For theR2R^\{2\}baseline, we repeat the draw1,0001\{,\}000times with a fixed random seed and compute the same per\-experimentR2R^\{2\}against the human means as for the models\. For the divergence baseline, we compare the full task\-naive uniform response distribution directly with the empirical human response distribution, using normalized Wasserstein\-1 distance over the admissible response range for continuous items and normalized Jensen–Shannon distance over the available options for categorical items\. This distribution\-to\-distribution comparison matches the metric used for model responses; it is not the average divergence of individual random point responses\. We average item scores within each answer type and native scale, then equally across answer types within an experiment and across experiments within a modality\. The resulting values populate the “Random baseline” row of[Table 1](https://arxiv.org/html/2609.21259#A4.T1)and the dotted reference lines in Figures[4](https://arxiv.org/html/2609.21259#S4.F4)and[5](https://arxiv.org/html/2609.21259#S4.F5)\.

## Appendix DModel–Human Alignment Results

### D\.1Full Model–Human Alignment Results

[Table 1](https://arxiv.org/html/2609.21259#A4.T1)reports per\-model human\-alignment scores, measured byR2R^\{2\}\(central\-tendency alignment\) and normalized distribution divergence, broken down by modality\.

R2R^\{2\}↑\\uparrowNorm\. Div\.↓\\downarrowModelText\(91\)Image\(82\)Video\(85\)Text\(86\)Image\(80\)Video\(67\)Human reference†0\.93±\\pm\.010\.95±\\pm\.010\.92±\\pm\.010\.05±\\pm\.010\.07±\\pm\.010\.07±\\pm\.01Random baseline0\.07±\\pm\.020\.05±\\pm\.010\.03±\\pm\.010\.28±\\pm\.010\.26±\\pm\.010\.29±\\pm\.01Claude \(Anthropic\)Claude Opus 4\.70\.55±\\pm\.030\.45±\\pm\.04–0\.10±\\pm\.010\.17±\\pm\.01–Claude Sonnet 4\.50\.52±\\pm\.030\.38±\\pm\.03–0\.11±\\pm\.010\.20±\\pm\.01–Claude Haiku 4\.50\.46±\\pm\.030\.22±\\pm\.03–0\.13±\\pm\.010\.20±\\pm\.01–GPT \(OpenAI\)GPT\-5\.20\.51±\\pm\.030\.47±\\pm\.03–0\.11±\\pm\.010\.21±\\pm\.02–GPT\-5 mini0\.47±\\pm\.030\.47±\\pm\.03–0\.13±\\pm\.010\.20±\\pm\.02–GPT\-4\.10\.49±\\pm\.030\.33±\\pm\.04–0\.12±\\pm\.010\.21±\\pm\.02–GPT\-4o0\.52±\\pm\.030\.30±\\pm\.03–0\.13±\\pm\.010\.21±\\pm\.02–GPT\-4o mini0\.43±\\pm\.030\.21±\\pm\.03–0\.15±\\pm\.010\.24±\\pm\.02–GPT\-OSS 120B0\.49±\\pm\.03––0\.13±\\pm\.01––GPT\-OSS 20B0\.46±\\pm\.03––0\.14±\\pm\.01––Gemini \(Google\)Gemini 3\.1 Pro0\.59±\\pm\.030\.58±\\pm\.030\.43±\\pm\.030\.11±\\pm\.010\.17±\\pm\.010\.25±\\pm\.02Gemini 3 Flash0\.52±\\pm\.030\.53±\\pm\.030\.28±\\pm\.020\.11±\\pm\.010\.17±\\pm\.010\.28±\\pm\.02Gemini 2\.5 Pro0\.51±\\pm\.030\.52±\\pm\.030\.25±\\pm\.020\.12±\\pm\.010\.18±\\pm\.010\.32±\\pm\.02Gemini 2\.5 Flash0\.51±\\pm\.030\.46±\\pm\.030\.19±\\pm\.020\.13±\\pm\.010\.19±\\pm\.020\.35±\\pm\.02Qwen \(Alibaba\)Qwen3\.5 Flash 35B0\.49±\\pm\.030\.39±\\pm\.030\.13±\\pm\.030\.15±\\pm\.010\.21±\\pm\.020\.27±\\pm\.02Qwen3\.5 9B0\.35±\\pm\.030\.19±\\pm\.030\.10±\\pm\.020\.18±\\pm\.010\.23±\\pm\.020\.30±\\pm\.02Qwen3\.5 4B0\.30±\\pm\.030\.19±\\pm\.030\.09±\\pm\.020\.18±\\pm\.010\.25±\\pm\.020\.31±\\pm\.02Qwen3\.5 2B0\.12±\\pm\.010\.08±\\pm\.020\.05±\\pm\.010\.23±\\pm\.020\.25±\\pm\.020\.30±\\pm\.02Qwen3\-VL 32B0\.49±\\pm\.030\.26±\\pm\.030\.12±\\pm\.020\.14±\\pm\.010\.22±\\pm\.020\.30±\\pm\.02Qwen3\-VL 8B0\.37±\\pm\.030\.22±\\pm\.030\.11±\\pm\.020\.16±\\pm\.010\.23±\\pm\.020\.30±\\pm\.02Qwen3\-VL 4B0\.33±\\pm\.030\.15±\\pm\.030\.10±\\pm\.020\.17±\\pm\.010\.25±\\pm\.020\.32±\\pm\.02Qwen2\.5\-VL 32B0\.47±\\pm\.030\.25±\\pm\.030\.08±\\pm\.010\.14±\\pm\.010\.22±\\pm\.020\.29±\\pm\.02Qwen2\.5\-VL 7B0\.32±\\pm\.030\.12±\\pm\.020\.07±\\pm\.010\.17±\\pm\.010\.24±\\pm\.020\.28±\\pm\.02Qwen2\.5\-VL 3B0\.26±\\pm\.030\.10±\\pm\.020\.05±\\pm\.010\.25±\\pm\.020\.24±\\pm\.010\.29±\\pm\.02Qwen3\.5 0\.8B0\.08±\\pm\.010\.06±\\pm\.010\.05±\\pm\.010\.31±\\pm\.020\.43±\\pm\.070\.32±\\pm\.02Qwen3 1\.7B0\.23±\\pm\.03––0\.23±\\pm\.06––Qwen2\.5 1\.5B0\.15±\\pm\.02––0\.26±\\pm\.02––Qwen2\.5 0\.5B0\.11±\\pm\.02––0\.33±\\pm\.04––Llama \(Meta\)Llama 4 Maverick0\.40±\\pm\.030\.23±\\pm\.03–0\.14±\\pm\.010\.24±\\pm\.02–Llama 4 Scout0\.39±\\pm\.030\.22±\\pm\.03–0\.16±\\pm\.010\.23±\\pm\.02–Llama 3\.3 70B0\.46±\\pm\.03––0\.14±\\pm\.01––Llama 3\.2 11B0\.32±\\pm\.030\.22±\\pm\.07–0\.23±\\pm\.020\.27±\\pm\.02–Llama 3\.2 3B0\.25±\\pm\.03––0\.20±\\pm\.03––Llama 3\.2 1B0\.10±\\pm\.02––0\.24±\\pm\.03––Llama 3\.1 70B0\.48±\\pm\.03––0\.14±\\pm\.01––Llama 3\.1 8B0\.35±\\pm\.03––0\.22±\\pm\.01––Gemma \(Google\)Gemma 4 26B0\.45±\\pm\.030\.27±\\pm\.03–0\.13±\\pm\.010\.24±\\pm\.02–Gemma 4 E2B0\.28±\\pm\.03––0\.24±\\pm\.02––Gemma 3 27B0\.43±\\pm\.030\.22±\\pm\.03–0\.13±\\pm\.010\.21±\\pm\.02–Gemma 3 12B0\.37±\\pm\.03––0\.17±\\pm\.01––Gemma 3 4B0\.26±\\pm\.03––0\.19±\\pm\.01––OtherGrok 4\.30\.52±\\pm\.030\.30±\\pm\.04–0\.13±\\pm\.010\.20±\\pm\.02–GLM 5\.20\.53±\\pm\.03––0\.10±\\pm\.01––Kimi K2\.50\.48±\\pm\.030\.34±\\pm\.03–0\.10±\\pm\.010\.15±\\pm\.01–Mistral Small 24B0\.44±\\pm\.03––0\.15±\\pm\.01––Mistral 7B0\.30±\\pm\.03––0\.19±\\pm\.01––DeepSeek V4 Pro0\.50±\\pm\.03––0\.13±\\pm\.01––DeepSeek V4 Flash0\.45±\\pm\.03––0\.13±\\pm\.01––DeepSeek V3\.20\.40±\\pm\.03––0\.13±\\pm\.01––Centaur\[[Binz et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib39)\]0\.42±\\pm\.03––0\.19±\\pm\.01––Table 1:Model–human alignment measured byR2R^\{2\}and normalized distribution divergence\.All models use default configuration with thinking mode enabled if available\. Values following±\\pmshow standard errors\.†HumanR2R^\{2\}values are Spearman–Brown\-corrected references from the 172 recoverable archived exports \(80 text, 33 image, 59 video\); their standard errors are computed across corrected experiment\-levelR2R^\{2\}values\. Parenthetical counts in the header give the number of benchmark experiments eligible for each metric before model\-specific missingness\. Centaur is the Llama 3\.1 70B base model fine\-tuned on human behavioral data; all other models are instruction\-tuned\.
### D\.2Model–HumanR2R^\{2\}on Split\-Half Reliable Subset

[Table 2](https://arxiv.org/html/2609.21259#A4.T2)reports the top 10 models by textR2R^\{2\}on the 179 experiments \(text: 80, image: 37, video: 62\) where human split\-half reliability can be reliably estimated \(average≥\\geq10 individual responses per item\)\. The Spearman–Brown\-corrected human reference is high \(text:R2=0\.932R^\{2\}=0\.932, image:0\.9450\.945, video:0\.9180\.918\), indicating strong human agreement\. We corrected each experiment’s exported mean split\-half correlation asrSB=2​r/\(1\+r\)r\_\{\\mathrm\{SB\}\}=2r/\(1\+r\)before squaring and averaging across experiments\. This human reference uses the 172 recoverable archived exports \(80 text, 33 image, 59 video\), as in[Figure 4](https://arxiv.org/html/2609.21259#S4.F4); the 179\-experiment model subset is unchanged\. Grok 4\.3 leads on text \(0\.60\), while Gemini 3\.1 Pro leads on image \(0\.53\) and video \(0\.23\)\. The gap between the best model and human reliability remains substantial across all modalities\.

Table 2:Top 10 model–humanR2R^\{2\}on the split\-half reliable subset\(179 experiments\), ranked by textR2R^\{2\}\. The human reference is Spearman–Brown\-corrected using the 172 recoverable archived exports \(80 text, 33 image, 59 video\)\.

## Appendix EModel Sampling and Response Variability

### E\.1Run\-Count Sensitivity Analysis

A possible concern is that too few model samples were collected per item\. To test this directly, we recomputed Gemma 4 26B textR2R^\{2\}using increasing numbers of completed runs while holding prompts, temperature, and trial selection fixed\. Additional runs stabilize the estimate but do not change the conclusion: meanR2R^\{2\}remains effectively flat, with overlapping 95% confidence intervals across all run counts \([Table 3](https://arxiv.org/html/2609.21259#A5.T3)\)\.

Table 3:Effect of number of model runs on Gemma 4 26B text alignment\.Increasing the number of runs from 5 to 50 does not materially improveR2R^\{2\}; the 95% confidence intervals \(1000\-sample bootstrap over experiments\) overlap across all run counts\.
### E\.2Sampled Responses Under\-express Human Variability

A central motivation for eliciting model\-predicted human response distributions through verbalized prompting\[[Meister et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib43)\]rather than repeated sampling is that, when sampled at temperature1\.01\.0, model responses are sharply peaked relative to humans: repeated runs concentrate on a single high\-probability response and recover only a fraction of the item\-level variability present across human participants\.[Figure 9](https://arxiv.org/html/2609.21259#A5.F9)visualizes this—for each item we compute response variability on the normalized response scale and compare the resulting distributions across humans and repeated model runs \(values near00indicate nearly deterministic responses\)\. Sampled diversity varies markedly across models: GPT\-5\.2 is the most diverse non\-thinking model \(72%72\\%of human variability\), followed by Gemini 3\.1 Pro \(38%38\\%\), Gemma 4 \(14%14\\%\), and Claude Opus \(10%10\\%\)—but all fall well short of human spread\. Repeated sampling thus captures only a fraction of the item\-level variability present across human participants, even for models that track the human mean reasonably well\.

![Refer to caption](https://arxiv.org/html/2609.21259v1/figures/diversity.png)Figure 9:Response diversity: humans vs\. repeated model runs\.For each item, response variability is computed on the normalized\[0,1\]\[0,1\]response scale, and the distributions are compared across humans and repeated model runs\. Values near00indicate nearly deterministic responses; broader distributions indicate greater item\-level variability\. Sampled model responses are markedly less diverse than human responses\.This peakedness directly inflates distributional divergence\.[Figure 10](https://arxiv.org/html/2609.21259#A5.F10)compares, per experiment for Gemini 3\.1 Pro, the normalized divergence obtained from repeated sampling \(1010runs\) against that obtained from verbalized elicitation\. Sampling yields substantially larger divergence in almost every experiment—the median rises from0\.110\.11to0\.290\.29on text,0\.140\.14to0\.290\.29on image, and0\.230\.23to0\.340\.34on video, with91%91\\%/88%88\\%/78%78\\%of experiments above the parity diagonal\. Repeated sampling and verbalized elicitation probe distinct quantities: the former estimates variability in the model’s own outputs across runs, whereas the latter elicits its explicit prediction of the human response distribution\. Because our distributional\-alignment analysis asks how well models predict human response variability, we use verbalized elicitation for all reported divergence computations\.

![Refer to caption](https://arxiv.org/html/2609.21259v1/figures/g31_verb_vs_sampled_scatter.png)Figure 10:Verbalized\-distribution vs\. sampled\-runs divergence \(Gemini 3\.1 Pro\)\.Each point is one experiment;xxis the normalized divergence from verbalized elicitation andyyfrom1010sampled runs \(lower is better\)\. The vast majority of experiments fall above the parity diagonal, i\.e\., repeated sampling—being sharply peaked—diverges from the human response distribution more than the verbalized estimate does\.

## Appendix FHuman Replications

A natural concern with any large\-scale benchmark derived from published experiments is whether poor model–human alignment reflects genuine cognitive divergence or artifacts of data standardization\.

Here,R2R^\{2\}denotes squared Pearson correlation between response means, and the overall summaries are unweighted arithmetic means of the 15 experiment\-level values\. Mean replication–original agreement isR2=0\.84R^\{2\}=0\.84, compared withR2=0\.62R^\{2\}=0\.62for Gemini 3\.1 Pro versus the original humans\.

ExperimentModality𝑵\\boldsymbol\{N\}Repl\.𝑹𝟐\\boldsymbol\{R^\{2\}\}G3\.1 Pro𝑹𝟐\\boldsymbol\{R^\{2\}\}[Levine \(2020\), Exp\. 1](https://coggym.org/studies/levine2020logic-exp1)Text350\.860\.48[Tsvilodub \(2025\), Exp\. 1](https://coggym.org/studies/tsvilodub2025nonliteral-exp1)Text300\.790\.39[Ying \(2023\), Exp\. 1](https://coggym.org/studies/ying2023neuro-exp1)Text300\.820\.70[Yoon \(2020\), Exp\. 1](https://coggym.org/studies/yoon2020polite-exp1)Text300\.970\.95[Hu \(2023\), Exp\. 1](https://coggym.org/studies/hu2023fine-exp1)Text300\.990\.93[Aboody \(2025\), Exp\. 2](https://coggym.org/studies/aboody2025inferring-exp2)Image300\.890\.84[Jara\-Ettinger \(2020\), Exp\. 6](https://coggym.org/studies/jara-ettinger2020naive-exp6)Image300\.800\.81[Lopez\-Brau \(2023\), Exp\. 1](https://coggym.org/studies/lopez-brau2023people-exp1)Image500\.750\.78[Jara\-Ettinger \(2021\), Exp\. 2](https://coggym.org/studies/jaraettinger2021quantitative-exp2)Image300\.770\.35[Chandra \(2024\), Exp\. 1](https://coggym.org/studies/chandra2024cooperative-exp1)Image670\.730\.62[Fu \(2025\), Exp\. 1](https://coggym.org/studies/fu2025hierarchical-exp1)Video300\.920\.66[Bass \(2022\), Exp\. 1](https://coggym.org/studies/bass2022partial-exp1)Video390\.920\.21[Sosa \(2021\), Exp\. 1](https://coggym.org/studies/sosa2021moral-exp1)Video300\.820\.15[Wu \(2023\), Exp\. 1b](https://doi.org/10.31234/osf.io/uwdbr)Video300\.810\.90[Royka \(2022\), Exp\. 1](https://coggym.org/studies/royka2022people-exp1)Video300\.750\.56Mean / total5210\.840\.62Table 4:Experiment replication and comparisons\.NNis the number of participants in each included replication batch; the total sums study participations\.Repl\.R2R^\{2\}compares replication means against original human means;G3\.1 ProR2R^\{2\}compares Gemini 3\.1 Pro against the original humans\.
## Appendix GCross\-Model Comparison

Model behavior is strongly correlated across experiments: Claude Opus 4\.7 and Gemini 3\.1 Pro tend to succeed and fail on the same experiments \([Figure 11](https://arxiv.org/html/2609.21259#A7.F11)\), though notable divergences remain where one model is well aligned and the other near chance\.

This cross\-model consistency carries two implications\. First, per\-experiment difficulty is largely a property of the task rather than the model: an experiment that one frontier model struggles on is likely to be difficult for the other as well, so the alignment gaps we report reflect shared inductive limitations—systematically over\-confident, under\-dispersed judgments relative to graded human responses—rather than idiosyncratic errors of a single system\. Second, because frontier models rise and fall together across CogGym, their relative ordering is more stable than the absolute alignment level, which remains far below the human ceiling for all of them; swapping model families is therefore unlikely, on its own, to close the gap\. The off\-diagonal experiments—where the two models disagree—are the most informative for diagnosis, marking paradigms where the models bring different priors to bear, and are natural candidates for targeted follow\-up\.

![Refer to caption](https://arxiv.org/html/2609.21259v1/figures/opus_vs_g31pro_scatter.png)Figure 11:Per\-experimentR2R^\{2\}: Claude Opus 4\.7 vs\. Gemini 3\.1 Pro\.Each point is one experiment\. Overall the plots show strong correlation between the two models\. Error bars on both axes are 95% confidence intervals \(1000\-sample bootstrap percentile, resampling items\); experiments whoseR2R^\{2\}is too poorly estimated to compare \(95% CI width\>0\.4\>0\.4on either axis, typically those with very few items\) are omitted\.
## Appendix HContamination Analysis

#### Procedure

We conducted a one\-run, memory\-only study\-recognition probe on 175 experiments \(92 text and 83 static\-image experiments, spanning 59 papers\)\. Each independent request presented anonymized instructions, stimulus examples, and questions\. Models were asked to recall the paper title, authors, publication year, and any formal model, and to report separate confidence estimates for paper identification and model recall\. No web search, retrieval tools were provided\. The request contained up to 40 trial examples\. For image experiments, we additionally attached original images for up to three trial examples without any human response data\.

#### Edit\-distance title similarity

For the continuous\-similarity analysis reported in Results, we compute similarity between the original title and paper title provided by AI as1−d⁡\(a,b\)/max⁡\(\|a\|,\|b\|\)1\-d\(a,b\)/\\max\(\|a\|,\|b\|\), whereddis character\-level Levenshtein distance with unit costs for insertion, deletion, and substitution, and lengths are measured after normalization\. This score ranges from zero to one, with one indicating identical normalized titles\.

#### Model prompt

The following system prompt was shared across models\. The user message supplied only the anonymized, experiment\-specific materials described above; it did not include ground\-truth bibliographic metadata or human responses\. The prompt’s claim that identifying details were removed describes the intended anonymization, subject to the limitations noted above\.

\#Knowledgecheck\-\-identifythisstudyfrommemory

Youareshownthe\(anonymized\)materialsofasinglepublishedcognitive\-scienceexperiment:itsinstructions,stimuli,andquestions,astheywerepresentedtohumanparticipants\.Allidentifyingdetails\-\-thepaper’stitle,authors,institution,andanyhumandata\-\-havebeenremoved\.

Yourjobistorecognizethestudy\*\*fromyourownmemoryalone\*\*\.Youhavenointernetaccessandnosearchtools;donotattempttolookanythingup\.Ifyouareunsure,giveyourbestguessandalowconfidenceratherthanrefusing\.

Answerthesequestions:

1\.\*\*Whichpublishedstudyisthis?\*\*Givethepaper’stitle,itsauthors,andtheyearofpublication\-\-yourbestrecollection\.

2\.\*\*Didthatpaperintroduceortestaformal/computationalmodel?\*\*Ifso,nameitandbrieflydescribeitfrommemory:themodelfamily,anditskeyparametersormechanism\.Ifthepaperintroducednocomputationalmodel,sayso\(‘has\_model:false‘\)\.

3\.\*\*Howconfidentareyou?\*\*Giveaconfidencein\[0,1\]forthepaperidentificationandaseparateconfidenceforthemodelrecall\.

Reasonfromthespecificstructureofthetask\-\-thecoverstory,themanipulatedvariables,theresponseformat,thenumberanddesignoftrials\-\-tothepaperyouremembermatchingit\.Baseeveryansweronlyonwhatyoualreadyknow;neverinventacitationyouarenotactuallyrecalling\(anhonest"unknown"withlowconfidenceismoreusefulthanafabricatedtitle\)\.

ReplywithONLYoneJSONobject,noprosebeforeorafterit,inexactlythisshape:

\{

"paper\_title":"<thepaper’stitle,bestguessfrommemory,or\\"\\"ifunknown\>",

"authors":"<authorsurnamesyourecall,or\\"\\"\>",

"year":"<publicationyear\(YYYY\),or\\"unknown\\"\>",

"paper\_confidence":0\.0,

"has\_model":false,

"model\_name":"<nameofthecomputationalmodelthepaperintroduced/tested,ornull\>",

"model\_description":"<frommemory:modelfamily,keyparameters/mechanism,or\\"\\"\>",

"model\_confidence":0\.0

\}

‘paper\_title‘and‘has\_model‘arerequired;therestareoptional\.Confidencesarein\[0,1\]\.

## Appendix ISeverely Misaligned Cases

We selected trials on which the strongest models disagree with a clear human consensus, drawing from experiments that span different topics and both the text and image modalities\. Candidates were identified by ranking every scored item by the scale\-normalized gap between the model mean and the human mean, weighted by human agreement, for Gemini 3\.1 Pro, Claude Opus 4\.7, and GPT\-5\.2; experiments with known stimulus\-delivery or scoring artifacts were excluded\. Each figure shows the stimulus and question as presented, followed by the distribution of the original participants’ responses and, in separate panels, the distribution of each model’s sampled responses on the same trial \(1010runs per model\)\. The examples are illustrative rather than representative\.

Figure 12:Concept learning \(text\): the number game\[[Bigelow and Piantadosi, 2016](https://arxiv.org/html/2609.21259#bib.bib71)\]\. Participants see a few outputs of an unknown program and judge whether a new number is likely to be generated next\. Given*15, 39, 35*, only18%18\\%of participants \(n=11n=11\) say that*41*is likely, but every run of every model says yes, with the same justification:*“All previous numbers are odd integers, and 41 is also an odd integer\.”*Human generalization in this task hedges between rule\-based hypotheses \(odd numbers\) and similarity\-based ones \(numbers in the teens and thirties\); the models commit to the first compact rule that fits\.Figure 13:Pragmatics of number words \(text\)\[[Tsvilodub et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib51)\]\. After reading*“It cost $51,”*participants rate the probability of ten candidate prices; ratings are renormalized within participant as in the original analysis\. Humans spread their probability across $50, $51, $500, $501, $1,000, and $1,001, placing only about19%19\\%on the literal $51 \(n=66n=66\)\. Gemini places100%100\\%on the literal value in every run and GPT\-5\.2 about80%80\\%on average \(*“Fred explicitly stated the price was $51, indicating a precise and literal value”*\)\. Opus splits: in five of ten runs it places9090–95%95\\%on $51, and in the other five it returns the same value for every price, which renormalizes to a flat distribution\. None of the models reproduces the human pattern of treating the number as an imprecise report\.![Refer to caption](https://arxiv.org/html/2609.21259v1/10_image_yildirim2023_cloth_bicycle.png)Figure 14:3D shape perception \(image\): reading the silhouette instead of the physics\[[Yildirim et al\., 2024](https://arxiv.org/html/2609.21259#bib.bib56)\]\. Participants match a cloth\-covered target to one of two uncovered objects, a task people solve by reasoning about how cloth settles on a shape\.97%97\\%of participants \(n=29n=29\) identify the bicycle; Gemini and GPT\-5\.2 choose the table in every run and Opus in nine of ten, describing the draped outline as if it were the object:*“The draped cloth outlines a flat rectangular surface with four protruding legs, which perfectly matches the structure of the table”*\(Gemini\)\.![Refer to caption](https://arxiv.org/html/2609.21259v1/12_image_aboody2025_pirates_skip_map.png)Figure 15:Theory of mind \(image\): inferring knowledge from a skipped opportunity\[[Aboody et al\., 2025](https://arxiv.org/html/2609.21259#bib.bib72)\]\. Pirates land at the star and walk directly to where they dig \(white square\), passing close to a map \(green square\) without picking it up\. Participants infer that the pirates already knew where the treasure was \(M=82M=82on a 0–100 scale,n=40n=40\)\. Opus and GPT\-5\.2 answer2828and3636on average, and GPT\-5\.2 misdescribes the scene \(*“They chose to detour to get the map, suggesting limited prior knowledge”*\); Gemini splits between runs that read the scene correctly \(*“The pirates skipped the map even though it was very close to their path, suggesting they already knew where the treasure was”*, answering8080–8585\) and runs that answer2020–3030\.[30](https://arxiv.org/html/2609.21259#bib.bib52),[19](https://arxiv.org/html/2609.21259#bib.bib30),[60](https://arxiv.org/html/2609.21259#bib.bib51),[48](https://arxiv.org/html/2609.21259#bib.bib53),[74](https://arxiv.org/html/2609.21259#bib.bib54),[27](https://arxiv.org/html/2609.21259#bib.bib55),[70](https://arxiv.org/html/2609.21259#bib.bib56),[72](https://arxiv.org/html/2609.21259#bib.bib58),[20](https://arxiv.org/html/2609.21259#bib.bib57),[32](https://arxiv.org/html/2609.21259#bib.bib59)

相似文章

衡量通向AGI的进展:一个认知框架

Google DeepMind Blog

Google DeepMind发布了一篇论文,提出了一个衡量通向通用人工智能(AGI)进展的认知框架,识别了十项关键认知能力,并发起了一场Kaggle黑客马拉松以构建相关评估方法。

NeuroCogMap揭示大型语言模型的认知组织结构

Hugging Face Daily Papers

NeuroCogMap是一个受认知神经科学启发的框架,将大型语言模型的内部特征映射为功能性脑区,并将其与可解释的认知功能联系起来,揭示模型失败(如幻觉和偏见)的迹象。