On the use of foundation models in cognitive science

arXiv cs.CL Papers

Summary

This perspective paper from arXiv articulates a four-stage inferential framework for evaluating foundation models as cognitive and developmental models, emphasizing that behavioral alignment alone is insufficient and must be embedded within theoretical commitments and contrastive evaluation.

arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory-driven and comparative evaluation. Throughout, we argue that behavioral fit alone is insufficient. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory-diagnostic tasks, and systematic contrastive evaluation across candidate models.
Original Article
View Cached Full Text

Cached at: 08/11/26, 08:05 AM

# On the use of foundation models in cognitive science
Source: [https://arxiv.org/html/2608.07812](https://arxiv.org/html/2608.07812)
\\startpage

1\\historydates\\doiheadtext

\\authormark

Shahet al\.\\titlemarkOn the use of Foundation models in cognitive science

\\corres

Corresponding author: Raj Sanjay Shah

Alex WarstadtMichael FrankSashank Varma\\orgdivInteractive Computing,\\orgnameGeorgia Institute of Technology,\\orgaddress\\stateGeorgia,\\countryUSA\\orgdivHalıcıoğlu Data Science Institute and Dept\. of Linguistics,\\orgnameUniversity of California, San Diego,\\orgaddress\\stateCalifornia,\\countryUSA\\orgdivDept\. of Psychology,\\orgnameStanford University,\\orgaddress\\stateCalifornia,\\countryUSA[rajsanjayshah@gatech\.edu](https://arxiv.org/html/2608.07812v1/mailto:[email protected])Raj Sanjay ShahAlex WarstadtMichael FrankSashank Varma

###### Abstract

\[Abstract\] A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models \(FMs\)\. These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children’s cognitive development\. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges\. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four\-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model\-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations\. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory\-driven and comparative evaluation\. Throughout, we argue that behavioral fit alone is insufficient\. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory\-diagnostic tasks, and systematic contrastive evaluation across candidate models\.

\\jnlcitation\\cname

and and and\\ctitleOn the use of foundation models in cognitive science\\cjournal\\cvol\.

###### keywords:

Cognitive modeling, Foundation models, Linking hypotheses, Cognitive science

††articletype:Perspective## 1Introduction

With the steadily improving performance of foundation models \(FMs\)anthropic2026claude,gemini31,gpt56, researchers have increasingly explored their use as computational models of cognitionpiantadosi2023modern,mahowald2024dissociating,warstadt2022what,binz2025foundation\. Here, we use FMs as an umbrella term for pretrained, adaptable models that can support a range of tasks and modalities\. Our scope includes large language models, multi\-modal models, reasoning\-oriented or post\-trained models, as well as models trained under constrained regimes, such as BabyLM models\. Across domains such as mathematical reasoningtestolin2020numerosity,ahn2024large,liucogmath, language comprehensionduan2024hlb,cog\_sci\_garden\_path,hu2024language, conceptual understandingcog\_sci\_typicallity,bhatia2022transformer, spatial reasoningramakrishnan2024does,wang2024picture, and analogical reasoningwebb2023emergent,hu2023context,geiger2023relational, FMs have been shown to reproduce behavioral signatures long documented by cognitive scientists\. For example, recent workshah2023numericfound that FMs exhibit the distance, size, and ratio effects characteristic of the human “mental number line”moyerTimeRequiredJudgements1967,parkman1971temporal,halberda2008individual\. Similarly,webb2023emergenttranslated Raven’s Progressive Matrices into symbolic digit matrices and reported FM performance comparable to humans\. These findings have motivated the development of benchmarks aimed at systematically evaluating Model\-human alignmentcoda2024cogbench,wang2025coglm\.

Researchers have also begun to use FMs to model cognitive development in childrenhosseini2022artificial,chang2022word,frank2023bridging,evanson2023language,ficarra2025distributional\. For example,Portelance2023PredictingAOshowed how language models can be used to predict the ages at which children acquire words\. Rather than examining only model end states,shah2024developmentevaluated intermediate training checkpoints to assess whether models’ developmental trajectories in numerical ability, linguistic ability, conceptual understanding, and fluid reasoning parallel the developmental patterns observed in children\. Similarly,tan2024devbenchexplored developmental parallels by comparing the learning trajectories of vision\-language models to both child and adult behavioral data\. These and many other studies reflect a broadening shift in research goals towards examining their learning dynamics\. Across both cognitive and developmental settings, we use ‘alignment’ to mean systematic correspondences between model outputs and human behavioral measures†††Our use of alignment is distinct from the one in AI safety, where it refers to aligning a system’s goals, rewards, or behavior with human values or intentions\.\.

![Refer to caption](https://arxiv.org/html/2608.07812v1/x1.png)Figure 1:A four\-stage framework for evaluating FMs as cognitive models\. \(A\) The inner alignment loop establishes whether a given model reproduces human performance under explicit task adaptations and linking hypotheses \(Stages 1\-3\)\. \(B\) The outer contrastive loop evaluates multiple models or manipulations to identify computational features necessary for alignment \(Stage 4\)\.In this paper, we ask: under what conditions does behavioral alignment justify treating foundation models as explanatory models of cognition? Because reproducing behavioral patterns does not by itself identify the mechanisms that generate them, researchers must distinguish explanatory cognitive models from behavioral proxiesgershman2024have,van2024reclaiming\. We propose an inferential framework for evaluating model\-human alignment that makes three contributions\. First, we present four evaluation stages \(see Figure[1](https://arxiv.org/html/2608.07812#S1.F1)\), separating task adaptation, linking hypotheses, behavioral evaluation, and contrastive model comparison, and delineate how each stage constrains the interpretation of alignment\. Second, we examine the interpretive role of linking hypotheses, the central theoretical commitment that connects model behavior to human performance measures\. Third, we distill our proposals and analyses into guidelines for designing evaluations of the explanatory value \(rather than the descriptive similarity\) of FMs\. Throughout, our goal is not to advocate uncritically for FMs as cognitive models but rather to clarify the conditions under which their alignment with human behavior becomes scientifically informative\.

## 2A four\-stage framework for evaluating FM alignment

Our proposed framework is directed toward research that uses computational models to test specific, theoretically motivated hypothesesfrank2025cognitive\. Such efforts can support explanatory claims only when researchers make explicit what claim is being tested and what counts as “alignment” between model and human behavior\. Without this structure, observed similarities risk being incidental\. More strongly, explanatory relevance is strengthened when \(1\) the model’s performance profiles match those of humans on the relevant dimensions specified by the hypothesis; \(2\) the model can predict human behavior in response to stimuli designed to test the hypothesis; and \(3\) the model’s internal representations and processes can be interpreted as theoretical proposals about the underlying mechanisms\. This perspective aligns with broader shifts toward model comparison and predictive adequacy in cognitive scienceyarkoni2017choosing, while emphasizing that predictive success alone does not constitute explanation\.

Our framework concerns the design and interpretation of model\-human comparisons\. It does not prescribe which architectures, objectives, or inductive biases constitute better cognitive models; these questions have been discussed elsewherewarstadt2022what\. Instead, it provides a structure for evaluating a theoretically motivated set of candidates\. To operationalize these principles, we propose a four\-stage framework for mapping between the performance of FMs and humans on cognitive tasks\.

1. 1\.Adapt the experimental stimuli and task\.Human experimental materials must be translated into the input\-output modalities of FMs so that the model and participants perform functionally equivalent tasks\.
2. 2\.Specify the linking hypothesis\.Researchers must define how model outputs correspond to measurable aspects of human behavior\. This mapping determines what counts as evidence of alignment, whether in terms of endpoint measures \(e\.g\., accuracies, reaction times, error patterns\) or developmental trajectories \(e\.g\., age\-of\-acquisition curves, learning curves\)\.
3. 3\.Evaluate correspondence with human performance:Model performance must be compared to human data using appropriate statistical measures of goodness of fit\. Importantly, alignment should be assessed not just in aggregate; emphasis must be placed on theoretically diagnostic variation, for example, across conditions, individuals, or stimuli\.
4. 4\.Compare across candidate models and manipulations:Demonstrating alignment in a single model provides, at most, proof of possibility\. Explanatory value increases through contrastive evaluation: comparing architectures, training regimes, scales, or ablated variants to identify which computational components are*necessary*to reproduce a behavioral signature\.

As depicted in Figure[1](https://arxiv.org/html/2608.07812#S1.F1), Stages 1\-3 form an inner alignment loop that establishes whether a model reproduces human performance under explicit task adaptations and linking hypotheses\. Stage 4 functions as an outer, contrastive loop: by comparing architectures, training regimes, or ablated variants, researchers can identify which computational features are necessary for the correspondence\. Related perspectives similarly emphasize that controlled manipulation and model comparison, rather than raw capability, are critical for making scientific progress with FMsong2024gpt,warstadt2022what,rozner2026perturbationsimpleefficientadversarial\. Our framework systematizes such proposals\.

Running Example: Digit\-Matrix Analogical Reasoningwebb2023emergent,webb2025evidencetranslated Raven’s Progressive Matrices \(RPM\) problems into text\-based “digit matrices” in which relational structure must be inferred to select the correct completion\. This symbolic \(vs\. visual\) representation \(arguably\) preserves the combinatorial and relational demands of fluid reasoning\. They showed that FMs can consistently select \(among multiple candidates\) the correct completion \(i\.e\., missing vector\)\. The question is how such performance should be mapped to human reasoning processes\.![[Uncaptioned image]](https://arxiv.org/html/2608.07812v1/x2.png)Figure 2:Adaptation of a visual RPM problem to a textual format\.Stage 1:To make this task compatible with the textual modality of FMs,webb2023emergentconverted each visual element into its symbolic attributes \(e\.g\., type, size, color; see Figure[2](https://arxiv.org/html/2608.07812#S2.F2)\), while preserving the relational structure of the original problem\. This enables the model to reason over the same combinatorial relationships present in the original visual problem\.Stage 2:Model behavior is then interpreted relative to human reasoning\. One option is a*prompting*strategy, in which the symbolic matrix and candidate answers are presented to the model exactly as they appear to human participants, and the model is asked to select the correct completion\. Alternatively, a*similarity\-based*linking hypothesis could analyze the model’s internal representations to test whether they encode the relational structure underlying the analogy\.Stage 3:Model predictions are then evaluated against human behavioral data using goodness\-of\-fit measures\. In this task, this involves comparing model accuracies, reaction times, and error patterns as a function of problem complexity to those observed in human participants\.Stage 4:Finally, contrastive comparisons identify which computational properties support alignment\. Whilewebb2023emergentprimarily demonstrate that alignment is possible, subsequent studies \(e\.g\.,hu2023context\) vary models, problem representations, and prompting strategies to examine how these changes influence the model’s ability to capture relational structure\.

## 3Linking Hypotheses: Mapping Model Performance to Human Performance

All three stages of the inner loop in Figure[1](https://arxiv.org/html/2608.07812#S1.F1)are necessary for establishing a principled alignment\. Task adaptation \(Stage 1\) and performance evaluation \(Stage 3\) are often straightforward to operationalize\. By contrast, the choice of linking hypothesis fundamentally determines what counts as evidence of alignment and how model behavior is interpretedwurgaftcontext\.

The “outputs” of FMs are often quite different from the behavioral measures that cognitive scientists collect in their experiments\. To bridge this gap, researchers use various*linking hypotheses*to map model performance measures to human performance measuresfrank2025cognitive\. We focus here on several widely used linking hypotheses that illustrate how different mapping assumptions shape claims of alignment\. These hypotheses are not neutral: each makes different theoretical commitments about what aspects of model computation map to cognition, thereby defining and potentially limiting how any observed alignment is interpreted\. A strong apparent fit under one linking hypothesis might weaken or vanish under another\. Consequently, alignment claims are not properties of models alone but of models evaluated under specific linking assumptions\.

### 3\.1Similarity

One common linking hypothesis treats the geometry of a model’s representational space as a proxy for the structure of human judgments or cognitive processes\. One form of this linking hypothesis concerns proximity: items that humans judge to be more similar or typical should occupy nearby regions of the model’s representational space\. Another concerns relational structure: stimuli or processes hypothesized to share a mechanism in humans might be represented through a common direction, rotation, or other transformation in the model’s representational spacehu2026representational\. Similarity\-based linking has been used to model category typicality effectsbhatia2022transformer,cog\_sci\_typicallityand numerical comparison phenomena such as distance, size, and ratio effectsshah2023numeric\. In these cases, distances in embedding space are mapped onto behavioral measures such as typicality ratings or response times\.

Despite its apparent simplicity, this linking hypothesis requires substantial assumptions\. It presupposes access to stable internal representations, which might not be available for closed\-weight, commercial models\. Computing a specific alignment score depends on a host of analytic choices, including layer selection, tokenization, and how category representations are defined \(e\.g\., via label embeddings or averaged exemplars\), each of which can materially alter similarity structuremisra2021language,bhatia2022transformer\. Results further hinge on how representational geometry is characterized, with no*a priori*guarantee that the chosen metric or transformation corresponds to psychologically meaningful structurerichie2021similarity\. However, this greater flexibility also introduces additional analytic choices and makes it more difficult to determine which characterization of the representational space provides the appropriate linking hypothesispiantadosi2024concepts,lampinen2025representation\.

### 3\.2Surprisal

One way of quantifying the uncertainty of an FM’s generation \(given context\) is in terms of its*surprisal*, or the negative logarithm of a token’s probability conditioned on the preceding tokens\. A common linking hypothesis is that higher surprisal values correspond to longer human response timeshale\-2001\-probabilistic,LEVY20081126\. FMs assign probabilities incrementally to text, enabling computation of surprisal for candidate continuations\. In the digit\-matrix analogy examplewebb2023emergent, surprisal could be used to quantify the model’s uncertainty over possible completions\. If the correct completion has lower surprisal than alternatives, this can be interpreted as evidence that the model encodes the relevant relational structure\. More generally, surprisal\-based linking has successfully predicted graded effects in word learning, grammaticality judgments, and incremental sentence processingwarstadt2020blimp,shain2024word\.

At the same time, adopting surprisal commits researchers to a particular relationship between predicted probability and processing difficulty\. However, surprisal can be used for other purposes, such as testing hypotheses about how information is distributed across an utterance, without treating it directly as a behavioral measure\. Moreover, surprisal is not model\-independent: different models produce substantially different estimates, and larger models do not necessarily better predict human reading timesoh2023transformer,kuribayashi2024psychometric\. It is also unclear where within a model surprisal should be measured, as estimates derived from intermediate layers can sometimes align better with human behavioral or neural responses than final\-layer estimateskuribayashi2025large\. FMs are also highly sensitive to contextual framing, and modest variations in input structure can substantially alter surprisal values, complicating interpretationjiang2024llms,oh2023transformer\. These considerations suggest that the appropriateness of surprisal as a linking hypothesis is ultimately empirical and task\-dependent\.

### 3\.3Prompting

In prompting, an FM is instructed to follow the same instructions as human participantsivanova2025evaluate\. Thus, prompting is quite a “proximal” linking hypothesis, enabling direct comparison of the generations of FMs with the productions of humans, reducing \(or even eliminating\) the need for relatively “distal” \(i\.e\., indirect\) linking hypothesespatel2021mapping,webb2023emergent\. Because FM generation is probabilistic, repeated prompting can produce a distribution over responses, enabling direct comparison to distributions of human productions or choice probabilities\. In the continuing digit\-matrix analogy examplewebb2023emergent, prompting would involve presenting the symbolic matrix and candidate completions exactly as shown to humans and asking the model to select the correct answer\. Under this mapping, successful performance indicates that the model reproduces the observable behavioral outcome without requiring further analysis of intermediate representations \(as with the similarity linking hypothesis\) or interpretation of probabilities \(as with surprisal\)\. Prompting has been used successfully in studies of incremental sentence interpretationcog\_sci\_garden\_pathand analogical reasoninglampinen2024language\.

Although prompting reduces the need for distal representational or probabilistic mappings, it shifts the interpretive risk to the elicitation procedure itself\. Model performance can be highly sensitive to prompt wording and formattingguo2024understanding,binz2023using,ullman2023large, raising questions about robustness\. Responses might simply reflect instruction\-following heuristics rather than the underlying cognitive mechanisms of interest, and systematic response biases \(e\.g\., yes\-response tendencies\) can inflate apparent alignmenthu2024language,dentella2023systematic\. Moreover, some research has questioned whether explicit metacognitive prompting \(e\.g\., asking models to judge grammatical acceptability\) faithfully captures the cognitive constructs of interest, especially for smaller\-scale modelshu2024auxiliary\. Finally, models can fail to respect experimentally defined response options \(e\.g\., producing explanations rather than forced\-choice responses\), complicating comparison to humanscai2024antagonistic\.

### 3\.4Process\-Trace Linking

A major recent development in foundation models is the rise of reasoning\-oriented systems that generate intermediate tokens before producing a final answer\. For many open\-weight systems, these traces and their token\-level properties are directly accessible, making them an increasingly practical target for analysis\. This has motivated a novel linking hypothesis: that properties of intermediate generation traces correspond to human processing dynamics\. For example, the length or structure of intermediate reasoning sequences \(e\.g\., number of tokens, number of reasoning steps, or branching in generated explanations\) might be mapped onto human response times or cognitive effort, with more extended traces reflecting increased difficulty\. In the digit\-matrix analogy examplewebb2023emergent, a model that generates longer intermediate reasoning sequences for more challenging Raven’s problems could be interpreted as exhibiting graded processing analogous to human analogical reasoning \(e\.g\., increasing response times\)\. This proposal aligns with long\-standing practices in cognitive psychology in which response times are treated as indirect measures of underlying cognitive processesluce1991response,anderson2009can,just2007organization\. More recently, research in large language models has proposed token\-level measures of inference\-time effort, such as reasoning\-token budgetsde2025cost,hu2026more, deep\-thinking ratioschen2026think, and sequential computation budgetsmadaan2025rethinking, as analogs of cognitive effort\.

Inferring cognitive processes from intermediate tokens raises interpretive challenges\. As with response\-time measures, differences in observable trace length do not necessarily imply differences in the targeted cognitive mechanismwhite2022need\. The causal impact of intermediate tokens on the final answer may be easy to overstate or misinterpret\. The literature on chain\-of\-thought faithfulness shows that generated rationales can systematically misrepresent the factors that drive model responses: models might rationalize biased or already\-determined answers, omit the true source of a decision, or expose answer\-relevant information in hidden states before the reasoning trace is producedturpin2023language,cox2026decoding\. Intermediate tokens might reflect instruction\-following conventions or stylistic artifacts \(e\.g\., requesting outputs as JSON files\) rather than intrinsic computational dynamicsjaroslawicz2025many\. Thus, process\-trace linking offers a promising yet still developing avenue for alignmentvankov2026correlations,hu2026thinking, and its validity hinges on whether intermediate token properties reliably capture meaningful cognitive computations\.

## 4Challenges of using FMs as cognitive models

Using FMs as cognitive models raises several inferential challenges\. The four\-stage framework helps make these challenges explicit by identifying where alignment claims can break down and what kinds of evidence can strengthen them\. These challenges do not undermine the potential scientific value of FMs; rather, they clarify the conditions under which model\-human correspondences can support explanatory claims\. In what follows, we articulate four challenges around a common structure: the inferential problem posed by FM\-human alignment and a corresponding way forward\.

### 4\.1Theoretical Underdetermination

Cognitive science experiments are not neutral measurement exercises; rather, they are designed to adjudicate between competing hypothesespopper2014conjectures,varadarajan2025capturing\. Experimental conditions are constructed to be diagnostically informative, with certain conditions carrying outsize theoretical weightporada2024controlled\. This challenge arises most directly inStages 2 and 3of the framework, where researchers specify linking hypotheses and evaluate behavioral correspondence\. An FM might show high aggregate goodness of fit while failing on the critical contrasts that motivated the original experiment\. Conversely, different linking hypotheses or analytic choices can yield different conclusions from the same model outputs\. In both cases, behavioral correspondence alone does not uniquely determine which cognitive mechanism, if any, the model instantiates\.

One way forward is to broaden the evidential base\. Alignment claims are stronger when a model explains not only one dataset or effect, but a broader set of theoretically related phenomena \(e\.g\., Centaurbinz2025foundation\), including diagnostic contrasts across different stimuli, tasks, populations, and developmental stages\. Broader coverage makes it harder for isolated empirical regularities, chance stimulus\-level patterns, or flexible linking hypothesis choices to explain the observed fit, strengthening alignment\.

### 4\.2Mechanistic Opacity

The second challenge is that FMs might reproduce human behavior without modeling the cognitive constructs that the experiment was designed to test\. In Marr’s terms, a model might align with humans at the computational level by solving the same task while diverging at the algorithmic level by relying on different representations, algorithms, or search strategiesmarr2010vision\. Mechanistic interpretability is a rapidly developing field attempting to identify functional subcomponents within modelskar2022interpretability,ferrando2024primer,milliere2025interventionist, and can therefore provide a secondary route throughStages 2 and 3: Researchers can ask whether model representations, activations, or processing dynamics correspond to theoretically relevant human measures\.Stage 4\-style comparisons, ablations, and causal interventions can then test whether those representational or computational features contribute to, and in some cases are necessary for, the observed behavioral alignment\.

FMs make this cross\-level inference difficult\. Their internal components, such as layers, attention heads, and activation patterns, do not map cleanly onto cognitive constructs such as working memory, parsing strategy, conceptual similarity, or analogical mappingcichy2019deep,mcgrath2023can,rumelhart1986pdp\. This difficulty reflects the longstanding principle of multiple realizability, and as a result, behavioral alignment can overstate explanatory equivalence: two systems might produce similar outputs while relying on different internal computationsputnam1967psychological,guest2023logical\. Although connecting lower\-level model mechanisms to higher\-level cognitive mechanisms remains difficultsharkey2025open, such analyses are essential for evaluating whether the model actually implements the constructs proposed by the cognitive theory\. Thus, mechanistic interpretability can contribute throughout the framework: it can support the formulation and evaluation of internal linking hypotheses inStages 2 and 3, whileStage 4comparisons and interventions can test how the identified mechanisms contribute causally to behavioral alignment\.

### 4\.3Training and Developmental Mismatch

The third challenge arises from differences between how FMs are trained and how humans learn\. FMs are typically optimized via next\-token prediction on large, static text corpora\. Human learning is much richer by contrast\. It unfolds interactively, under biological and environmental constraints, and is scaffolded by social interactions\. Thus, differences in learning objectives and inductive biases might shape model representations in ways that diverge from human developmental trajectorieswarstadt2022what,sutton2019bitter,oh2025model\.

Developmental alignment presents an additional challenge\. In principle, intermediate training checkpoints could be used to examine whether improvements in model performance track the developmental progressions observed in childrenfrank2023bridging,shah2024development,evanson2023language\. Open training efforts that release these checkpointsbiderman2023pythia,liu2023llm360,OLMocreate new opportunities for such analyses\. However, differences in the quantity, modality, and ordering of training data complicate establishing such mappingscuskley2024limitations,charpentier\-etal\-2025\-findings\. Even when model learning curves resemble developmental trajectories, such correspondences remain correlational and, by themselves, do not establish shared mechanisms of change\.

Thus, claims of developmental alignment must be evaluated with particular caution\. Stronger evidence for such claims requires causal manipulations of training regimes, data ordering, or architectural constraints to test whether such perturbations produce theoretically predicted changes in developmental trajectories\. Recent efforts such as the BabyLM challengechoshen2026babylmillustrate this opportunity by introducing developmentally motivated constraints on training data, exposure, and learning dynamics\. These perturbations extendStage 4\-style contrastive evaluation beyond architectural ablations to test how training conditions shape developmental trajectories\.

### 4\.4Variability and Population\-Level Alignment

A final challenge concerns variability\. FMs are typically trained on aggregate data and evaluated against average human performance\. This emphasis mirrors the*experimental*tradition in psychology, which often treats variability as noise\. By contrast, the*differential*tradition emphasizes individual differences as theoretically importantcarpenter1990one,Cronbach1957TheTD,underwood1975individual\. A growing body of work has examined whether FMs capture structured variation in human behaviormcduff2024cognitive,frank2025cognitive,fung2026individual\. Demonstrating alignment only at the group level means ignoring important dimensions of cognitive variability\. At the same time, recent work has begun exploring the simulation of behavioral heterogeneity via persona\-based promptingpark2022social,tseng2024two,aher2023using\. Such research is still in its infancy, and conflicting findings show that further exploration is requiredmilivcka2024large,salewski2024context\. Alignment claims that ignore individual differences in cognition risk overstating the generality of model\-human correspondences\.

More broadly, recent critiques caution that AI\-based simulations might create illusions of generalizability, as the behavior of models can reflect the demographic and cultural biases of their training data rather than the diversity of human populationscrockett2025ai\. A more informative approach is to treat variability as an observation to be explained rather than as noise to be averaged over\. Alignment claims should ask whether models capture structured variation across individuals, groups, developmental stages, or contexts, and whether this variation follows theoretically predicted patterns\. This reframes population\-level heterogeneity as a target of explanation, not merely a source of error\.

## 5Guidelines for Using FMs in Cognitive Science Research

The four\-stage framework presents a structured approach for evaluating FMs as candidate cognitive models\. Here, we distill this framework into operational principles for research practice\. These guidelines are not an exhaustive checklist of best practices, but a way of organizing inquiry so that alignment claims are developed with appropriate theoretical and empirical guardrails\.

#### 1\. Design evaluations around diagnostic measurement contrasts

The goal of alignment is not simply to reproduce average performance, but to test whether models capture theoretically diagnostic distinctions\. Experiments in cognitive science are typically designed to distinguish among competing hypotheses, with particular conditions carrying outsize explanatory weightpopper2014conjectures\. When evaluating FMs, researchers should therefore prioritize contrasts that adjudicate among alternative accounts\. As important as a model performing well overall \(Stage 3\) is whether its pattern of successes and failures discriminates among competing theories of the underlying cognitive mechanisms\.

#### 2\. Ensure that task adaptations preserve the diagnostic structure

When translating human experiments into model\-compatible formats, it is essential to preserve the logical distinctions of the original task\. Adaptations should retain the combinatorial structure across conditions that makes an experiment theoretically informative \(Stage 1\)\. For example, in adapting a visual reasoning task into text, the adapted version should preserve the relational dependencies, distractor options, and graded difficulty that make the original task diagnostic\. Superficially similar tasks might show low observed alignment because of inadvertently introduced formatting artifacts, interface constraints, or surface cues that change what the task measures\.

#### 3\. Triangulate across linking assumptions and behavioral measures

Because alignment depends on how model outputs are mapped onto human data, when possible, researchers should evaluate performance under multiple relevant linking hypotheses \(Stage 2\)\. This corresponds to the experimental strategy of triangulating a phenomenon by investigating it using multiple dependent measures: response times, error rates, verbal self\-reports, neuroimaging measures, etc\. Different linking hypotheses reveal different aspects of model behavior\. Convergent findings across linking strategies strengthen alignment claims, particularly when combined withStage 4\-style comparative evaluation across models and manipulations\.

#### 4\. Evaluate model alignment comparatively, not in isolation

Model\-human correspondence gains explanatory force when evaluated relative to alternative models and manipulated variants of the same model\. Comparisons across architectures, training regimes, scales, or ablations help identify which computational features are necessary for reproducing behavioral signatures \(Stage 4\)\. In addition, comparisons with simpler statistical baselines \(e\.g\.,nn\-gram modelsmichaelov2026n, regression modelsboehm2018using\), and established cognitive modelsbinz2025foundationclarify what FMs uniquely contribute\. Alignment is therefore inherently contrastive\.

#### 5\. Treat alignment as a starting point for causal or mechanistic investigation

Close behavioral fit \(Stage 3\) is necessary but not sufficient for explanatory adequacy\. Once correspondence has been established, the next step is to investigate whether the model’s internal representations and computations support the same mechanisms posited by the target theory\. This involves moving from whether the model aligns to how alignment is achieved within a model\. Such analysis might include ablation studies \(Stage 4\), representational probing, or examination of training dynamics to test whether disrupting specific components alters the observed behavior in theoretically predictable ways\. Demonstrating alignment should thus be viewed as a starting point for more stringent investigation, rather than an endpoint\.

## 6Conclusion

This paper has advocated for the careful use of Foundation Models as tools for investigating human cognition and its development\. The scientific value of FMs lies not in outperforming humans or achieving leaderboard success, but in their potential to illuminate which computational principles are sufficient, and perhaps necessary, for reproducing characteristic patterns of human cognition\. Here we distill and expand guidance from prior commentariesmcgrath2023can,frank2023bridging,warstadt2022what,mahowald2024dissociatinginto a framework for establishing the alignment of FMs and humans on cognitive tasks \(Figure[1](https://arxiv.org/html/2608.07812#S1.F1)\)\. Within this framework, alignment is not a property of models in isolation, but a claim embedded within explicit theoretical commitments, task adaptations, linking hypotheses, and comparative evaluation\. We offer the four\-stage framework not as a final, fixed prescription\. As models evolve, training regimes diversify, and multimodal systems become more prominent, our framework will have to be adapted\. This is important because FMs and their successors offer not just new tools for modeling behavior, but a new opportunity to sharpen our understanding of the structure, limits, and development of human thought\.

\\bmsection

\*Acknowledgments We thank Cory Shain, Abhijit Mahabal, Ali Emami, Harsh Lalai, and Carrie Bruce for thoughtful feedback and helpful discussions on earlier versions of this manuscript\.

## References

Similar Articles

Small Foundation Models of Human Cognition and Behaviour

arXiv cs.AI

This paper investigates whether small foundation models fine-tuned on human behavioral data can serve as cognitive proxies, finding that scale matters little in-distribution but larger models generalize better out-of-distribution.

Collective Intelligence with Foundation Models

arXiv cs.CL

This paper presents a multi-agent reasoning framework where multiple foundation models collaborate through structured critique and aggregation, demonstrating that model heterogeneity significantly improves step-wise reasoning accuracy and reduces variance across domains.

AI for Games in the Foundation Model Era

Hugging Face Daily Papers

This paper surveys the use of foundation models in game AI across roles like playing, modeling, design, and evaluation, highlighting transferability challenges and the need for game-specific validation.

Causal Foundation Models

Hugging Face Daily Papers

This paper introduces Causal Foundation Models, which use pretrained neural networks to estimate causal effects on new datasets via in-context learning without fine-tuning, providing a practical guide to this emerging field.