Small Foundation Models of Human Cognition and Behaviour

Hugging Face Daily Papers Papers

Summary

This paper trains 14 small language models (135M to 14B parameters) on Psych-101, a dataset of 10.7 million trial-level human choices, finding that small models suffice for in-distribution matching while larger models generalize better out-of-distribution. Diagnostics show that masking stimuli and feedback destroys most learned information, indicating choice history alone is insufficient.

Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.
Original Article
View Cached Full Text

Cached at: 08/10/26, 02:15 PM

Paper page - Small Foundation Models of Human Cognition and Behaviour

Source: https://huggingface.co/papers/2608.05224

Abstract

Largelanguagemodelsfine-tunedonhumanbehaviouraldatahaveemergedasgeneral-purposecognitiveproxies,butthescalethisrequires,andwhetherthesemodelsprocesstaskstructureorexploitstatisticalshortcuts,remainopenquestions.Wetrainfourteenmodelsfrom135Mto14BparametersacrossfourarchitecturefamiliesonPsych-101,adatasetof10.7milliontrial-levelchoicesfrom160experiments.In-distribution,scalebarelymatters.Themodelsfallwithinanarrowband,asthoughagainstaceiling,and0.6Bto1Bparameterssufficetomatcha70Bbaselineonheld-outparticipants.Out-of-distribution,thatbandopensintoamarkedlysteeperscalinggradient,withlargermodelsclearlyadvantagedingeneralisationtonoveltaskstructure.Todeterminewhatinformationthesemodelsuse,weruntwodiagnostics.Weprogressivelystripfourpromptchannels--taskinstructions,experimentalstimuli,outcomefeedback,andchoicehistory--across27experiments,andpermutetrialorder.Maskingthecontentofstimuliandfeedbackdestroys75.7%oflearnedinformationandpushesmodelsbelowchance,demonstratingthatchoicehistoryalonedoesnotaccountforperformance.Permutationrevealsinvarianceontaskswithindependenttrialsbutsensitivitywheretrialorderisdeterminedbypriorresponses.Smallcognitivelyfine-tunedmodelsthereforeshowpromiseasnoiseceilingestimatorsforpsychologicalexperiments,thoughtheirscoperemainsboundedbytheparadigmsseenintraining.

View arXiv pageView PDFGitHub3Add to collection

Similar Articles

Small Foundation Models of Human Cognition and Behaviour

arXiv cs.AI

This paper investigates whether small foundation models fine-tuned on human behavioral data can serve as cognitive proxies, finding that scale matters little in-distribution but larger models generalize better out-of-distribution.

Are small local models for automation a thing?

Reddit r/LocalLLaMA

A Reddit user discusses the potential of small local language models (1B-4B parameters) for automation and scripting, and asks for resources focused on this use case.