Small Foundation Models of Human Cognition and Behaviour
Summary
This paper trains 14 small language models (135M to 14B parameters) on Psych-101, a dataset of 10.7 million trial-level human choices, finding that small models suffice for in-distribution matching while larger models generalize better out-of-distribution. Diagnostics show that masking stimuli and feedback destroys most learned information, indicating choice history alone is insufficient.
View Cached Full Text
Cached at: 08/10/26, 02:15 PM
Paper page - Small Foundation Models of Human Cognition and Behaviour
Source: https://huggingface.co/papers/2608.05224
Abstract
Largelanguagemodelsfine-tunedonhumanbehaviouraldatahaveemergedasgeneral-purposecognitiveproxies,butthescalethisrequires,andwhetherthesemodelsprocesstaskstructureorexploitstatisticalshortcuts,remainopenquestions.Wetrainfourteenmodelsfrom135Mto14BparametersacrossfourarchitecturefamiliesonPsych-101,adatasetof10.7milliontrial-levelchoicesfrom160experiments.In-distribution,scalebarelymatters.Themodelsfallwithinanarrowband,asthoughagainstaceiling,and0.6Bto1Bparameterssufficetomatcha70Bbaselineonheld-outparticipants.Out-of-distribution,thatbandopensintoamarkedlysteeperscalinggradient,withlargermodelsclearlyadvantagedingeneralisationtonoveltaskstructure.Todeterminewhatinformationthesemodelsuse,weruntwodiagnostics.Weprogressivelystripfourpromptchannels--taskinstructions,experimentalstimuli,outcomefeedback,andchoicehistory--across27experiments,andpermutetrialorder.Maskingthecontentofstimuliandfeedbackdestroys75.7%oflearnedinformationandpushesmodelsbelowchance,demonstratingthatchoicehistoryalonedoesnotaccountforperformance.Permutationrevealsinvarianceontaskswithindependenttrialsbutsensitivitywheretrialorderisdeterminedbypriorresponses.Smallcognitivelyfine-tunedmodelsthereforeshowpromiseasnoiseceilingestimatorsforpsychologicalexperiments,thoughtheirscoperemainsboundedbytheparadigmsseenintraining.
Similar Articles
Small Foundation Models of Human Cognition and Behaviour
This paper investigates whether small foundation models fine-tuned on human behavioral data can serve as cognitive proxies, finding that scale matters little in-distribution but larger models generalize better out-of-distribution.
Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games
This paper proposes Equation-to-Behavior Prompting and reinforcement learning to guide large language models to simulate diverse human decision-making patterns in persuasion games, showing improved belief accuracy and training outcomes.
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
This paper evaluates nine open-weight small language models (135M to 3B parameters) on a structured benchmark and shows that parameter-efficient fine-tuning significantly improves accuracy, making them viable for local deployment in structured niche workloads.
Are small local models for automation a thing?
A Reddit user discusses the potential of small local language models (1B-4B parameters) for automation and scripting, and asks for resources focused on this use case.
@rohanpaul_ai: New Meta paper shows, small models may not be bad predictors of scale; they may just be getting under-tuned. Finds scal…
A new Meta paper reveals that small models can accurately predict scaling laws but require more extensive hyperparameter tuning. The study finds scaling laws emerge around 4M parameters and become clearer with proper tuning.