Children, but not language models, show accelerating returns in word learning
Summary
The research demonstrates that children show accelerating returns in word learning, whereas language models exhibit constant proportional returns, explaining the data gap in training efficiency between humans and AI.
View Cached Full Text
Cached at: 08/19/26, 09:49 AM
# Children, but not language models, show accelerating returns in word learning
Source: [https://arxiv.org/html/2608.17120](https://arxiv.org/html/2608.17120)
###### Abstract
Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed\. Prior models describe vocabulary growth as evidence accumulation over time\. Here we show that the process is best characterized as*accelerating*accumulation: children learn more from each additional unit of linguistic experience than they did from the one before\. In contrast to children, language models – even those trained on child\-directed speech – do not accelerate\. Instead, they show constant proportional returns on new data, consistent with scaling laws\. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation\.
Children’s language grows rapidly during early childhood \(*1*,*2*\), and the “explosion” of expressive vocabulary is a signature human cognitive achievement \(*3*,*4*\)\. Providing a quantitative characterization of this process is an important scientific goal \(*5*–*9*\), but until recently, language acquisition could only be studied in human learners\. Language models \(LMs\) show increasing linguistic sophistication, however, allowing them to become a model system for the study of human language \(*10*,*11*\)\.
The most successful LMs differ from human learners in an important way: they are trained on vastly more data than human learners receive during childhood\. While even open\-source models are routinely trained on trillions of words of data \(*12*\), children hear on the order of 100k – 1M words of input per month, with 30\-40M words as an approximate upper bound by age three \(*13*,*14*\)\. Why do models require so many more orders of magnitude of training data to achieve basic linguistic competence?
Here we study this “data gap” by comparing word learning as a function of training data in both children and LMs\. Vocabulary is a good target for comparison because the total number of different words a child produces is an important measure of expressive abilities that is also tightly coupled to grammatical development and other aspects of early language \(*1*,*15*\)\. Vocabulary size varies widely between children of the same age \(*1*\), and relative vocabulary size is a predictor of later academic success \(*16*,*17*\)\. Large longitudinal vocabulary datasets using the MacArthur\-Bates Communicative Development Inventories \(CDIs\), a popular parent\-report form, make it possible to characterize this growth even in young children \(*1*,*18*,*19*\)\.
We compare how vocabulary grows with input data in children and LMs\. We begin by developing a psychometric model of vocabulary growth in children \(an “accelerating accumulator”\), which shows a good fit to population variation across children\. This model suggests that children’s accumulation of evidence for words, not just their observed vocabulary size, accelerates across early childhood\. This acceleration leads to an increasing “return” on learning from linguistic experience: the same experience counts for more later in development\. In contrast to children, LMs do not accelerate: they show constant proportional returns on additional data\. This analysis exposes a critical disanalogy between LMs and children and suggests new avenues for the study of both\.
## An accelerating accumulator model of vocabulary growth
Recent models describe children’s vocabulary growth as a process of accumulation \(*8*,*9*,*20*,*21*\)\. Each distinct word \(type\) can be conceptualized as a “bucket,” and every token of language input then is a “drop” that falls into its respective bucket\. When each bucket is full, the word is learned \(*9*,*20*\)\. Accumulator models have been used to recover interpretable differences in “bucket size,” often based on word properties including concreteness and syntactic category \(*8*,*9*\)\. They also provide a graded explanation for the widely observed “vocabulary spurt” \(*20*,*22*,*23*\)\. Children’s*observed vocabulary*\(in terms of count\) grows supralinearly with age \(*1*,*18*\), but this qualitative pattern is consistent with many possible underlying mechanisms \(*20*\)\.
We begin an accumulator model that operates over ratio\-scale variables with real units: words \(tokens heard\) per hour and total words \(types\) in the child’s vocabulary \(*9*\)\. The presence of a true zero in each of these variables means that they can be log transformed, which is critical for capturing the shape of their growth over development \(*24*,*25*\)\. This accumulator can be written as a Rasch model from psychometrics \(*26*\), which assumes that each childiihas a learning abilityθi\\theta\_\{i\}and each wordjjhas a difficultyδj\\delta\_\{j\}\.111Prior work has focused on estimating properties of individual words \(*8*,*9*,*27*\), revealing that there is predictable word\-level difficulty information but that only a modest portion of the variation between words is carried by frequency information\. Words can be hard for a child to acquire because they vary also in the complexity and concreteness of their meanings, their syntactic category and syntactic context, their phonology, and many other factors\. In the current work, we estimate word difficulty as a free parameter, and focus instead on variation between children\.The probability that childiiproduces wordjjis:
P\(wi,j=1∣θi,δj\)=exp\(θi−δj\)1\+exp\(θi−δj\)\.P\(w\_\{i,j\}=1\\mid\\theta\_\{i\},\\delta\_\{j\}\)=\\frac\{\\exp\(\\theta\_\{i\}\-\\delta\_\{j\}\)\}\{1\+\\exp\(\\theta\_\{i\}\-\\delta\_\{j\}\)\}\.
We then describe the child’s ability as
θi\(t\)=ξi\+κilog\(t/a0\)\+log\(H\),\\theta\_\{i\}\(t\)=\\xi\_\{i\}\+\\kappa\_\{i\}~\\log\(t/a\_\{0\}\)\+\\log\(H\),
which contains three terms: an interceptξi\\xi\_\{i\}describing the child’s baseline learning efficiency, an acceleration term,κilog\(t/a0\)\\kappa\_\{i\}~\\log\(t/a\_\{0\}\)that describes the rate of accumulation across developmental time \(tt, relative to an anchora0a\_\{0\}\), and a constantHHfor the number of waking hours over which accumulation occurs\.
This model is an “accelerating accumulator” model, in which accumulating language “counts for more” later in development;κ\\kappais the key exponent describing developmental scaling\. In contrast, an accumulator withκ=1\\kappa=1would beθi\(t\)=ξ\+log\(t/a0\)\+log\(H\)\\theta\_\{i\}\(t\)=\\xi\+log\(t/a\_\{0\}\)\+log\(H\)\. This “pure accumulator” yields linear returns on experience, whileκ\>1\\kappa\>1shows acceleration\. Underκ=1\\kappa=1, the size of the “drops” filling the buckets remains constant across development; ifκ\>1\\kappa\>1, later “drops” are bigger\.
We posit that children vary on both their efficiency of accumulationξi\\xi\_\{i\}and their accelerationκi\\kappa\_\{i\}, creating a variant of a standard longitudinal growth model in which children vary on their slope and intercept \(*28*\) \(Figure[1](https://arxiv.org/html/2608.17120#Sx2.F1)A\)\. Our model is close to the formulations of \(*8*\) and \(*9*\), but has not yet been proposed \(supplementary text\)\.
## Fitting the accumulator model
CDIs are a widely used, reliable, and valid method for taking an in\-depth snapshot of children’s early vocabulary \(*1*,*18*,*19*\)\. Our analysis depends on the availability of longitudinal CDI data at the individual word level\. Summed vocabulary scores do not allow for the estimation of word difficultiesδ\\deltaand cross\-sectional data do not allow for model identification \(fig\. S1\)\. Thus, we use item\-level longitudinal data from children with three or more longitudinal administrations\. We analyze reported production, which is both measured over a wider age range and tends to be a more reliable measure than reported comprehension \(*1*\)\.
We use five longitudinal datasets of monolingual, typically developing children ages 8–36 months \(three from American English and one each from Norwegian and Japanese,Ntotal=1841N\_\{total\}=1841; see fig\. S2 and table S1\), downloaded from Wordbank \(*29*\)\. We fit the accelerating accumulator model \(M3\) to each dataset using Bayesian inference\. To test whether each model component was necessary, we additionally fit three model ablations, removing individual variation in acceleration \(withoutκi\\kappa\_\{i\}; M2\), removing individual variation in efficiency \(withoutξi\\xi\_\{i\}; M1\), and finally removing acceleration entirely \(withoutκ\\kappa; M0\)\.
The accelerating accumulator provided the best fit to children’s growth trajectories \(Figure[1](https://arxiv.org/html/2608.17120#Sx2.F1)B, see table S2\), with substantial improvements in item\-level expected log predictive density between the accelerating accumulator and the next best model \(LOO ELPD\>760\>760for all five datasets\)\. Indeed, fits from M3 corresponded closely to curves derived from a non\-parametric GAMLSS beta regression, a flexible model class used to produce percentile norms for the CDI instruments \(*19*\) \(fig\. S3\)\.
Estimates of acceleration were high for all datasets, robustly rejecting the pure accumulator \(κ=1\\kappa=1\) in every case: 10\.6 \[10\.5, 10\.8\] – 13\.3 \[13\.2, 13\.3\]\.κ\\kappaestimates were essentially unchanged using a larger dataset and in a hierarchical model pooling across all datasets \(table S5\)\. A 2PL IRT model achieved a better fit butκ\\kappaestimates remained far above the pure accumulator and word\-level age of acquisition estimates were very similar \(table S8\); the specifics of item selection on CDI forms had only minimal effects onκ\\kappaestimates \(fig\. S4\)\. Linear \(rather than logarithmic\) growth models did not provide better fit \(table S7\)\.
Variability is a ubiquitous feature of early language \(*1*,*2*\)\. The between\-child standard deviation in acceleration \(σb\\sigma\_\{b\}\) ranged from 3\.2 to 6\.6 across the five samples; the fitted population distribution placed 99\.5% of children above the pure\-accumulator value ofκ=1\\kappa=1\. Although the population fits showed substantial slope variation, typical CDI series are too sparse to predict individual children’s acceleration reliably beyond the population mean \(fig\. S5 and tables S3, S4 and S9\)\. Nevertheless, the addition of a population acceleration parameter improved individual trajectory prediction in 91\.9% of children when five or more prior CDIs were available\.
Figure 1:The accelerating accumulator provides a psychometric model of vocabulary growth, describing the children’s word learning across datasets and languages\.\(A\) Model predictions in latent\-ability \(θ\\theta\) space and, projected through the logistic item response model, in words\-produced \(CDI\) space\. Both pure accumulator \(κ=1\\kappa=1, dashed\) and accelerating accumulator predictions are plotted\. Greater efficiency \(ξ\\xi\) lifts the curve’s level, while greater acceleration \(κ\\kappa\) fans it open\. \(B\) Individual children’s observed trajectories \(gray\) for each of the five longitudinal datasets, with quantiles computed across the per\-child parameters of the fitted models \(colored percentile bands for M2 and M3; black lines for M0 and M1, which do not vary across individuals\)\.
## Accelerating accumulation implies less exposure is needed for later\-acquired words
Figure 2:Under the accelerating accumulator model, the number of exposures needed to learn a word decreases with age\.Plots show estimated cumulative exposures to a word, plotted by the age at which 50% of the children in our three English\-language datasets were estimated to produce the word\. Subplots show different word classes, and solid and dashed lines show simple linear fits from the accelerating \(M3\) and pure accumulator \(M0\) models, respectively; gray dots show the full distribution of words across all word classes\.The accelerating accumulator model captures a salient fact about language acquisition: the number of exposures necessary to learn a word decreases dramatically over childhood\. A toddler may have heard a word like “book” thousands of times before they say it; a few years later a preschooler may hear “ibex” a handful of times on a trip to the zoo and repeat it the next day\.
To quantify this phenomenon, we computed each word’s age of acquisition \(AoA\)AAfrom our fitted models, with AoA defined as the point at which 50% of children produce the word – in other words, the age at which child ability and word difficulty balance:A\(j\)=a0exp\(\(δj−logH−μξ\)/κ\)A\(j\)=a\_\{0\}\\exp\\\!\\big\(\(\\delta\_\{j\}\-\\log H\-\\mu\_\{\\xi\}\)/\\kappa\\big\)\. Then for each word, we estimated the cumulative word frequency for an individual child based on transcripts of speech to children \(*30*\)\.
Figure[2](https://arxiv.org/html/2608.17120#Sx3.F2)shows the relationship between cumulative exposures and age of acquisition for the accelerating accumulator\. Within each word class \(e\.g\., nouns, verbs, and function words\), the number of examples experienced before learning drops exponentially with development\.222Indeed, this analysis likely understates the degree of change in efficiency as it assumes that each exposure to a word is equally meaningful\. In fact, individual exposures vary widely in how much information they give\. Sometimes words are overheard rather than directed to the child, or lack physical context to allow for clear inference \(*21*,*31*\)\.In contrast, the fitted pure accumulator predicts a far shallower decline; it cannot reproduce the observed range of acquisition ages, predicting that many early acquired words will not be learned until adolescence \(fig\. S7 and table S10\)\.
## Language models do not accelerate
Figure 3:Language models \(LMs\) show no acceleration in their vocabulary growth\.\(A\) Density of the acceleration exponentκ\\kappafor both children \(per child, English and Norwegian\) and LMs \(by word\)\. \(B\) Variation in acceleration \(kappa\) across children and LMs\. Numbers show coefficients of variation and 95% CIs\. \(C\) Proportional return on training \(the ratio of decrease in loss to excess loss\) for both children and LMs\. Blue lines show LM trajectories; red lines show children’s fitted model trajectories from M3\. Three children are highlighted in dark red with horizontal uncertainty bands indicating \+/\- 1 SD in calibration of their input\. Logistic acquisition alone can produce some increase in proportional return due to ceiling effects; these asymptote at 1 \(gray dashed line\)\.Do LMs show the same signatures of acceleration and variability as children? To evaluate this question, we begin by building a bridge between well\-known “scaling laws” for large language models \(*32*\) and the model we describe here for vocabulary growth\. The lossLLof a neural network language model on a dataset sizeDDisL=E\+BDβL=E\+\\frac\{B\}\{D^\{\\beta\}\}, whereBBandβ\\betaare the free parameters of a power law andEEis the entropy of the text being modeled \(*32*\)\. This power law scaling behavior can emerge from a set of independently learned “quanta” that are themselves power\-law distributed \(*33*\); the words of natural language follow such a power law and are potential quanta of this type\.
To trace the emergence of individual words in LMs, we fit sigmoidal curves to the trajectory of surprisal values \(negative log probability of individual words\) across training \(following*34*\)\. These sigmoids provide a link to the logistic curves in the Rasch model we used with children\. The slope of the sigmoid for each word has the same units as our acceleration parameterκ\\kappa– it captures how fast evidence for that particular word accumulates with training data\.
We computed the distribution of these slopes on GPT\-2 models \(*35*\) trained on 24M words of English child\-directed speech from the CHILDES archive \(*10*,*30*,*36*\), approximating a three\-year\-old child’s total language input \(fig\. S9\)\. Per\-wordκ\\kappaestimates computed across training checkpoints clustered around 1 \(see Figure[3](https://arxiv.org/html/2608.17120#Sx4.F3)A, “training” sequence\): knowledge about individual words accumulated gradually\. This analysis computes learning over multiple training passes through the same dataset, however\. We thus trained “developmental” sequences of GPT\-2 models\. Each model in a sequence was trained to convergence on an increasingly larger subsample of the CHILDES data, ranging from \.5M words to 24M words\.κ\\kappavalues computed across these runs also showed no evidence of acceleration\. In contrast, when we applied the same by\-word estimator to data from children, we found highκ\\kappavalues \(fig\. S8 and table S11\)\.κ\\kappavariation was far higher in children than LMs, even when training data were varied \(Figure[3](https://arxiv.org/html/2608.17120#Sx4.F3)B\)\.
To make a holistic comparison between the growth of individual models \(rather than words\) and individual children, we computed learners’ return on new data\. Scaling law analyses describe the change in a model’s lossLLwith respect to the fundamental entropyEEof the training data\. These can be manipulated to compute a*proportional return*, describing the total amount of remaining possible loss that has been removed by a particular amount of data\. To compute this quantity in children, we computed each child’s “loss” across all words in the vocabulary, which we aligned to children’s estimated input rate over time \(fig\. S6\)\. On this analysis, LMs had constant proportional returns, while individual children’s proportional returns increased \(Figure[3](https://arxiv.org/html/2608.17120#Sx4.F3)C\)\.
## Discussion
Children’s vocabulary growth over early childhood can be described as a process of accelerating accumulation\. Knowledge of individual words accumulates over time, but later in development, fewer exposures are needed to acquire a new word\. In contrast, LMs – even trained on data from children – show no such acceleration\. Converted to the same metric, children get increasing returns from additional experience, whereas conventional LMs show constant proportional returns\.
What explains the acceleration in children’s word learning? During the first three years, children’s learning, memory, and speed of processing all change substantially \(*24*,*37*–*39*\)\. In the language of our model, these changes would mean that older children’s “buckets” require fewer drops to fill them\. Changes in the speed and accuracy of word recognition, in particular, are deeply related to children’s growing vocabulary; faster processing speed allows children to extract more signal from each incoming utterance \(*24*,*39*\)\. All of these changes could lead to some degree of acceleration\.
A second source of acceleration comes from children “learning to learn”: using the language they know to become more effective learners\. Their increasingly sophisticated inferential abilities also allow them to extract progressively more information from the linguistic signal \(*40*,*41*\)\. They learn generalizations that guide their inferences based on regularities in how their language works \(*42*\)\. They also use the language they already know in the moment to extract more information from new utterances – leveraging word meanings \(*40*\), syntactic knowledge \(*43*\), and pragmatic inference \(*41*\) to make sophisticated inferences about gaps in their knowledge \(*3*,*5*\)\.
Simple changes in the quantity or ordering of linguistic input are likely not responsible for acceleration\. First, parents do not typically “fine\-tune” their language to children’s developmental level \(*44*,*45*\)\. Second, curricularization of children’s language input, either by age or by other heuristics for simplicity fails to increase LM performance \(*36*,*46*\)\. Still, broader changes in the richness of input or child\-caregiver interaction could facilitate acceleration; these possibilities should be explored\.
In contrast to children, LMs do not mature developmentally, and it is unclear if standard pre\-training provides an analogue to children’s growing inferential abilities\. While in\-context learning abilities emerge, they typically do so only with vastly more data than children receive \(*47*,*48*\)\. And few\-shot rapid word learning may emerge LMs via meta\-learning, rather than pre\-training \(*49*\)\. Establishing linking hypotheses for comparison and alignment of children and LMs is thus a critical goal for further work \(*50*\)\.
The accelerating accumulator is a parsimonious description of variation in children’s vocabulary growth and it accords with the general structure of the learning problem faced by children\. But our evidence is based on statistical fit and model comparison, rather than a causal manipulation\. More work is required to test the accumulator as a mechanistic account\. Further, here we relied on parent report data from the CDI\. While CDIs have been extensively validated as a measure of language ability \(*1*,*18*,*19*\), they include variation due to changes in overall verbal production ability as well as vocabulary size\. Linking our model to direct measures of vocabulary is an important next step\.
In sum, our study identifies acceleration in learning from experience as a fundamental difference between children and LMs\. Characterizing the nature of children’s accelerating returns and understanding the circumstances under which it can arise in artificial systems should thus be a focus for future investigation\.
## Acknowledgments
Funding: Computing for this project was performed on the Sherlock and Marlowe clusters\. We would like to thank Stanford University and Stanford Research Computing for providing computational resources and support that contributed to these research results\. Funds for computational experiments were also provided by a gift from Meta to support research on children’s home language environment\.Author contributions: M\.C\.F designed and performed all research and wrote the paper\. Claude Code was used for code generation, visualization, and simulation infrastructure\.Competing interests: The author declares no competing interests\.Data, code, and materials availability: Reproducible code for this paper is available at[https://github\.com/mcfrank/acceleration](https://github.com/mcfrank/acceleration)\. All model fits are archived at[https://redivis\.com/datasets/datapages\.acceleration:a1c7](https://redivis.com/datasets/datapages.acceleration:a1c7)\. Data were retrieved from Wordbank using thewordbankrpackage \(*29*\) and CHILDES using thechildesrpackage \(*51*\)\. All trained LMs are available at[https://huggingface\.co/mcxfrank/childes\-gpt2\-ladder](https://huggingface.co/mcxfrank/childes-gpt2-ladder)and[https://huggingface\.co/mcxfrank/gpt2\-composition\-control](https://huggingface.co/mcxfrank/gpt2-composition-control)\.
## List of Supplementary Materials
Materials and Methods
Supplementary Text
Figs\. S1 to S9
Tables S1 to S11
## References
## References
- 11\.
- 2M\. C\. Frank, M\. Braginsky, D\. Yurovsky, V\. A\. Marchman,*Variability and Consistency in Early Language Learning: The Wordbank Project*\(MIT Press, 2021\)\.
- 2\. E\. Kidd, S\. Donnelly, Individual differences in first language acquisition\.*Annual review of linguistics*6, 319–340 \(2020\)\.
- 3\. P\. Bloom,*How Children Learn the Meanings of Words*\(MIT press, 2002\)\.
- 4\. E\. S\. Spelke, What makes us smart? Core knowledge and natural language\.*Language in mind: Advances in the study of language and thought*277, 311 \(2003\)\.
- 5\. M\. C\. Frank, N\. D\. Goodman, J\. B\. Tenenbaum, Using speakers’ referential intentions to model early cross\-situational word learning\.*Psychological science*20, 578–585 \(2009\)\.
- 6\. B\. McMurray, J\. S\. Horst, L\. K\. Samuelson, Word learning emerges from the interaction of online referent selection and slow associative learning\.*Psychological review*119, 831 \(2012\)\.
- 7\. J\. C\. Trueswell, T\. N\. Medina, A\. Hafri, L\. R\. Gleitman, Propose but verify: Fast mapping meets cross\-situational word learning\.*Cognitive psychology*66, 126–156 \(2013\)\.
- 8\. S\. Hidaka, A computational model associating learning process, word attributes, and age of acquisition\.*PloS one*8\(2013\)\.
- 9\. G\. Kachergis, V\. A\. Marchman, M\. C\. Frank, Toward a “standard model” of early language learning\.*Current Directions in Psychological Science*31, 20–27 \(2022\)\.
- 10\. P\. A\. Huebner, E\. Sulem, C\. Fisher, D\. Roth, “[BabyBERTa: Learning more grammar with small\-scale child\-directed language](https://doi.org/10.18653/v1/2021.conll-1.49)” in*Proceedings of the 25th Conference on Computational Natural Language Learning \(CoNLL\)*\(2021\), pp\. 624–646\.
- 11\. R\. Futrell, K\. Mahowald, How linguistics learned to stop worrying and love the language models\.*Behavioral and Brain Sciences*, 1–98 \(2025\)\.
- 12\. L\. Gao, S\. Biderman, S\. Black, L\. Golding, T\. Hoppe, C\. Foster, J\. Phang, H\. He, A\. Thite, N\. Nabeshima, others, The pile: An 800gb dataset of diverse text for language modeling\.*arXiv preprint arXiv:2101\.00027*\(2020\)\.
- 13\. M\. C\. Frank, Bridging the data gap between children and large language models\.*Trends in Cognitive Sciences*27, 990–992 \(2023\)\.
- 14\. A\. Warstadt, A\. Mueller, L\. Choshen, E\. Wilcox, C\. Zhuang, J\. Ciro, R\. Mosquera, B\. Paranjabe, A\. Williams, T\. Linzen, others, “Findings of the BabyLM challenge: Sample\-efficient pretraining on developmentally plausible corpora” in*Proceedings of the Babylm Challenge at the 27th Conference on Computational Natural Language Learning*\(2023\), pp\. 1–34\.
- 15\. E\. Bates, V\. Marchman, D\. Thal, L\. Fenson, P\. Dale, J\. S\. Reznick, J\. Reilly, J\. Hartung, Developmental and stylistic variation in the composition of early vocabulary\.*Journal of child language*21, 85–123 \(1994\)\.
- 16\. V\. A\. Marchman, A\. Fernald, Speed of word recognition and vocabulary knowledge in infancy predict cognitive and language outcomes in later childhood\.*Developmental science*11, F9–F16 \(2008\)\.
- 17\. H\. W\. Catts, M\. E\. Fey, J\. B\. Tomblin, X\. Zhang, A longitudinal investigation of reading outcomes in children with language impairments\.*Journal of speech, Language, and hearing Research*45, 1142–1157 \(2002\)\.
- 18\. L\. Fenson, P\. S\. Dale, J\. S\. Reznick, E\. Bates, D\. J\. Thal, S\. J\. Pethick, M\. Tomasello, C\. B\. Mervis, J\. Stiles, Variability in early communicative development\.*Monographs of the society for research in child development*, i–185 \(1994\)\.
- 19\. V\. A\. Marchman, P\. Dale, L\. Fenson,*MacArthur\-Bates Communicative Development Inventories User’s Guide and Technical Manual, Third Edition*\(Brookes Publishing, 2021\)\.
- 20\. B\. McMurray, Defusing the childhood vocabulary explosion\.*Science*317, 631–631 \(2007\)\.
- 21\. F\. Mollica, S\. T\. Piantadosi, How data drive early word learning: A cross\-linguistic waiting time analysis\.*Open Mind*1, 67–77 \(2017\)\.
- 22\. J\. Ganger, M\. R\. Brent,[Reexamining the vocabulary spurt](https://doi.org/10.1037/0012-1649.40.4.621)\.*Developmental Psychology*40, 621–632 \(2004\)\.
- 23\. M\. Gómez Díaz, L\. Fibla, R\. K\.\-Y\. Tsui, K\. Byers\-Heinlein,[Testing theories of the vocabulary spurt with monolingual and bilingual infants](https://doi.org/10.1037/dev0001777)\.*Developmental Psychology*60, 1357–1371 \(2024\)\.
- 24\. M\. C\. Frank, V\. A\. Marchman, C\. A\. Bergey, V\. Boyce, M\. Braginsky, G\. Kachergis, J\. Mankewitz, S\. Meylan, B\. Prystawski, N\. Ram, others, Continuous developmental changes in word recognition support language learning across early childhood\.*eLife*14\(2026\)\.
- 25\. R\. Kail, Processing time declines exponentially during childhood and adolescence\.*Developmental psychology*27, 259 \(1991\)\.
- 26\. G\. Rasch,*Probabilistic Models for Some Intelligence and Attainment Tests*\(Danish Institute for Educational Research, Copenhagen, Denmark, 1960\)\.
- 27\. J\. C\. Goodman, P\. S\. Dale, P\. Li, Does frequency count? Parental input and the acquisition of vocabulary\.*Journal of child language*35, 515–531 \(2008\)\.
- 28\. K\. J\. Grimm, N\. Ram, R\. Estabrook,*Growth Modeling: Structural Equation and Multilevel Modeling Approaches*\(Guilford Publications, 2016\)\.
- 29\. M\. C\. Frank, M\. Braginsky, D\. Yurovsky, V\. A\. Marchman, Wordbank: An open repository for developmental vocabulary data\.*Journal of Child Language*44, 677–694 \(2017\)\.
- 30\. B\. MacWhinney,*The CHILDES Project: Tools for Analyzing Talk*\(Psychology Press, 2000\)vol\. 1\.
- 31\. E\. A\. Cartmill, B\. F\. Armstrong III, L\. R\. Gleitman, S\. Goldin\-Meadow, T\. N\. Medina, J\. C\. Trueswell, Quality of early parent input predicts child vocabulary 3 years later\.*Proceedings of the National Academy of Sciences*110, 11278–11283 \(2013\)\.
- 32\. J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, others, Training compute\-optimal large language models\.*arXiv preprint arXiv:2203\.15556*10\(2022\)\.
- 33\. E\. J\. Michaud, Z\. Liu, U\. Girit, M\. Tegmark, The quantization model of neural scaling \(2024\)\.[https://arxiv\.org/abs/2303\.13506](https://arxiv.org/abs/2303.13506)\.
- 34\. T\. A\. Chang, B\. K\. Bergen,[Word acquisition in neural language models](https://doi.org/10.1162/tacl_a_00444)\.*Transactions of the Association for Computational Linguistics*10, 1–16 \(2022\)\.
- 35\. A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, I\. Sutskever,[Language models are unsupervised multitask learners](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)\.*OpenAI*\(2019\)\.
- 36\. S\. Y\. Feng, N\. D\. Goodman, M\. C\. Frank, “Is child\-directed speech effective training data for language models?” in*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, Y\. Al\-Onaizan, M\. Bansal, Y\.\-N\. Chen, Eds\. \(Association for Computational Linguistics, Miami, Florida, USA, 2024;[https://aclanthology\.org/2024\.emnlp\-main\.1231/](https://aclanthology.org/2024.emnlp-main.1231/)\), pp\. 22055–22071\.
- 37\. K\. Hartshorn, C\. Rovee\-Collier, P\. Gerhardstein, R\. S\. Bhatt, T\. L\. Wondoloski, P\. Klein, J\. Gilch, N\. Wurtzel, M\. Campos\-de\-Carvalho, The ontogeny of long\-term memory over the first year\-and\-a\-half of life\.*Developmental Psychobiology: The Journal of the International Society for Developmental Psychobiology*32, 69–89 \(1998\)\.
- 38\. P\. J\. Bauer, J\. A\. Wenner, P\. L\. Dropik, S\. S\. Wewerka, M\. L\. Howe, Parameters of remembering and forgetting in the transition from infancy to early childhood\.*Monographs of the Society for Research in Child Development*, i–213 \(2000\)\.
- 39\. A\. Fernald, V\. A\. Marchman, Individual differences in lexical processing at 18 months predict vocabulary growth in typically developing and late\-talking toddlers\.*Child development*83, 203–222 \(2012\)\.
- 40\. E\. M\. Markman, G\. F\. Wachtel, Children’s use of mutual exclusivity to constrain the meanings of words\.*Cognitive Psychology*20, 121–157 \(1988\)\.
- 41\. M\. Bohn, M\. C\. Frank, The pervasive role of pragmatics in early language\.*Annual Review of Developmental Psychology*1, 223–249 \(2019\)\.
- 42\. L\. B\. Smith, S\. S\. Jones, B\. Landau, L\. Gershkoff\-Stowe, L\. Samuelson,[Object name learning provides on\-the\-job training for attention](https://doi.org/10.1111/1467-9280.00403)\.*Psychological Science*13, 13–19 \(2002\)\.
- 43\. L\. Gleitman, The structural sources of verb meanings\.*Language acquisition*1, 3–55 \(1990\)\.
- 44\. E\. L\. Newport, H\. Gleitman, L\. R\. Gleitman, Mother, i’d rather do it myself\.*Sentence first, arguments afterward: Essays in language and learning*141\(2020\)\.
- 45\. D\. P\. Hayes, M\. G\. Ahrens, Vocabulary simplification for children: A special case of “motherese”?*Journal of child language*15, 395–410 \(1988\)\.
- 46\. R\. D\. Martinez, Z\. Goriely, H\. McGovern, C\. Davis, A\. Caines, P\. Buttery, L\. Beinborn, “CLIMB–curriculum learning for infant\-inspired model building” in*Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning*\(2023\), pp\. 112–127\.
- 47\. J\. Wei, Emergent abilities of large language models\.*arXiv preprint arXiv:2206\.07682*\(2022\)\.
- 48\. Q\. Dong, L\. Li, D\. Dai, C\. Zheng, J\. Ma, R\. Li, H\. Xia, J\. Xu, Z\. Wu, B\. Chang, others, “A survey on in\-context learning” in*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*\(2024\), pp\. 1107–1128\.
- 49\. W\. Wang, G\. Jiang, T\. Linzen, B\. Lake, “Rapid word learning through meta in\-context learning” in*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, V\. Peng, Eds\. \(Association for Computational Linguistics, Suzhou, China, 2025;[https://aclanthology\.org/2025\.emnlp\-main\.1631/](https://aclanthology.org/2025.emnlp-main.1631/)\), pp\. 32038–32073\.
- 50\. M\. C\. Frank, N\. D\. Goodman, Cognitive modeling using artificial intelligence\.*Annual Review of Psychology*77\(2025\)\.
- 51\. A\. Sanchez, S\. C\. Meylan, M\. Braginsky, K\. E\. MacDonald, D\. Yurovsky, M\. C\. Frank, Childes\-db: A flexible and reproducible interface to the child language data exchange system\.*Behavior Research Methods*51, 1928–1941 \(2019\)\.
- 52\. D\. J\. Thal, V\. A\. Marchman, J\. B\. Tomblin, “Late talking toddlers: Characterization and prediction of continued delay” in*Late Talkers: Language Development, Interventions, and Outcomes*, L\. Rescorla, P\. Dale, Eds\. \(Brookes Publishing, Baltimore, MD, 2013\)\.
- 53\. H\. G\. Simonsen, K\. E\. Kristoffersen, D\. Bleses, S\. Wehberg, R\. N\. Jørgensen,[The Norwegian Communicative Development Inventories: Reliability, main developmental trends and gender differences](https://doi.org/10.1177/0142723713510997)\.*First Language*34, 3–23 \(2014\)\.
- 54\. H\. Hagihara, M\. Barbir, M\. Ishibashi, Y\. Kanakogi, M\. Kato, I\. Lovcevic, Y\. Lu, Y\. Minagawa, Y\. Moriguchi, M\. Sakagami, Y\. Shinya, H\. Yamamoto, S\. Tsuji, A sharable merged dataset on Japanese children’s vocabulary measures using the Japanese MacArthur–Bates Communicative Development Inventory \(2023\)\.[https://doi\.org/10\.17605/osf\.io/s5ydw](https://doi.org/10.17605/osf.io/s5ydw)\.
- 55\. B\. Carpenter, A\. Gelman, M\. D\. Hoffman, D\. Lee, B\. Goodrich, M\. Betancourt, M\. Brubaker, J\. Guo, P\. Li, A\. Riddell, Stan: A probabilistic programming language\.*Journal of statistical software*76, 1–32 \(2017\)\.
- 56\. D\. E\. Sperry, L\. L\. Sperry, P\. J\. Miller,[Reexamining the verbal environments of children from different socioeconomic backgrounds](https://doi.org/10.1111/cdev.13072)\.*Child Development*90, 1303–1318 \(2019\)\.
- 57\. B\. Hart, T\. R\. Risley,*Meaningful Differences in the Everyday Experience of Young American Children*\(Brookes, 1995\)\.
- 58\. M\. Y\. Hu, A\. Mueller, C\. Ross, A\. Williams, T\. Linzen, C\. Zhuang, R\. Cotterell, L\. Choshen, A\. Warstadt, E\. G\. Wilcox, “Findings of the second BabyLM challenge: Sample\-efficient pretraining on developmentally plausible corpora” in*The 2nd BabyLM Challenge at the 28th Conference on Computational Natural Language Learning*\(2024\), pp\. 1–21\.
- 59\. S\. Diao, Y\. Yang, Y\. Fu, X\. Dong, D\. Su, M\. Kliegl, Z\. Chen, P\. Belcak, Y\. Suhara, H\. Yin, M\. Patwary, Y\. Lin, J\. Kautz, P\. Molchanov, Nemotron\-CLIMB: CLustering\-based iterative data mixture bootstrapping for language model pre\-training \(2025\)\.[https://arxiv\.org/abs/2504\.13161](https://arxiv.org/abs/2504.13161)\.
- 60\. C\. Mitchell, B\. McMurray, On leveraged learning in lexical acquisition and its relationship to acceleration\.*Cognitive Science*33, 1503–1523 \(2009\)\.
- 61\. G\. Kachergis, V\. A\. Marchman, P\. S\. Dale, J\. Mankewitz, M\. C\. Frank, Online computerized adaptive tests of children’s vocabulary development in english and mexican spanish\.*Journal of Speech, Language, and Hearing Research*65, 2288–2308 \(2022\)\.
- 62\. K\. A\. Adams, V\. A\. Marchman, E\. C\. Loi, M\. D\. Ashland, A\. Fernald, H\. M\. Feldman, Caregiver talk and medical risk as predictors of language outcomes in full term and preterm toddlers\.*Child Development*89, 1674–1690 \(2018\)\.
- 63\. B\. Long, R\. Z\. Sparks, V\. Xiang, S\. Stojanov, Z\. Yin, G\. Keene, A\. W\. M\. Tan, S\. Y\. Feng, A\. Nag, C\. Zhuang, V\. A\. Marchman, D\. L\. K\. Yamins, M\. C\. Frank, “[The BabyView dataset: High\-resolution egocentric videos of infants’ and young children’s everyday experiences](https://doi.org/10.32470/vbbjtb0)” in*Proceedings of the Conference on Cognitive Computational Neuroscience*\(2025\)\.
- 64\. A\. Fernald, V\. A\. Marchman, A\. Weisleder, SES differences in language processing skill and vocabulary are evident at 18 months\.*Developmental science*16, 234–248 \(2013\)\.
- 65\. S\. Egan\-Dailey, E\. Bergelson, Early child measures outpredict input measures of preschool language skills in US english learners\.*Developmental psychology*\(2025\)\. Supplemental Information
## Materials and Methods
### Vocabulary data
All longitudinal datasets with three or more observations per child were included\. English \(Thal\) data are from \(*52*\), Norwegian is from \(*53*\), and Japanese is from \(*54*\); English \(Marchman\) and English \(Smith\) are unpublished datasets available through Wordbank \(*29*\)\. We applied a quality filter to exclude administrations that showed either decreases of more than 25% of productive vocabulary or increases that exceeded 40% of total vocabulary per month\. This excluded 36 children in total \(1\.9%\), nearly all from the English \(Marchman\) dataset \(16%; Norwegian 0\.5%, none from the others\) \(SI: Data Exclusions\)\.
### Bayesian models
Using Stan \(*55*\), Bayesian models were fit separately to each dataset per the specification described in the text\. The model was parameterized to capture individual children’s variation from the population mean such thatξi=μξ\+ai\\xi\_\{i\}=\\mu\_\{\\xi\}\+a\_\{i\}andκi=μκ\+bi\\kappa\_\{i\}=\\mu\_\{\\kappa\}\+b\_\{i\}, where\(ai,bi\)∼MVN\(𝟎,Σ\)\(a\_\{i\},b\_\{i\}\)\\sim\\mathrm\{MVN\}\(\\mathbf\{0\},\\Sigma\)\. To facilitate interpretation of the accumulator mechanism, we included the constantlogH=log\(365\)≈5\.90\\log H=\\log\(365\)\\approx 5\.90\(waking hours/month\) and the anchor agea0=18a\_\{0\}=18months \(solog\(t/a0\)=0\\log\(t/a\_\{0\}\)=0at 18 months\)\.
We chose weakly informative priors throughout, including on the efficiency and acceleration parametersμξ∼𝒩\(−6,5\)\\mu\_\{\\xi\}\\sim\\mathcal\{N\}\(\-6,\\ 5\)andμκ∼1\+𝒩\(0,5\)\\mu\_\{\\kappa\}\\sim 1\+\\mathcal\{N\}\(0,\\ 5\)\(centered around the pure accumulator model\) as well as the individual child standard deviationsσa∼Half\-Normal\(0,3\)\\sigma\_\{a\}\\sim\\mathrm\{Half\\text\{\-\}Normal\}\(0,\\ 3\)andσb∼Half\-Normal\(0,5\)\\sigma\_\{b\}\\sim\\mathrm\{Half\\text\{\-\}Normal\}\(0,\\ 5\)and the item difficultiesτδ∼Half\-Normal\(0,5\)\\tau\_\{\\delta\}\\sim\\mathrm\{Half\\text\{\-\}Normal\}\(0,\\ 5\)\. We also chose a prior that mildly regularized the correlations betweenaia\_\{i\}andbib\_\{i\}\(𝐋∼LKJ\(2\)\\mathbf\{L\}\\sim\\mathrm\{LKJ\}\(2\)\)\.
All models were fit using cmdstanr, with four independent chains of 1,000 warmup and 1,000 sampling iterations each \(4,000 post\-warmup draws\) and a target acceptance rate of 0\.9\. The primary series of models was additionally refit using 2,000 warmup and 8,000 sampling iterations\. We assessed convergence using rank\-normalized split\-R^\\hat\{R\}and bulk effective sample size\. For parameters we interpret \(primarilyκ\\kappa,σa\\sigma\_\{a\},σb\\sigma\_\{b\}\), we requiredR^<1\.01\\hat\{R\}<1\.01and bulk ESS above 400\.
### Learning efficiency
Age of acquisition for each word was computed based on the weighted average of word difficultiesδj\\delta\_\{j\}from the three fitted M3 models for English\. For the population level ofξ\\xiandκ\\kappa, we computed the earliest age at which there was at least a \.5 probability of producing the word\. Average exposures per month to each word were computed asNj=R⋅t50\(j\)⋅pjN\_\{j\}=R\\cdot t\_\{50\}\(j\)\\cdot p\_\{j\}, whereRRis the rate of input \(words/month\),t50\(j\)t\_\{50\}\(j\)is the word\-level age of acquisition computed above, andpjp\_\{j\}is the word’s corpus probability\. We used average input estimates derived from \(*56*\) and \(*57*\) for our estimate ofRR\(see SI: Input Estimation\), and computedpjp\_\{j\}from empirical word frequencies in all North American English corpora in childes\-db release 2021\.1\.
### Language model training
Following \(*36*\), we fit randomly\-initializedGPT\-2\-small\(124M parameters\) models, using the 24M\-word \(~47M\-token\) preparation of English\-language CHILDES data from that study\. Consistent with the previous work, we included children’s own utterances in the training data\. Children do in fact hear their own speech; more importantly, removing one side of every interaction would create disjoint, unnatural training data\. We also fit models to the 2nd BabyLM challenge dataset \(downloaded from[osf\.io/ad7qg](https://osf.io/ad7qg), with CHILDES data removed\) \(*58*\) and the ClimbMix dataset \(available at[https://huggingface\.co/datasets/karpathy/climbmix\-400b\-shuffle](https://huggingface.co/datasets/karpathy/climbmix-400b-shuffle)\) \(*59*\)\. We constructed three disjoint subsets of 24M words from each\.
For comparability, all data were tokenized using a standard BPE tokenizer trained on the full CHILDES dataset \(all 24M\)\. Surprisal was computed as the mean per\-word negative log likelihood over ~50 held\-out validation contexts \(last 128 tokens each\) for the 609 CDI words that were used in \(*34*\) and that appeared in the training data with sufficient frequency in the CHILDES held\-out contexts\. Surprisal values for sigmoid fits were read out from the epoch with the lowest validation loss\.
For the computation of surprisal across training runs \(“training”\), we trained models using 3 random seeds on the full dataset\. For the computation across different dataset sizes \(“development”\), the data were segmented into 18 nested datasets ranging in size from \.5M to the full 24M\. We then chose 10 random seed values and trained 10 models on each dataset for a total of 180 models\. For the computation of variability across model runs \(“composition”\), we ran 8 random seed values across 12 nested dataset sizes for three disjoint subsets of both BabyLM and ClimbMix, resulting in 8 \(seeds\) x 12 \(sizes\) x 3 \(subsets\) x 2 \(datasets\) = 576 distinct models\.
Models were trained for a total of 20 epochs with sequence length 1024 and batch size 8, using AdamW for optimization and learning rate of 1e\-4 with a linear decay, no warmup, weight decay of 0, and no in\-epoch shuffling\. Training was performed on NVIDIA A40 GPUs\.
### Proportional Return on Input
Proportional return on input was computed for models based on mean surprisalL\(D\)L\(D\)of the CDI words after training onDDword tokens of input\. This follows the form from \(*32*\) ofL\(D\)=E\+BD−βL\(D\)=E\+B\\,D^\{\-\\beta\}, withEEthe irreducible entropy floor andB,βB,\\betafit across budgets \(hereE≈2\.94E\\approx 2\.94nats,B≈312B\\approx 312,β≈0\.32\\beta\\approx 0\.32,R2=0\.998R^\{2\}=0\.998\)\. The excess lossL\(D\)−E=BD−βL\(D\)\-E=B\\,D^\{\-\\beta\}is the reducible part that training removes; it falls as a power law of slope−β\-\\beta\. Marginal return is the loss removed peree\-fold of data,
R\(D\)=−dLdlnD=β\(L\(D\)−E\),R\(D\)=\-\\frac\{\\mathrm\{d\}L\}\{\\mathrm\{d\}\\ln D\}=\\beta\\,\\big\(L\(D\)\-E\\big\),
obtained by differentiating the scaling law\. We compute the ratio of this quantity to the excess loss, obtaining the proportional return: the share of the remaining reducible loss removed peree\-fold,γLM\(D\)=R\(D\)L\(D\)−E=β\.\\gamma\_\{\\text\{LM\}\}\(D\)=\\frac\{R\(D\)\}\{L\(D\)\-E\}=\\beta\.For a power law, this is constant:γLM=β≈0\.32\\gamma\_\{\\text\{LM\}\}=\\beta\\approx 0\.32at every budget, recapitulating the scaling law\.
We can compute the same quantity for children\. We define a vocabulary production loss,Li\(t\)=meanj\[−logpij\(t\)\]L\_\{i\}\(t\)=\\operatorname\{mean\}\_\{j\}\[\-\\log p\_\{ij\}\(t\)\], whose floor is zero \(a known word is produced,p→1p\\to 1\)\. Its marginal return isRi\(t\)=−dLi/dlogt=κimeanj\(1−pij\)R\_\{i\}\(t\)=\-\\mathrm\{d\}L\_\{i\}/\\mathrm\{d\}\\log t=\\kappa\_\{i\}\\,\\operatorname\{mean\}\_\{j\}\(1\-p\_\{ij\}\), so the proportional return is
γi\(t\)=Ri\(t\)Li\(t\)=κimeanj\(1−pij\)meanj\(−logpij\),\\gamma\_\{i\}\(t\)=\\frac\{R\_\{i\}\(t\)\}\{L\_\{i\}\(t\)\}=\\kappa\_\{i\}\\,\\frac\{\\operatorname\{mean\}\_\{j\}\(1\-p\_\{ij\}\)\}\{\\operatorname\{mean\}\_\{j\}\(\-\\log p\_\{ij\}\)\},which rises toward the child’s ownκi\\kappa\_\{i\}as the vocabulary saturates\.
### Computing Acceleration in LMs
To compare acceleration in LMs and children, we link theκ\\kappaparameter from our accelerating accumulator model to the word\-by\-word learning performance of LMs\. We first focus on the slope for individual words rather than individual children\. Starting with the M3, differentiatingηij\(t\)\\eta\_\{ij\}\(t\)\(the probability of producing a particular word\) with respect tologt\\log tgivesκi\\kappa\_\{i\}: the logit\-probability that a child produces a word rises linearly in log\-age with slope equal to the child’s scaling exponent\. Thus, the per\-child sigmoid slope*is*κi\\kappa\_\{i\}, by construction\.
Next, to estimate the slope of acquisition of individual words, \(*34*\) compute surprisal for each word across LM training checkpoints and then fit a four\-parameter logistic,sw\(x\)=ℓw\+\(uw−ℓw\)/\(1\+exp\(\(x−mw\)/scalew\)\)s\_\{w\}\(x\)=\\ell\_\{w\}\+\(u\_\{w\}\-\\ell\_\{w\}\)/\(1\+\\exp\(\(x\-m\_\{w\}\)/\\mathrm\{scale\}\_\{w\}\)\), inx=log10x=\\log\_\{10\}training data\. Letpw=\(sw−ℓw\)/\(uw−ℓw\)p\_\{w\}=\(s\_\{w\}\-\\ell\_\{w\}\)/\(u\_\{w\}\-\\ell\_\{w\}\)be the fraction of that word’s surprisal reduction still outstanding\. The LM slope is thereforedlogit\(pw\)/dlnD=1/\(scalewln10\)d\\,\\mathrm\{logit\}\(p\_\{w\}\)/d\\ln D=1/\(\\mathrm\{scale\}\_\{w\}\\ln 10\)— a logit per unit of natural\-log input\. That is the same asκi\\kappa\_\{i\}for children, where the logit probability of producing a word rises in log\-age with slopeκi\\kappa\_\{i\}and the outstanding fraction is1−P\(produce\)1\-P\(\\text\{produce\}\)\.
## Supplementary Text
### Comparison with Prior Models
Our accelerating accumulator model has not been described in the literature in the exact form we use, but it is closely related to other models\. Here we describe these models as special cases of our model\. We write the latent ability of childiion wordjjat agettas
ηi,j\(t\)=logri\+logαi\+κilogt−δj,\\eta\_\{i,j\}\(t\)=\\log r\_\{i\}\+\\log\\alpha\_\{i\}\+\\kappa\_\{i\}\\log t\-\\delta\_\{j\},whererir\_\{i\}andαi\\alpha\_\{i\}decompose the termξi\\xi\_\{i\}in our model:rir\_\{i\}is the input rate for the child andαi\\alpha\_\{i\}is their learning efficiency\. Several prior accumulator models can then be seen as special cases in which one or more of these components is fixed:
- •\(*20*\) is a pure accumulator in whichκ=1\\kappa=1and children do not differ from one another \(σα=σκ=0\\sigma\_\{\\alpha\}=\\sigma\_\{\\kappa\}=0\), so both input rate and efficiency reduce to a constant, corresponding to our M0\.
- •\(*8*\) introduces a model in which ability grows astD\+1t^\{D\+1\}, creating acceleration\. This exponent is exactly ourκ\\kappa, and this model is equivalent to our M1\.
- •\(*9*\) present our M2: they allowα\\alphato vary across individuals, but notκ\\kappa\.
In addition, several other models in the literature can be described as variants of this framework:
- •\(*60*\) augment the pure M0 accumulator with a “leverage” term, in which words learned earlier facilitate the learning of later words\. Their central result is a negative one: leverage changes the shape and timing of vocabulary growth, but it cannot create acceleration that is not already present in the distribution of word difficulties\.
- •\(*21*\) define a pure accumulator model withκ=1\\kappa=1, but with an additional parameter relating to the speed of per\-word accumulation\. This model maps onto a combination of our M0 and the 2\-parameter logistic variant that we fit below\.
### Non\-identifiability in Cross\-Sectional Data
Since acceleration is a property of change over time, a single cross\-sectional observation from an individual child cannot identify the child’s acceleration\. As Figure[4](https://arxiv.org/html/2608.17120#Sx10.F4)shows, a single observation from a child could be accounted for by variation in eitherξi\\xi\_\{i\}\(efficiency\) orκi\\kappa\_\{i\}\(acceleration\)\.
Figure 4:Non\-identifiability of growth models in cross\-sectional data\. Panels show the same simulated data, reflecting the possibility of stable between\-child intercept differences \(A\) vs\. varying growth slopes \(B\)\.
### Datasets
Table[1](https://arxiv.org/html/2608.17120#Sx10.T1)gives the characteristics of each dataset used\. In the main text we focus on children with three or more administrations \(for more precise estimates ofκ\\kappa\), but below we also confirm the robustness of our claims including children with two or more administrations\.
Table 1:Characteristics of all datasets used in the paper\. N and ages reflect the analysis samples \(post\-exclusion\) for the two exclusion criteria \(3\+ and 2\+ administrations\)\. Adm/Chi gives the median administrations per child\.Age \(months\)CitationLanguageN \(3\+/2\+\)AdminsMeanMinMaxAdm/ChiThal \(Wordbank\)English \(American\)639 / 6531,91919\.012293Smith \(Wordbank\)English \(American\)152 / 31660722\.216304Marchman \(Wordbank\)English \(American\)162 / 2,13649920\.38303Simonsen et al\. \(2014\)Norwegian792 / 1,6304,77024\.18365Hagihara et al\. \(2023\)Japanese96 / 18728916\.511333Fig\.[5](https://arxiv.org/html/2608.17120#Sx10.F5)shows the full spaghetti plots for each dataset, highlighting in red those observations that failed our quality check and were excluded\. We speculate that these observations are due to either data entry issues or faulty correspondences between observations\. Such issues occur frequently in archival datasets and affect only a small portion of our data\.
Figure 5:Data\-quality exclusions\. Red lines show children excluded from the analysis entirely, with red dots marking flagged administrations; gray lines are a random sample of retained children, shown for reference\.
### Full Model Comparison Results
Table[2](https://arxiv.org/html/2608.17120#Sx10.T2)gives the full leave\-one\-out model comparison results across models for both the dataset with 3\+ administrations/child and the larger dataset with 2\+ administrations/child\.
Table 2:Bayesian leave\-one\-out \(LOO\) model comparison\. Each cell is the difference in expected log predictive density \(ELPD\) relative to the best model \(M3\), with the standard error of the difference in parentheses; more negative is worse\.Δ\\DeltaELPD vs M3 \(SE\)ModelThalSmithMarchmanNorwegianJapanese≥\\geq3 administrations \(main text\)M0\. accumulator \(κ=1\\kappa=1\)\-154,678 \(350\)\-102,675 \(340\)\-83,954 \(285\)\-130,396 \(369\)\-24,628 \(186\)M1\. \+ acceleration\-46,214 \(289\)\-62,282 \(309\)\-40,868 \(258\)\-70,888 \(336\)\-7,923 \(126\)M2\. \+ efficiency var\.\-7,880 \(133\)\-6,542 \(121\)\-4,878 \(110\)\-5,702 \(111\)\-761 \( 40\)M3\. \+ acceleration var\.00000≥\\geq2 administrations \(full longitudinal\)M0\. accumulator \(κ=1\\kappa=1\)\-155,002 \(351\)\-120,701 \(372\)\-125,833 \(355\)\-131,793 \(373\)\-53,239 \(260\)M1\. \+ acceleration\-45,613 \(287\)\-75,934 \(343\)\-66,770 \(332\)\-75,942 \(346\)\-21,922 \(197\)M2\. \+ efficiency var\.\-7,723 \(132\)\-7,795 \(129\)\-14,004 \(195\)\-6,047 \(115\)\-1,919 \( 65\)M3\. \+ acceleration var\.00000
### Individual\-level Cross\-validation
The model comparisons above report ELPD LOO values for held\-out data from individual item\-responses, which are nested within children\. To assess the value of M3 \(the accelerating accumulator\) in making fully out\-of\-sample predictions, we designed a prospective test\. As our data, we selected a subset of the Norwegian sample that included 371 children with≥6\\geq 6longitudinal CDIs\. We then fit models to the child’s most recentkkobservations and used these to predict the child’s next administration \(3 months later\)\. This design avoids confounding data density with the gap between the fitted and predicted observations\.
Results from this analysis are shown in Table[3](https://arxiv.org/html/2608.17120#Sx10.T3)\. The result is very sensitive to the value ofkkused\. Withk=2k=2, the fixed slope model \(M2\) clearly predicts better because the per\-child slope model \(M3\) fully parameterizes the two observations per child and absorbs measurement error as a result\. This effect decreases askkincreases until the two models are tied byk=5k=5\. We expect that with even denser data, a separate slope parameter would lead to individually better fits, but in practice the single slope model makes acceptable predictions\.
Importantly, because both M2 and M3 contain a population acceleration parameter, this analysis does not test for the presence of acceleration\. Instead it suggests that measuring the precise amount of acceleration each child is undergoing requires sustained, precise longitudinal measurement\.
Table 3:Prospective cross\-validation for Norwegian children with at least six administrations \(n = 371\)\. Each row describes a model fit to k administrations and scored on the child’s held\-out final sitting\. dELPD/child is the per\-child difference in expected log predictive density, full model minus shared\-slope model; negative favors the shared slope\. sigma\_b is the between\-child acceleration SD in the training fit, against a full\-data value of 5\.65\.kkhorizon \(mo\)Δ\\DeltaELPD/childSEzz% children betterσb\\sigma\_\{b\}23\-32\.416\.77\-4\.83912\.5633\-5\.324\.10\-1\.3497\.8643\-0\.333\.64\-0\.1486\.2053\+0\.023\.51\+0\.0515\.17We next tested whether acceleration plays an identifiable role in predicting individual children’s trajectories at all\. We therefore fit M20, a variant of M2 with the acceleration exponent constrained to the pure accumulator value \(κ=1\\kappa=1\) — rather than the population acceleration rate – while retaining the per\-child efficiency termξi\\xi\_\{i\}\. The comparison of M2 and M20thus tests the hypothesis that acceleration is predictive of individual trajectories\. Table[4](https://arxiv.org/html/2608.17120#Sx10.T4)shows the results of this comparison\. M2, the accelerating model, predicted held\-out administrations better at every depth\.
Table 4:Prospective cross\-validation of acceleration, on the same Norwegian children and held\-out administrations as Table[3](https://arxiv.org/html/2608.17120#Sx10.T3)\. M2 retains a free population acceleration exponent; M2\-zero removes it\. dELPD/child is the per\-child difference in expected log predictive density, M2 minus M2\-zero; positive favors acceleration\.kkhorizon \(mo\)Δ\\DeltaELPD/childSEzz% children better23\+91\.876\.97\+13\.28433\+143\.218\.90\+16\.18843\+197\.2710\.33\+19\.19053\+248\.1011\.22\+22\.192
### Comparison to a Non\-Parametric Quantile Model
To test whether the accelerating accumulator appropriately fits the distribution of growth curves across children, we compared against a strong baseline: the flexible non\-parametric GAMLSS family that is used to produce normative curves for the CDI \(*19*\)\. Specifically, for each dataset we fit a GAMLSS beta\-regression of the proportion of words produced by the child’s age\. We used a penalized spline function on the mean and on the scales\.
Fig\.[6](https://arxiv.org/html/2608.17120#Sx10.F6)shows that across all five datasets the M3 fit \(blue\) tracks both the non\-parametric GAMLSS fit \(orange\) and the empirical quantiles \(circles\)\. Thus, the accelerating accumulator is a faithful, low\-dimensional compression of the data’s quantile structure\.
Figure 6:M3 fits \(blue, drawn from fitted per\-child acceleration\) vs\. non\-parametric GAMLSS beta\-regression \(orange\), with empirical 10th/50th/90th percentiles \(open circles\)\. Lines are the 10/50/90 percentiles of each model\. Color denotes model\.
### Robustness Across Model and Data Settings
In the main text, we report fits to children with 3\+ administrations with separate models fit to each dataset, but our parameter estimates were substantially similar across two other settings: first, separate fits to children with 2\+ administrations, and second, a full hierarchical model fit to children with 3\+ administrations\. Table[5](https://arxiv.org/html/2608.17120#Sx10.T5)shows LOO fits and the robustness of theκ\\kappaestimates across these different settings\.
The pooled model was structured similarly to the by\-dataset models, but children and words were modeled as nested within their own dataset\. Each dataset had its own meanξ\\xiandκ\\kappavalues\. Dataset\-level standard deviations on efficiency and acceleration \(σa\\sigma\_\{a\}andσb\\sigma\_\{b\}\) were distributed aslogσa\[d\]∼𝒩\(ma,sa\)\\log\\sigma\_\{a\[d\]\}\\sim\\mathcal\{N\}\(m\_\{a\},s\_\{a\}\)andlogσb\[d\]∼𝒩\(mb,sb\)\\log\\sigma\_\{b\[d\]\}\\sim\\mathcal\{N\}\(m\_\{b\},s\_\{b\}\), so a dataset with few children had its estimated spread drawn toward the typical value across datasets, while a dataset with many children was left essentially unconstrained\. The item\-difficulty distribution was estimated separately within each dataset, since the different datasets reflect data from different languages\. Overall, the pooled model reproduced the parameters of the independent fits closely\.
Table 5:For each dataset and setting: number of children \(N\), population acceleration \(κ\\kappa\), between\-child acceleration SD \(σb\\sigma\_\{b\}\), and the per\-observation LOO advantage of the full model M3 over M2 \(Δ\\DeltaELPD per item response\)\.SettingNκ\\kappaσb\\sigma\_\{b\}Δ\\DeltaELPD/obs \(M2→\\toM3\)English \(Thal\)≥\\geq2 separate65311\.53\.190\.0154≥\\geq3 separate \(main text\)63911\.53\.190\.0158pooled \(≥\\geq3\)63911\.53\.22 \[3\.07, 3\.38\]—English \(Smith\)≥\\geq2 separate31612\.88\.060\.0156≥\\geq3 separate \(main text\)15212\.95\.140\.0159pooled \(≥\\geq3\)15212\.85\.13 \[4\.67, 5\.63\]—English \(Marchman\)≥\\geq2 separate2,13610\.66\.890\.0280≥\\geq3 separate \(main text\)16210\.66\.600\.0164pooled \(≥\\geq3\)16210\.46\.51 \[5\.87, 7\.30\]—Norwegian≥\\geq2 separate1,63012\.97\.630\.0121≥\\geq3 separate \(main text\)79213\.35\.650\.0114pooled \(≥\\geq3\)79213\.25\.65 \[5\.40, 5\.93\]—Japanese≥\\geq2 separate18712\.05\.460\.0074≥\\geq3 separate \(main text\)9611\.65\.450\.0059pooled \(≥\\geq3\)9611\.55\.40 \[4\.74, 6\.25\]—
### Sensitivity to Data\-Quality Exclusions
To ensure that the exclusion filter we used in the main text does not result in significant changes to our estimates, we refit the model at four settings: with the filter disabled entirely, at a loose threshold, at the reported threshold, and at a tight one \(Table[6](https://arxiv.org/html/2608.17120#Sx10.T6)\)\. These refits apply only to two datasets, English \(Marchman\) and Norwegian\. The exclusion filter does not result in major changes toκ\\kappaestimates in either case\.
Table 6:Population acceleration and between\-child acceleration SD by data\-quality filter setting, for the two datasets in which the filter removes observations\. The loose filter removes children with a decline of more than 40% below a child’s running peak or a rise of more than 60 percentage points per month; the main filter uses a decline of 25% and a rise of 40%, as reported in the Methods; and the tight filter uses 15% and 25%\.FilterRemovedκ\\kappaσb\\sigma\_\{b\}English \(Marchman\)none09\.378\.33loose2911\.178\.30main3210\.656\.60tight3410\.445\.55Norwegiannone013\.055\.77loose3613\.195\.64main4613\.255\.65tight5513\.225\.34
### Linear vs\. Logarithmic Age
We additionally fit M3 variants with linear, rather than logarithmic, growth in acceleration over time\. Table[7](https://arxiv.org/html/2608.17120#Sx10.T7)shows that the linear model was dispreferred across all five datasets\.
Table 7:Log\-age vs linear\-age accumulation\. LOO ELPD advantage of the log\-age model \(M3\) over its linear\-age counterpart, with SE; positive favors log\-age\.DatasetΔ\\DeltaELPD \(log−\-linear\)Thal267 \(31\)Smith249 \(23\)Marchman281 \(26\)Norwegian323 \(39\)Japanese39 \( 8\)
### 2PL Model Comparison
In the main analysis we use a Rasch \(1PL\) item model, in which every word has a difficulty but all words have equal discrimination between children\. \(*61*\) found that a 2PL model, which adds a per\-word discriminationλj\\lambda\_\{j\}, was preferred over the 1PL for CDI data\. We therefore refit M3 with a per\-word discrimination parameter\.
Table[8](https://arxiv.org/html/2608.17120#Sx10.T8)shows the results of this comparison\. The 2PL does genuinely fit better, with substantial differences in LOO ELPD for every language\. However, model results are otherwise unchanged\. Estimates ofκ\\kappaincrease somewhat, but they are not directly comparable across item models\. In normal IRT models, ability scores are constrained to a standard normal distribution, but in our model, abilities grow over time so there is no single standard scale, and the addition of the discrimination parameter can change the overall scale\. For this reason, we reportlogκ\\log\\kappa, since a pure change of units would be a constant shift\. In addition, across individual words, we find that the predicted age of acquisition \(the age at which the word is predicted to have a 50% probability of being known\) is closely aligned, suggesting that the models are highly comparable\.
Critically,δ\\delta\(difficulty\) andλ\\lambda\(discrimination\) correlations were small to modest\. If harder words were substantially more discriminating, this correlation could inflate estimates ofκ\\kappain the 1PL model; in practice we find little support for this hypothesis\.
Table 8:Rasch \(1PL\) versus two\-parameter logistic \(2PL\) item models, M3 refit on the same data\. Discrimination: SD oflogλj\\log\\lambda\_\{j\}, its 10th\-90th percentile range on the natural scale, and its correlation with item difficulty \(90% interval\)\. Acceleration: populationκ\\kappaunder each item model and the difference inlogκ\\log\\kappa\. Age of acquisition: correlation and median absolute difference in per\-wordt50t\_\{50\}, which is invariant to the discrimination scale\. LOO: expected log predictive density of the 1PL relative to the 2PL, with SE; negative favors the 2PL\.DiscriminationAccelerationκ\\kappaWord slopesλjκi\\lambda\_\{j\}\\kappa\_\{i\}Age of acquisitionLOODatasetSDlogλ\\log\\lambdaλ\\lambdap10–p90r\(λ,δ\)r\(\\lambda,\\delta\)1PL2PLΔlogκ\\Delta\\log\\kappamed\.%\>1\>1r\(t50\)r\(t\_\{50\}\)med\.\|Δt50\|\|\\Delta t\_\{50\}\|Δ\\DeltaELPD \(SE\)English \(Thal\)0\.250\.73–1\.34\+0\.13 \[\+0\.11, \+0\.15\]11\.512\.4\+0\.07212\.299\.80\.9990\.13\-2,897 \( 82\)English \(Smith\)0\.270\.71–1\.37\+0\.10 \[\+0\.07, \+0\.13\]12\.913\.7\+0\.05312\.999\.30\.9960\.13\-2,726 \( 79\)English \(Marchman\)0\.240\.74–1\.33\+0\.05 \[\+0\.02, \+0\.09\]10\.611\.6\+0\.08911\.198\.40\.9980\.25\-1,412 \( 59\)Norwegian0\.290\.70–1\.36\+0\.22 \[\+0\.21, \+0\.23\]13\.314\.1\+0\.06513\.799\.40\.9950\.16\-4,426 \(108\)Japanese0\.340\.65–1\.51\+0\.07 \[\+0\.02, \+0\.12\]11\.615\.3\+0\.27914\.7100\.00\.9861\.47\-641 \( 41\)
### Sensitivity to Item Selection
One potential concern about ourκ\\kappaestimates is that they could be influenced by the composition of the CDI instrument\. Could a broader spread of item difficulties result in a largerκ\\kappaestimate? To address this issue, we refit M3 on the same children and same administrations, but with the item set narrowed\. We fit models under three sets of conditions: refitting the middle 50% and middle 25% of items \(to narrow the width of the distribution\), refitting the easiest and hardest halves separately \(to shift the difficulty distribution\), and subsetting randomly to 50% and 25% of items\. Ifκ\\kappascales as a function of the difficulty distribution \(σδ\\sigma\_\{\\delta\}\), recoveredκ\\kappavalues should scale with the standard deviation of the retained items; if not,κ\\kappashould be relatively invariant\. Fig\.[7](https://arxiv.org/html/2608.17120#Sx10.F7)shows the results of this analysis\. Overall,κ\\kappavalues were invariant to retained item variability\.
Figure 7:Acceleration under deliberately narrowed item sets, plotted as the ratio to the full\-data fit against the ratio of retained item\-difficulty spread\. Each point is one refit on the same children and administrations\. If acceleration rescaled with the instrument, points would fall on the diagonal; if it is a property of the children, they fall on the horizontal line at one\. Open points are random subsets of matched size, which discard the same number of items without narrowing the spread\.
### Variation in Acceleration
To explore the distribution of individual children’s fitted parameters, we show the best linear unbiased predictor \(BLUP\) for each child \(Figure[8](https://arxiv.org/html/2608.17120#Sx10.F8)\) as well as the random effects of the fitted models \(Table[9](https://arxiv.org/html/2608.17120#Sx10.T9)\)\. While the distribution appears generally normal, we do not give a strong interpretation as this shape may in part be due to the normal form of the model’s random effect structure\. However, across all datasets, acceleration appeared right\-skewed\.
Figure 8:Histograms of per\-child BLUP estimates, z\-scored within each dataset, for both efficiency and acceleration parameters\. Dashed lines indicate the standard normal reference distribution\.Table 9:Random\-effect parameters for M3 fits, by dataset: population accelerationκ\\kappa, between\-child SD in efficiency \(σa\\sigma\_\{a\}\) and acceleration \(σb\\sigma\_\{b\}\), their correlation \(ρ\\rho\), and the percentage of children estimated above the pure\-accumulator valueκ=1\\kappa=1\. Brackets give 90% posterior intervals\.DatasetNκ\\kappaσa\\sigma\_\{a\}σb\\sigma\_\{b\}ρ\\rho%κi\>1\\kappa\_\{i\}\>1English \(Thal\)63911\.5 \[11\.5, 11\.6\]1\.56 \[1\.49, 1\.64\]3\.19 \[3\.05, 3\.35\]\+0\.10 \[\+0\.03, \+0\.17\]99\.8English \(Smith\)15212\.9 \[12\.8, 13\.1\]2\.08 \[1\.89, 2\.31\]5\.14 \[4\.69, 5\.68\]\-0\.26 \[\-0\.39, \-0\.13\]99\.3English \(Marchman\)16210\.6 \[10\.5, 10\.8\]1\.80 \[1\.65, 1\.99\]6\.60 \[5\.90, 7\.32\]\+0\.17 \[\+0\.04, \+0\.29\]97\.5Norwegian79213\.3 \[13\.2, 13\.3\]2\.71 \[2\.60, 2\.83\]5\.65 \[5\.41, 5\.88\]\-0\.62 \[\-0\.66, \-0\.58\]99\.5Japanese9611\.6 \[11\.1, 12\.1\]1\.95 \[1\.71, 2\.25\]5\.45 \[4\.74, 6\.32\]\+0\.26 \[\+0\.07, \+0\.43\]100\.0
### Input Estimation
We make use of estimates of children’s language input in both Fig\. 2 and Fig\. 3C\. We calculate average language input per month by estimating words per hour and multiply by an estimated number of waking hours per month\. Word tokens per hour estimates come from 42 children with manually transcribed data in two prior studies \(*56*,*57*, estimates used also in*9*\)\. Across these studies, the pooled geometric mean isexp\(7\.338\)≈1,537\\exp\(7\.338\)\\approx 1\{,\}537word tokens per hour of child\-directed \(any\-adult\) speech, with between\-child standard deviation of log rateσr=0\.534\\sigma\_\{r\}=0\.534\. We multiply these estimates by an assumed 12 waking exposure\-hours per day to get a monthly average number of hours \(12×30\.4≈36512\\times 30\.4\\approx 365\), giving1,537×365≈561K1\{,\}537\\times 365\\approx 561\\text\{K\}tokens per month\. A child’s average cumulative input at agettmonths is thent×561Kt\\times 561\\text\{K\}under a constant\-rate assumption\.
In Fig\. 3C, we give uncertainty bands on representative children based on±1\\pm 1SD of the between\-child log input rate\. A child one standard deviation above \(below\) the median rate reaches the same age having heard×1\.71\\times 1\.71\(÷1\.71\\div 1\.71\) as many tokens\. This results in a horizontal shift of the child’s curve along the log\-token axis\.
Our calibration computations assume that input is constant across age\. To provide empirical verification, we used data from four pre\-existing datasets: \(*62*\), \(*63*\) \(BabyView\), \(*64*\), \(*65*\) \(SEEDLings\)\. These studies measured children’s input using either LENA recorders or head\-mounted cameras; we computed total words heard per hour in each dataset for each recording\. Fig\.[9](https://arxiv.org/html/2608.17120#Sx10.F9)shows speech rate per hour in these datasets\. Slopes over age for each were shallow and non\-significant\.
Figure 9:Input rate does not change appreciably across development\. Each panel plots log adult input rate against child age, with one point per recording, faint gray lines joining recordings from the same child, and the fitted within\-child trend \(black\) with its 95% confidence band\. Dashed lines show zero slope\. Trends are fitted over 8\-36 months \(shaded\), the range over which the constant\-rate assumption is applied; gray points fall outside that window and are excluded from the fit\.
### Age of Acquisition in the Pure Accumulator Model
Figure 2 in the main text converts each word’s fitted difficulty into an estimated number of cumulative exposures at the age it is acquired\. Since later\-acquired words are also rarer, here we ask whether that frequency change would be sufficient to produce the observed relationship between cumulative exposure and age of acquisition\.
Fig\.[10](https://arxiv.org/html/2608.17120#Sx10.F10)shows the same relationship, but comparing age of acquisition estimates between M3 \(the accelerating accumulator\) and M0 \(the pure accumulator\)\. The pure accumulator does not reproduce either the scale or the slope of the relationship \(Table[10](https://arxiv.org/html/2608.17120#Sx10.T10)\)\. Within every word class, the slope of the relationship is substantially shallower in M0 than M3, and the M0 forces the fitted difficulties to spread acquisition across 5 to 254 months, against 14 to 35 months for the accelerating model and 8 to 36 months of actual observation\.
Figure 10:Estimated cumulative exposures at age of acquisition under the accelerating accumulator \(M3, left, reproduced from main text\) and the fitted pure accumulator \(M0, right\), using identical corpus frequencies\. Points are words, colored by class, with per\-class linear fits\.Table 10:Slope of estimated cumulative exposures with respect to age of acquisition, by word class, under each item model\. Negative values mean fewer exposures are needed for later\-acquired words\.M3 \(accelerating\)M0 \(pure accumulator\)Word classnnslopeR2R^\{2\}slopeR2R^\{2\}nouns293\-4\.770\.38\-0\.670\.11action words103\-3\.470\.07\-0\.060\.00descriptive words62\-1\.680\.03\+0\.490\.03function words90\-3\.180\.13\-0\.220\.01
### Matched Estimation of Acceleration
One potential objection to the comparison of children and LMs in the main text is that the estimators are different in two ways\. First, the estimator for children is Rasch\-based, while the model from \(*34*\) is a four\-parameter logistic\. Second, and perhaps more saliently, the child estimator is forκ\\kappaacross children, while the LM estimator is across words\.
To address these issues, we fit the LM four\-parameter estimator directly to data on children’s word acquisition, averaging across children to produce a proportion of children producing each word at each age\. \(The Thal and Japanese datasets had fewer children and less age diversity, so not all words produced a viable curve estimate, but most words could be estimated for the Smith, Marchman, and Norwegian datasets\)\. Fig\.[11](https://arxiv.org/html/2608.17120#Sx10.F11)shows the results\. Across all datasets,κ\\kappaestimates were robustly higher than LMs, though somewhat attenuated relative to the M3 fits\.
One possible source of this attenuation is the averaging of many different acquisition curves across children\. To estimate this attenuation, we sampled from the fitted M3 parameters forξi\\xi\_\{i\},κi\\kappa\_\{i\}andδj\\delta\_\{j\}, and ran the identical pipeline over the data to recover a correction factor \(the ratio of the estimated by\-wordκ\\kappato the populationκ\\kappaused to generate the data\)\. We then applied this correction factor to our empirical estimates\. Table[11](https://arxiv.org/html/2608.17120#Sx10.T11)shows these corrected estimates, which match very closely to the M3 values\.
In sum, even with a completely matched estimator, acceleration diverges substantially between children and LMs\.
Figure 11:Per\-word acceleration for children and language models, both estimated with the same four\-parameter logistic\. Points are words, with the median marked; language model values are medians across the ten training seeds\. The dashed line at one is the pure\-accumulator value\. Child estimates are attenuated by pooling across children and so understate the separation; the correction factors are given in the table\.Table 11:Per\-word acceleration estimated with the language models’ own estimator\. For each word we fit the same four\-parameter logistic used on LM surprisal to the proportion of children not yet producing that word, and take kappa = 0\.434/scale\. Recovery is the fraction of the population kappa that the same pipeline returns when run on data generated from the fitted M3, and is the factor by which pooling across children attenuates the estimate; the corrected estimate is the raw estimate divided by this factor\. The final column is the kappa M3 model reports for that dataset\.Samplewords fittedκ\\kapparaw \[IQR\]recoveryκ\\kappacorrectedκ\\kappa\(M3\)English \(Thal\)135/68110\.18 \[4\.32, 14\.84\]98%10\.4211\.51English \(Smith\)673/6808\.21 \[6\.04, 10\.67\]62%13\.2112\.95English \(Marchman\)624/6816\.34 \[3\.12, 9\.29\]59%10\.6610\.65Norwegian713/7337\.87 \[6\.97, 8\.69\]57%13\.8113\.25Japanese343/44710\.22 \[7\.66, 15\.56\]86%11\.9211\.57Language models609/6091\.16 \[0\.93, 1\.46\]
### Other Architectures
The acceleration results we report in the main text are not specific to our choice of the GPT\-2 architecture\. We re\-analyze the published fits of \(*34*\)\. Fig\.[12](https://arxiv.org/html/2608.17120#Sx10.F12)shows a similar distribution of per\-wordκ\\kappavalues for the four different architectures reported in that study, which were trained on standard LM corpora rather than child\-directed input\.
Figure 12:Per\-word acquisition slopes for the four architectures studied in Chang & Bergen \(2022\) – BERT, GPT\-2, BiLSTM, LSTM – trained on BookCorpus \+ WikiText\-103\. CHILDES\-trained GPT\-2 models \(training axis\) are overlaid in gray\.Similar Articles
Kids outlearn AI—and we still don’t know why
Children learn language with far less data than AI models like LLMs, and understanding this data efficiency gap could lead to more efficient AI systems and insights into human cognition.
AI Isn’t Smarter Than a Baby—Yet
A new benchmark test, EgoBabyVLM, challenges AI vision-language models to learn from video footage captured from baby head-cameras, revealing that current AI models fail to match the learning efficiency of infants and suggesting that baby-like learning architectures could lead to more efficient AI.
@_jasonwei: When language models first started using tools well, I was sympathetic to the narrative that instead of scaling up lang…
The author argues that while tool use allows smaller language models to perform tasks effectively, larger models remain crucial for speed, reliability, and internalized knowledge, emphasizing the ongoing need for scaling in AI.
The Download: kids outlearning AI, and space travel agents
The article highlights the data efficiency gap where children outperform AI models in learning, explores research to close this gap, and covers emerging careers in space travel alongside other tech developments like AI data center backlash and humanoid robot records.
When More Becomes Less: Position-Dependent Repetition Effects in Language Models
This paper shows that repetition effects in language models depend on readout position: adjacent repetition boosts target probability, while displaced repetition produces an inverted-U curve. The finding challenges assumptions in cloze-style probing and is validated across multiple models and languages.