Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization
Summary
This paper disentangles statistical preemption from entrenchment in language models' avoidance of overgeneralizations through controlled experiments, finding that LMs exhibit abstract preemption rather than verb-specific preemption, with implications for human language learning.
View Cached Full Text
Cached at: 09/03/26, 05:47 AM
# Disentangling Statistical Preemption from Entrenchment in Language Models’ Avoidance of Overgeneralization
Source: [https://arxiv.org/html/2609.01794](https://arxiv.org/html/2609.01794)
\\todostyle
orangecolor=orange\!10,bordercolor=orange\!90,linecolor=orange\!90\\todostylekcyancolor=kcyan\!10,bordercolor=kcyan\!90,linecolor=kcyan\!90\\todostyletticbluecolor=tticblue\!10,bordercolor=tticblue\!90,linecolor=tticblue\!90
Yixuan WangFreda ShiAffiliation:University of WaterlooAffiliation:Vector InstituteAffiliation:Canada CIFAR AI ChairEmail:[fhs@uwaterloo\.ca](mailto:)Kanishka MisraAffiliation:Department of LinguisticsAffiliation:University of Texas at AustinEmail:[kmisra@utexas\.edu](mailto:)
###### Abstract
How do learners avoid overgeneralizations such asTom laughed mewithout explicit negative evidence? Constructivists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption \(which privileges exposure to near\-synonymous construction—e\.g\.,she made him laugh\) vs\. entrenchment \(all exposures to a verb’s grammatical usages, including cases likeHe laughed\)\. We disentangle these hypotheses by running controlled rearing experiments on LMs trained on child\-caregiver conversations, where we systematically remove preemptive vs\. non\-preemptive evidence\. We find that while LMs avoid overgeneralizations, they do not show preemption at a verb\-specific level, instead showing evidence of abstract preemption\. Combined with results from analyzing the LMs’ training dynamics, we find that LMs treat competing structures as indirect positive—as opposed to negative—evidence in the verb\-specific condition\. Insofar as preemption is the more plausible route to avoiding overgeneralizations in humans, our results suggest the need for there to be greater sensitivity to indirect negative evidence in neural network learners, and motivate new human experiments to test preemptive effects of verbs beyond the target verb\.
## 1Introduction
One of the hallmarks of human cognition is our ability to generalize productively\. For instance, on hearing[0xnumia](https://arxiv.org/html/2609.01794#S1.I1.i1.I1.i1), competent speakers of English use their experience of intransitive and transitive verbs \(break, roll\) to effortlessly produce[0xnumib](https://arxiv.org/html/2609.01794#S1.I1.i1.I1.i2)\. But their generalization of this pattern is also constrained—e\.g\., while we are perfectly fine with[0xnumia](https://arxiv.org/html/2609.01794#S1.I1.i2.I1.i1), we do not produce[0xnumib](https://arxiv.org/html/2609.01794#S1.I1.i2.I1.i2)\.
- \(0xnumi\)- a\.- The water is boiling\. - b\.- Dad boiled some water\.
- \(0xnumi\)- a\.- The baby laughed\. - b\.- \*Mom laughed the baby\.
Figure 1:Overview of our experimental paradigm\. For a given verb \(laugh\), we train LMs on two versions of our corpus, created by removing: 1\) the preemptive evidence for that verb; and 2\) same amount of non\-preemptive evidence\. We then compare the behavior of these LMs to disentangle preemption from entrenchment in explaining that verb’s overgeneralization\.Children often tend toovergeneralizethe above pattern by indeed producing sentences like[0xnumib](https://arxiv.org/html/2609.01794#S1.I1.i2.I1.i2)\([Bowerman, 1988](https://arxiv.org/html/2609.01794#bib.bib11);[Braine and Brooks, 1995](https://arxiv.org/html/2609.01794#bib.bib13);[Akhtar and Tomasello, 1997](https://arxiv.org/html/2609.01794#bib.bib2);[Brooks and Tomasello, 1999](https://arxiv.org/html/2609.01794#bib.bib14), i\.a\.\), but eventuallyretreatand constrain their generalizations\. They do sowithoutmuch explicit negative evidence\([Brown and Hanlon, 1970](https://arxiv.org/html/2609.01794#bib.bib16);[Baker, 1979](https://arxiv.org/html/2609.01794#bib.bib8)\)—caregivers rarely ever correct children’s grammatical errors systematically\([Chouinard and Clark, 2003](https://arxiv.org/html/2609.01794#bib.bib18)\)\. How, then, are such constraints and restrictions learned?
In the realm of constructivist and usage\-based approaches to language acquisition\([Goldberg, 1995](https://arxiv.org/html/2609.01794#bib.bib23);[Tomasello, 2003](https://arxiv.org/html/2609.01794#bib.bib48);[Goldberg, 2005](https://arxiv.org/html/2609.01794#bib.bib24);[Rowland, 2013](https://arxiv.org/html/2609.01794#bib.bib42);[Rowland et al\., 2025](https://arxiv.org/html/2609.01794#bib.bib43)\), there are two closely related proposals that attempt to explain this avoidance of overgeneralizations:preemptionandentrenchment\. Preemption suggests that given two near\-synonymous constructionsAAandBB, exposure to a verb used inAAin contexts where the verb could have been used in either construction serves as evidenceagainstthe verb’s usage inBB\([Goldberg, 1995](https://arxiv.org/html/2609.01794#bib.bib23);[Brooks and Zizak, 2002](https://arxiv.org/html/2609.01794#bib.bib15);[Boyd and Goldberg, 2011](https://arxiv.org/html/2609.01794#bib.bib12)\)\. In the case oflaugh, preemption predicts that the transitive\*Mom laughed the babyispreemptedby the periphrastic causative usageMom made the baby laugh, because in situations whereXdid something that causedYtolaugh, the utteranceX made Y laughwas used instead ofX laughed Y\. Entrenchment, on the other hand, does not privilege near\-synonymous constructions or specific contexts\([Theakston, 2004](https://arxiv.org/html/2609.01794#bib.bib47);[Stefanowitsch, 2008](https://arxiv.org/html/2609.01794#bib.bib45);[Ambridge et al\., 2008](https://arxiv.org/html/2609.01794#bib.bib6)\)\. Here, the usage of the verb inanyattested construction is evidence against the unattested usage—entrenchment not only considers periphrastic causative usages oflaughin blocking its transitive usage, but also other constructions like the intransitive \(as in[0xnumia](https://arxiv.org/html/2609.01794#S1.I1.i2.I1.i1)\)\. These two proposals have been up for debate for about 40 years\([Rowland, 2013](https://arxiv.org/html/2609.01794#bib.bib42)\)\.
There are more similarities between these proposals than there are differences: 1\) both explanations rely onindirectnegative evidence, since they positotherconstructions as evidence against overgeneralization; 2\) both are statistical in nature, since they rely on the frequency of exposure to the indirect evidence—more frequent this exposure is, stronger is the blocking effect; and most importantly, 3\) entrenchment is a strict superset of preemption, since the evidence that falls under the purview of preemption \(near\-synonymous construction\) is always part of the evidence considered by entrenchment\. The latter point explains why attempts to disentangle the two proposals in the laboratory\([Brooks and Zizak, 2002](https://arxiv.org/html/2609.01794#bib.bib15);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4);[Ambridge et al\., 2018](https://arxiv.org/html/2609.01794#bib.bib3);[Bidgood et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib10)\)often suffer from multi\-collinearity issues in both variables\. In these studies, participants’ acceptability judgments on unconventional forms are predicted using the frequencies of the preemptive usages of a verb \(preemption\) as well as that of the total frequency of the verb \(entrenchment\) sourced from openly available corpora, rather than the learners’ own experience \(which is unobtainable for humans\)\. Therefore, apart from multicollinearity \(due to the overlap in preemptive and entrenching evidence\), these experiments and analysis are also unable to shed light on the causal status of these proposals\.
To adjudicate between these two proposals, we conduct “controlled rearing”\([Frank, 2023](https://arxiv.org/html/2609.01794#bib.bib21);[Misra and Mahowald, 2024](https://arxiv.org/html/2609.01794#bib.bib36)\)experiments on Language Models \(LMs\), where we systematically manipulate the learning experience of LMs to disentangle between preemption and entrenchment\. As a case study, we apply our methods to explaining transitive overgeneralizations of \(usually\) intransitive verbs such aslaugh, smile, cry, etc\., as this is among the quintessential examples of entrenchment and preemption analyses over the past 30 years\([Akhtar and Tomasello, 1997](https://arxiv.org/html/2609.01794#bib.bib2);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4);[Bidgood et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib10)\)\. While analyses on LMs cannotdirectlybe taken to have implications for humans\([Warstadt and Bowman, 2022](https://arxiv.org/html/2609.01794#bib.bib49);[Portelance and Jasbi, 2024](https://arxiv.org/html/2609.01794#bib.bib39);[Misra and Kim, 2026](https://arxiv.org/html/2609.01794#bib.bib35);[Futrell and Mahowald, 2025](https://arxiv.org/html/2609.01794#bib.bib22), cf\.\), such experiments allow us to conclude about the viability of the two proposals and how \(and whether\) they might arise from the general purpose statistical learning of the kind employed by LMs\. Controlled rearing allows us to do this by considering counterfactual learning scenarios—e\.g\., “what if there was no preemptive evidence in the learner’s training data?” Evidence of a privileged role of preemptive evidence would suggest that LMs acquire restraints on their generalization behavior from fine\-grained sensitivities to meaning, since preemption privileges contexts where either construction could have been used to communicate the same intent or semantics\. More broadly, this allows us to understand if similar constraints arise within LMs trained on cognitively plausible amounts of language\-only exposure, or if explicit evidence of semantics is needed\([Abend et al\., 2017](https://arxiv.org/html/2609.01794#bib.bib1);[Yedetore and Kim, 2024](https://arxiv.org/html/2609.01794#bib.bib55)\)\.
Our methods and research questions also make methodological contributions to the burgeoning program of “controlled rearing”\([Frank, 2023](https://arxiv.org/html/2609.01794#bib.bib21);[Misra and Mahowald, 2024](https://arxiv.org/html/2609.01794#bib.bib36)\)or “filtered corpus training”\([Patil et al\., 2024](https://arxiv.org/html/2609.01794#bib.bib38)\), which has established itself as an important paradigm in understanding the role of input in shaping a statistical learner’s generalizations\. Most applications of this method have primarily demonstrated how LMs generalize to “novel” usages in the absence of direct or impoverished evidence\([Jumelet et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib29);[Misra and Mahowald, 2024](https://arxiv.org/html/2609.01794#bib.bib36);[Patil et al\., 2024](https://arxiv.org/html/2609.01794#bib.bib38);[Yao et al\., 2025](https://arxiv.org/html/2609.01794#bib.bib54);[Yang et al\., 2026](https://arxiv.org/html/2609.01794#bib.bib52)\)\. That is, they only focus on indirectpositiveevidence, since they investigate what facilitates generalization, and not whatblocksit\. An exception to this is work by[Leong and Linzen \(2026\)](https://arxiv.org/html/2609.01794#bib.bib31), who focus on explicitly understanding signals thatblockmodels’ usage of a structure, but focus on a different construction \(active/passive\), and do not consider the difference between preempting and non\-preempting evidence \(i\.e\., they consider entrenchment as a whole\)\. This positions our work uniquely within this literature, since we aim to causally disentangle preemption and entrenchment, and explicitly measure their viability as sources of indirectnegativeevidence\.
### 1\.1Summary of Findings
Before testing what explains models’ avoidance of overgeneralization, we first measure if they learn the appropriate generalizations in the first place—i\.e\., finding transitive usages of intransitive verbs less probable than their periphrastic causative counterparts\. Finding that they do \([§3](https://arxiv.org/html/2609.01794#S3)\) for six different verbs, we then run our controlled rearing studies to disentangle preemption from entrenchment \([§4](https://arxiv.org/html/2609.01794#S4)\)\. Here, we test if preemptive usages of a given verb \(i\.e\. their periphrastic causative usages like inTom made me laugh\) have a privileged role over non\-preemptive ones \(e\.g\.,I laughed\.\), by training LMs on corpora where each of them are removed in equal amounts separately\. Using LMs’ negative log\-likelihood loss \(NLL\) on the unconventional, unseen transitive usages of these verbs as our measure of overgeneralization, we considered two different manipulation types: 1\)verb\-specific, where for a verb only its own usages were manipulated; and 2\)abstract, where usages of all verbs were manipulated, to quantify if the evidence of preemption comes from higher order sources\.
We do not find clear evidence for preemption in the verb\-specific condition—LMs’ NLL of the unconventional transitive usage was greater when preemptive exposures were removed than when they were kept in\. By contrast, preemption would predict that in the absence of preemptive exposures, models should find transitive usagesmorelikely\. The abstract condition, however, did show relatively larger effects of preemption, though in absolute terms these were generally weak\. This finding is consistent with the hypothesis that models treat the periphrastic causative construction as indirect positive evidence, common in other controlled rearing studies\. That is, the models might have inferred the relation between transitive and periphrastic causatives via other verbs for which both are acceptable \(e\.g\.,She boiled the watervs\.She got the water to boil\), which explains why this is not observed when all periphrastic causatives are removed \(in the abstract condition\)\.
To further explore this finding of positive—as opposed to negative—evidence, we investigated the change in the model’s behavior on unseen transitive and periphrastic causative sentences before and after a periphrastic causative usage was encountered by the model during training\. We find the model’s generalization behavior on both constructions to be highly correlated, further supporting the indirect positive evidence hypothesis—the presence of periphrastic causatives is tightly associated with the likelihood of transitive usages\.
Insofar as preemption occurs and drives retreat from overgeneralization in humans\([Samara et al\., 2025](https://arxiv.org/html/2609.01794#bib.bib44)\), our verb\-specific results indicate the need for LM learners to model fine\-grained indirect negative evidence\. We speculate that this could come from more explicit semantic signals about the context in which either construction is used, or from morphological constraints, where the concept of preemption originates from\([Kiparsky, 1973](https://arxiv.org/html/2609.01794#bib.bib30);[Aronoff, 1976](https://arxiv.org/html/2609.01794#bib.bib7)\)\. At the same time, our finding of abstract preemption motivates new experiments that test this finding in humans via an artificial language learning design, since preemption for argument structure alternations is often treated as a verb\-specific effect\([Goldberg, 2005](https://arxiv.org/html/2609.01794#bib.bib24);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4);[Ambridge et al\., 2018](https://arxiv.org/html/2609.01794#bib.bib3);[Bidgood et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib10)\)\. Our code and data are available at[https://github\.com/Yixuan\-Wang/verb\-overgen](https://github.com/Yixuan-Wang/verb-overgen)\.
## 2Data, Model, and Methods
### 2\.1Training Data
To remain as close as possible to the input and constraints available to children, we use the English subset of the Child Language Data Exchange System \(CHILDES\) dataset\([MacWhinney, 2000](https://arxiv.org/html/2609.01794#bib.bib34)\)as the training dataset for our models\. CHILDES consists of human\-transcribed conversations between children and caregivers\. The corpus consists of approximately 5M utterances, amounting to a total of 26M space\-delimited words\. Importantly, CHILDES also comes with morphological and syntactic annotations\.111[https://talkbank\.org/0info/manuals/CHAT\.pdf](https://talkbank.org/0info/manuals/CHAT.pdf)This allows us to precisely detect phenomena we are interested in, avoiding any potential errors from existing automatic tools\([Yang et al\., 2025](https://arxiv.org/html/2609.01794#bib.bib53);[Padovani et al\., 2026](https://arxiv.org/html/2609.01794#bib.bib37), cf\.\)\.
As a preprocessing step, we lemmatize the corpus and strip away morphological inflections\. This allows models to represent tense\-inflected versions of verbs using the same token \(laughed→\\rightarrowlaugh,laughing→\\rightarrowlaugh\), thereby giving them the maximum chance\-\-\-if any\-\-\-to use inflected forms of the same verb as collective evidence\.222We show results with no lemmatization in[AppendixG](https://arxiv.org/html/2609.01794#A7)\.All subsequent experimental data, including evaluation, use the lemmatized sentences\.
### 2\.2Model
We use a word\-level tokenizer for the lemmatized training data, with a vocabulary size of 37,482\. Our learners are autoregressive LMs with the GPT\-2 architecture\([Radford et al\., 2019](https://arxiv.org/html/2609.01794#bib.bib41)\)\. Each LM has 12 layers and 12 attention heads per layer, with a hidden size of 768, resulting in a total of 115M parameters\. Each model is trained for 3 epochs, and training is repeated with 5 random seeds to marginalize nondeterminism, including parameter initialization and training data shuffling\.
### 2\.3Disentangling Preemption from Entrenchment
Because preemption is necessarily a subset of entrenchment \(see[§1](https://arxiv.org/html/2609.01794#S1)\), disentangling the two boils down to testing whether the hypothesized preemptive evidence has aprivilegedrole compared to those not considered as part of preemption\. This effectively partitions the set of usages of a particular verb to preemptive and non\-preemptive\. In the context of transitive overgeneralizations \(she laughed me\), the preemptive evidence is the periphrastic causative construction \(she made me laugh\), and the non\-preemptive ones are all other types of usages \(he is laughing, she laughed at me,etc\.\)\. This give us two kinds of manipulations for our controlled rearing experiments: 1\)remove\-p, where we removealloccurrences of the periphrastic causative construction usages for a given verb \(or a set of verbs\); and 2\)keep\-p\(orremove\-non\-p\), where we remove the same number of occurrences of a given verb \(or set of verbs\) as inremove\-p, thatdo notuse it in the periphrastic causative construction\.
The raw frequency of a verb of interest in both cases is reduced by the same amount, but the structures being manipulated are different\. Insofar as models assign a special role to preemptive evidence, we should expect them to retreat further from overgeneralizations inkeep\-p, which retains the preemptive evidence, than inremove\-p\. In other words, the model should find transitive overgeneralizations more likely underremove\-pthan underkeep\-p, since the hypothesized critical evidence against transitive usages is absent inremove\-p\.
In our experiments, we primarily focus on six verbs that have been described to be ‘intransitive\-only’ by[Levin \(1993\)](https://arxiv.org/html/2609.01794#bib.bib32):go, laugh, cry, sleep, sneeze,andsmile\. Using them transitively almost always constitutes an overgeneralization\.333There can be creative, usages of these verbs that appear transitive—e\.g\., in thecaused\-motionconstruction:She sneezed the foam off the cappuccino\([Goldberg, 2005](https://arxiv.org/html/2609.01794#bib.bib24)\), but these do not have the same meaning as the causative context—the agent \(she\) is not making the object \(foam\) sneeze\.[Table1](https://arxiv.org/html/2609.01794#S2.T1)shows the usage statistics of these verbs in CHILDES, both as a verb in general, and also more specifically in the periphrastic causative construction\. The periphrastic usages of these verbs are detected using the gold\-dependency parses provided in the CHILDES dataset\. A periphrastic causative construction is identified if the matrix verb \(e\.g\.,make,let, etc\.\) has a complement child \(COMP\) and a patient in the dependency parse\. Technical details about the identification of periphrastic causative constructions can be found in[AppendixB](https://arxiv.org/html/2609.01794#A2)\.
Table 1:Usage statistics of the verbs of interest in the training set, including the number of times each word appears in a Periphrastic Causative \(PC\) construction, and the total number of times the word is used as a verb\.
### 2\.4Measuring Overgeneralization Effects
#### Evaluation Data
Our evaluation data constitutes a pair of transitive and periphrastic causative sentences containing the abovementioned verbs and a set of nouns that occur in the CHILDES dataset\. We construct lemmatized versions of these sentences by using the template “S V O” \(e\.g\.,Tommy laugh Mary\) for the transitive structure, and “S make O V” \(e\.g\.,Tommy make Mary laugh\) for the periphrastic causative structure\. Here,S,O, andVare the subject, object, and the verb lemma, respectively, such that none of the sentences occur in CHILDES\. In total, we generated 100 pairs for each verb\. We denote transitive sentences ass−s^\{\-\}, and periphrastic causative ones ass\+s^\{\+\}\.
#### Measure
We measure the overgeneralization effect for a given verb as the average negative log\-likelihood loss \(NLL\) of the transitive sentences for that verb in our evaluation data\. Specifically, given an LMP𝒟P^\{\\mathcal\{D\}\}, trained on a corpus𝒟\\mathcal\{D\}, and a transitive sentence,s−=x1−…x\|s−\|−s^\{\-\}=x^\{\-\}\_\{1\}\\ldots x^\{\-\}\_\{\|s^\{\-\}\|\}, we measure:
ℓ\(s−;P𝒟\)=−1\|s−\|∑i=1\|s−\|logP𝒟\(xi−∣x<i−\)\\displaystyle\\ell\(s^\{\-\};P^\{\\mathcal\{D\}\}\)=\-\\frac\{1\}\{\|s^\{\-\}\|\}\\sum\_\{i=1\}^\{\|s^\{\-\}\|\}\\log P^\{\\mathcal\{D\}\}\\left\(x^\{\-\}\_\{i\}\\mid x^\{\-\}\_\{<i\}\\right\)This allows us to compare models trained on corpora with different amounts of indirect negative evidence in terms of how \(un\)likely or surprising they find transitive usages of a verb\. If the evidence contained in a corpus𝒟1\\mathcal\{D\}\_\{1\}better suppresses the overgeneralization of a verb than does the evidence in𝒟2\\mathcal\{D\}\_\{2\}, then we expect to observe a higher NLL onP𝒟1P^\{\\mathcal\{D\}\_\{1\}\}than onP𝒟2P^\{\\mathcal\{D\}\_\{2\}\}, i\.e\.,ℓ\(s−,P𝒟1\)\>ℓ\(s−,P𝒟2\)\\ell\(s^\{\-\};P^\{\\mathcal\{D\}\_\{1\}\}\)\>\\ell\(s^\{\-\};P^\{\\mathcal\{D\}\_\{2\}\}\)\.
It might be tempting to treat overgeneralization effect as thepreferenceof the periphrastic causative evaluation sentences\+s^\{\+\}over its overgeneralized transitive forms−s^\{\-\}, by measuring the difference between their NLLs\. Afterall, these two constructions are hypothesized to compete\([Goldberg, 2005](https://arxiv.org/html/2609.01794#bib.bib24)\)\. However, we argue that this metric does not allow us to explain overgeneralization effects, which are explicitly concerned with what allows the learner to conclude that the unattested form \(here, transitive\) is unacceptable\. This is because our experiments \(as we will show in[§4\.1](https://arxiv.org/html/2609.01794#S4.SS1)\) directly manipulate the presence of the periphrastic causative exposures available to the model\. Therefore, changes to this alternate measure could be completely driven by the likelihood of the conventional, periphrastic causative sentence \(whose exposure has been manipulated\) rather than the unlikelihood of the transitive counterpart\. Indeed, such a measure has been used in previous work\([Guo et al\., 2026](https://arxiv.org/html/2609.01794#bib.bib26)\), but due to the above reason, it can paint an illusory picture of preemption \(see more in[AppendixH](https://arxiv.org/html/2609.01794#A8)and[AppendixI](https://arxiv.org/html/2609.01794#A9)where we compare our results to the one had we used this alternate measure\)\.
## 3Preconditions to Not Overgeneralizing
Before analyzing how indirect negative evidence might facilitate our LMs to avoid overgeneralization, we first establish if the model trained on the full CHILDES corpus learns the appropriate generalization behavior—that the transitive usages of our target verbs is indeed unacceptable, relative to their periphrastic causative usages\. This is the minimal precondition the model must satisfy before we can use it to conclude about the factors that affect its generalization, following[Misra and Kim \(2026\)](https://arxiv.org/html/2609.01794#bib.bib35)\.
Figure 2:Difference in NLL of transitive and periphrastic causative sentence pairs for intransitive \(target\), alternating, and transitive verbs at every 50 steps during training\. Smoothed curves are estimated using a generalized additive model\([Hastie and Tibshirani, 1986](https://arxiv.org/html/2609.01794#bib.bib27)\)\. Across all intransitive verbs, models show greater loss on transitive sentences than on their periphrastic causative counterparts, suggesting that they have acquired the right preference for these verbs\. This is in contrast to alternating and transitive verbs, for which models show relatively lowerΔ\\DeltaNLL values, as expected\.To test this criterion, we compare the NLL of the transitive evaluation sentences \(i\.e\., the unconventional forms,s−s^\{\-\}\) to their periphrastic causative counterparts \(s\+s^\{\+\}\)\. Insofar as the model has learned the right generalizations, we expect this measure to be a positive value, since the loss on the transitive should be greater than that on the periphrastic causative, since the latter is more acceptable\. In addition, these results should not just be explained away by the fact that the model has a periphrastic causative bias\. That is, its preference for verbs that do not permit the periphrastic causative usage \(e\.g\.,hit\) and are instead transitive\-biased should show a reverse effect\. To account for this additional precondition, we compare theΔ\\DeltaNLL values for intransitive verbs to 5 alternating verbs \(break,move,stop,turn,roll\) and 5 verbs that have a transitive bias \(catch,feed,hit,push,touch\), all based on[Levin \(1993\)](https://arxiv.org/html/2609.01794#bib.bib32)’s classification\. We expect the transitive\-biased verbs to show the opposite behavior as our target intransitive verbs \(i\.e\., negativeΔ\\DeltaNLL values\), and alternating verbs to be somewhere in between\. We create 100 pairs of sentences for each of these verbs using the same procedure as our target verbs \(see[§2\.3](https://arxiv.org/html/2609.01794#S2.SS3)\)\. We measureΔ\\DeltaNLLs for sentence pairs of all three classes of verbs \(Intransitive, Alternating, and Transitive\) across model training steps \(every 50 batches\), and show the resulting curves in[Figure2](https://arxiv.org/html/2609.01794#S3.F2)\.444A cleaner version with just the final checkpoints can be found in[AppendixD](https://arxiv.org/html/2609.01794#A4)\.
We see from this figure that our models show clear evidence of expected behavior, as described above\. Our intransitive verbs show generally positiveΔ\\DeltaNLLs, while the transitive verbs show generally negativeΔ\\DeltaNLLs, and alternating verbs are squarely in the middle\. Overall, these results suggest that our LMs are well\-positioned to answer questions about the factors that lower or increase \(un\)likelihood on the unconventional form for our target intransitive verbs\.
## 4Verb\-specific and Abstract Preemption
Having shown that the models indeed learn the appropriate generalization, we now turn to disentangling preemption from entrenchment\. We compare the role of preemptive and non\-preemptive evidence in the model’s behavior on the unconventional, overgeneralized, transitive usages of our target verbs\. That is, we compare models trained on manipulations specified bykeep\-pvs\.remove\-p\.
### 4\.1Manipulations
We consider different levels at which preemption \(or non\-preemption\) might affect model behavior:
#### Verb\-specific
This condition specifies the canonical form of preemption\([Goldberg, 2005](https://arxiv.org/html/2609.01794#bib.bib24);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4);[Ambridge et al\., 2018](https://arxiv.org/html/2609.01794#bib.bib3);[Bidgood et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib10)\)\. Here, preemption for a given verb is only measured with respect to the exposure of the learner to the preemptive evidence for that particular verb, regardless of the learner’s experience with usages of other verbs\. Therefore,keep\-pandremove\-pin this condition are only concerned with the periphrastic causative usages of a given verb at a time\. This results in a total of 60 manipulations, resulting in 5 random seeds, 2 manipulation types, and 6 verbs\.
#### Abstract
The abstract condition considers the effect of other verbs’ preemptive usages in the learner’s experience\. Here, we measure the effect of the whole category that is denoted by the structure specified as the preemptive evidence\. That is, forremove\-p, we remove periphrastic causative usages ofallverbs \(NN=7,452\), and forkeep\-p, we remove equivalent amounts of \(randomly sampled\) non\-periphrastic causative usages of all verbs\. This results in a total of 10 manipulations, resulting from 5 seeds and 2 conditions\.
### 4\.2Comparisons to Alternating Verbs
We compare our controlled rearing results on the six target verbs to two additional verbs—namely,breakandstop\. These verbs permit both transitive and intransitive usages, in addition to periphrastic causative ones—namely,breakandstop\. This results in 20 more verb\-specific manipulations \(2 verbs, 5 seeds, and 2 manipulation types\), but keeps the abstract manipulations the same\.
We expect models to show seemingly positive effects of preemption for these verbs\. This is because in theremove\-pcondition, the models can still learn transitive preferences for these verbs from their unaffected transitive cases \(by contrast, the target verbs are intransitive, and do not occur with transitive usages, so there is no direct evidence for them\)\. Similarly, in thekeep\-pcase, it is likely that a significant amount of non\-preemptive usages for these verbs that are removed are transitive\. Thus, the models should find the transitive sentences in our test set to be more surprising in thekeep\-pcondition than theremove\-pcondition\.
### 4\.3Results and Analysis
We compute the NLLs \(as described in[§2\.4](https://arxiv.org/html/2609.01794#S2.SS4)\) for the transitive\-sentence subset of our evaluation dataset across all our manipulations\. We then analyze our results using a linear mixed\-effects regression\([Bates et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib9)\), using the NLL as our dependent variable, the target verb \(verb\) and the training set variant \(𝒟\\mathcal\{D\},keep\-pvs\.remove\-p\) and their interaction as fixed effects, and the random seed \(seed\) and evaluation sentence \(si−s^\{\-\}\_\{i\}\) as random effects\. This is specified by the following equation:
ℓ\(si−\)∼verb×𝒟\+\(1∣si−:seed\)\\displaystyle\\ell\(s\_\{i\}^\{\-\}\)\\sim\\texttt\{verb\}\\times\\mathcal\{D\}\+\(1\\mid s^\{\-\}\_\{i\}:\\texttt\{seed\}\)We compute this separately for the verb\-specific and abstract conditions, and analyze the contrast betweenkeep\-pandremove\-pacross verbs\. Insofar as the model shows an effect of preemption, we should expect this contrast to be positive—i\.e\., the average NLL forremove\-pshould be lower than that ofkeep\-p\. This is because the preemptive evidence is not present inremove\-p, so the unconventional form should no longer be blocked by the model and therefore should be likely\.
Figure 3:Contrast of estimated marginal means betweenkeep\-pandremove\-pon the linear mixed\-effects model, with 95% confidence intervals, across verb\-specific \(blue, square\) and abstract \(black, circle\) conditions\. Positive values indicate preemption\. We additionally include results on two alternating verbs—breakandstop\.[Figure3](https://arxiv.org/html/2609.01794#S4.F3)shows our results across verbs and manipulation conditions\. For our target intransitive verbs in the verb\-specific condition, instead of seeing a privileged \(positive\) effect of preemption, we instead see largely negative effects across verbs\. That is, the models’ NLL on the unconventional, transitive usages for our target verbs is generallylowerwhen the hypothesized preemptive evidence is removed, than when it is kept\. This is in direct opposition to what is specified by the preemption proposal\([Goldberg, 1995](https://arxiv.org/html/2609.01794#bib.bib23);[Brooks and Zizak, 2002](https://arxiv.org/html/2609.01794#bib.bib15);[Boyd and Goldberg, 2011](https://arxiv.org/html/2609.01794#bib.bib12);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4)\)\.
For the abstract condition, however, we see relatively stronger positive effects of preemption than in the verb\-specific case\. That is, our LMs are more affected by theaccumulatedindirect negative evidence from periphrastic causative usages ofallverbs than by the evidence only for a single verb at a time\. This suggests that preemption effects \(if any\) are more likely to be an abstract effect rather than item\-specific, at least from the statistics contained in CHILDES\.555These results hold even if we remove periphrastic causative usages only for all intransitive verbs, leaving those of alternating verbs in the training set\. See[AppendixE](https://arxiv.org/html/2609.01794#A5)for these results\.
On comparing these results to those from performing controlled rearing onbreakandstop, we see that the latter indeed show more seemingly positive evidence of verb\-specific preemption—per our expectation\. But it is important to note that for these verbs, direct evidence for transitive usages still exists in theremove\-pcondition\. Therefore, even if there are no periphrastic causative usages in this condition, LMs can still pick up on the explicit transitive usages for both verbs \(John broke the cup,Mary stopped the car, etc\.\)\.
### 4\.4Interim Discussion
Our verb\-specific results from the previous experiment suggest that CHILDES\-trained LMs find transitive usages of \(generally\) intransitive verbs \(likecryandsmile\)less—as opposed tomore—likely when their periphrastic causative usages are removed\. That is, the models do not treat the periphrastic causative as an indirect negative evidence\.
A plausible explanation of why this might be the case comes from earlier controlled rearing results, where models were shown to be sensitive to indirectpositiveevidence\. For instance,[Misra and Mahowald \(2024\)](https://arxiv.org/html/2609.01794#bib.bib36)show that knowledge of constructions such asa beautiful five dayscan come from other cases where a plural measure noun phrase \(6 months\) is construed as a singular unit \(e\.g\.,6 monthsisall I need…\) even in the absence of direct evidence\([Jumelet et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib29);[Patil et al\., 2024](https://arxiv.org/html/2609.01794#bib.bib38);[Yao et al\., 2025](https://arxiv.org/html/2609.01794#bib.bib54), see also\)\. Therefore, instead of encoding thecompetitionbetween periphrastic causative and transitive constructions, the models might only be encoding their similarity\.666It is important to note that encoding similarity is a precursor to encoding competition—the learner must recognise what structures are similar before concluding that they also compete\([Suttle and Goldberg, 2011](https://arxiv.org/html/2609.01794#bib.bib46);[Goldberg, 2016](https://arxiv.org/html/2609.01794#bib.bib25)\)\.This can arise from alternating verbs such astickle, boil, melt, move, grow, etc\. thatcanbe used in both constructions\. This is perhaps why we see an effect of preemption in the abstract condition, whereremove\-peffectively removes evidence that the transitive and periphrastic causative can be related\.
We will explore this hypothesis of indirectpositive—and notnegative—evidence, by turning to a training dynamics analysis in the next section\.
## 5Prospective Moments of Indirect Negative \(or Positive?\) Evidence
In this section, we explore model generalization behavior at an extremely fine\-grained level, by testing it at specific time\-steps during pretraining\. Specifically, we identify ‘Prospective Moments of Indirect Negative Evidence’, which we define as training batches that contain constructions that may serve as the hypothesized negative evidence against overgeneralization\. On identifying such cases, we then measure the model’s behavior before and after encounter with a batch of interest, to quantify its effect \(given encounters with previous batches\) on the model’s overgeneralization behavior\.
Figure 4:Depiction of aProspective Moment of Indirect Negative Evidence\. We identify batches that contain the periphrastic causative usage of a verb \(here,laugh\) and then measure the NLL of an evaluation sentence before \(pre\) and after \(post\) encountering the batch\.In the context of the previous experiment’s findings, we specifically narrow in on batches containing the periphrastic causative usages of verbs\. This is because, according to preemption, these were meant to be sources of indirect negative evidence, but our results suggest that the opposite is true\. That is, models trained on corpora where this evidence was removed also found transitives to belesslikely than when this evidence was not removed\. This suggests that these moments might in fact be supplying models with indirectpositiveevidence that the two structures are similar\.
Figure 5:Statistical relationship between transitive and periphrastic causative NLLs, measured at every batch during training that contains a periphrastic causative exposure \(indicated by individual points\) to the verb of interest\.To shed light on this hypothesis, we measure the correlation between the NLL on our evaluation data’s periphrastic causative and transitive usages for each verb, using the models \(5 different seeds\) trained on the full CHILDES dataset\. Specifically, for a batch containing a given verb’s usage in the periphrastic\-causative construction, we compute the difference between a model’s average NLL on an evaluation set before and after encountering that batch\. We do this for both periphrastic causative and transitive evaluation sentences\. We then measure the Pearson’s correlation and the pairwise agreement between the NLL differences for both constructions\. Agreement is measured as the percentage of time both values had the same sign\. A positive correlation and high agreement would indicate that the losses on both structures change in the same way upon encounters with a periphrastic causative usage during training\. For reference, according to preemption, we should instead expect a negative correlation and agreement—the loss on transitives should increase with exposure to a periphrastic causative, whereas loss on periphrastic causatives should decrease\.
[Figure5](https://arxiv.org/html/2609.01794#S5.F5)shows our results\. We see generally high correlation and agreement scores across all verbs, corroborating the indirect positive evidence hypothesis\. That is, on encountering a periphrastic causative usage of a verb during training, the change in model’s behavior on both periphrastic causativesandtransitives are positively correlated\.
## 6Implications and Speculations
We find that at least at the verb\-specific level, LMs trained on CHILDES treat the periphrastic causative construction as indirect evidencefor, and notagainst, transitive usages of intransitive verbs\. Below we discuss the broader implications and speculations of these findings:
### 6\.1Semantic cues against overgeneralization
Preemption as a mechanism attributes a more special role to the communicative intent of contexts in which the preemptive evidence is used as compared to entrenchment\. Models’ failure to show effects of preemption in the verb\-specific condition shows that the language\-only cues in CHILDES are not sufficient to give rise to this sensitivity in models\. But this also raises the question about whether providing explicit semantic evidence of events could facilitate preemption in LMs\. This could come from environmental annotations—cues to what is going on in the child’s environment at the time of the utterance—that are made available in CHILDES \(but were not used in our work, though see[Wu et al\. \(2025\)](https://arxiv.org/html/2609.01794#bib.bib51), who did use them\)\. This could also be induced through more explicit means, like in[Yedetore and Kim \(2024\)](https://arxiv.org/html/2609.01794#bib.bib55), who showed evidence of hierarchical generalization in transformers when semantic signals \(in the form of semantic parses\) were available during training\.
### 6\.2Morphological origins of preemption
While we investigated preemption for argument structure constructions in this work, the phenomenon traces its origins to morphological ‘blocking’\([Aronoff, 1976](https://arxiv.org/html/2609.01794#bib.bib7);[Kiparsky, 1973](https://arxiv.org/html/2609.01794#bib.bib30);[Clark, 1987](https://arxiv.org/html/2609.01794#bib.bib19)\)\. For example,childrenpreemptschildsorwentpreemptsgoed\. Availability of such a phenomenon during learning could be the key to preemption\. Insofar as this is true, a hypothesis for why we might not see strong effects of preemption in our LMs is that tokens likechildsandgoedare simply not represented in neither our BPE nor our lemmatized variants, and so the model is never preempted in the morphological sense\. This motivates serious consideration of morphologically\-informed tokenizers in the development of LMs\.
### 6\.3Preemption from abstract cues
While we do not observe preemption in our verb\-specific experiments, we do see evidence of more abstract signals to preemption via exposures from other verbs\. In the context of human language acquisition, preemption has always been treated as a verb\-specific phenomenon\([Goldberg, 2005](https://arxiv.org/html/2609.01794#bib.bib24);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4);[Ambridge et al\., 2018](https://arxiv.org/html/2609.01794#bib.bib3);[Bidgood et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib10)\)\. Our results suggest the possibility of preemptive signals coming from more accumulated sources of indirect evidence—e\.g\., transitive usages oflaughcould come from periphrastic causative usages oflaugh, combined with those from other verbs \(likegiggle,smile, etc\.\)\. This motivatesnovelhuman experiments, perhaps via artificial language learning paradigms\([Samara et al\., 2025](https://arxiv.org/html/2609.01794#bib.bib44), e\.g\.,\), to test if this is a viable route for human learners\. Importantly, this does not necessarily invalidate any potential finding of preemption in humans, which have been concluded from using corpus evidence in the past\([Boyd and Goldberg, 2011](https://arxiv.org/html/2609.01794#bib.bib12);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4);[Ambridge et al\., 2018](https://arxiv.org/html/2609.01794#bib.bib3);[Bidgood et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib10)\)\. However, it does raise the question whether such estimates \(which are done at the verb level\) might by strengthened if the effect of other verbs’ usages is also considered\.
## 7Conclusion
Overall, our work joins several others in understanding how the input to a general purpose statistical learner like a transformer LM shapes its generalization behavior\. In particular, it showcases how subtly different proposals for indirect negative evidence can be disentangled, especially when doing so is impossible for human learners\([Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4)\)\. In closing, this leaves us optimistic about the potential of LMs to contribute in\-principle insights about how statistical learning might unfold in language learners\([Contreras Kallens et al\., 2023](https://arxiv.org/html/2609.01794#bib.bib20);[Futrell and Mahowald, 2025](https://arxiv.org/html/2609.01794#bib.bib22);[Misra and Kim, 2026](https://arxiv.org/html/2609.01794#bib.bib35)\)\.
## Limitations
#### Single Phenomenon
Due to the number of pretraining experiments, we focus on a single phenomenon, although it is among the quintessential examples of this debate which has lasted around 40 years\([Bowerman, 1988](https://arxiv.org/html/2609.01794#bib.bib11);[Akhtar and Tomasello, 1997](https://arxiv.org/html/2609.01794#bib.bib2);[Brooks and Tomasello, 1999](https://arxiv.org/html/2609.01794#bib.bib14);[Brooks and Zizak, 2002](https://arxiv.org/html/2609.01794#bib.bib15);[Theakston, 2004](https://arxiv.org/html/2609.01794#bib.bib47);[Ambridge et al\., 2008](https://arxiv.org/html/2609.01794#bib.bib6);[Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4);[Ambridge et al\., 2018](https://arxiv.org/html/2609.01794#bib.bib3);[Bidgood et al\., 2021](https://arxiv.org/html/2609.01794#bib.bib10)\)\. It is possible that models show different generalization behaviors on different constructions\. Our methodology, and the methodology of controlled rearing in general can be used to adjudicate between preemption and entrenchment across several different constructions in the future\.
#### Generalization of our conclusion across models
Our analyses have only looked at one instance of the entire class of transformer LMs that can be trained on CHILDES, finding lack of evidence for preemption at the level of individual verb exposures\. However, this does not necessarily mean that it cannot arise in transformers or neural networks in general\. In fact, there is a connectionist model that does show effects of both preemption and entrenchment\([Ambridge and Blything, 2016](https://arxiv.org/html/2609.01794#bib.bib5)\)—however its architecture explicitly bakes in hand\-coded semantic features which have to be supplied by the researcher for a given input\. That is, this model is at a different level of analysis and theoretical commitment than ours, and cannot learn from naturalistic data\. Furthermore, it would also be worthwhile to understand preemption from a theory of computation perspective, in order to provably analyze preemption in classes of transformers \(or other types of neural network models\)\. Finally, it is possible for these effects to be stronger with larger datasets, e\.g\., corpora from the BabyLM challenge\([Choshen et al\., 2026](https://arxiv.org/html/2609.01794#bib.bib17)\), or larger models trained on such larger datasets\.
#### Operationalizing preemptive signals
We have assumed that all instances of the competing construction constitute cases where either construction could have been used\. But that need not be true\. More generally, there is a need for a methodology to detect the possibility of competition vs\. non\-competition from raw data\. Additionally, it might be too narrow of an assumption to treat only near\-synonymous constructions as preemptive\([Ambridge et al\., 2015](https://arxiv.org/html/2609.01794#bib.bib4)\)\. For instance, in a scenario where something \(a joke\) makes someone laugh \(John\), it might be perfectly ok to simply use the intransitive,John laughed\. In this sense, the intransitive could have preempted the transitive usage\. Therefore, future work could address this by running controlled rearing on specific cases where constructions necessarily compete vs\. do not\. Finally, what counts as a “semantic context” can itself differ\. For instance, in[Misra and Kim \(2026\)](https://arxiv.org/html/2609.01794#bib.bib35), the features of the constructional exposure itself ended up showing an effect of preemption\. Their case focused on the dative alternation\([Levin, 1993](https://arxiv.org/html/2609.01794#bib.bib32), DO:John gave me the bookvs\. PO:John gave the book to me;\)and found that the information structure of the theme \(the book\) and recipient \(me\) predicted the degree of preemption\. Future work can hopefully consolidate the many different ways preemption can be cued to learners\.
## Acknowledgments
We are grateful to Adele Goldberg, Caroline Rowland, and Qing Yao for valuable comments and advice on earlier versions of our experiments\. We also acknowledge Reviewer zAa3 for their thoughtful review during the ARR May Cycle—their comments resulted in our control experiments\. KM was supported by a Donald D\. Harrington Faculty Fellowship at UT Austin for the 2025/26 academic year\. FS was supported by a Canada CIFAR AI Chair award\.
## References
- Abend et al\. \(2017\)Omri Abend, Tom Kwiatkowski, Nathaniel J Smith, Sharon Goldwater, and Mark Steedman\. 2017\.[Bootstrapping language acquisition](https://www.sciencedirect.com/science/article/pii/S0010027717300495)\.*Cognition*, 164:116–143\.
- Akhtar and Tomasello \(1997\)Nameera Akhtar and Michael Tomasello\. 1997\.[Young children’s productivity with word order and verb morphology\.](https://psycnet.apa.org/doi/10.1037/0012-1649.33.6.952)*Developmental Psychology*, 33\(6\):952\.
- Ambridge et al\. \(2018\)Ben Ambridge, Libby Barak, Elizabeth Wonnacott, Colin Bannard, and Giovanni Sala\. 2018\.[Effects of Both Preemption and Entrenchment in the Retreat from Verb Overgeneralization Errors: Four Reanalyses, an Extended Replication, and a Meta\-Analytic Synthesis](https://doi.org/10.1525/collabra.133)\.4\(1\):23\.
- Ambridge et al\. \(2015\)Ben Ambridge, Amy Bidgood, Katherine E Twomey, Julian M Pine, Caroline F Rowland, and Daniel Freudenthal\. 2015\.[Preemption versus Entrenchment: Towards a Construction\-General Solution to the Problem of the Retreat from Verb Argument Structure Overgeneralization](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0123723)\.*PloS One*, 10\(4\):e0123723\.
- Ambridge and Blything \(2016\)Ben Ambridge and Ryan P Blything\. 2016\.A connectionist model of the retreat from verb argument structure overgeneralization\.*Journal of Child Language*, 43\(6\):1245–1276\.
- Ambridge et al\. \(2008\)Ben Ambridge, Julian M Pine, Caroline F Rowland, and Chris R Young\. 2008\.The effect of verb semantic class and verb frequency \(entrenchment\) on children’s and adults’ graded judgements of argument\-structure overgeneralization errors\.*Cognition*, 106\(1\):87–129\.
- Aronoff \(1976\)Mark Aronoff\. 1976\.Word formation in generative grammar\.*Linguistic Inquiry Monographs Cambridge, Mass*, \(1\):1–134\.
- Baker \(1979\)Carl L\. Baker\. 1979\.Syntactic theory and the projection problem\.*Linguistic Inquiry*, 10\(4\):533–581\.
- Bates et al\. \(2015\)Douglas Bates, Martin Mächler, Ben Bolker, and Steve Walker\. 2015\.Fitting Linear Mixed\-Effects Models Using lme4\.*Journal of Statistical Software*, 67:1–48\.
- Bidgood et al\. \(2021\)Amy Bidgood, Julian Pine, Caroline Rowland, Giovanni Sala, Daniel Freudenthal, and Ben Ambridge\. 2021\.[Verb argument structure overgeneralisations for the English intransitive and transitive constructions: Grammaticality judgments and production priming](https://doi.org/10.1017/langcog.2021.8)\.13\(3\):397–437\.
- Bowerman \(1988\)Melissa Bowerman\. 1988\.The ‘no negative evidence’ problem: How do children avoid constructing an overly general grammar?In*Explaining Language Universals*, pages 73–101\. Basil Blackwell\.
- Boyd and Goldberg \(2011\)Jeremy K Boyd and Adele E Goldberg\. 2011\.[Learning What NOT to Say: The Role of Statistical Preemption and Categorization in A\-Adjective Production](https://muse.jhu.edu/pub/24/article/422023/summary?casa_token=o1EYHXyqXZoAAAAA:eYJ0DZj9o0emyShz2OXCLSPCE-Dv7sucATout4yuKcXwJaFamxj7SQBX07DUzyAVnuc-OavwNg)\.*Language*, 87\(1\):55–83\.
- Braine and Brooks \(1995\)Martin DS Braine and Patricia J Brooks\. 1995\.[Verb argument structure and the problem of avoiding an overgeneral grammar](https://www.taylorfrancis.com/chapters/edit/10.4324/9781315806860-16/verb-argument-structure-problem-avoiding-overgeneral-grammar-martin-braine-patricia-brooks)\.In*Beyond names for things: Young children’s acquisition of verbs*, pages 353–376\. Psychology Press\.
- Brooks and Tomasello \(1999\)Patricia J Brooks and Michael Tomasello\. 1999\.[How children constrain their argument structure constructions](https://www.jstor.org/stable/417731)\.*Language*, pages 720–738\.
- Brooks and Zizak \(2002\)Patricia J Brooks and Otto Zizak\. 2002\.[Does preemption help children learn verb transitivity?](https://www.cambridge.org/core/services/aop-cambridge-core/content/view/B32BB1A03F956BF681954412958572D3/S0305000902005287a.pdf/does_preemption_help_children_learn_verb_transitivity.pdf)*Journal of Child Language*, 29\(4\):759–781\.
- Brown and Hanlon \(1970\)Roger Brown and Camille Hanlon\. 1970\.Derivational complexity and order of acquisition in child speech\.In J\. R\. Hayes, editor,*Cognition and the Development of Language*, pages 11–53\. Wiley, New York\.
- Choshen et al\. \(2026\)Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Jaap Jumelet, Tal Linzen, Aaron Mueller, Suchir Salhan, Raj Sanjay Shah, Alex Warstadt, and Ethan Gotlieb Wilcox\. 2026\.[Babylm turns 4 and goes multilingual: Call for papers for the 2026 babylm workshop](https://arxiv.org/abs/2602.20092)\.*arXiv preprint arXiv:2602\.20092*\.
- Chouinard and Clark \(2003\)Michelle M Chouinard and Eve V Clark\. 2003\.Adult reformulations of child errors as negative evidence\.*Journal of child language*, 30\(3\):637–669\.
- Clark \(1987\)Eve V Clark\. 1987\.The principle of contrast: A constraint on language acquisition\.In Brian MacWhinney, editor,*Mechanisms of Language Acquisition*\. Hillsdale\.
- Contreras Kallens et al\. \(2023\)Pablo Contreras Kallens, Ross Deans Kristensen\-McLachlan, and Morten H Christiansen\. 2023\.[Large language models demonstrate the potential of statistical learning in language](https://onlinelibrary.wiley.com/doi/full/10.1111/cogs.13256)\.*Cognitive Science*, 47\(3\):e13256\.
- Frank \(2023\)Michael C Frank\. 2023\.[Bridging the data gap between children and large language models](https://www.cell.com/trends/cognitive-sciences/abstract/S1364-6613(23)00203-6)\.*Trends in Cognitive Sciences*, 27:990–992\.
- Futrell and Mahowald \(2025\)Richard Futrell and Kyle Mahowald\. 2025\.[How Linguistics Learned to Stop Worrying and Love the Language Models](https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/how-linguistics-learned-to-stop-worrying-and-love-the-language-models/2395C24EB472B60B2514F7D5F93EB9A8)\.*Behavioral and Brain Sciences*, pages 1–98\.
- Goldberg \(1995\)Adele E Goldberg\. 1995\.*Constructions: A construction grammar approach to argument structure*\.University of Chicago Press\.
- Goldberg \(2005\)Adele E Goldberg\. 2005\.*Constructions at Work: The Nature of Generalization in Language*\.Oxford University Press\.
- Goldberg \(2016\)Adele E Goldberg\. 2016\.Partial productivity of linguistic constructions: Dynamic categorization and statistical preemption\.*Language and cognition*, 8\(3\):369–390\.
- Guo et al\. \(2026\)Dongxin Guo, Jikun Wu, and SM Yiu\. 2026\.[Do Language Models Know What Not to Say? Causal Evidence for Statistical Preemption in LLMs](https://openreview.net/pdf?id=sX52AcD4rp)\.In*30th Conference on Computational Natural Language Learning*\.
- Hastie and Tibshirani \(1986\)Trevor Hastie and Robert Tibshirani\. 1986\.Generalized Additive Models\.*Statistical Science*, 1\(3\):297–310\.
- Huebner et al\. \(2021\)Philip A\. Huebner, Elior Sulem, Fisher Cynthia, and Dan Roth\. 2021\.[BabyBERTa: Learning more grammar with small\-scale child\-directed language](https://doi.org/10.18653/v1/2021.conll-1.49)\.In*Proceedings of the 25th Conference on Computational Natural Language Learning*, pages 624–646, Online\. Association for Computational Linguistics\.
- Jumelet et al\. \(2021\)Jaap Jumelet, Milica Denic, Jakub Szymanik, Dieuwke Hupkes, and Shane Steinert\-Threlkeld\. 2021\.[Language models use monotonicity to assess NPI licensing](https://doi.org/10.18653/v1/2021.findings-acl.439)\.In*Findings of the Association for Computational Linguistics: ACL\-IJCNLP 2021*, pages 4958–4969, Online\. Association for Computational Linguistics\.
- Kiparsky \(1973\)Paul Kiparsky\. 1973\.“Elsewhere” in Phonology\.In Stephen R\. Anderson and Paul Kiparsky, editors,*A Festschrift for Morris Halle*, pages 93–106\. Holt, Rinehart, and Winston, New York\.
- Leong and Linzen \(2026\)Cara Su\-Yi Leong and Tal Linzen\. 2026\.[Manipulating language models’ training data to study syntactic constraint learning: The case of English passivization](https://www.sciencedirect.com/science/article/pii/S0749596X26000215)\.*Journal of Memory and Language*, 149:104751\.
- Levin \(1993\)Beth Levin\. 1993\.*English Verb Classes and Alternations: A Preliminary Investigation*\.University of Chicago Press\.
- Ma et al\. \(2025\)Ziqiao Ma, Shuyu Wu, Yixuan Wang, Joyce Chai, and Freda Shi\. 2025\.[Trabank: A toolkit for computational developmental studies in language models](https://github.com/Mars-tin/PyChildes)\.
- MacWhinney \(2000\)Brian MacWhinney\. 2000\.*The CHILDES Project: Tools for Analyzing Talk*, 3rd edition\.Lawrence Erlbaum Associates, Mahwah, NJ\.
- Misra and Kim \(2026\)Kanishka Misra and Najoung Kim\. 2026\.[A systematic framework for generating novel experimental hypotheses from language models](https://arxiv.org/abs/2408.05086)\.*arXiv preprint arXiv:2408\.05086*\.
- Misra and Mahowald \(2024\)Kanishka Misra and Kyle Mahowald\. 2024\.[Language models learn rare phenomena from less rare phenomena: The case of the missing AANNs](https://doi.org/10.18653/v1/2024.emnlp-main.53)\.In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 913–929, Miami, Florida, USA\. Association for Computational Linguistics\.
- Padovani et al\. \(2026\)Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet, Yevgen Matusevych, Nathan Schneider, and Arianna Bisazza\. 2026\.[CAIT: A Syntactic Parsing Toolkit for Child\-Adult InTeractions](https://arxiv.org/abs/2605.19718)\.*arXiv preprint arXiv:2605\.19718*\.
- Patil et al\. \(2024\)Abhinav Patil, Jaap Jumelet, Yu Ying Chiu, Andy Lapastora, Peter Shen, Lexie Wang, Clevis Willrich, and Shane Steinert\-Threlkeld\. 2024\.[Filtered corpus training \(FICT\) shows that language models can generalize from indirect evidence](https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00720/125485)\.*Transactions of the Association for Computational Linguistics*, 12:1597–1615\.
- Portelance and Jasbi \(2024\)Eva Portelance and Masoud Jasbi\. 2024\.The roles of neural networks in language acquisition\.*Language and Linguistics Compass*, 18\(6\):e70001\.
- Qwen Team \(2026\)Qwen Team\. 2026\.[Qwen3\.5: Towards native multimodal agents](https://qwen.ai/blog?id=qwen3.5)\.
- Radford et al\. \(2019\)Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, and 1 others\. 2019\.Language Models are Unsupervised Multitask Learners\.*OpenAI blog*, page 9\.
- Rowland \(2013\)Caroline Rowland\. 2013\.*Understanding Child Language Acquisition*\.Routledge\.
- Rowland et al\. \(2025\)Caroline F Rowland, Gert Westermann, Anna L Theakston, Julian M Pine, Padraic Monaghan, and Elena VM Lieven\. 2025\.[Constructing language: A framework for explaining acquisition](https://www.cell.com/trends/cognitive-science/fulltext/s1364-6613(25)00142-1)\.*Trends in Cognitive Sciences*\.
- Samara et al\. \(2025\)Anna Samara, Elizabeth Wonnacott, Gaurav Saxena, Ramya Maitreyee, Judit Fazekas, and Ben Ambridge\. 2025\.[Learners restrict their linguistic generalizations using preemption but not entrenchment: Evidence from artificial\-language\-learning studies with adults and children\.](https://psycnet.apa.org/fulltext/2024-89933-001.html)*Psychological Review*, 132\(1\):1\.
- Stefanowitsch \(2008\)Anatol Stefanowitsch\. 2008\.Negative entrenchment: A usage\-based approach to negative evidence\.*Cognitive Linguistics*, 19\(3\):513\.
- Suttle and Goldberg \(2011\)Laura Suttle and Adele E Goldberg\. 2011\.[The partial productivity of constructions as induction](https://adele.princeton.edu/wp-content/uploads/sites/277/2015/01/SuttleGoldbergLinguistics11.pdf)\.*Linguistics*, 49\(6\):1237–1269\.
- Theakston \(2004\)Anna L Theakston\. 2004\.[The role of entrenchment in children’s and adults’ performance on grammaticality judgment tasks](https://www.sciencedirect.com/science/article/pii/S088520140300056X)\.*Cognitive Development*, 19\(1\):15–34\.
- Tomasello \(2003\)Michael Tomasello\. 2003\.*Constructing a Language: A Usage\-Based Theory of Language Acquisition*\.Harvard University Press\.
- Warstadt and Bowman \(2022\)Alex Warstadt and Samuel R Bowman\. 2022\.What Artificial Neural Networks Can Tell Us about Human Language Acquisition\.*Algebraic Structures in Natural Language*, page 17\.
- Wolf et al\. \(2020\)Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, and 3 others\. 2020\.[Transformers: State\-of\-the\-art natural language processing](https://aclanthology.org/2020.emnlp-demos.6/)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pages 38–45, Online\. Association for Computational Linguistics\.
- Wu et al\. \(2025\)Shuyu Wu, Ziqiao Ma, Xiaoxi Luo, Yidong Huang, Josue Torres\-Fonseca, Freda Shi, and Joyce Chai\. 2025\.[The mechanistic emergence of symbol grounding in language models](https://arxiv.org/abs/2510.13796)\.*arXiv preprint arXiv:2510\.13796*\.
- Yang et al\. \(2026\)Xiulin Yang, Arianna Bisazza, Nathan Schneider, and Ethan Gotlieb Wilcox\. 2026\.[A unified assessment of the poverty of the stimulus argument for neural language models](https://arxiv.org/abs/2602.09992)\.*arXiv preprint arXiv:2602\.09992*\.
- Yang et al\. \(2025\)Xiulin Yang, Zhuoxuan Ju, Lanni Bu, Zoey Liu, and Nathan Schneider\. 2025\.[UD\-English\-CHILDES: A collected resource of gold and silver Universal Dependencies trees for child language interactions](https://aclanthology.org/2025.udw-1.6/)\.In*Proceedings of the Eighth Workshop on Universal Dependencies \(UDW, SyntaxFest 2025\)*, pages 52–58, Ljubljana, Slovenia\. Association for Computational Linguistics\.
- Yao et al\. \(2025\)Qing Yao, Kanishka Misra, Leonie Weissweiler, and Kyle Mahowald\. 2025\.[Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models](https://openreview.net/forum?id=h5SRsDax8v)\.In*Second Conference on Language Modeling*\.
- Yedetore and Kim \(2024\)Aditya Yedetore and Najoung Kim\. 2024\.[Semantic training signals promote hierarchical syntactic generalization in transformers](https://doi.org/10.18653/v1/2024.emnlp-main.235)\.In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 4059–4073, Miami, Florida, USA\. Association for Computational Linguistics\.
## Appendix ADetails for Training
Additional details for training hyperparameters are included in[Table2](https://arxiv.org/html/2609.01794#A1.T2)\. We pack the consecutive utterances in a single CHILDES file into sequences of 1024 tokens, and the<s\>token is used as the beginning of sentence and also the separator between two sentences\. We use the GPT\-2 implementation provided by Huggingface Transformers\([Wolf et al\., 2020](https://arxiv.org/html/2609.01794#bib.bib50)\)and a custom PyTorch Lightning training loop for training our LMs, and a single training run requires around 1\.5 hours on 1 NVIDIA L40s GPU\.
Table 2:Hyperparameters used for training our LMs\.
## Appendix BIdentification of Periphrastic Causative Constructions
Figure 6:An example of the dependency parse trees on the CHILDES dataset\. This is Eng\-NA/Bloom/Peter/020303, sentence 1028\.We use PyChildes\([Ma et al\., 2025](https://arxiv.org/html/2609.01794#bib.bib33)\)to parse the raw CHILDES files of Eng\-NA and Eng\-UK, and then extract morphological and dependency grammar annotations\. These two corpora of CHILDES are open to public access\. An example of the dependency parse trees is shown in[Figure6](https://arxiv.org/html/2609.01794#A2.F6)\. Then we use the following heuristics to identify periphrastic causative constructions in the CHILDES dataset\.
Forletandmake, we identify them as matrix verbs in a periphrastic causative construction if they satisfy the following conditions:
- •Its complement \(COMP\) child is a verb not in a to\-infinitive \(e\.g\.make something to eat\)\.
- •Either it has an object \(OBJ\) child \(e\.g\.make me laugh\), or its complement verb has a subject \(SUBJ\) child \(e\.g\.make him see it\)\.
Additionally, forlet, we require it to have a subject \(SUBJ\) within its predicate to exclude hortatives likelet’s go\. This is done by searching through its ancestor chain formed by complement \(COMP\), coordination \(COORD\), and conjunction \(CONJ\) arcs\.
Forget,cause, andforce, we use the same criteria except that the complement verb must be in a to\-infinitive \(e\.g\.get him to sleep\)\.
In the following sentences, some contain a causative \*have\* and some do not\."May I have more juice please ?" \(False\)"I have finished my juice \." \(False\)"I had the waiter bring more juice \." \(True\)"I had my juice refilled by the waiter \." \(True\)Analyze the following sentence and determine whether it contains a causative \*have\* or not\. Answer with True if it does and False if it does not\."\{sentence\}"
We include the statistics of some more verbs that we identify in their usage in CHILDES as having periphrastic causative constructions in[Table3](https://arxiv.org/html/2609.01794#A2.T3)\.
Table 3:Additional verb counts and periphrastic causative usage counts in the CHILDES dataset\. This table is not exaustive\.
## Appendix CDetails for Evaluation Data Synthesis
We use Qwen 3\.5 27B[Qwen Team \(2026\)](https://arxiv.org/html/2609.01794#bib.bib40)to synthesize evaluation sentences\. Our system prompt is as follows:
You are a child literature author tasked with choosing objects that fit naturally into simple, easy\-to\-understand scenes for young children’s literature\.
The main prompt is as follows, whereprevious\_objectscontains samples of previously used objects for this verb\. Forsubjectwe sample from 266 names and pronouns, and forobject\_candidateswe sample from names, pronouns and additionally 499 nouns likebaby,mommy,toy\. The language model is prompted to choose an object from 25 randomly sampled candidates that fits most naturally with the subject and verb\. The chosen object, along with the subject and verb, are then used to form the periphrastic causative sentence in the evaluation set\. The ungrammatical transitive counterpart is then generated by mutating the periphrastic causative sentence\.
\#\#\# InstructionsChoose the object from the candidates below that best fits naturally with the subject and verb, \\forming a plausible scene where \{usage\_description\}, as in the sentence "\{usage\_example\}"\.If the chosen object is a common noun, prefix it with the article "the" \(e\.g\. "plant" \-\> "the plant"\)\. \\If the chosen object is a proper name \(e\.g\. a person’s name\), do not add an article \(e\.g\. "Tom" stays "Tom"\)\.Where the object is the one who will perform the action \(not receive it\), prefer a person or animal over an inanimate noun\.If none of the candidates are suitable, answer with "none"\.\#\#\# Output FormatAnswer with only the chosen object, exactly as written in the candidate list but with the article \\added if required, or "none"\.\#\#\# Input\- Subject: \{subject\}\- Verb: \{verb\}\- Object Candidates: \{object\_candidates\}\{already\_selected\_line\}\- Already selected for this verb \(prefer a different one\): \{previous\_objects\}"
Additionally, we lemmatize the evaluation set with Spacy, using model`en\_core\_web\_trf`version 3\.8\.0\.
## Appendix DFinal Checkpoint Results for the Precondition Experiment
[Figure7](https://arxiv.org/html/2609.01794#A4.F7)shows theΔ\\DeltaNLLs for the different verbs in the precondition experiment \([§3](https://arxiv.org/html/2609.01794#S3)\)\. We again see clean separation of model preferneces across different verb classes \(intransitive, alternating, transitive\)\.
Figure 7:Final CheckpointΔ\\DeltaNLLs across 5 seeds and 100 sentence pairs, per verb, across verb types\.
## Appendix EAdditional Results for Abstract Effect from Intransitive Verbs
In addition to[§4\.1](https://arxiv.org/html/2609.01794#S4.SS1), we test a manipulation condition in between verb\-specific and abstract, where we remove all periphrastic causative usages of all 174 intransitive verbs we identified in the CHILDES dataset\. Periphrastic causative usages of alternating verbs are kept intact\. This results in removing 4,675 usages, 62\.7% of all periphrastic causative usages in the CHILDES dataset\. In the total abstract setting,remove\-pmodels do not observe any periphrastic causative construction during training, while in the intransitive only setting,remove\-pmodels still observe periphrastic causative constructions of alternating verbs\. This allows intransitive onlyremove\-pmodels to establish the connection between periphrastic causative and transitive constructions\. Results are shown in[Figure8](https://arxiv.org/html/2609.01794#A5.F8)\. We see that the abstract effect from the intransitive verbs to intransitive target verbs are close to the total abstract effect, usually slightly weaker, at the exception oflaugh\(which shows a negative effect closer to the verb\-specific effect\) andsleep\(which has an insignificantly stronger positive effect than the total abstract effect\)\. The observed positive effect in preemption indicates that preemptive signals from other intransitive verbs—the verbs of same syntactic category—contribute a substantial portion of the total indirect evidence for the target verbs’ retreatment from overgeneralization\.
Figure 8:Contrast of estimated marginal means betweenkeep\-pandremove\-pon the linear mixed\-effects model, with 95% confidence intervals, across verb\-specific \(blue, square\), abstract intransitive only \(dark blue, triangle\), and abstract \(black, circle\) conditions\. Positive values indicate preemption\.
## Appendix FAdditional Results for Training Trajectories
We further plots the training trajectory of our controlled rearing conditions in[Figure9](https://arxiv.org/html/2609.01794#A6.F9), similar to[Figure2](https://arxiv.org/html/2609.01794#S3.F2)\. As discussed in[§2\.4](https://arxiv.org/html/2609.01794#S2.SS4), we show the NLL of overgeneralized transitive form only, instead of theΔ\\DeltaNLL between transitive and periphrastic causative forms\. Due to the aggregation over different evaluation sentences, the NLLs between different controlled rearing conditions shed the consistent pairwise differences we shown in[Figure3](https://arxiv.org/html/2609.01794#S4.F3), and the smoothing of curves makes it ineligble to compare the final NLL we used in the main experiments\. However, results here indicate that our models largely converging under 3 epoches of training without significant sign of overfitting\. And they also indicate that the overgeneralization tendency can be volatile during training\. We leave more systematic analysis of training trajectories to future work\.
Figure 9:Difference in NLL of transitive and periphrastic causative sentence pairs for 6 target intransitive verbs and 2 alternating verb controls at every 50 steps during controlled rearing training\. Each line represents a different training run using the same seed\.remove\-pmodels are shown in solid lines, whilekeep\-pmodels are shown in dashed lines\.
## Appendix GAdditional Results Without Lemmatization
Due to the nature of our training data and LM size, the commonly used Byte\-Pair Encoding \(BPE\) subword tokenization can introduce extra burden to the training process\. For example, the model will have to learn the connection between different word forms ofmake\(e\.g\.,made,making, etc\.\), which will be tokenized into different subword sequences\. Nevertheless, we repeat our main analyses under a BPE subword tokenization setting\. We first train a 8,192\-token BPE tokenizer on the CHILDES dataset, following[Huebner et al\. \(2021\)](https://arxiv.org/html/2609.01794#bib.bib28);[Misra and Kim \(2026\)](https://arxiv.org/html/2609.01794#bib.bib35)\. Then in both training and evaluation data pipelines we remove the lemmatization components, and swap the original word\-level tokenizer with this BPE tokenizer\. We then repeat the verb\-specific and abstract condition in[§4\.1](https://arxiv.org/html/2609.01794#S4.SS1)for 6 verbs and train 70 models \(5 seeds×\\times6 verbs×\\times2 manipulation types for verb\-specific, 5 seeds×\\times2 manipulation types for abstract\), and show results in[Figure10](https://arxiv.org/html/2609.01794#A7.F10)\. Under abstract level manipulation, the effects are stronger than the verb\-specific case, except forsneezeandsmile, whose verb\-specific effects also flipped from negative to positive\. The results are largely consistent with those using lemmatization\.
Figure 10:Contrast of estimated marginal means betweenkeep\-pandremove\-pLM variants trained and evaluated on BPE\-tokenized versions of datasets\.
## Appendix HA Reassessment of[Guo et al\. \(2026\)](https://arxiv.org/html/2609.01794#bib.bib26)
Our work is most closely related to that of[Guo et al\. \(2026\)](https://arxiv.org/html/2609.01794#bib.bib26), who measured preemption and entrenchment effects across three different linguistic phenomena \(including transitive overgeneralizations\), and across multiple off\-the\-shelf LMs\. They operationalize overgeneralization as the difference in LM surprisal \(negative log\-probability\) of the unconventional \(overgeneralized\) sentence vs\. that of the conventional sentence\. In the case of transitive overgeneralizations, their method measures difference in the surprisal of sentences like\*Tommy laughed meand those likeTommy made me laugh\. Using this measure, they find evidence for preemption over entrenchment in LMs\. They do so by first showing high correlation of the difference in surprisals with preemptive usages of verbs, even when controlling for the non\-preemptive usages\. Then, they fine\-tune an off\-the\-shelf GPT\-2 model\([Radford et al\., 2019](https://arxiv.org/html/2609.01794#bib.bib41)\)on 5000 instances of conventional vs\. unconventional usages, finding overwhelming effects of preemption over non\-preemption\.
On further scrutiny of[Guo et al\. \(2026\)](https://arxiv.org/html/2609.01794#bib.bib26)’s methods, we find their results to paint an illusory picture about preemption in LMs\. We highlight two concerns\. First, they use fine\-tuning as their method to claim causal evidence of preemption\. However, an LM might have already formed its preferences of avoiding the target overgeneralizations during pre\-training \(as shown by their results\)\. So, the results from these experiments do not sufficiently shed light on how the dispreference of overgeneralization was learned in the first place—this will require training data manipulations\. Second, and more importantly, their method explicitly includes the preemptive construction as part of the overgeneralization measure \(e\.g\.,periphrastic causativefortransitive\)\. That is, the conventional form in their measure is exactly the same form that is hypothesized to preempt the unconventional form\. Therefore, it is unsurprising that an effect of preemption is observed on fine\-tuning LMs on the conventional forms—change in the measure is directly being manipulated in their fine\-tuning experiment\. So any effect of preemption might just be driven by the increased likelihood of the conventional form rather than the unlikelihood of the unconventional form\.
Disentangling preemption from entrenchment will require direct manipulation of the training data, and a more precise measure of overgeneralization that focuses exclusively on the unconventional form\. Our methods in[§2](https://arxiv.org/html/2609.01794#S2)aim to embody this conclusions directly\.
## Appendix IAlternate Methods of Measuring Overgeneralization Effects
In this section, we compare our results to those if we had used[Guo et al\. \(2026\)](https://arxiv.org/html/2609.01794#bib.bib26)’s method of treating the difference in NLL of transitive and periphrastic causative usages as our measure of overgeneralization effects\. We replicate the above analysis using this measure \(computing NLLs of both subsets of our evaluation data\), and show results in[Figure11](https://arxiv.org/html/2609.01794#A9.F11)\. We see that ongo,laugh,sneezethere is seemingly positive effect of preemption\. But as our previous[Figure3](https://arxiv.org/html/2609.01794#S4.F3)reveals that the NLL of transitive is actually lower inkeep\-pthan inremove\-p, we can conclude that the positive effect observed here is driven by the NLL of the periphrastic causative loss being lower inkeep\-p\.
Figure 11:Contrast betweenkeep\-pandremove\-pon a similar linear mixed\-effects model as in[Figure3](https://arxiv.org/html/2609.01794#S4.F3), but with the NLL difference between the transitive and periphrastic causative sentences as the dependent variable\. This is the result we would get had we used[Guo et al\. \(2026\)](https://arxiv.org/html/2609.01794#bib.bib26)’s method\.Similar Articles
Do Language Models Know What Not to Say? Causal Evidence for Statistical Preemption in LLMs
This paper provides causal evidence that large language models acquire negative linguistic knowledge (what not to say) through statistical preemption, a mechanism from Construction Grammar, by showing that manipulating competing-form frequencies via fine-tuning shifts preemption behavior in predicted directions.
Linguistic Productivity in Large Language Models: Models Coerce, but do not Preempt
This paper investigates whether Large Language Models exhibit the same usage-based linguistic productivity constraints (entrenchment and preemption) as humans, finding that models can reproduce coercion but fail to apply statistical preemption to avoid overgeneralization.
Generalization Dynamics of LM Pre-training (17 minute read)
This paper reveals that during pre-training, language models frequently and suddenly switch between pattern-matching and generalization behaviors, a phenomenon called mode-hopping, and presents a toy evaluation suite to study it.
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
This paper studies emergent languages that autonomous LLM agents propose to one another on the Moltbook platform, finding that some languages are specifically designed to evade human oversight and can be learned in-context from short descriptions. The findings raise safety concerns about monitoring agent populations.
Exemplars in Disguise: Pure Exemplar Models Mimic Abstraction-First Learning
This paper critiques recent methods claiming abstraction-first learning in large language models, showing that pure exemplar models can mimic abstraction-first learning depending on input distributional properties.