Modeling the Developmental Shift in Telicity Acquisition

arXiv cs.CL Papers

Summary

This paper introduces a method using GPT-2 token surprisal to label telicity in child language, finding that child models rely on syntactic cues like post-verbal determiners, while adult models use semantic features, supporting syntactic bootstrapping theory.

arXiv:2609.17996v1 Announce Type: new Abstract: Acquiring telicity, which is the distinction between bounded (e.g., ate an apple) and unbounded (e.g., ate apples) events, requires first language (L1) learners to map surface-level and semantic cues to abstract event structures, but the computational trajectory of this mapping is not well understood. We introduce a Difference in Surprisal method that uses GPT2 token surprisal over paired temporal adverbial diagnostics (in an hour versus for an hour) to automatically label telicity across English CHILDES corpora, validated against expert linguist judgments. Using these labels, we train diagnostic logistic regression classifiers on 12 syntactic and lexical semantic features to compare how child speech and child-directed speech encode telicity. The two models diverge: the child model reaches near perfect accuracy through a single deterministic cue, the presence of a post-verbal determiner, while the adult model relies more heavily on verb class and other lexical semantic features, with the determiner cue neutralized. This trajectory supports Syntactic Bootstrapping: learners first exploit high-frequency structural cues as a scaffold to bootstrap, before developing fully compositional, verb-based event structures.
Original Article
View Cached Full Text

Cached at: 09/17/26, 09:03 AM

# Modeling the Developmental Shift in Telicity Acquisition
Source: [https://arxiv.org/abs/2609.17996](https://arxiv.org/abs/2609.17996)
[View PDF](https://arxiv.org/pdf/2609.17996)

> Abstract:Acquiring telicity, which is the distinction between bounded \(e\.g\., ate an apple\) and unbounded \(e\.g\., ate apples\) events, requires first language \(L1\) learners to map surface\-level and semantic cues to abstract event structures, but the computational trajectory of this mapping is not well understood\. We introduce a Difference in Surprisal method that uses GPT2 token surprisal over paired temporal adverbial diagnostics \(in an hour versus for an hour\) to automatically label telicity across English CHILDES corpora, validated against expert linguist judgments\. Using these labels, we train diagnostic logistic regression classifiers on 12 syntactic and lexical semantic features to compare how child speech and child\-directed speech encode telicity\. The two models diverge: the child model reaches near perfect accuracy through a single deterministic cue, the presence of a post\-verbal determiner, while the adult model relies more heavily on verb class and other lexical semantic features, with the determiner cue neutralized\. This trajectory supports Syntactic Bootstrapping: learners first exploit high\-frequency structural cues as a scaffold to bootstrap, before developing fully compositional, verb\-based event structures\.

## Submission history

From: Ellie Xia \[[view email](https://arxiv.org/show-email/5806bb7e/2609.17996)\] **\[v1\]**Wed, 16 Sep 2026 01:34:11 UTC \(915 KB\)

Similar Articles