Early Language Learning via Spreading Activation and Category Exploration in Complex Networks

arXiv cs.CL Papers

Summary

This paper models early language acquisition as a search on a graph-based mental lexicon using spreading activation and category exploration, outperforming a shortest path baseline in simulating normative word acquisition across four languages.

arXiv:2607.06258v1 Announce Type: new Abstract: Is word acquisition in children uneven with respect to semantic and lexical categories? To answer this question, we model early language learning as a search on a graph-based mental lexicon, driven by two interacting processes: spreading activation and an enforced exploration (rather than exploitation) of lexical categories. We evaluate model performance on four languages (German, English, Dutch, and Rioplatense Spanish), using CDIs as ground-truth data for lexical categories, normative ages derived from the Wordbank repository, and state-of-the-art resources for reconstructing graphs of word similarities. We find that spreading activation outperforms a shortest path baseline in simulating normative word acquisition. At the category level, we highlight complex transitions between CDIs. By studying their sequences in terms of burstiness and average persistence time within the same CDI, we find that spreading activation better captures the exploration dynamics observed empirically. Overall, our findings suggest that vocabulary development can be understood through the non-trivial interplay between activation dynamics and some degree of constraints regulating the visiting of lexical categories in complex networks.
Original Article
View Cached Full Text

Cached at: 07/08/26, 04:42 AM

# Early Language Learning via Spreading Activation and Category Exploration in Complex Networks
Source: [https://arxiv.org/abs/2607.06258](https://arxiv.org/abs/2607.06258)
[View PDF](https://arxiv.org/pdf/2607.06258)[HTML \(experimental\)](https://arxiv.org/html/2607.06258v1)

> Abstract:Is word acquisition in children uneven with respect to semantic and lexical categories? To answer this question, we model early language learning as a search on a graph\-based mental lexicon, driven by two interacting processes: spreading activation and an enforced exploration \(rather than exploitation\) of lexical categories\. We evaluate model performance on four languages \(German, English, Dutch, and Rioplatense Spanish\), using CDIs as ground\-truth data for lexical categories, normative ages derived from the Wordbank repository, and state\-of\-the\-art resources for reconstructing graphs of word similarities\. We find that spreading activation outperforms a shortest path baseline in simulating normative word acquisition\. At the category level, we highlight complex transitions between CDIs\. By studying their sequences in terms of burstiness and average persistence time within the same CDI, we find that spreading activation better captures the exploration dynamics observed empirically\. Overall, our findings suggest that vocabulary development can be understood through the non\-trivial interplay between activation dynamics and some degree of constraints regulating the visiting of lexical categories in complex networks\.

## Submission history

From: Salvatore Citraro \[[view email](https://arxiv.org/show-email/6fdc9121/2607.06258)\] **\[v1\]**Tue, 7 Jul 2026 13:25:18 UTC \(12,927 KB\)

Similar Articles

Cross-Lingual Exploration for Parametric Knowledge

arXiv cs.CL

This paper explores cross-lingual prompting strategies to improve access to parametric knowledge in large language models, demonstrating significant gains in knowledge transfer and factual recall across 17 languages on multilingual benchmarks.

Language Acquisition Device in Large Language Models

arXiv cs.CL

This paper proposes LAD-inspired pre-pretraining using a formal language called MP-Struct that encodes natural-language-like structures. It shows that this approach improves token efficiency and imparts human-like resistance to structurally implausible languages, challenging prior hypotheses about effective pre-pretraining languages.