Tag
A book examining the role of algebraic models versus deep learning in natural language acquisition, featuring perspectives from leading researchers in computational linguistics, psychology, and mathematical linguistics.
This paper introduces Semantic Field Theory (SFT), a computational model for lexical semantics that models meaning through semantic fields, contextual deformation, interaction terms, and energy minimization. It provides formal elements including Gaussian product closure, Möbius inversion for higher-order interactions, and stability conditions.
This study demonstrates that contextual semantic relevance, measuring how strongly an incoming word relates to its recent semantic context, reliably predicts fMRI BOLD responses during naturalistic speech comprehension across two datasets, whereas surprisal (local probabilistic expectation) does not. The findings support that slow hemodynamic responses are especially sensitive to contextual semantic integration rather than local prediction.
This paper models early language acquisition as a search on a graph-based mental lexicon using spreading activation and category exploration, outperforming a shortest path baseline in simulating normative word acquisition across four languages.
Emily Bender clarifies the original meaning of 'stochastic parrots' from her 2021 paper, debunking common misconceptions about LLMs and critiquing the term 'artificial intelligence' for overselling technology.
Svarna is an open-source web-based corpus workbench for Modern Greek, integrating multiple databases with over 507 million words and providing various linguistic analysis tools, released under MIT license.
This paper investigates how ethos and pathos appeals in social media messages resonate with silent readers, finding that rhetorical content leads to greater interpretive divergence and can predict audience attitudes toward the author.
This paper introduces UD_Czech-PDTC, a large and genre-diverse treebank for Czech in the Universal Dependencies framework, derived from the Prague Dependency Treebank-Consolidated. It describes the conversion process and differences between annotation schemes.
This paper proposes a benchmark suite grounded in Pāṇinian grammar to unify Indic language processing across languages, aiming to improve accuracy, data efficiency, and transferability.
The third edition of the Speech and Language Processing textbook by Jurafsky and Martin was released in January 2026, featuring a clear explanation of Transformers and various updates including new chapters on ASR, TTS, and DPO.
Dango is a 1.8B-parameter LLM trained strictly on Japanese (L1) then fine-tuned on English (L2) to study language transfer effects in second language acquisition. The model filters English contamination from the pretraining corpus and shows human-like L2 production patterns.
This paper presents multi-agent simulations of the emergence of morphological alternation patterns (like 'go/went') in language, using an AI Historical Linguist (LLM-driven) to evaluate plausibility of evolved morphologies against real languages.
CAF-Gen is a multi-agent LLM-driven framework that enriches shallow argument structures into formal Carneades Argumentation Framework models using an iterative Creator-Reviewer pipeline, achieving improved structural alignment and quality.
GlossAssist is a tool for creating interlinear glossed text (IGT) corpora in low-resource language documentation settings, built around the CWoMP retrieval-based architecture with an active learning feedback loop that improves predictions as annotators make corrections without retraining the model.
This article evaluates the integration of data from the French syntactic lexicon Lexicon-Grammar into a probabilistic parser, using word clustering methods on verbs to improve parsing accuracy for French.
This paper presents a modular framework for generating artificial lexicons that are pronounceable, typologically plausible, and semantically structured, using phoneme inventories from PHOIBLE and probabilistic grammars, outperforming deterministic baselines.
This paper proposes Scene Abstraction, a framework for constructing structured representations of the interpretive scenes that words evoke in context, using few-shot prompting of large language models. The authors introduce COCA-Scenes, a dataset of 520 usage instances, and provide empirical evidence that scenes are reliably identifiable and align better with human interpretation than alternatives.
Presents a novel pattern-and-root model for describing Arabic noun inflection, focusing on broken plurals, with a taxonomy of 160 classes and an encoding scheme applied to 3,200 entries, aiming to improve computational language resources.
This paper presents a data-driven analysis of multi-word expressions (MWEs) based on 16 theoretical criteria, annotated by linguistics experts, finding that no expressions are absolutely idiomatic and that lexical criteria are most influential.
The paper introduces IMLJD, a computational dataset designed for analyzing Indian matrimonial litigation, supporting natural language processing and legal analytics research.