Tag
This paper develops a vector logic for formal semantics, characterizing when compressed vector representations allow exact linear or affine readouts of truth conditions, with experiments on GloVe and word2vec embeddings.
This paper introduces a method using GPT-2 token surprisal to label telicity in child language, finding that child models rely on syntactic cues like post-verbal determiners, while adult models use semantic features, supporting syntactic bootstrapping theory.
FrameBench is a new benchmark for evaluating whether large language models can distinguish context-dependent frame-semantic interpretations of verbs, constructed for English and Japanese using FrameNet resources and released with code.
This paper proposes a layered taxonomy for annotating grammatical errors in Chinese learner writing, combining computational and pedagogical perspectives, and evaluates it through coverage analysis and consistency studies with language models.
This study uses multilingual encoders to measure framing differences across Wikipedia's language editions, finding that religious and political concepts diverge more than scientific ones.
SynFlow is an open-source toolkit for multidimensional diachronic semantic analysis, applying a shared workflow to linguistic representations for studying lexical semantic change across syntactic, morphological, and other dimensions.
This paper examines how shared language technology contributes to asymmetric discourse homogenization, using empirical evidence from Reddit discussions.
The paper presents an unsupervised method to extract linguistic metaphors and group them into conceptual metaphors, applying the approach to analyze framing differences in left- vs. right-leaning podcasts.
This paper introduces an activation-guided neuron intervention framework using Qwen3-8B to induce Alzheimer's-related computational language phenotypes, demonstrating that amplifying AD-associated neurons produces graded impairments in multiple cognitive domains.
This paper tests whether statistical measures used to argue that undeciphered scripts like the Indus script encode language are specific to language, using a generative emblem system called Sigil. It finds that these measures detect organization without being specific to language, undermining claims that they establish encoded speech.
This paper rethinks the linguistic notion of 'state' as a systemic morphosyntactic mechanism across synthetic languages, formalizing it as a set-valued function over grammatical templates within a Template-Based Modular Cognitive framework and offering a unified computational learning theory account.
This study applies computational linguistics and supervised machine learning to predict early-stage startup exits from textual descriptors alone, finding that founder narratives carry predictive signal and introducing a quantifiable Hyping Score.
The authors use an iterated learning model to study how meaning frequency affects the emergence of compositionality, expressivity, and stability in evolved languages. They find that high-frequency whole meanings escape grammatical pressure, but frequency applied to sub-meaning parts leads to transmission failure.
A two-dialect finite-state morphological analyzer for the Dungan language is presented, with a multi-genre evaluation measuring inflection, ambiguity, and lexical coverage.
A book examining the role of algebraic models versus deep learning in natural language acquisition, featuring perspectives from leading researchers in computational linguistics, psychology, and mathematical linguistics.
This paper introduces Semantic Field Theory (SFT), a computational model for lexical semantics that models meaning through semantic fields, contextual deformation, interaction terms, and energy minimization. It provides formal elements including Gaussian product closure, Möbius inversion for higher-order interactions, and stability conditions.
This study demonstrates that contextual semantic relevance, measuring how strongly an incoming word relates to its recent semantic context, reliably predicts fMRI BOLD responses during naturalistic speech comprehension across two datasets, whereas surprisal (local probabilistic expectation) does not. The findings support that slow hemodynamic responses are especially sensitive to contextual semantic integration rather than local prediction.
This paper models early language acquisition as a search on a graph-based mental lexicon using spreading activation and category exploration, outperforming a shortest path baseline in simulating normative word acquisition across four languages.
Emily Bender clarifies the original meaning of 'stochastic parrots' from her 2021 paper, debunking common misconceptions about LLMs and critiquing the term 'artificial intelligence' for overselling technology.
Svarna is an open-source web-based corpus workbench for Modern Greek, integrating multiple databases with over 507 million words and providing various linguistic analysis tools, released under MIT license.