Tag
This article analyzes BPE tokenization in Polish, highlighting its limits in inflectional languages and proposing grammatical form anchoring to improve language modeling.
This paper proposes a layered taxonomy for annotating grammatical errors in Chinese learner writing, combining computational and pedagogical perspectives, and evaluates it through coverage analysis and consistency studies with language models.
The article describes a tool that distinguishes AI-generated comments from human-written ones in code using linguistic features like character frequency, and discusses its methodology and implications.
SynFlow is an open-source toolkit for multidimensional diachronic semantic analysis, applying a shared workflow to linguistic representations for studying lexical semantic change across syntactic, morphological, and other dimensions.
This paper investigates whether text generated by multilingual large language models exhibits signs of translationese, and compares it to human-written translations to assess linguistic naturalness.
This paper evaluates Wiktionary as an ethically crowdsourced lexicon for English dialects, demonstrating its coverage compared to traditional dictionaries and its performance with geo-referenced social media data.
This paper analyzes statistical patterns in LLM-generated text using n-gram distributions, revealing stylistic deficiencies and showing that style and semantics are not separable.
This paper presents a multimodal emotion recognition module for proactive conversational agents, using facial recognition and linguistic analysis. A user study with 20 participants reveals a 'poker face' effect where visual cues are unreliable, while linguistic analysis proves more accurate; the study also shows agents can elicit emotions through conversational adaptation.
A preprint proposes a 33-feature quantitative linguistic framework that distinguishes professionally edited from self-published books and outperforms existing story-level evaluation metrics.
This paper systematically evaluates the applications of large language models in low-resource language research, analyzing opportunities and challenges across linguistic variation, historical documentation, cultural expressions, and literary analysis. The study emphasizes interdisciplinary collaboration and customized model development to preserve linguistic and cultural heritage while addressing issues of data accessibility, model adaptability, and cultural sensitivity.