Beyond "AI Language": The case for the idiolectal nature of LLM output

arXiv cs.CL Papers

Summary

This paper argues that LLM outputs exhibit model-specific linguistic signatures akin to human idiolects, analyzing 2024 and 2026 corpora to show both generational shifts and stable individual model profiles.

arXiv:2608.06589v1 Announce Type: new Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects. We analyse two datasets of LLM-generated texts on societal topics: a 2024 corpus of six models (Improta et al. 2024) and a newly generated 2026 corpus using the same prompts featuring six contemporary models. Our findings, utilising computational descriptors and stylometric principal component analysis reveal a generational shift between the style of the 2024 and 2026 cohorts, while demonstrating that each individual model maintains a unique linguistic profile. This multi-layered interplay is illustrated by contraction frequencies, which vary from over 1,200 to over 30,000 per million words within the same cohort of models (2026). Ultimately, we conclude that treating LLM output as idiolectal in nature provides a valuable framework with potential implications for research on variation and change, LLM-generated text detection, forensic linguistics and usage-based approaches to language.
Original Article
View Cached Full Text

Cached at: 08/10/26, 08:02 AM

# Beyond "AI Language": The case for the idiolectal nature of LLM output
Source: [https://arxiv.org/abs/2608.06589](https://arxiv.org/abs/2608.06589)
[View PDF](https://arxiv.org/pdf/2608.06589)

> Abstract:While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model\-specific linguistic signatures akin to human idiolects\. We analyse two datasets of LLM\-generated texts on societal topics: a 2024 corpus of six models \(Improta et al\. 2024\) and a newly generated 2026 corpus using the same prompts featuring six contemporary models\. Our findings, utilising computational descriptors and stylometric principal component analysis reveal a generational shift between the style of the 2024 and 2026 cohorts, while demonstrating that each individual model maintains a unique linguistic profile\. This multi\-layered interplay is illustrated by contraction frequencies, which vary from over 1,200 to over 30,000 per million words within the same cohort of models \(2026\)\. Ultimately, we conclude that treating LLM output as idiolectal in nature provides a valuable framework with potential implications for research on variation and change, LLM\-generated text detection, forensic linguistics and usage\-based approaches to language\.

## Submission history

From: Thomas Stephan Juzek \[[view email](https://arxiv.org/show-email/309f2afd/2608.06589)\] **\[v1\]**Thu, 6 Aug 2026 21:03:13 UTC \(1,061 KB\)

Similar Articles

Linguistic Monoculture in LLM-Assisted Language Use

arXiv cs.AI

This paper introduces a mathematical framework to study how reliance on shared LLMs for writing may reduce population-level linguistic diversity, analyzing fixed, recursive, and personalized interaction mechanisms and characterizing equilibria and convergence rates.

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

arXiv cs.CL

This paper introduces a register-aware linguistic evaluation framework to assess how human-like large language models (LLMs) are by comparing the distribution of 67 lexico-grammatical features between human and LLM-generated texts using Maximum Mean Discrepancy. Experiments across seven instruction-tuned open-source models and five registers show that no model perfectly matches human baselines, and closeness to human language varies by register rather than model size.

Content for Content’s Sake

Armin Ronacher

The author investigates how LLMs are influencing word usage in coding and everyday language, finding that words favored by LLMs show increased frequency in both coding sessions and Google Trends, raising concerns about humans adopting LLM writing styles.