When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era
Summary
This paper introduces AVShift, the first German benchmark for authorship verification under distribution shifts in genre, time, and AI-era. It evaluates feature-based, embedding-based, and LLM-based approaches, finding that temporal drift significantly impacts performance while no measurable AI-era shift is detected.
View Cached Full Text
Cached at: 08/19/26, 10:11 AM
# When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era
Source: [https://arxiv.org/html/2608.17979](https://arxiv.org/html/2608.17979)
Brisca BalthesChristoph LeiterYamen AjjourElena SchmidtSteffen EgerAffiliation:lotta\.kiefer@utn\.de steffen\.eger@utn\.de
###### Abstract
Authorship verification \(AV\) assumes that an author’s writing style remains sufficiently stable to distinguish it from that of other writers\. In practice, however, this assumption is challenged by distribution shifts caused by changes in genre, time, and AI\-assisted writing\. Existing AV benchmarks typically study these factors in isolation and focus predominantly on English, limiting our understanding of model robustness under realistic conditions\. We introduce AVShift, the first German benchmark for systematically evaluating AV under multiple distribution shifts\. AVShift comprises over 150,000 text pairs spanning three genres and 21 years, enabling controlled evaluation of cross\-genre, temporal, and AI\-era shifts within a unified framework\. We benchmark representative feature\-based, embedding\-based, and LLM\-based approaches\. Our experiments show that fine\-tuned LLMs generalize best across genres and benefit substantially from stylistically diverse training data\. We further demonstrate that temporal drift is one of the strongest factors affecting AV, with performance degrading significantly as the time gap between documents increases\. In contrast, we find no evidence of a measurable AI\-era distribution shift within AVShift\. Finally, our feature analysis reveals stylistic features that remain stable across genres, while their relative importance varies depending on the specific genre transition\. We release AVShift and our code for future research\.
## 1Introduction
Figure 1:Overview of the stylistic distribution shifts covered by AVShift: cross\-genre variation, temporal evolution, and AI\-era changes before and after the widespread adoption of generative AI\.Authorship analysis aims to identify the author of a text based on characteristic patterns of individual language use, commonly referred to as an*idiolect*\([9](https://arxiv.org/html/2608.17979#bib.bib20);[35](https://arxiv.org/html/2608.17979#bib.bib21)\)\. Numerous benchmark datasets spanning books, blogs, news, emails, reviews, social media, and darknet forums have established authorship analysis as an important tool in literary studies, plagiarism detection, and forensic linguistics\([51](https://arxiv.org/html/2608.17979#bib.bib35);[5](https://arxiv.org/html/2608.17979#bib.bib36);[30](https://arxiv.org/html/2608.17979#bib.bib22);[32](https://arxiv.org/html/2608.17979#bib.bib23);[38](https://arxiv.org/html/2608.17979#bib.bib8);[48](https://arxiv.org/html/2608.17979#bib.bib14)\)\.
This work focuses on authorship verification \(AV\), which determines whether two texts were written by the same author\. Unlike authorship attribution that relies on a predefined set of candidates, AV naturally generalizes to previously unseen authors without requiring retraining, making it particularly attractive for forensic investigations involving new suspects\.
The concept of linguistic individuality is sometimes described as alinguistic fingerprint\([27](https://arxiv.org/html/2608.17979#bib.bib25);[10](https://arxiv.org/html/2608.17979#bib.bib24)\)\. However, this analogy can be misleading, giving the wrong impression that writing style is an immutable concept like a biological fingerprint[9](https://arxiv.org/html/2608.17979#bib.bib20);[35](https://arxiv.org/html/2608.17979#bib.bib21)\. Instead, linguistic style is influenced by numerous contextual factors, including the intended audience, topic, communicative situation, medium, and evolves over time\.
We therefore view AV as a learning problem that is inherently exposed to distribution shifts\. In practice, both training and test data, as well as the paired texts themselves, may differ in genre, time, or broader language use\. Existing benchmarks typically investigate these challenges in isolation and focus predominantly on English\. However, AV relies on language\-specific stylistic cues, making it unclear whether findings obtained on English generalize to other languages\. This lack of diverse benchmarks limits our understanding of AV robustness under realistic conditions and across languages\.
To address this gap, we introduceAVShift, the first German benchmark for systematically evaluating AV under three key distribution shifts: \(1\) cross\-genre shift across forum posts, reviews, and fanfiction, \(2\) temporal shifts spanning more than two decades of writing, and \(3\) AI\-era shifts before and after the widespread adoption of generative AI \(genAI\) writing assistants\. Figure[1](https://arxiv.org/html/2608.17979#S1.F1)provides an overview of these evaluation settings\.
We make the following contributions:
- •We introduce AVShift, the first German benchmark comprising more than 150,000 text pairs for evaluating AV under cross\-genre, temporal, and AI\-era distribution shifts\.
- •We compare feature\-based, embedding\-based, and LLM\-based approaches across all benchmark settings\.
- •We show that verification performance varies substantially across genres and that LLMs trained on mixed\-domain data exhibit strong cross\-genre performance reaching an F1 of up to 0\.77\.
- •We introduce a feature stability score, revealing that the robustness of stylistic features strongly depends on the genre pair\.
- •We show that temporal shifts substantially degrade verification performance by up to 0\.21 F1, whereas no significant AI\-era degradation is observed in our benchmark\.
## 2Related Work
### Authorship Analysis Methods
AV methods can broadly be categorized into feature\-based, embedding\-based, and LLM\-based approaches\.
Feature\-based methods represent documents using handcrafted stylistic features, which are compared using similarity measures or statistical or neural classifiers\. While these approaches remain attractive due to their competitive performance, efficiency and interpretability\([49](https://arxiv.org/html/2608.17979#bib.bib37);[16](https://arxiv.org/html/2608.17979#bib.bib38)\), they have been shown to underperform modern neural methods[58](https://arxiv.org/html/2608.17979#bib.bib39)\.
Embedding\-based approaches learn dense stylistic representations directly from text, either through task\-specific neural architectures\([17](https://arxiv.org/html/2608.17979#bib.bib40);[40](https://arxiv.org/html/2608.17979#bib.bib41);[4](https://arxiv.org/html/2608.17979#bib.bib42)\)or extracted from pretrained transformer models\([45](https://arxiv.org/html/2608.17979#bib.bib43);[31](https://arxiv.org/html/2608.17979#bib.bib7);[12](https://arxiv.org/html/2608.17979#bib.bib44)\)\.
More recently, LLMs have emerged as a promising approach to authorship analysis by jointly learning stylistic representations and the verification task\. Although zero\- and few\-shot prompting with closed\-source models such as GPT\-3\.5\([36](https://arxiv.org/html/2608.17979#bib.bib48)\)and GPT\-4\([37](https://arxiv.org/html/2608.17979#bib.bib47)\)achieves competitive performance, their reliance on online APIs limits deployment in forensic and privacy\-sensitive settings\([23](https://arxiv.org/html/2608.17979#bib.bib45);[42](https://arxiv.org/html/2608.17979#bib.bib46)\)\. More recent work shows that fine\-tuned open\-source LLMs outperform both prompting\-based approaches and previous AV methods while remaining practical to deploy\([22](https://arxiv.org/html/2608.17979#bib.bib59);[42](https://arxiv.org/html/2608.17979#bib.bib46);[25](https://arxiv.org/html/2608.17979#bib.bib5)\)\.
### Out\-of\-Distribution Evaluation
Cross\-domain AVencompasses evaluation settings in which data are drawn from different distributions\. This may refer either to domain transfer between training and test data\([45](https://arxiv.org/html/2608.17979#bib.bib43)\)or to the more challenging setting, where the two compared texts originate from different topics, genres, or platforms\([38](https://arxiv.org/html/2608.17979#bib.bib8);[31](https://arxiv.org/html/2608.17979#bib.bib7)\)\. While cross\-topic AV has received considerable attention, with numerous methods proposed to reduce topic bias\([50](https://arxiv.org/html/2608.17979#bib.bib10);[18](https://arxiv.org/html/2608.17979#bib.bib3);[21](https://arxiv.org/html/2608.17979#bib.bib15)\), cross\-platform and cross\-genre settings remain comparatively underexplored\. Existing work consistently reports substantial performance degradation under these shifts\([47](https://arxiv.org/html/2608.17979#bib.bib13);[48](https://arxiv.org/html/2608.17979#bib.bib14);[3](https://arxiv.org/html/2608.17979#bib.bib1);[24](https://arxiv.org/html/2608.17979#bib.bib4);[31](https://arxiv.org/html/2608.17979#bib.bib7)\)\. Recent work has explored domain\-adaptive style representations\([59](https://arxiv.org/html/2608.17979#bib.bib16)\)and training with hard negative and cross\-domain positive pairs\([52](https://arxiv.org/html/2608.17979#bib.bib11)\), improving robustness without eliminating the performance degradation caused by certain domain shifts\.
Temporal shiftrefers to changes in an author’s writing style over time, raising the question of how well systems can recognize authors across substantial time gaps\. Despite its practical relevance, temporal effects in AV have received limited attention\. An existing study on six French novelists shows that performance can degrade as temporal distance increases\([6](https://arxiv.org/html/2608.17979#bib.bib2)\)\. However, the small number of authors limits generalizability, and it remains unclear whether findings from literary texts transfer to other domains, where stylistic choices may be less deliberate\. Approaches addressing temporal variation are similarly scarce\.[56](https://arxiv.org/html/2608.17979#bib.bib12)model changes in authors’ interests over time through topic drift, while[2](https://arxiv.org/html/2608.17979#bib.bib17)estimate temporal changes in lexical style and show improvements for authorship attribution on tweets and emails\.
AI\-era shiftshift refers to stylistic changes introduced by genAI writing assistance\. Initial work suggests that AI\-assisted writing can increase stylistic similarity between authors and thereby affect verification performance, particularly by increasing false positives\([44](https://arxiv.org/html/2608.17979#bib.bib9)\)\. Studying AI shift remains challenging due to the diverse forms of human\-AI collaboration, including text generation, revision, and multi\-turn interaction[34](https://arxiv.org/html/2608.17979#bib.bib18);[29](https://arxiv.org/html/2608.17979#bib.bib19)\. Many benchmarks disregard the possible influence arising from documents sampled from the AI\-era\.
### Multilingual Evaluation
Despite the language\-dependent nature of stylistic features, authorship analysis remains heavily focused on English\. Recent work has introduced multilingual embedding models capable of transferring across languages\([41](https://arxiv.org/html/2608.17979#bib.bib49);[26](https://arxiv.org/html/2608.17979#bib.bib6)\), and several multilingual benchmarks\([24](https://arxiv.org/html/2608.17979#bib.bib4);[33](https://arxiv.org/html/2608.17979#bib.bib50);[19](https://arxiv.org/html/2608.17979#bib.bib51)\)\. German\-specific benchmarks have also become available\([5](https://arxiv.org/html/2608.17979#bib.bib36);[25](https://arxiv.org/html/2608.17979#bib.bib5)\), but these primarily focus on in\-domain or cross\-topic evaluation\. To our knowledge, no benchmark systematically combines different distribution shifts in a multilingual setting\.
Overall, prior work has investigated individual distribution shifts largely in isolation and predominantly on English datasets\. AVShift addresses this gap by providing the first German benchmark that systematically evaluates AV under cross\-genre, temporal, and AI\-era distribution shifts within a unified evaluation framework\.
## 3Data Curation
The goal of the data collection process is to construct a German corpus for systematically evaluating AV under distribution shifts\. After evaluating several candidate sources, we selected[www\.fanfiktion\.de](https://www.fanfiktion.de/), which offers three complementary writing environments within a single platform: fanfiction stories, reviews, and forum posts\. This allows us to study distribution shifts while minimizing platform\-specific confounds\.
We use the forum section, specificallyAllgemeines Geplauder\(“General Chit Chat”\), as the entry point of our scraping pipeline\. Using BeautifulSoup\([43](https://arxiv.org/html/2608.17979#bib.bib26)\), we extract author profile links from forum threads and subsequently collect all available forum posts, reviews, and fanfiction stories for each author\. This produces three aligned corpora containing texts from the same authors across two to three genres\.
To capture temporal variation, we traverse the forum archive back to its earliest available entries from 2004\. The resulting corpus spans 21 years \(2004–2025\), enabling both long\-term temporal analyses and comparisons between texts written before and after the public release of ChatGPT\([36](https://arxiv.org/html/2608.17979#bib.bib48)\)in November 2022\.
### Preprocessing
We remove HTML tags, normalize whitespace, and replace all URLs with a shared token\. German message closings are removed together with subsequent content, and remaining self\-identifying information is removed through fuzzy username matching\.
Texts shorter than 50 words are discarded to ensure sufficient stylistic content, while texts longer than 3,000 words are truncated to reduce computational cost\. Unlike previous work, we deliberately retain topical vocabulary and instead control for topic\-bias effects during pair construction\. Since the website exclusively hosts German\-language content, no language filtering is required\.
### AVShift Benchmark
AVShift consists of three sub\-benchmarks that evaluate complementary distribution shifts: GenreShift, TimeShift, and AIShift\. We provide an example in Appendix[A](https://arxiv.org/html/2608.17979#A1)\.
GenreShiftevaluates cross\-genre generalization using seven AV datasets: three in\-domain datasets \(Forum, Review, Story\), three cross\-genre datasets \(Review\-Forum, Story\-Forum, Review–Story\), and one Mixed dataset obtained by uniformly sampling from the remaining six datasets\.
Authors are split uniformly into 80% training, 10% validation, and 10% test partitions across all datasets, ensuring that no author appears in multiple splits, allowing cross\-dataset comparison\.
Each story sample consists of a single chapter, and each review and forum sample consists of one respective post\. To reduce topic leakage, positive pairs are sampled under genre\-specific constraints: story pairs originate from different fanfiction works, review pairs review different stories, and forum pairs come from different discussion threads\. Positive pairs contain texts by the same author, while negative pairs contain texts by different authors\. All datasets are class\-balanced, and we sample at most five positive and five negative pairs per author to prevent highly active users from dominating the benchmark\. Table[1](https://arxiv.org/html/2608.17979#S3.T1)summarizes the resulting GenreShift datasets\.
DatasetSample NumberUnique PostsUnique UsersMean SampleLen \(in Words\)Forum15,27216,5162,104173Review22,22823,2422,466184Story40,32059,0074,6271,358Review\-Forum18,49018,2741,849179Review\-Story26,09035,9222,609786Story\-Forum32,58037,2333,258787Mixed30,00043,1564,806576
Table 1:Statistical Comparison: All GenreShift datasetsTimeShiftevaluates robustness to temporal shift\. We construct ten dataset slices covering temporal gaps of 12 months each, ranging from 0\-12 months to 108\-120 months\. All text pairs are sampled such that the publication dates of the two texts in each pair fall within the corresponding temporal interval of a given slice \(e\.g\., in the 12\-24 month slice, the publication dates of Text A and Text B differ by at least 12 and less than 24 months in both positive and negative pairs\)\. This procedure is applied independently to the Forum, Review, and Story corpora\. To ensure comparability, all datasets are downsampled to the size of the smallest subset within each genre \(Forum: 670 pairs, Review: 446 pairs, Story: 4,466 pairs\)\.
AIShiftevaluates the impact of the widespread adoption of genAI on AV\. The corpus is partitioned into four periods:Early \(2004–2010\),Mid \(2011–2016\),Pre\-AI \(2017–2022\), andAI \(2023–2025\)\. Following the same pair construction procedure as GenreShift, we construct separate datasets for each genre and period, enabling controlled comparisons of AV before and after the emergence of AI\-assisted writing\.
## 4Experimental Setup
To evaluate AV under different distribution shifts, we compare three representative approaches covering the dominant AV paradigms: feature\-based, embedding\-based, and LLM\-based methods\.
### AV Models
As a representative feature\-based approach, we train an XGBoost classifier\([8](https://arxiv.org/html/2608.17979#bib.bib34)\)on handcrafted stylometric features\. We extract a comprehensive set of more than 4,000 distinct stylistic features \(see Appendix[B](https://arxiv.org/html/2608.17979#A2)for details\)\. Feature vectors are extracted independently for both texts and combined using their element\-wise difference before classification\.
As an embedding\-based method, we use the Multilingual Style Representation \(MSR\) model proposed by[26](https://arxiv.org/html/2608.17979#bib.bib6)\. MSR learns language\-agnostic stylistic embeddings from 36 languages and 13 domains and achieves strong performance, even on unseen languages such as German\.
For the LLM\-based approach, we follow the framework of[25](https://arxiv.org/html/2608.17979#bib.bib5), which fine\-tunes instruction\-tuned LLMs with LoRA\([20](https://arxiv.org/html/2608.17979#bib.bib52)\)to answer the binary question of whether two texts were written by the same author\. We replace their best\-performing model, Gemma\-3\-12B\-it\([14](https://arxiv.org/html/2608.17979#bib.bib29)\), with the more recent Gemma\-4\-31B\-it\([13](https://arxiv.org/html/2608.17979#bib.bib30)\), while keeping the training procedure unchanged\.
### Evaluation Metrics
We report macro F1\-score as the primary evaluation metric throughout the paper, as it is well suited for our binary, balanced classification setting\. We additionally report Accuracy scores in the Appendix\.
### GenreShift Evaluation
We train and evaluate all three models on each of the seven AVShift datasets \(see Appendix[C](https://arxiv.org/html/2608.17979#A3)for the training setup\)\. Note that for MSR, the embeddings remain unchanged, and only the decision threshold is tuned on AVShift using Youden’sJJstatistic\([57](https://arxiv.org/html/2608.17979#bib.bib53)\)\. This evaluation setup assesses both generalization to unseen domains and within\-sample cross\-genre AV performance\. Statistical significance of performance differences between models is assessed using paired bootstrap testing with 10,000 resamples\([11](https://arxiv.org/html/2608.17979#bib.bib54)\)\.
Because AV performance is strongly influenced by text length and can be affected by training set size\([10](https://arxiv.org/html/2608.17979#bib.bib24);[25](https://arxiv.org/html/2608.17979#bib.bib5)\), we additionally construct standardized GenreShift datasets\. Every document is normalized to 500 words, extending shorter texts by concatenating additional texts from the same author and truncating longer texts\. We further downsample all datasets to the size of the smallest genre split \(3,640 samples\), allowing us to isolate the effect of genre independently of text length and dataset size\. The standardized GenreShift datasets are marked by an additional500\.
To further investigate stylistic variation across genres, we exploit the interpretability of the handcrafted feature representation used in the XGB approach\. We first concatenate all texts written by each author into a single document and balance the three resulting genre corpora by truncating them to the same total token count\.
A shared feature vectorizer is then fitted on the combined author corpora of stories, reviews, and forum posts, enabling direct comparison of feature distributions across genres\. This analysis allows us to identify stylistic features that remain stable across genres as well as those that are strongly genre\-dependent\.
### Time Shift Evaluation
For each genre, we evaluate the best\-performing in\-domain model on the TimeShift benchmark\. Performance is measured on ten datasets with temporal gaps ranging from 0–12 months to 108–120 months\. We report Pearson\([39](https://arxiv.org/html/2608.17979#bib.bib32)\)and Spearman\([46](https://arxiv.org/html/2608.17979#bib.bib31)\)correlation coefficients between temporal distance and verification performance to quantify the effect of temporal drift\.
### AIShift Evaluation
We evaluate AIShift using a leave\-one\-era\-out protocol\. For each experiment, one temporal period is held out for testing while the remaining three periods are combined for training\. For each split, we uniformly sample the same number of pairs \(train: 4,122, test: 1,374\) to ensure comparable dataset sizes across all experiments\. This setup enables a controlled evaluation of AV performance before and after the emergence of AI\-assisted writing across different genres\.
### Crossnews
To assess the generalizability of our findings beyond German, we additionally evaluate our methods on the English CrossNews benchmark introduced by[31](https://arxiv.org/html/2608.17979#bib.bib7)\. CrossNews links news articles and tweets written by the same author, yielding two in\-domain datasets \(Article and Tweet\) and one cross\-genre dataset \(Article–Tweet\)\. We compare our methods against the two best\-performing approaches reported by the authors: \(1\) a prompting\-based method using LLaMA\-3\-70B[15](https://arxiv.org/html/2608.17979#bib.bib27), and \(2\) SELMA, which performs AV using embedding distances obtained with e5\-mistral\-7b\-instruct\([54](https://arxiv.org/html/2608.17979#bib.bib28)\)\.
## 5Results
Figure 2:F1 scores of all models trained and evaluated on all GenreShift train and test splits\. Model names are shown on the y\-axis and test dataset names on the x\- axis\. Each model name is followed by its training or calibration dataset\. The best score for each test set is shown in bold\. Statistically significant superiority over all other models is indicated by an asterisk \(\*; p < 0\.05\)\.### How robust are models across genres?
Our results on the GenreShift benchmark are summarized in Figure[2](https://arxiv.org/html/2608.17979#S5.F2)\(Appendix[D](https://arxiv.org/html/2608.17979#A4)for full results\)\. Models are on the y\-axis followed by the name of the training set and test datasets on the x\-axis with in\-domain datasets referred by the genre name \(e\.g\. Review\) and cross\-genre datasets by both genres the text pairs are drawn from \(e\.g\. Review\-Forum where a pair consists of one text from Review and one text from Forum\)\. Across all seven datasets Gemma consistently outperforms both the feature\-based XGB classifier and the embedding\-based MSR model, achieving the highest F1 score on every test set \(p<0\.05p<0\.05\)\. While models trained on the Review and Story datasets perform best on their respective in\-domain tasks, the Mixed Gemma model achieves the strongest performance on all remaining datasets\. This demonstrates that exposing LLMs to stylistically diverse training data substantially improves robustness to distribution shifts\.
Among the three in\-domain datasets, Review consistently reveals to be the easiest genre for AV, followed by Story and Forum\. The best\-performing in\-domain models achieve F1 scores of 0\.89, 0\.80, and 0\.78, respectively\. This suggests that reviews contain the strongest and most consistent authorial signal, whereas forum posts represent the most challenging writing style\.
Both Gemma and XGB degrade substantially under cross\-genre transfer\. For example, Gemma achieves an F1 score of 0\.89 when trained and tested on reviews but loses approximately 0\.2 F1 when trained on stories or forum posts\. The Forum dataset is an exception, where training on reviews or stories yields better performance than training on forum data, suggesting that forum posts provide weaker supervision for learning robust stylistic representations\. In contrast, MSR is less sensitive to the calibration genre, likely because its embedding model remains fixed and only the verification threshold is adapted to AVShift\.
The three cross\-genre datasets also differ substantially in difficulty\. Review\-Forum is consistently the easiest cross\-genre setting, whereas Review–Story and Story\-Forum are considerably more challenging\. Surprisingly, the Mixed Gemma model outperforms models trained directly on the corresponding cross\-genre datasets in every setting\. This indicates that exposing the model to a broad range of stylistic variation is more beneficial than specializing on a single genre transition\. Notably, its performance on Review\-Forum approaches the in\-domain performance obtained on the Forum dataset \(0\.77 F1\), demonstrating that robust cross\-genre AV is achievable when sufficient stylistic diversity is observed during training\. While previous work consistently reported substantial performance degradation under cross\-genre evaluation\([31](https://arxiv.org/html/2608.17979#bib.bib7);[24](https://arxiv.org/html/2608.17979#bib.bib4);[48](https://arxiv.org/html/2608.17979#bib.bib14)\), our results suggest that much of this degradation can be mitigated through sufficiently diverse training data\.
To determine whether these differences are genuinely caused by genre rather than confounding factors such as text length or training set size, we repeat the experiments on our standardized AVShift datasets in which both factors are controlled\. The relative difficulty of the three genres remains unchanged: Review \(0\.84\) continues to outperform Story and Forum \(both 0\.74\) \(see Appendix[E](https://arxiv.org/html/2608.17979#A5)\), confirming that the observed ranking is intrinsic to the writing genres rather than an artifact of dataset construction\. In contrast, the ranking of the models changes considerably\. Under these controlled conditions, Gemma is no longer consistently superior\. MSR achieves the highest performance on the standardized Review and Story datasets, while XGB performs best on Forum, although the best\-performing Gemma models follow closely and do not differ significantly \(p≥0\.05p\\geq 0\.05\)\. At the same time, cross\-genre performance deteriorates for all models, which we attribute primarily to the substantially smaller training sets resulting from the controlled sampling procedure\. Together, these findings indicate that genre itself is the dominant source of difficulty in AVShift, while the relative performance of different AV approaches depends strongly on the characteristics of the training data\. In particular, XGB and MSR benefit from the standardized setting and appear less sensitive to the reduced training size, whereas Gemma benefits more from the larger and stylistically more diverse training data available in the original benchmark\.
### Do our results generalize to English?
To assess whether the trends observed on AVShift generalize beyond German, we evaluate all three approaches on the English CrossNews benchmark\. Table[2](https://arxiv.org/html/2608.17979#S5.T2)compares the best\-performing models reported by[31](https://arxiv.org/html/2608.17979#bib.bib7)with the strongest model results we achieved on their benchmark\. Overall, our results closely mirror the findings on AVShift\. Gemma achieves the best performance on the Tweet–Tweet and Article–Tweet datasets, improving upon the previously reported state of the art by 0\.11 and 0\.07 F1, respectively\. On the Article–Article dataset, MSR achieves the highest performance, slightly outperforming SELMA\. These results demonstrate that the strong cross\-genre generalization of Gemma is not limited to German but also transfers to English\.
Overall, CrossNews remains consistently easier than AVShift, with substantially higher F1 scores across datasets\. While this may partly reflect the predominantly English pretraining of models such as Gemma, the benchmarks differ in several other aspects, preventing attribution of the performance gap to language alone\.
ModelArticle\-ArticleTweet\-TweetArticle\-Tweet[31](https://arxiv.org/html/2608.17979#bib.bib7)LLaMA Prompting0\.77±0\.089\\pm 0\.0890\.79±0\.048\\pm 0\.0480\.40±0\.064\\pm 0\.064[31](https://arxiv.org/html/2608.17979#bib.bib7)SELMA0\.86±0\.018\\pm 0\.0180\.75±0\.020\\pm 0\.0200\.80±0\.023\\pm 0\.023Gemma\-AT0\.830\.880\.87Gemma\-TT0\.830\.900\.86MSR\-AA0\.880\.840\.69
Table 2:Results on the Crossnews benchmark\. The two models scoring best for the three benchmark subsets as reported by[31](https://arxiv.org/html/2608.17979#bib.bib7)are shown alongside the best\-scoring models from our analysis\.
### Do Stylistic Features Survive Genre Shifts?
To better understand why some genre transitions are more challenging than others, we analyze the shared handcrafted feature space described in Section[4](https://arxiv.org/html/2608.17979#S4)\. Figure[3](https://arxiv.org/html/2608.17979#S5.F3)shows a t\-SNE\([7](https://arxiv.org/html/2608.17979#bib.bib33)\)projection of the resulting author representations\. Rather than clustering primarily by author, the vectors are largely separated by genre, indicating that genre exerts a strong influence on the handcrafted feature representation even for texts written by the same individual\.
To quantify how well individual features preserve authorial style across genres, we compute a cross\-genre feature stability score
Stability\(f\)=1−σwithin\(f\)σbetween\(f\)\+ε\\text\{Stability\}\(f\)=1\-\\frac\{\\sigma\_\{\\mathrm\{within\}\}\(f\)\}\{\\sigma\_\{\\mathrm\{between\}\}\(f\)\+\\varepsilon\}\(1\)which compares within\-author variation across genres to between\-author variation\. High stability indicates that a feature remains consistent for the same author while discriminating between different authors and representing robust indicators of authorial style\.
Across all three genres, stability scores range from−1\.70\-1\.70to 0\.90, with a mean of 0\.22 and a median of 0\.25 \(Table[3](https://arxiv.org/html/2608.17979#S5.T3)\)\. Overall, 80% of the handcrafted features exhibit positive stability, indicating that most stylistic features remain relatively consistent across genres despite the strong genre separation observed in the t\-SNE projection\. This suggests that successful cross\-genre AV remains feasible\. We provide the thirty most and least stable features in Appendix[G](https://arxiv.org/html/2608.17979#A7)\.
Feature stability further reflects differences in cross\-genre verification difficulty\. Review\-Forum exhibits the highest average stability \(0\.31 and 85% positive features\), consistent with the results showing highest classification performance\. Even though Story\-Forum ranks second in performance it shows the lowest stability scores in our analysis \(−0\.06\-0\.06; 48% positive features\)\. Thus, feature stability does not perfectly predict verification performance, but the overall trend suggests that it is a useful indicator of how well authorial style is preserved across genres and, consequently, of the expected difficulty of cross\-genre AV\.
Finally, we investigate whether the same features remain stable across different genre transitions\. We look at the overlap among the 100 most stable and 100 least stable features and find that overlap is generally low, ranging from 2% to 43%, indicating that different genre transitions affect different subsets of stylistic features\. Nevertheless, the overall feature rankings remain moderately correlated \(Spearman’sρ=0\.42−0\.72\\rho=0\.42\-0\.72\), suggesting that while genre shifts change which features are most informative, the broader ordering of feature stability is largely preserved \(see Appendix[H](https://arxiv.org/html/2608.17979#A8)\)\. Taken together, these findings indicate that feature stability should be analyzed separately for each genre transition rather than assuming a universal set of robust stylistic features\.
Figure 3:t\-SNE visualization of feature author vectors from different genres\.DatasetMean StabilityMedian StabilityPercentage StableFeaturesOverall0\.22±0\.28\\pm 0\.280\.250\.2580%Review\-Forum0\.31±0\.27\\pm 0\.270\.340\.3485%Story\-Forum0\.19±0\.37\\pm 0\.370\.250\.2575%Review\-Story−0\.06\-0\.06±0\.38\\pm 0\.38−0\.01\-0\.0148%
Table 3:Mean and median stability feature scores alongside the percentage of features with positive score for the whole dataset next to each genre\-pair individually\.
### How does authorial style change over time?
Figure[4](https://arxiv.org/html/2608.17979#S5.F4)shows the performance of the best\-performing in\-domain model for each genre on the TimeShift benchmark\. Across all genres, AV performance decreases steadily as the temporal gap between two texts increases\. Pearson and Spearman correlation analyses reveal a strong and statistically significant negative relationship between temporal distance and F1\-score for all genres \(p<0\.05p<0\.05\)\.
The magnitude of this degradation differs across genres\. Review shows the largest decline, with the F1\-score dropping from 0\.90 for text pairs separated by 0–12 months to 0\.69 after 9–10 years\. In contrast, Forum shows the smallest decrease \(0\.76 to 0\.68\), although this may partly reflect its lower initial performance, leaving less room for degradation\. Despite these differences, the relative difficulty of the three genres remains largely unchanged \(Review \> Story \> Forum\) across temporal intervals, suggesting that the higher performance on Review is unlikely to be attributable to potential differences in the time spans across genres\.
These findings demonstrate that authorial style is not static but evolves continuously over time, substantially reducing verification performance even for state\-of\-the\-art models\. Notably, a large decline occurs already after the first year, showing that temporal drift emerges very early\. Since all three genres exhibit the same overall trend, temporal variation should be considered an important factor when constructing and evaluating AV benchmarks\. For high\-stakes real\-world applications in particular, our results suggest that verification is only reliable when comparing documents written within relatively short time windows without further model modifications\.
Figure 4:Correlation of increasing time spans \(x\-axis\) and F1\-score \(y\-axis\) alongside Pearson \(r\) and Spearman \(ρ\\rho\) correlation coefficients
### Has the emergence of genAI changed AV?
We find no evidence that the emergence of genAI has introduced a systematic distribution shift for AV \(see Appendix[F](https://arxiv.org/html/2608.17979#A6)\)\. While performance differs significantly between individual hold\-out eras, these differences do not follow a consistent chronological pattern\. In particular, AI\-era texts are not systematically more difficult to verify than earlier texts\. For example, the Story dataset exhibits its largest performance drop when the Middle era is held out, whereas the Review dataset achieves its highest performance on the AI era\. Overall, the observed variation is more likely explained by dataset\-specific factors, such as text length, topic, or temporal sampling, than by the widespread adoption of genAI\.
This finding should, however, be interpreted with caution, as our corpus was not annotated for AI\-assisted writing\. Consequently, we cannot determine the prevalence of genAI within AVShift\. Future work should validate these findings using datasets with controlled levels of AI use, enabling a more direct assessment of its impact on AV\.
## 6Conclusion
In this work, we introduced AVShift, the first German benchmark for systematically evaluating AV under realistic distribution shifts\. AVShift unifies cross\-genre, temporal, and AI\-era evaluation, enabling a comprehensive assessment of robustness beyond conventional in\-domain settings\. We used it to compare feature\-based, embedding\-based, and LLM\-based approaches under challenging real\-world conditions\.
Our experiments reveal three main findings\. First, although cross\-genre AV remains challenging, fine\-tuned LLMs perform particularly well, benefiting from stylistically diverse training data and even matching in\-domain performance in one setting\. Second, temporal drift is one of the strongest factors affecting AV, with performance consistently declining as the time gap between documents increases\. Third, we find no evidence that the widespread adoption of genAI has introduced a measurable distribution shift in AVShift, although this should be revisited using datasets with controlled AI\-assisted writing\.
Beyond benchmarking, our analyses provide new insights into authorial style\. Although handcrafted features are strongly influenced by genre, many stylistic features remain stable enough to support reliable cross\-genre AV\. Together with the superior performance of models trained on stylistically diverse data, these findings suggest that robustness is achieved not by eliminating stylistic variation, but by learning representations that capture author\-specific characteristics despite changes in writing context\.
We hope AVShift will serve as a valuable resource for robust AV research, particularly for non\-English languages where benchmark datasets remain scarce\. More broadly, our results suggest that future progress should be measured not only by in\-domain performance but also by robustness to realistic distribution shifts, providing a more reliable assessment for forensic and other real\-world applications\.
## Limitations
This work has several limitations that should be taken into consideration in the interpretation of the results\. First, our evaluation covers a representative but limited selection of AV approaches\. While we select three models representing feature\-based, embedding\-based, and LLM\-based methods, other approaches may yield different results\. Moreover, within each model category, design choices such as LLM architecture, fine\-tuning strategy, threshold calibration, or feature selection may influence absolute performance and model rankings\. A broader evaluation across additional methods and configurations remains an important direction for future work\.
Second, AVShift is constructed from a single German platform, which enables controlled comparisons across genres while reducing platform\-specific confounding factors\. However, this also limits the diversity of writing environments represented in the benchmark\. Future extensions incorporating additional platforms and languages could further assess the generalizability of the observed findings\.
Third, while AVShift introduces realistic distribution shifts, measuring some forms of shift remains challenging\. In particular, AI\-era shift depends on the extent and type of human\-AI interaction, which we did not try to quantify in this work\. Future work could investigate controlled settings with known levels of AI assistance\.
Finally, this work focuses on analyzing robustness under distribution shifts rather than developing methods to improve style shift robustness\. We provide AVShift and our analyses as a foundation for future research on adaptive training strategies, robust representations, and methods specifically designed to address distribution shifts in AV\.
## Ethical Considerations
We aim to minimize the environmental impact of our experiments by restricting GPU usage to the resources required for model training and evaluation\.
AV has potential societal benefits in applications such as forensic investigations and plagiarism detection\. However, we acknowledge that these technologies may also be misused, for example to deanonymize individuals or undermine legitimate privacy protections\. Furthermore, AV systems are inherently imperfect, and their predictions should not be interpreted as definitive evidence in real\-world applications\. In particular, we cannot fully exclude the influence of demographic, social, or other contextual factors that may introduce biases against specific groups\.
We provide details on model and dataset licenses in Appendix[I](https://arxiv.org/html/2608.17979#A9)\. The released dataset will use pseudonymized usernames and will be made available exclusively for research purposes\. The underlying platform provides publicly accessible content without requiring user authentication; nevertheless, we recognize that publicly available data may still carry privacy considerations and encourage responsible use of the resource\.
## Acknowledgements
We gratefully acknowledge the support that made this work possible\. This work was supported by the German Federal Ministry of Research, Technology and Space \(BMFTR\) through the ALiAS research project \(grants 13N17272 and 13N17273\) within the security research program, and by the German Research Foundation \(DFG\) under the Heisenberg Grant EG 375/5\-1\.
## References
- Anselet al\.\(2024\)J\. Ansel, E\. Yang, H\. He, N\. Gimelshein, A\. Jain, M\. Voznesensky, B\. Bao, P\. Bell, D\. Berard, E\. Burovski, G\. Chauhan, A\. Chourdia, W\. Constable, A\. Desmaison, Z\. DeVito, E\. Ellison, W\. Feng, J\. Gong, M\. Gschwind, B\. Hirsh, S\. Huang, K\. Kalambarkar, L\. Kirsch, M\. Lazos, M\. Lezcano, Y\. Liang, J\. Liang, Y\. Lu, C\. Luk, B\. Maher, Y\. Pan, C\. Puhrsch, M\. Reso, M\. Saroufim, M\. Y\. Siraichi, H\. Suk, M\. Suo, P\. Tillet, E\. Wang, X\. Wang, W\. Wen, S\. Zhang, X\. Zhao, K\. Zhou, R\. Zou, A\. Mathews, G\. Chanan, P\. Wu, and S\. ChintalaPyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation\.In29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 \(ASPLOS ’24\),External Links:[Document](https://dx.doi.org/10.1145/3620665.3640366),[Link](https://docs.pytorch.org/assets/pytorch2-2.pdf)Cited by:[Appendix C](https://arxiv.org/html/2608.17979#A3.p1.1)\.
- Azarbonyadet al\.\(2015\)H\. Azarbonyad, M\. Dehghani, M\. Marx, and J\. KampsTime\-aware authorship attribution for short text streams\.SIGIR ’15,New York, NY, USA,pp\. 727–730\.External Links:ISBN 9781450336215,[Link](https://doi.org/10.1145/2766462.2767799),[Document](https://dx.doi.org/10.1145/2766462.2767799)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p2.1)\.
- Barlas and Stamatatos \(2020\)G\. Barlas and E\. StamatatosCross\-domain authorship attribution using pre\-trained language models\.InArtificial Intelligence Applications and Innovations,I\. Maglogiannis, L\. Iliadis, and E\. Pimenidis \(Eds\.\),Cham,pp\. 255–266\.External Links:ISBN 978\-3\-030\-49161\-1Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- Boenninghoffet al\.\(2019\)B\. Boenninghoff, S\. Hessler, D\. Kolossa, and R\. M\. NickelExplainable authorship verification in social media via attention\-based similarity learning\.In2019 IEEE International Conference on Big Data \(Big Data\),pp\. 36–45\.External Links:[Document](https://dx.doi.org/10.1109/BigData47090.2019.9005650)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p3.1)\.
- Boenninghoffet al\.\(2024\)B\. Boenninghoff, H\. Hosseini, R\. M\. Nickel, and D\. KolossaWho wrote when? author diarization in social media discussions\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 15721–15734\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.922/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.922)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px3.p1.1)\.
- Cafieroet al\.\(2025\)F\. Cafiero, L\. Ing, S\. Gabay, and T\. Clérice“I am too old for this style\!” a stylometric benchmark of age effect on authorship attribution\.InComputational Humanities Research 2025,T\. Arnold, M\. Fantoli, and R\. Ros \(Eds\.\),pp\. 1248–1260\.External Links:[Document](https://dx.doi.org/10.63744/By09x5ZX3yWX)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p2.1)\.
- Cai and Ma \(2022\)T\. T\. Cai and R\. MaTheoretical foundations of t\-sne for visualizing high\-dimensional clustered data\.J\. Mach\. Learn\. Res\.23\(1\)\.External Links:ISSN 1532\-4435Cited by:[§5](https://arxiv.org/html/2608.17979#S5.SS0.SSS0.Px3.p1.1)\.
- Chen and Guestrin \(2016\)T\. Chen and C\. GuestrinXGBoost: a scalable tree boosting system\.KDD ’16,New York, NY, USA,pp\. 785–794\.External Links:ISBN 9781450342322,[Link](https://doi.org/10.1145/2939672.2939785),[Document](https://dx.doi.org/10.1145/2939672.2939785)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px1.p1.1)\.
- Coulthard \(2004\)M\. CoulthardAuthor identification, idiolect, and linguistic uniqueness\.Applied Linguistics25\(4\),pp\. 431–447\.External Links:ISSN 0142\-6001,[Document](https://dx.doi.org/10.1093/applin/25.4.431),[Link](https://doi.org/10.1093/applin/25.4.431),https://academic\.oup\.com/applij/article\-pdf/25/4/431/480988/250431\.pdfCited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1),[§1](https://arxiv.org/html/2608.17979#S1.p3.1)\.
- Eder \(2011\)M\. EderStyle\-markers in authorship attribution: a cross\-language study of the authorial fingerprint\.Studies in Polish Linguistics6\(1\),pp\. 99–114\.External Links:[Link](https://ejournals.eu/en/journal_article_files/full_text/018ecec1-46ce-719e-af43-df8ff955dba5/download)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p3.1),[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px3.p2.1)\.
- Efron \(1979\)Bradley\. EfronBootstrap Methods: Another Look at the Jackknife\.The Annals of Statistics7\(1\),pp\. 1 – 26\.External Links:[Document](https://dx.doi.org/10.1214/aos/1176344552),[Link](https://doi.org/10.1214/aos/1176344552)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px3.p1.1)\.
- Fabienet al\.\(2020\)M\. Fabien, E\. Villatoro\-Tello, P\. Motlicek, and S\. ParidaBertAA : BERT fine\-tuning for authorship attribution\.InProceedings of the 17th International Conference on Natural Language Processing \(ICON\),P\. Bhattacharyya, D\. M\. Sharma, and R\. Sangal \(Eds\.\),Indian Institute of Technology Patna, Patna, India,pp\. 127–137\.External Links:[Link](https://aclanthology.org/2020.icon-main.16/)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p3.1)\.
- Gemma Teamet al\.\(2026\)Gemma Team, S\. E\. Abd, V\. Aggarwal, R\. Algayres, A\. Andreev, O\. Bachem, I\. Ballantyne, C\. Brick, V\. Cărbune, M\. Casbon, M\. Chaturvedi, A\. Chawla, V\. Cotruta, A\. Coucke, P\. Culliton, R\. Dadashi, L\. Dixon, M\. Elhawaty, U\. Evci, C\. Farabet, J\. Ferret, F\. Galgani, S\. Girgin, J\. Grill, M\. Grootendorst, J\. Guo, C\. Hardin, Y\. He, S\. M\. Hernandez, O\. Homburger, L\. Hussenot, J\. Ji, A\. Joulin, A\. Kamath, P\. Kassraie, O\. Lacombe, P\. Lahoti, G\. Liu, G\. Martins, L\. Martins, T\. Matejovicova, R\. Merhej, N\. Momchev, S\. Mondal, R\. Mullins, S\. R\. Panyam, S\. Pathak, S\. Perrin, A\. S\. Pinto, E\. Pot, A\. Pouget, A\. Ramé, S\. Ramos, D\. Reid, D\. Rim, M\. Rivière, K\. Roth, L\. Rouillard, O\. Sanseviero, P\. G\. Sessa, S\. Settle, D\. Sinopalnikov, S\. Smoot, P\. Stanczyk, A\. Steiner, L\. Stewart, I\. Tolstikhin, M\. Tschannen, A\. Tsitsulin, N\. Vieillard, R\. Wu, P\. Xu, H\. Yang, E\. Yvinec, B\. Zhang, L\. Zhang, J\. Zou, N\. Aagnes, A\. Abdelhamed, J\. Adamek, S\. Agrawal, S\. Agrawal, I\. Alabdulmohsin, J\. B\. Alayrac, U\. Alon, C\. Amarnath, A\. Anand, C\. Anastasiou, S\. Ariafar, F\. Aubet, K\. Axiotis, F\. Barbero, J\. Barral, A\. Bendebury, U\. Bergmann, S\. Bileschi, K\. Black, M\. Blondel, S\. Borgeaud, A\. Bražinskas, R\. Burnell, R\. Busa\-Fekete, M\. Cai, D\. Calandriello, G\. Cameron, C\. Caucheteux, R\. Chaabouni, G\. Chadha, J\. Chan, B\. J\. Chen, J\. Chen, L\. Chen, X\. Chen, D\. Cheng, T\. Chien, N\. Chinaev, Y\. Chou, Z\. Chu, B\. Coleman, P\. Consul, S\. Conway\-Rahman, S\. Crowell, D\. Cutler, V\. Dani, S\. Daruki, A\. Das, D\. Deutsch, N\. Dikkala, L\. Ding, Q\. Ding, S\. Dodhia, K\. Donhauser, T\. Doshi, A\. Dragan, A\. Druinsky, S\. Dua, Z\. Egyed, D\. Eisenbud, D\. Eppens, C\. Fan, B\. Fatemi, Y\. Fathullah, V\. Feinberg, M\. Ferev, S\. Flennerhag, T\. Fujimoto, J\. G\. Oliveira, I\. Galatzer\-Levy, J\. Gante, S\. Geisler, S\. Ghosal, A\. M\. Girgis, T\. von Glehn, A\. Go, A\. Gokhale, A\. Grills, Y\. Gu, M\. Gupta, P\. Gupta, G\. Guruganesh, R\. Hadsell, H\. Harkous, J\. Harlalka, D\. Hassabis, A\. Hauth, J\. Heyward, A\. Hosseini, C\. Hsia, I\. Hsu, X\. Huang, Y\. Huang, K\. Hui, A\. Hutter, T\. I, F\. Iliopoulos, A\. Jain, G\. Jawahar, Z\. Ji, Q\. Jin, M\. Johnson, K\. Joshi, A\. Kandoor, W\. Kang, K\. Kavukcuoglu, M\. Kazemi, K\. Kenealy, A\. Khalifa, P\. Kirk, I\. Korotkov, S\. Kothawade, V\. Kovalev, N\. Kovelamudi, A\. Kraft, R\. Kumar, V\. Kumar, H\. Kuppam, J\. Lannin, C\. Lee, S\. Lee, D\. Lepikhin, A\. Levkovitch, D\. Li, Q\. Li, V\. Liévin, E\. Lin, Z\. Lin, C\. Liu, T\. Liu, T\. Liu, X\. Liu, I\. Lobov, M\. Lunayach, M\. Ma, G\. Madan, A\. Maksai, E\. Malmi, M\. Matuszak, D\. McDuff, G\. Menghani, M\. Mikuła, D\. Mirylenka, K\. Misiunas, V\. Misra, A\. Mitran, K\. Mohamed, M\. Mukha, E\. Noland, J\. O’Donnell, B\. O’Donoghue, K\. Olszewska, B\. Orlando, W\. Pan, R\. Panigrahy, U\. Parekh, N\. Perez\-Nieves, C\. Park, E\. Paskie, L\. Peng, B\. Petrini, S\. Petrov, J\. Pfeiffer, B\. Piot, M\. Plomecka, S\. Poder, O\. Ponce, A\. Pramanik, D\. Racz, A\. Rajan, M\. Ramanovich, A\. Rao, M\. Ritter, V\. Rodrigues, E\. Rosen, M\. Rybiński, N\. Sachdeva, M\. E\. Sander, R\. Sathyanarayana, S\. Savla, S\. Schmidgall, T\. Schuster, G\. Scrivener, B\. Seguin, A\. Sellergren, A\. Severyn, I\. Shafran, D\. Shah, B\. Shahriari, Y\. Shangguan, A\. Shenoy, P\. Shenoy, R\. Shivanna, P\. Sho, L\. Spangher, W\. Stokowiec, T\. Strother, Y\. Su, Y\. Sun, M\. Sundararajan, A\. Tacchetti, M\. H\. Taege, P\. Tafti, J\. Tarbouriech, C\. Tekur, S\. Thakoor, R\. Thapa, M\. Traverse, L\. Treven, T\. Tu, C\. T\. Tung, Ç\. Ünlü, P\. Veličković, M\. P\. Venkat, S\. G\. Venkatesh, V\. Venkiteswaran, F\. Visin, A\. Vitvitskyi, K\. Vodrahalli, W\. Wang, X\. Wang, T\. Warkentin, J\. Wassenberg, J\. Wieting, C\. Wu, L\. Xiao, H\. Xu, Y\. Xu, F\. Xue, A\. Yadav, J\. Yan, A\. Yang, L\. Yang, M\. Yang, Z\. Ying, J\. H\. Yoo, M\. Zadimoghaddam, S\. Zafar, F\. Zhang, J\. Zhang, J\. Zhang, X\. Zhang, C\. Zhao, D\. Zhou, and C\. ZouGemma 4 technical report\.External Links:2607\.02770,[Link](https://arxiv.org/abs/2607.02770)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px1.p3.1)\.
- Gemma Teamet al\.\(2025\)Gemma Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. HussenotGemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px1.p3.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px6.p1.1)\.
- Grieve \(2007\)J\. GrieveQuantitative authorship attribution: an evaluation of techniques\.Literary and linguistic computing22\(3\),pp\. 251–270\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1093/llc/fqm020)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p2.1)\.
- Guptaet al\.\(2019\)S\. T\. Gupta, J\. K\. Sahoo, and R\. K\. RoulAuthorship identification using recurrent neural networks\.InProceedings of the 2019 3rd International Conference on Information System and Data Mining,ICISDM ’19,New York, NY, USA,pp\. 133–137\.External Links:ISBN 9781450366359,[Link](https://doi.org/10.1145/3325917.3325935),[Document](https://dx.doi.org/10.1145/3325917.3325935)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p3.1)\.
- Halvani and Graner \(2021\)O\. Halvani and L\. GranerPOSNoise: an effective countermeasure against topic biases in authorship analysis\.InProceedings of the 16th International Conference on Availability, Reliability and Security,ARES ’21,New York, NY, USA\.External Links:ISBN 9781450390514,[Link](https://doi.org/10.1145/3465481.3470050),[Document](https://dx.doi.org/10.1145/3465481.3470050)Cited by:[Appendix G](https://arxiv.org/html/2608.17979#A7.p1.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- Halvaniet al\.\(2016\)O\. Halvani, C\. Winter, and A\. PflugAuthorship verification for different languages, genres and topics\.Digital Investigation16,pp\. S33–S43\.Note:DFRWS 2016 EuropeExternal Links:ISSN 1742\-2876,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.diin.2016.01.006),[Link](https://www.sciencedirect.com/science/article/pii/S1742287616000074)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px3.p1.1)\.
- Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px1.p3.1)\.
- Huet al\.\(2023\)X\. Hu, W\. Ou, S\. Acharya, S\. H\.H\. Ding, R\. D’Gama, and H\. YuTDRLM: stylometric learning for authorship verification by topic\-debiasing\.Expert Systems with Applications233,pp\. 120745\.External Links:ISSN 0957\-4174,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.eswa.2023.120745),[Link](https://www.sciencedirect.com/science/article/pii/S0957417423012472)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- Huet al\.\(2024\)Y\. Hu, Z\. Hu, C\. Seah, and R\. K\. LeeInstructAV: instruction fine\-tuning large language models for authorship verification\.External Links:2407\.12882,[Link](https://arxiv.org/abs/2407.12882)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p4.1)\.
- Huanget al\.\(2024\)B\. Huang, C\. Chen, and K\. ShuCan large language models identify authorship?\.pp\. 445–460\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.26/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.26)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p4.1)\.
- Israeliet al\.\(2025\)A\. Israeli, S\. Liu, J\. May, and D\. JurgensThe million authors corpus: a cross\-lingual and cross\-domain Wikipedia dataset for authorship verification\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 25997–26017\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1335/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1335),ISBN 979\-8\-89176\-256\-5Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2608.17979#S5.SS0.SSS0.Px1.p4.1)\.
- Kieferet al\.\(2026\)L\. Kiefer, C\. Leiter, S\. Takeshita, E\. Schmidt, and S\. EgerGerAV: towards new heights in German authorship verification using fine\-tuned LLMs on a new benchmark\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 40050–40069\.External Links:[Link](https://aclanthology.org/2026.findings-acl.1991/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1991),ISBN 979\-8\-89176\-395\-1Cited by:[Appendix C](https://arxiv.org/html/2608.17979#A3.p1.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p4.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px1.p3.1),[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px3.p2.1)\.
- Kimet al\.\(2025\)J\. Kim, H\. Zhang, and D\. JurgensLeveraging multilingual training for authorship representation: enhancing generalization across languages and domains\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 34867–34892\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1766/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1766),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px1.p2.1)\.
- Kreuz \(2023\)R\. KreuzLinguistic fingerprints: how language creates and reveals identity\.Simon and Schuster\.External Links:ISBN 978\-1\-63388\-897\-5Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p3.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with pagedattention\.InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles,External Links:[Document](https://dx.doi.org/10.1145/3600006.3613165)Cited by:[Appendix C](https://arxiv.org/html/2608.17979#A3.p1.1)\.
- Leeet al\.\(2022\)M\. Lee, P\. Liang, and Q\. YangCoAuthor: designing a human\-ai collaborative writing dataset for exploring language model capabilities\.InCHI Conference on Human Factors in Computing Systems,CHI ’22,pp\. 1–19\.External Links:[Link](http://dx.doi.org/10.1145/3491102.3502030),[Document](https://dx.doi.org/10.1145/3491102.3502030)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p3.1)\.
- Lewiset al\.\(2004\)D\. D\. Lewis, Y\. Yang, T\. G\. Rose, and F\. LiRCV1: a new benchmark collection for text categorization research\.J\. Mach\. Learn\. Res\.5,pp\. 361–397\.External Links:[Link](https://api.semanticscholar.org/CorpusID:11027141)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1)\.
- Maet al\.\(2025\)M\. Ma, D\. M\. Le, J\. Kang, Y\. Dou, J\. Cadigan, D\. Freitag, A\. Ritter, and W\. XuCROSSNEWS: a cross\-genre authorship verification and attribution benchmark\.Proceedings of the AAAI Conference on Artificial Intelligence39\(23\),pp\. 24777–24785\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/34659),[Document](https://dx.doi.org/10.1609/aaai.v39i23.34659)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p3.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px6.p1.1),[§5](https://arxiv.org/html/2608.17979#S5.SS0.SSS0.Px1.p4.1),[§5](https://arxiv.org/html/2608.17979#S5.SS0.SSS0.Px2.p1.1),[Table 2](https://arxiv.org/html/2608.17979#S5.T2),[Table 2](https://arxiv.org/html/2608.17979#S5.T2.2.1.1.1.2.1),[Table 2](https://arxiv.org/html/2608.17979#S5.T2.2.1.1.1.3.1)\.
- Manolacheet al\.\(2022\)A\. Manolache, F\. Brad, A\. Barbalau, R\. T\. Ionescu, and M\. PopescuVeriDark: a large\-scale benchmark for authorship verification on the dark web\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 15574–15588\.External Links:[Document](https://dx.doi.org/10.52202/068431-1133),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/64008fa30cba9b4d1ab1bd3bd3d57d61-Paper-Datasets_and_Benchmarks.pdf)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1)\.
- Murauer and Specht \(2019\)B\. Murauer and G\. SpechtGenerating cross\-domain text classification corpora from social media comments\.InExperimental IR Meets Multilinguality, Multimodality, and Interaction,F\. Crestani, M\. Braschler, J\. Savoy, A\. Rauber, H\. Müller, D\. E\. Losada, G\. Heinatz Bürki, L\. Cappellato, and N\. Ferro \(Eds\.\),Cham,pp\. 114–125\.External Links:ISBN 978\-3\-030\-28577\-7,[Document](https://dx.doi.org/https%3A//doi.org/10.1007/978-3-030-28577-7%5F7)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px3.p1.1)\.
- Mysoreet al\.\(2025\)S\. Mysore, D\. Das, H\. Cao, and B\. SarrafzadehPrototypical human\-AI collaboration behaviors from LLM\-assisted writing in the wild\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,Suzhou, China,pp\. 16819–16846\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.852/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.852),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p3.1)\.
- Nini \(2023\)A\. NiniA theory of linguistic individuality for authorship analysis\.Elements in Forensic Linguistics,Cambridge University Press\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1017/9781108974851)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1),[§1](https://arxiv.org/html/2608.17979#S1.p3.1)\.
- OpenAI \(2022\)OpenAIGPT\-3\.5\.Technical reportOpenAI\.External Links:[Link](https://platform.openai.com/docs/models/gpt-3.5-turbo)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p4.1),[§3](https://arxiv.org/html/2608.17979#S3.p3.1)\.
- OpenAI \(2023\)OpenAIGPT\-4\.Technical reportOpenAI\.External Links:[Link](https://cdn.openai.com/papers/gpt-4-system-card.pdf)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p4.1)\.
- Overdorf and Greenstadt \(2016\)R\. Overdorf and R\. GreenstadtBlogs, twitter feeds, and reddit comments: cross\-domain authorship attribution\.Proceedings on Privacy Enhancing Technologies2016,pp\.\.External Links:[Document](https://dx.doi.org/10.1515/popets-2016-0021)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- Pearson and Galton \(1895\)K\. Pearson and F\. GaltonVII\. note on regression and inheritance in the case of two parents\.Proceedings of the Royal Society of London58\(347\-352\),pp\. 240–242\.External Links:[Document](https://dx.doi.org/10.1098/rspl.1895.0041),[Link](https://royalsocietypublishing.org/doi/abs/10.1098/rspl.1895.0041),https://royalsocietypublishing\.org/doi/pdf/10\.1098/rspl\.1895\.0041Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px4.p1.1)\.
- Qianet al\.\(2017\)C\. Qian, T\. He, and R\. ZhangDeep learning based authorship identification\.Report, Stanford University,pp\. 1–9\.External Links:[Link](https://api.semanticscholar.org/CorpusID:42982101)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p3.1)\.
- Qiuet al\.\(2025\)J\. Qiu, J\. Zhu, A\. Patel, M\. Apidianaki, and C\. Callison\-BurchMStyleDistance: multilingual style embeddings and their evaluation\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 16917–16931\.External Links:[Link](https://aclanthology.org/2025.findings-acl.869/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.869),ISBN 979\-8\-89176\-256\-5Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px3.p1.1)\.
- Ramnathet al\.\(2025\)S\. Ramnath, K\. Pandey, E\. Boschee, and X\. RenCAVE: controllable authorship verification explanations\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 8939–8961\.External Links:[Link](https://aclanthology.org/2025.naacl-long.451/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.451),ISBN 979\-8\-89176\-189\-6Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p4.1)\.
- Richardson \(2025\)Beautiful soup documentationExternal Links:[Link](https://beautiful-soup.readthedocs.io/en/latest/)Cited by:[§3](https://arxiv.org/html/2608.17979#S3.p2.1)\.
- Richburget al\.\(2024\)A\. Richburg, C\. Bao, and M\. CarpuatAutomatic authorship analysis in human\-AI collaborative writing\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 1845–1855\.External Links:[Link](https://aclanthology.org/2024.lrec-main.165/)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p3.1)\.
- Rivera\-Sotoet al\.\(2021\)R\. A\. Rivera\-Soto, O\. E\. Miano, J\. Ordonez, B\. Y\. Chen, A\. Khan, M\. Bishop, and N\. AndrewsLearning universal authorship representations\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 913–919\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.70/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.70)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p3.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- Spearman \(1904\)C\. Spearman’General intelligence,’ objectively determined and measured\.The American Journal of Psychology15\(2\),pp\. 201–293\.External Links:[Document](https://dx.doi.org/10.2307/1412107),[Link](https://doi.org/10.2307/1412107)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px4.p1.1)\.
- Stamatatoset al\.\(2022\)E\. Stamatatos, M\. Kestemont, K\. Kredens, P\. Pezik, A\. Heini, J\. Bevendorff, B\. Stein, and M\. PotthastOverview of the authorship verification task at pan 2022\.InCEUR workshop proceedings,Vol\.3180,pp\. 2301–2313\.Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- Stamatatoset al\.\(2023\)E\. Stamatatos, K\. Kredens, P\. Pezik, A\. Heini, J\. Bevendorff, B\. Stein, and M\. PotthastOverview of the authorship verification task at pan 2023\.InConference and Labs of the Evaluation Forum,External Links:[Link](https://api.semanticscholar.org/CorpusID:264441636)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1),[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.17979#S5.SS0.SSS0.Px1.p4.1)\.
- Stamatatos \(2006\)E\. StamatatosEnsemble\-based author identification using character n\-grams\.InProceedings of the 3rd International Workshop on Text\-based Information Retrieval,Vol\.36,pp\. 41–46\.External Links:[Link](https://api.semanticscholar.org/CorpusID:4632801)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p2.1)\.
- Stamatatos \(2018\)E\. StamatatosMasking topic\-related information to enhance authorship attribution\.Journal of the Association for Information Science and Technology69\(3\),pp\. 461–473\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/asi.23968),[Link](https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.23968),https://asistdl\.onlinelibrary\.wiley\.com/doi/pdf/10\.1002/asi\.23968Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- Tyoet al\.\(2023\)J\. Tyo, B\. Dhingra, and Z\. C\. LiptonValla: standardizing and benchmarking authorship attribution and verification through empirical evaluation and comparative analysis\.InProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),J\. C\. Park, Y\. Arase, B\. Hu, W\. Lu, D\. Wijaya, A\. Purwarianti, and A\. A\. Krisnadhi \(Eds\.\),Nusa Dua, Bali,pp\. 649–660\.External Links:[Link](https://aclanthology.org/2023.ijcnlp-main.43/),[Document](https://dx.doi.org/10.18653/v1/2023.ijcnlp-main.43)Cited by:[§1](https://arxiv.org/html/2608.17979#S1.p1.1)\.
- van Leeuwenet al\.\(2026\)B\. van Leeuwen, S\. Bhulai, and R\. van der MeiCross\-domain authorship verification with feature interaction networks: evaluating no\-holdout and holdout protocols\.Machine Learning with Applications25,pp\. 100943\.External Links:ISSN 2666\-8270,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.mlwa.2026.100943),[Link](https://www.sciencedirect.com/science/article/pii/S2666827026001088)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
- von Werraet al\.\(2020\)TRL: Transformers Reinforcement LearningExternal Links:[Link](https://github.com/huggingface/trl)Cited by:[Appendix C](https://arxiv.org/html/2608.17979#A3.p1.1)\.
- Wanget al\.\(2024\)L\. Wang, N\. Yang, X\. Huang, L\. Yang, R\. Majumder, and F\. WeiImproving text embeddings with large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 11897–11916\.External Links:[Link](https://aclanthology.org/2024.acl-long.642/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.642)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px6.p1.1)\.
- Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. M\. RushTransformers: State\-of\-the\-Art Natural Language Processing\.pp\. 38–45\.External Links:[Link](https://www.aclweb.org/anthology/2020.emnlp-demos.6)Cited by:[Appendix C](https://arxiv.org/html/2608.17979#A3.p1.1)\.
- Yanget al\.\(2017\)M\. Yang, D\. Zhu, Y\. Tang, and J\. WangAuthorship attribution with topic drift model\.Proceedings of the AAAI Conference on Artificial Intelligence31\(1\)\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/11062),[Document](https://dx.doi.org/10.1609/aaai.v31i1.11062)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p2.1)\.
- Youden \(1950\)W\. J\. YoudenIndex for rating diagnostic tests\.Cancer3\(1\),pp\. 32–35\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/1097-0142%281950%293%3A1%3C32%3A%3AAID-CNCR2820030106%3E3.0.CO%3B2-3)Cited by:[§4](https://arxiv.org/html/2608.17979#S4.SS0.SSS0.Px3.p1.1)\.
- Zenget al\.\(2025\)P\. Zeng, P\. Alipoormolabashi, J\. Mun, G\. Dey, N\. Soni, N\. Balasubramanian, O\. Rambow, and H\. SchwartzResidualized similarity for faithfully explainable authorship verification\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 15824–15837\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.856/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.856),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px1.p2.1)\.
- Zhanget al\.\(2021\)Y\. Zhang, D\. Boumber, M\. Hosseinia, F\. Yang, and A\. MukherjeeImproving authorship verification using linguistic divergence\.InROMCIR@ECIR,External Links:[Link](https://api.semanticscholar.org/CorpusID:232223310)Cited by:[§2](https://arxiv.org/html/2608.17979#S2.SS0.SSS0.Px2.p1.1)\.
## Appendix AAVShift Example
We present document examples from each genre written by a single author in Table[4](https://arxiv.org/html/2608.17979#A1.T4), together with their English translations\. These examples provide an impression of how writing style varies across genres\. The forum text is relatively informal, for instance using emojis such as "xD", whereas the story example differs substantially by employing a more literary style\. The review is again more informal but has the distinctive characteristic of directly addressing the author of the story\.
ForumStoryReviewIm Sommer saß ich mal mit einer Freundin in meinem Zimmer am Fußboden\. Wir haben gemalt, da fiel auf einmal eine riesige Raupe/Larve von der Decke\! Wir wissen bis heute nicht, wie sie da hin gekommen ist\.Das andere ist im Sommer am Schulfest passiert\. Ich saß im Gras, als mir plötzlich ein Vogel was auf’s Knie fallen ließ \- genau so wie letztes Jahr\. xDMegatron bewegte sich durch die dunklen Gänge, doch seine Gedanken waren ferner denn je\. Jeder Schritt quälte die Ruhe, wie der Ton eines fallenden Tropfens die endliche Stille\. Seine Präsenz füllte den Ort wie das Summen einer Stimmgabel, das in jede Ecke drang, sich selbst in jenen Flächen nieder ließ, die den Raum begrenzten, um ihn aus seinem Frieden zu reißen\. \[…\]Hi, hui ja, das ist wirklich ein ungewöhnliches Pairing xD Aber nicht schlecht, dein Stil ist eigentlich richtig gut und es lässt sich alles flüssig lesen\. Ich schließe mich Hera an und meine, dass Absätze im Text nicht schlecht gewesen wären\. Du musst bedenken, dass die Geschichte am Bildschirm gelesen wird, was anstrengend für die Augen ist\. \[…\]TranslationOne summer, I was sitting on the floor in my room with a friend\. We were drawing when suddenly a huge caterpillar/larva fell from the ceiling\! To this day, we still don’t know how it got there\. The other thing happened at the school festival that summer\. I was sitting in the grass when suddenly a bird dropped something on my knee \- just like last year\. xDMegatron moved through the dark corridors, yet his thoughts were farther away than ever\. Each step disturbed the stillness, like the sound of a falling drop breaking the finite silence\. His presence filled the place like the hum of a tuning fork, penetrating every corner, settling even into the surfaces that bounded the room, to tear it from its peace\. \[…\]Hi, wow, yeah, that’s really an unusual pairing xD But not bad, your writing style is actually really good, and it all reads smoothly\. I agree with Hera \- I think some paragraphs in the text wouldn’t have been a bad idea\. You have to keep in mind that the story is being read on a screen, which can be hard on the eyes\. \[…\]Table 4:Example of texts from one user writing in all three genres alongside English translations\. Usernames occurring in the examples have been pseudonymized\.
## Appendix BXGB Feature Configuration
Table[5](https://arxiv.org/html/2608.17979#A2.T5)presents the feature configuration used in the feature\-based approach\. It lists each feature abbreviation together with a brief description\.
FeatureDescriptionmfw2Normalized frequencies of the 1,000 most frequent word bigrams\.mftNormalized frequencies of the 1,000 most frequent POS trigrams\.mfcNormalized frequencies of the 2,500 most frequent character 4\-grams\.mfeNormalized frequencies of the 100 most frequent emojis\.wordLenDistriDistribution of word lengths from 1 to 20 characters\.wordLenAverage word length\.messageLenAverage document length\.nrPunctuationNormalized frequencies of individual punctuation symbols\.nrOOVProportion of out\-of\-vocabulary words\.Table 5:Handcrafted features used to train the feature\-based XGB AV model\.
## Appendix CTraining Setup and Hyperparameters
To support LoRA fine\-tuning of the Gemma\-4\-31B\-it model, we conduct our experiments on a system equipped with four H200 GPUs\. We follow the hyperparameter configuration proposed by[25](https://arxiv.org/html/2608.17979#bib.bib5)and update the software stack to support the newer Gemma version\. Specifically, we use PyTorch 2\.12\.1\([1](https://arxiv.org/html/2608.17979#bib.bib55)\), Transformers 5\.12\.1\([55](https://arxiv.org/html/2608.17979#bib.bib57)\), TRL 0\.21\.0\([53](https://arxiv.org/html/2608.17979#bib.bib56)\), and vLLM 0\.24\.0\([28](https://arxiv.org/html/2608.17979#bib.bib58)\)\.
For the feature\-based XGBoost model, we use an environment with Transformers 4\.36\.2 and PyTorch 2\.5\.1\. We perform a dedicated hyperparameter grid search for each training setting, tuning the maximum tree depth \(3, 6, 10, 15\), minimum child weight \(0, 2, 4, 5\), regularization parameterα\\alpha\(0, 1\),γ\\gamma\(0, 1\), and the learning rate \(0\.01, 0\.1, 0\.3\)\.
The MSR model is evaluated using Transformers 5\.12\.1, Sentence\-Transformers 5\.2\.2, and PyTorch 2\.12\.1\. We use the original sentence embeddings without modification and tune only the verification threshold for each dataset\.
We release complete environment and training setups within our GitHub repository for reproducibility\.
## Appendix DGenreShift Full Results
Figure[5](https://arxiv.org/html/2608.17979#A4.F5)reports accuracy scores to complement the F1 results presented in the main results section\. The overall findings remain unchanged, with Gemma consistently outperforming the other models and the review genre reaching highest performance\.
Figure 5:Accuracy scores of all models trained and evaluated on all GenreShift train and test splits\. Model names are shown on the y\-axis and test dataset names on the x\-axis\. Each model name is followed by its training or calibration dataset\. The best score for each test set is shown in bold\. Statistically significant superiority over all other models is indicated by an asterisk \(\*; p < 0\.05\)\.
## Appendix EGenreShift Standardized Full Results
Figures[6](https://arxiv.org/html/2608.17979#A5.F6)and[7](https://arxiv.org/html/2608.17979#A5.F7)present the full F1 and accuracy results, respectively, on the standardized GenreShift benchmark\. Compared to the unstandardized benchmark, Gemma no longer consistently outperforms the other approaches but shares the best performance with either MSR or XGB across all datasets\. Performance in cross\-genre settings drops significantly to a consistent F1 score of 0\.67\-0\.68 across all genre pairs, indicating that larger training sizes are required for this more challenging setting\. Furthermore, mixed training no longer provides a performance benefit in the standardized setting\. However, the performance differences between genres in the in\-domain setting remain consistent: review data again achieves substantially higher performance, despite document lengths and training sample sizes being standardized across all three genres\.
Figure 6:F1 scores of all models trained and evaluated on all standardized GenreShift train and test splits\. Model names are shown on the y\-axis and test dataset names on the x\-axis\. Each model name is followed by its training or calibration dataset\. The best score for each test set is shown in bold\. Statistically significant superiority over all other models is indicated by an asterisk \(\*; p < 0\.05\)\.Figure 7:Accuracy scores of all models trained and evaluated on all standardized GenreShift train and test splits\. Model names are shown on the y\-axis and test dataset names on the x\-axis\. Each model name is followed by its training or calibration dataset\. The best score for each test set is shown in bold\. Statistically significant superiority over all other models is indicated by an asterisk \(\*; p < 0\.05\)\.
## Appendix FAI\-Era Results
Figure 8:Gemma performance under different era evaluations for each genre\. Datasets on the x\-axis refer to the held\-out test era with F1\-scores on the y\-axis\.Figure[8](https://arxiv.org/html/2608.17979#A6.F8)shows the F1 scores for the leave\-one\-era\-out evaluation across all three genres\. While several significant differences between hold\-out eras are observed, no consistent pattern emerges\. For Forum, the AI era yields the lowest performance and is significantly outperformed by all other eras \(p<0\.05p<0\.05\), although the Pre\-AI era is likewise significantly outperformed by the Early and Mid eras\. In contrast, Review achieves its highest F1 score \(0\.89\) on the AI\-era test set, significantly outperforming all other eras\. For Story, the AI era is significantly outperformed only by the Early era, whereas the Mid era performs significantly worse than all remaining periods\.\.
## Appendix GMost and Least Stable Features
Table[6](https://arxiv.org/html/2608.17979#A7.T6)presents the top and bottom 30 features ranked by stability score across all three genre transfers \(see Table[5](https://arxiv.org/html/2608.17979#A2.T5)for feature descriptions\)\. Among the most stable features, we find several word bigrams consisting of function words, as well as part\-of\-speech \(POS\) bigrams, confirming earlier findings that function words and POS sequences represent stable indicators of writing style\([18](https://arxiv.org/html/2608.17979#bib.bib3)\)\. We further identify characteristic punctuation usage patterns, such as repeated exclamation marks\.
Among the least stable features, we find, for example, average message length, reflecting the challenge of varying document lengths across different text genres\. However, we also observe individual word bigrams, such as "als er", and POS trigrams, such as "punct\-propn\-verb", indicating that not all function word or POS sequences constitute stable style predictors\. Furthermore, many character 4\-grams appear among the least stable features, suggesting that character\-level patterns may be more sensitive to genre\-specific variations\.
Top FeaturesBottom Featuresmfw2\_aber inmfc\_ konmfw2\_so alsmfc\_hiennrPunct\|mfc\_seinmfw2\_mal einemfc\_onntmfw2\_sein ichmfc\_ blimfw2\_sollte manmft\_punct propn verbmfw2\_nicht alsmfw2\_als ernrPunct\#mfc\_Blicmfw2\_ist dochmfc\_ Ermfw2\_von sichmfc\_ sahmfw2\_als dasmfc\_konnmfw2\_als diemfc\_sichmfw2\_nur dassmfc\_ ermft\_sconj adj nounmfc\_„IchnrPunct\\mfc\_ „Icmfw2\_nach einemmfc\_hielmfw2\_ist wiewordLenDistri5mft\_part punct auxmessageLen0mfw2\_nicht gerademfc\_gtemft\_sconj det detmfc\_ertemfw2\_zum beispielmft\_punct punct verbmfw2\_mit einermfc\_egtemfc\_\!\!\!\!mfc\_eltemfw2\_doch auchmfc\_attemfw2\_nicht fürmfc\_te\.nrPunct\+mfc\_hattmfw2\_und allesmfc\_ktemfw2\_du siemfc\_cktemfc\_ Selmfc\_tetemfw2\_wir sindmfc\_ete
Table 6:Top 30 most and least stable features across all genre transfers\.
## Appendix HPairwise Feature Analysis
Table[7](https://arxiv.org/html/2608.17979#A8.T7)reports the overlap of the 100 most and least stable features across different genre transitions\. The overlap is generally low, ranging from 5% to 27% for the most stable features and from 2% to 43% for the least stable features\. This indicates that no universally stable feature set exists across genre shifts, suggesting that feature stability should be analyzed separately for each genre pair, particularly when performing explicit feature selection\. In contrast, Spearman rank correlations of feature stability remain moderate \(0\.42\-0\.72\), indicating that while the most and least stable features vary substantially between genre pairs, the overall ranking of feature stability is comparatively consistent\.
Pair1Pair 2Shared Top 100Shared Bottom 100Spearman Rank CorrelationReview\-ForumStory\-Forum27%3%0\.63Review\-ForumReview\-Story8%2%0\.42Story\-ForumReview\-Story5%43%0\.72
Table 7:Differences in overlap of the 100 most and least stable features and relative stability feature ranking for each genre\-pair combination\.
## Appendix IModel and Data Licences
We use all models in this work in accordance with their intended use as specified by their respective licences\. Specifically, Gemma\-4\-31B\-it is released under the Apache 2\.0 licence111[https://ai\.google\.dev/gemma/apache\_2](https://ai.google.dev/gemma/apache_2), which permits modification, including fine\-tuning, redistribution of fine\-tuned models, and publication of research results\.
To support the reproducibility of our experiments and encourage future research on AV under distribution shifts while protecting the privacy of the original authors, we will provide access to the AVShift benchmark with pseudonymized usernames for academic research purposes only\. In addition, we will release the complete preprocessing pipeline and the scraping code to facilitate reproducibility and provide transparency regarding the dataset construction process\.Similar Articles
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Introduces ARB, a matched authorship-rewriting benchmark for evaluating AI-text detectors, showing that detector performance drops significantly when human text is rewritten by an LLM despite high recall on direct LLM-generated text.
@emollick: There is a lot being written about the stylistic tells of AI writing (em-dashes, etc.) but this paper looks at AI narra…
This paper introduces StoryScope, a pipeline that analyzes discourse-level narrative features to distinguish AI-generated fiction from human-written stories. It achieves high accuracy and reveals distinct narrative fingerprints for different LLMs like Claude, GPT, and Gemini.
Are AI writing tools quietly flattening how everyone writes into one voice?
The article explores the hypothesis that AI writing tools are subtly flattening individual writing voices into a shared average style, potentially reducing stylistic diversity over time.
Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift
This paper proposes a test-time adaptation approach using semi-supervised learning for AI text detection that adapts to continual distribution shifts from new LLMs, adversarial humanization, and temporal drift, outperforming state-of-the-art supervised detectors.
AI can write prize-winning fiction. Now what?
An article discussing the controversy over a prize-winning short story that was accused of being generated by AI, and the broader implications for authorship and detection in the age of large language models.