Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation

arXiv cs.CL Papers

Summary

Proposes LiSCP, a lightweight stylistic consistency profiling method for robust detection of LLM-generated textual content, focusing on feature stability under adversarial manipulation. Achieves superior performance on in-domain and cross-domain detection with notable robustness.

arXiv:2605.05950v1 Announce Type: new Abstract: The increasing prevalence of Large Language Models (LLMs) in content creation has made distinguishing human-written textual content from LLM-generated counterparts a critical task for multimedia moderation. Existing detectors often rely on statistical cues or model-specific heuristics, making them vulnerable to paraphrasing and adversarial manipulations, and consequently limiting their robustness and interpretability. In this work, we proposeLiSCP , a novel lightweight stylistic consistency profiling method for robust detection of LLM-generated textual content, focusing on feature stability under adversarial manipulation. Our approach constructs a consistency profile that combines discrete stylistic features with continuous semantic signals, leveraging stylistic stability across multimodal-guided paraphrased text variants. Experiments spanning real-world multimedia news and movie datasets and conventional text domains demonstrate that LiSCP achieves superior performance on in-domain detection and outperforms existing approaches by up to 11.79% in cross-domain settings. Additionally,it demonstrates notable robustness under adversarial scenarios, including adversarial attacks and hybrid human-AI settings.
Original Article
View Cached Full Text

Cached at: 05/08/26, 06:49 AM

# Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
Source: [https://arxiv.org/html/2605.05950](https://arxiv.org/html/2605.05950)
Siyuan Li,Aodu WulianghaiSchool of Computer Science, Shanghai Jiao Tong UniversityShanghaiChina[melusine\.wlhad@sjtu\.edu\.cn](https://arxiv.org/html/2605.05950v1/mailto:[email protected]),Xi LinSchool of Computer Science, Shanghai Jiao Tong UniversityShanghaiChina[linxi234@sjtu\.edu\.cn](https://arxiv.org/html/2605.05950v1/mailto:[email protected]),Xibin YuanSchool of Computer Science, Shanghai Jiao Tong UniversityShanghaiChina[2022yxb@sjtu\.edu\.cn](https://arxiv.org/html/2605.05950v1/mailto:[email protected]),Qinghua MaoSchool of Computer Science, Shanghai Jiao Tong UniversityShanghaiChina[mmmm2018@sjtu\.edu\.cn](https://arxiv.org/html/2605.05950v1/mailto:[email protected]),Guangyan LiInstitute of Automation, Chinese Academy of SciencesBeijingChina[liguangyan2022@ia\.ac\.cn](https://arxiv.org/html/2605.05950v1/mailto:[email protected]),Xiang ChenCollege of Computer Science and Technology, Zhejiang UniversityHangzhouChina[wasdnsxchen@gmail\.com](https://arxiv.org/html/2605.05950v1/mailto:[email protected]),Jun WuSchool of Computer Science, Shanghai Jiao Tong UniversityShanghaiChina[junwuhn@sjtu\.edu\.cn](https://arxiv.org/html/2605.05950v1/mailto:[email protected])andJianhua LiSchool of Computer Science, Shanghai Jiao Tong UniversityShanghaiChina[lijh888@sjtu\.edu\.cn](https://arxiv.org/html/2605.05950v1/mailto:[email protected])

\(2026\)

###### Abstract\.

The increasing prevalence of Large Language Models \(LLMs\) in content creation has made distinguishing human\-written textual content from LLM\-generated counterparts a critical task for multimedia moderation\. Existing detectors often rely on statistical cues or model\-specific heuristics, making them vulnerable to paraphrasing and adversarial manipulations, and consequently limiting their robustness and interpretability\. In this work, we proposeLiSCP, a novel lightweight stylistic consistency profiling method for robust detection of LLM\-generated textual content, focusing on feature stability under adversarial manipulation\. Our approach constructs a consistency profile that combines discrete stylistic features with continuous semantic signals, leveraging stylistic stability across multimodal\-guided paraphrased text variants\. Experiments spanning real\-world multimedia news and movie datasets and conventional text domains demonstrate that LiSCP achieves superior performance on in\-domain detection and outperforms existing approaches by up to 11\.79% in cross\-domain settings\. Additionally, it demonstrates notable robustness under adversarial scenarios, including adversarial attacks and hybrid human\-AI settings\.

LLM\-generated content detection, Multimedia content Moderation, Stylistic consistency profiling

††copyright:none††journalyear:2026††conference:the 34th ACM International Conference on Multimedia; November 10–14, 2026; Rio de Janeiro, Brazil††ccs:Computing methodologies Artificial intelligence††ccs:Computing methodologies Machine learning## 1\.Introduction

Large Language Models \(LLMs\) have become a default tool for open\-domain content creation, enabling high\-quality generation across various domains such as news articles, product reviews, academic essays, technical documentation, and even multimedia\-associated textual content \(e\.g\., image captions, video subtitles, and multimodal platform reviews\)\(Dubeyet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib55); Chenet al\.,[2026](https://arxiv.org/html/2605.05950#bib.bib11); Suet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib10); Hurstet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib56); Liet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib9); Xuet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib59)\)\. Widely deployed in both everyday tools and professional workflows, these models increasingly blur the distinction between human\-written and machine\-generated textual content, raising serious concerns regarding content authenticity, attribution, and accountability, which are particularly acute in multimedia content moderation scenarios where text often interacts with visual or audio modalities\(Sadasivanet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib27); Wuet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib52); Yuet al\.,[2025b](https://arxiv.org/html/2605.05950#bib.bib53); Liet al\.,[2026b](https://arxiv.org/html/2605.05950#bib.bib12); Arevaloet al\.,[2017](https://arxiv.org/html/2605.05950#bib.bib47); Nguyenet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib57)\)\. Consequently, the reliable detection of LLM\-generated textual content has become critical for applications including academic integrity enforcement, auditing of high\-stakes decisions, maintaining trust in digital communication, and ensuring the credibility of multimedia\(Liet al\.,[2025b](https://arxiv.org/html/2605.05950#bib.bib48); Cao,[2025](https://arxiv.org/html/2605.05950#bib.bib58)\)\.

Despite significant advances, robust detection under realistic conditions—especially in multimedia\-derived scenarios—remains an open challenge\. Current detectors often rely on token\-level statistics \(e\.g\., likelihood\- or rank\-based features\) or model\-specific heuristics\(Hanset al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib17); Gehrmannet al\.,[2019](https://arxiv.org/html/2605.05950#bib.bib3); Liet al\.,[2026a](https://arxiv.org/html/2605.05950#bib.bib60); Koikeet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib49); Abdelnabi and Fritz,[2021](https://arxiv.org/html/2605.05950#bib.bib36); Kirchenbaueret al\.,[2023](https://arxiv.org/html/2605.05950#bib.bib37)\), which tend to degrade when the generative model changes, the domain shifts, the text undergoes post\-editing, or the textual content is adjusted to align with accompanying visual elements \(e\.g\., edited image captions for misinformation propagation\)\(Chen and Wang,[2025](https://arxiv.org/html/2605.05950#bib.bib14); Chenget al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib51); Zhouet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib54); Guoet al\.,[2023](https://arxiv.org/html/2605.05950#bib.bib61); Krishnaet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib35)\)\. Although recent paraphrase\- or re\-query\-based approaches reduce the need for supervision, they frequently treat paraphrasing merely as a means of score aggregation and remain tightly coupled to specific models or prompting strategies\. This leads to unpredictable performance under distribution shifts, stylistic variations, hybrid human–AI compositions, or multimodal context changes \(e\.g\., text reused across unrelated images\)\(Leiet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib13); Baoet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib32); Zhanget al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib33)\)\. Moreover, most existing methods evaluate a text as a monolithic unit, offering limited insight into how stylistic patterns behave under meaning\-preserving manipulations—an issue that is exacerbated when text is part of a broader multimedia ecosystem requiring cross\-modal consistency\(Chen and Wang,[2025](https://arxiv.org/html/2605.05950#bib.bib14); Baoet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib32)\)\.

To focus on the above problems, this work is guided by the following key research questions \(RQs\):

- •RQ1:How to design a detection framework that remains robust under adversarial manipulations without relying on heavyweight models?
- •RQ2:How to enhance the detection generalization across domains and multimedia\-derived textual scenarios, especially when statistical features become unreliable under semantic\-preserving transformations?

In response to these two questions, we proposeLiSCP, alightweight stylistic consistency profilingmethod for LLM\-generated textual content detection\. Our method builds on a simple yet powerful principle:LLM authorship can be inferred from the stability of a text’s stylistic patterns under meaning\-preserving manipulation\. Instead of depending on a single text instance or aggregated detection scores, LiSCP explicitly profiles stylistic behavior by generating multiple multimodal\-aligned paraphrased variants of the input \(ensuring consistency with accompanying multimodal context\) and measuring consistency across them\. Specifically, our method constructs a stylistic consistency profile that integrates: \(i\)discrete stylistic consistency signalsthat capture surface\-level invariances, and \(ii\)continuous semantic signalsthat quantify semantic alignment across variants, including implicit alignment with the underlying multimodal context\. This profiling perspective shifts the focus from fragile, wording\-specific artifacts to stability patterns that persist under paraphrasing and multimodal context adaptations, thereby aiming to addressRQ2\.

To answerRQ1, LiSCP is designed to be both lightweight and robust in real\-world detection deployment\. Final decisions are derived from a compact consistency profile, rather than large end\-to\-end models or detector\-specific heuristics, which improves efficiency and reduces dependence on any particular generator\. By aggregating stability signals over meaning\-preserving variants, LiSCP emphasizes transformation\-consistent signals that are difficult to remove through post\-editing or rewriting\. In summary, this work makes the following contributions:

- •Lightweight style profiling framework\.We propose theLiSCP, a framework that profiles stylistic consistency by aggregating stability signals across multimodal\-aligned paraphrased variants, robust against post\-edits, model shifts\.
- •Multi\-level stylistic\-semantic integration\.In this work, detection is reformulated asstylistic consistency inference under manipulation\. Our profile combines discrete stylistic and continuous semantic signals for stable patterns to enhance robustness and generalization across diverse scenarios\.
- •Empirical validation across challenging settings\.Across diverse domains and real\-world multimedia scenarios \(e\.g\., image\-text pair verification, video subtitle authentication\), LiSCP achieves state\-of\-the\-art performance and exhibits strong robustness against adversarial attacks and hybrid human\-AI compositions\.

## 2\.Related Works

#### Statistical and Model\-Based Detectors\.

Early efforts on machine\-generated text detection rely on statistical irregularities between human\-written and machine\-generated content, such as perplexity, entropy, or likelihood\-based measures\(Lavergneet al\.,[2008](https://arxiv.org/html/2605.05950#bib.bib1); Hashimotoet al\.,[2019](https://arxiv.org/html/2605.05950#bib.bib2); Gehrmannet al\.,[2019](https://arxiv.org/html/2605.05950#bib.bib3)\)\. These approaches identify anomalies at the token or sequence level, forming the foundation of many modern detectors\. However, as LLMs have become increasingly fluent, such surface\-level statistics are often insufficient to capture the subtle stylistic patterns exhibited by contemporary machine\-generated text\(Jawaharet al\.,[2020](https://arxiv.org/html/2605.05950#bib.bib4)\)\. Notably, in multimedia content scenarios, e\.g\., image\-text pairs, video subtitles, and multimodal reviews, these statistical methods face additional challenges: they fail to leverage cross\-modal semantic alignment cues and often degrade when text is paraphrased to adapt to multimedia context\(Yuet al\.,[2025a](https://arxiv.org/html/2605.05950#bib.bib62); Zhanget al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib63); Sadanandan and Behzadan,[2026](https://arxiv.org/html/2605.05950#bib.bib64)\)\. More recent work has explored supervised training\-based detectors, typically fine\-tuning large models to distinguish human\-written and machine\-generated textual content\(Yanget al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib25)\)\. Representative systems such as GPTZero and OpenAI’s classifier train RoBERTa\-style models on labeled corpora to learn discriminative representations\(Tian,[2023](https://arxiv.org/html/2605.05950#bib.bib20); Solaimanet al\.,[2019](https://arxiv.org/html/2605.05950#bib.bib21)\)\. While effective in controlled settings, these detectors frequently suffer from domain shift, limited cross\-model generalization, and sensitivity to post\-editing or rewriting\(Huet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib6); Maoet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib18); Daiet al\.,[2026](https://arxiv.org/html/2605.05950#bib.bib69)\)\. In contrast, our work avoids reliance on heavyweight classifiers or task\-specific supervision, instead focusing on stylistic stability and multimodal\-guided semantic consistency\.

#### Paraphrase and Perturbation\-Based Detection\.

To mitigate overfitting to a single text instance, several methods leverage perturbations or paraphrasing to improve robustness\. DetectGPT\(Mitchellet al\.,[2023](https://arxiv.org/html/2605.05950#bib.bib15)\)exploits likelihood curvature by comparing model scores before and after perturbations, based on the observation that LLM\-generated text often lies near local likelihood maxima and becomes unstable under controlled rewrites\. Fast\-DetectGPT\(Baoet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib7)\)improves efficiency through a more lightweight perturbation routine\. Other approaches, such as Binoculars\(Hanset al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib17)\)and BiScope\(Guoet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib8)\), contrast scores across different language models to enhance generalization beyond a single generator\. Despite their effectiveness, most perturbation\-based methods treat paraphrasing as a mechanism for score aggregation rather than a signal in its own right\(Shportko and Verbitsky,[2025](https://arxiv.org/html/2605.05950#bib.bib67); Mao and others,[2025](https://arxiv.org/html/2605.05950#bib.bib68); Liet al\.,[2025a](https://arxiv.org/html/2605.05950#bib.bib66),[b](https://arxiv.org/html/2605.05950#bib.bib48)\)\. As a result, they continue to rely on global likelihood statistics and offer limited insight into fine\-grained stylistic properties, especially in multimedia scenarios where text stylistic patterns are often constrained by paired visual content\(Cao,[2025](https://arxiv.org/html/2605.05950#bib.bib58)\)\. Our work differs fundamentally by explicitly modeling stylistic consistency across multimodal\-guided paraphrased variants, shifting the focus from scores to stability patterns that persist under adversarial manipulation and multimedia context adaptation\.

#### Robust Detection in Real\-World Settings\.

Recent studies have increasingly emphasized real\-world constraints such as hybrid human\-AI authorship, partial access to proprietary models, and streaming text scenarios\. PALD\(Leiet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib13)\)estimates the proportion of machine\-generated content at the sentence level, enabling partial authorship analysis in mixed documents\. GLIMPSE\(Baoet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib32)\)bridges white\-box and black\-box settings by reconstructing probability distributions from limited observations, improving robustness across proprietary models\. Additional work explores non\-parametric distribution comparison\(Songet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib16)\)or sequential hypothesis testing for online detection\(Chen and Wang,[2025](https://arxiv.org/html/2605.05950#bib.bib14)\)\. Existing robust detection methods improve applicability, but they often output a single global score, remain sensitive to paraphrasing or light editing, and overlook cross\-modal semantic constraints\. In contrast, our method explicitly profiles stylistic behavior under meaning\-preserving manipulations and multimodal alignment, enabling robust detection across domains and adversarial settings without model\-specific assumptions\.

## 3\.Lightweight Stylistic Consistency Profiling for LLM\-Generated Content Detection

In this section, we formally develop our lightweight framework for detecting LLM\-generated textual content, a critical task in multimedia content moderation \(e\.g\., verifying text authenticity in image\-text reviews, video subtitles, and multimodal academic papers\)\. We first define the paraphrase\-induced text space and stylistic stability, then introduce a multi\-level consistency profile derived from discrete and continuous feature mappings\. Based on this formulation, we present the detection rule and efficient algorithmic realization, which can be seamlessly integrated into real\-world multimedia moderation pipelines\.

### 3\.1\.Detection Problem Formulation and Paraphrase Space

Let𝒳\\mathcal\{X\}denote the space of all texts \(the core object of detection in multimedia content\)\. Given an input textx∈𝒳x\\in\\mathcal\{X\}\(often paired with visual/audio content in multimedia scenarios\), we aim to predict its authorship labely∈\{0,1\}y\\in\\\{0,1\\\}\(human vs\. LLM\-generated\)\.

We leverage transformation stability to distinguish authorship\. Under a predefined prompt set𝒫\\mathcal\{P\}, a multimodal\-guided paraphrasing operatorMIM\_\{I\}maps the input pair\(I,x\)\(I,x\)to a set of meaning\-preserving rewrites:

\(1\)𝒫​\(x;I\)=\{x^k∣x^k=MI​\(I,x,pk\),pk∈𝒫,k=1,…,K\}\.\\mathcal\{P\}\(x;I\)=\\\{\\hat\{x\}\_\{k\}\\mid\\hat\{x\}\_\{k\}=M\_\{I\}\(I,x,p\_\{k\}\),\\;p\_\{k\}\\in\\mathcal\{P\},\\;k=1,\\dots,K\\\}\.We enforce semantic preservation by filtering out variants with semantic similarity toxxbelow a thresholdδ\\delta, ensuring rewrites remain consistent with the original text’s core meaning, which is an essential property for text detection in multimedia content \(e\.g\., avoiding off\-topic rewrites in image captions\)\. Given the constructed paraphrase set\{x\}∪𝒫​\(x;I\)\\\{x\\\}\\cup\\mathcal\{P\}\(x;I\), the detection objective is to learn a stylistic consistency profile that captures invariant patterns across rewrites, enabling robust discrimination even in multimedia scenarios with noisy or manipulated text\.

### 3\.2\.Multi\-Level Stylistic Consistency Profiling

We formulate detection as an inference problem over transformation stability patterns\. Instead of analyzing a single text instance, we construct a structured profile that captures how stylistic signals behave across paraphrased variants\.

###### Definition 0 \(Stylistic Consistency Profile\)\.

A stylistic consistency profile is an aggregated vector representation

\(2\)v:𝒳→ℝd,v:\\mathcal\{X\}\\to\\mathbb\{R\}^\{d\},wherev​\(x\)v\(x\)is constructed from the stability measurements betweenxxand its paraphrased variantsx^∈𝒫​\(x;I\)\\hat\{x\}\\in\\mathcal\{P\}\(x;I\)\.

The profilev​\(x\)v\(x\)is constructed by integrating surface\-level discrete consistency signals and semantic\-level continuous consistency signals, as detailed below\.

![Refer to caption](https://arxiv.org/html/2605.05950v1/x1.png)Figure 1\.Overview of the LiSCP\.\(a\) Multimodal\-Guided Paraphrase Generation:Given an input pair\(I,x\)\(I,x\), a rewrite modelℳI\\mathcal\{M\}\_\{I\}generates semantically consistent variants\{x^i\}\\\{\\hat\{x\}\_\{i\}\\\}\.\(b\) Multi\-Level Stylistic Consistency Profiling:Discrete featuressD​\(x,x^\)s\_\{D\}\(x,\\hat\{x\}\)and continuous featuressC​\(x,x^\)s\_\{C\}\(x,\\hat\{x\}\)are extracted to capture linguistic stability\.\(c\) Consistency\-Based Detection:Features are aggregated intov​\(x\)v\(x\)and fed into a gradient\-boosted classifier to predict whetherxxis LLM\-generated\.#### Surface\-Level Discrete Consistency Profiling\.

For the tokenized textxx, its n\-gram set is defined as𝒩​\(n,x\)=\{\(wi,…,wi\+n−1\)\}i=1\|x\|−n\+1\\mathcal\{N\}\(n,x\)=\\\{\(w\_\{i\},\\dots,w\_\{i\+n\-1\}\)\\\}\_\{i=1\}^\{\|x\|\-n\+1\}\. The normalized n\-gram stability across a range\[n1,n2\]\[n\_\{1\},n\_\{2\}\]is:

\(3\)sN​\(x,x^\)=1n2−n1\+1​∑n=n1n2\|𝒩​\(n,x\)∩𝒩​\(n,x^\)\|\|𝒩​\(n,x\)∪𝒩​\(n,x^\)\|\.s\_\{N\}\(x,\\hat\{x\}\)=\\frac\{1\}\{n\_\{2\}\-n\_\{1\}\+1\}\\sum\_\{n=n\_\{1\}\}^\{n\_\{2\}\}\\frac\{\|\\mathcal\{N\}\(n,x\)\\cap\\mathcal\{N\}\(n,\\hat\{x\}\)\|\}\{\|\\mathcal\{N\}\(n,x\)\\cup\\mathcal\{N\}\(n,\\hat\{x\}\)\|\}\.To further capture fine\-grained lexical perturbations, we incorporate edit\-based consistency\. Let𝒟​\(x,x^\)\\mathcal\{D\}\(x,\\hat\{x\}\)denote the Levenshtein distance defined via dynamic programming\. We define the normalized edit stability as:

\(4\)sE​\(x,x^\)=1−𝒟​\(x,x^\)max⁡\(\|x\|,\|x^\|\)\.s\_\{E\}\(x,\\hat\{x\}\)=1\-\\frac\{\\mathcal\{D\}\(x,\\hat\{x\}\)\}\{\\max\(\|x\|,\|\\hat\{x\}\|\)\}\.Furthermore, the discrete consistency vector betweenxxandx^\\hat\{x\}is𝐬D​\(x,x^\)=\[sN​\(x,x^\),sE​\(x,x^\)\]⊤\\mathbf\{s\}\_\{D\}\(x,\\hat\{x\}\)=\\bigl\[s\_\{N\}\(x,\\hat\{x\}\),\\;s\_\{E\}\(x,\\hat\{x\}\)\\bigr\]^\{\\top\}, where human content typically exhibits lower stability \(flexible rewrites\) and LLM\-generated content exhibits higher stability \(structural invariance\)—a pattern that holds even in multimedia\-derived texts\.

#### Semantic\-Level Continuous Consistency Profiling\.

We capture continuous consistency using a shared text encoderξ\\xi\. Givenxxandx^\\hat\{x\}, their contextual embeddings areh​\(x\)=Pool​\(ξ​\(x\)\)h\(x\)=\\mathrm\{Pool\}\(\\xi\(x\)\)andh​\(x^\)=Pool​\(ξ​\(x^\)\)h\(\\hat\{x\}\)=\\mathrm\{Pool\}\(\\xi\(\\hat\{x\}\)\)\. We then compute a normalized angular consistency score

\(5\)sC​\(x,x^\)=1−1π​arccos⁡\(h​\(x\)⊤​h​\(x^\)‖h​\(x\)‖2​‖h​\(x^\)‖2\)\.s\_\{C\}\(x,\\hat\{x\}\)=1\-\\frac\{1\}\{\\pi\}\\arccos\\left\(\\frac\{h\(x\)^\{\\top\}h\(\\hat\{x\}\)\}\{\\\|h\(x\)\\\|\_\{2\}\\\|h\(\\hat\{x\}\)\\\|\_\{2\}\}\\right\)\.AlthoughsC​\(x,x^\)s\_\{C\}\(x,\\hat\{x\}\)is defined over textual representations, both paraphrasesx^\\hat\{x\}and the resulting consistency profile are obtained under the paired image contextIIthrough the multimodal\-guided rewrite process\. This measure captures semantic stability beyond surface\-level wording in multimedia scenarios\.

Given the demands of real\-time multimedia content moderation, this profile is engineered to be lightweight and compact, ensuring efficient inference\. For the textxxwith its paraphrase set𝒫​\(x;I\)\\mathcal\{P\}\(x;I\), the final consistency profile is derived by aggregating pairwise stylistic and semantic stability features:

\(6\)v​\(x\)=1\|𝒫​\(x;I\)\|​∑x^∈𝒫​\(x;I\)\[α⋅𝐬D​\(x,x^\)⊕β⋅sC​\(x,x^\)\],v\(x\)=\\frac\{1\}\{\|\\mathcal\{P\}\(x;I\)\|\}\\sum\_\{\\hat\{x\}\\in\\mathcal\{P\}\(x;I\)\}\\Bigl\[\\alpha\\cdot\\mathbf\{s\}\_\{D\}\(x,\\hat\{x\}\)\\;\\oplus\\;\\beta\\cdot s\_\{C\}\(x,\\hat\{x\}\)\\Bigr\],whereα,β\>0\\alpha,\\beta\>0are scaling coefficients and⊕\\oplusdenotes the feature vector concatenation\.

Algorithm 1Multimodal\-Guided Stylistic Consistency Detection1:Input: multimodal input pair

\(I,x\)\(I,x\), multimodal paraphraser

MIM\_\{I\}, prompt set

𝒫\\mathcal\{P\}, encoder

ξ\\xi, classifier

fθf\_\{\\theta\}
2:Parameters: number of paraphrases

KK, fusion weights

α,β\\alpha,\\beta, decision threshold

τ\\tau
3:Output: Predicted label

y^\\hat\{y\}
4:

𝒳p←∅,SN,SE,SC←∅\\mathcal\{X\}\_\{p\}\\leftarrow\\varnothing,\\kern 5\.0ptS\_\{N\},S\_\{E\},S\_\{C\}\\leftarrow\\varnothing
5:for

k=1,…,Kk=1,\\dots,Kdo

6:

x^k←MI​\(I,x,pk\)\\hat\{x\}\_\{k\}\\leftarrow M\_\{I\}\(I,x,p\_\{k\}\)
7:

𝒳p←𝒳p∪\{x^k\}\\mathcal\{X\}\_\{p\}\\leftarrow\\mathcal\{X\}\_\{p\}\\cup\\\{\\hat\{x\}\_\{k\}\\\}
8:endfor

9:foreach

x^∈𝒳p\\hat\{x\}\\in\\mathcal\{X\}\_\{p\}do

10:

sN←NgramStability​\(x,x^\)s\_\{N\}\\leftarrow\\textsc\{NgramStability\}\(x,\\hat\{x\}\)
11:

SN←SN∪\{sN\}S\_\{N\}\\leftarrow S\_\{N\}\\cup\\\{s\_\{N\}\\\}
12:

sE←EditStability​\(x,x^\)s\_\{E\}\\leftarrow\\textsc\{EditStability\}\(x,\\hat\{x\}\)
13:

SE←SE∪\{sE\}S\_\{E\}\\leftarrow S\_\{E\}\\cup\\\{s\_\{E\}\\\}
14:

𝐡x←ξ​\(x\),𝐡x^←ξ​\(x^\)\\mathbf\{h\}\_\{x\}\\leftarrow\\xi\(x\),\\kern 5\.0pt\\mathbf\{h\}\_\{\\hat\{x\}\}\\leftarrow\\xi\(\\hat\{x\}\)
15:

sC←SemanticConsistency​\(𝐡x,𝐡x^\)s\_\{C\}\\leftarrow\\textsc\{SemanticConsistency\}\(\\mathbf\{h\}\_\{x\},\\mathbf\{h\}\_\{\\hat\{x\}\}\)
16:

SC←SC∪\{sC\}S\_\{C\}\\leftarrow S\_\{C\}\\cup\\\{s\_\{C\}\\\}
17:endfor

18:

s¯N←1\|SN\|​∑s∈SNs\\bar\{s\}\_\{N\}\\leftarrow\\frac\{1\}\{\|S\_\{N\}\|\}\\sum\\limits\_\{s\\in S\_\{N\}\}s,s¯E←1\|SE\|​∑s∈SEs\\bar\{s\}\_\{E\}\\leftarrow\\frac\{1\}\{\|S\_\{E\}\|\}\\sum\\limits\_\{s\\in S\_\{E\}\}s,s¯C←1\|SC\|​∑s∈SCs\\bar\{s\}\_\{C\}\\leftarrow\\frac\{1\}\{\|S\_\{C\}\|\}\\sum\\limits\_\{s\\in S\_\{C\}\}s

19:

vD​\(x\)←\(s¯N,s¯E\)v\_\{D\}\(x\)\\leftarrow\(\\bar\{s\}\_\{N\},\\bar\{s\}\_\{E\}\)
20:

v​\(x\)←α⋅vD​\(x\)⊕β⋅s¯Cv\(x\)\\leftarrow\\alpha\\cdot v\_\{D\}\(x\)\\oplus\\beta\\cdot\\bar\{s\}\_\{C\}
21:

z←fθ​\(v​\(x\)\)z\\leftarrow f\_\{\\theta\}\(v\(x\)\),

y^←𝕀​\[σ​\(z\)≥τ\]\\hat\{y\}\\leftarrow\\mathbb\{I\}\[\\sigma\(z\)\\geq\\tau\]
22:return

y^\\hat\{y\}

### 3\.3\.Consistency\-based Detection Rule

Letv​\(x\)∈ℝdv\(x\)\\in\\mathbb\{R\}^\{d\}denote the aggregated consistency profile\. We model the conditional likelihood of authorship via a parametric decision functionfθ:ℝd→ℝf\_\{\\theta\}:\\mathbb\{R\}^\{d\}\\rightarrow\\mathbb\{R\}\. The detection score is defined as:z​\(x\)=fθ​\(v​\(x\)\)z\(x\)=f\_\{\\theta\}\(v\(x\)\), and the predicted label is obtained via thresholding:y^=𝕀​\[σ​\(z​\(x\)\)≥τ\]\\hat\{y\}=\\mathbb\{I\}\\\!\\left\[\\sigma\(z\(x\)\)\\geq\\tau\\right\], whereσ​\(⋅\)\\sigma\(\\cdot\)is the sigmoid function andτ\\tauis a fixed decision threshold\. The theoretical guarantee for the separability is as follows:

###### Theorem 2 \(Expected Separation of Stability Features\)\.

To analyze the separability of human\-written and LLM\-generated textual content, letv​\(x\)∈ℝdv\(x\)\\in\\mathbb\{R\}^\{d\}denote the stylistic consistency profile induced by an input textxxand its multimodal\-guided paraphrased variants\. Assume that there exists a constantϵ\>0\\epsilon\>0such that the class\-conditional expectations satisfy

\(7\)𝔼​\[v​\(x\)∣y=1\]−𝔼​\[v​\(x\)∣y=0\]⪰ϵ​1,\\mathbb\{E\}\[v\(x\)\\mid y=1\]\-\\mathbb\{E\}\[v\(x\)\\mid y=0\]\\succeq\\epsilon\\,\\mathbf\{1\},where𝟏∈ℝd\\mathbf\{1\}\\in\\mathbb\{R\}^\{d\}denotes the all\-ones vector,y∈\{0,1\}y\\in\\\{0,1\\\}is the binary authorship label, and⪰\\succeqdenotes element\-wise inequality\. Then there exists a linear scoring functionfθ​\(v\)=θ⊤​vf\_\{\\theta\}\(v\)=\\theta^\{\\top\}vthat achieves positive expected margin separation between the two classes\.

###### Proof\.

Letμ1=𝔼​\[v​\(x\)∣y=1\],μ0=𝔼​\[v​\(x\)∣y=0\]\\mu\_\{1\}=\\mathbb\{E\}\[v\(x\)\\mid y=1\],\\qquad\\mu\_\{0\}=\\mathbb\{E\}\[v\(x\)\\mid y=0\]denote the class\-conditional mean consistency profiles\. By assumption, we haveμ1−μ0⪰ϵ​1\.\\mu\_\{1\}\-\\mu\_\{0\}\\succeq\\epsilon\\,\\mathbf\{1\}\.Consider the linear scoring function defined byθ=𝟏∈ℝd\.\\theta=\\mathbf\{1\}\\in\\mathbb\{R\}^\{d\}\.Then

\(8\)θ⊤​μ1−θ⊤​μ0=∑k=1d\(μ1,k−μ0,k\)≥d​ϵ\>0\.\\theta^\{\\top\}\\mu\_\{1\}\-\\theta^\{\\top\}\\mu\_\{0\}=\\sum\_\{k=1\}^\{d\}\(\\mu\_\{1,k\}\-\\mu\_\{0,k\}\)\\geq d\\epsilon\>0\.Consequently, there exists a constantΔ=d​ϵ\>0\\Delta=d\\epsilon\>0such that

\(9\)𝔼​\[θ⊤​v​\(x\)∣y=1\]−𝔼​\[θ⊤​v​\(x\)∣y=0\]≥Δ\.\\mathbb\{E\}\[\\theta^\{\\top\}v\(x\)\\mid y=1\]\-\\mathbb\{E\}\[\\theta^\{\\top\}v\(x\)\\mid y=0\]\\geq\\Delta\.Therefore,fθ​\(v\)=θ⊤​vf\_\{\\theta\}\(v\)=\\theta^\{\\top\}vachieves a strictly positive expected separation margin between the two classes\. ∎

The theorem confirms our stylistic consistency profile’s discriminative capability for computationally constrained multimedia content moderation pipelines without heavy neural networks\.

Algorithm[1](https://arxiv.org/html/2605.05950#alg1)explicitly decouples text paraphrase set generation, stylistic stability measurement, and final decision making into three modular stages\. In the first stage, the paraphrasing component explores a local semantic neighborhood of the multimodal input pair\(I,x\)\(I,x\)by generating a paraphrase set𝒫​\(x;I\)=\{x^k\}k=1K\\mathcal\{P\}\(x;I\)=\\\{\\hat\{x\}\_\{k\}\\\}\_\{k=1\}^\{K\}, which provides multiple meaning\-preserving variants for subsequent analysis\. In the second stage, the consistency extraction module maps each original\-paraphrase pair\(x,x^k\)\(x,\\hat\{x\}\_\{k\}\)into a set of stability signals𝐬​\(x,x^k\)=\[𝐬D​\(x,x^k\),sC​\(x,x^k\)\]\\mathbf\{s\}\(x,\\hat\{x\}\_\{k\}\)=\\bigl\[\\mathbf\{s\}\_\{D\}\(x,\\hat\{x\}\_\{k\}\),\\,s\_\{C\}\(x,\\hat\{x\}\_\{k\}\)\\bigr\], capturing both surface\-level stylistic invariance and continuous semantic consistency\. In the final stage, all pairwise stability signals are aggregated into a fixed\-dimensional consistency profilev​\(x\)=1K​∑k=1K𝐬​\(x,x^k\)v\(x\)=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\mathbf\{s\}\(x,\\hat\{x\}\_\{k\}\), which is independent of the paraphrase countKKand thus enables efficient downstream inference without increasing model capacity\.

## 4\.Experiments

We evaluate the proposed LiSCP from five complementary perspectives: in\-domain detection performance, cross\-domain generalization, robustness under adversarial and hybrid settings, interpretability of the learned stability profile, and sensitivity to the choice of semantic encoder\.

### 4\.1\.Experimental Setup

#### Datasets\.

We evaluate LiSCP on datasets selected from two complementary perspectives: widely\-adopted benchmarks in the MGT detection literature for comparability with prior work, and datasets sourced from multimedia platforms to evaluate applicability in real\-world multimedia content moderation scenarios\. The former includes five conventional text domains and the large\-scaleRAIDbenchmark; the latter includesVisualNewsandMM\-IMDb, where textual content is inherently paired with visual media\.

News Domain \(Reuter News\)\.UsingChatGPT\(davinci\)\(Vermaet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib19)\), we generate LLM\-written news articles paired with human\-written news from theReuter\_50\_50dataset\(Houvardas and Stamatatos,[2006](https://arxiv.org/html/2605.05950#bib.bib41)\)\.

Essay Domain \(Student Essay\)\.Human\-written essays are sourced from IvyPanda\(IvyPanda,[2022](https://arxiv.org/html/2605.05950#bib.bib42)\), a repository of student\-written essays, while LLM\-generated essays are produced usingChatGPT\(Vermaet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib19)\)\.

Code Domain \(HumanEval Code\)\.We use theHumanEval Codedataset\(Maoet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib18)\)for human\-written code and useGPT\-3\.5\-Turboto generate machine\-written code\. This domain is included to test whether our method remains effective on highly structured content\.

Review Domain \(Yelp Review\)\.UsingGPT\-3\.5\-Turbo\(Maoet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib18)\), we generate LLM\-written reviews and pair them with human\-written reviews from Yelp\(Zhanget al\.,[2015](https://arxiv.org/html/2605.05950#bib.bib43)\)\.

Paper Abstract Domain \(Paper Abstract\)\.We sample 500 human\-written abstracts from ACL 2023, 2024 papers and useGPT\-3\.5\-Turboto generate LLM\-written paper abstracts\.

Visual News Domain \(VisualNews\)\.Human\-written articles are fromVisualNews\(Liuet al\.,[2021](https://arxiv.org/html/2605.05950#bib.bib46)\), collected from four major multimedia news outlets, with each article paired with a news image\. LLM\-generated counterparts are produced usingGPT\-3\.5\-Turbo\.

Movie Description Domain \(MM\-IMDb\)\.Human\-written plot descriptions are fromMM\-IMDb\(Arevaloet al\.,[2017](https://arxiv.org/html/2605.05950#bib.bib47)\), a multimodal benchmark pairing movie posters with editorial synopses, while LLM\-generated descriptions are produced usingGPT\-3\.5\-Turbo\.

RAIDBenchmark\.The officialRAIDbenchmark\(Duganet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib44)\)includes over 10 million documents from 11 LLMs, testing generalization across generators, decoding strategies, and attack conditions\.

#### Baselines\.

We compare the proposed LiSCP against several representative baseline detectors from multiple categories:

GPTZero\(Tian,[2023](https://arxiv.org/html/2605.05950#bib.bib20)\)\.GPTZero is a commercial classifier that relies on handcrafted features and shallow syntactic heuristics\. We use its official API for implementation\.

DetectGPT\(Mitchellet al\.,[2023](https://arxiv.org/html/2605.05950#bib.bib15)\)\.DetectGPT identifies LLM\-generated content by examining changes in the curvature of log\-probability under small input perturbations\.

Ghostbuster\(Vermaet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib19)\)\.Ghostbuster is a black\-box detector that enforces cross\-domain generalization by ensembling features from multiple weaker models\.

RAIDAR\(Maoet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib18)\)\.RAIDAR detects machine\-generated texts by rewriting the input and comparing the resulting differences to identify discrepancies between human and LLMs\.

Fast\-DetectGPT\(Baoet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib7)\)\.Fast\-DetectGPT is a more efficient zero\-shot detector than DetectGPT, which approximates probability curvature signals via conditional sampling\.

R\-Detect\(Songet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib16)\)\.R\-Detect applies a nonparametric kernel relative test to determine whether a test text is statistically closer to a human or a machine distribution\.

#### Implementation Details\.

We use AUROC as the primary metric for ranking quality, and we also report the best F1 score obtained by sweeping the decision threshold\. ForFast\-DetectGPT,Binoculars,R\-Detect, andDetectGPT, we use their official implementations but re\-evaluate them under a common protocol: AUROC is computed from raw scores, and the F1 score is obtained via threshold sweeping on the same split as our method\. ForGhostbuster, the original work relies onGPT\-AdaandGPT\-Davinci, which are now deprecated\. We replace them withGPT\-3\.5\-Turboas drop\-in substitutes, keeping all other hyperparameters unchanged\. ForDetectGPT, we follow the original paper and useT5\-3Bas the perturbation model\. All LLM calls are made in a batchified manner to control variance across methods\. For LiSCP, we useGPT\-3\.5\-Turboas the paraphrase modelℳ\\mathcal\{M\}and SBERT as the default encoderξ\\xi\. The classifierffis instantiated as a gradient\-boosted tree with early stopping based on validation AUROC\. To ensure fairness, F1 scores reported for all baselines in subsequent tables and figures are computed by threshold sweeping on the same held\-out validation splits, and RAID configurations follow\(Songet al\.,[2025](https://arxiv.org/html/2605.05950#bib.bib16)\)unless otherwise noted\.

Table 1\.Main detection AUROC of the LLM\-generated content acrossReuter News,HumanEval Code,Student Essay,Yelp Review,VisualNews, andMM\-IMDbdatasets\.Boldindicates the best performance\.MethodNewsCodeEssayYelpVisualNewsMM\-IMDbEntropy0\.42460\.43060\.48080\.46970\.47460\.4358Rank0\.65600\.53480\.68490\.68190\.54120\.5292LogRank0\.74380\.53500\.67580\.52940\.67120\.5068RoBERTa\-base0\.70240\.42170\.63170\.47230\.53580\.4076RoBERTa\-large0\.73010\.46920\.33250\.40610\.56740\.4371DetectGPT0\.82130\.52670\.64100\.63420\.81240\.6058Ghostbuster0\.64010\.53780\.57980\.66910\.71220\.5684Fast\-DetectGPT0\.94860\.66790\.92060\.62300\.91730\.6446RAIDAR0\.89560\.81730\.90910\.86160\.93120\.8246R\-Detect0\.98170\.64900\.76290\.71210\.96500\.7048LiSCP \(Ours\)0\.93560\.81080\.94550\.87180\.97460\.9576

### 4\.2\.Detection Performance

#### In\-Domain Detection Performance\.

We first examine the core detection performance of LiSCP on both conventional text domains and multimedia\-associated datasets\.[Table 1](https://arxiv.org/html/2605.05950#S4.T1)reports AUROC scores across six representative domains, spanning conventional text domains and multimedia content scenarios\. LiSCP achieves the best average performance and either outperforms or closely matches the strongest baseline in each individual domain\. Notably, LiSCP demonstrates strong performance in domains with structurally diverse content such asStudent EssayandHumanEval Code, significantly outperforming likelihood\-based detectors and supervised classifiers that rely on surface\-level statistics\. Furthermore, LiSCP achieves particularly strong results on multimedia content domains,VisualNewsandMM\-IMDb, outperforming all baselines by a clear margin\. This confirms the domain\-agnostic nature of stylistic consistency as a detection signal, extending naturally to multimedia content moderation without additional adaptation\.

#### Evaluation on the RAID Benchmark\.

Table 2\.Main detection AUROC on the RAID benchmark under six mixed data and adversarial attack configurations\.MethodMix1Mix2Mix3Att1Att2Att3DetectGPT0\.64370\.66320\.49870\.59310\.51110\.4554Ghostbuster0\.70130\.66430\.53880\.66450\.64650\.6356Fast\-DetectGPT0\.75960\.79010\.76200\.73240\.84100\.7129RAIDAR0\.80900\.68750\.65000\.78760\.64760\.7112R\-Detect0\.86430\.76560\.76500\.78550\.78290\.7163LiSCP \(Ours\)0\.89570\.79580\.78130\.82680\.77140\.7608

We further evaluate our method on the RAID benchmark, which consists of multiple generators, genres, decoding strategies, and adversarial attacks\. Following prior work\(Duganet al\.,[2024](https://arxiv.org/html/2605.05950#bib.bib44)\), we test the model’s ability to generalize across unseen generator\-attack configurations\. As shown in[Table 2](https://arxiv.org/html/2605.05950#S4.T2), LiSCP performs competitively with the strongest baselines in clean settings and often surpasses them when evaluated on attacked configurations\. In particular, detectors that rely heavily on raw likelihoods experience a sharp performance drop under paraphrasing and corruption\. In contrast, our stability\-based approach remains effective and provides informative signals even in the presence of such adversarial perturbations\.

### 4\.3\.Cross\-Domain Generalization

Table 3\.Cross\-domain generalization F1 score on ID\-OOD splits\. Each row group represents a source domain \(used for training\), while each column shows the target domain used for evaluation\.OOD\-Avgdenotes the average F1 score on out\-of\-domain targets\.Boldindicates the best performance, andunderlineindicates the second best\.Paper AbstractHumanEval CodeMethodPaper AbstractHumanEvalStudent EssayReuter NewsYelp ReviewPaper AbstractHumanEvalStudent EssayReuter NewsYelp ReviewEntropy56\.75±\\pm0\.98428\.14±\\pm0\.71936\.52±\\pm0\.50748\.14±\\pm0\.81341\.83±\\pm1\.23126\.10±\\pm0\.75358\.23±\\pm1\.09335\.68±\\pm0\.72436\.17±\\pm0\.14840\.22±\\pm1\.062Rank54\.60±\\pm0\.21830\.97±\\pm0\.41335\.62±\\pm0\.70638\.79±\\pm1\.03852\.63±\\pm1\.07135\.11±\\pm1\.20248\.51±\\pm0\.30840\.10±\\pm0\.65028\.31±\\pm0\.82842\.23±\\pm0\.391LogRank64\.05±\\pm0\.55429\.10±\\pm0\.63243\.71±\\pm0\.75242\.35±\\pm1\.03651\.72±\\pm0\.95342\.89±\\pm0\.86552\.09±\\pm0\.73333\.67±\\pm1\.03239\.16±\\pm1\.02635\.31±\\pm0\.973RoBERTa\-base57\.78±\\pm0\.86738\.72±\\pm0\.84435\.91±\\pm0\.59248\.02±\\pm0\.50856\.78±\\pm1\.21539\.01±\\pm0\.29546\.74±\\pm1\.05935\.10±\\pm0\.52039\.80±\\pm0\.76947\.90±\\pm0\.695RoBERTa\-large63\.40±\\pm0\.85041\.48±\\pm0\.92150\.81±\\pm0\.90555\.27±\\pm0\.59661\.82±\\pm1\.01039\.15±\\pm0\.93558\.10±\\pm0\.99547\.92±\\pm1\.00339\.58±\\pm1\.25652\.09±\\pm0\.958DetectGPT83\.33±\\pm0\.98333\.64±\\pm1\.27153\.96±\\pm0\.59264\.14±\\pm0\.91559\.24±\\pm1\.03741\.39±\\pm0\.85456\.23±\\pm0\.47069\.55±\\pm0\.79246\.95±\\pm0\.84947\.05±\\pm0\.384Ghostbuster74\.52±\\pm0\.85031\.20±\\pm0\.51940\.31±\\pm1\.24862\.75±\\pm0\.49764\.81±\\pm0\.58839\.65±\\pm0\.55061\.27±\\pm0\.68171\.90±\\pm0\.67946\.71±\\pm1\.26051\.38±\\pm1\.054Fast\-DetectGPT86\.50±\\pm1\.32036\.48±\\pm0\.94256\.07±\\pm1\.41965\.12±\\pm0\.92063\.80±\\pm1\.16145\.05±\\pm1\.00762\.48±\\pm1\.25670\.37±\\pm1\.31058\.45±\\pm0\.85263\.40±\\pm0\.956RAIDAR78\.34±\\pm0\.52637\.05±\\pm0\.54353\.86±\\pm0\.65273\.62±\\pm1\.05960\.85±\\pm0\.47846\.28±\\pm0\.62875\.01±\\pm0\.34347\.24±\\pm0\.93859\.42±\\pm0\.60979\.51±\\pm1\.203R\-Detect85\.94±\\pm0\.95040\.53±\\pm1\.03454\.13±\\pm1\.19264\.05±\\pm0\.80462\.16±\\pm0\.62142\.15±\\pm0\.45276\.50±\\pm0\.92569\.72±\\pm0\.80656\.10±\\pm0\.76364\.38±\\pm0\.629LiSCP \(Ours\)91\.54±\\pm0\.34045\.14±\\pm0\.46267\.16±\\pm0\.21566\.56±\\pm0\.40970\.09±\\pm0\.31746\.43±\\pm0\.68383\.33±\\pm0\.42970\.79±\\pm0\.15562\.70±\\pm0\.51767\.48±\\pm0\.438Reuter NewsYelp ReviewPaper AbstractHumanEvalStudent EssayReuter NewsYelp ReviewPaper AbstractHumanEvalStudent EssayReuter NewsYelp ReviewEntropy36\.08±\\pm0\.84539\.51±\\pm0\.57323\.58±\\pm0\.79446\.70±\\pm0\.44038\.95±\\pm0\.12536\.06±\\pm0\.60727\.85±\\pm1\.19437\.80±\\pm1\.53837\.13±\\pm0\.78948\.10±\\pm0\.949Rank42\.79±\\pm1\.45431\.06±\\pm0\.64439\.18±\\pm1\.36549\.40±\\pm1\.57347\.80±\\pm0\.76032\.88±\\pm0\.41349\.87±\\pm0\.65440\.56±\\pm0\.74335\.63±\\pm1\.18240\.02±\\pm0\.119LogRank49\.07±\\pm1\.12745\.97±\\pm0\.19848\.33±\\pm1\.10356\.15±\\pm0\.32746\.36±\\pm0\.70929\.62±\\pm0\.98344\.50±\\pm0\.35232\.01±\\pm0\.91534\.22±\\pm0\.38753\.25±\\pm0\.726RoBERTa\-base50\.27±\\pm1\.16041\.50±\\pm0\.58236\.16±\\pm0\.62854\.59±\\pm1\.09431\.29±\\pm1\.06943\.76±\\pm0\.53950\.66±\\pm1\.46645\.30±\\pm1\.42345\.64±\\pm1\.30662\.62±\\pm1\.416RoBERTa\-large57\.32±\\pm1\.35648\.03±\\pm1\.75245\.71±\\pm1\.72063\.25±\\pm1\.34941\.63±\\pm1\.21848\.13±\\pm0\.72051\.18±\\pm1\.20632\.36±\\pm1\.02359\.31±\\pm1\.66270\.23±\\pm1\.230DetectGPT52\.57±\\pm0\.76532\.19±\\pm0\.84656\.44±\\pm1\.36386\.70±\\pm0\.83837\.09±\\pm0\.72841\.63±\\pm0\.74537\.51±\\pm0\.54559\.32±\\pm1\.15061\.08±\\pm1\.34468\.10±\\pm1\.223Ghostbuster58\.71±\\pm1\.38839\.02±\\pm0\.41160\.61±\\pm0\.59682\.56±\\pm0\.38641\.30±\\pm1\.74940\.12±\\pm0\.31332\.10±\\pm0\.26757\.42±\\pm1\.45144\.28±\\pm0\.65166\.34±\\pm0\.268Fast\-DetectGPT62\.75±\\pm1\.10745\.73±\\pm1\.21668\.14±\\pm0\.82181\.73±\\pm0\.59345\.18±\\pm0\.81846\.60±\\pm0\.98850\.22±\\pm1\.36166\.80±\\pm0\.86362\.08±\\pm0\.39068\.03±\\pm1\.103RAIDAR60\.37±\\pm0\.44133\.07±\\pm1\.71159\.74±\\pm0\.30089\.67±\\pm0\.37232\.23±\\pm0\.75943\.09±\\pm1\.73842\.31±\\pm1\.79567\.79±\\pm1\.65147\.36±\\pm0\.39771\.88±\\pm0\.227R\-Detect64\.15±\\pm0\.81845\.42±\\pm1\.19871\.09±\\pm1\.03177\.42±\\pm1\.06746\.31±\\pm0\.99244\.54±\\pm0\.95852\.01±\\pm0\.81067\.51±\\pm0\.60563\.15±\\pm0\.96772\.40±\\pm0\.585LiSCP \(Ours\)69\.81±\\pm0\.33646\.35±\\pm0\.91482\.61±\\pm0\.84788\.21±\\pm0\.84549\.17±\\pm0\.55846\.67±\\pm0\.67253\.43±\\pm0\.63469\.03±\\pm0\.31669\.88±\\pm0\.53880\.83±\\pm0\.305Student EssayOOD\-AvgPaper AbstractHumanEvalStudent EssayReuter NewsYelp ReviewPaper AbstractHumanEvalStudent EssayReuter NewsYelp ReviewEntropy48\.15±\\pm0\.56526\.93±\\pm1\.39736\.68±\\pm0\.75341\.05±\\pm1\.10231\.16±\\pm0\.76640\.63±\\pm0\.75136\.13±\\pm0\.99534\.05±\\pm0\.86341\.84±\\pm0\.65840\.05±\\pm0\.827Rank51\.23±\\pm0\.87526\.18±\\pm0\.54051\.75±\\pm0\.31043\.22±\\pm0\.28245\.61±\\pm0\.47243\.32±\\pm0\.83237\.32±\\pm0\.51241\.44±\\pm0\.75539\.07±\\pm0\.98145\.66±\\pm0\.563LogRank48\.86±\\pm0\.52731\.23±\\pm1\.20259\.10±\\pm1\.08046\.77±\\pm1\.80041\.67±\\pm0\.31946\.90±\\pm0\.81140\.58±\\pm0\.62343\.36±\\pm0\.97643\.73±\\pm0\.91545\.66±\\pm0\.736RoBERTa\-base58\.55±\\pm0\.55532\.32±\\pm1\.37945\.14±\\pm0\.66334\.26±\\pm0\.38050\.52±\\pm0\.84249\.87±\\pm0\.68341\.99±\\pm1\.06639\.52±\\pm0\.76544\.46±\\pm0\.81149\.82±\\pm1\.047RoBERTa\-large53\.10±\\pm1\.14528\.66±\\pm1\.06661\.93±\\pm0\.58337\.82±\\pm1\.80647\.04±\\pm0\.73652\.22±\\pm1\.00145\.49±\\pm1\.18847\.75±\\pm1\.04751\.05±\\pm1\.33454\.56±\\pm1\.030DetectGPT50\.41±\\pm0\.97631\.40±\\pm0\.76170\.56±\\pm1\.59764\.34±\\pm1\.04341\.54±\\pm0\.59053\.87±\\pm0\.86538\.19±\\pm0\.77961\.97±\\pm1\.09964\.64±\\pm0\.99850\.60±\\pm0\.792Ghostbuster63\.97±\\pm0\.40435\.63±\\pm1\.33071\.47±\\pm1\.50370\.18±\\pm0\.87438\.73±\\pm0\.35755\.39±\\pm0\.70139\.84±\\pm0\.64260\.34±\\pm1\.09561\.30±\\pm0\.73452\.51±\\pm0\.803Fast\-DetectGPT55\.29±\\pm1\.10734\.16±\\pm0\.98373\.52±\\pm1\.21066\.59±\\pm0\.79250\.06±\\pm0\.38559\.24±\\pm1\.10645\.81±\\pm1\.15266\.98±\\pm1\.12566\.79±\\pm0\.70958\.09±\\pm0\.885RAIDAR64\.75±\\pm1\.43234\.69±\\pm0\.66072\.02±\\pm0\.99474\.88±\\pm1\.54636\.97±\\pm0\.32358\.57±\\pm0\.95344\.43±\\pm1\.01060\.13±\\pm0\.90768\.99±\\pm0\.79756\.29±\\pm0\.598R\-Detect66\.89±\\pm0\.27133\.06±\\pm0\.93976\.44±\\pm1\.09572\.33±\\pm0\.42649\.05±\\pm0\.71460\.73±\\pm0\.69049\.50±\\pm0\.98167\.78±\\pm0\.94666\.61±\\pm0\.80658\.86±\\pm0\.708LiSCP \(Ours\)71\.31±\\pm0\.89238\.52±\\pm0\.33489\.27±\\pm1\.53276\.07±\\pm0\.93053\.15±\\pm0\.57065\.15±\\pm0\.58553\.35±\\pm0\.55575\.77±\\pm0\.61372\.68±\\pm0\.64864\.14±\\pm0\.438

While the in\-domain results demonstrate strong discriminative ability, practical deployment also requires detectors to transfer across domains with different writing styles and content distributions\. We therefore next evaluate cross\-domain generalization\. We perform experiments using ID\-OOD splits, where detectors are trained on a source domain and evaluated on unseen target domains\. As summarized in[Table 3](https://arxiv.org/html/2605.05950#S4.T3), the results show that LiSCP consistently achieves the highest OOD\-Avg F1 score across all source domains, highlighting its strong ability to generalize to new, unseen domains\. When trained on formal domains such asPaper AbstractandStudent Essay, our method generalizes well to informal targets likeReuter NewsandYelp Review, demonstrating its versatility across different writing styles\. Similarly, when trained on casual domains likeYelp Review, it maintains stable performance even when tested on more formal domains likePaper AbstractandStudent Essay\. Although cross\-domain detection remains a challenging task for all methods, the proposed stability\-based approach consistently outperforms other methods, showing stronger transferability under domain shifts\. Notably, we observe that machine\-generated content consistently exhibits higher mean values than human\-written content across each component of the consistency profile\. This observation aligns with the consistency dominance assumption stated in[Theorem 2](https://arxiv.org/html/2605.05950#S3.Thmtheorem2), which posits a coordinate\-wise separation of class\-conditional expectations in the consistency profile space\.

### 4\.4\.Robustness Analysis

Beyond domain transfer, a robust detector should also remain reliable under post\-editing and mixed\-authorship settings\. We therefore further evaluate LiSCP under adversarial perturbations and hybrid human\-LLM composition\.

#### Robustness to Adversarial Manipulation\.

![Refer to caption](https://arxiv.org/html/2605.05950v1/x2.png)Figure 2\.Evaluation of detection performance degradation under adversarial perturbations: Comparative analysis of original AUROC, AUROC after perturbation, and relative drop rate acrossReuter News,HumanEval Code,Student Essay,Yelp Review,VisualNews, andMM\-IMDbdatasets, highlighting robustness under word\-level attacks\.To assess robustness against meaning\-preserving edits, we adopt TextAttack\-style perturbations and introduce character swaps/insertions, synonym\-level word substitutions, and sentence\-level paraphrases, with a maximum modification rate of 20% tokens per sample\. All detectors are trained on clean data and tested directly on perturbed sets without adaptation, reflecting real\-world deployment where edited or partially rewritten text is common\. As illustrated in[Figure 2](https://arxiv.org/html/2605.05950#S4.F2), we report original AUROC, post\-perturbation AUROC, and the relative performance drop\. Across domains, LiSCP consistently incurs substantially smaller degradation than likelihood\-based or probability–dependent baselines\. Notably, detectors such as DetectGPT and Ghostbuster exhibit pronounced drops, especially inYelp ReviewandHumanEval Code, where perturbations disrupt probability curvature or token statistics, whereas LiSCP maintains high accuracy with only minor fluctuations\. This pattern holds consistently across both conventional text domains and multimedia content domains \(VisualNewsandMM\-IMDb\), demonstrating that stylistic consistency remains a reliable signal regardless of content modality\. This addresses our goal of robust detection \(RQ1\), confirming the resilience of stylistic consistency under adversarial manipulation\.

#### Robustness to Hybrid Human\-LLM Composition\.

![Refer to caption](https://arxiv.org/html/2605.05950v1/x3.png)Figure 3\.Evaluation of detection performance degradation under adversarial mixed text: Comparative analysis of original AUROC, AUROC after mixing, and relative performance drop acrossReuter News,HumanEval Code,Student Essay,Yelp Review,VisualNewsandMM\-IMDbdatasets, characterizing stability under hybrid human–LLM composition\.Beyond local perturbations, we further evaluate robustness under global content mixing, where human\-written and LLM\-generated segments are interleaved within a single document\. Following standard protocol, we construct hybrid samples by concatenating segments at a 4:1 ratio and assign labels based on dominant authorship\. This setup reflects situations where users revise LLM outputs or insert generated paragraphs into human writing\. The results in[Figure 3](https://arxiv.org/html/2605.05950#S4.F3)show that LiSCP continues to outperform all baselines on mixed inputs and maintains the smallest AUROC drop across domains\. While detectors relying on perplexity or representation distance degrade significantly under blending, often losing the authorship signal once machine spans are surrounded by human context, our stylistic consistency profile remains discriminative even without span\-level supervision\. For example, onYelp Review, most baselines experience severe degradation, while LiSCP drops only modestly, indicating strong resilience to human–AI hybridization\.

Combined with perturbation experiments, this confirms that LiSCP supports detection not only under local edits but also under mixed scenarios, addressingRQ2and highlighting its practicality across diverse real\-world multimedia content moderation scenarios\.

### 4\.5\.Interpretability Analysis

Besides the robustness of LiSCP, we next analyze whether the learned stability profile also yields an interpretable feature\-space structure for more transparency\.

#### Visualizing Explainability through UMAP

![Refer to caption](https://arxiv.org/html/2605.05950v1/x4.png)Figure 4\.Explainable in\-domain detection visualization: UMAP projections of stability signatures separating human\-written and LLM\-generated textual content across six domains \(Yelp Review,Reuter News,HumanEval Code,Student Essay,VisualNews, andMM\-IMDb\)\.To highlight the explainability of our proposed method, LiSCP, we utilize UMAP\(McInneset al\.,[2020](https://arxiv.org/html/2605.05950#bib.bib45)\)to visualize the distribution of feature vectors extracted from various datasets\. As shown in[Figure 4](https://arxiv.org/html/2605.05950#S4.F4), UMAP provides a two\-dimensional projection that facilitates the analysis of high\-dimensional data, offering a feature\-space sanity check on how our stability profile organizes texts from different sources rather than serving as a decision tool\. As shown in[Figure 4](https://arxiv.org/html/2605.05950#S4.F4), the red points represent machine\-generated content and the green points correspond to human\-written content\. Across all domains, the two groups are clearly separated, with limited overlap in the projected space\. Notably,Student Essayexhibits a particularly clean separation, indicating that LiSCP captures style\-based differences effectively even in relatively complex content domains\. This pattern also extends to multimedia content domains:MM\-IMDbandVisualNewsboth display clear cluster boundaries, suggesting that the learned consistency signatures remain informative across heterogeneous settings and supporting the generalizability of LiSCP in real\-world content moderation scenarios\.

#### Quantitative Evaluation

![Refer to caption](https://arxiv.org/html/2605.05950v1/x5.png)Figure 5\.Explainability experiments using KL Divergence and Hellinger Distance for feature and classifier\-based methods across multiple datasetsWhile the UMAP visualization in the previous section provided a qualitative view of the feature space separation, we quantify the model’s ability to distinguish between human and machine content using two distribution divergence metrics: KL Divergence and Hellinger Distance\. These metrics measure the divergence between the distributions of human\-written and LLM\-generated textual content, providing a deeper understanding of how the model distinguishes between the two\. The results are presented in[Figure 5](https://arxiv.org/html/2605.05950#S4.F5), where two methods are used: the first derives distributions from the featurev​\(x\)v\(x\)extracted by LiSCP, and the second uses classifier prediction scoresσ​\(y^\)\\sigma\(\\hat\{y\}\)\. The results consistently show that the classifier\-based method \(usingσ​\(y^\)\\sigma\(\\hat\{y\}\)\) outperforms the feature\-based method \(usingv​\(x\)v\(x\)\) in both KL Divergence and Hellinger Distance across all datasets\. In particular, the KL Divergence is significantly higher for the distributionσ​\(y^\)\\sigma\(\\hat\{y\}\), especially in domains likeYelp Review, suggesting that the classifier captures more distinct stylistic differences\. Similarly, higher Hellinger Distance values for the distributionσ​\(y^\)\\sigma\(\\hat\{y\}\)indicate clearer separability between human\-written and LLM\-generated textual content\. This gap betweenv​\(x\)v\(x\)andσ​\(y^\)\\sigma\(\\hat\{y\}\)is expected: the stability profile provides a compact, interpretable representation, while the classifierfffurther amplifies the separation by learning non\-linear decision boundaries\. Together, they support both transparent feature\-space inspection and strong end\-to\-end detection performance\. These findings demonstrate that LiSCP not only achieves high detection accuracy but also enhances transparency and explainability\.

### 4\.6\.Ablation Study

Finally, to understand how much the framework depends on the specific continuous representation module, we conduct an ablation study over different semantic encoders\. Specifically, we conduct a component ablation study by replacing the continuous representation moduleξ\\xiwith different feature extractors ranging from lightweight statistical vectors to deep contextual encoders\. As shown in[Table 4](https://arxiv.org/html/2605.05950#S4.T4), LiSCP maintains consistently high performance across encoder choices\. This evaluates the plug\-and\-play compatibility of LiSCP and verifies that the stylistic consistency profiling remains effective when using weaker or stronger semantic embeddings\.

Interestingly, even when using TF\-IDF, a purely statistical and non\-contextual representation without deep semantics, our method still delivers competitive results\. Replacing TF\-IDF with pretrained distributed embeddings \(Word2Vec/GloVe\), LiSCP obtains improved results through richer lexical features, suggesting that capturing global lexical semantics benefits profile construction\. Among these semantic encoders, Contextual encoders \(BERT, SBERT\) are particularly effective, with SBERT providing the best overall average AUROC when plugged into LiSCP\. These results confirm that our mechanism is encoder\-agnostic, providing flexibility for various deployment scenarios\.

Table 4\.AUROC results of replacing the semantic encoder with different representations across domains\. LiSCP remains effective across feature extractors, confirming the plug\-and\-play capability\.EncoderReuter NewsHumanEvalEssayYelp ReviewTF\-IDF0\.86830\.77090\.92850\.8114Word2Vec/GloVe0\.88640\.80870\.93690\.8309BERT0\.93000\.78120\.94780\.8566SBERT \(Default\)0\.93560\.81080\.94550\.8718

## 5\.Conclusion

In this work, we proposed LiSCP, a lightweight framework for detecting LLM\-generated textual content through stylistic consistency profiling\. By combining discrete stylistic signals with continuous semantic consistency, LiSCP provides a compact representation that remains effective under paraphrase\-based variation and multimodal\-guided rewriting\. Experiments across conventional and multimedia\-associated domains show that LiSCP achieves strong detection performance and robust behavior under adversarial perturbations and hybrid human\-LLM composition\. It also consistently outperforms existing methods in cross\-domain evaluations, demonstrating strong generalization under distribution shift\. Further analyses show that LiSCP remains effective across different semantic encoders, supporting flexible deployment\.

## References

- S\. Abdelnabi and M\. Fritz \(2021\)Adversarial watermarking transformer: towards tracing text provenance with data hiding\.In2021 IEEE Symposium on Security and Privacy \(SP\),pp\. 121–140\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- J\. Arevalo, T\. Solorio, M\. Montes\-y\-Gómez, and F\. A\. González \(2017\)Gated multimodal units for information fusion\.arXiv preprint arXiv:1702\.01992\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p8.1)\.
- G\. Bao, Y\. Zhao, J\. He, and Y\. Zhang \(2025\)Glimpse: enabling white\-box methods to use proprietary models for zero\-shot llm\-generated text detection\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1),[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px3.p1.1)\.
- G\. Bao, Y\. Zhao, Z\. Teng, L\. Yang, and Y\. Zhang \(2024\)Fast\-detectgpt: efficient zero\-shot detection of machine\-generated text via conditional probability curvature\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px2.p6.1.1)\.
- L\. Cao \(2025\)A practical synthesis of detecting ai\-generated textual, visual, and audio content\.arXiv preprint arXiv:2504\.02898\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1),[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1)\.
- B\. Chen, G\. Li, J\. Wu, J\. Li, M\. Chen, and J\. Wang \(2026\)AgentChain: blockchain\-empowered multi\-agent coordination for trustworthy llm question\-answering systems\.IEEE Transactions on Dependable and Secure Computing\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- C\. Chen and J\. Wang \(2025\)Online detection of llm\-generated texts via sequential hypothesis testing by betting\.InInternational Conference on Machine Learning,pp\. 9231–9276\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1),[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px3.p1.1)\.
- Z\. Cheng, L\. Zhou, F\. Jiang, B\. Wang, and H\. Li \(2025\)Beyond binary: towards fine\-grained llm\-generated text detection via role recognition and involvement measurement\.InProceedings of the ACM on Web Conference 2025,pp\. 2677–2688\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- F\. Dai, X\. Jiang, and Z\. Deng \(2026\)HLPD: aligning llms to human language preference for machine\-revised text detection\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 30440–30448\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan,et al\.\(2024\)The llama 3 herd of models\.arXiv e\-prints,pp\. arXiv–2407\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- L\. Dugan, A\. Hwang, F\. Trhlik, J\. M\. Ludan, A\. Zhu, H\. Xu, D\. Ippolito, and C\. Callison\-Burch \(2024\)Raid: a shared benchmark for robust evaluation of machine\-generated text detectors\.arXiv preprint arXiv:2405\.07940\.Cited by:[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p9.1),[§4\.2](https://arxiv.org/html/2605.05950#S4.SS2.SSS0.Px2.p1.1)\.
- S\. Gehrmann, H\. Strobelt, and A\. M\. Rush \(2019\)GLTR: statistical detection and visualization of generated text\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations,pp\. 111–116\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1),[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Guo, S\. Cheng, X\. Jin, Z\. Zhang, K\. Zhang, G\. Tao, G\. Shen, and X\. Zhang \(2024\)Biscope: ai\-generated text detection by checking memorization of preceding tokens\.Advances in Neural Information Processing Systems37,pp\. 104065–104090\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1)\.
- R\. Guo, J\. Wei, L\. Sun, B\. Yu, G\. Chang, D\. Liu, S\. Zhang, Z\. Yao, M\. Xu, and L\. Bu \(2023\)A survey on image\-text multimodal models\.arXiv preprint arXiv:2309\.15857\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- A\. Hans, A\. Schwarzschild, V\. Cherepanova, H\. Kazemi, A\. Saha, M\. Goldblum, J\. Geiping, and T\. Goldstein \(2024\)Spotting llms with binoculars: zero\-shot detection of machine\-generated text\.InInternational Conference on Machine Learning,pp\. 17519–17537\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1),[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1)\.
- T\. B\. Hashimoto, H\. Zhang, and P\. Liang \(2019\)Unifying human and statistical evaluation for natural language generation\.arXiv preprint arXiv:1904\.02792\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Houvardas and E\. Stamatatos \(2006\)N\-gram feature selection for authorship identification\.InInternational conference on artificial intelligence: Methodology, systems, and applications,pp\. 77–86\.Cited by:[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p2.1)\.
- X\. Hu, P\. Chen, and T\. Ho \(2024\)Radar: robust ai\-text detection via adversarial learning\.Advances in Neural Information Processing Systems36\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.\(2024\)Gpt\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- IvyPanda \(2022\)IvyPanda essays dataset\.Note:Hugging FaceExternal Links:[Link](https://huggingface.co/datasets/qwedsacf/ivypanda-essays)Cited by:[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p3.1)\.
- G\. Jawahar, M\. Abdul\-Mageed, and V\. Laks Lakshmanan \(2020\)Automatic detection of machine generated text: a critical survey\.InProceedings of the 28th International Conference on Computational Linguistics,pp\. 2296–2309\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Kirchenbauer, J\. Geiping, Y\. Wen, J\. Katz, I\. Miers, and T\. Goldstein \(2023\)A watermark for large language models\.InInternational Conference on Machine Learning,pp\. 17061–17084\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- R\. Koike, M\. Kaneko, A\. Niwa, P\. Nakov, and N\. Okazaki \(2025\)ExaGPT: example\-based machine\-generated text detection for human interpretability\.arXiv preprint arXiv:2502\.11336\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- K\. Krishna, Y\. Song, M\. Karpinska, J\. Wieting, and M\. Iyyer \(2024\)Paraphrasing evades detectors of ai\-generated text, but retrieval is an effective defense\.Advances in Neural Information Processing Systems36\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- T\. Lavergne, T\. Urvoy, and F\. Yvon \(2008\)Detecting fake content with relative entropy scoring\.InProceedings of the 2008 International Conference on Uncovering Plagiarism, Authorship and Social Software Misuse\-Volume 377,pp\. 27–31\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Lei, H\. Hsu, and C\. Chen \(2025\)PaLD: detection of text partially written by large language models\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1),[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Li, X\. Lin, G\. Li, Z\. Liu, A\. Wulianghai, L\. Ding, J\. Wu, and J\. Li \(2026a\)Model\-agnostic sentiment distribution stability analysis for robust llm\-generated texts detection\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 35608–35616\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- S\. Li, X\. Lin, Y\. Liu, and J\. Li \(2024\)Trustworthy ai\-generative content in intelligent 6g network: adversarial, privacy, and fairness\.arXiv preprint arXiv:2405\.05930\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- S\. Li, A\. Wulianghai, G\. Li, X\. Lin, Q\. Mao, Y\. Chen, J\. Wu, and J\. Li \(2026b\)DSIPA: detecting llm\-generated texts via sentiment\-invariant patterns divergence analysis\.arXiv preprint arXiv:2604\.26328\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- S\. Li, A\. Wulianghai, X\. Lin, G\. Li, X\. Chen, J\. Wu, and J\. Li \(2025a\)StyleDecipher: robust and explainable detection of llm\-generated texts with stylistic analysis\.arXiv preprint arXiv:2510\.12608\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1)\.
- X\. Li, Z\. Yin, H\. Tan, S\. Jing, D\. Su, Y\. Cheng, H\. Shen, and F\. Sun \(2025b\)PRDetect: perturbation\-robust llm\-generated text detection based on syntax tree\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 8290–8301\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1),[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1)\.
- F\. Liu, Y\. Wang, T\. Wang, and V\. Ordonez \(2021\)Visual news: benchmark and challenges in news image captioning\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 6761–6771\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.542/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.542)Cited by:[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p7.1)\.
- C\. Maoet al\.\(2025\)Learning to rewrite: generalized llm\-generated text detection\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5897–5912\.External Links:[Link](https://aclanthology.org/2025.acl-long.322/)Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1)\.
- C\. Mao, C\. Vondrick, H\. Wang, and J\. Yang \(2024\)Detecting generated text via rewriting\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p4.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p5.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px2.p5.1.1)\.
- L\. McInnes, J\. Healy, and J\. Melville \(2020\)UMAP: uniform manifold approximation and projection for dimension reduction\.External Links:1802\.03426,[Link](https://arxiv.org/abs/1802.03426)Cited by:[§4\.5](https://arxiv.org/html/2605.05950#S4.SS5.SSS0.Px1.p1.1)\.
- E\. Mitchell, Y\. Lee, A\. Khazatsky, C\. D\. Manning, and C\. Finn \(2023\)Detectgpt: zero\-shot machine\-generated text detection using probability curvature\.InInternational Conference on Machine Learning,pp\. 24950–24962\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px2.p3.1.1)\.
- D\. T\. Nguyen, N\. H\. Lam, A\. H\. Nguyen, and T\. Do \(2025\)MTikGuard system: a transformer\-based multimodal system for child\-safe content moderation on tiktok\.arXiv preprint arXiv:2511\.17955\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- B\. Sadanandan and V\. Behzadan \(2026\)PSF\-med: measuring and explaining paraphrase sensitivity in medical vision language models\.arXiv preprint arXiv:2602\.21428\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- V\. S\. Sadasivan, A\. Kumar, S\. Balasubramanian, W\. Wang, and S\. Feizi \(2025\)Can ai\-generated text be reliably detected? stress testing ai text detectors under various attacks\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- A\. Shportko and I\. Verbitsky \(2025\)Paraphrasing attack resilience of various machine\-generated text detection methods\.InProceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 4: Student Research Workshop\),pp\. 450–456\.External Links:[Link](https://aclanthology.org/2025.naacl-srw.46/)Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px2.p1.1)\.
- I\. Solaiman, M\. Brundage, J\. Clark, A\. Askell, A\. Herbert\-Voss, J\. Wu, A\. Radford, G\. Krueger, J\. W\. Kim, S\. Kreps,et al\.\(2019\)Release strategies and the social impacts of language models\.arXiv preprint arXiv:1908\.09203\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Song, Z\. Yuan, S\. Zhang, Z\. Fang, J\. Yu, and F\. Liu \(2025\)Deep kernel relative test for machine\-generated text detection\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px2.p7.1.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px3.p1.3)\.
- X\. Su, Q\. Mao, Z\. Wu, X\. Lin, S\. You, Y\. Liao, and C\. Xu \(2025\)Large language models driven neural architecture search for universal and lightweight disease diagnosis on histopathology slide images\.npj Digital Medicine8\(1\),pp\. 682\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- E\. Tian \(2023\)External Links:[Link](https://gptzero.me/)Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px2.p2.1.1)\.
- V\. Verma, E\. Fleisig, N\. Tomlin, and D\. Klein \(2024\)Ghostbuster: detecting text ghostwritten by large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 1702–1717\.Cited by:[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p2.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p3.1),[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px2.p4.1.1)\.
- J\. Wu, S\. Yang, R\. Zhan, Y\. Yuan, L\. S\. Chao, and D\. F\. Wong \(2025\)A survey on llm\-generated text detection: necessity, methods, and future directions\.Computational Linguistics51\(1\),pp\. 275–338\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- Q\. Xu, W\. Mu, J\. Li, T\. Sun, and X\. Jiang \(2025\)Advancements in ai\-generated content forensics: a systematic literature review\.ACM Computing Surveys58\(3\),pp\. 1–36\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- X\. Yang, L\. Pan, X\. Zhao, H\. Chen, L\. Petzold, W\. Y\. Wang, and W\. Cheng \(2024\)A survey on detection of llms\-generated content\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 9786–9805\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Yu, M\. Zhao, J\. Lu, K\. Niu, Y\. Wang, W\. Yin, W\. Jia, T\. Fu, Y\. Liu, J\. Liu,et al\.\(2025a\)Eve: towards end\-to\-end video subtitle extraction with vision\-language models\.arXiv preprint arXiv:2503\.04058\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- X\. Yu, Y\. Yu, D\. Liu, K\. Chen, W\. Zhang, N\. Yu, and J\. Shao \(2025b\)EvoBench: towards real\-world llm\-generated text detection benchmarking for evolving large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 14605–14620\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p1.1)\.
- Q\. Zhang, C\. Gao, D\. Chen, Y\. Huang, Y\. Huang, Z\. Sun, S\. Zhang, W\. Li, Z\. Fu, Y\. Wan, and L\. Sun \(2024\)LLM\-as\-a\-coauthor: can mixed human\-written and machine\-generated text be detected?\.pp\. 409–436\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.
- X\. Zhang, J\. Zhao, and Y\. LeCun \(2015\)Character\-level convolutional networks for text classification\.Advances in neural information processing systems28\.Cited by:[§4\.1](https://arxiv.org/html/2605.05950#S4.SS1.SSS0.Px1.p5.1)\.
- Z\. Zhang, Y\. Wang, L\. Cheng, Z\. Zhong, D\. Guo, and M\. Wang \(2025\)Asap: advancing semantic alignment promotes multi\-modal manipulation detecting and grounding\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 4005–4014\.Cited by:[§2](https://arxiv.org/html/2605.05950#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Zhou, J\. Zhu, P\. Su, K\. Ye, Y\. Yang, S\. A\. Gavioli\-Akilagun, and C\. Shi \(2025\)AdaDetectGPT: adaptive detection of llm\-generated text with statistical guarantees\.arXiv preprint arXiv:2510\.01268\.Cited by:[§1](https://arxiv.org/html/2605.05950#S1.p2.1)\.

Similar Articles

Literary Non-Style in LLM-Generated Text

arXiv cs.CL

This paper analyzes statistical patterns in LLM-generated text using n-gram distributions, revealing stylistic deficiencies and showing that style and semantics are not separable.

A Linguistics-Aware LLM Watermarking via Syntactic Predictability

arXiv cs.CL

This paper introduces STELA, a linguistics-aware watermarking framework for LLMs that leverages syntactic predictability via POS n-grams to balance text quality and detection robustness. The method enables publicly verifiable watermark detection without requiring access to model logits, demonstrating superior performance across typologically diverse languages (English, Chinese, Korean).