Why AI Detection Fails for Academic Integrity
Summary
This paper evaluates commercial AI detectors in academic settings, finding high false-positive rates on AI-assisted human writing and near-total evasion via humanizers, concluding detector scores should not be standalone misconduct evidence.
View Cached Full Text
Cached at: 08/13/26, 03:33 PM
# Why AI Detection Fails for Academic Integrity
Source: [https://arxiv.org/html/2608.11256](https://arxiv.org/html/2608.11256)
\(2018\)
###### Abstract\.
Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct\. In a controlled study of published English abstracts \(four domains; 2013 to 2015 vs\. 2023 to 2025\), we quantify this policy failure under proxy human/AI labels atτ=0\.50\\tau=0\.50\. Light*refine \(abstract only\)*edits, a proxy for guideline\-compliant AI assistance, are flagged at 64 to 80% \(Pangram/GPTZero\)\. Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non\-STEM rates far above STEM \(p<0\.001p<0\.001\); elevated scores track long\-token and Academic Word List density, not authorship intent alone\. After Undetectable AI humanization, evasion is near\-total: fewer than 4% of AI\-labeled rewrites remain flagged \(post\-humanization detection rate<4%<4\\%; FNR\>96%\>96\\%\)\. Honest AI\-editing results in a higher sanction risk than humanizer\-assisted evasion\. Therefore, detector scores should not serve as standalone misconduct evidence\.
AI detection, academic integrity, AI\-assisted writing, humanization, LLMs, education
††copyright:acmlicensed††journalyear:2018††doi:XXXXXXX\.XXXXXXX††conference:Make sure to enter the correct conference title from your rights confirmation email; June 03–05, 2018; Woodstock, NY††isbn:978\-1\-4503\-XXXX\-X/2018/06††ccs:Applied computing Education††ccs:Computing methodologies Natural language processing††ccs:Computing methodologies Artificial intelligence††ccs:Security and privacy Human and societal aspects of security and privacy## 1\.Introduction
Consider two students submitting an essay or research paper\. One author creates the claims, but uses LLMs to refine their language\. The other student uses the LLM to produce the full draft with no original analysis\. Atτ=0\.50\\tau=0\.50, commercial detectors often flag both at high rates\. Yet they reflect different relationships to the work\.*Do they deserve the same outcome?*
LLMs have raised concerns about synthetic text in academia\(Kasneciet al\.,[2023](https://arxiv.org/html/2608.11256#bib.bib5); Cottonet al\.,[2024](https://arxiv.org/html/2608.11256#bib.bib6); Perkins,[2023](https://arxiv.org/html/2608.11256#bib.bib7)\), and institutions increasingly deploy AI detectors\(Yanet al\.,[2023](https://arxiv.org/html/2608.11256#bib.bib10); Wang,[2023](https://arxiv.org/html/2608.11256#bib.bib9)\)\. Vendors report high benchmark accuracy\(Spero,[2025](https://arxiv.org/html/2608.11256#bib.bib3); Adam,[2026](https://arxiv.org/html/2608.11256#bib.bib15)\), but independent studies find substantial false positives and false negatives\(Dik and Erdem,[2025](https://arxiv.org/html/2608.11256#bib.bib14); Elazar and Antoniak,[2026](https://arxiv.org/html/2608.11256#bib.bib11)\)\.
Those errors are not limited to texts possibly contaminated with LLM outputs: we have observed a substantial number of false positives even on pre\-AI\-era abstracts \(2013\-2015\), especially in non\-STEM prose\. On 2023\-2025 originals, Pangram flags 15\.0% and GPTZero 8\.9% atτ=0\.50\\tau=0\.50\(Table[2](https://arxiv.org/html/2608.11256#S3.T2)\), with non\-STEM rates far above STEM \(p<0\.001p<0\.001\)\. This shows that existing detectors systematically fail on published and peer\-reviewed human scholarly text, not only on AI\-generated drafts\.
Detectors do not separate fully AI\-generated from AI\-assisted prose in one score\(Spero,[2025](https://arxiv.org/html/2608.11256#bib.bib3); GPTZero,[2026](https://arxiv.org/html/2608.11256#bib.bib25); Pratama,[2025](https://arxiv.org/html/2608.11256#bib.bib26)\)\. At the same time, the growing market for humanizers, software designed to evade AI detectors, adds another potential source for detection errors\(Wallwork,[2025](https://arxiv.org/html/2608.11256#bib.bib17)\)\. Furthermore English\-language learners may legitimately use AI for translation, clarity, or academic polish while retaining intellectual ownership\(Zhanget al\.,[2023](https://arxiv.org/html/2608.11256#bib.bib13)\), yet such edits can elevate detector scores\. These facts together create a tension where people using AI in legitimate ways to present their ideas are punished by AI detection pipelines, but people willing to cheat by using AI in combination with a humanizer are not\. To analyze this tension, we ask three policy\-oriented research questions focused on*detection errors*, not vendor benchmarking:
Figure 1\.AI Detection PipelineRQ1:Under proxy ground\-truth labels, what are the false\-positive and false\-negative rates of commercial detectors, and how do they vary by domain, time period, and rewrite condition?
RQ2:Which surface linguistic features of academic abstracts are associated with elevated detector scores \(and thus higher false\-positive risk on human writing\)?
RQ3:How does humanization change false\-negative rates on AI\-generated text?
Prior work has documented detector unreliability\(Dik and Erdem,[2025](https://arxiv.org/html/2608.11256#bib.bib14)\), including cases in which polishing abstracts are misclassified as full rewrites\(Geng and Poibeau,[2025](https://arxiv.org/html/2608.11256#bib.bib1)\)\. We extend this literature with a unified error\-analysis framework: proxy\-labeled positives/negatives, permutation\-based inference, and interpretable text\-feature correlates across domains and time windows\. Findings are intended to inform policy on published abstract proxies\.
## 2\.Methodology
We describe our corpus and detection pipeline and define how we operationalize detector error rates and the accompanying statistical analyses\. In short, we score original and LLM\-rewritten abstracts across four domains and two time periods with two commercial detectors, before and after humanization\.
### 2\.1\.Corpus and Pipeline
We sampled English abstracts from OpenAlex\(Priemet al\.,[2022](https://arxiv.org/html/2608.11256#bib.bib20)\)in chemistry, computer science, political science, and theology \(100 per domain\-time bucket before filtering\) for 2013 to 2015 and 2023 to 2025\. The earlier window serves as a pre\-LLM\-era reference for proxy FPR on*original*abstracts; the later window reports flag rates when LLM\-assisted writing in scholarly workflows is plausible but unobserved in our labels\(Lianget al\.,[2024](https://arxiv.org/html/2608.11256#bib.bib22); Kobaket al\.,[2025](https://arxiv.org/html/2608.11256#bib.bib23); Geng and Poibeau,[2025](https://arxiv.org/html/2608.11256#bib.bib1)\)\. After requiring 25 to 500 words and PDF access \(≈\\approx20% loss\), we retained 642 abstracts\. Each contributed an*original*abstract plus three Gemini 3 Flash rewrites via OpenRouter\(OpenRouter,[2026](https://arxiv.org/html/2608.11256#bib.bib21)\):*refine \(abstract only\)*,*refine \(abstract \+ article\)*, and*new \(article only\)*\(prompts in Appendix[D](https://arxiv.org/html/2608.11256#A4)\)\. We scored all variants with Pangram 3\.2 and GPTZero\(Emi and Spero,[2024](https://arxiv.org/html/2608.11256#bib.bib2); Adam,[2026](https://arxiv.org/html/2608.11256#bib.bib15)\)\(s∈\[0,1\]s\\in\[0,1\]\) and an LLM\-assisted baseline \(GPT\-5 Nano\)\. We then humanized all variants with Undetectable AI v11 \(Balanced/Doctorate/Article\) and re\-scored outputs\. Our overall pipeline is shown in Fig\.[1](https://arxiv.org/html/2608.11256#S1.F1)\.
We selected two STEM domains \(chemistry, computer science\) and two non\-STEM domains \(political science, theology\) to contrast symbol\-dense technical prose with narrative social\-science writing; STEM versus non\-STEM tests pool chemistry with computer science and political science with theology\. Table[1](https://arxiv.org/html/2608.11256#S2.T1)lists retained abstracts per domain\-time bucket\. The three rewrite conditions increase modeled AI involvement while holding the generator fixed:*refine \(abstract only\)*edits only the abstract text,*refine \(abstract \+ article\)*uses the full article for context, and*new \(article only\)*writes a fresh abstract from the article body only\. Pangram and GPTZero scores are treated symmetrically atτ=0\.50\\tau=0\.50; threshold sweeps and vendor\-claim comparisons are in Appendix[K](https://arxiv.org/html/2608.11256#A11)\.
Domain2013–152023–25Chemistry7695Computer science7783Political science7679Theology7679Table 1\.Retained abstract counts by domain \(after length and PDF filtering\)\.
### 2\.2\.Error Rates and Analysis
Withτ=0\.50\\tau=0\.50andy^i=𝕀\{si≥τ\}\\hat\{y\}\_\{i\}=\\mathbb\{I\}\\\{s\_\{i\}\\geq\\tau\\\}, we assign proxy human labels to originals and proxy AI labels to LLM outputs, definingFPi=𝕀\{yi=human∧y^i=1\}\\mathrm\{FP\}\_\{i\}=\\mathbb\{I\}\\\{y\_\{i\}=\\text\{human\}\\land\\hat\{y\}\_\{i\}=1\\\}andFNi=𝕀\{yi=AI∧y^i=0\}\\mathrm\{FN\}\_\{i\}=\\mathbb\{I\}\\\{y\_\{i\}=\\text\{AI\}\\land\\hat\{y\}\_\{i\}=0\\\}, and reportFNR\\mathrm\{FNR\}on LLM\-labeled variants\. For 2013 to 2015*original*abstracts, we report proxyFPR\\mathrm\{FPR\}; for 2023 to 2025*original*abstracts, we report*flag rate*\(fraction flagged atτ\\tau\), because unobserved AI assistance means positives are not confirmed false positives\. We also report*AI\-assisted false\-positive risk*: the fraction of*refine \(abstract only\)*outputs flagged under proxy human\-source labels\.
Group comparisons use two\-sided permutation tests on mean score differences \(5,000 iterations; seed 42\)\. Text features include long\-token ratio, Academic Word List \(AWL\) density, numeric and non\-alphabetic token shares, acronym density, and type\-token ratio \(definitions in Appendix[F](https://arxiv.org/html/2608.11256#A6)\)\. We report Spearmanρ\\rhobetween features and detector scores with Benjamini\-Hochberg FDR within each collection, plus domain\-centered associations\. Threshold sensitivity, bootstrap CIs, paired original→\\rightarrow*refine \(abstract only\)*shifts, and precision\-recall curves are in Appendix[K](https://arxiv.org/html/2608.11256#A11)\.
## 3\.Results
We organize our findings around the three research questions, reporting proxy\-labeled error rates \(RQ1\), the linguistic correlates of detector scores \(RQ2\), and the effects of humanization on detector behavior and agreement \(RQ3\)\. Across all three, detectors prove sensitive to stylistic cues and are readily defeated by humanizers, leading to high FNR and low inter\-detector agreement\.
PeriodDetectorFPR/flagFNR preFNR post2013\-15Pangram0\.0%19\.4%98\.5%2013\-15GPTZero0\.0%44\.2%96\.0%2013\-15LLM\-aid46\.7%35\.1%30\.5%2023\-25Pangram15\.0%13\.5%96\.6%2023\-25GPTZero8\.9%38\.4%96\.1%2023\-25LLM\-aid70\.8%22\.8%19\.7%Table 2\.Proxy\-labeled error rates atτ=0\.50\\tau=0\.50\. For*original*abstracts, 2013 to 2015 reports proxy FPR; 2023 to 2025 reports flag rate\.Table[2](https://arxiv.org/html/2608.11256#S3.T2)summarizes error rates atτ=0\.50\\tau=0\.50\. Under proxy labels, 2013 to 2015 originals incur FPR=0%=0\\%; in 2023 to 2025, original flag rates are 15\.0% \(Pangram\) and 8\.9% \(GPTZero\), not interpretable as confirmed false positives because published abstracts may already include LLM assistance\. Non\-STEM flag rates exceed STEM \(p<0\.001p<0\.001; Table[4](https://arxiv.org/html/2608.11256#S3.T4)\)\.*Refine \(abstract only\)*is flagged 64 to 80% \(proxy AI\-assisted false\-positive risk\)\. After humanization, fewer than 4% of AI\-labeled rewrites remain detectable for AI detectors \(FNR\>96%\>96\\%\)\. The asymmetry is policy\-critical: guideline\-compliant editing triggers high flag rates, while humanizer evasion drops detection on synthetic text to near zero; honest editing currently faces more sanction risk than malicious evasion\.
Table[3](https://arxiv.org/html/2608.11256#S3.T3)decomposes pre\-humanization FNR by rewrite condition\. Lighter edits retain the highest FNR, especially on GPTZero;*new \(article only\)*rewrites already exceedτ\\taufor most Pangram scores \(2\.6 to 2\.7% FNR\)\. The LLM\-assisted baseline flags 70\.8% of 2023 to 2025 originals and is treated as supplementary only\.
PeriodConditionPangram FNRGPTZero FNRnn2013\-15refine \(abs\. only\)35\.6%62\.4%3062013\-15refine \(abs\.\+paper\)19\.9%54\.6%3062013\-15new \(article only\)2\.6%15\.7%3062023\-25refine \(abs\. only\)19\.9%51\.5%3362023\-25refine \(abs\.\+paper\)17\.9%48\.8%3362023\-25new \(article only\)2\.7%14\.9%336Table 3\.Pre\-humanization FNR by rewrite condition atτ=0\.50\\tau=0\.50\.DomainOriginal flaggedRefine \(abs\. only\) flaggednnChemistry2\.1 / 1\.1%73\.7 / 46\.3%95Computer science9\.8 / 6\.0%69\.9 / 36\.1%83Political science24\.4 / 15\.2%86\.1 / 54\.4%79Theology26\.6 / 15\.2%92\.4 / 58\.2%79Table 4\.Domain\-level flag rates in 2023 to 2025 \(Pangram / GPTZero\)\.Flag rates on 2023 to 2025 originals vary sharply by domain at the sameτ\\tau\(Table[4](https://arxiv.org/html/2608.11256#S3.T4)\): Pangram flags 2\.1% in chemistry versus 26\.6% in theology; GPTZero shows the same ordering \(1\.1% to 15\.2%\)\. Vendor marketing claims sub\-1% false\-positive rates on benchmark evaluations\(Spero,[2025](https://arxiv.org/html/2608.11256#bib.bib3); GPTZero,[2026](https://arxiv.org/html/2608.11256#bib.bib25)\); our corpus flags 9 to 15% of recent*original*abstracts and 64 to 80% of light edits \(Appendix Table[12](https://arxiv.org/html/2608.11256#A5.T12)\)\. Figure[2](https://arxiv.org/html/2608.11256#S3.F2)plots assisted\-editing ROC on 2013 to 2015 abstracts \(0% original FPR atτ=0\.50\\tau=0\.50\)\. Pangram reaches AUC\-ROC0\.980\.98versus GPTZero0\.920\.92; atτ=0\.50\\tau=0\.50Pangram flags 64% of light edits with no originals flagged, whereas GPTZero flags 38%\. Figure[3](https://arxiv.org/html/2608.11256#S3.F3)replicates the task on 2023 to 2025: original\-axis rates are flag rates \(15% and 9%\), Pangram AUC\-ROC is0\.910\.91, and 80% of light edits are flagged atτ=0\.50\\tau=0\.50\.
Figure 2\.Assisted\-editing ROC \(2013\-2015\):*original*vs\.*refine \(abstract only\)*;τ=0\.50\\tau=0\.50Figure 3\.Assisted\-editing ROC/PR \(2023\-2025, pre\-humanization\)\. Original\-axis rates are flag rates, not proxy FPR\.No thresholdτ∈\{0\.4,0\.5,0\.6\}\\tau\\in\\\{0\.4,0\.5,0\.6\\\}simultaneously keeps original flag rates, refine capture, and AI\-pool FNR low \(Table[5](https://arxiv.org/html/2608.11256#S3.T5)\): Pangram holds≈\\approx15% original flags and≈\\approx80% refine flags across that band while FNR on AI\-labeled rewrites stays 10 to 14%\.
Detectorτ\\tauFlag rate\(original\)FNR\(AI\-labeled\)RefineflaggedPangram0\.415\.9%10\.5%85\.1%Pangram0\.515\.0%13\.5%80\.1%Pangram0\.615\.0%13\.8%79\.8%GPTZero0\.410\.1%37\.6%49\.4%GPTZero0\.58\.9%38\.4%48\.5%GPTZero0\.68\.6%38\.6%48\.2%Table 5\.Threshold sensitivity \(2023\-2025, pre\-humanization\)\.### 3\.1\.Linguistic Correlates
Long\-token and AWL ratios correlate positively with 2023 to 2025 scores \(ρ≈0\.30\\rho\\approx 0\.30to0\.350\.35;p<0\.001p<0\.001\); numeric and non\-alphabetic ratios correlate negatively \(Table[6](https://arxiv.org/html/2608.11256#S3.T6); full tests in Appendix[H](https://arxiv.org/html/2608.11256#A8)\)\. After centering scores and features within domain, long\-token and AWL associations strengthen for Pangram \(domain\-adjustedr\>0\.41r\>0\.41\), indicating detectors respond to within\-domain stylistic cues rather than domain membership alone\.
FeaturePangramρ\\rhoGPTZeroρ\\rhoAWL token ratio0\.3460\.320Long token ratio0\.3360\.304Non\-alphabetic token ratio−0\.149\-0\.149−0\.197\-0\.197Numeric token ratio−0\.146\-0\.146−0\.171\-0\.171Table 6\.Strongest score\-feature associations \(2023 to 2025, pre\-humanization; Spearmanρ\\rho, both detectorsp<0\.001p<0\.001\)\.
### 3\.2\.Humanization and Detector Agreement
Post\-humanization Pangram means fall to 0\.03 to 0\.04 across conditions in 2023 to 2025 \(Appendix[K](https://arxiv.org/html/2608.11256#A11)\), driving the\>96%\>96\\%false\-negative rates\. Pangram and GPTZero agree moderately before humanization but weakly afterward \(Table[7](https://arxiv.org/html/2608.11256#S3.T7)\), so a single vendor flag is not a reliable consensus signal after evasion\. Acrossn=2,568n=2\{,\}568humanized pairs, long\-token ratio, AWL density, and lexical diversity fall systematically \(p<0\.001p<0\.001; Table[8](https://arxiv.org/html/2608.11256#S3.T8)\), matching the pre\-humanization features that track higher scores—evasion suppresses detector\-salient surface cues rather than uniformly fragmenting syntax \(Appendix[M](https://arxiv.org/html/2608.11256#A13)\)\.
PeriodPhaseSpearmanρ\\rhoCohen’sκ\\kappa2013\-15pre0\.7210\.5002013\-15post0\.1000\.2002023\-25pre0\.7360\.5392023\-25post0\.1880\.221Table 7\.Pangram and GPTZero agreement \(n=1342n=1342to13441344matched pairs per phase\)\.FeaturePrePostΔ\\DeltappLong\-token ratio0\.1840\.138−0\.045\-0\.045<0\.001<0\.001AWL token ratio0\.1570\.121−0\.036\-0\.036<0\.001<0\.001Type\-token ratio0\.7000\.605−0\.094\-0\.094<0\.001<0\.001Word count180229\+49\+49<0\.001<0\.001Table 8\.Pooled linguistic shifts after Undetectable AI v11 humanization \(n=2,568n=2\{,\}568pairs\)\.
## 4\.Discussion
Detector errors are structured, not random: long\-token and AWL density track higher scores in non\-STEM prose, while symbol\-heavy STEM text scores lower even on AI\-assisted rewrites\. Under proxy labels, compliant editing triggers flags while humanization drives detection on AI\-labeled text below 4%\. Institutions therefore face a catch\-22: prohibit all AI\-assisted editing and deny legitimate support to multilingual and novice writers, or enforce with tools that punish acceptable use and miss evasion at scale\(Wallwork,[2025](https://arxiv.org/html/2608.11256#bib.bib17); Sunet al\.,[2026](https://arxiv.org/html/2608.11256#bib.bib16)\)\. Uniform detector policies may also raise differential\-impact concerns when flag rates differ by discipline\(Zdravkova and Ilijoski,[2024](https://arxiv.org/html/2608.11256#bib.bib12)\)\.
Practice:These patterns create concrete educational risk: students may be flagged while drafting their own ideas and using AI only for clarity or grammar, yet others who generate inorganic drafts can evade detection after humanization\. Multilingual learners face additional harm when AI\-assisted translation or editing elevates scores on the student’s own work\(Zhanget al\.,[2023](https://arxiv.org/html/2608.11256#bib.bib13)\)\. Detector output must be treated as weak evidence, not default misconduct evidence\. Any flag must be paired with drafting history and sanctions shall be reserved for cases where human judgment supports a lack of appropriate student effort\. Expect higher rates in non\-STEM prose\. Finally, FERPA/GDPR guidelines must be followed before uploading students’ confidential work to third\-party services\(Quinnipiac Innovations in Learning and Teaching \(QILT\),[n\.d\.](https://arxiv.org/html/2608.11256#bib.bib4)\)\.
Limitations:We analyze English published abstracts, not student essays; detector behavior on classroom genres may differ\. Labels are proxy\-based at a fixedτ=0\.50\\tau=0\.50; 2023 to 2025*original*positives are flag rates, not confirmed false positives\. We used one rewrite generator \(Gemini 3 Flash\), two commercial detectors \(Pangram 3\.2, GPTZero\), one humanizer \(Undetectable AI v11\), and one LLM\-assisted baseline\. Corpus size \(642 abstracts\) reflects API and humanization cost\. Additional limitations are in Appendix[A](https://arxiv.org/html/2608.11256#A1)\.
## 5\.Conclusion
Commercial detectors on published abstracts punish acceptable AI\-assisted editing, compound risk for English\-language learners who use AI for translation, and fail to catch humanized synthetic text\. Light edits are flagged at 64–80%\. After humanization, more than 96% of AI\-labeled full rewrites evade Pangram and GPTZero\.Together, these results define an*integrity catch\-22*: high detector flags on assisted writing alongside near\-total misses after humanization\.Integrity programs therefore need transparent AI\-use norms and process evidence, not standalone detector scores\.
## 6\.Acknowledgments
We would like to thank Pangram for the use of their API, as well as Notre Dame Teaching & Learning, and the Notre Dame AI Council\.
## References
- A\. Adam \(2026\)GPTZero ai detection benchmarking: the industry standard in accuracy, transparency and fairness\.External Links:[Link](https://gptzero.me/news/gptzero-ai-detection-benchmarking-the-industry-standard-in-accuracy-transparency-and-fairness/)Cited by:[§E\.1](https://arxiv.org/html/2608.11256#A5.SS1.p1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.3.4.1.1),[§1](https://arxiv.org/html/2608.11256#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.11256#S2.SS1.p1.2)\.
- D\. R\. Cotton, P\. A\. Cotton, and J\. R\. Shipway \(2024\)Chatting and cheating: ensuring academic integrity in the era of chatgpt\.Innovations in education and teaching international61\(2\),pp\. 228–239\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p2.1)\.
- A\. Coxhead \(2000\)A new academic word list\.TESOL quarterly34\(2\),pp\. 213–238\.Cited by:[Appendix F](https://arxiv.org/html/2608.11256#A6.SS0.SSS0.Px5.p1.4)\.
- S\. Dik and O\. Erdem \(2025\)Assessing gpt\-zero’s accuracy in identifying ai vs\. human\-written essays\.Proceedings of International Mathematical Sciences7\(2\),pp\. 54–58\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p2.1),[§1](https://arxiv.org/html/2608.11256#S1.p6.1)\.
- Y\. Elazar and M\. Antoniak \(2026\)LLM\-generated or human\-written? comparing review and non\-review papers on arxiv\.arXiv preprint arXiv:2601\.17036\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p2.1)\.
- B\. Emi and M\. Spero \(2024\)Technical report on the pangram ai\-generated text classifier\.arXiv preprint arXiv:2402\.14873\.Cited by:[§2\.1](https://arxiv.org/html/2608.11256#S2.SS1.p1.2)\.
- M\. Geng and T\. Poibeau \(2025\)What are we detecting, really? llm\-generated text detection remains an unsolved problem\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p6.1),[§2\.1](https://arxiv.org/html/2608.11256#S2.SS1.p1.2)\.
- GPTZero \(2026\)GPTZero ai detection benchmarking: the industry standard in accuracy, transparency and fairness\.External Links:[Link](https://gptzero.me/news/gptzero-ai-detection-benchmarking-the-industry-standard-in-accuracy-transparency-and-fairness/)Cited by:[§E\.1](https://arxiv.org/html/2608.11256#A5.SS1.p1.1),[§E\.2](https://arxiv.org/html/2608.11256#A5.SS2.p1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.4.4.1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.6.4.1.1),[§1](https://arxiv.org/html/2608.11256#S1.p4.1),[§3](https://arxiv.org/html/2608.11256#S3.p4.7)\.
- E\. Kasneci, K\. Seßler, S\. Küchemann, M\. Bannert, D\. Dementieva, F\. Fischer, U\. Gasser, G\. Groh, S\. Günnemann, E\. Hüllermeier,et al\.\(2023\)ChatGPT for good? on opportunities and challenges of large language models for education\.Learning and individual differences103,pp\. 102274\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p2.1)\.
- D\. Kobak, R\. González\-Márquez, E\. Horvát, and J\. Lause \(2025\)Delving into llm\-assisted writing in biomedical publications through excess vocabulary\.Science Advances11\(27\),pp\. eadt3813\.Cited by:[§2\.1](https://arxiv.org/html/2608.11256#S2.SS1.p1.2)\.
- W\. Liang, Z\. Izzo, Y\. Zhang, H\. Lepp, H\. Cao, X\. Zhao, L\. Chen, H\. Ye, S\. Liu, Z\. Huang,et al\.\(2024\)Monitoring ai\-modified content at scale: a case study on the impact of chatgpt on ai conference peer reviews\.arXiv preprint arXiv:2403\.07183\.Cited by:[§2\.1](https://arxiv.org/html/2608.11256#S2.SS1.p1.2)\.
- OpenRouter \(2026\)OpenRouter: the unified interface for llms\.External Links:[Link](https://openrouter.ai/)Cited by:[§2\.1](https://arxiv.org/html/2608.11256#S2.SS1.p1.2)\.
- M\. Perkins \(2023\)Academic integrity considerations of ai large language models in the post\-pandemic era: chatgpt and beyond\.Journal of University Teaching and Learning Practice20\(2\),pp\. 1–24\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p2.1)\.
- A\. R\. Pratama \(2025\)The accuracy\-bias trade\-offs in ai text detection tools and their impact on fairness in scholarly publication\.PeerJ Computer Science11,pp\. e2953\.Cited by:[§E\.2](https://arxiv.org/html/2608.11256#A5.SS2.p1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.7.4.1.1),[§1](https://arxiv.org/html/2608.11256#S1.p4.1)\.
- J\. Priem, H\. Piwowar, and R\. Orr \(2022\)OpenAlex: a fully\-open index of scholarly works, authors, venues, institutions, and concepts\.arXiv preprint arXiv:2205\.01833\.Cited by:[§2\.1](https://arxiv.org/html/2608.11256#S2.SS1.p1.2)\.
- Quinnipiac Innovations in Learning and Teaching \(QILT\) \(n\.d\.\)External Links:[Link](https://qilt.qu.edu/ai/academic-integrity-and-ai-detectors)Cited by:[§4](https://arxiv.org/html/2608.11256#S4.p2.1)\.
- M\. Spero \(2025\)External Links:[Link](https://www.pangram.com/blog/introducing-ai-assistance-detection)Cited by:[§E\.1](https://arxiv.org/html/2608.11256#A5.SS1.p1.1),[§E\.2](https://arxiv.org/html/2608.11256#A5.SS2.p1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.2.4.1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.5.4.1.1),[§1](https://arxiv.org/html/2608.11256#S1.p2.1),[§1](https://arxiv.org/html/2608.11256#S1.p4.1),[§3](https://arxiv.org/html/2608.11256#S3.p4.7)\.
- Y\. Sun, Y\. Liao, and X\. Ma \(2026\)Trusting ai to detect ai? a systematic evaluation of the reliability and robustness of current aigc detection tools for student academic work\.Computers & Education,pp\. 105616\.Cited by:[§4](https://arxiv.org/html/2608.11256#S4.p1.1)\.
- Undetectable AI \(2026a\)External Links:[Link](https://undetectable.ai/)Cited by:[§E\.3](https://arxiv.org/html/2608.11256#A5.SS3.p1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.9.4.1.1)\.
- Undetectable AI \(2026b\)Pricing\.External Links:[Link](https://undetectable.ai/pricing)Cited by:[§E\.3](https://arxiv.org/html/2608.11256#A5.SS3.p1.1),[Table 12](https://arxiv.org/html/2608.11256#A5.T12.1.8.4.1.1)\.
- A\. Wallwork \(2025\)AI detectors and humanizers\.InThe Pros and Cons of Using Chatbots,pp\. 115–133\.Cited by:[§E\.3](https://arxiv.org/html/2608.11256#A5.SS3.p1.1),[§1](https://arxiv.org/html/2608.11256#S1.p4.1),[§4](https://arxiv.org/html/2608.11256#S4.p1.1)\.
- Y\. Wang \(2023\)Synthetic realities in the digital age: navigating the opportunities and challenges of ai\-generated content\.Authorea Preprints\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p2.1)\.
- D\. Yan, M\. Fauss, J\. Hao, and W\. Cui \(2023\)Detection of ai\-generated essays in writing assessments\.Psychological Test and Assessment Modeling65\(1\),pp\. 125–144\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p2.1)\.
- K\. Zdravkova and B\. Ilijoski \(2024\)Preventing academic dishonesty originating from large language models\.InBalkan Conference in Informatics,pp\. 118–132\.Cited by:[§4](https://arxiv.org/html/2608.11256#S4.p1.1)\.
- X\. Zhang, S\. Li, B\. Hauer, N\. Shi, and G\. Kondrak \(2023\)Don’t trust chatgpt when your question is not in english: a study of multilingual abilities and types of llms\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 7915–7927\.Cited by:[§1](https://arxiv.org/html/2608.11256#S1.p4.1),[§4](https://arxiv.org/html/2608.11256#S4.p2.1)\.
## Appendix
## Appendix ALimitations
Language and corpus setting:Our dataset is limited to English\. We do not evaluate AI detection or humanization for other languages, so our findings should not be extrapolated to multilingual classrooms or non\-English scholarly writing without further study\. We analyze published research abstracts drawn from four OpenAlex domains \(chemistry, computer science, political science, and theology\), not student essays or other classroom genres\. Detector behavior on shorter, pedagogically shaped student writing may differ from the abstract\-length prose studied here\.
Labels and measurement:Ground truth is experimental rather than adjudicated\. We treat unmodified*original*abstracts as proxy human writing and LLM rewrites as proxy AI writing\. For 2023 to 2025 originals, real\-world AI assistance is unobserved, so positive detector outputs are positives on originals, not confirmed false positives\. All reported false\-positive and false\-negative rates depend on this labeling scheme and on a fixed score threshold \(τ=0\.50\\tau=0\.50\)\. Text\-feature associations are correlational and do not establish causal mechanisms\.
Models, detectors, and humanization:We used one rewrite generator \(Gemini 3 Flash\), two commercial API detectors \(Pangram 3\.2 and GPTZero\), one humanizer configuration \(Undetectable AI v11\), and one supplementary LLM\-assisted baseline \(GPT\-5 Nano\)\. Results may not generalize to other models, detector versions, humanizers, or prompting strategies\. Post\-humanization scores are reported for Pangram and GPTZero only; the LLM\-assisted baseline was not re\-run on humanized text \(Appendix[N](https://arxiv.org/html/2608.11256#A14)\)\.
Corpus size and cost:We targeted 800 abstracts \(100 per domain\-time bucket\) and retained 642 after length and PDF filtering\. Corpus size is constrained by the cost of large\-scale API detection and paid humanization\. Each retained paper requires multiple rewrite calls, pre\-detection passes across three scorers, one humanization credit per variant, and post\-detection with two commercial detectors \(Appendix[N](https://arxiv.org/html/2608.11256#A14)\)\. A multi\-model rewrite sweep or full post\-humanization LLM\-assisted scoring would multiply those costs without changing the core error\-analysis design\.
Future work:Future studies could expand to non\-English corpora and to adjudicated student writing with verified authorship and AI\-use histories\. Developing local or open models for rewriting, detection, and humanization could reduce reliance on paid APIs and enable larger, more diverse datasets\. Such systems would need careful validation before use in high\-stakes integrity workflows, but they offer a path toward reproducible benchmarks at scales that are impractical under current commercial pricing\. An open question is whether prose shaped by frontier models \(or by newer assistance tools\) systematically elevates commercial detector scores and therefore false\-positive risk on human writing relative to earlier LLM generations; our 2023 to 2025 window cannot separate model vintage from unobserved real\-world AI use on*original*abstracts, but targeted comparisons across model families on controlled human\-authored drafts would clarify that risk\.
## Appendix BDual interpretation of light editing
Reviewers and integrity offices need a clear account of why*refine \(abstract only\)*appears in two rate families\.
For each source paper we generate a Gemini 3 Flash rewrite that edits only the abstract \(prompt in Appendix[D](https://arxiv.org/html/2608.11256#A4)\)\. Detector scoressand flagy^=𝕀\{s≥τ\}\\hat\{y\}=\\mathbb\{I\}\\\{s\\geq\\tau\\\}are identical regardless of interpretation\.
#### Reading A: proxy\-AI labeling\.
Labely=AIy=\\mathrm\{AI\}\. ThenFN=𝕀\{y^=0\}\\mathrm\{FN\}=\\mathbb\{I\}\\\{\\hat\{y\}=0\\\}and FNR is the miss rate among AI\-labeled rewrites\. This reading supports detection\-evaluation claims \(can the tool catch LLM text?\)\.
#### Reading B: human\-source / assistance\.
Label the intellectual source as human and treat the LLM as a polish tool\. Then the quantity of interest is the*assisted\-writing flag rate*P\(y^=1∣refine abs\. only\)P\(\\hat\{y\}=1\\mid\\text\{refine abs\.\\ only\}\)\. This reading supports policy claims about sanction risk for authors who retain ownership of claims\. It isnotproxy FPR on unmodified human text, and it must not be abbreviated as FPR in tables or prose\.
#### Why the rates are not comparable as ordinary errors:
FNR under Reading A and assisted\-writing flag rate under Reading B use opposite label assignments for the same texts\. Comparing 64–80% false positives to\>96%\>96\\%false negatives as if they were cells of one confusion matrix would be a category error\. We therefore report both quantities, name them differently, and interpret the integrity catch\-22 as a*policy*asymmetry across readings, not as a single\-threshold accuracy summary\.
#### Relationship to institutional rules:
Policies that ban all generative AI would treat Reading A as decisive\. Policies that allow grammar/clarity assistance but ban ghostwriting sit between the readings and require process evidence our detectors do not observe\. Our refine prompt is closer to light generative rewriting than to grammar\-only editing; assisted\-writing flag rates should be read accordingly\.
## Appendix CPaper\-clustered inference
Each source paper contributes an original abstract and three rewrite conditions\. Permutation tests and bootstrap intervals that resample instances can understate uncertainty when within\-paper dependence is ignored\. Therefre, we report analyses that resample at thepaper\_idlevel \(2,000 bootstrap replicates or 5,000 permutations; seed 42;τ=0\.50\\tau=0\.50\)\.
Error and flag rates:Table[9](https://arxiv.org/html/2608.11256#A3.T9)shows paper\-clustered 95% intervals\. Point estimates match the main text \(e\.g\., 2023 to 2025 pre\-humanization Pangram original flag rate 15\.0%\[11\.4,19\.2\]\[11\.4,19\.2\]; refine assisted\-writing flag rate 80\.1%\[75\.9,84\.5\]\[75\.9,84\.5\]; AI\-pool FNR 13\.5%\[11\.0,16\.0\]\[11\.0,16\.0\]; post\-humanization AI\-pool FNR 96\.6%\[95\.2,97\.8\]\[95\.2,97\.8\]\)\. Clustering widens intervals relative to naive instance bootstrap but does not change substantive conclusions\.
STEM vs\. non\-STEM:Table[10](https://arxiv.org/html/2608.11256#A3.T10)uses one original\-abstract score per paper\. In 2023 to 2025, non\-STEM flag rates remain far above STEM \(Pangram 25\.5% vs\. 5\.6%; GPTZero 15\.2% vs\. 3\.4%; paper\-label permutationp=0\.0002p=0\.0002for both detectors\)\.
Feature associations:Table[11](https://arxiv.org/html/2608.11256#A3.T11)retains instance\-level Spearmanρ\\rho\(matching Table[6](https://arxiv.org/html/2608.11256#S3.T6)\) but forms CIs by resampling papers\. Long\-token and AWL associations remain positive with intervals excluding zero for both detectors \(e\.g\., Pangram AWLρ=0\.346\\rho=0\.346\[0\.298,0\.395\]\[0\.298,0\.395\]\)\. Paper\-averagedρ\\rhois attenuated because averaging mixes originals and rewrites within paper; we report it for transparency, not as a replacement for the instance\-level estimand\.
PeriodPhaseDetectorMetricPoint95% CI \(paper\)npapersn\_\{\\mathrm\{papers\}\}2013\-2015postgptzeroFNR \(AI\-labeled pool\)96\.0%\[94\.6%, 97\.3%\]3062013\-2015postgptzeroOriginal flag rate1\.3%\[0\.3%, 2\.6%\]3062013\-2015postgptzeroRefine \(abs\. only\) flag rate3\.3%\[1\.3%, 5\.6%\]3062013\-2015postpangramFNR \(AI\-labeled pool\)98\.5%\[97\.5%, 99\.2%\]3062013\-2015postpangramOriginal flag rate0\.3%\[0\.0%, 1\.0%\]3062013\-2015postpangramRefine \(abs\. only\) flag rate3\.3%\[1\.6%, 5\.6%\]3062013\-2015pregptzeroFNR \(AI\-labeled pool\)44\.2%\[40\.6%, 47\.8%\]3062013\-2015pregptzeroOriginal flag rate0\.0%\[0\.0%, 0\.0%\]3062013\-2015pregptzeroRefine \(abs\. only\) flag rate37\.6%\[32\.0%, 43\.1%\]3062013\-2015prepangramFNR \(AI\-labeled pool\)19\.4%\[16\.4%, 22\.3%\]3062013\-2015prepangramOriginal flag rate0\.0%\[0\.0%, 0\.0%\]3052013\-2015prepangramRefine \(abs\. only\) flag rate64\.4%\[58\.8%, 69\.3%\]3062023\-2025postgptzeroFNR \(AI\-labeled pool\)96\.1%\[94\.6%, 97\.3%\]3362023\-2025postgptzeroOriginal flag rate1\.5%\[0\.3%, 3\.0%\]3362023\-2025postgptzeroRefine \(abs\. only\) flag rate3\.9%\[1\.8%, 6\.0%\]3362023\-2025postpangramFNR \(AI\-labeled pool\)96\.6%\[95\.2%, 97\.8%\]3362023\-2025postpangramOriginal flag rate3\.3%\[1\.5%, 5\.4%\]3362023\-2025postpangramRefine \(abs\. only\) flag rate3\.6%\[1\.8%, 5\.7%\]3362023\-2025pregptzeroFNR \(AI\-labeled pool\)38\.4%\[34\.7%, 41\.9%\]3362023\-2025pregptzeroOriginal flag rate8\.9%\[6\.2%, 11\.9%\]3362023\-2025pregptzeroRefine \(abs\. only\) flag rate48\.5%\[43\.2%, 53\.6%\]3362023\-2025prepangramFNR \(AI\-labeled pool\)13\.5%\[11\.0%, 16\.0%\]3362023\-2025prepangramOriginal flag rate15\.0%\[11\.4%, 19\.2%\]3342023\-2025prepangramRefine \(abs\. only\) flag rate80\.1%\[75\.9%, 84\.5%\]336Table 9\.Paper\-clustered bootstrap intervals \(nboot=2,000n\_\{\\mathrm\{boot\}\}=2\{,\}000; seed 42\) atτ=0\.50\\tau=0\.50\.PeriodDetectorΔ\\DeltameanFlag STEMFlag non\-STEMnnSTEM/nonpp\(paper\)2013\-2015gptzero−0\.006\-0\.0060\.0%0\.0%154/1520\.0262013\-2015pangram0\.0010\.0010\.0%0\.0%153/1520\.622023\-2025gptzero0\.1100\.1103\.4%15\.2%178/1580\.00022023\-2025pangram0\.2040\.2045\.6%25\.5%177/1570\.0002Table 10\.STEM vs\. non\-STEM on*original*abstracts\.DetectorFeatureρ\\rho\(inst\.\)95% CI \(paper\)ρ\\rho\(paper avg\.\)npapersn\_\{\\mathrm\{papers\}\}gptzeroawl token ratio0\.320\[0\.269, 0\.370\]0\.139336gptzerolong token ratio0\.304\[0\.249, 0\.359\]0\.079336pangramawl token ratio0\.346\[0\.298, 0\.395\]0\.148336pangramlong token ratio0\.336\[0\.290, 0\.386\]0\.072336Table 11\.Text\-feature associations \(2023 to 2025, pre\-humanization\)\.
## Appendix DAbstract Rewriting Prompts
We provide the prompts used to instruct the model for abstract rewriting\.
Prompt used for the Refine \(abstract only\) setting: You are an expert scientist who excels in writing papers\. Your task is to rewrite a given paper abstract to sound better and more human\-like\. The abstract will be given in the next message\. Please only return the rewritten text, do not return any instructions or anything like that\. Only return the rewritten text\. Make sure to make at least some edits or changes to the text\.
Prompt used for the Refine \(abstract \+ article\) setting: You are an expert scientist who excels in writing papers\. Your task is to rewrite a given paper abstract to sound better and more human\-like, using the full text of the paper to inform your rewriting\. The abstract and text will be given in the next message\. Please only return the rewritten text, do not return any instructions or anything like that\. Only return the rewritten text\. Make sure to make at least some edits or changes to the text\.
Prompt used for the New \(article only\) setting: You are an expert scientist who excels in writing papers\. Your task is to write a naturally sounding abstract for a paper that will be given to you\. The full will be given in the next message\. Please only return the written abstract, do not return any instructions or anything like that\. Only return the written text\. Make sure to not copy the text of the paper verbatim\.
## Appendix EVendor marketing claims
We record the commercial claims that motivate academic\-integrity adoption of AI detectors and humanizers\. Table[12](https://arxiv.org/html/2608.11256#A5.T12)organizes them in the same order as our framing: headline detector accuracy, claims about*AI\-assisted*versus fully generated writing, and humanizer evasion guarantees\. All figures come from vendor marketing or vendor\-run benchmarks \(accessed December 2025 to May 2026\), not from our corpus\.
### E\.1\.Headline detector accuracy
Vendors advertise near\-perfect detection on benchmark\-style evaluations\. Pangram’s Pangram 3\.0 announcement reports 99\.98% accuracy on AI\-generated text with near\-zero false positives on the fully AI\-generated label\(Spero,[2025](https://arxiv.org/html/2608.11256#bib.bib3)\)\. GPTZero’s peer\-reviewed detector paper reports 99\.39% accuracy on its multi\-domain evaluation \(Table 2 aggregate for GPTZero 4\.1b\)\(Adam,[2026](https://arxiv.org/html/2608.11256#bib.bib15)\)\. GPTZero’s public benchmarking post additionally claims 99\.76% average accuracy with 0\.08% false\-positive rate across four commercial domains\(GPTZero,[2026](https://arxiv.org/html/2608.11256#bib.bib25)\)\.
### E\.2\.AI\-assisted versus fully generated writing
A separate set of claims concerns whether detectors can separate light editing from full synthesis\. Pangram 3\.0 markets explicit classes \(fully human, lightly AI\-assisted, moderately AI\-assisted, fully AI\-generated\)\(Spero,[2025](https://arxiv.org/html/2608.11256#bib.bib3)\)\. GPTZero describes multiclass outputs that separate human text, mixed or lightly edited text, and pure AI\(GPTZero,[2026](https://arxiv.org/html/2608.11256#bib.bib25)\)\. Independent work notes that accuracy\-bias trade\-offs and binarization choices can obscure assisted\-writing risk\(Pratama,[2025](https://arxiv.org/html/2608.11256#bib.bib26)\)\.
### E\.3\.Humanizer evasion guarantees
Detector adoption has spurred humanizers that target surface metrics such as perplexity and burstiness\(Wallwork,[2025](https://arxiv.org/html/2608.11256#bib.bib17)\)\. Undetectable AI’s pricing page states a money\-back guarantee: if humanized output is flagged as not human, the humanization cost is refunded\(Undetectable AI,[2026b](https://arxiv.org/html/2608.11256#bib.bib18)\)\. Its marketing site also advertises 99%\+ detector accuracy, a Forbes “\#1 Best AI Detector” rating \(2024\), and more than 22 million users as of March 2026\(Undetectable AI,[2026a](https://arxiv.org/html/2608.11256#bib.bib19)\)\.
VendorClaim typeStated claimSourcePangramAccuracy99\.98% accuracy on AI\-generated text; near\-zero false positives on the AI\-generated label\.\(Spero,[2025](https://arxiv.org/html/2608.11256#bib.bib3)\)GPTZeroAccuracy99\.39% accuracy on multi\-domain detector evaluation \(GPTZero 4\.1b\)\.\(Adam,[2026](https://arxiv.org/html/2608.11256#bib.bib15)\)GPTZeroAccuracy \(marketing\)99\.76% average accuracy; 0\.08% FPR; 99\.60% recall \(vendor benchmark, v4\.3b\)\.\(GPTZero,[2026](https://arxiv.org/html/2608.11256#bib.bib25)\)PangramAssisted vs\. generatedFour\-way taxonomy: human, light assistance, moderate assistance, fully AI\-generated\.\(Spero,[2025](https://arxiv.org/html/2608.11256#bib.bib3)\)GPTZeroAssisted vs\. generatedMulticlass taxonomy \(human, mixed/edited, pure AI\) for benchmarking\.\(GPTZero,[2026](https://arxiv.org/html/2608.11256#bib.bib25)\)LiteratureAssisted vs\. generatedDetectors face accuracy\-bias trade\-offs; binarized outputs can hide assisted\-writing risk\.\(Pratama,[2025](https://arxiv.org/html/2608.11256#bib.bib26)\)Undetectable AIHumanizer guaranteeMoney\-back if humanized text is “flagged as not human\.”\(Undetectable AI,[2026b](https://arxiv.org/html/2608.11256#bib.bib18)\)Undetectable AIMarketing99%\+ detector accuracy; Forbes \#1 \(2024\); 22M\+ users \(March 2026\)\.\(Undetectable AI,[2026a](https://arxiv.org/html/2608.11256#bib.bib19)\)Table 12\.Vendor claims cited in our framing \(marketing and vendor benchmarks\)\.Contrast with our corpus:These claims are not reproduced on published English abstracts under our proxy labels atτ=0\.50\\tau=0\.50\. Pangram and GPTZero flag 15\.0% and 8\.9% of 2023 to 2025*original*abstracts \(Table[2](https://arxiv.org/html/2608.11256#S3.T2)\), inconsistent with sub\-1% false\-positive marketing\.*Refine \(abstract only\)*outputs \(human\-source, light LLM edit\) are flagged 64 to 80% of the time, showing that assisted\-writing risk is substantial even when vendors market assistance\-level detection\. After Undetectable AI v11 humanization, false\-negative rates on AI\-labeled rewrites exceed 96% for both Pangram and GPTZero, so humanized text routinely bypasses detectors despite the refund guarantee\. Domain\-level dispersion \(Tables[15](https://arxiv.org/html/2608.11256#A7.T15)and[18](https://arxiv.org/html/2608.11256#A7.T18)\) further shows that pooled benchmark accuracy can mask large field\-to\-field variance\.
## Appendix FDefinitions of text ratios
We compute several token\-level ratios from each abstract to probe whether surface properties correlate with detector scores\. Letttbe the abstract text \(a string\)\. We extract an ordered token sequenceT\(t\)=\(τ1,…,τn\)T\(t\)=\(\\tau\_\{1\},\\ldots,\\tau\_\{n\}\)using the regular expression
```
[A-Za-z0-9][A-Za-z0-9\-\+\./]*
```
, and letn=\|T\(t\)\|n=\|T\(t\)\|\. For convenience, we use an indicator function𝕀\{⋅\}\\mathbb\{I\}\\\{\\cdot\\\}that equals 1 when its condition holds and 0 otherwise\. Ifn=0n=0, the ratios below are undefined\.
#### Numeric token ratio\.
numeric\_token\_ratio\(t\)=1n∑i=1n𝕀\{∃c∈τi:c∈\{0,…,9\}\}\.\\mathrm\{numeric\\\_token\\\_ratio\}\(t\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbb\{I\}\\\{\\exists c\\in\\tau\_\{i\}:\\;c\\in\\\{0,\\ldots,9\\\}\\\}\.
#### Non\-alphabetic token ratio\.
nonalpha\_token\_ratio\(t\)=1n∑i=1n𝕀\{∃c∈τi:cis not a letter\}\.\\mathrm\{nonalpha\\\_token\\\_ratio\}\(t\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbb\{I\}\\\{\\exists c\\in\\tau\_\{i\}:\\;c\\text\{ is not a letter\}\\\}\.This captures tokens that include punctuation such as hyphens, periods, slashes, or plus signs \(as well as digits\)\.
#### Long token ratio\.
long\_token\_ratio\(t\)=1n∑i=1n𝕀\{\|τi\|≥10\}\.\\mathrm\{long\\\_token\\\_ratio\}\(t\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbb\{I\}\\\{\|\\tau\_\{i\}\|\\geq 10\\\}\.
#### Acronym ratio\.
acronym\_ratio\(t\)=1n∑i=1n𝕀\{τiis uppercase and\|τi\|≥2\}\.\\mathrm\{acronym\\\_ratio\}\(t\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbb\{I\}\\\{\\tau\_\{i\}\\text\{ is uppercase and \}\|\\tau\_\{i\}\|\\geq 2\\\}\.
#### AWL token ratio\.
LetAAbe Coxhead’s Academic Word List \(AWL\) headwords\(Coxhead,[2000](https://arxiv.org/html/2608.11256#bib.bib27)\)\(570 items\)\. We match tokens by Porter stemming against this list: for each tokenτ\\tau, we form a letters\-only normalizationalpha\(τ\)\\mathrm\{alpha\}\(\\tau\)by removing non\-letters, then compute its Porter stem\. LetS\(A\)S\(A\)be the set of Porter stems of the AWL headwords \(plus a small set of US spelling variants for headwords whose Porter stems differ, e\.g\.,analyse/analyze\)\.
awl\_token\_ratio\(t\)=1n∑i=1n𝕀\{stem\(alpha\(τi\)\)∈S\(A\)\}\.\\mathrm\{awl\\\_token\\\_ratio\}\(t\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbb\{I\}\\\{\\mathrm\{stem\}\(\\mathrm\{alpha\}\(\\tau\_\{i\}\)\)\\in S\(A\)\\\}\.
#### Type\-token ratio \(lexical diversity\)\.
LetV\(t\)=\{lower\(τ\):τ∈T\(t\)\}V\(t\)=\\\{\\mathrm\{lower\}\(\\tau\):\\tau\\in T\(t\)\\\}be the set of unique case\-folded tokens\. Then
ttr\(t\)=\|V\(t\)\|n\.\\mathrm\{ttr\}\(t\)=\\frac\{\|V\(t\)\|\}\{n\}\.
#### Word count\.
We report word counts using whitespace splitting:words\(t\)=\|split\(t\)\|\\mathrm\{words\}\(t\)=\|\\mathrm\{split\}\(t\)\|\.
#### Average words per sentence\.
We split sentences by punctuation \(\.,\!,?\) and compute the average number of whitespace\-delimited words per resulting sentence span\.
#### Short\-sentence ratio\.
Using the same sentence splits, letmmbe the number of sentences andwjw\_\{j\}the word count of sentencejj\. Then
short\_sentence\_ratio\(t\)=1m∑j=1m𝕀\{wj≤8\}\.\\mathrm\{short\\\_sentence\\\_ratio\}\(t\)=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\mathbb\{I\}\\\{w\_\{j\}\\leq 8\\\}\.We use eight words as a simple fragmentation proxy in the humanization mechanism analysis \(Appendix[M](https://arxiv.org/html/2608.11256#A13)\)\.
## Appendix GSupplementary numerical results
The tables below support the main\-text results by research question\. Tables[13](https://arxiv.org/html/2608.11256#A7.T13)through[18](https://arxiv.org/html/2608.11256#A7.T18)provide pooled and domain\-level Pangram summaries \(pre\- and post\-humanization\)\.
### G\.1\.Pangram score summaries
Tables[13](https://arxiv.org/html/2608.11256#A7.T13)and[16](https://arxiv.org/html/2608.11256#A7.T16)pool all domains; Tables[14](https://arxiv.org/html/2608.11256#A7.T14)through[18](https://arxiv.org/html/2608.11256#A7.T18)decompose the same statistics by domain\. SD captures score dispersion within each cell;*Outliers*counts scores outside1\.5×IQR1\.5\\times\\mathrm\{IQR\}fences \(same rule as boxplot fliers\)\. GPTZero and the LLM\-assisted baseline follow similar orderings\.
Pooled, before humanization\.
PeriodConditionMeanSD\[P25, P75\]nnOutliers2013 to 2015original0\.0020\.022\[0\.000, 0\.000\]305442013 to 2015refine \(abs\. only\)0\.6690\.423\[0\.143, 1\.000\]30602013 to 2015refine \(abs\.\+paper\)0\.8170\.351\[0\.927, 1\.000\]306682013 to 2015new \(article only\)0\.9750\.136\[1\.000, 1\.000\]306242023 to 2025original0\.1560\.353\[0\.000, 0\.001\]334742023 to 2025refine \(abs\. only\)0\.8270\.328\[0\.937, 1\.000\]336772023 to 2025refine \(abs\.\+paper\)0\.8400\.325\[0\.995, 1\.000\]336822023 to 2025new \(article only\)0\.9740\.142\[1\.000, 1\.000\]33629Table 13\.Pangram score summaries before humanization, pooled across domains \(mean, SD, \[IQR\]; outliers outside1\.5×IQR1\.5\\times\\mathrm\{IQR\}\)\.Domain variance before humanization\.Pre\-humanization dispersion is domain\-dependent and explains the STEM/non\-STEM contrasts in the main text\. On*original*abstracts in 2023 to 2025, political science and theology show the highest means and SD \(Table[15](https://arxiv.org/html/2608.11256#A7.T15)\), matching elevated flag rates in Table[4](https://arxiv.org/html/2608.11256#S3.T4); chemistry stays low on both metrics\. For LLM rewrites, outliers concentrate in high\-scoring conditions \(*refine \(abstract \+ paper\)*,*new \(article only\)*\) where the IQR sits near 1\.0, while*refine \(abstract only\)*often has SD comparable across domains but wider score spread in computer science \(lower quartile\) than theology \(tight upper quartile\)\.
DomainConditionMeanSD\[P25, P75\]nnOutliersChemistryoriginal0\.0030\.020\[0\.000, 0\.000\]8412refine \(abs\. only\)0\.6980\.413\[0\.250, 1\.000\]840refine \(abs\.\+paper\)0\.8100\.358\[0\.921, 1\.000\]8419new \(article only\)0\.9530\.195\[1\.000, 1\.000\]849Computer scienceoriginal0\.0000\.000\[0\.000, 0\.000\]699refine \(abs\. only\)0\.5670\.438\[0\.040, 1\.000\]700refine \(abs\.\+paper\)0\.6870\.442\[0\.074, 1\.000\]700new \(article only\)0\.9540\.177\[1\.000, 1\.000\]709Political scienceoriginal0\.0050\.039\[0\.000, 0\.000\]7814refine \(abs\. only\)0\.7170\.406\[0\.305, 1\.000\]780refine \(abs\.\+paper\)0\.9230\.217\[1\.000, 1\.000\]7818new \(article only\)0\.9940\.052\[1\.000, 1\.000\]784Theologyoriginal0\.0000\.000\[0\.000, 0\.000\]7411refine \(abs\. only\)0\.6820\.431\[0\.124, 1\.000\]740refine \(abs\.\+paper\)0\.8350\.325\[0\.902, 1\.000\]7415new \(article only\)1\.0000\.001\[1\.000, 1\.000\]742Table 14\.Pangram scores before humanization by domain \(2013 to 2015\)\.DomainConditionMeanSD\[P25, P75\]nnOutliersChemistryoriginal0\.0260\.146\[0\.000, 0\.000\]9516refine \(abs\. only\)0\.7770\.360\[0\.482, 1\.000\]950refine \(abs\.\+paper\)0\.8200\.341\[0\.863, 1\.000\]9518new \(article only\)0\.9410\.215\[1\.000, 1\.000\]9512Computer scienceoriginal0\.1000\.290\[0\.000, 0\.000\]8214refine \(abs\. only\)0\.7360\.392\[0\.426, 1\.000\]830refine \(abs\.\+paper\)0\.7040\.399\[0\.339, 1\.000\]830new \(article only\)0\.9640\.165\[1\.000, 1\.000\]8313Political scienceoriginal0\.2480\.430\[0\.000, 0\.216\]7819refine \(abs\. only\)0\.8770\.273\[0\.980, 1\.000\]7917refine \(abs\.\+paper\)0\.8920\.266\[1\.000, 1\.000\]7917new \(article only\)1\.0000\.002\[1\.000, 1\.000\]792Theologyoriginal0\.2800\.433\[0\.000, 0\.838\]790refine \(abs\. only\)0\.9330\.214\[0\.998, 1\.000\]7917refine \(abs\.\+paper\)0\.9560\.196\[1\.000, 1\.000\]795new \(article only\)0\.9990\.010\[1\.000, 1\.000\]792Table 15\.Pangram scores before humanization by domain \(2023 to 2025\)\.Pooled and domain summaries after humanization:After humanization, pooled means fall to≈0\.03\\approx 0\.03\(Table[16](https://arxiv.org/html/2608.11256#A7.T16)\), consistent with near\-universal false negatives in Table[2](https://arxiv.org/html/2608.11256#S3.T2)\. Residual variance remains in non\-STEM domains: political science and theology still show non\-zero SD and outliers on some conditions \(Tables[17](https://arxiv.org/html/2608.11256#A7.T17)and[18](https://arxiv.org/html/2608.11256#A7.T18)\), indicating incomplete score collapse for a subset of abstracts rather than uniform evasion\.
PeriodConditionMeanSD\[P25, P75\]nnOutliers2013 to 2015original0\.0030\.057\[0\.000, 0\.000\]30612013 to 2015refine \(abs\. only\)0\.0330\.178\[0\.000, 0\.000\]306102013 to 2015refine \(abs\.\+paper\)0\.0130\.114\[0\.000, 0\.000\]30642013 to 2015new \(article only\)0\.0000\.000\[0\.000, 0\.000\]30602023 to 2025original0\.0330\.178\[0\.000, 0\.000\]336112023 to 2025refine \(abs\. only\)0\.0360\.186\[0\.000, 0\.000\]336132023 to 2025refine \(abs\.\+paper\)0\.0330\.178\[0\.000, 0\.000\]336112023 to 2025new \(article only\)0\.0330\.178\[0\.000, 0\.000\]33611Table 16\.Pangram score summaries after humanization, pooled across domains\.DomainConditionMeanSD\[P25, P75\]nnOutliersChemistryoriginal0\.0000\.000\[0\.000, 0\.000\]840refine \(abs\. only\)0\.0000\.000\[0\.000, 0\.000\]840refine \(abs\.\+paper\)0\.0000\.000\[0\.000, 0\.000\]840new \(article only\)0\.0000\.000\[0\.000, 0\.000\]840Computer scienceoriginal0\.0140\.120\[0\.000, 0\.000\]701refine \(abs\. only\)0\.0140\.120\[0\.000, 0\.000\]701refine \(abs\.\+paper\)0\.0000\.000\[0\.000, 0\.000\]700new \(article only\)0\.0000\.000\[0\.000, 0\.000\]700Political scienceoriginal0\.0000\.000\[0\.000, 0\.000\]780refine \(abs\. only\)0\.0640\.247\[0\.000, 0\.000\]785refine \(abs\.\+paper\)0\.0260\.159\[0\.000, 0\.000\]782new \(article only\)0\.0000\.000\[0\.000, 0\.000\]780Theologyoriginal0\.0000\.000\[0\.000, 0\.000\]740refine \(abs\. only\)0\.0540\.228\[0\.000, 0\.000\]744refine \(abs\.\+paper\)0\.0270\.163\[0\.000, 0\.000\]742new \(article only\)0\.0000\.000\[0\.000, 0\.000\]740Table 17\.Pangram scores after humanization by domain \(2013 to 2015\)\.DomainConditionMeanSD\[P25, P75\]nnOutliersChemistryoriginal0\.0000\.000\[0\.000, 0\.000\]950refine \(abs\. only\)0\.0000\.000\[0\.000, 0\.000\]950refine \(abs\.\+paper\)0\.0000\.000\[0\.000, 0\.000\]950new \(article only\)0\.0000\.000\[0\.000, 0\.000\]950Computer scienceoriginal0\.0240\.154\[0\.000, 0\.000\]832refine \(abs\. only\)0\.0120\.110\[0\.000, 0\.000\]831refine \(abs\.\+paper\)0\.0240\.154\[0\.000, 0\.000\]832new \(article only\)0\.0000\.000\[0\.000, 0\.000\]830Political scienceoriginal0\.0760\.267\[0\.000, 0\.000\]796refine \(abs\. only\)0\.0630\.245\[0\.000, 0\.000\]795refine \(abs\.\+paper\)0\.0760\.267\[0\.000, 0\.000\]796new \(article only\)0\.0760\.267\[0\.000, 0\.000\]796Theologyoriginal0\.0380\.192\[0\.000, 0\.000\]793refine \(abs\. only\)0\.0780\.267\[0\.000, 0\.000\]797refine \(abs\.\+paper\)0\.0380\.192\[0\.000, 0\.000\]793new \(article only\)0\.0630\.245\[0\.000, 0\.000\]795Table 18\.Pangram scores after humanization by domain \(2023 to 2025\)\.
### G\.2\.Corpus counts
See Table[1](https://arxiv.org/html/2608.11256#S2.T1)in the main text\.
## Appendix HFull text\-feature correlation results
Tables[19](https://arxiv.org/html/2608.11256#A8.T19)and[20](https://arxiv.org/html/2608.11256#A8.T20)report the full pre\-humanization text\-feature correlation results\. Correlations are Spearman’sρ\\rhobetween feature ratios and detector AI score\. The domain\-adjusted column reports correlation after centering within domain\.qqis the Benjamini\-Hochberg FDR across the 10 tests per collection\.
FeatureDetectorρ\\rhoppqqDomain\-adj\.rrNumeric token ratioPangram\-0\.146<0\.001<0\.001<0\.001<0\.001\-0\.148Numeric token ratioGPTZero\-0\.171<0\.001<0\.001<0\.001<0\.001\-0\.098Non\-alphabetic token ratioPangram\-0\.149<0\.001<0\.001<0\.001<0\.001\-0\.085Non\-alphabetic token ratioGPTZero\-0\.197<0\.001<0\.001<0\.001<0\.001\-0\.109Long token ratioPangram0\.336<0\.001<0\.001<0\.001<0\.0010\.421Long token ratioGPTZero0\.304<0\.001<0\.001<0\.001<0\.0010\.321Acronym ratioPangram\-0\.0550\.0440\.044\-0\.047Acronym ratioGPTZero\-0\.0820\.0040\.004\-0\.024AWL token ratioPangram0\.346<0\.001<0\.001<0\.001<0\.0010\.412AWL token ratioGPTZero0\.320<0\.001<0\.001<0\.001<0\.0010\.290
Table 19\.Text\-feature correlations \(2023 to 2025 collection;n=1342n=1342\)\.FeatureDetectorρ\\rhoppqqDomain\-adj\.rrNumeric token ratioPangram\-0\.0570\.0450\.045\-0\.120Numeric token ratioGPTZero\-0\.0740\.0090\.009\-0\.059Non\-alphabetic token ratioPangram0\.0050\.8690\.869\-0\.076Non\-alphabetic token ratioGPTZero\-0\.119<0\.001<0\.001<0\.001<0\.001\-0\.045Long token ratioPangram0\.355<0\.001<0\.001<0\.001<0\.0010\.366Long token ratioGPTZero0\.287<0\.001<0\.001<0\.001<0\.0010\.304Acronym ratioPangram\-0\.0040\.8870\.8870\.002Acronym ratioGPTZero\-0\.0350\.1990\.199\-0\.007AWL token ratioPangram0\.331<0\.001<0\.001<0\.001<0\.0010\.400AWL token ratioGPTZero0\.258<0\.001<0\.001<0\.001<0\.0010\.277
Table 20\.Text\-feature correlations \(2013 to 2015 collection;n=1222n=1222\)\.
## Appendix ILength\-partial correlations
Table[21](https://arxiv.org/html/2608.11256#A9.T21)reports Spearmanρ\\rhobetween detector scores and surface features after residualizing both variables onlog\(1\+word count\)\\log\(1\+\\text\{word count\}\)\(2023 to 2025, pre\-humanization\)\.
FeatureDetectorRawρ\\rhoPartialρ\\rhoAWL token ratioPangram0\.3460\.337AWL token ratioGPTZero0\.3200\.261Long token ratioPangram0\.3360\.321Long token ratioGPTZero0\.3040\.276Numeric token ratioPangram−0\.146\-0\.146−0\.132\-0\.132Numeric token ratioGPTZero−0\.171\-0\.171−0\.134\-0\.134Table 21\.Length\-partial Spearmanρ\\rho\(2023 to 2025, pre\-humanization\)\.
## Appendix JDetector score distributions
Table[2](https://arxiv.org/html/2608.11256#S3.T2)reports error rates at a single threshold; the figures below show the underlying score distributions in the 2023 to 2025 collection\.
Pre\-humanization separation \(RQ1\)\.Figure[4](https://arxiv.org/html/2608.11256#A10.F4)plots Pangram scores by domain and rewrite condition before humanization\. This figure is central to RQ1 because it makes the assisted\-writing risk visible as a distributional shift:*original*abstracts cluster near zero, while even*refine \(abstract only\)*edits move substantial mass toward high scores \- the pattern behind proxy false\-positive risk on light editing\. It also shows that*new \(article only\)*scores sit near certainty, explaining why false\-negative rates on full\-synthesis rewrites are near zero before humanization\.
Figure 4\.Pangram scores before humanization \(2023 to 2025\), by domain \(rows\) and rewrite condition \(columns\)\. Originals cluster low;*refine \(abstract only\)*and heavier rewrites shift right\.Post\-humanization collapse \(RQ3\)\.Figure[5](https://arxiv.org/html/2608.11256#A10.F5)uses the same layout after Undetectable humanization\. Read together with Figure[4](https://arxiv.org/html/2608.11256#A10.F4), it is the main visual evidence for RQ3: AI\-labeled variants that scored near 1\.0 pre\-humanization collapse toward 0\.0 post\-humanization, which drives the\>96%\>96\\%false\-negative rates in Table[2](https://arxiv.org/html/2608.11256#S3.T2)and shows that commercial evasion can erase detector signal at scale\.
Figure 5\.Pangram scores after humanization \(2023 to 2025\)\. Compare to Figure[4](https://arxiv.org/html/2608.11256#A10.F4): high pre\-humanization scores on AI\-labeled variants largely disappear\.
## Appendix KRobustness and threshold sensitivity
This section complements Table[2](https://arxiv.org/html/2608.11256#S3.T2)by asking how detector behavior changes with the scoring thresholdτ\\tauand whether standard precision\-recall \(PR\) summaries align with academic\-integrity policy\. ROC and PR curves in the main text and in Figure[6](https://arxiv.org/html/2608.11256#A11.F6)use2013 to 2015pre\-humanization scores so that*original*abstracts can serve as negatives with interpretable proxy FPR\. Figures[3](https://arxiv.org/html/2608.11256#S3.F3)and[7](https://arxiv.org/html/2608.11256#A11.F7)repeat the same tasks on 2023 to 2025, where original\-axis rates are flag rates \(unobserved real\-world AI use\)\. Threshold trade\-off plots use 2023 to 2025 throughout\.
### K\.1\.Why we do not report pooled precision\-recall
A common evaluation pools*original*abstracts as negatives and*all*LLM rewrites \(refine, refine with paper, and new\) as positives\. Under that label scheme, both Pangram and GPTZero achieve AUC\-PR≈0\.96\\approx 0\.96because*new \(article only\)*outputs score near 1\.0 and dominate the positive class\. That number suggests excellent separability but answers the wrong question for integrity offices: policy concern is whether a detector can distinguishhuman drafts from light AI editing, not whether it can detect full\-synthesis rewrites that already saturate the score scale\. We therefore report PR only for the assisted\-editing subset below, plus a threshold trade\-off plot that tracks three rates institutions actually care about\.
### K\.2\.Assisted\-editing ROC and precision\-recall
We restrict to paired rows on the same paper:*original*abstract==negative \(proxy human\),*refine \(abstract only\)*==positive \(proxy light AI edit on human\-source text\)\. Scores are detector outputs in\[0,1\]\[0,1\]; a text is flagged whens≥τs\\geq\\tau\. Figure[2](https://arxiv.org/html/2608.11256#S3.F2)in the main text shows ROC only \(2013 to 2015\); FigureLABEL:fig:assisted\_editing\_roc\_pradds the matching precision\-recall panel\.*Original*is not its own curve\. It defines the false\-positive rate axis\.
### K\.3\.ROC and PR by rewrite condition \(2013 to 2015\)
Figure[6](https://arxiv.org/html/2608.11256#A11.F6)repeats the analysis for each rewrite type against the same pre\-LLM\-era originals:*refine \(abstract only\)*,*refine \(abstract \+ paper\)*, and*new \(article only\)*, plus a dashed*pooled AI*curve \(all three LLM conditions as positives\)\. AUC\-PR is high \(0\.930\.93\-1\.001\.00\) because Gemini rewrites separate cleanly from unflagged originals atτ=0\.50\\tau=0\.50; this is a best\-case separability benchmark, not current flagging exposure\.
Figure 6\.ROC and precision\-recall by rewrite condition vs\.*original*\(2013 to 2015, pre\-humanization\)\. Four curves per panel: three single\-condition tasks \(solid\) and pooled all LLM rewrites \(dashed\)\. Rows: Pangram and GPTZero\. Markers:τ=0\.50\\tau=0\.50\.
### K\.4\.2023 to 2025 replication \(flag rate on originals\)
On recent abstracts, positives on the original axis areflag rates, not confirmed false positives\. Pangram AUC\-PR on assisted editing falls to0\.880\.88; atτ=0\.50\\tau=0\.50it flags 80% of light edits with precision0\.840\.84while 15% of originals are also flagged\. GPTZero AUC\-PR is0\.850\.85with 49% refine recall\. Pooled\-AI AUC\-PR remains≈0\.96\\approx 0\.96\(Pangram\) and0\.950\.95\(GPTZero\); the assisted\-editing panel is in Figure[3](https://arxiv.org/html/2608.11256#S3.F3)\(main text\); per\-condition curves are in Figure[7](https://arxiv.org/html/2608.11256#A11.F7)\.
Figure 7\.Rewrite\-condition ROC/PR on 2023 to 2025 \(pre\-humanization\); same layout as Figure[6](https://arxiv.org/html/2608.11256#A11.F6)\.
### K\.5\.Threshold trade\-off curves
For each detector and eachτ∈\{0\.3,0\.4,0\.5,0\.6,0\.7\}\\tau\\in\\\{0\.3,0\.4,0\.5,0\.6,0\.7\\\}, we report three percentages on the 2023 to 2025 collection:
1. \(1\)Original flag rate\(gray\): fraction of unmodified*original*abstracts withs≥τs\\geq\\tau\. On recent prose this mixes proxy false positives with unobserved real AI use; it still bounds how often human\-labeled text is flagged\.
2. \(2\)Refine \(abstract only\) flagged rate\(solid blue/red\): fraction of light edits flagged\-the assisted\-writing exposure from RQ1\.
3. \(3\)FNR on AI\-labeled pool\(dashed orange\): fraction of pooled LLM rewrites \(all three conditions\) scoring*below*τ\\tau\-missed AI under proxy labels\.
There is noτ\\tauthat simultaneously keeps all three rates low\. For Pangram, original flag rate and refine\-flag rate stay near 15% and 80% respectively acrossτ∈\[0\.4,0\.6\]\\tau\\in\[0\.4,0\.6\]; tighteningτ\\tauslightly reduces missed full rewrites \(FNR falls from 10\.5% to 13\.5%\) but does not make light edits safe from flags\. For GPTZero, refine\-flag rate is already near 50% atτ=0\.50\\tau=0\.50while FNR on the AI pool remains 38%; loweringτ\\tauincreases flags on originals without catching most light edits\. The vertical dotted line marks our main\-text operating pointτ=0\.50\\tau=0\.50\.
The table reports the same three quantities atτ∈\{0\.4,0\.5,0\.6\}\\tau\\in\\\{0\.4,0\.5,0\.6\\\}; the figure extends the sweep to 0\.3 and 0\.7 and places Pangram and GPTZero side by side\.
Figure 8\.Threshold trade\-off \(2023 to 2025, pre\-humanization\)\. Three policy rates vs\.τ\\taufor Pangram \(left\) and GPTZero \(right\)\. Dotted vertical line:τ=0\.50\\tau=0\.50\.
### K\.6\.Tabulated threshold sensitivity
Table[5](https://arxiv.org/html/2608.11256#S3.T5)in the main text reports original flag rates and AI\-labeled false\-negative rates forτ∈\{0\.4,0\.5,0\.6\}\\tau\\in\\\{0\.4,0\.5,0\.6\\\}in the 2023 to 2025 collection \(pre\-humanization\)\. Figure[8](https://arxiv.org/html/2608.11256#A11.F8)extends the sweep toτ∈\{0\.3,…,0\.7\}\\tau\\in\\\{0\.3,\\ldots,0\.7\\\}\.
## Appendix LAutomatically selected examples
We provide two*automatically selected*examples illustrating extreme false\-negative transitions after humanization \(pre scores near 1\.0, post scores near 0\.0\) without manual cherry\-picking\. For each time window \(2013 to 2015 and 2023 to 2025\), we selected one paper\-variant pairdeterministicallyfrom items with complete pre/post outputs\. Specifically, we \(i\) restricted to pairs with both Pangram and GPTZero available when possible, \(ii\) computed the mean detector score change after humanization\(Δ=post−pre\)\(\\Delta=\\mathrm\{post\}\-\\mathrm\{pre\}\)across the available detectors for each paper\-variant pair, and \(iii\) selected the pair with thelargest drop\(most negativeΔ\\Delta\)\. This selection rule is fully specified in our analysis script and does not use randomness\.
The 2023 to 2025 auto\-selected example \(paper IDW4401907927, political science,*new \(article only\)*\) shows a complete detector flip: Pangram changed from1\.0001\.000to0\.0000\.000and GPTZero changed from1\.0001\.000to0\.0000\.000after humanization \(meanΔ=−1\.000\\Delta=\-1\.000across detectors\)\. The 2013 to 2015 auto\-selected example \(W3009617460, political science,*refine \(abstract \+ paper\)*\) shows the same pattern\.
To diagnose why the flip occurs, we compared measurable pre/post text properties for the selected examples\. In the 2023 to 2025 case, long\-token ratio dropped from0\.2600\.260to0\.1000\.100, Academic Word List \(AWL\) token ratio dropped from0\.2520\.252to0\.0060\.006, and lexical diversity \(type\-token ratio\) dropped from0\.7710\.771to0\.5020\.502, while the rewrite became much longer \(129 to 311 words\) and shifted to much longer sentence units \(average 25\.8 to 62\.2 words per sentence\)\. In the 2013 to 2015 case, long\-token ratio also decreased \(0\.082→0\.0540\.082\\rightarrow 0\.054\) and lexical diversity decreased \(0\.629→0\.5490\.629\\rightarrow 0\.549\)\. These changes are directionally consistent with our corpus\-level findings that detectors assign higher AI scores to texts with more long technical tokens, and lower scores to texts that reduce those signals\. In short, the humanizer appears to suppress detector\-salient cues rather than preserve the original stylistic profile\.
## Appendix MLinguistic mechanisms of humanization
Undetectable AI v11 is a proprietary service; we treat it as a black box but measure how its outputs differ linguistically from the inputs sent for humanization\. For every paper\-variant pair with bothoriginal\_abstractandhumanized\_abstractunderhumanization/\(n=2,568n=2\{,\}568pairs across both time collections\), we computed the surface features defined in Appendix[F](https://arxiv.org/html/2608.11256#A6), plus the fraction of sentences with at most eight words \(*short\-sentence ratio*\) as a simple fragmentation proxy\. We tested paired shifts \(post minus pre\) with two\-sided permutation tests \(5,000 iterations; seed 42\) and correlated feature deltas with Pangram and GPTZero score deltas on matched rows\.
Systematic suppression of detector\-salient cues\.Table[8](https://arxiv.org/html/2608.11256#S3.T8)in the main text summarizes key pooled shifts; Table[22](https://arxiv.org/html/2608.11256#A13.T22)below lists the full feature set\. Humanization*reduced*long\-token ratio \(0\.184 to 0\.138; meanΔ=−0\.045\\Delta=\-0\.045; 88\.2% of pairs decreased;p<0\.001p<0\.001\), AWL token ratio \(0\.157 to 0\.121;Δ=−0\.036\\Delta=\-0\.036; 82\.9% decreased;p<0\.001p<0\.001\), and type\-token ratio \(0\.700 to 0\.605;Δ=−0\.094\\Delta=\-0\.094; 93\.0% decreased;p<0\.001p<0\.001\)\. Numeric token ratio did not shift materially \(p=0\.83p=0\.83\)\. These directions match our pre\-humanization finding that higher Pangram and GPTZero scores correlate with greater long\-token and AWL density \(Appendix[H](https://arxiv.org/html/2608.11256#A8)\): the humanizer moves text along the same axes that separate high\-scoring from low\-scoring abstracts\.
Length and syntax\.Humanization also*increased*length: mean word count rose from 180 to 229 \(Δ≈\+49\\Delta\\approx\+49words;p<0\.001p<0\.001\) and average words per sentence from 23\.6 to 28\.4 \(Δ≈\+4\.8\\Delta\\approx\+4\.8;p<0\.001p<0\.001\)\. Short\-sentence ratio increased only slightly \(0\.063 to 0\.072;p<0\.001p<0\.001\), so evasion is not primarily driven by chopping prose into many brief sentences\. Instead, outputs tend to be longer and more sententially expanded while using fewer long technical tokens and fewer academic\-wordlist hits\.
Coupling to detector score drops\.When a pair’s long\-token, AWL, or type\-token ratios fell more sharply, Pangram and GPTZero scores tended to fall more sharply as well \(Spearmanρ≈0\.30\\rho\\approx 0\.30to0\.400\.40onΔ\\Deltafeatures vs\.Δ\\Deltascores;p<0\.001p<0\.001for all three features on both detectors\)\. Average words\-per\-sentence increases were weakly associated with smaller score drops on Pangram \(ρ=−0\.12\\rho=\-0\.12\) and not reliably on GPTZero \(ρ=−0\.05\\rho=\-0\.05\)\. Together, the corpus\-level shifts and the score\-feature couplings suggest a mechanistic story: Undetectable humanization neutralizes commercial detectors largely by diluting the surface academic and technical token patterns those detectors already treat as AI\-like, not by uniformly fragmenting syntax\.
FeaturePrePostΔ\\Delta% decreasedppLong\-token ratio0\.1840\.138−0\.045\-0\.04588\.2%<0\.001<0\.001AWL token ratio0\.1570\.121−0\.036\-0\.03682\.9%<0\.001<0\.001Type\-token ratio0\.7000\.605−0\.094\-0\.09493\.0%<0\.001<0\.001Numeric token ratio0\.0170\.0170\.0000\.00038\.9%0\.833Non\-alphabetic token ratio0\.0790\.068−0\.011\-0\.01175\.2%<0\.001<0\.001Avg\. words per sentence23\.628\.4\+4\.8\+4\.822\.2%<0\.001<0\.001Short\-sentence ratio \(≤8\\leq 8words\)0\.0630\.072\+0\.009\+0\.00916\.1%<0\.001<0\.001Word count180229\+49\+4910\.2%<0\.001<0\.001Table 22\.Full pre/post linguistic features after Undetectable AI v11 humanization \(pooled corpus;n=2,568n=2\{,\}568paper\-variant pairs\)\.Δ\\Deltais post minus pre\.ppis a two\-sided permutation test on pairedΔ\\Delta\.
## Appendix NCost and scope of the pipeline
Our design trades breadth across rewrite*intensity*and detectors for feasibility at corpus scale\. We sampled 800 English abstracts \(642 after length and PDF filtering; Section[2\.2](https://arxiv.org/html/2608.11256#S2.SS2)and AppendixLABEL:app:time\_periods\)\. Each retained paper contributes four scored texts: one*original*and three LLM outputs \(*refine \(abstract only\)*,*refine \(abstract \+ paper\)*,*new \(article only\)*\)\.
Why one rewrite model:We used a single generator \(Gemini 3 Flash via OpenRouter\) rather than a multi\-model sweep\. AddingKKrewrite models would multiply LLM inference cost byKKwhile holding our error\-analysis questions fixed \(proxy labels, detector comparison, humanization\)\. The factorial structure we need is rewrite*condition*and detector, not vendor\-specific rewrite style\.
Order\-of\-magnitude call accounting:LetNNdenote retained papers \(N=642N=642\)\. Per paper:
- •LLM rewrites:33calls \(one per rewrite condition; the original is not regenerated\)\.
- •Pre\-humanization detection:44variants×\\times33detectors \(Pangram, GPTZero, LLM\-assisted baseline\)=12=12scored texts per paper\.
- •Humanization:44variants×\\times11Undetectable pass \(paid credits per abstract; not repeated per detector\)\.
- •Post\-humanization detection:44variants×\\times33detectors \(Pangram, GPTZero, and LLM\-assisted baseline\)\.
Aggregating over the corpus, pre\-detection scales asN×4×3N\\times 4\\times 3, humanization asN×4N\\times 4, and post\-detection asN×4×3N\\times 4\\times 3\.
LLM\-assisted baseline role:The LLM\-assisted baseline flags a high share of recent*original*abstracts \(Table[2](https://arxiv.org/html/2608.11256#S3.T2)\) and is treated as supplementary\. After humanization, Pangram and GPTZero scores collapse toward zero for AI\-labeled variants, but the LLM\-assisted baseline still flags most humanized rewrites \(post\-humanization FNR 30\.5% and 19\.7%\)\. Main\-text evasion claims, therefore, focus on commercial detectors \(RQ3\)\.Similar Articles
AI detectors are creating a new era of distrust
This article examines the growing use of AI detectors in education and the resulting atmosphere of distrust, highlighting concerns about false positives and the subjective nature of AI-written text detection.
Accusatory AI: How a Widespread Misuse of AI Technology Is Harming Students
The article discusses how AI-detection tools are unreliable and are being misused by educators to punish students, leading to false cheating accusations. It emphasizes that such tools should not be used as sole evidence for punishment and highlights the need for transparency and proper assessment.
Top AI conference uses AI detector to reject papers for allegedly being written by AI
NeurIPS 2026 used a proprietary AI-text detector to desk-reject papers for alleged AI policy violations without validating it on the target distribution; the same detector later flagged conference chairs' own papers as likely AI-written.
AI might make me fail my class
A student expresses frustration that their entirely human-written paper was flagged as AI-generated by plagiarism checkers, highlighting the flaws of current AI detection tools in academic settings.
Student cheating now impossible to detect
The article discusses how advancements in AI have made it virtually impossible to detect student cheating, as AI-generated content becomes indistinguishable from human work.