A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID

arXiv cs.CL Papers

Summary

The paper investigates neurons in frozen BERT that drive AI-text detection using sparse probing and activation patching on the RAID benchmark, identifying a small set of causally relevant neurons that generalize across generator families.

arXiv:2609.30287v1 Announce Type: new Abstract: AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers x 768), which we call neurons. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full-feature detection accuracy. Bidirectional activation patching confirms this set's causal relevance: in both directions it flips predictions an order of magnitude more often than size-matched random sets. Mean-ablating the same neurons leaves accuracy largely intact; the signal is therefore redundantly distributed. Cross-generator analysis reveals a bipartite structure: instruction-tuned generators concentrate 30-36% of stable neurons in BERT's final layer while both base generators fall below 14%, consistent with a layer-12 footprint of post-training alignment. Leave-one-family-out evaluation shows the selected neurons retain 86-94% of the full-feature ceiling on unseen generator families, so a detector can operate on a small fixed subspace without re-identifying neurons per generator.
Original Article
View Cached Full Text

Cached at: 09/28/26, 09:32 AM

# A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT:Sparse Probing and Activation Patching on RAID
Source: [https://arxiv.org/html/2609.30287](https://arxiv.org/html/2609.30287)
Miłosz GrunwaldAffiliation:Gradient PGAffiliation:Gdańsk University of Technology

###### Abstract

AI\-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood\. We study which neurons in afrozenBERT\-base\-uncased encoder support AI\-text detection, using the RAID benchmark across six generators spanning pure\-base and instruction\-tuned models\. We apply the L1→\\toL2 sparse\-probing protocol of[Gurnee et al\. \(2023\)](https://arxiv.org/html/2609.30287#bib.bib14)to all 9,216 CLS hidden\-state dimensions \(12 layers×\\times768\), which we callneurons\. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full\-feature detection accuracy\. Bidirectional activation patching confirms this set’s causal relevance: in both directions it flips predictions an order of magnitude more often than size\-matched random sets\. Mean\-ablating the same neurons leaves accuracy largely intact; the signal is therefore redundantly distributed\. Cross\-generator analysis reveals a bipartite structure: instruction\-tuned generators concentrate 30–36% of stable neurons in BERT’s final layer while both base generators fall below 14%, consistent with a layer\-12 footprint of post\-training alignment\. Leave\-one\-family\-out evaluation shows the selected neurons retain 86–94% of the full\-feature ceiling on unseen generator families, so a detector can operate on a small fixed subspace without re\-identifying neurons per generator\.

## 1Introduction

Input textshuman / AIBERT\-base\-uncased\(frozen\)𝐳∈ℝ9,216\\mathbf\{z\}\\\!\\in\\\!\\mathbb\{R\}^\{9\{,\}216\}CLS, layers 1–12L1 probeC=0\.005C\\\!=\\\!0\.005Stable set𝒮∗\\mathcal\{S\}^\{\*\}45–62 neurons≥80%\{\\geq\}80\\%of 15 cells5\-fold×\\times3 seeds = 15 cellsMean ablationz𝒮∗←μ¯trainz\_\{\\mathcal\{S\}^\{\*\}\}\\leftarrow\\bar\{\\mu\}\_\{\\mathrm\{train\}\}L2 eval probe\(frozen\)Acc\. drop≤0\.11\\leq\\\!0\.11pp5/6; 1\.06 pp Cohere§[5\.2](https://arxiv.org/html/2609.30287#S5.SS2)necessity testActivation patchingbidirectionalL2 eval probe\(frozen\)Fwd: 1–8% \(10–16×\\times\)Rev: 0\.65–5\.74% \(7–12×\\times\)§[5\.1](https://arxiv.org/html/2609.30287#S5.SS1)sufficiency test

Figure 1:Overview of the experimental pipeline\. Frozen BERT\-base\-uncased CLS activations are concatenated across all 12 layers, yielding𝐳∈ℝ9,216\\mathbf\{z\}\\in\\mathbb\{R\}^\{9\{,\}216\}\. An L1\-regularised probe selects a candidate set per \(fold, seed\) cell; neurons stable in≥80%\{\\geq\}80\\%of all 15 cells become thestable set𝒮∗\\mathcal\{S\}^\{\*\}\. An independent L2 probe trained on all 9,216 neurons is held frozen as the decision function for both experiments\.Left branch\(§[5\.2](https://arxiv.org/html/2609.30287#S5.SS2)\): mean ablation\.Right branch\(§[5\.1](https://arxiv.org/html/2609.30287#S5.SS1)\): bidirectional activation patching, forward \(AI→\\tohuman\) and reverse \(human→\\toAI\)\. Numerical results are reported in the cited sections\.The rapid development of Large Language Models \(LLMs\) has made AI\-generated text increasingly difficult to distinguish from human writing[Dugan et al\. \(2024\)](https://arxiv.org/html/2609.30287#bib.bib8)\. This poses a significant challenge across domains ranging from academic integrity to social media moderation and journalistic authenticity\. Although detection accuracy has improved substantially, with encoder\-based classifiers reaching high performance on established benchmarks[Dugan et al\. \(2024\)](https://arxiv.org/html/2609.30287#bib.bib8), a fundamental question remains unanswered:What internal representations drive these predictions?Mechanistic interpretability\([Elhage et al\., 2021](https://arxiv.org/html/2609.30287#bib.bib10);[Bereska and Gavves, 2024](https://arxiv.org/html/2609.30287#bib.bib2)\)has seen little application to AI\-text detection, leaving the field without a causal explanation of how detection works\.

Existing interpretability work on AI\-text detection falls into five strands: \(a\)black\-box feature attributionvia SHAP\([Lundberg and Lee, 2017](https://arxiv.org/html/2609.30287#bib.bib26)\)or LIME\([Ribeiro et al\., 2016](https://arxiv.org/html/2609.30287#bib.bib35)\), which identifies predictive tokens without inspecting model internals[Pudasaini et al\. \(2026\)](https://arxiv.org/html/2609.30287#bib.bib32); \(b\)confounder removal, which suppresses neurons correlated with domain or style artefacts in fine\-tuned encoders to improve out\-of\-distribution accuracy[Borile and Abrate \(2025\)](https://arxiv.org/html/2609.30287#bib.bib3); \(c\)sparse autoencoder feature analysisapplied to decoder\-based models[Kuznetsov et al\. \(2025\)](https://arxiv.org/html/2609.30287#bib.bib22); \(d\)probing on fine\-tuned detectors, which reveals reliance on narrow lexical or structural cues\([Choi and Zu, 2025](https://arxiv.org/html/2609.30287#bib.bib4)\); and \(e\)topological analysisof attention maps, characterising surface and syntactic properties that differentiate AI from human text\([Kushnareva et al\., 2021](https://arxiv.org/html/2609.30287#bib.bib21)\)\. These approaches share a fundamental limitation documented across neuron\-level analysis broadly\([Sajjad et al\., 2022](https://arxiv.org/html/2609.30287#bib.bib37)\): they areobservational\. They identify neurons that correlate with detector behaviour or remove neurons that hurt generalisation, but none directly tests whether a small set of internal representations iscausally sufficientto produce AI\-text predictions\.

We address this gap with a causal intervention study on afrozenBERT\-base\-uncased; Figure[1](https://arxiv.org/html/2609.30287#S1.F1)summarises the experimental pipeline\. Rather than fine\-tuning the model for detection, we train sparse linear probes on pretrained CLS\-token activations concatenated across all 12 layers \(9,216 neurons total\), following the sparse\-probing protocol of[Gurnee et al\. \(2023\)](https://arxiv.org/html/2609.30287#bib.bib14)\. An L1\-regularised logistic regression selects a per\-cell candidate set; neurons that survive across folds and seeds form the aggregated stable set𝒮∗\\mathcal\{S\}^\{\*\}\. An independent L2 probe serves as the frozen evaluator for all subsequent experiments\. We then apply two causal interventions to the same selected neuron set:mean ablation, which replaces these neurons with their training\-set means to test necessity; andsame\-domain bidirectional activation patching, run in both the forward direction \(AI donor→\\tohuman target, testing sufficiency\) and the reverse direction \(human donor→\\toAI target, testing partial necessity on\-manifold\)\. The intervention orientation is the key distinction from prior work: rather than suppressing suspected confounders, we inject donor activations and measure whether predictions flip\.

Across six generators spanning two pure\-base and four instruction\-tuned families, bidirectional activation patching shows the selected neurons are sufficient to drive the probe’s AI predictions \(forward, 10–16×\\timesabove random\) and partially necessary to preserve them \(reverse, 7–12×\\timesabove random\)\. Mean\-ablation of the same neurons leaves accuracy largely intact\. The detection signal is thereforeredundantly distributed: easy to inject, hard to erase\. Across generators, the stable sets are generator\-specific yet overlap∼\\sim24×\\timesabove chance, with a layer\-12 concentration in instruction\-tuned generators not present in either base model\. Leave\-one\-family\-out evaluation confirms that the selected neurons retain most of the full\-feature detection signal on entirely unseen generator families\.

Our work is most directly related to[Borile and Abrate \(2025\)](https://arxiv.org/html/2609.30287#bib.bib3), who suppress confounding neurons in a fine\-tuned encoder to improve out\-of\-distribution detection, an observational approach that identifies which neurons hurt generalisation rather than which neurons causally drive predictions\. More broadly,[Antverg and Belinkov \(2022\)](https://arxiv.org/html/2609.30287#bib.bib1)show that individual\-neuron analysis in language models is prone to false positives without causal validation \(a pitfall observed in recent neuron\-level analysis of AI\-generated text\([Enkhbayar, 2025](https://arxiv.org/html/2609.30287#bib.bib11)\)\)\. We address the lack of a sufficiency test for AI\-detection neurons directly via activation patching, and discuss the relationship to this and other prior work in Section[2](https://arxiv.org/html/2609.30287#S2)\.

##### Contributions\.

We make four contributions:

1. 1\.A sparse\-probing characterisation of AI\-detection signal in frozen BERT\-base\-uncased across six RAID generators\. Stable sets of<<1% of CLS dimensions are sufficient for near\-ceiling detection accuracy with high cross\-fold stability\.
2. 2\.The first application of bidirectional activation patching to AI\-text detection neurons\. Forward patching establishes that the selected neurons are sufficient to drive the probe’s AI predictions \(10–16×\\timesabove random\); reverse patching establishes that they are partially necessary to preserve them \(7–12×\\timesabove random\)\.
3. 3\.A cross\-generator analysis\. Generator\-specific stable sets nonetheless overlap∼24×\{\\sim\}24\\timesabove chance, with a layer\-12 footprint concentrated in instruction\-tuned generators and absent in both pure\-base models\.
4. 4\.A leave\-one\-family\-out \(LOGO\) evaluation\. The selected neurons retain 86–94% of the full\-feature detection signal on entirely unseen generator families\.

## 2Related Work

### 2\.1AI\-Generated Text Detection

Supervised fine\-tuned classifiers\([Solaiman et al\., 2019](https://arxiv.org/html/2609.30287#bib.bib38);[Zellers et al\., 2019](https://arxiv.org/html/2609.30287#bib.bib45)\)achieve high in\-distribution accuracy but degrade substantially on unseen generators or domains\. Zero\-shot methods such as DetectGPT\([Mitchell et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib28)\)and Binoculars\([Hans et al\., 2024](https://arxiv.org/html/2609.30287#bib.bib15)\)exploit log\-probability statistics without labelled data but require whitebox or API access to a reference model\. Systematic benchmarking via RAID\([Dugan et al\., 2024](https://arxiv.org/html/2609.30287#bib.bib8)\), spanning 11 generators and 8 domains, confirms that no single method generalises reliably across generators or domains\. This motivates our focus on what internal representationsdrivedetection, rather than improving detection accuracy itself\.

### 2\.2Interpretability of Text Detectors

The simplest form of interpretability for AI\-text detectors uses input\-level attribution\.[Pudasaini et al\. \(2026\)](https://arxiv.org/html/2609.30287#bib.bib32)apply SHAP to explain classifier decisions, finding that the most predictive tokens diverge substantially across domains\. This approach characterises the model’s decision surface but cannot speak to what internal representations underlie it\.

More recent work examines representations inside the encoder\.[Borile and Abrate \(2025\)](https://arxiv.org/html/2609.30287#bib.bib3)identify neurons in a fine\-tuned BERT encoder whose activations correlate with domain or style artefacts rather than AI\-authorship signal, and suppress them to improve out\-of\-distribution accuracy\.[Kuznetsov et al\. \(2024\)](https://arxiv.org/html/2609.30287#bib.bib23)take a geometric approach, removing a learned embedding subspace to produce a detector more robust to lexical variation\. Both works operate on fine\-tuned models and ask which neurons hurt OOD generalisation, not which neurons are sufficient to drive predictions\. Our work differs from[Borile and Abrate \(2025\)](https://arxiv.org/html/2609.30287#bib.bib3)on four axes: \(i\) we use afrozenrather than fine\-tuned encoder; \(ii\) we probe CLS\-token representations rather than layer\-specific FFN sublayer activations; \(iii\) we perform both ablation to measurenecessityand activation patching to test causalsufficiency, rather than ablating to identify and suppress confounders; and \(iv\) we evaluate on RAID rather than DAIGT, M4\([Wang et al\., 2024](https://arxiv.org/html/2609.30287#bib.bib43)\), or HC3\([Guo et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib13)\)\. The terminological consequence is that[Borile and Abrate \(2025\)](https://arxiv.org/html/2609.30287#bib.bib3)identify confounding neurons in the MLP intermediate layers of a fine\-tuned BERT detector \(H=3,072H=3\{,\}072units per layer\); we instead probe the residual\-stream output at the CLS position \(768 dimensions per layer\), following the broader[Sajjad et al\. \(2022\)](https://arxiv.org/html/2609.30287#bib.bib37)convention\.

[Kuznetsov et al\. \(2025\)](https://arxiv.org/html/2609.30287#bib.bib22)apply sparse autoencoders to Gemma\-2 decoder activations, extracting semantically interpretable features correlated with AI authorship\. This work is complementary but distinct\. The model class \(autoregressive decoder\), representation type \(token\-position activations rather than CLS pooling\), and analysis method \(SAE reconstruction\) all differ from ours\. The approach also remains observational: features are identified by correlation, not by causal intervention\. None of the above provides direct evidence that a small identified set of neurons issufficientto produce AI\-text predictions when injected into human\-text activations; that is the question we address\.

### 2\.3Sparse Probing and Activation Patching

Earlier work identified task\-relevant neurons in pretrained models via correlation\-based selection, and mapped salient neurons to linguistic properties across layers and model families\([Durrani et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib9)\)\. Sparse probing uses L1\-regularised linear classifiers to identify a small set of neurons in a pretrained model that carry task\-relevant information\.[Gurnee et al\. \(2023\)](https://arxiv.org/html/2609.30287#bib.bib14)introduced and systematically evaluated the L1\-then\-L2 protocol: an L1 probe selects candidate neurons, and an L2 probe re\-fitted on those neurons provides unbiased accuracy estimates\. Their study shows that individual neurons in GPT\-2 and LLaMA encode grammatical, factual, and linguistic properties, and that the selected sets are stable across seeds and data subsets\. We adopt this protocol directly, applying it to BERT\-base\-uncased CLS\-token activations concatenated across all 12 layers\.

Activation patching replaces internal representations in a target forward pass with those from a donor pass, then measures the effect on model outputs\([Vig et al\., 2020](https://arxiv.org/html/2609.30287#bib.bib42);[Meng et al\., 2022](https://arxiv.org/html/2609.30287#bib.bib27)\)\.[Heimersheim and Nanda \(2024\)](https://arxiv.org/html/2609.30287#bib.bib17)recommend donor patching with real inputs \(rather than zero or mean vectors\) to avoid out\-of\-distribution artefacts, and mean ablation as a conservative necessity baseline; we follow both\. The closest methodological precedent for our flip\-rate framing is[Ravindran \(2025\)](https://arxiv.org/html/2609.30287#bib.bib34)\(donor patching, flip\-rate metric; autoregressive decoder, safety alignment\)\. Activation patching has also been applied to BERT cross\-encoders for IR\([Lu et al\., 2025](https://arxiv.org/html/2609.30287#bib.bib25)\), but to our knowledge no prior work applies it to encoder\-based AI\-text detection\.

## 3Experimental Setup

### 3\.1Data: RAID

RAID\([Dugan et al\., 2024](https://arxiv.org/html/2609.30287#bib.bib8)\)provides human\-written and machine\-generated texts across 11 generators and 8 domains\. We use six of the eight RAID domains: abstracts, books, news, Reddit posts, reviews, and Wikipedia\. We exclude poetry and recipes, as both have very low sample counts in RAID and risk introducing genre\-specific artefacts that confound cross\-domain analysis\. We select six generators spanning five model families: GPT\-4\([OpenAI, 2023](https://arxiv.org/html/2609.30287#bib.bib30)\), GPT\-2\([Radford et al\., 2019](https://arxiv.org/html/2609.30287#bib.bib33)\), MPT\-30B\([MosaicML NLP Team, 2023](https://arxiv.org/html/2609.30287#bib.bib29)\), Mistral\-7B\-Instruct\-v0\.1\([Jiang et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib18)\), LLaMA\-2\-Chat\([Touvron et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib41)\), and Cohere\-Command\.111Cohere Command does not have a canonical reference paper; see the model documentation at[https://cohere\.com/command](https://cohere.com/command)\.This covers two pure\-base generators \(GPT\-2, MPT\-30B\), one SFT\-only generator \(Mistral\-7B\-Instruct\-v0\.1, fine\-tuned on instruction data without RLHF\([Jiang et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib18)\)\), and three SFT\+RLHF generators \(GPT\-4, LLaMA\-2\-Chat, Cohere\-Command\)\. The six were chosen to span two axes rather than to cover RAID exhaustively: model family \(one generator per family\) and post\-training regime \(every regime present in the benchmark\)\. This design enables the base vs\. instruction\-tuned comparison in §[6](https://arxiv.org/html/2609.30287#S6)\. The remaining five RAID generators \(GPT\-3, ChatGPT, Mistral\-7B base, MPT\-Chat, Cohere\-Chat\) are same\-family duplicates; none contributes an architecture family or a post\-training regime that the six do not already cover\. All 11 generators are used in the leave\-one\-family\-out evaluation of §[7](https://arxiv.org/html/2609.30287#S7)\. We subsample 7,500 examples per generator \(3,750 human, 3,750 AI\) balanced across domains, yielding 45,000 samples in total\.

### 3\.2Model and Activations

We use BERT\-base\-uncased\([Devlin et al\., 2019](https://arxiv.org/html/2609.30287#bib.bib7)\)with all weights frozen throughout\. For each input text we extract the CLS\-token hidden state at every transformer layerl∈\{1,…,12\}l\\in\\\{1,\\ldots,12\\\}, giving a per\-layer vector𝐡\(l\)∈ℝ768\\mathbf\{h\}^\{\(l\)\}\\in\\mathbb\{R\}^\{768\}\. The full hidden\-state representation is the concatenation

𝐳=\[𝐡\(1\);…;𝐡\(12\)\]∈ℝ9,216\.\\mathbf\{z\}=\\bigl\[\\mathbf\{h\}^\{\(1\)\};\\,\\ldots;\\,\\mathbf\{h\}^\{\(12\)\}\\bigr\]\\in\\mathbb\{R\}^\{9\{,\}216\}\.We refer to individual dimensions of𝐳\\mathbf\{z\}asneurons; neuron indexiiresides in layer⌊i/768⌋\+1\\lfloor i/768\\rfloor\+1and positionimod768i\\bmod 768within that layer\.222The termneuroncarries two meanings in the literature\. The NLP probing tradition\([Dalvi et al\., 2019](https://arxiv.org/html/2609.30287#bib.bib5);[Sajjad et al\., 2022](https://arxiv.org/html/2609.30287#bib.bib37);[Durrani et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib9)\)uses it for any scalar coordinate of a hidden representation; some mechanistic\-interpretability work\([Elhage et al\., 2021](https://arxiv.org/html/2609.30287#bib.bib10);[Gurnee et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib14)\)restricts it to MLP post\-activation units, where the elementwise nonlinearity yields a privileged basis absent from the residual stream\. We follow the former; individual residual\-stream coordinates in BERT are known to carry meaningful signal on their own\([Kovaleva et al\., 2021](https://arxiv.org/html/2609.30287#bib.bib20);[Timkey and van Schijndel, 2021](https://arxiv.org/html/2609.30287#bib.bib40)\)\.Using the frozen pretrained model \(rather than a fine\-tuned detector\) ensures that any informative neurons we identify are present in general\-purpose BERT representations, not artefacts of task\-specific fine\-tuning\.

### 3\.3Protocol: L1 Selection, L2 Evaluation

Following[Gurnee et al\. \(2023\)](https://arxiv.org/html/2609.30287#bib.bib14), for each \(seed, fold\) cell we fit two independent probes on the training split\. \(1\)L1 selection probe: L1\-regularised logistic regression on all 9,216 neurons, fitted with scikit\-learn’s\([Pedregosa et al\., 2011](https://arxiv.org/html/2609.30287#bib.bib31)\)liblinearsolver at inverse regularisation strengthC=0\.005C=0\.005\(smallerCC⇒\\Rightarrowstronger L1 penalty; see Appendix[A](https://arxiv.org/html/2609.30287#A1)for the exact loss\)\. Neurons with non\-zero coefficients form the per\-cellcandidate set𝒮\\mathcal\{S\}; neurons appearing in𝒮\\mathcal\{S\}in at least 80% of the 15 cells form the aggregatedstable set𝒮∗\\mathcal\{S\}^\{\*\}\. Interventions \(patching, ablation\) target the per\-fold𝒮\\mathcal\{S\};𝒮∗\\mathcal\{S\}^\{\*\}is reserved for cross\-generator analyses \(§[6](https://arxiv.org/html/2609.30287#S6)\) and LOGO\. \(2\)L2 evaluation probe: L2\-regularised logistic regression \(lbfgs\) on the full𝐳\\mathbf\{z\}\. This probe is held frozen for all downstream experiments to ensure that intervention effects are measured against a fixed decision function\.

We use 5\-fold cross\-validation stratified by \(label×\\timesdomain\) with seeds\{42,123,456\}\\\{42,123,456\\\}, yielding 15 cells per \(generator, experiment\)\. All reported metrics are 15\-cell means±\\pmstandard deviation unless otherwise noted; per\-cell values are tabulated in Appendix[C](https://arxiv.org/html/2609.30287#A3)\.

##### Hyperparameter selection\.

C=0\.005C=0\.005andN=7,500N=7\{,\}500were chosen via a multi\-generator stability sweep run prior to any results analysis, maximising the minimum\-over\-generators pairwise Jaccard subject to≥50\\geq 50neurons per draw\([Gurnee et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib14)\)\. The chosen operating point achieves a minimum Jaccard of 0\.676 \(GPT\-2 bottleneck; other five generators 0\.707–0\.751\); full grid available in Appendix[B](https://arxiv.org/html/2609.30287#A2)\.

## 4Sparse Detection Neurons in Frozen BERT

Table[1](https://arxiv.org/html/2609.30287#S4.T1)reports sparse\-probe results across all six generators\. Between 45 and 62 neurons \(0\.49–0\.67% of the 9,216\-dimension CLS representation\) are stable across at least 12 of the 15 cells, yet the L2 evaluation probe achieves 90\.5–98\.8% validation accuracy \(AUC\-ROC 0\.968–0\.999\)\. A probe restricted to the stable set alone achieves 86\.5–97\.2% \(see Appendix[E](https://arxiv.org/html/2609.30287#A5)\); the stable neurons are therefore themselves sufficient for most of the detection accuracy\. Selection is highly stable: pairwise fold Jaccard ranges 0\.67–0\.76 across 105 fold\-pairs per generator, compared to a random\-null expectation of∼\\sim0\.003–0\.005 for sets of this size drawn from 9,216 neurons, placing within\-generator stability at 150–230×\\timesabove chance\. The selection pool is narrow: between 91 neurons ever selected \(LLaMA\) and 146 \(GPT\-2\)\. The probe consistently selects from a small region of the feature space\. Layer\-by\-layer neuron distributions and the base/instruction\-tuned asymmetry are examined in Section[6](https://arxiv.org/html/2609.30287#S6)\.

Table 1:Sparse\-probe results across six generators: 15\-cell means±\\pmstd \(5 folds×\\times3 seeds\)\. Val acc: L2 evaluation probe accuracy \(all 9,216 neurons\)\. AUC: area under the ROC curve\.nseln\_\{\\text\{sel\}\}: mean candidate set size per cell\. Jaccard: mean pairwise fold Jaccard over all 105 fold\-pairs \(C\(15,2\)\), computed from per\-fold selection sets\.nstablen\_\{\\text\{stable\}\}: neurons in≥\\geq80% of cells;nobsn\_\{\\text\{obs\}\}: total distinct neurons selected at least once\. The six generators span one model family each and all post\-training regimes present in RAID \(two pure\-base, one SFT\-only, three SFT\+RLHF\); the remaining five RAID generators are same\-family, same\-regime duplicates \(§[3\.1](https://arxiv.org/html/2609.30287#S3.SS1)\)\.
## 5Causal Role of the Selected Neurons

Having identified a sparse set of neurons, we ask two complementary causal questions: are these neuronssufficientto drive the probe’s AI predictions, and are theynecessaryto preserve them? Section[5\.1](https://arxiv.org/html/2609.30287#S5.SS1)addresses both via bidirectional activation patching: transplanting AI activations into human targets \(forward; sufficiency\) and human activations into AI targets \(reverse; partial necessity\)\. Section[5\.2](https://arxiv.org/html/2609.30287#S5.SS2)complements this with mean ablation\. We provide a restricted\-probe localisation analysis in Appendix[E](https://arxiv.org/html/2609.30287#A5); it reinforces the base/instruction\-tuned split\.

### 5\.1Causal Role via Bidirectional Activation Patching

#### 5\.1\.1Patching Protocol

We run patchingbidirectionallyto assess both causal directions: \(i\)forward\(AI donor→\\tohuman target\): transplanting the selected neurons’ activations from AI samples into human targets, testing whether the selected set issufficientto drive the probe’s AI predictions; and \(ii\)reverse\(human donor→\\toAI target\): transplanting human activations into AI targets, testing whether the selected set is partiallynecessaryfor maintaining AI predictions\.

For each \(seed, fold\) cell and direction we identifyvalid targetsamples: test\-fold samples of the target class that share a domain with at least one test\-fold donor sample\. For each ofnshuffles=20n\_\{\\text\{shuffles\}\}=20random within\-domain donor permutations we construct a patched activation by replacing the selected\-neuron activations of the target with those of the donor \(𝐳~𝒮←𝐳𝒮donor\\tilde\{\\mathbf\{z\}\}\_\{\\mathcal\{S\}\}\\\!\\leftarrow\\\!\\mathbf\{z\}^\{\\text\{donor\}\}\_\{\\mathcal\{S\}\}\), leaving non\-selected neurons unchanged\. Aflipoccurs when the frozen evaluation probe’s predicted class changes after patching \(human→\\toAI in the forward direction; AI→\\tohuman in the reverse\)\. The flip rate is the fraction of valid target samples that flip, averaged over the 20 donor permutations and then over the 15 cells\. Donor sampling is restricted to the same domain as the target, following the on\-manifold recommendation of[Heimersheim and Nanda \(2024\)](https://arxiv.org/html/2609.30287#bib.bib17)to avoid injecting out\-of\-distribution activations\.

Therandom\-kkbaselinedraws 20 random neuron subsets of sizekkand runs the identical procedure\. This serves as a null that controls for the number of patched dimensions\. We sweepk∈\{1,5,10,20,50,full\}k\\in\\\{1,5,10,20,50,\\ \\text\{full\}\\\}for both the selected\-neuron patching and the random\-k baseline\.k=fullk=\\text\{full\}refers to the per\-fold L1\-selected set𝒮\\mathcal\{S\}\(53–92 neurons depending on generator and fold; see §[3\.3](https://arxiv.org/html/2609.30287#S3.SS3)for the per\-fold𝒮\\mathcal\{S\}vs aggregated𝒮∗\\mathcal\{S\}^\{\*\}distinction\)\.

#### 5\.1\.2Flip\-Rate Results \(Forward and Reverse\)

##### Forward \(AI→\\tohuman\)\.

Flip rates atk=fullk=\\text\{full\}are 7–18×\\timeshigher than atk=5k=5across all generators \(Figure[2](https://arxiv.org/html/2609.30287#S5.F2); full sweep in Table[13](https://arxiv.org/html/2609.30287#A10.T13), Appendix[J](https://arxiv.org/html/2609.30287#A10); per\-domain breakdown in Appendix[G](https://arxiv.org/html/2609.30287#A7)\), indicating that the detection signal is distributed across the selected set rather than concentrated in a small subset\.

Table 2:Forward\-direction \(AI→\\tohuman\) selected and random\-kkflip rates \(%\) atk=fullk=\\text\{full\}\(15\-cell means±\\pmstd\)\. Ratio = selected / random \(rounded to nearest integer; precise values 9\.7–15\.7×\\times\)\. Fullkk\-sweep in Appendix[J](https://arxiv.org/html/2609.30287#A10)\. Reverse\-direction \(human→\\toAI\) results are reported in the text\.Generators nearer the decision boundary flip more: Cohere flips most readily \(8\.15%, lowest baseline at 90\.5%\), LLaMA least \(1\.07%, baseline 98\.6%\) \(Table[2](https://arxiv.org/html/2609.30287#S5.T2)\)\.

##### Reverse \(human→\\toAI\)\.

Transplanting human\-donor activations into AI\-target samples tests whether the selected neurons are partially necessary for maintaining AI\-class predictions\. Atk=fullk=\\text\{full\}, 0\.65–5\.74% of AI samples were reclassified as human \(per\-generator breakdown in Table[13](https://arxiv.org/html/2609.30287#A10.T13), Appendix[J](https://arxiv.org/html/2609.30287#A10)\), consistently 7–12×\\timesabove the random\-neuron baseline\. Reverse flip rates are lower than forward \(0\.65–5\.74% vs\. 1\.07–8\.15%\): the remaining 9,000\+\+unpatched neurons carry enough redundant AI signal to partially preserve the prediction even when the selected set is overwritten\. The same\-domain donor constraint eliminates domain shift by construction; mean\-ablation caveats are discussed in §[5\.2](https://arxiv.org/html/2609.30287#S5.SS2)\. Appendix[L](https://arxiv.org/html/2609.30287#A12)reports one patched sample in full\. Absolute flip rates are low because a flip requires overturning a probe that reads all 9,216 neurons while fewer than 1% of them are edited, so the informative quantity is the ratio to the size\-matched random\-kkbaseline \(10–16×\\timesforward, 7–12×\\timesreverse\) rather than the raw percentage\. That the selected neurons carry the detection signal, rather than merely perturbing it, is established separately by the restricted probe, which reaches 86\.5–97\.2% accuracy on those neurons alone \(Appendix[E](https://arxiv.org/html/2609.30287#A5)\)\.

##### Domain structure\.

Flip rates are not uniform across RAID domains, and the pattern differs by generator: Cohere is uniformly high across all six domains while GPT\-2 is elevated on news, Reddit, and Wikipedia relative to books and reviews\. Appendix[G](https://arxiv.org/html/2609.30287#A7)gives the per\-domain breakdown for every generator \(Table[10](https://arxiv.org/html/2609.30287#A7.T10)\) together with a dispersion analysis showing that domain sensitivity does not follow the base/instruction\-tuned split that organises the layer distributions \(§[6\.1](https://arxiv.org/html/2609.30287#S6.SS1)\)\.

Figure 2:Forward flip rate \(%\) vs\. number of patched neuronskkon a log\-scale x\-axis, averaged over 15 cells \(5 folds×\\times3 seeds\)\. Instruction\-tuned generators are shown in blue, base generators in red\. Numbers to the right of each line are selected/random ratios atk=50k=50\(8–14×\\times\); headline ratios atk=fullk=\\text\{full\}\(the per\-fold L1\-selected set, 53–92 neurons\) are 10–16×\\timesacross generators \(Table[13](https://arxiv.org/html/2609.30287#A10.T13)\)\.

### 5\.2Necessity via Mean Ablation

To complement the sufficiency test we ask the converse question: is the selected setnecessary? We answer this via mean ablation, replacing the selected neurons in the test\-fold activations with their training\-fold mean \(computed on the training split to avoid leakage\) and re\-evaluating the frozen L2 probe\. Table[3](https://arxiv.org/html/2609.30287#S5.T3)reports accuracy drop atk=fullk=\\text\{full\}alongside the random\-kkbaseline\.

For five of six generators the result is negative\. Atk=fullk=\\text\{full\}, ablating the selected set causes accuracy to drop by only 0\.03–0\.11 pp, within noise of the 0\.00–0\.06 pp random\-kkbaseline; the drops for GPT\-4, GPT\-2, MPT, and Mistral are within one standard deviation of zero \(Table[3](https://arxiv.org/html/2609.30287#S5.T3)\), so no reliable necessity is detectable for these generators\. The selected neurons are not individually necessary: the L2 evaluation probe, which uses all 9,216 neurons, retains thousands of redundant correlates that compensate for the removed subset\.

Cohere is the exception\. Its full\-set ablation drop is 1\.06±\\pm0\.52 pp \(well outside the 0\.06 pp random baseline and the per\-cell noise floor\), and the drop grows monotonically fromk=5k=5upward, unlike the other five generators where the curve is flat\. Its other signals point the same way: lowest baseline accuracy, smallest mean\-difference norm, and largest probe weight norm \(CAV diagnostics, Appendix[F](https://arxiv.org/html/2609.30287#A6); AUC\-ROC\-vs\-L1 reversal in Appendix[D](https://arxiv.org/html/2609.30287#A4)\)\. Together these point to a detection representation that ispartiallylocalised rather than fully distributed\.

We treat the necessity result as a conservative lower bound: mean ablation moves activations off\-manifold\([Heimersheim and Nanda, 2024](https://arxiv.org/html/2609.30287#bib.bib17)\), and the on\-manifold reverse\-patching test \(§[5\.1\.2](https://arxiv.org/html/2609.30287#S5.SS1.SSS2)\) confirms partial necessity\. If L1 had selected only high\-variance correlates, donor injection would not flip human predictions at10–16×10\\text\{\-\-\}16\\timesrandom\. The selected set is one sufficient subspace among others, singled out by stable recovery across folds, seeds, and \(via LOGO\) generator families\.

Table 3:Mean\-ablation accuracy drop atk=fullk=\\text\{full\}, 15\-cell means±\\pmstd \(positive = drop\)\. Sel\. drop: ablating the per\-fold L1\-selected set \(k=fullk=\\text\{full\}, 53–92 neurons\); Rand\. drop: mean over 20 randomkk\-subsets\. Cohere is the only generator substantially above the random baseline\.

## 6Cross\-Generator Representation Geometry

### 6\.1Layer Distribution of Stable Neurons

Figure[3](https://arxiv.org/html/2609.30287#S6.F3)shows the layer\-by\-layer distribution of stable neurons across all six generators \(per\-layer counts in Appendix[H](https://arxiv.org/html/2609.30287#A8)\)\. The four instruction\-tuned generators and the two pure\-base generators differ sharply\. For GPT\-4, Mistral, LLaMA, and Cohere, the single most populated layer is layer 12, which accounts for 30–36% of each generator’s stable set \(16–19 neurons\)\. Both pure\-base generators are exceptions: GPT\-2 peaks at layer 11 with 11 stable neurons \(18\.3% of its stable set\) and has only 13\.3% in layer 12; MPT peaks earlier with just 2 stable neurons in layer 12 \(3\.8%\) and distributes its signal more broadly across earlier layers \(layers 1, 8, and 9 dominate\)\.

Across both pure\-base generators, the shared pattern is markedly low layer\-12 concentration \(3\.8–13\.3%\) versus any instruction\-tuned generator \(30–36%\)\. The L12\-avoidance pattern is not a rigid L11\-peaking rule \(GPT\-2 peaks at L11 while MPT peaks at L1/L8\)\. On this evidence, BERT’s layer 12 acts as a task\-linked readout that instruction\-tuned outputs activate more strongly than base\-model outputs\([Tenney et al\., 2019](https://arxiv.org/html/2609.30287#bib.bib39);[Rogers et al\., 2020](https://arxiv.org/html/2609.30287#bib.bib36);[Geva et al\., 2021](https://arxiv.org/html/2609.30287#bib.bib12)\)\. Because this reading rests onN=2N=2base generators \(GPT\-2, MPT\) and describes the selected\-neuron geometry rather than a mechanistic claim about RLHF/SFT\-induced circuitry, we treat it as a working hypothesis \(see[Limitations](https://arxiv.org/html/2609.30287#Sx1)\)\. Cohere’s partial localisation \(§[5\.2](https://arxiv.org/html/2609.30287#S5.SS2)\) fits this reading: an IT generator encoding the L12 signal at lower per\-neuron magnitude, with the probe compensating via larger weights\.

Figure 3:Proportion of stable neurons in each layer group per generator\. Layer 12 percentage annotated in white\. Instruction\-tuned generators \(IT\) concentrate 30–36% of stable neurons in layer 12; base generators retain≤\\leq14%\. Dashed line separates the two groups\. The generators shown are one per RAID model family, selected to place both pure\-base and both instruction\-tuned post\-training regimes on this axis; see §[3\.1](https://arxiv.org/html/2609.30287#S3.SS1)for the selection and exclusion criteria\.
### 6\.2Stable\-Set Overlap Across Generators

Pairwise Jaccard similarity between stable sets \(Table[12](https://arxiv.org/html/2609.30287#A9.T12)in Appendix[I](https://arxiv.org/html/2609.30287#A9)\) ranges from 0\.00 to 0\.24\. Stable sets are therefore generator\-specific\. The appropriate baseline is not zero but the expected Jaccard for two random sets of the same sizes drawn from 9,216 neurons, which is approximately 0\.003 per pair\. Relative to this null, the mean observed Jaccard of 0\.073 is∼\\sim24×\\timesabove chance, indicating a partial shared substrate despite the generator\-specific structure\.

The data reveal a clear bipartite structure \(Figure[5](https://arxiv.org/html/2609.30287#A9.F5)in Appendix[I](https://arxiv.org/html/2609.30287#A9)\)\. Among instruction\-tuned generators, pairwise Jaccard ranges 0\.08–0\.24 \(27–86×\\timeschance\), with the strongest pairs being Cohere–Mistral \(86×\\times\) and LLaMA–Mistral \(64×\\times\)\. Among the two pure\-base generators, MPT and GPT\-2 share 12 neurons \(J = 0\.120, 40×\\timeschance\), more than several instruction–instruction pairs, so the base models form a coherent sub\-cluster\. In contrast, every base\-vs\-instruction\-tuned pair has near\-zero overlap \(J≤\\leq0\.05\), with three pairs at exactly zero \(full breakdown in Table[12](https://arxiv.org/html/2609.30287#A9.T12)\)\. This base\-vs\-instruction\-tuned partition explains the bulk of the cross\-generator geometry\.

### 6\.3A Cross\-Generator Core Concentrated in Layer 12

Neurons appearing in the stable sets of at least 3 of the 6 generators \(minimal majority threshold\) form acoreset of 17 neurons\. Of these, 10 \(59%\) reside in layer 12, far exceeding any individual generator’s L12 share \(3\.8–36%\); the remaining 7 are scattered across layers 1, 2, 5, 7, 8, and 11\. This concentration suggests BERT’s layer 12 is the primary cross\-generator site for AI\-detection signal: a majority of neurons shared across half or more of the studied generators converge there, even though no individual generator devotes a majority of its own stable set to L12\. Appendix[K](https://arxiv.org/html/2609.30287#A11)examines what these neurons respond to by correlating the highest\-weight stable neurons with four surface text statistics; the strongest associations are moderate \(\|ρ\|=0\.41\|\\rho\|=0\.41–0\.540\.54\) and vary in sign across generators, so the selected subspace does not reduce to a single lexical cue\.

## 7Leave\-One\-Family\-Out Generalisation

The experiments above characterise stable neurons for each generator*independently*\. A leave\-one\-family\-out \(LOGO\) evaluation tests whether neurons identified fromothergenerators transfer to a held\-out generator family never seen during neuron selection\.

##### Protocol\.

We partition all 11 RAID generators into five families:gpt,cohere,llama,mistral, andmpt\. For each held\-out family we run the sparse\-probe pipeline on the remaining generators and evaluate two probes on the held\-out set: \(i\)Full L2on all 9,216 neurons, and \(ii\)Sel L2on only the stable neurons identified from training families\. To handle the larger combined training pool, LOGO uses a tighter regularisation \(C=0\.001C=0\.001\) and caps each generator’s contribution at 5,000 samples \(compared toC=0\.005C=0\.005,N=7,500N=7\{,\}500in the per\-generator headline runs\)\. Family sizes are unequal \(1–4 generators per family, following RAID’s taxonomy\), so training pool size and test\-set size vary across folds\. The Sel/Full ratio in Table[4](https://arxiv.org/html/2609.30287#S7.T4)normalises for this by expressing each family’s held\-out accuracy as a fraction of its own full\-feature ceiling\.

##### Results\.

Table[4](https://arxiv.org/html/2609.30287#S7.T4)reports the outcome\. Full L2 generalises strongly in all five conditions \(88\.6–99\.7% accuracy\), showing that frozen BERT representations are themselves informative across the RAID generator space\. The Sel L2 column shows the main result: 91–125 stable neurons \(at most 1\.4% of 9,216 neurons\) achieve 76\.3–93\.7% accuracy on generators from entirely held\-out families\. Expressed as a fraction of the full\-feature ceiling for each family, the selected set retains 86–94% of the full\-feature signal \(Table[4](https://arxiv.org/html/2609.30287#S7.T4), Sel/Full column\)\.

The worst\-case family iscohere\(Sel L2 = 76\.3%, 86% of ceiling\), matching its low pairwise Jaccard with other generators; the best isllama\(93\.7%, 94%\), which benefits from high overlap within the instruction\-tuned group \(§[6\.2](https://arxiv.org/html/2609.30287#S6.SS2)\)\.mptandmistralfall in the middle tier \(85\.6–85\.7%\)\.

We interpret the 6–14 pp residual gap between Sel L2 and Full L2 as evidence that a cross\-family core of neurons carries most of the detection signal that transfers, while the remainder of the full\-feature ceiling is supplied byfamily\-specificneurons that the LOGO procedure cannot observe by construction\. This tracks the bipartite Jaccard structure \(§[6\.2](https://arxiv.org/html/2609.30287#S6.SS2)\), where each generator’s stable set is dominated by a generator\-specific periphery layered on a small shared substrate\.

##### Practical implication\.

Because the stable neurons are identified without observing the held\-out family, a detector built on frozen BERT can operate on at most 1\.4% of the representation and still generalise to generator families that were unavailable when the neuron set was chosen, without re\-running selection for each new generator\. Sel L2 is the leave\-one\-family\-out counterpart of the restricted probe of Appendix[E](https://arxiv.org/html/2609.30287#A5), which reaches 86\.5–97\.2% when trained and tested on the same generator; the 76\.3–93\.7% here is the price of never seeing the target family\. We report this as a property of the representation rather than as a detection system; the[Ethical considerations](https://arxiv.org/html/2609.30287#Sx2)section discusses the symmetric risk that the same localisation informs evasion\.

Figure 4:Leave\-one\-family\-out generalisation\. For each held\-out RAID family, the full\-feature L2 probe \(all 9,216 neurons\) and the sparse L2 probe \(only the 91–125 neurons selected from the four training families\) are evaluated on the held\-out family\. Values above each bar are accuracies; the percentage inside each sparse bar is its retention relative to that family’s own full\-feature ceiling\. The sparse probe never sees the held\-out family during selection or fitting\. Numerical values, including AUC, in Table[4](https://arxiv.org/html/2609.30287#S7.T4)\.Table 4:Leave\-one\-family\-out \(LOGO\) generalisation\.nstablen\_\{\\text\{stable\}\}: stable neurons identified from the four training families\. Full L2: L2\-probe accuracy using all 9,216 BERT neurons\. Sel L2: L2\-probe accuracy restricted tonstablen\_\{\\text\{stable\}\}neurons\. Sel/Full: ratio expressing how much of the full\-feature ceiling the selected neurons retain\. Full AUC: area under the ROC curve \(AUC\) for the full\-feature L2 probe\. Sel AUC: AUC\-ROC for the Sel L2 probe, ranging from 0\.863 \(cohere\) to 0\.986 \(llama\)\. All metrics are means over 3 seeds×\\times5 folds = 15 cells\. Test set sizes: 5,000 \(llama, 1 generator\) to 20,000 \(gpt, 4 generators\)\.

## 8Conclusion

An L1\-regularised probe stably selects 45–62 neurons \(<<1% of 9,216 CLS dimensions\) across six RAID generators\. Bidirectional activation patching shows these neurons are sufficient to drive the probe’s AI predictions \(forward, 10–16×\\timesabove random\) and partially necessary to preserve them \(reverse, 7–12×\\timesabove random\); mean ablation costs≤\\leq0\.11 pp on five of six generators \(1\.06 pp on Cohere\)\. The detection signal iseasy to inject, hard to erase: redundantly distributed yet concentrated enough to be sufficient\([Dalvi et al\., 2020](https://arxiv.org/html/2609.30287#bib.bib6)\)\.

The same data expose a bipartite structure: instruction\-tuned generators concentrate stable neurons in BERT’s final layer; base generators do not \(see[Limitations](https://arxiv.org/html/2609.30287#Sx1)\)\. Leave\-one\-family\-out evaluation confirms 86–94% of the full\-feature ceiling is retained on held\-out families, with the residual gap reflecting family\-specific neurons\.

## Limitations

##### Single encoder architecture\.

All experiments use frozen BERT\-base\-uncased\. Whether the stable\-neuron structure and patching effects extend to larger encoders \(RoBERTa\-large\([Liu et al\., 2019](https://arxiv.org/html/2609.30287#bib.bib24)\), DeBERTa\-v3\([He et al\., 2023](https://arxiv.org/html/2609.30287#bib.bib16)\)\) or encoder\-decoder architectures is an open question\. The frozen setting is a deliberate design choice: it isolates what the pretrained representation already encodes, but it means results cannot be directly compared to fine\-tuned detectors without controlling for the frozen\-vs\-fine\-tuned distinction\.

##### English\-only evaluation\.

We evaluate exclusively on the English\-language RAID benchmark\. Whether the layer\-12 footprint of post\-training alignment, the bipartite base/instruction\-tuned representation structure, and the redundancy patterns we observe generalise to languages with richer morphology, different syntactic structure \(e\.g\., agglutinative or pro\-drop languages\), or non\-Latin scripts requires separate evaluation\. Multilingual encoder variants \(e\.g\., mBERT, XLM\-R\) would be a natural starting point\.

##### Mean ablation is off\-manifold\.

Replacing selected neurons with their training\-mean pushes activations off the natural data manifold, which can introduce artefacts unrelated to the detection signal\([Heimersheim and Nanda, 2024](https://arxiv.org/html/2609.30287#bib.bib17)\)\. The reverse\-patching experiment \(§[5\.1](https://arxiv.org/html/2609.30287#S5.SS1)\) provides a complementary on\-manifold necessity test using real donor activations; the mean\-ablation result should be read alongside it rather than in isolation\. We treat the combined evidence as a conservative lower bound on necessity\.

##### Layer\-asymmetry claim across two base generators\.

The observation that instruction\-tuned generators concentrate 30–36% of stable neurons in layer 12 while pure\-base generators have≤\\leq14% is supported by two independent base models \(GPT\-2 and MPT\-30B\)\. However, the two base generators differ in their peak layer \(GPT\-2: L11; MPT\-30B: L1/L8\), so the claim is best stated as L12\-avoidance rather than a strict L11\-peaking rule\. We only claim lower L12 mass in base models; the specific lower\-layer peak is not consistent across base models\.

##### Adversarial robustness not evaluated\.

RAID includes eleven adversarial rephrasing attacks \(paraphrase, homoglyph, synonym substitution, etc\.\) that substantially reduce detector accuracy\([Dugan et al\., 2024](https://arxiv.org/html/2609.30287#bib.bib8)\)\. Whether the stable neurons identified here remain causal under these attacks is left to future work; the present analysis covers only the non\-adversarial RAID subset\.

## Ethical considerations

AI\-text detection is a dual\-use capability: detectors can support platform integrity and journalistic verification, but can also be used to penalise legitimate AI\-assisted writing or to guide adversarial evasion\. This work is mechanistic\-interpretability research: we study which neurons drive an existing frozen encoder’s representations, not deploy a new detection system\. All data come from the publicly released RAID benchmark\([Dugan et al\., 2024](https://arxiv.org/html/2609.30287#bib.bib8)\); no human subjects were involved and no personally identifiable information was collected or processed\. The stable neuron sets we identify could in principle inform evasion strategies \(by targeting the identified neuron sets\), but the same information could equally guide more robust detector design\. We release all experiment code to support reproducibility \(provided as supplementary material\)\.

## References

- Antverg and Belinkov \(2022\)Omer Antverg and Yonatan Belinkov\. 2022\.[On the pitfalls of analyzing individual neurons in language models](https://openreview.net/forum?id=8uz0EWPQIMu)\.In*International Conference on Learning Representations*\.
- Bereska and Gavves \(2024\)Leonard Bereska and Stratis Gavves\. 2024\.[Mechanistic interpretability for AI safety \- a review](https://openreview.net/forum?id=ePUVetPKu6)\.*Transactions on Machine Learning Research*\.
- Borile and Abrate \(2025\)Claudio Borile and Carlo Abrate\. 2025\.[How to generalize the detection of AI\-generated text: Confounding neurons](https://doi.org/10.18653/v1/2025.findings-emnlp.1388)\.In*Findings of the Association for Computational Linguistics: EMNLP 2025*, pages 25461–25476\.
- Choi and Zu \(2025\)Ikkyu Choi and Jiyun Zu\. 2025\.[Exploring the interpretability of AI\-generated response detection with probing](https://aclanthology.org/2025.aimecon-sessions.12/)\.In*Proceedings of the Artificial Intelligence in Measurement and Education Conference \(AIME\-Con\): Coordinated Session Papers*, pages 99–106\.
- Dalvi et al\. \(2019\)Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass\. 2019\.[What is one grain of sand in the desert? analyzing individual neurons in deep NLP models](https://doi.org/10.1609/aaai.v33i01.33016309)\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 33, pages 6309–6317\.
- Dalvi et al\. \(2020\)Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov\. 2020\.[Analyzing redundancy in pretrained transformer models](https://doi.org/10.18653/v1/2020.emnlp-main.398)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 4908–4926, Online\. Association for Computational Linguistics\.
- Devlin et al\. \(2019\)Jacob Devlin, Ming\-Wei Chang, Kenton Lee, and Kristina Toutanova\. 2019\.[BERT: Pre\-training of deep bidirectional transformers for language understanding](https://doi.org/10.18653/v1/N19-1423)\.In*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pages 4171–4186, Minneapolis, Minnesota\. Association for Computational Linguistics\.
- Dugan et al\. \(2024\)Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, and Chris Callison\-Burch\. 2024\.[RAID: A shared benchmark for robust evaluation of machine\-generated text detectors](https://doi.org/10.18653/v1/2024.acl-long.674)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 12463–12492, Bangkok, Thailand\. Association for Computational Linguistics\.
- Durrani et al\. \(2023\)Nadir Durrani, Fahim Dalvi, and Hassan Sajjad\. 2023\.[Discovering salient neurons in deep NLP models](http://jmlr.org/papers/v24/23-0074.html)\.*Journal of Machine Learning Research*, 24\(362\):1–40\.
- Elhage et al\. \(2021\)Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield\-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, and 6 others\. 2021\.[A mathematical framework for transformer circuits](https://transformer-circuits.pub/2021/framework/index.html)\.Transformer Circuits Thread\.
- Enkhbayar \(2025\)Tsogt\-Ochir Enkhbayar\. 2025\.[Atomic literary styling: Mechanistic manipulation of prose generation in neural language models](https://arxiv.org/abs/2510.17909)\.*Preprint*, arXiv:2510\.17909\.
- Geva et al\. \(2021\)Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy\. 2021\.[Transformer feed\-forward layers are key\-value memories](https://doi.org/10.18653/v1/2021.emnlp-main.446)\.In*Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing*, pages 5484–5495\.
- Guo et al\. \(2023\)Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu\. 2023\.[How close is ChatGPT to human experts? comparison corpus, evaluation, and detection](https://arxiv.org/abs/2301.07597)\.*Preprint*, arXiv:2301\.07597\.
- Gurnee et al\. \(2023\)Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas\. 2023\.[Finding neurons in a haystack: Case studies with sparse probing](https://openreview.net/forum?id=JYs1R9IMJr)\.*Transactions on Machine Learning Research*\.
- Hans et al\. \(2024\)Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein\. 2024\.[Spotting LLMs with binoculars: Zero\-shot detection of machine\-generated text](https://proceedings.mlr.press/v235/hans24a.html)\.In*Proceedings of the 41st International Conference on Machine Learning*, pages 17519–17537\.
- He et al\. \(2023\)Pengcheng He, Jianfeng Gao, and Weizhu Chen\. 2023\.[DeBERTaV3: Improving DeBERTa using ELECTRA\-style pre\-training with gradient\-disentangled embedding sharing](https://openreview.net/forum?id=sE7-XhLxHA)\.In*The Eleventh International Conference on Learning Representations*\.
- Heimersheim and Nanda \(2024\)Stefan Heimersheim and Neel Nanda\. 2024\.[How to use and interpret activation patching](https://arxiv.org/abs/2404.15255)\.*Preprint*, arXiv:2404\.15255\.
- Jiang et al\. \(2023\)Albert Q\. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie\-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed\. 2023\.[Mistral 7B](https://arxiv.org/abs/2310.06825)\.*Preprint*, arXiv:2310\.06825\.
- Kim et al\. \(2018\)Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres\. 2018\.[Interpretability beyond feature attribution: Quantitative testing with concept activation vectors \(TCAV\)](https://proceedings.mlr.press/v80/kim18d.html)\.In*Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10\-15, 2018*, volume 80 of*Proceedings of Machine Learning Research*, pages 2668–2677\. PMLR\.
- Kovaleva et al\. \(2021\)Olga Kovaleva, Saurabh Kulshreshtha, Anna Rogers, and Anna Rumshisky\. 2021\.[BERT busters: Outlier dimensions that disrupt transformers](https://doi.org/10.18653/v1/2021.findings-acl.300)\.In*Findings of the Association for Computational Linguistics: ACL\-IJCNLP 2021*, pages 3392–3405\. Association for Computational Linguistics\.
- Kushnareva et al\. \(2021\)Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, and Evgeny Burnaev\. 2021\.[Artificial text detection via examining the topology of attention maps](https://doi.org/10.18653/v1/2021.emnlp-main.50)\.In*Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing*, pages 635–649\.
- Kuznetsov et al\. \(2025\)Kristian Kuznetsov, Laida Kushnareva, Anton Razzhigaev, Polina Druzhinina, Anastasia Voznyuk, Irina Piontkovskaya, Evgeny Burnaev, and Serguei Barannikov\. 2025\.[Feature\-level insights into artificial text detection with sparse autoencoders](https://doi.org/10.18653/v1/2025.findings-acl.1321)\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 25727–25748, Vienna, Austria\. Association for Computational Linguistics\.
- Kuznetsov et al\. \(2024\)Kristian Kuznetsov, Eduard Tulchinskii, Laida Kushnareva, German Magai, Serguei Barannikov, Sergey Nikolenko, and Irina Piontkovskaya\. 2024\.[Robust AI\-generated text detection by restricted embeddings](https://doi.org/10.18653/v1/2024.findings-emnlp.992)\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 17036–17055, Miami, Florida, USA\. Association for Computational Linguistics\.
- Liu et al\. \(2019\)Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov\. 2019\.[RoBERTa: A robustly optimized BERT pretraining approach](https://arxiv.org/abs/1907.11692)\.*Preprint*, arXiv:1907\.11692\.
- Lu et al\. \(2025\)Meng Lu, Catherine Chen, and Carsten Eickhoff\. 2025\.[Pathway to relevance: How cross\-encoders implement a semantic variant of BM25](https://doi.org/10.18653/v1/2025.emnlp-main.1297)\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 25525–25547\.
- Lundberg and Lee \(2017\)Scott M\. Lundberg and Su\-In Lee\. 2017\.[A unified approach to interpreting model predictions](https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html)\.In*Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4\-9, 2017, Long Beach, CA, USA*, pages 4765–4774\.
- Meng et al\. \(2022\)Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov\. 2022\.[Locating and editing factual associations in GPT](https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems*\.
- Mitchell et al\. \(2023\)Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D\. Manning, and Chelsea Finn\. 2023\.[DetectGPT: Zero\-shot machine\-generated text detection using probability curvature](https://proceedings.mlr.press/v202/mitchell23a.html)\.In*Proceedings of the 40th International Conference on Machine Learning*, pages 24950–24962\.
- MosaicML NLP Team \(2023\)MosaicML NLP Team\. 2023\.MPT\-30B: Raising the bar for open\-source foundation models\.[https://www\.databricks\.com/blog/mpt\-30b](https://www.databricks.com/blog/mpt-30b)\.
- OpenAI \(2023\)OpenAI\. 2023\.[GPT\-4 technical report](https://arxiv.org/abs/2303.08774)\.*Preprint*, arXiv:2303\.08774\.
- Pedregosa et al\. \(2011\)Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay\. 2011\.[Scikit\-learn: Machine learning in Python](https://www.jmlr.org/papers/v12/pedregosa11a.html)\.*Journal of Machine Learning Research*, 12:2825–2830\.
- Pudasaini et al\. \(2026\)Shushanta Pudasaini, Luis Miralles\-Pechuán, David Lillis, and Marisa Llorens Salvador\. 2026\.[Why AI\-generated text detection fails: Evidence from explainable AI beyond benchmark accuracy](https://arxiv.org/abs/2603.23146)\.*Preprint*, arXiv:2603\.23146\.
- Radford et al\. \(2019\)Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever\. 2019\.[Language models are unsupervised multitask learners](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)\.OpenAI Technical Report\.
- Ravindran \(2025\)Santhosh Kumar Ravindran\. 2025\.[Adversarial activation patching: A framework for detecting and mitigating emergent deception in safety\-aligned transformers](https://arxiv.org/abs/2507.09406)\.*Preprint*, arXiv:2507\.09406\.
- Ribeiro et al\. \(2016\)Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin\. 2016\.["why should I trust you?": Explaining the predictions of any classifier](https://doi.org/10.1145/2939672.2939778)\.In*Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 13\-17, 2016*, pages 1135–1144\. ACM\.
- Rogers et al\. \(2020\)Anna Rogers, Olga Kovaleva, and Anna Rumshisky\. 2020\.[A primer in BERTology: What we know about how BERT works](https://doi.org/10.1162/tacl_a_00349)\.*Transactions of the Association for Computational Linguistics*, 8:842–866\.
- Sajjad et al\. \(2022\)Hassan Sajjad, Nadir Durrani, and Fahim Dalvi\. 2022\.[Neuron\-level interpretation of deep NLP models: A survey](https://aclanthology.org/2022.tacl-1.74/)\.*Transactions of the Association for Computational Linguistics*, 10:1285–1303\.
- Solaiman et al\. \(2019\)Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert\-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang\. 2019\.[Release strategies and the social impacts of language models](https://arxiv.org/abs/1908.09203)\.*Preprint*, arXiv:1908\.09203\.
- Tenney et al\. \(2019\)Ian Tenney, Dipanjan Das, and Ellie Pavlick\. 2019\.[BERT rediscovers the classical NLP pipeline](https://doi.org/10.18653/v1/P19-1452)\.In*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics*, pages 4593–4601, Florence, Italy\. Association for Computational Linguistics\.
- Timkey and van Schijndel \(2021\)William Timkey and Marten van Schijndel\. 2021\.[All bark and no bite: Rogue dimensions in transformer language models obscure representational quality](https://doi.org/10.18653/v1/2021.emnlp-main.372)\.In*Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing*, pages 4527–4546\. Association for Computational Linguistics\.
- Touvron et al\. \(2023\)Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, and 1 others\. 2023\.[Llama 2: Open foundation and fine\-tuned chat models](https://arxiv.org/abs/2307.09288)\.*Preprint*, arXiv:2307\.09288\.
- Vig et al\. \(2020\)Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber\. 2020\.[Investigating gender bias in language models using causal mediation analysis](https://proceedings.neurips.cc/paper/2020/hash/92650b2e92217715fe312e6fa7b90d82-Abstract.html)\.In*Advances in Neural Information Processing Systems*\.
- Wang et al\. \(2024\)Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Toru Sasaki, Thomas Arnold, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, and Preslav Nakov\. 2024\.[M4: Multi\-generator, multi\-domain, and multi\-lingual black\-box machine\-generated text detection](https://doi.org/10.18653/v1/2024.eacl-long.83)\.In*Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 1369–1407\.
- Wolf et al\. \(2020\)Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, and 3 others\. 2020\.[Transformers: State\-of\-the\-art natural language processing](https://doi.org/10.18653/v1/2020.emnlp-demos.6)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pages 38–45, Online\. Association for Computational Linguistics\.
- Zellers et al\. \(2019\)Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi\. 2019\.[Defending against neural fake news](https://proceedings.neurips.cc/paper/2019/hash/3e9f0fc9b2f89e043bc6233994dfcf76-Abstract.html)\.In*Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8\-14, 2019, Vancouver, BC, Canada*\.

## Appendix APipeline Details

Figure[1](https://arxiv.org/html/2609.30287#S1.F1)in Section[1](https://arxiv.org/html/2609.30287#S1)provides a visual overview of the pipeline described below\.

##### Data sampling\.

For each generator we drawN=7,500N=7\{,\}500balanced samples from the RAID benchmark: equal numbers of AI\-generated and human\-written texts, stratified across the six RAID domains \(abstracts, books, news, Reddit, reviews, Wikipedia\) using stratified sampling without replacement\. Texts shorter than 50 whitespace\-delimited tokens are excluded\. Human texts for each generator are the original source passages from which the AI texts were generated; this pairs each AI sample with a human sample from the same domain and source\.

##### Activation extraction\.

We encode each text with BERT\-base\-uncased using HuggingFace Transformers\([Wolf et al\., 2020](https://arxiv.org/html/2609.30287#bib.bib44)\)\. The input is truncated to 512 WordPiece tokens\. We extract the CLS\-token hidden state from*all*12 transformer layers and concatenate them into a single vector𝐳∈ℝ9,216\\mathbf\{z\}\\in\\mathbb\{R\}^\{9\{,\}216\}\(12×76812\\times 768\)\. Gradients are disabled throughout; BERT weights are never updated\.

##### Normalisation\.

Before any probe training, each of the 9,216 neurons iszz\-scored with aStandardScalerfit on the training fold and applied to the test fold\. The same scaler is reused for all downstream experiments \(patching, ablation\) within a given cell so that scale is held constant across interventions\.

##### Cross\-validation protocol\.

We use 5\-fold stratified cross\-validation repeated with 3 random seeds \(42, 123, 456\), giving 15 independent evaluation cells per generator\. The stratification is over the binary AI/human label; domain balance is not enforced at the fold level but is approximately preserved because the initial sample is domain\-stratified\.

##### L1 selection\.

The*selector*is aℓ1\\ell\_\{1\}\-regularised logistic regression \(liblinearsolver,C=0\.005C=0\.005,max\_iter=1,000=1\{,\}000,class\_weight=balanced\) fit on the training fold\. Concretely,liblinearminimises

12​C​‖𝐰‖1\+∑i=1nlog⁡\(1\+exp⁡\(−yi​𝐱i⊤​𝐰\)\)\\tfrac\{1\}\{2C\}\\\|\\mathbf\{w\}\\\|\_\{1\}\+\\sum\_\{i=1\}^\{n\}\\log\\\!\\bigl\(1\+\\exp\(\-y\_\{i\}\\,\\mathbf\{x\}\_\{i\}^\{\\top\}\\mathbf\{w\}\)\\bigr\)over training\-fold samples\(𝐱i,yi\)\(\\mathbf\{x\}\_\{i\},y\_\{i\}\)withyi∈\{−1,\+1\}y\_\{i\}\\in\\\{\-1,\+1\\\}; smallerCCinflates theℓ1\\ell\_\{1\}coefficient and therefore enforces a sparser solution\. The selected set𝒮\\mathcal\{S\}for a given cell is the set of neurons with non\-zero weight\. All 9,216 neurons are provided as input; the regulariser drives most to zero\.

##### Stable neuron aggregation\.

A neuron is*stable*if it appears in𝒮\\mathcal\{S\}in at least 80% of the 15 cells \(≥12\\geq 12out of 15\)\. This threshold was chosen to require cross\-seed and cross\-fold consistency simultaneously: a neuron must survive in the large majority of both fold and seed resamplings, not merely in a bare majority of cells\. Neurons appearing in 0–11 cells are discarded; those in 12–15 cells form the stable set𝒮∗\\mathcal\{S\}^\{\*\}\.

##### L2 evaluation probe\.

The*evaluator*is a separateℓ2\\ell\_\{2\}\-regularised logistic regression \(lbfgssolver,C=1\.0C=1\.0,max\_iter=1,000=1\{,\}000\) trained on each fold’s training split using the same scaler\. For all main experiments \(sparse\-probe accuracy, activation patching, mean ablation\), it is trained once on all 9,216 neurons per fold and held frozen for the corresponding downstream interventions; the interventions modify input activations, not the probe\. The restricted\-probe analysis \(Appendix[E](https://arxiv.org/html/2609.30287#A5)\) is the only setting in which the L2 probe is re\-trained on a reduced input dimension \(the selected columns only\)\.

##### Patching protocol\.

Patching uses same\-domain donor matching: for each human target sample, a donor AI sample from the same domain is drawn without replacement\. The selectedkkneuron activations of the target are replaced with the corresponding values from the donor\. Thekkneurons are the firstkkmembers of the per\-fold candidate set𝒮\\mathcal\{S\}, ranked by descending absolute L1 coefficient \(most important first\)\. For the full\-kksweep, all\|𝒮\|\|\\mathcal\{S\}\|per\-fold neurons are patched \(53–92 depending on generator and fold\)\.

##### Software and licensing\.

We use HuggingFace Transformers v5\.3\.0\([Wolf et al\., 2020](https://arxiv.org/html/2609.30287#bib.bib44)\)for BERT inference andscikit\-learnv1\.8\.0\([Pedregosa et al\., 2011](https://arxiv.org/html/2609.30287#bib.bib31)\)for all linear probes\. RAID is released under the MIT license; BERT\-base\-uncased is released under the Apache 2\.0 license\. We use both within their permitted scope \(non\-commercial research analysis\); no data or model weights are redistributed\.

##### Computational budget\.

BERT\-base\-uncased has 110M parameters and is used in inference\-only mode \(no fine\-tuning\) throughout\. CLS activations for all 45,000 samples are encoded once using the GPU \(∼30\{\\sim\}30minutes\) and cached; all downstream experiments \(L1/L2 probes, stability sweep, mean ablation, bidirectional patchingkk\-sweep, LOGO evaluation\) operate on the cached representations\. The complete pipeline runs in approximately 7–8 wall\-clock hours on a laptop \(AMD Ryzen 5 5600H, NVIDIA GeForce RTX 3060 Laptop GPU, 6 GB VRAM\)\. The remaining stages are all CPU\-bound: the multi\-generator stability grid takes∼1\{\\sim\}1hour, the six per\-generator main runs including thekk\-sweep withnshuffles=20n\_\{\\text\{shuffles\}\}=20donor permutations take∼5\{\\sim\}5hours total, and the LOGO evaluation across five family folds takes∼15\{\\sim\}15minutes\. L1/L2 logistic regression, mean ablation, and activation patching with cached representations do not benefit from GPU acceleration\. The pipeline is therefore I/O\- rather than compute\-bound after the initial encoding pass\.

## Appendix BStability Sweep

Section[3\.3](https://arxiv.org/html/2609.30287#S3.SS3.SSS0.Px1)justifies the choice ofC=0\.005C=0\.005andN=7,500N=7\{,\}500by reference to this appendix\. The stability sweep evaluates mean pairwise Jaccard similarity between selection sets produced by 15 independent draws of the L1 selector \(varying random seed and fold assignment\) for all six RAID generators across a grid of regularisation strengthsC∈\{0\.001,0\.002,0\.005,0\.01,0\.02,0\.05\}C\\in\\\{0\.001,0\.002,0\.005,0\.01,0\.02,0\.05\\\}and sample sizesN∈\{500,1000,2000,3500,5000,7500\}N\\in\\\{500,1000,2000,3500,5000,7500\\\}\(K=15K=15draws per cell\)\. Table[5](https://arxiv.org/html/2609.30287#A2.T5)reports the full grid\. Higher Jaccard means the selection is more stable\. Values in parentheses are the mean number of neurons selected per drawn¯\\bar\{n\}\.

Table 5:Mean Jaccard stability \(and meann¯\\bar\{n\}selected\) across 15 random draws of the L1 selector\. Rows areCCvalues; columns areNNvalues\. Larger Jaccard is better\.Bold: chosen operating point \(C=0\.005C=0\.005,N=7,500N=7\{,\}500\)\. Dagger \(†\\dagger\): degenerate cell \(n¯<1\\bar\{n\}<1, Jaccard trivially 1\.0 or unstable\)\.N=10,000N=10\{,\}000omitted \(onlyK=1K=1draw possible; pool exhausted\)\.Four observations drive the choice ofC=0\.005C=0\.005,N=7,500N=7\{,\}500\.\(1\) Sparse\-end cut\-off:C≤0\.002C\\leq 0\.002yields high or trivial Jaccard at smallNNbut selects fewer than 22 neurons atN=7,500N=7\{,\}500across all generators: too sparse to support the interventions we run\. Among operating points selecting≥50\\geq 50neurons per draw,C=0\.005C=0\.005achieves the highest minimum\-over\-6\-generators Jaccard\.\(2\) Regularisation:at any fixedN≥2,000N\\geq 2\{,\}000in theC≥0\.005C\\geq 0\.005region,C=0\.005C=0\.005dominates all weaker regularisation values \(largerCC\) on every generator\.\(3\) Sample size:forC=0\.005C=0\.005Jaccard increases monotonically inNN, with the largest gains betweenN=1,000N=1\{,\}000andN=3,500N=3\{,\}500\. TheN=5,000→7,500N=5\{,\}000\\to 7\{,\}500step adds 0\.016–0\.103 per generator;N=10,000N=10\{,\}000cannot be evaluated because the pool is exhausted \(K=1K=1\)\. The chosen operating point achieves a minimum\-over\-6\-generators Jaccard of 0\.676, with GPT\-2 as the bottleneck \(0\.676\) and the remaining five generators ranging 0\.707–0\.751\.\(4\) Degenerate region:atN≤500N\\leq 500andC=0\.005C=0\.005, every generator selects zero neurons \(Jaccard=1\.0=1\.0trivially\);N=7,500N=7\{,\}500reliably produces non\-empty sets across all six generators\.

## Appendix CPer\-Cell Metrics

Table[6](https://arxiv.org/html/2609.30287#A3.T6)reports L2\-probe accuracy, AUC\-ROC, F1, and number of L1\-selected neurons for every evaluation cell \(seed×\\timesfold\)\.

Table 6:Per\-cell L2\-probe metrics for all six generators\. Each block of 15 rows is one generator; columns are seed and fold index\. Acc: accuracy; AUC: AUC\-ROC;nseln\_\{\\text\{sel\}\}: neurons selected by L1 in this cell\.

## Appendix DAUC\-ROC vs\. L1 Selection

In this appendix, AUC\-ROC serves as a univariate neuron\-ranking criterion: each neuron is scored by its individual AUC\-ROC against the binary label on the training fold, rather than as a probe performance metric\.

We compare L1\-regularised selection \(the method used throughout this paper\) to a univariate AUC\-ROC\-based ranking: for each neuron, we compute its individual AUC\-ROC against the binary label on the training fold, together with a Mann–WhitneyUUtest with Bonferroni correction atα=0\.001\\alpha=0\.001; retain neurons that are both statistically significant and exceed a per\-neuron AUC\-ROC threshold of0\.70\.7\(or fall below0\.30\.3\)\. Selected set sizes therefore vary by generator rather than being constrained to match the L1 count\. We evaluate the informativeness of each selected set by*complement ablation*: the L2 probe is re\-evaluated after mean\-ablating all neurons*not*in the selected set \(i\.e\., only the selected neurons are retained\)\. A smaller drop from the full\-probe baseline indicates the selected set is more self\-contained\.

Table 7:L1 vs\. AUC\-ROC selection: mean across 15 cells\.nn: neurons selected\. Jaccard: overlap between the two selection sets\. Drop: accuracy drop from baseline when complement is ablated \(positive = degradation\)\. Three regimes are visible: parsimony \(GPT\-4, MPT, Mistral, LLaMA: L1 selects far fewer neurons with equivalent or smaller drop\), convergence \(GPT\-2, similarnnand similar drop\), and reversal \(Cohere: AUC\-ROC selects fewer with smaller drop\)\.The Jaccard overlap between L1 and AUC\-ROC\-selected sets is low \(0\.05–0\.19; Table[7](https://arxiv.org/html/2609.30287#A4.T7)\) despite both being applied to the same data\. The two criteria thus select largely non\-overlapping neurons\. For GPT\-4, MPT, Mistral, and LLaMA, L1 achieves equivalent or better task performance with far fewer neurons \(4×4\\timesto13×13\\timesfewer\): ablating L1’s 58–74 neurons costs only 0\.03–0\.11pp, while ablating AUC\-ROC’s 238–784 neurons costs 0\.52–6\.90pp\. This parsimony makes L1 preferable for mechanistic analysis: a smaller identified set is easier to interpret and less likely to include redundant or near\-duplicated neurons\.

For GPT\-2, the two methods select nearly the same number of neurons \(85 vs\. 86\) with similar ablation cost; neither dominates\. For Cohere, AUC\-ROC selects*fewer*neurons \(38 vs\. 72\) and its ablation cost is lower \(\+0\.28pp vs\. \+1\.06pp\): Cohere’s detection signal is concentrated in a small number of individually highly\-discriminative neurons that univariate ranking recovers more directly than joint L1 regularisation\. This reversal fits Cohere’s distinct geometry \(§[5\.2](https://arxiv.org/html/2609.30287#S5.SS2), Appendix[F](https://arxiv.org/html/2609.30287#A6)\): the model with the lowest baseline probe accuracy is also the one where AUC\-ROC beats L1 on parsimony\.

## Appendix ESignal Localisation: Restricted Probe

The necessity result \(§[5\.2](https://arxiv.org/html/2609.30287#S5.SS2)\) uses the full\-feature L2 probe, which retains thousands of redundant correlates\. A complementary question is whether the stable neurons alone are sufficient for afreshly trainedprobe, and how this compares to ablating the complement\. We measure thelocalisation gap: the accuracy difference between a probe retrained on only the\|𝒮∗\|\|\\mathcal\{S\}^\{\*\}\|stable neurons \(smaller model\) and the original L2 probe evaluated with non\-stable neurons mean\-ablated \(ablate complement\)\. A large positive gap means the stable set can support an independent probe on its own, but the full\-feature probe leaned on non\-stable neurons to define its decision boundary\.

Table 8:Restricted probe results\.Smaller model: L2 probe retrained on only the\|𝒮∗\|\|\\mathcal\{S\}^\{\*\}\|stable neurons\.Gap: smaller\-model accuracy minus ablate\-complement accuracy \(positive = stable neurons self\-sufficient; larger = more dependence on non\-stable neurons in the full\-feature probe\)\. 15\-cell means±\\pmstd\.The pattern \(Table[8](https://arxiv.org/html/2609.30287#A5.T8)\) splits cleanly along the base/instruction\-tuned axis\. Instruction\-tuned generators \(GPT\-4, Mistral, LLaMA\) show gaps of≤\\leq1\.2 pp: their stable neurons are nearly self\-sufficient and the full\-feature probe does not rely heavily on non\-stable dimensions\. The two pure\-base generators and Cohere show gaps of 3\.6–4\.7 pp: their full\-feature probes lean more on non\-stable neurons, making complement ablation more costly\. This points to a more distributed encoding structure for these three generators, in line with the layer\-12 avoidance seen in base generators and the large ablation sensitivity of Cohere documented in the main body\.

## Appendix FDetection Signal Geometry \(CAV Diagnostics\)

To characterise the geometry of the detection signal within each stable set we compute a concept activation vector \(CAV\)\([Kim et al\., 2018](https://arxiv.org/html/2609.30287#bib.bib19)\)as the mean difference between AI and human activations restricted to the stable neurons, and compare it to the L1 probe weight vector\. Table[9](https://arxiv.org/html/2609.30287#A6.T9)reports cosine similarity between the CAV and the probe weight, logistic\-regression accuracy on the CAV projection, and the norms of both vectors\.

Cohere has the lowest CAV cosine similarity \(0\.057±0\.0010\.057\\pm 0\.001\), the lowest LR accuracy on the CAV projection \(0\.910±0\.0080\.910\\pm 0\.008\), and by far the largest probe weight norm \(108\.3±2\.1108\.3\\pm 2\.1\); the detection signal is diffuse and the probe compensates for it with large weights\. MPT\-30B is a notable exception: it has a low cosine \(0\.065±0\.0000\.065\\pm 0\.000, close to Cohere’s\) but thelargestmean\-difference norm \(‖Δ​μ‖=9\.52\\\|\\Delta\\mu\\\|=9\.52\)\. The AI−\-human class separation is therefore geometrically strong per\-axis yet misaligned with the probe direction\. GPT\-2 also shows a low cosine \(0\.073±0\.0020\.073\\pm 0\.002\) and elevated weight norm \(69\.9±2\.969\.9\\pm 2\.9\), placing both base generators at the low\-cosine end\.

Instruction\-tuned generators cluster at higher cosine \(0\.1100\.110–0\.1380\.138\) with the smallest weight norms\. GPT\-4 has the most compact weight vector \(45\.8±1\.745\.8\\pm 1\.7\); its detection signal is both geometrically aligned and concentrated\. Each metric is mean±\\pmstd across three LR train/test split seeds; the mean\-difference norm is deterministic \(no split randomness\)\. These per\-generator geometric differences track the intervention asymmetries documented in §[5](https://arxiv.org/html/2609.30287#S5)\.

Table 9:CAV diagnostics on per\-generator stable sets\. All values mean±\\pmstd across 3 random seeds\. Cosine: similarity between mean\-difference CAV and L1 probe weight\. LR acc: logistic\-regression accuracy on the CAV projection\.‖Δ​μ‖\\\|\\Delta\\mu\\\|: AI−\-human mean\-difference norm \(deterministic; bold = largest\)\.‖w‖\\\|w\\\|: L1 weight norm \(bold = smallest\)\.
## Appendix GPer\-Domain Patching Breakdown

Table[10](https://arxiv.org/html/2609.30287#A7.T10)reports selected flip rates \(%\) atk=fullk=\\text\{full\}broken down by RAID domain, averaged over all 15 cells \(5 folds×\\times3 seeds\)\. Values are 15\-cell means; per\-cell standard deviations ranged 0\.6–2\.6 pp across all generator–domain combinations\.

Table 10:Per\-domain selected flip rate \(%\) atk=fullk=\\text\{full\}, 15\-cell means\. Domain means may differ slightly from the overall mean due to variation in the number of valid\-human pairs per domain\. GPT\-2 shows elevated flip rates on news, Reddit, and wiki \(3\.1–4\.9%\) relative to books and reviews \(1\.9–2\.0%\)\. Cohere is uniformly high across all six domains \(5\.5–10\.1%\); its probe operates near the decision boundary throughout\. Mistral shows a notable wiki spike \(3\.54%\) while reviews are weak \(0\.62%\)\. CV: coefficient of variation over each generator’s six domain flip rates\. See §[5\.1\.2](https://arxiv.org/html/2609.30287#S5.SS1.SSS2)for interpretation\.##### Domain sensitivity across generators\.

The bottom row of Table[10](https://arxiv.org/html/2609.30287#A7.T10)summarises each generator’s spread across domains as a coefficient of variation\. Mistral \(0\.55\), MPT \(0\.42\), and GPT\-2 \(0\.33\) are domain\-sensitive; GPT\-4 \(0\.18\), LLaMA \(0\.22\), and Cohere \(0\.22\) are close to domain\-invariant\. Cohere’s flip rates are high everywhere, which is consistent with a probe operating near its decision boundary regardless of content, whereas GPT\-2 separates news, Reddit, Wikipedia, and abstracts \(3\.1–4\.9%\) from books and reviews \(1\.9–2\.0%\)\. Both pure\-base generators fall in the sensitive group, but Mistral does as well, so domain sensitivity does not follow the base/instruction\-tuned split that organises the layer distributions \(§[6\.1](https://arxiv.org/html/2609.30287#S6.SS1)\)\. The two axes are related but distinct, and the present data do not identify what drives the difference\.

## Appendix HStable\-Neuron Layer Distributions \(Full Table\)

Table[11](https://arxiv.org/html/2609.30287#A8.T11)reports the per\-layer distribution of stable neurons for all six generators, summarised in §[6\.1](https://arxiv.org/html/2609.30287#S6.SS1)\.

Table 11:Stable\-neuron counts by layer group across six generators\. Bold indicates the peak layer group for each generator\. Instruction\-tuned generators \(GPT\-4, Mistral, LLaMA, Cohere\) concentrate 30–36% of stable neurons in layer 12; both pure\-base generators \(GPT\-2, MPT\) have≤\\leq14% in layer 12\.
## Appendix IPairwise Jaccard Similarity \(Full Table\)

Table[12](https://arxiv.org/html/2609.30287#A9.T12)reports pairwise Jaccard similarity between stable sets for all 15 generator pairs, discussed in §[6\.2](https://arxiv.org/html/2609.30287#S6.SS2)\. Figure[5](https://arxiv.org/html/2609.30287#A9.F5)visualises the same data as a heatmap with generators reordered to expose the base/instruction\-tuned partition\.

![Refer to caption](https://arxiv.org/html/2609.30287v1/fig_jaccard.png)Figure 5:Pairwise Jaccard similarity heatmap \(6×\\times6\)\. Generators reordered: instruction\-tuned \(IT\) first, base last\. Dashed blue box: IT×\\timesIT block\. Dashed red box: base×\\timesbase block\. The bipartite structure \(higher within\-group similarity, near\-zero cross\-group overlap\) is immediately visible\.Table 12:Pairwise Jaccard similarity between stable sets \(15 generator pairs\)\.×\\timeschance is observed Jaccard divided by the analytic random\-null Jaccard for two sets of the same sizes drawn from 9,216 neurons \(≈\\approx0\.003 per pair\)\. The mpt–gpt2 pair \(bold\) is the highest base×\\timesbase overlap; all base×\\timesinstruction\-tuned pairs cluster at the bottom\.
## Appendix JFullkk\-Sweep Flip Rates

Table[13](https://arxiv.org/html/2609.30287#A10.T13)reports selected flip rates \(%\) across allkkvalues for each generator, referenced in §[5\.1\.2](https://arxiv.org/html/2609.30287#S5.SS1.SSS2)\.

Table 13:Flip rate \(%\) for both patching directions, 15\-cell means \(5 folds×\\times3 seeds\)\.k=fullk=\\text\{full\}equals each cell’s actual\|𝒮\|\|\\mathcal\{S\}\|\(53–92 neurons\); std shown only for selectedk=fullk=\\text\{full\}\.Forward direction\(AI donor→\\tohuman target\): selected flip rate is monotonically increasing inkkfor every generator\. Selected/random ratios atk=fullk=\\text\{full\}are 10–16×\\times\(precise values 9\.7–15\.7×\\times; see Table[2](https://arxiv.org/html/2609.30287#S5.T2)\)\.Reverse direction\(human donor→\\toAI target\) is likewise monotonically increasing inkkfor every generator; selected/random ratios atk=fullk=\\text\{full\}are 7–12×\\times, and absolute rates are below the forward direction at everykk\. See §[5\.1\.2](https://arxiv.org/html/2609.30287#S5.SS1.SSS2)for the partial\-redundancy reading\.
## Appendix KWhat the Selected Neurons Track

Section[6](https://arxiv.org/html/2609.30287#S6)characteriseswherethe stable neurons sit; this appendix examines what they respond to\. We correlate neuron activations with four surface statistics of the input text: type–token ratio, mean word length, mean sentence length, and punctuation density\. These statistics are simple and model\-independent; the question they address is how much of a stable neuron’s behaviour a single surface cue accounts for\.

##### Protocol\.

For each generator we rank its stable set𝒮∗\\mathcal\{S\}^\{\*\}by mean\|w\|\|w\|in the L1 selection probe, averaged over the 15 cells, and keep the ten highest\-weight neurons\. Each neuron’s activation is Spearman\-correlated with each of the four features over allN=9,996N=9\{,\}996cached samples, giving 40 tests per generator, to which we apply a Benjamini–Hochberg correction atq=0\.05q=0\.05within the generator\. Feature values are computed on the token sequence the encoder actually saw, reconstructed by decoding the cached inputs, so features and activations are index\-aligned by construction\.

##### Results\.

Table[14](https://arxiv.org/html/2609.30287#A11.T14)reports the strongest surviving association per generator\. The associations are consistent but partial: 31–37 of the 40 tests survive correction for every generator, while the strongest effect per generator is\|ρ\|=0\.41\|\\rho\|=0\.41–0\.540\.54, or roughly 16–29% of variance explained\. Most of the variance in these neurons is therefore unaccounted for by all four features combined, and the selected set does not reduce to a single surface cue\.

Mean word length is the top\-associated feature for five of six generators \(type–token ratio for MPT\), but thesignof that association is inconsistent: positive for GPT\-4, GPT\-2, and LLaMA, negative for Mistral and Cohere\. A neuron whose activation increases with longer words under one generator decreases with them under another\. This is consistent with the largely disjoint stable sets of §[6\.2](https://arxiv.org/html/2609.30287#S6.SS2), and it argues against reading these neurons as a generator\-independent detector of lexical complexity\.

The neurons carrying the strongest surface association are also not always the layer\-12 neurons that dominate the instruction\-tuned distributions\. They sit in layer 12 for GPT\-4 and LLaMA, but in layers 1 and 5 for GPT\-2, MPT, Mistral, and Cohere\. Two neurons recur across generators \(L12/308 for GPT\-4 and LLaMA; L5/531 for Mistral and Cohere\), and both belong to the 17\-neuron cross\-generator core of §[6\.3](https://arxiv.org/html/2609.30287#S6.SS3)\.

Surface lexical statistics are thus measurably present in the selected subspace and are the strongest single\-feature account we obtained, but they remain an incomplete one\. Identifying the remaining structure would require learned feature dictionaries rather than hand\-chosen statistics, which we leave to future work \(see[Limitations](https://arxiv.org/html/2609.30287#Sx1)\)\.

Table 14:Strongest surviving neuron–feature association per generator, over the ten highest\-weight stable neurons×\\timesfour surface features \(40 Spearman tests per generator, Benjamini–Hochberg atq=0\.05q=0\.05; 31–37 tests survive per generator\)\.ρ2\\rho^\{2\}is the approximate share of variance explained\. “word length” is mean word length, “type–token” the type–token ratio\. Neuron is given as layer/unit\. All associations are moderate and none approaches full localisation of the detection signal to a surface cue\.

## Appendix LA Worked Patching Example

The flip rates of §[5\.1\.2](https://arxiv.org/html/2609.30287#S5.SS1.SSS2)are aggregates over 15 cells and 20 donor permutations\. This appendix reports a single intervention in full\.

We take the Cohere generator, seed 42, fold 0, and patch the full stable set𝒮∗\\mathcal\{S\}^\{\*\}\(62 neurons, of which 19, or 31%, lie in layer 12, matching the 30–36% band reported for instruction\-tuned generators in Figure[3](https://arxiv.org/html/2609.30287#S6.F3)\)\. The target is a human\-written news article; the donor is a Cohere\-generated news article from the same test fold, so the same\-domain constraint of §[5\.1\.1](https://arxiv.org/html/2609.30287#S5.SS1.SSS1)is respected\. Both are roughly 1,030 characters long\.

Before patching, the frozen L2 probe classifies the target as human withP⁡\(AI\)=0\.007P\(\\text\{AI\}\)=0\.007, a confident prediction\. Replacing the 62 selected activations with the donor’s values, and leaving the other 9,154 neurons unchanged, moves the probe toP⁡\(AI\)=0\.654P\(\\text\{AI\}\)=0\.654, so the prediction flips to AI\. Editing 0\.67% of the representation is sufficient to reverse a decision the probe held with high confidence\.

The same donor draw flips 47 of 750 valid targets \(6\.3%\), close to Cohere’s reported forward flip rate of 8\.15% atk=fullk=\\text\{full\}\(Table[2](https://arxiv.org/html/2609.30287#S5.T2)\), so the example is representative rather than extreme\. Most targets donotflip, which is the redundancy reported in §[5\.2](https://arxiv.org/html/2609.30287#S5.SS2), observed at the level of a single draw\.

Similar Articles

Paraphrasing Attack Resilience of Various AI-Generated Text Detection Methods

arXiv cs.LG

This paper investigates the resilience of AI-generated text detection methods (fine-tuned RoBERTa, Binoculars, text feature analysis, and ensembles) against paraphrasing attacks, finding that Binoculars-inclusive ensembles are most effective but also most vulnerable to attacks, highlighting a dichotomy between performance and resilience.

MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text

arXiv cs.CL

This paper introduces MELD, a detector for AI-generated text that uses multi-task learning with auxiliary heads for generator family, attack type, and source domain to improve robustness. MELD achieves strong performance on the RAID benchmark and maintains low false-positive rates under adversarial attacks.