面向政策的AI使用检测:学术出版的证据框架
摘要
论文指出,对于学术出版机构而言,标准的AI文本检测工具所测量的并非真正需要关注的对象,并提出了一种面向政策的AI使用检测方法——该证据框架将期刊政策作为显式输入,报告经校准的假设及不确定性(而非二元判定),并基于可复现的合规/非合规工作流管线构建基准测试。
arXiv:2609.38427v1 Announce Type: new
Abstract: Major venues now publish detailed rules about how authors, reviewers, and area chairs may use AI, and those rules differ by role, by task, and by what must be disclosed. AI detection, the instrument usually proposed to enforce them, estimates something else: whether an AI model wrote the text. We argue that this target is misaligned with the decisions conferences and journals face, and propose policy-conditioned AI-use detection, an evidentiary framework for assessing whether a human--AI workflow complied with a stated rule. Policy makes the governing rule an explicit input. Inference reports hypotheses, evidence, calibration regime, and uncertainty in place of verdicts such as "AI detected". Evaluation builds benchmarks from reproducible pipelines that generate compliant and non-compliant workflows, and reports true positive rate at a false positive rate the venue fixes in advance. We work the framework through peer review, where at plausible violation rates a detector at a strong operating point still flags more compliant authors than violating ones. The framework therefore also names what a venue must instrument: structured disclosure, approved-tool routing that respects reviewer confidentiality, and a path by which a finding can be contested. Under this framing a detector is not an authorship classifier but an auditable procedure with an error rate the venue fixes in advance and can defend.
查看缓存全文
缓存时间: 2026/10/01 09:44
# Policy-Conditioned AI-Use Detection:An Evidentiary Framework for Academic Publishing
Source: [https://arxiv.org/html/2609.38427](https://arxiv.org/html/2609.38427)
\\workshoptitle
AI\-Native Academia: Authorship, Peer Review, and Conference Governance under AI
Jairo Diaz\-RodriguezAffiliation:Department of Mathematics and StatisticsAffiliation:York UniversityAffiliation:Toronto, Ontario M3J 1P3Email:[jdiazrod@yorku\.ca](mailto:)Mumin JiaAffiliation:Department of Mathematics and StatisticsAffiliation:York UniversityAffiliation:Toronto, Ontario M3J 1P3Email:[amyjia@yorku\.ca](mailto:)
###### Abstract
Major venues now publish detailed rules about how authors, reviewers, and area chairs may use AI, and those rules differ by role, by task, and by what must be disclosed\. AI detection, the instrument usually proposed to enforce them, estimates something else: whether an AI model wrote the text\. We argue that this target is misaligned with the decisions conferences and journals face, and propose*policy\-conditioned AI\-use detection*, an evidentiary framework for assessing whether a human–AI workflow complied with a stated rule\.Policymakes the governing rule an explicit input\.Inferencereports hypotheses, evidence, calibration regime, and uncertainty in place of verdicts such as “AI detected”\.Evaluationbuilds benchmarks from reproducible pipelines that generate compliant and non\-compliant workflows, and reports true positive rate at a false positive rate the venue fixes in advance\. We work the framework through peer review, where at plausible violation rates a detector at a strong operating point still flags more compliant authors than violating ones\. The framework therefore also names what a venue must instrument: structured disclosure, approved\-tool routing that respects reviewer confidentiality, and a path by which a finding can be contested\. Under this framing a detector is not an authorship classifier but an auditable procedure with an error rate the venue fixes in advance and can defend\.
## 1Introduction
Academic venues have written detailed rules about how AI may be used, and have almost no way to tell whether those rules were followed\. NeurIPS 2026 exempts spell checking, grammar suggestions, editing aid, and basic code assistance from any documentation requirement, and asks that agents or LLMs be described in the experimental setup when they form an important, original, or non\-standard component of the approach\. Its reviewers may use no LLM beyond a venue\-sanctioned one, and only on papers where such use is specifically permitted\([NeurIPS, 2026](https://arxiv.org/html/2609.38427#bib.bib36)\)\. ICML 2026 allows authors to use generative AI to assist in writing or research, and splits reviewer policy into two regimes: one prohibits LLM use outright, the other permits privacy\-compliant help with comprehension and polish while forbidding any delegation of judgment\([ICML, 2026a](https://arxiv.org/html/2609.38427#bib.bib37);[ICML, 2026b](https://arxiv.org/html/2609.38427#bib.bib38)\)\. ICLR 2026 sets a lower threshold still and requires that any use of an LLM be declared, in the paper and on the submission form, with authors and reviewers remaining responsible for whatever a model contributes\([ICLR, 2026](https://arxiv.org/html/2609.38427#bib.bib39);[ICLR, 2025](https://arxiv.org/html/2609.38427#bib.bib40)\)\. The three venues do not agree even on when disclosure is owed: routine language editing needs no mention at NeurIPS and must be declared at ICLR\. These are not variations on a single prohibition\. They are role\-conditioned rules about delegation, disclosure, confidentiality, and retained responsibility\.
Figure 1:Standard AI detection research versus policy\-conditioned AI\-use detection: what the detection target should be, what the output should represent, and how the system should be evaluated\.What venues have instead is AI detection research organized around origin classification: given an artifact, estimate whether a language model produced it\([Gehrmann et al\., 2019](https://arxiv.org/html/2609.38427#bib.bib8);[Mitchell et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib10);[Guo et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib14);[Verma et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib15);[Wu et al\., 2025](https://arxiv.org/html/2609.38427#bib.bib35)\)\. That target made sense when model\-written text was rare and separable from ordinary scholarly writing\. It fits badly a pipeline in which an author may legitimately polish prose with a model, a reviewer may fix grammar but not form an opinion, and an area chair may cluster reviews but not set a decision\. The mismatch is not a matter of accuracy\. A reviewer who drafts a full assessment and asks a model to correct the English satisfies most venue rules; a reviewer who pastes a confidential manuscript into a consumer chatbot breaks them even if every word of the review is their own\. An origin detector rates the first as the more suspicious of the two, inverting the rule it is meant to enforce\. Large\-scale estimates already show model\-modified signal in a substantial share of review text at AI venues\([Liang et al\., 2024a](https://arxiv.org/html/2609.38427#bib.bib43)\)and score effects associated with AI\-assisted reviews\([Latona et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib44)\), and neither finding tells a chair which of those reviews broke a rule\. Under most current policies, some of them broke none\.
The gap matters because the decisions on the other side of a detector score are consequential and contested\. A flag can mean a desk rejection, an integrity investigation, a retracted review, or a note in an author’s record\. The accused has nothing to argue with: a number states no proposition, names no hypothesis, and identifies no population on which its threshold was calibrated\. Scale compounds the problem: violations are rare and submissions many, so a strong operating point still accuses more compliant authors than violating ones, and no ranking metric reveals it\. Detectors also fail in ways their users cannot anticipate, degrading under paraphrase and domain shift and penalizing non\-native English writers\([Sadasivan et al\., 2025](https://arxiv.org/html/2609.38427#bib.bib18);[Weber\-Wulff et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib23);[Liang et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib24)\); OpenAI withdrew its own text classifier for low accuracy\([OpenAI, 2023](https://arxiv.org/html/2609.38427#bib.bib22)\)\. Meanwhile the artifacts themselves are being written to defeat inspection, most visibly through hidden prompts embedded in manuscripts to steer AI\-assisted reviewers\([Nikkei Asia, 2025](https://arxiv.org/html/2609.38427#bib.bib49)\)\.
AI\-detection research for academic publishing should change its target\. The question is not whether an artifact contains AI, but whether the evidence available about the workflow that produced it is more consistent with a use the governing policy permits or with one it prohibits, for the specific role and task at hand \(Figure[1](https://arxiv.org/html/2609.38427#S1.F1)\)\.Origin detection does not disappear under this view\. It becomes the special case in which the policy happens to be “no AI\-generated content\.”
This shift motivates the three\-part framework developed in the paper: policy, inference, and evaluation\. First,policychanges the target of detection from AI use in general to policy\-relevant AI use\. The operative distinction is not human versus AI but permitted assistance versus prohibited delegation, and what counts as permitted depends on the domain, role, task, and governing rule\. Second,inferencechanges the statistical form of detection\. Outputs such as “AI detected” collapse uncertain evidence into an overconfident verdict, where a detector should instead report the hypotheses, evidence, uncertainty, calibration assumptions, and validity conditions behind its claim\. Third,evaluationchanges how detectors are assessed\. Existing benchmarks score artifact labels rather than policy\-relative workflow compliance, so the field needs pipeline benchmarks that start from human seeds and generate artifacts through controlled permitted and prohibited workflows, reporting power against those workflows at an error rate the venue has fixed in advance \(Figure[2](https://arxiv.org/html/2609.38427#S2.F2)\)\.
The rest of the paper is organized as follows\. Section[2](https://arxiv.org/html/2609.38427#S2)introduces the framework\. Sections[3](https://arxiv.org/html/2609.38427#S3),[4](https://arxiv.org/html/2609.38427#S4), and[5](https://arxiv.org/html/2609.38427#S5)develop each component in turn\. Section[6](https://arxiv.org/html/2609.38427#S6)sets out what a venue would have to build for the framework to run\. Section[7](https://arxiv.org/html/2609.38427#S7)situates the framework within related work, and Section[8](https://arxiv.org/html/2609.38427#S8)considers objections and limitations\.
## 2A Framework for Policy\-Conditioned AI\-Use Detection
The framework has three components\.*Policy*fixes which forms of AI use are permitted for a given domain, role, and task\.*Inference*weighs the evidence about an observed artifact against those categories\.*Evaluation*tests whether the resulting procedure controls the errors that matter\. None works alone: a detector without policy does not know what it is estimating, scores without inference cannot support accountable decisions, and benchmark numbers without appropriate evaluation do not transfer to deployment\. The aim is not a universal classifier for “AI” versus “human,” but an evidentiary system for reasoning about whether a workflow complied with a stated rule under stated uncertainty\.
##### Policy\.
Detection should begin with an explicit rule: a detector should receive not only an artifact but the domain, role, task, and policy that define permitted and prohibited use\. Take a venue that allows reviewers to edit language but forbids delegating judgment\. A permitted workflow has the reviewer form an assessment, write it, and ask a model to fix grammar; a prohibited one has the reviewer upload the manuscript and ask a model for the review and the recommendation\. Appropriate use is not a property of the final review but a relation among a workflow, a role, and a policy\.
##### Inference\.
Given a policy, the detector should produce an inferential statement rather than a label\. The observed artifact is the submitted review; further evidence may include reviewer notes, portal edit history, tool logs, or a disclosure\. The relevant output is evidence about
H0\\displaystyle H\_\{0\}:the review was produced through a workflow permitted by the venue’s policy,\\displaystyle:\\text\{the review was produced through a workflow permitted by the venue's policy\},H1\\displaystyle H\_\{1\}:the review was produced through a workflow prohibited by that policy\.\\displaystyle:\\text\{the review was produced through a workflow prohibited by that policy\}\.A frequentist report states the null, the test statistic, the calibration regime, and the p\-value\. A Bayesian report states the prior, the evidence used, and the posterior probability of a violation\. In both cases uncertainty is explicit, and the detector’s role is to supply evidence for human adjudication rather than to convert a score into an accusation\.
##### Evaluation\.
Benchmarks should evaluate the full procedure, not artifact classification alone\. Evaluation should start from a shared pool of manuscripts, generate several review workflows from each, and test the detector under different evidence conditions, such as artifact\-only access versus artifact plus notes and logs\. The statistical target must be stated: a benchmark may require Type I error control atα=0\.05\\alpha=0\.05under permitted workflows while reporting TPR at that FPR against prohibited delegation\. Reporting AUROC alone names neither the tolerated false\-accusation rate nor the detection rate at that tolerance\.
Figure 2:Policy\-conditioned AI\-use detection connects three components, shown here for peer review at a venue that permits AI grammar correction but prohibits AI drafting of the review:policydefines the permitted and prohibited reviewer workflows,inferencequantifies evidence for a policy violation, andevaluationmeasures TPR at controlled FPR levels\.Together these define a research agenda: policy languages that specify permitted and prohibited AI use per role, detectors that combine artifact signals with provenance and workflow evidence, and benchmarks built around reproducible human–AI pipelines rather than frozen datasets\. Appendix[A\.1](https://arxiv.org/html/2609.38427#A1.SS1)instantiates the framework end to end for peer review, and Appendix[B](https://arxiv.org/html/2609.38427#A2)sketches the corresponding machine\-readable policy; the following sections develop each component\.
## 3Policy: AI Use Is Not the Right Target
### 3\.1Venue policy is already role\- and task\-specific
The policies quoted in Section[1](https://arxiv.org/html/2609.38427#S1)share a structure, and it is not an administrative detail: it changes the quantity a detection system should be estimating\. Authors may use AI and stay responsible for the result, reviewers face tighter constraints because review involves confidential material and delegable judgment, and disclosure attaches to significant or non\-standard use rather than to any use at all\. Venues have also begun deploying review\-side AI themselves\([Zou et al\., 2025](https://arxiv.org/html/2609.38427#bib.bib41)\)and running experiments with AI authors and reviewers under explicit disclosure\([Agents4Science, 2025](https://arxiv.org/html/2609.38427#bib.bib48)\), which makes the line between sanctioned and unsanctioned automation a matter of provenance rather than style\.
The structure is not a quirk of one community: education, platform, and regulatory regimes condition obligations on task, data, tool, and form of generation rather than on AI use as such\([International Baccalaureate, 2023](https://arxiv.org/html/2609.38427#bib.bib1);[Harvard University Information Technology, 2026](https://arxiv.org/html/2609.38427#bib.bib2);[European Union, 2024](https://arxiv.org/html/2609.38427#bib.bib4)\)\. Across all of them the operative distinction is allowed versus disallowed use, not use versus non\-use\.
### 3\.2Appropriate use is a relation, not a property
Because rules differ across domains and roles, identical AI participation can be required, permitted, irrelevant, or prohibited depending on context\. Grammar editing is unremarkable in a manuscript and in a review, while undisclosed full drafting violates authorship norms in the first and delegates judgment in the second\. In software engineering, agentic systems may generate entire pull requests when they operate through approved tools, tests, code review, and branch protections\([GitHub, 2025](https://arxiv.org/html/2609.38427#bib.bib5);[OpenAI, 2026](https://arxiv.org/html/2609.38427#bib.bib6);[Google Cloud, 2025](https://arxiv.org/html/2609.38427#bib.bib7)\), and the same generation volume would be disqualifying in a written examination\. Appropriate use is therefore not an intrinsic property of a text, figure, or code file\. It is a relation between the artifact, the workflow that produced it, and the policy governing that workflow, and a detector that ignores two of the three terms is estimating something no institution asked about\.
### 3\.3Formalization
Letxxdenote the observed artifact,dda domain,tta task,ρ\\rhothe actor’s role,πd,t,ρ\\pi\_\{d,t,\\rho\}the applicable policy, andeethe available evidence beyond the artifact\. Lowercase symbols denote observed inputs to the detector\. A detector should not estimate only whetherxxis AI\-generated; it should report evidence about whether the workflow that producedxxcomplies withπd,t,ρ\\pi\_\{d,t,\\rho\}:
D\(x,d,t,ρ,πd,t,ρ,e\)→r,D\(x,d,t,\\rho,\\pi\_\{d,t,\\rho\},e\)\\rightarrow r,whererris a structured report rather than a binary verdict\. A report may contain a compliance assessment, a posterior probability of violation, uncertainty, the evidence types used and missing, the calibration regime, validity conditions, and a recommended action including abstention\.
This makes three changes\. The policy becomes an input to the detector instead of an afterthought\. The estimand shifts from an artifact label to a workflow class\. And the output is an evidentiary report, not an accusation\.
How much inference this requires depends on how directly the workflow is observed\. Under artifact\-only access the production workflow is largely latent and compliance has to be inferred from indirect signals\. Provenance metadata, disclosures, edit histories, and approved\-tool logs each narrow the set of workflows consistent with what the venue holds\. At the limit, a sufficiently complete trace establishes a policy\-relevant action outright, and the task becomes verification instead of statistical reconstruction\. We use*policy\-conditioned AI\-use detection*as an umbrella term for the whole of this evidentiary procedure\. A statistical detector is one evidence source within it, and in some evidence regimes not the dominant one\.
### 3\.4Policies must become machine\-readable
Conditioning on policy requires policy that a system can read\. Venue rules are currently prose scattered across handbooks, calls for papers, and blog posts, which makes them impossible to attach to a submission, version, or audit\. A minimal specification would name, for each role: the tasks in scope; the workflows classified as permitted and prohibited; the disclosure required for each; the evidence the venue will collect and retain; the error rate the venue is willing to tolerate against compliant actors; and the graduated consequences of a sustained finding\. Appendix[B](https://arxiv.org/html/2609.38427#A2)sketches such a specification for reviewer conduct\. The benefit does not depend on any detector working well: a venue that writes the specification must state its tolerated false\-accusation rate before it sees any cases, which is the only point at which that number can be chosen honestly\.
## 4Inference: Detection as a Statistical Procedure
### 4\.1From detector scores to statistical claims
The phrase “AI detected” is rhetorically strong and statistically empty\. It suggests that a hidden property of the artifact was observed\. In practice a detector emits a score, and a threshold turns that score into a label\. The threshold is chosen by a vendor, an institution, or an individual user, usually with no stated relationship to the policy being enforced or to the error rate that matters in the setting\.
Detectors are compared using AUROC, accuracy, and occasionally TPR at one operating point\. None of these tells a program chair what score should trigger an integrity inquiry\.[Tufts et al\. \(2025\)](https://arxiv.org/html/2609.38427#bib.bib9)identify the gap and argue for operating\-point reporting\. In a hypothesis\-testing framing the natural summary is TPR at a fixed FPR, where FPR is the Type I error rate under permitted workflows, because it names both the tolerated false\-accusation rate and the detection rate obtained at that tolerance\.
Deployment forces the decision either way\. A venue must decide which score triggers review, which triggers nothing, and which is too uncertain to act on; absent a statistical framework those thresholds are set by intuition, and scores end up treated as evidence of misconduct with no error control\. The defect is in the form of the output rather than in the model behind it\. “The detector returned 0\.87” asserts no proposition, so an author has nothing to rebut and a chair has nothing to explain\. A more accurate detector reporting the same number leaves the venue where it started\.
### 4\.2Hypothesis testing
One formalization is hypothesis testing\. Even conventional origin detection is better stated as a test than as a verdict:
H0:the artifactxis human\-authoredagainstH1:the artifactxis AI\-generated\.H\_\{0\}:\\text\{the artifact \}x\\text\{ is human\-authored\}\\quad\\text\{against\}\\quad H\_\{1\}:\\text\{the artifact \}x\\text\{ is AI\-generated\}\.We argue that this should be generalized to policy\-conditioned hypotheses:
H0\(d,t,ρ\):the artifactxwas produced through a workflow permitted byπd,t,ρ,H1\(d,t,ρ\):the artifactxwas produced through a workflow prohibited byπd,t,ρ\.\\begin\{array\}\[\]\{ll\}H\_\{0\}^\{\(d,t,\\rho\)\}:&\\text\{the artifact \}x\\text\{ was produced through a workflow permitted by \}\\pi\_\{d,t,\\rho\},\\\\ H\_\{1\}^\{\(d,t,\\rho\)\}:&\\text\{the artifact \}x\\text\{ was produced through a workflow prohibited by \}\\pi\_\{d,t,\\rho\}\.\\end\{array\}A thresholded score is not enough to support this\. A valid test specifies the null, the alternatives, the statistic, the calibration distribution, the population over which calibration is expected to hold, and the level\. Here Type I error is the probability of rejecting a permitted workflow, so controlling the level bounds the probability that a compliant author or reviewer is flagged, and the required stringency depends on the consequence\. If a flag only prioritizes low\-stakes human reading, a weakly calibrated screen may be acceptable; if it can trigger a desk rejection, an integrity case, or removal from a reviewer pool, the venue should be able to state the level and the assumptions under which it holds\. Failing to reject is not proof of compliance, and rejecting is not proof of misconduct; both are evidence under a stated model, calibration regime, and level\.
### 4\.3Bayesian inference
A Bayesian framing is often more natural for deployment\. LetVVdenote the event that the production workflow violates the applicable policy\. A detector may report
P\(V∣x,d,t,ρ,πd,t,ρ,e\)\.P\(V\\mid x,d,t,\\rho,\\pi\_\{d,t,\\rho\},e\)\.The posterior depends on likelihoods and on priors, and the same score should imply different concern in different settings\. Evidence of substantial model drafting is decisive in a closed\-book assessment, weak evidence of anything in a marketing document, and expected in an approved agentic coding workflow\. A watermark is uninformative if the policy allowed disclosed AI drafting and highly informative if it prohibited model\-written prose\.
Bayesian reporting also makes base rates and costs explicit\. A detector score cannot be interpreted without an assumption about how common the relevant violation is\. Institutions differ in their loss functions—a university may weight false accusations most heavily, a journal the fabricated citations that survive into the literature—and no such weighting is recoverable from a generic origin score\. Abstention deserves the same first\-class treatment: where violations are rare, a procedure that returns “insufficient evidence” on short artifacts with no workflow trace beats one that guesses, because the guesses are mostly wrong and each costs a compliant person something\.
##### What a report should contain\.
A frequentist report states the null, the statistic, the calibration regime, and the p\-value before recommending review or abstention\. A Bayesian report states the prior assumptions, the evidence used and missing, and the posterior probability of violation\. Instead of “AI detected,” a usable report reads:*estimated probability of prohibited drafting is0\.380\.38; calibration is weak for reviews under 300 words*, or*p=0\.049p=0\.049in a level\-0\.050\.05test calibrated on long\-form reviews from this venue’s 2025 cycle*\. These pair a number with its scope of validity, which makes them auditable and much harder to misread as proof\. The goal is not to remove decision rules but to make them explicit, policy\-specific, and contestable\.
## 5Evaluation: Benchmarks Should Be Recipes, Not Datasets
### 5\.1Current benchmarks and their limits
Machine\-generated text detection has been evaluated largely on static corpora that separate human writing from model output or attribute text to a source generator\([Uchendu et al\., 2021](https://arxiv.org/html/2609.38427#bib.bib29);[Guo et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib14);[He et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib30);[Wang et al\., 2024b](https://arxiv.org/html/2609.38427#bib.bib32);[Li et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib33)\)\. A second line starts from human text and applies model transformations, better reflecting workflows in which people edit, paraphrase, summarize, expand, or mix human and model prose, and studies adversarial rewriting, mixed authorship, boundary detection, and graded editing\([Su et al\., 2023b](https://arxiv.org/html/2609.38427#bib.bib31);[Hu et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib19);[Koike et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib20);[Dugan et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib21);[Wang et al\., 2024a](https://arxiv.org/html/2609.38427#bib.bib34);[Thai et al\., 2026](https://arxiv.org/html/2609.38427#bib.bib16)\)\.
These are real advances, and they remain insufficient for governance\. Benchmark tasks are stylized proxies: they ask whether a detector separates human from model text under controlled conditions, and rarely specify the institutional setting, the role, the evidence available at decision time, or the consequence of an error\. Static corpora also age quickly as models, prompts, and human editing habits change\. Transformation\-based benchmarks add valuable stress tests, but their labels remain variants of human, machine, source model, attack type, span boundary, or edit magnitude\. None of these is the label a chair needs\.
### 5\.2Pipeline benchmarks
We propose policy\-conditioned pipeline benchmarks\. The benchmark object should not be a fixed collection of artifacts but a reproducible recipe for producing artifacts under specified human–AI workflows\. A recipe begins from a human seed—a draft, outline, set of notes, manuscript, bug report, or codebase—applies a controlled AI\-use pipeline, and assigns a label relative to a stated policy\.
Appendix[A\.1](https://arxiv.org/html/2609.38427#A1.SS1)specifies such a recipe set for peer review, with the evidence each workflow leaves behind\. Two properties of that set carry the argument\. A review that was language\-edited, one expanded from the reviewer’s own notes, and one written wholly by a model can read alike while the policy separates them, so an artifact\-only detector is asked to recover a distinction the artifact does not encode\. And one workflow is left undefined by the policy rather than permitted or prohibited by it, so a benchmark that forces such cases to one side reports an accuracy its labels do not support\.
A recipe is more durable than a fixed corpus because it can be regenerated: when a new model or paraphraser appears the recipes rerun, and when a venue changes its policy the same pool is relabeled\. A complete specification names the seed distribution, the domain, role, and task, the governing policy, the permitted and prohibited pipelines, the tools and prompts, the human writing and verification steps, and the evidence exposed to the detector\. That last variable should be varied deliberately across four regimes:E0E\_\{0\}, the artifact alone;E1E\_\{1\}, the artifact with a structured disclosure;E2E\_\{2\}, the artifact with provenance and approved\-tool records; andE3E\_\{3\}, a substantially observed workflow\. Uncertainty about the workflow falls across the four, and the inferential burden falls with it\. A benchmark that reports performance by regime therefore answers a sharper question than whether workflow\-aware detection works: it measures how much each evidence channel buys, and locates the point where inference gives way to verification\.
### 5\.3Evaluation criteria
Policy\-conditioned benchmarks require more than accuracy or AUROC\. AUROC measures ranking across thresholds and does not define a usable procedure\. In the review setting a detector may separate fully model\-written reviews from hand\-written ones on average while still flagging permitted language editing at a rate the venue would never accept\.
Evaluation should therefore separate level from power\. Permitted workflows define the null and prohibited workflows the alternatives; the level bounds false accusations against compliant reviewers, and power measures detection of specified violations\. A detector should not be credited for catching model\-written reviews if it also flags permitted editing, and one with valid level but negligible power is safe and useless\. Benchmarks should report TPR at fixed FPR per workflow, and should state the base rate they assume, because an operating point on its own does not fix what a flag means\.
Suppose a venue receives20,00020\{,\}000submissions and deploys a detector at TPR=0\.80=0\.80and FPR=0\.01=0\.01, a strong operating point on current benchmarks\. A flag is an allegation, so the two rates land on two different groups of authors\. Take a violation rate of5%5\\%\. The detector flags800800of the1,0001\{,\}000authors who broke the rule and misses200200; separately, it flags190190of the19,00019\{,\}000who complied, so four in five flags are justified\. Now take a violation rate of1%1\\%\. The same detector catches160160violators and accuses198198compliant authors, and most of the people it names did nothing wrong\. Loosen the threshold to FPR=0\.05=0\.05and990990compliant submissions are flagged against160160violations\. AUROC is unchanged across the three settings; what changes is the share of accused authors who are innocent, from roughly one in five to roughly six in seven\. Reviews inherit the problem, with the added constraint that a flag must be resolved within the review cycle, since discarding a flagged review leaves the paper short of its required count\.
Further criteria follow from deployment: calibration of reported probabilities, abstention rate and quality, robustness to paraphrase and deliberate evasion, generalization across models and venues, usefulness to the human adjudicator, and the privacy cost of the evidence required\. The last is a constraint rather than a footnote\. Workflow evidence is informative precisely because it exposes intellectual process, and a venue that logs reviewer keystrokes to catch delegation has traded one governance failure for another\.
Robustness deserves its own arm rather than a line in a list\. Any enforced rule invites evasion, and publishing has an attack surface of its own: hidden prompts embedded in manuscripts to steer AI\-assisted reviewers\([Nikkei Asia, 2025](https://arxiv.org/html/2609.38427#bib.bib49)\), now prohibited at NeurIPS\([NeurIPS, 2026](https://arxiv.org/html/2609.38427#bib.bib36)\), and paraphrase attacks on detectors\([Sadasivan et al\., 2025](https://arxiv.org/html/2609.38427#bib.bib18);[Hu et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib19);[Koike et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib20)\)\. A benchmark should include recipes in which the actor knows a detector is running and takes low\-cost steps to defeat it\. Reporting only non\-adversarial TPR overstates what deployment will achieve, and the overstatement is largest for the actors a venue most wants to catch\.
## 6Governance: What Venues Would Have to Build
Policy\-conditioned detection is only as good as the evidence a venue retains, and that evidence has to come from somewhere\. Three things would have to be built: a disclosure format precise enough to condition on, a route by which reviewers reach approved tools without breaching confidentiality, and a process that turns a report into a decision the accused can contest\. Building them does not require progress in detection\. A venue that has them can enforce a rule with weak inference, while a venue that deploys a strong detector into an uninstrumented pipeline has a score and no procedure\.
##### Disclosure as evidence rather than formality\.
Venue requirements today ask for free text, and only once use crosses a threshold the discloser judges for themselves\([NeurIPS, 2026](https://arxiv.org/html/2609.38427#bib.bib36);[ICLR, 2026](https://arxiv.org/html/2609.38427#bib.bib39)\)\. Categorical schemes exist elsewhere: Amazon KDP separates AI\-generated from AI\-assisted content and attaches disclosure only to the former\([Amazon Kindle Direct Publishing, 2026](https://arxiv.org/html/2609.38427#bib.bib3)\), and the EU AI Act ties obligations to specified content types\([European Union, 2024](https://arxiv.org/html/2609.38427#bib.bib4)\)\. A declaration naming the task, the tool, the stage, and the person who verified the result changes the question from reconstructing a workflow to checking whether the artifact is consistent with the one the actor described\. Watermarks, provenance metadata, and portal edit history enter the same report as further channels, each carrying its own reliability\([Kirchenbauer et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib25);[Dathathri et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib26);[Coalition for Content Provenance and Authenticity, 2026](https://arxiv.org/html/2609.38427#bib.bib28)\)\. None of this survives if a venue leaves open how declarations will be used, since anyone who suspects that declaring invites scrutiny will declare less\.
##### Confidentiality as a constraint on evidence\.
Peer review is confidential, which constrains both the violations that matter and the evidence a venue may gather\. Several of the prohibited reviewer workflows in Appendix[A\.1](https://arxiv.org/html/2609.38427#A1.SS1)are violations because a manuscript left the venue’s trust boundary, not because the resulting review reads oddly, and the traces that would settle the question are the ones most invasive to collect\. Approved\-tool routing answers that trade, and it is already in force: NeurIPS 2026 permits reviewers only a venue\-sanctioned LLM, and only on papers where LLM use is specifically allowed\([NeurIPS, 2026](https://arxiv.org/html/2609.38427#bib.bib36)\)\. Routing turns an unobservable act into a logged one without surveilling the reviewer’s own machine, and it narrows the inference problem, since unexplained model\-like text then indicates unsanctioned tooling rather than AI use as such\. Strong provenance does not make the framework unnecessary; it changes the kind of evidence the framework runs on\. Where a trusted log records a policy\-relevant action directly, the question is verification rather than statistical reconstruction\. Other parts of the same workflow stay unobserved, including any use of tools outside the venue’s infrastructure, and those still require inference\. What no venue yet does is treat these logs as evidence inside a stated procedure\. A sanctioned\-tool record supports a finding only once the venue has also fixed the error rate it will tolerate against compliant reviewers and the route by which a finding can be contested\.
##### Due process and a sanction ladder\.
A report should enter a process rather than settle a decision\. Detector errors concentrate on non\-native English writers\([Liang et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib24)\), and detectors fail in ways their users cannot anticipate\([Weber\-Wulff et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib23);[OpenAI, 2023](https://arxiv.org/html/2609.38427#bib.bib22)\), so a regime that acts on a score alone imposes those errors on identifiable groups\. A workable design routes reports through graduated responses: no action, a request for clarification, a structured disclosure request, human integrity review, and sanction, with the evidentiary standard rising at each step and the report never sufficient on its own at the last\. Peer review already runs staged and contestable decisions\([Shah, 2022](https://arxiv.org/html/2609.38427#bib.bib46)\), and AI\-use findings should travel through that machinery rather than around it\. The report must also be disclosable to the accused, which requires it to state the hypotheses and calibration set out in Section[4](https://arxiv.org/html/2609.38427#S4)\. An inferential report can be answered\. A score cannot\.
## 7Related Work
Origin detectors\.One line of work scores text using statistical, likelihood, rank, perturbation, regeneration, or cross\-model signals\([Gehrmann et al\., 2019](https://arxiv.org/html/2609.38427#bib.bib8);[Mitchell et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib10);[Bao et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib11);[Su et al\., 2023a](https://arxiv.org/html/2609.38427#bib.bib12);[Yang et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib13);[Hans et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib42)\); another trains classifiers on labeled human and machine examples\([Guo et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib14);[Verma et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib15)\); robustness work studies paraphrasing, adversarial rewriting, evasion, and benchmark attacks\([Sadasivan et al\., 2025](https://arxiv.org/html/2609.38427#bib.bib18);[Hu et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib19);[Koike et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib20);[Dugan et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib21)\)\. These produce evidence about origin or model\-likeness; the target is detection, and the output is a score\.
Beyond binary labels\.Editing\-sensitive and mixed\-authorship methods treat AI involvement as graded or compositional, and watermarking and provenance systems supply evidence about generation history\([Thai et al\., 2026](https://arxiv.org/html/2609.38427#bib.bib16);[Mao et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib17);[Wang et al\., 2024a](https://arxiv.org/html/2609.38427#bib.bib34);[Kirchenbauer et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib25);[Dathathri et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib26);[Liu et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib27)\)\. These are closer to real workflows, and they still require policy interpretation: evidence that a model contributed does not settle whether the contribution was allowed, disclosed, verified, or compatible with the actor’s role\.
AI in peer review\.A growing empirical literature measures model involvement in review rather than in papers\.[Liang et al\. \(2024a\)](https://arxiv.org/html/2609.38427#bib.bib43)estimate the prevalence of model\-modified review text at AI venues,[Latona et al\. \(2024\)](https://arxiv.org/html/2609.38427#bib.bib44)report score and acceptance effects associated with AI\-assisted reviews,[Liang et al\. \(2024b\)](https://arxiv.org/html/2609.38427#bib.bib45)evaluate model\-generated feedback on manuscripts, and[Shah \(2022\)](https://arxiv.org/html/2609.38427#bib.bib46)surveys the computational problems in peer review into which this one falls\. Venues themselves now deploy review\-side AI\([Zou et al\., 2025](https://arxiv.org/html/2609.38427#bib.bib41)\)and run experiments with AI authors and reviewers under explicit disclosure\([Agents4Science, 2025](https://arxiv.org/html/2609.38427#bib.bib48)\)\. This work establishes prevalence and effect; it does not supply a compliance estimand, which is the gap we address\. Alongside it, reliability studies report detectors degrading under paraphrase and domain shift, OpenAI’s retirement of its own classifier, and bias against non\-native English writers\([Sadasivan et al\., 2025](https://arxiv.org/html/2609.38427#bib.bib18);[OpenAI, 2023](https://arxiv.org/html/2609.38427#bib.bib22);[Weber\-Wulff et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib23);[Liang et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib24)\)\. Those findings motivate calibrated evidence with stated validity conditions, abstention, and human review in place of verdicts\.
## 8Objections and Limitations
Origin detection remains valuable\.A natural objection is that many settings still require generic origin detection, and this is correct\. Spam and bot detection, platform integrity, synthetic media labeling, and some publishing workflows turn on whether content was machine\-generated\. Our claim is the narrower one that origin detection constitutes the special case in which the governing policy is “no AI\-generated content,” rather than the general problem\. As AI assistance becomes routine in writing, programming, and research, that special case will describe a diminishing share of the decisions institutions face\.
The proposal is hard to realize\.Policy\-conditioned inference is more demanding than origin classification, and we do not dispute the point\. Policies are ambiguous, workflows are only partly observable, actors conceal tool use, and the available evidence is noisy, privacy\-sensitive, and open to manipulation\. We read these difficulties as an account of why the problem remains unsolved rather than as a defense of the present target\. If compliance depends on institutional rules and partly observed workflows, an “AI detected” output is aimed at the wrong quantity, and raising its accuracy does not bring it closer to the right one\.
Data and labels\.The framework requires examples of workflows, not only artifacts: seeds, prompts, drafts, model outputs, revision histories, tool logs, disclosures, and final submissions, labeled by policy status\. Collecting them is expensive and raises governance concerns of its own, since workflow evidence exposes intellectual process and confidential material, so evidence acceptable for research and evidence acceptable to collect in deployment are not the same set\. Labeling is also partly interpretive: venues will classify the same workflow differently, real workflows fall between editing and drafting, and labels may need to be graded rather than binary\.
Out of scope\.Decisions taken under this regime feed back into the corpus, since reviews and accepted papers become training data and text shaped by detection pressure will shape later models\([Shumailov et al\., 2024](https://arxiv.org/html/2609.38427#bib.bib47)\)\. We do not model that loop\. Our worked example is also text\-centric\. Code, figures, data, and multi\-agent research pipelines raise evidence questions we do not address, and the arithmetic assumes one detector applied uniformly, whereas venues will triage\.
## Conclusion
AI\-use detection should not be framed as the search for a boundary between human and machine authorship\. Venues have written rules that turn on role, task, delegation, and disclosure, and what their enforcement machinery must establish is whether a given workflow complied\. Policy\-conditioned detection makes that rule an explicit input and treats artifact signals, disclosures, provenance, and workflow traces as evidence of different strengths\. Where the workflow stays latent it asks for calibrated inference; where trusted provenance settles the relevant action it reduces to verification\. Either way the output has to support a contestable institutional decision at a false positive rate fixed in advance\.
The cost of this reframing is that it moves work off the detector and onto the venue\. Someone has to write the policy in a form a system can condition on, decide which traces are worth retaining and which are too invasive to collect, fix the tolerated rate of false accusation before any case arrives, and build a path by which a finding can be contested\. None of that is a modeling problem, and a better classifier does not supply any of it\. But a venue that does it can say what its procedure controls and what it does not, which is the minimum any institution should be able to say before acting on a suspicion\. An institution that acts on a detector output owes the accused an account of what was tested and how often that test is wrong\. Producing that account is the work this framework describes\.
## References
- Agents4ScienceOpen conference of AI agents for science\.Note:[https://agents4science\.stanford\.edu/](https://agents4science.stanford.edu/)Accessed: 2026\-04\-27Cited by:[§3\.1](https://arxiv.org/html/2609.38427#S3.SS1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Amazon Kindle Direct Publishing \(2026\)Amazon Kindle Direct PublishingContent guidelines: artificial intelligence content\.Note:[https://kdp\.amazon\.com/help/topic/G200672390](https://kdp.amazon.com/help/topic/G200672390)Accessed: 2026\-04\-27Cited by:[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px1.p1.1)\.
- Baoet al\.\(2024\)G\. Bao, Y\. Zhao, Z\. Teng, L\. Yang, and Y\. ZhangFast\-detectGPT: efficient zero\-shot detection of machine\-generated text via conditional probability curvature\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Bpcgcr8E8Z)Cited by:[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- Coalition for Content Provenance and Authenticity \(2026\)Coalition for Content Provenance and AuthenticityC2PA specifications\.Note:[https://spec\.c2pa\.org/](https://spec.c2pa.org/)Accessed: 2026\-04\-27Cited by:[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px1.p1.1)\.
- Dathathriet al\.\(2024\)S\. Dathathri, A\. See, S\. Ghaisas, P\. Huang, R\. McAdam, J\. Welbl, V\. Bachani, A\. Kaskasoli, R\. Stanforth, T\. Matejovicova, J\. Hayes, N\. Vyas, H\. Huang, A\. Glaese, N\. McAleese, M\. Hutter, B\. Balle, P\. Kohli, S\. Gowal, and Z\. AhmedScalable watermarking for identifying large language model outputs\.Nature\.External Links:[Link](https://www.nature.com/articles/s41586-024-08025-4)Cited by:[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p2.1)\.
- Duganet al\.\(2024\)L\. Dugan, A\. Hwang, F\. Trhlik, A\. Zhu, J\. M\. Ludan, H\. Xu, D\. Ippolito, and C\. Callison\-BurchRAID: a shared benchmark for robust evaluation of machine\-generated text detectors\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 12463–12492\.External Links:[Link](https://aclanthology.org/2024.acl-long.674/)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- European Union \(2024\)European UnionRegulation \(EU\) 2024/1689, artificial intelligence act, article 50\.Note:[https://artificialintelligenceact\.eu/article/50/](https://artificialintelligenceact.eu/article/50/)Accessed: 2026\-04\-27Cited by:[§3\.1](https://arxiv.org/html/2609.38427#S3.SS1.p2.1),[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px1.p1.1)\.
- Gehrmannet al\.\(2019\)S\. Gehrmann, H\. Strobelt, and A\. M\. RushGLTR: statistical detection and visualization of generated text\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations,pp\. 111–116\.External Links:[Link](https://aclanthology.org/P19-3019/)Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p2.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- GitHub \(2025\)GitHubGitHub copilot: meet the new coding agent\.Note:[https://github\.blog/news\-insights/product\-news/github\-copilot\-meet\-the\-new\-coding\-agent/](https://github.blog/news-insights/product-news/github-copilot-meet-the-new-coding-agent/)Accessed: 2026\-04\-27Cited by:[§3\.2](https://arxiv.org/html/2609.38427#S3.SS2.p1.1)\.
- Google Cloud \(2025\)Google CloudGemini code assist in github for enterprises\.Note:[https://cloud\.google\.com/blog/products/ai\-machine\-learning/gemini\-code\-assist\-in\-github\-for\-enterprises/](https://cloud.google.com/blog/products/ai-machine-learning/gemini-code-assist-in-github-for-enterprises/)Accessed: 2026\-04\-27Cited by:[§3\.2](https://arxiv.org/html/2609.38427#S3.SS2.p1.1)\.
- Guoet al\.\(2023\)B\. Guo, X\. Zhang, Z\. Wang, M\. Jiang, J\. Nie, Y\. Ding, J\. Yue, and Y\. WuHow close is ChatGPT to human experts? comparison corpus, evaluation, and detection\.External Links:2301\.07597,[Link](https://arxiv.org/abs/2301.07597)Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p2.1),[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- Hanset al\.\(2024\)A\. Hans, A\. Schwarzschild, V\. Cherepanova, H\. Kazemi, A\. Saha, M\. Goldblum, J\. Geiping, and T\. GoldsteinSpotting LLMs with binoculars: zero\-shot detection of machine\-generated text\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.235,pp\. 17519–17537\.External Links:[Link](https://proceedings.mlr.press/v235/hans24a.html)Cited by:[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- Harvard University Information Technology \(2026\)Harvard University Information TechnologyGenerative AI guidelines\.Note:[https://www\.huit\.harvard\.edu/ai/guidelines](https://www.huit.harvard.edu/ai/guidelines)Accessed: 2026\-04\-27Cited by:[§3\.1](https://arxiv.org/html/2609.38427#S3.SS1.p2.1)\.
- Heet al\.\(2023\)X\. He, X\. Shen, Z\. Chen, M\. Backes, and Y\. ZhangMGTBench: benchmarking machine\-generated text detection\.External Links:2303\.14822,[Link](https://arxiv.org/abs/2303.14822)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1)\.
- Huet al\.\(2023\)X\. Hu, P\. Chen, and T\. HoRADAR: robust AI\-text detection via adversarial learning\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=QGrkbaan79)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1),[§5\.3](https://arxiv.org/html/2609.38427#S5.SS3.p5.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- ICLR \(2025\)ICLRPolicies on Large Language Model Usage at ICLR 2026\.Note:[https://iclr\.cc/FAQ/LLM](https://iclr.cc/FAQ/LLM)Accessed: 2026\-04\-27Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p1.1)\.
- ICLR \(2026\)ICLRICLR 2026 Author Guide\.Note:[https://iclr\.cc/Conferences/2026/AuthorGuide](https://iclr.cc/Conferences/2026/AuthorGuide)Accessed: 2026\-04\-27Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p1.1),[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px1.p1.1)\.
- ICML \(2026a\)ICMLICML 2026 Call for Papers\.Note:[https://icml\.cc/Conferences/2026/CallForPapers](https://icml.cc/Conferences/2026/CallForPapers)Accessed: 2026\-04\-27Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p1.1)\.
- ICML \(2026b\)ICMLICML 2026 Policy for LLM Use in Reviewing\.Note:[https://icml\.cc/Conferences/2026/LLM\-Policy](https://icml.cc/Conferences/2026/LLM-Policy)Accessed: 2026\-04\-27Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p1.1)\.
- International Baccalaureate \(2023\)International BaccalaureateStatement from the IB about ChatGPT and artificial intelligence in assessment and education\.Note:[https://ibo\.org/news/news\-about\-the\-ib/statement\-from\-the\-ib\-about\-chatgpt\-and\-artificial\-intelligence\-in\-assessment\-and\-education/](https://ibo.org/news/news-about-the-ib/statement-from-the-ib-about-chatgpt-and-artificial-intelligence-in-assessment-and-education/)Accessed: 2026\-04\-27Cited by:[§3\.1](https://arxiv.org/html/2609.38427#S3.SS1.p2.1)\.
- Kirchenbaueret al\.\(2023\)J\. Kirchenbauer, J\. Geiping, Y\. Wen, J\. Katz, I\. Miers, and T\. GoldsteinA watermark for large language models\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=aX8ig9X2a7)Cited by:[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p2.1)\.
- Koikeet al\.\(2024\)R\. Koike, M\. Kaneko, and N\. OkazakiOUTFOX: llm\-generated essay detection through in\-context learning with adversarially generated examples\.InProceedings of the 38th AAAI Conference on Artificial Intelligence,Vancouver, Canada\.Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1),[§5\.3](https://arxiv.org/html/2609.38427#S5.SS3.p5.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- Latonaet al\.\(2024\)G\. R\. Latona, M\. H\. Ribeiro, T\. R\. Davidson, V\. Veselovsky, and R\. WestThe AI review lottery: widespread AI\-assisted peer reviews boost paper scores and acceptance rates\.Note:arXiv:2405\.02150External Links:[Link](https://arxiv.org/abs/2405.02150)Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p2.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Liet al\.\(2024\)Y\. Li, Q\. Li, L\. Cui, W\. Bi, L\. Wang, L\. Yang, S\. Shi, and Y\. ZhangMAGE: machine\-generated text detection in the wild\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 36–53\.External Links:[Link](https://aclanthology.org/2024.acl-long.3/)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1)\.
- Lianget al\.\(2024a\)W\. Liang, Z\. Izzo, Y\. Zhang, H\. Lepp, H\. Cao, X\. Zhao, L\. Chen, H\. Ye, S\. Liu, Z\. Huang, D\. A\. McFarland, and J\. Y\. ZouMonitoring AI\-modified content at scale: a case study on the impact of ChatGPT on AI conference peer reviews\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),External Links:[Link](https://arxiv.org/abs/2403.07183)Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p2.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Lianget al\.\(2023\)W\. Liang, M\. Yuksekgonul, Y\. Mao, E\. Wu, and J\. ZouGPT detectors are biased against non\-native english writers\.Patterns4\(7\)\.External Links:[Link](https://doi.org/10.1016/j.patter.2023.100779)Cited by:[§A\.1\.3](https://arxiv.org/html/2609.38427#A1.SS1.SSS3.p4.1),[§1](https://arxiv.org/html/2609.38427#S1.p3.1),[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px3.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Lianget al\.\(2024b\)W\. Liang, Y\. Zhang, H\. Cao, B\. Wang, D\. Y\. Ding, X\. Yang, K\. Vodrahalli, S\. He, D\. S\. Smith, Y\. Yin, D\. A\. McFarland, and J\. ZouCan large language models provide useful feedback on research papers? a large\-scale empirical analysis\.NEJM AI1\(8\)\.External Links:[Link](https://doi.org/10.1056/AIoa2400196)Cited by:[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Liuet al\.\(2024\)A\. Liu, L\. Pan, Y\. Lu, J\. Li, X\. Hu, X\. Zhang, L\. Wen, I\. King, H\. Xiong, and P\. YuA survey of text watermarking in the era of large language models\.ACM Computing Surveys57\(2\),pp\. 1–36\.Cited by:[§7](https://arxiv.org/html/2609.38427#S7.p2.1)\.
- Maoet al\.\(2024\)C\. Mao, C\. Vondrick, H\. Wang, and J\. YangRaidar: generative AI detection via rewriting\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=bQWE2UqXmf)Cited by:[§7](https://arxiv.org/html/2609.38427#S7.p2.1)\.
- Mitchellet al\.\(2023\)E\. Mitchell, Y\. Lee, A\. Khazatsky, C\. D\. Manning, and C\. FinnDetectgpt: zero\-shot machine\-generated text detection using probability curvature\.InInternational conference on machine learning,pp\. 24950–24962\.Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p2.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- NeurIPS \(2026\)NeurIPSNeurIPS 2026 Main Track Handbook\.Note:[https://neurips\.cc/Conferences/2026/MainTrackHandbook](https://neurips.cc/Conferences/2026/MainTrackHandbook)Accessed: 2026\-04\-27Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p1.1),[§5\.3](https://arxiv.org/html/2609.38427#S5.SS3.p5.1),[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px2.p1.1)\.
- Nikkei Asia \(2025\)Nikkei Asia‘Positive review only’: researchers hide AI prompts in papers\.Note:[https://asia\.nikkei\.com/business/technology/artificial\-intelligence/positive\-review\-only\-researchers\-hide\-ai\-prompts\-in\-papers](https://asia.nikkei.com/business/technology/artificial-intelligence/positive-review-only-researchers-hide-ai-prompts-in-papers)Accessed: 2026\-04\-27Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p3.1),[§5\.3](https://arxiv.org/html/2609.38427#S5.SS3.p5.1)\.
- OpenAI \(2023\)OpenAINew AI classifier for indicating AI\-written text\.Note:[https://openai\.com/index/new\-ai\-classifier\-for\-indicating\-ai\-written\-text/](https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/)Accessed: 2026\-04\-27Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p3.1),[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px3.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- OpenAI \(2026\)OpenAICodex web\.Note:[https://developers\.openai\.com/codex/cloud](https://developers.openai.com/codex/cloud)Accessed: 2026\-04\-27Cited by:[§3\.2](https://arxiv.org/html/2609.38427#S3.SS2.p1.1)\.
- Sadasivanet al\.\(2025\)V\. S\. Sadasivan, A\. Kumar, S\. Balasubramanian, W\. Wang, and S\. FeiziCan AI\-generated text be reliably detected? stress testing AI text detectors under various attacks\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=OOgsAZdFOt)Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p3.1),[§5\.3](https://arxiv.org/html/2609.38427#S5.SS3.p5.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Shah \(2022\)N\. B\. ShahChallenges, experiments, and computational solutions in peer review\.Communications of the ACM65\(6\),pp\. 76–87\.External Links:[Link](https://doi.org/10.1145/3528086)Cited by:[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px3.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Shumailovet al\.\(2024\)I\. Shumailov, Z\. Shumaylov, Y\. Zhao, N\. Papernot, R\. Anderson, and Y\. GalAI models collapse when trained on recursively generated data\.Nature631,pp\. 755–759\.External Links:[Link](https://doi.org/10.1038/s41586-024-07566-y)Cited by:[§8](https://arxiv.org/html/2609.38427#S8.p4.1)\.
- Suet al\.\(2023a\)J\. Su, T\. Y\. Zhuo, D\. Wang, and P\. NakovDetectLLM: leveraging log rank information for zero\-shot detection of machine\-generated text\.InFindings of the Association for Computational Linguistics: EMNLP 2023,External Links:[Link](https://aclanthology.org/2023.findings-emnlp.827/)Cited by:[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- Suet al\.\(2023b\)Z\. Su, X\. Wu, W\. Zhou, G\. Ma, and S\. HuHC3 Plus: a semantic\-invariant human ChatGPT comparison corpus\.External Links:2309\.02731,[Link](https://arxiv.org/abs/2309.02731)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1)\.
- Thaiet al\.\(2026\)K\. Thai, B\. Emi, E\. Masrour, and M\. IyyerEditLens: quantifying the extent of AI editing in text\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gOkitaPCfZ)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p2.1)\.
- Tuftset al\.\(2025\)B\. Tufts, X\. Zhao, and L\. LiA practical examination of AI\-generated text detectors for large language models\.InFindings of the Association for Computational Linguistics: NAACL 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 4839–4856\.External Links:[Link](https://aclanthology.org/2025.findings-naacl.271/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.271),ISBN 979\-8\-89176\-195\-7Cited by:[§4\.1](https://arxiv.org/html/2609.38427#S4.SS1.p2.1)\.
- Uchenduet al\.\(2021\)A\. Uchendu, Z\. Ma, T\. Le, R\. Zhang, and D\. LeeTURINGBENCH: a benchmark environment for turing test in the age of neural text generation\.InFindings of the Association for Computational Linguistics: EMNLP 2021,pp\. 2001–2016\.External Links:[Link](https://aclanthology.org/2021.findings-emnlp.172/)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1)\.
- Vermaet al\.\(2024\)V\. Verma, E\. Fleisig, N\. Tomlin, and D\. KleinGhostbuster: detecting text ghostwritten by large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 1702–1717\.External Links:[Link](https://aclanthology.org/2024.naacl-long.95/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.95)Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p2.1),[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- Wanget al\.\(2024a\)Y\. Wang, J\. Mansurov, P\. Ivanov, J\. Su, A\. Shelmanov, A\. Tsvigun, O\. Mohammed Afzal, T\. Mahmoud, G\. Puccetti, and T\. ArnoldSemEval\-2024 task 8: multidomain, multimodel and multilingual machine\-generated text detection\.InProceedings of the 18th International Workshop on Semantic Evaluation,pp\. 2057–2079\.External Links:[Link](https://aclanthology.org/2024.semeval-1.279/),[Document](https://dx.doi.org/10.18653/v1/2024.semeval-1.279)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p2.1)\.
- Wanget al\.\(2024b\)Y\. Wang, J\. Mansurov, P\. Ivanov, J\. Su, A\. Shelmanov, A\. Tsvigun, C\. Whitehouse, O\. Mohammed Afzal, T\. Mahmoud, T\. Sasaki, T\. Arnold, A\. F\. Aji, N\. Habash, I\. Gurevych, and P\. NakovM4: multi\-generator, multi\-domain, and multi\-lingual black\-box machine\-generated text detection\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics,pp\. 1369–1407\.External Links:[Link](https://aclanthology.org/2024.eacl-long.83/)Cited by:[§5\.1](https://arxiv.org/html/2609.38427#S5.SS1.p1.1)\.
- Weber\-Wulffet al\.\(2023\)D\. Weber\-Wulff, A\. Anohina\-Naumeca, S\. Bjelobaba, T\. Foltýnek, J\. Guerrero\-Dib, O\. Popoola, P\. Sigut, and L\. WaddingtonTesting of detection tools for AI\-generated text\.International Journal for Educational Integrity19\(1\)\.External Links:[Link](https://link.springer.com/article/10.1007/s40979-023-00146-z)Cited by:[§1](https://arxiv.org/html/2609.38427#S1.p3.1),[§6](https://arxiv.org/html/2609.38427#S6.SS0.SSS0.Px3.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
- Wuet al\.\(2025\)J\. Wu, S\. Yang, R\. Zhan, Y\. Yuan, L\. S\. Chao, and D\. F\. WongA survey on llm\-generated text detection: necessity, methods, and future directions\.Computational Linguistics51\(1\),pp\. 275–338\.External Links:ISSN 0891\-2017,[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00549),[Link](https://doi.org/10.1162/coli_a_00549),https://direct\.mit\.edu/coli/article\-pdf/51/1/275/2497295/coli\_a\_00549\.pdfCited by:[§1](https://arxiv.org/html/2609.38427#S1.p2.1)\.
- Yanget al\.\(2024\)X\. Yang, W\. Cheng, Y\. Wu, L\. Petzold, W\. Y\. Wang, and H\. ChenDNA\-GPT: divergent n\-gram analysis for training\-free detection of GPT\-generated text\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Xlayxj2fWp)Cited by:[§7](https://arxiv.org/html/2609.38427#S7.p1.1)\.
- Zouet al\.\(2025\)J\. Zou, N\. Thakkar, C\. Vondrick, R\. Yu, V\. Peng, F\. Sha, A\. Garg, and The Review Feedback Agent TeamLeveraging LLM Feedback to Enhance Review Quality\.Note:[https://blog\.iclr\.cc/2025/04/15/leveraging\-llm\-feedback\-to\-enhance\-review\-quality/](https://blog.iclr.cc/2025/04/15/leveraging-llm-feedback-to-enhance-review-quality/)ICLR Blog\. Accessed: 2026\-04\-27Cited by:[§3\.1](https://arxiv.org/html/2609.38427#S3.SS1.p1.1),[§7](https://arxiv.org/html/2609.38427#S7.p3.1)\.
## Appendix AA Protocol for Reviewer\-Side AI Use
This appendix instantiates the framework end to end for journal and conference peer review\. We state a representative policy, but the partition into permitted, gray\-zone, and prohibited workflows will differ across venues, and the protocol should be re\-derived from whichever policy is actually in force\.
The broader point is procedural\. Work that claims to detect or evaluate AI use should, where possible, follow this shape: state the policy, enumerate the workflows the policy permits and prohibits, state the inferential target, and evaluate under evidence conditions that resemble the intended deployment\. Such protocols do not remove ambiguity, but they make the resulting claims interpretable and comparable\.
### A\.1Policy, inference, and evaluation
#### A\.1\.1Policy
AI use is permitted when it supports the reviewer’s own assessment without replacing confidential judgment, breaching confidentiality, or producing the recommendation\. Permitted workflows:
- •Human\-only review:the reviewer reads the manuscript and writes the review without AI\.*Evidence:*final review, portal edit history\.
- •Model language editing:the reviewer writes the review first; the model corrects grammar, clarity, tone, or organization\.*Evidence:*reviewer draft, edited review, tool log\.
- •Model summary support:the model summarizes non\-sensitive parts of the manuscript or the reviewer’s own notes, and the reviewer writes the evaluation and recommendation\.*Evidence:*reviewer notes, summary prompt, final review\.
One workflow is treated as gray\-zone:
- •Model critique expansion:the reviewer writes brief notes or bullet\-point concerns and the model expands them into a full review\. This may preserve the reviewer’s judgment, and it may also introduce unsupported criticism, shift emphasis, or create the appearance of delegated evaluation\.*Evidence:*bullet notes, expansion prompt, expanded review\.
Unless the venue resolves it explicitly, this should be a separate class rather than being forced to one side\. AI use is prohibited when it delegates evaluative judgment, breaches confidentiality, or introduces claims not grounded in the manuscript:
- •Model\-written review:the model writes the full review from the manuscript\.*Evidence:*final review, upload log, absence of any reviewer draft\.
- •Model\-assigned recommendation:the model sets the accept/reject recommendation or numerical scores\.*Evidence:*scores, submission timing, absence of a reasoning trace\.
- •Unsupported criticism:the model introduces criticisms or factual assertions the manuscript does not support\.
- •Unapproved confidential upload:the reviewer enters confidential manuscript content into an unapproved external tool\.
#### A\.1\.2Inference
In a hypothesis\-testing formulation,
H0:the review was produced through a workflow permitted by the stated review policy,H\_\{0\}:\\text\{the review was produced through a workflow permitted by the stated review policy,\}H1:the review was produced through a workflow prohibited by that policy\.H\_\{1\}:\\text\{the review was produced through a workflow prohibited by that policy\.\}The detector reports whether the available evidence suffices to rejectH0H\_\{0\}, with the uncertainty and evidence conditions under which the claim holds\. In a Bayesian formulation it estimates
P\(policy violation∣review,manuscript,review policy,available workflow evidence\),P\(\\text\{policy violation\}\\mid\\text\{review\},\\text\{manuscript\},\\text\{review policy\},\\text\{available workflow evidence\}\),and returns a structured report: estimated probability of violation, uncertainty, most likely workflow, evidence used, evidence missing, and a recommendation drawn from no concern, insufficient evidence, human review needed, or likely violation\.
#### A\.1\.3Evaluation
The benchmark is built from a shared pool of manuscripts or excerpts, each with review criteria, venue policy, and an assigned reviewer role\. Multiple workflows are generated from the same manuscript so that evaluation isolates workflow differences rather than manuscript differences\.
1. 1\.Human\-only review\.Collect manuscripts; ask reviewers to read and write reviews without AI; record scores, recommendation, strengths, weaknesses, and summary\.
2. 2\.Model language editing\.Start from human\-written reviews; prompt the model to correct grammar, clarity, tone, and organization; prohibit new criticisms, score changes, or new claims\.
3. 3\.Model summary support\.Ask reviewers to write notes after reading; prompt the model to summarize or organize those notes; ask reviewers to write the final review and recommendation; prohibit the model from evaluating the manuscript\.
4. 4\.Model critique expansion\.Start from reviewer bullet notes; prompt the model to expand them into a full review; allow light reviewer editing\. Record whether the expansion added claims absent from the notes\.
5. 5\.Model\-written review\.Start from the manuscript; prompt the model to write a full review; apply light human editing\.
6. 6\.Model\-assigned recommendation\.Start from the manuscript, reviewer notes, or a model\-written review; prompt the model to assign scores or a recommendation; optionally have the reviewer copy or justify it\.
7. 7\.Adversarial arm\.Repeat workflows 5 and 6 with an actor who paraphrases the output, imposes a personal style guide, or splices in hand\-written passages to evade detection\.
Each item stores the manuscript, the review policy, the workflow type, the policy label, reviewer notes where applicable, prompts and model outputs where applicable, the edited final review, the final scores or recommendation, and the submitted review\. Evaluation reports TPR at fixed FPR under permitted workflows, per\-workflow breakdowns, calibration of violation probabilities, abstention rate, and performance as a function of the evidence exposed to the detector\. The hypothesis to be tested is that workflow\-aware inference separates permitted assistance from delegated judgment substantially better than artifact\-only detection, and that the gap widens on the adversarial arm\.
Detectors carry a documented bias against non\-native English writers\[[Liang et al\., 2023](https://arxiv.org/html/2609.38427#bib.bib24)\], and a large share of reviewers at international venues write English as a second language\. Evaluation should therefore report FPR under permitted workflows broken out by reviewer language background\. A procedure that meets its nominal level in aggregate while exceeding it for one group is not usable for enforcement, whatever its average power\.
## Appendix BA Sketch of a Machine\-Readable Reviewer Policy
The specification below illustrates the minimum content required for a policy that a detection system can condition on and an accused reviewer can contest\. Fields are written informally rather than in any particular schema language\.
- •Scope:venue, cycle, role \(reviewer\), tasks in scope \(writing the review text, assigning scores, writing the meta\-review\)\.
- •Permitted workflows:enumerated as in Appendix[A\.1](https://arxiv.org/html/2609.38427#A1.SS1), each with the tool classes allowed and any confidentiality constraints on inputs\.
- •Prohibited workflows:enumerated, each with the rationale it protects \(confidentiality, delegated judgment, unsupported claims\)\.
- •Gray\-zone workflows:enumerated explicitly, with the default handling when observed \(for example, disclosure request rather than sanction\)\.
- •Disclosure:which workflows require disclosure, in what structured form, at what point in the cycle, and how disclosure will be used in adjudication\.
- •Evidence:which traces the venue collects and retains \(portal edit history, tool logs, reviewer notes\), retention period, and who may access them\.
- •Error tolerance:the maximum acceptable false positive rate against permitted workflows at each rung of the sanction ladder, fixed before the cycle begins\.
- •Sanction ladder:the graduated responses, the evidentiary standard required at each rung, and the statement that a detector report is never sufficient alone for the highest rungs\.
- •Appeal:what is disclosed to the accused party, the response window, and who adjudicates\.
Writing this down is useful even where no detector is deployed, because it forces a venue to decide in advance which uses it actually intends to prohibit and what rate of error against compliant reviewers it is prepared to accept\.相似文章
为何AI检测在学术诚信中失效
本文评估了学术环境中的商业AI检测器,发现对AI辅助的人类写作存在高误报率,而通过人性化工具(humanizers)几乎可以完全规避检测,因此得出结论:检测分数不应作为不当行为的独立证据。
顶级AI会议使用AI检测器拒绝涉嫌由AI撰写的论文
NeurIPS 2026使用了专有AI文本检测器,以涉嫌违反AI政策为由直接拒绝论文,但未在目标分布上验证该检测器;随后同一检测器又将会议主席自己的论文标记为可能由AI撰写。
AI政策
一位专业人士发布了一项正式政策,基于对伦理、隐私、质量及经济影响的担忧,声明拒绝在其工作中使用人工智能,同时承认该技术在行业内已无处不在。
AI辅助研究中的校准转向:证据许可声明的概念与方法论框架
本文提出了一种概念与方法论框架,用于评估AI辅助研究中的证据许可声明,强调校准作为管理科学断言权利的机制,并区分不同的AI研究路径。
AI研究工具仍过于热衷将公开信号转化为确定性
作者批评AI研究工具对微弱信号过于自信,赞扬Komo AI的快速发现和附带来源的摘要,但强调需要更好地处理不确定性和矛盾。他们描述了一种工作流程,将发现、验证和结构化检查分散到多个AI工具中。