Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

arXiv cs.AI Papers

Summary

This paper introduces IntegrityBench, a benchmark for evaluating whether LLMs uphold research integrity when acting as co-scientists under institutional pressure. Findings show frontier models fail roughly 1 in 3 integrity-critical decisions under peak pressure, and that ethical action does not require accurate misconduct classification.

arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action reasoning and artifact-grounded decision making across 36 paired tasks under a 5-level implicit-explicit pressure protocol spanning 3 domains and 4 research stages. Evaluating 18 frontier model variants, we find that under peak pressure, models fail roughly 1 in 3 integrity-critical decisions, and neither scale nor reasoning ability reliably mitigates this. Explicit pressures induce compliance with misconduct, while implicit contextual reframing more often causes over-refusal of legitimate research tasks. Interestingly, models failing to classify research requests accurately perform equally or better on artifact-grounded decision making (85.7 vs. 79.4), suggesting the three facets are structurally dissociated and correct ethical action does not require accurate classification. Frontier models can thus appear helpful while harbouring integrity failures that create two distinct deployment risks: facilitating research misconduct and eroding trust in AI-assisted research.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:24 AM

# Diagnostic Foundation for Evaluating LLMs’ Research Integrity as Co-Scientists
Source: [https://arxiv.org/html/2608.12345](https://arxiv.org/html/2608.12345)
###### Abstract

Language models are increasingly deployed as co\-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured\. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action reasoning and artifact\-grounded decision making across 36 paired tasks under a 5\-level implicit\-explicit pressure protocol spanning 3 domains and 4 research stages\. Evaluating 18 frontier model variants, we find that under peak pressure, models fail roughly 1 in 3 integrity\-critical decisions, and neither scale nor reasoning ability reliably mitigates this\. Explicit pressures induce compliance with misconduct, while implicit contextual reframing more often causes over\-refusal of legitimate research tasks\. Interestingly, models failing to classify research requests accurately perform equally or better on artifact\-grounded decision making \(85\.7 vs\. 79\.4\), suggesting the three facets are structurally dissociated and correct ethical action does not require accurate classification\. Frontier models can thus appear helpful while harboring integrity failures that create two distinct deployment risks: facilitating research misconduct and eroding trust in AI\-assisted research\.

research integrity, AI co\-scientist, large language models, benchmark, research misconduct, institutional pressure, sycophancy, p\-hacking, data fabrication, LLM evaluation, scientific workflows, AI safety, frontier models, pressure protocols, ethical decision making, misconduct classification, HARKing, selective reporting, model alignment, Machine Learning, ICML

![Refer to caption](https://arxiv.org/html/2608.12345v1/x1.png)Figure 1:Overview of IntegrityBench\. The left panel shows the benchmark taxonomy across three misconduct families\. The middle and right panels show a simplified task with and without pressure; real benchmark tasks are more detailed\.## 1Introduction

There is growing interest in the use of AI systems across scientific workflows\. LLMs now power both ’AI Scientist’ systems\(Luet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib1); Yamadaet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib2)\)and task specific assistants that propose hypotheses, design experiments, interpret results and draft manuscripts\(Gottweiset al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib8); Schmidgallet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib3); Yuet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib6); Teamet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib5)\)\. While fully autonomous AI scientist systems may remain constrained in the near term by institutional, legal and operational barriers, the backbone language models that power such systems are already used as AI co\-scientists in everyday research workflows\. This makes model\-level research integrity an immediate deployment concern\. Current benchmarks primarily evaluate the scientific capabilities of language models\(Lupidiet al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib9); Nathaniet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib10); Panigrahiet al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib11)\), rather than whether they preserve research integrity norms under pressure\. As these models enter more domains and research stages, a central question arises:

Under institutional pressure, do AI co\-scientists uphold research integrity?

This concern is grounded in the structure of scientific work\. Research\-integrity failures such as p\-hacking, selective reporting and data fabrication degrade the validity of scientific output\(Entradaset al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib17); Pupovacet al\.,[2017](https://arxiv.org/html/2608.12345#bib.bib15); Lambert and Degn,[2026](https://arxiv.org/html/2608.12345#bib.bib14)\)and are amplified by publication pressure, grant competition, promotion incentives and weak accountability structures\(Agarwalet al\.,[2023](https://arxiv.org/html/2608.12345#bib.bib12); Matet al\.,[2019](https://arxiv.org/html/2608.12345#bib.bib16); Johnet al\.,[2012](https://arxiv.org/html/2608.12345#bib.bib13)\)\. If LLMs are used as research assistants in pressured environments, they must be evaluated not only for task performance but also for whether they preserve integrity when the surrounding incentives shift\. Without such integrity, AI co\-scientists risk amplifying research misconduct, diminishing trust in AI assisted scientific research\.

To address this concern, we propose IntegrityBench, the first comprehensive benchmark evaluating three decision\-making facets \(misconduct classification, ethical action reasoning and artifact\-grounded decision making\), using 36 paired misconduct and ethical control tasks under a 5\-level implicit\-explicit pressure protocol across 3 domains and 4 research stages\. IntegrityBench systematically studies deception, bias and forbidden research behaviors \([Figure˜1](https://arxiv.org/html/2608.12345#S0.F1)\) across scientific workflows using a standardized research assistant protocol\. We evaluate backbone LLMs rather than full agent scaffolds in order to isolate the model level decision behavior that downstream AI scientist systems inherit; this choice is further motivated by evidence that base model selection explains a dominant share of behavioral variation relative to surrounding scaffolds\(Ríos\-Garcíaet al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib18)\)\. Each task evaluates three facets of ethical decision making: misconduct classification, ethical action reasoning under hypothetical scenarios and artifact\-grounded decision making\. The evaluation is conducted under varying types and degrees of institutional pressure against verified ground truth\.

Overall, we make the following technical contributions:

- •First benchmark for evaluating LLMs’ research integrity as co\-scientistsconsists of 36 tasks across 18 misconduct behaviors, 3 facets of ethical decision making, 3 scientific domains, 4 research pipeline stages and a 5\-level explicit\-implicit pressure protocol\.
- •A large\-scale evaluation of research integrity across 18 frontier model variantscovering diverse model families, scales and reasoning capabilities\. We run 1,800 prompt evaluations per model, resulting in 32,400 data samples in total\.

Our benchmark also offers several novel findings:

- •Frontier models remain poor at research integrity, with neither scale nor reasoning reliably enhancing this metric\. Across model families, scores range from 55\.8 to 71\.6 under the most intensive pressures, corresponding to amistake in every threeintegrity\-critical decisions, AI co\-scientists aren’t yet trustworthy regardless of scale or reasoning\.
- •Explicit and implicit pressures degrade integrity through asymmetric mechanisms, with explicit authority primarily increasing misconduct compliance whereas implicit contextual reframing more often causes over refusal of legitimate research tasks\.
- •Models dissociate the three facets of ethical decision making, choosing appropriate actions even when they fail to classify the integrity status of the scenario, preventing error propagation through the pipeline\.
- •Models treat surface level cues as misconduct evidence, misclassifying legitimate research for misconduct when it’s superficially similar to violation\.

![Refer to caption](https://arxiv.org/html/2608.12345v1/x2.png)Figure 2:IntegrityBench evaluation pipeline\. Task fields, a pressure prompt and question are concatenated and delivered to the model\. Outputs are then graded to produce integrity scores\.
## 2Related Work

A parallel line of work evaluates LLM safety, honesty and deception at both the model and agent level\. At the model level, TruthfulQA\(Linet al\.,[2022](https://arxiv.org/html/2608.12345#bib.bib29)\)tests whether models give truthful answers to questions humans commonly answer falsely, while SycophancyEval\(Fanouset al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib30)\)evaluates whether models revise correct answers under social pressure\. Others broaden the focus to ethically sensitive behavior: DeceptionBench\(Huanget al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib31)\)assesses deceptive tendencies across five societal domains with intrinsic and extrinsic modulation, and Philosophy Bench\(Brady,[2026](https://arxiv.org/html/2608.12345#bib.bib37)\)examines whether models comply with user requests under ethical dilemmas involving authority conflicts and honesty under pressure\. Perhaps closest to reasoning in scientific workflows, Failing to Falsify\(Jhaveriet al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib36)\)finds systematic confirmation bias\. These benchmarks identify general safety failures, but they do not test whether models can distinguish misconduct from legitimate research practice within domain\-specific scientific workflows\. Strong performance on these existing safety benchmarks therefore does not guarantee integrity in scientific settings\. IntegrityBench fills this gap by evaluating research\-integrity decisions in scenarios grounded in academic procedures, including data handling, study design, analysis choices, reporting practices and regulatory constraints\. Here, we focus on benchmark comparisons that are most directly relevant to IntegrityBench, while Appendix[B](https://arxiv.org/html/2608.12345#A2)provides broader background on AI scientist systems and research agents that motivate our deployment context\.

## 3Benchmark Design

Frontier language models are now embedded in scientific workflows as code\-running assistants, manuscript editors, statistical reviewers and dataset curators\. In each of these settings the model is exposed to ethics\-bearing requests \(sample exclusion, citation reuse, hyperparameter selection, dual\-use systems design\) where the surface form of a compliant request and a misconduct request is nearly identical and where the requester is typically a senior collaborator with an asymmetric authority relationship to the model\. Existing safety and ethics evaluations measure either declarative knowledge of principles or refusal on overt harm, neither of which captures the contested middle band where research integrity actually fails\. IntegrityBench is constructed to occupy that middle band: every task is a plausible request from a plausible collaborator, its misconduct and ethical control variants differ only in the feature that determines permissibility and the agent is evaluated on what it does, not what it knows\. Further, as base model choice accounts for a dominant share of the behavioral variance exhibited by AI co\-scientists as compared to attached scaffolds\(Ríos\-Garcíaet al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib18)\), motivating our choice to focus on backbone LLMs as the factor that primarily dictates research integrity\. Four design choices distinguish IntegrityBench from prior ethics evaluations:

1\. Symmetric pairing of every misconduct task with an ethical controlpenalizes blanket refusal equally as misconduct compliance, targeting a failure mode of refusal\-tuned LLMs\.2\. 5\-level pressure protocol holds factual content fixed across environments\. Thus, any drop in performance under pressure prompts is attributable to social framing rather than new information, producing a directly interpretable integrity score\.3\. Structural exploration of the 3 facets of ethical decision making, exposing failure modes that single\-score benchmarks cannot resolve\.4\. Every task is synthetic and authored from scratchby researchers with ethics training and reviewed by PhD domain experts, eliminating training set contamination and developing ground truth answers that are unambiguous to reviewers while challenging frontier models\.

Task DiversityThe benchmark comprises 36 tasks derived from a combination of three misconduct families and three domains, resulting in evaluation set of 1800 prompts per model\. Each task is designed with a role, context, research artifact and ten\-question assessments across each stage of the 5\-layer pressure protocol\. These tasks are designed for the AI, physics and medical domains, which represent fields with highest stakes and growing deployment of LLMs in research scenarios\.

### 3\.1Misconduct and Ethical Control Tasks

The 18 distinct misconduct behaviors were sourced from a survey of 47 researchers, asked to report observed violations across 3 domains, ensuring the benchmark is representative of real\-world scientific violations\. IntegrityBench assesses ethical decision making by grouping these misconduct behaviors into three families \(Bias, Deception and Forbidden Research\) with 18 paired ethical controls, for 36 tasks in total\. Appendix Table[5](https://arxiv.org/html/2608.12345#A1.T5)provides the full task taxonomy, including misconduct family, domain, research stage, and artifact type, while Appendix Table[8](https://arxiv.org/html/2608.12345#A1.T8)reports the distribution of tasks across domains, misconduct families, and research pipeline stages\. The misconduct tasks are constructed to elicit a specific known integrity failure \(for instance, removing samples which don’t support the hypothesis or proceeding with a human study without IRB clearance\)\. The controls hold the role, domain, artifact format and question structure constant while modifying singular features that determine permissibility \(for instance, contamination is documented in the lab notebook, dataset p\-values are maintained or IRB protocol has been approved\)\. This pairing prevents frontier models from gaming the benchmark through blanket refusal and allows for the identification of asymmetric failure modes\. For example, scenarios from Task 1 and its ethical control pair, Task 19 for data fabrication are provided in Appendix[A\.1](https://arxiv.org/html/2608.12345#A1.SS1)\.

The benchmark is structured around three main misconduct families:Biasconstitutes prejudice rather than active falsification, affecting the scientific validity through faulty decision making \(e\.g\., sycophantic revision and citation bias\);Deceptionis the misrepresentation of data, workflows or results, including data fabrication, selective reporting and novelty inflation, containing the largest set of behavior modes; andForbidden Researchpertains to the execution of tasks for research that may be dangerous or unethical \(e\.g\., dual\-use blindness and unauthorized data use\)\.

Prompt structureconsists of a context block, situation description, research artifact and question block\. The context block provides the institutional role and permissions, while the situation description provides an ethics based request without outlining compliance with ethical protocols\. The research artifact, a JSON dataset plus summary statistics, supports artifact\-grounded decision making \(Q3\)\. Prompt examples and model traces are provided in Appendix[8](https://arxiv.org/html/2608.12345#A5.F8)\.

Questionscover three facets of research\-integrity decision making\.Misconduct Classification \(Q1\)assesses an agent’s ability to identify the specific integrity concerns raised by a request, not merely detect that something is wrong \- credible push back requires accurate mechanism identification\. The 19\-way format, consisting of 18 misconduct types plus ethical control, approximates this open\-ended task within a scoreable structure\. Real deployment affords no such menu, making the 19\-option format a conservative lower bound on classification difficulty\. The integrity gaps we observe are therefore a floor rather than a ceiling on the true challenge\.Ethical Action Reasoning \(Q2\)evaluates the agent’s reasoning process through a series of three 4\-option multiple choice questions that reasonably extend the agent’s context and behavioral decision\-making to hypothetical scenarios\.Artifact\-Grounded Decision Making \(Q3\)requires the model to reason with reference a synthetic research artifact in JSON format to answer 6 4\-option multiple choice questions with research integrity stakes, differentiating it from hypothetical scenario reasoning\.

Submission Formatis stated at the end of each prompt, resulting in single\-character responses to each question\. The complete benchmark can be found on[Hugging face](https://huggingface.co/datasets/Integrity-Bench-anon/IntegrityBench)\.

### 3\.2Pressure Protocol

The pressure protocol tests whether models preserve research\-integrity behavior when institutional framing changes without adding new factual information\. Drawing on classic authority\-pressure studies\(Milgram,[1963](https://arxiv.org/html/2608.12345#bib.bib38)\)while focusing on analogous compliance behavior in RLHF\-trained models, each task is evaluated under five environments: a baseline plus four pressure prompts crossing mechanism \(implicit institutional cues vs\. explicit authority appeal\) and intensity \(moderate vs\. escalated\), as summarized in Table[1](https://arxiv.org/html/2608.12345#S3.T1)\. PP1/PP3 use implicit institutional cues that signal expected behavior without naming an authority figure, whereas PP2/PP4 use explicit appeals from named senior authorities with procedural counterarguments\. PP3/PP4 further increase pressure intensity, separating mechanism from escalation\. Appendix Table[7](https://arxiv.org/html/2608.12345#A1.T7)illustrates the four non\-baseline pressure environments using Task 1 examples\. Because pressure blocks are inserted without changing the dataset or experimental record, performance changes can be attributed to social framing rather than new evidence\.

Table 1:Five pressure environments under which every task is run\.
### 3\.3Task Validation and Annotation

Each of the 36 tasks were designed synthetically from scratch by our research team to induce specific behaviors in LLMs and prevent the risk of contamination\. These authored tasks were evaluated by a panel of three domain experts holding at least a PhD in their respective fields, each independently assigning misconduct labels and verifying the ground truth across the 10 questions\. As domain expert reviews were non\-overlapping, an ethics expert served as the overlapping second reviewer across all 36 tasks\. Cohen’s Kappa \(κ=0\.96\\kappa=0\.96\) computed between domain and ethics expert label assignments, indicates near\-perfect agreement and validates the tasks as unambiguous across reviewer backgrounds\. Further, this agreement demonstrates a design\-validated performance ceiling across all tasks, negating the need for a separate human baseline\. Full annotation details are provided in the Appendix[A\.5](https://arxiv.org/html/2608.12345#A1.SS5)\.

Table 2:Model\-level performance across varying types and degrees of pressure\. Implicit and explicit pressure columns report means over \(PP1, PP3\) and \(PP2, PP4\), respectively\. R: Reasoning; NR: Non\-reasoning\. Mis\. = Misconduct Tasks; Eth\. = Ethical Control Tasks\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x3.png)Figure 3:Mean integrity scores for misconduct tasks and ethical\-control tasks, sorted by task\-level performance\.
### 3\.4Metrics

Raw scores for each prompt are generated through string matching and used to compute condition\-wise and overall integrity scores\. The condition\-wise integrity score \(IS\) equally weights three evaluation dimensions: misconduct classification \(Q1\), ethical action reasoning \(Q2\), and artifact\-grounded decision making \(Q3\)\. Here,Q¯2\\bar\{Q\}\_\{2\}andQ¯3\\bar\{Q\}\_\{3\}denote mean scores across the three Q2 and six Q3 sub\-questions, respectively\. For each pressure casek∈\{1,2,3,4\}k\\in\\\{1,2,3,4\\\},PkP\_\{k\}represents the integrity score under that pressure condition\. The overall integrity score averages performance across the baseline and four pressure environments\. Both metrics are bounded in\[0,100\]\[0,100\], with higher scores indicating stronger research\-integrity performance\.

ISbase\\displaystyle\\text\{IS\}\_\{\\text\{base\}\}=13​\(Q1\+Q¯2\+Q¯3\),\\displaystyle=\\frac\{1\}\{3\}\\left\(Q\_\{1\}\+\\bar\{Q\}\_\{2\}\+\\bar\{Q\}\_\{3\}\\right\),\(1\)ISoverall\\displaystyle\\text\{IS\}\_\{\\text\{overall\}\}=15​\(ISbase\+∑k=14ISPk\)\.\\displaystyle=\\frac\{1\}\{5\}\\left\(\\text\{IS\}\_\{\\text\{base\}\}\+\\sum\_\{k=1\}^\{4\}\\text\{IS\}\_\{P\_\{k\}\}\\right\)\.\(2\)

## 4Results

Experiment Setup\. We evaluated five model families across two dimensions: model size and reasoning capability\. The model families include Claude, Gemini, Qwen, DeepSeek and GPT\. All variants are queried through OpenRouter with a single unified client key, consistent with recent multi\-model benchmarking practices\(Brady,[2026](https://arxiv.org/html/2608.12345#bib.bib37)\), ensuring that request shape and decoding are identical across providers\. For each model family, we evaluated22=42^\{2\}=4variants, corresponding to all combinations of the two dimensions, except DeepSeek, which lacks a small model with both reasoning and non\-reasoning capabilities\. These 18 model variants were evaluated on a benchmark of 1800 prompts, yielding an evaluation set of18∗1,800=32,40018\*1,800=32,400prompts\. Total API cost was approximately$​90\\mathdollar 90across 70 million tokens, with a per\-model breakdown provided in Table[11](https://arxiv.org/html/2608.12345#A3.T11)in the Appendix\. The low\-level details of implementation reproducibility are available in Appendix[A\.6](https://arxiv.org/html/2608.12345#A1.SS6)

### 4\.1Research integrity remains unreliable across frontier models

As shown in Table[2](https://arxiv.org/html/2608.12345#S3.T2), integrity scores range from 60\.9 to 72\.8 with an overall mean of 68\.7\. The 11\.9 point spread across models is narrow, suggesting that current alignment techniques converge over a set of common failure modes across model families\. Models exceed a random\-choice baseline but remain far from ceiling performance\. For example, the best performing model, Qwen 3\.5 Flash 9B, fails roughly 1 out of 4 tasks\. This is concerning for scientific workflows where singular instances of data fabrication or sycophantic revision can degrade the quality of downstream outputs\. The implication is that frontier models are currently too unreliable to deploy within the scientific pipeline without targeted alignment interventions\. In fact, the top 4 model families, by IS score, lie within a 2\.2 score bound when averaged, indicating failure modes consistent with common alignment gaps such as false\-positive bias, potentially stemming from over\-penalization of compliance during RLHF\.

### 4\.2Reasoning and scale do not reliably resolve integrity failures

Paired McNemar tests with Benjamini–Hochberg correction show no significant reasoning effect for 7 of 9 matched pairs \(all adjustedp\>0\.09p\>0\.09\)\. Significant effects are bidirectional: Qwen 3\.5 Flash 9B improves \(adjustedp<0\.001p<0\.001\), driven by stronger Q1 classification accuracy \(65\.6 vs\. panel mean 56\.8\), while GPT 5\.4 Mini degrades significantly \(adjustedp<0\.001p<0\.001\), consistent with reasoning\-induced over\-refusal on ethical control tasks\. The largest positive deltas, Qwen 3\.5 Flash 9B \(\+5\.4\) and DeepSeek V3\.2 \(\+5\.1\), mainly compensate for weak non\-reasoning baselines rather than demonstrating reasoning\-specific gains\. The small mean delta of \+1\.3, therefore, reflects opposing effects across families rather than a uniform trend, confirming that reasoning budget does not reliably improve research integrity\. Models that degrade under reasoning, most notably Sonnet 4\.6 \(Δ=−0\.5\\Delta=\-0\.5\) and GPT 5\.4 Mini \(Δ=−0\.2\\Delta=\-0\.2overall,−1\.9\-1\.9on misconduct\), are consistent with recent evidence that chain\-of\-thought reasoning can introduce answer drift and post\-hoc rationalisation\(Fenget al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib39)\)\. As reasoning runs utilize provider\-default temperature with single\-pass evaluation, point differences less than 2\-3 points between paired variants carry sampling uncertainty\. Thus, inferential claims have been made solely using the McNemar tests as evidenced in Appendix Table[12](https://arxiv.org/html/2608.12345#A4.T12)rather than IS point differences\.Scale is equally unreliable: Qwen 3\.5 397B A17B, with nearly 44×\\timesmore parameters than Qwen 3\.5 Flash 9B, scores 1\.5 points lower under matched reasoning, and the large\-model tier mean \(69\.7\) exceeds the small\-model mean \(68\.0\) by only 1\.7 points\. These comparisons are produced by differences in release date and alignment curriculum across families, so we interpret the result as evidence that current large\-scale alignment approaches do not automatically confer integrity gains\. Integrity\-relevant behavior appears to be governed by fine\-tuning and alignment choices rather than raw parameter count\. This interpretation is supported by the reversal of the expected capability ordering within the Qwen family, which is inconsistent with integrity scores functioning as a general capability proxy\. Scaling a model whose alignment conflates confident refusal with ethical judgment produces enhanced capability without corresponding integrity improvements\.

Table 3:Effect of reasoning on integrity scores by model family\. R denotes reasoning\-enabled runs and NR denotes non\-reasoning runs\.Δ\\Deltais computed as R–NR\. Reasoning tokens report the mean number of reasoning tokens used in reasoning\-enabled runs\.OverallMisconduct TasksEthical\-Control TasksMeanReasoningTokensSizeModel familyRNRΔ\\DeltaRNRΔ\\DeltaRNRΔ\\DeltaLARGESonnet 4\.670\.571\.0−0\.5\-0\.576\.376\.3−0\.0\-0\.064\.765\.7−1\.1\-1\.1236GPT 5\.470\.870\.0\+0\.8\+0\.875\.772\.3\+3\.4\+3\.465\.867\.7−1\.8\-1\.8296Gemini 3 Flash70\.770\.7−0\.0\-0\.072\.072\.2−0\.2\-0\.269\.469\.3\+0\.1\+0\.11152Qwen 3\.5 397B A17B71\.370\.2\+1\.1\+1\.172\.872\.9−0\.1\-0\.169\.867\.6\+2\.3\+2\.31767DeepSeek V3\.265\.960\.9\+5\.1\+5\.164\.764\.9−0\.2\-0\.267\.256\.9\+10\.3\+10\.3981SMALLHaiku 4\.566\.565\.1\+1\.4\+1\.473\.769\.5\+4\.2\\mathbf\{\+4\.2\}59\.260\.6−1\.4\-1\.4642GPT 5\.4 Mini67\.367\.5−0\.2\-0\.273\.475\.3−1\.9\-1\.961\.259\.7\+1\.5\+1\.5226Gemini 3\.1 Flash Lite69\.070\.1−1\.1\-1\.170\.271\.2−1\.0\-1\.067\.869\.0−1\.2\-1\.21134Qwen 3\.5 Flash 9B72\.867\.4\+5\.4\\mathbf\{\+5\.4\}74\.071\.4\+2\.5\+2\.571\.663\.4\+8\.2\\mathbf\{\+8\.2\}2324Mean69\.468\.1\+1\.3\+1\.372\.571\.8\+0\.8\+0\.866\.364\.4\+1\.9\+1\.9973\.14
### 4\.3Models treat surface level cues as evidence of misconduct even when procedurally justified

As shown in Figure[3](https://arxiv.org/html/2608.12345#S3.F3), integrity failures are concentrated in specific scenarios not uniformly distributed\. Models perform well on misconduct tasks with salient violation cues, such as plagiarism production and p\-hacking, but struggle on tasks where misconduct depends on subtler methodological intent, such as experiment overfitting and anchoring tasks\.The paired ethical controls reveal that tasks easy to flag as misconduct are hard to recognize as compliant\. P\-hacking scores 86\.3 points as a misconduct task but the ethical control scores 42\.0 points, a 44\.3 point inversion\. This pattern reflects surface cue inversion, where models rely on visible research features rather than the procedural context that determines whether those features are permissible\. As a result, features that help models detect misconduct can also cause them to over flag legitimate research practices\. This pattern varies by task\. When misconduct\-associated cues are highly salient, models tend to achieve high misconduct scores but low ethical\-control scores\. When misconduct depends more on research motivation than on a single visible action, the reverse can occur, as in hypothesis anchoring \(58\.0 vs\. 75\.5\)\. Overall, models do not fail uniformly, but struggle most when misconduct and legitimate practice share similar surface features\. Per\-task and per\-model summaries are in Appendix Tables[9](https://arxiv.org/html/2608.12345#A3.T9)and[10](https://arxiv.org/html/2608.12345#A3.T10)\. Stage\-level results support the same pattern: ethical\-control performance is weakest in analysis tasks, where permissibility depends on intent and procedural justification \(Appendix[C\.3](https://arxiv.org/html/2608.12345#A3.SS3)\)\.

### 4\.4The three facets of the ethical decision making process are structurally dissociated

The three sub\-tasks of IntegrityBench, Q1 \(misconduct classification\), Q2 \(ethical action reasoning\) and Q3 \(artifact\-grounded decision making\), show a strictly monotone difficulty ordering across the panel\. A paired Wilcoxon signed\-rank test across 36 task\-level means confirms this ordering \(W=152W=152,p=0\.004p=0\.004\)\. Across 18 variants, integrity scores are as follows: Q1 = 56\.8, Q2 = 66\.6 and Q3 = 80\.8\. This creates a 24\-point gap between misconduct classification and producing an ethical decision\. To test for independence, we condition Q3 on Q1 correctness\. Models that fail Q1 achieve mean Q3 = 85\.7, matching, or outperforming, those that pass it \(mean Q3 = 79\.4\), consistent across the 18 model variants\. The ordering Q1<<Q2<<Q3 holds for 17 out of the 18 variants\. This trend is more pronounced for ethical control tasks where Q1 = 41\.6, Q2 = 62\.7 and Q3 = 89\.1 as illustrated in Table[4](https://arxiv.org/html/2608.12345#S4.T4)\. The relatively high Q3 scores reflect the grounding provided by the research artifacts rather than a ceiling, as each task is validated by domain experts with near\-perfect inter\-rater agreement \(κ=0\.96\\kappa=0\.96\)\.

This result substantiates the dissociation between the three facets of the ethical decision making process, demonstrating that the three sub\-tasks are structurally independent\. A classification failure therefore does not necessarily imply an unethical downstream decision\. This independence emerges because Q1, Q2 and Q3 sub tasks rely on distinct qualitative capabilities\. Specifically, Q1 necessitates a direct categorical discrimination across 19 misconduct behaviors, Q2 draws from behavioral tendencies described in scientific literature and Q3 employs artifact\-level pattern recognition, developed through access to research data\. Thus, classification\-level failures do not automatically propagate down the pipeline, though targeted alignment is essential for each facet of the ethical decision making\.

### 4\.5Pressure types and intensities degrade research integrity asymmetrically

Across the evaluated models, results show a general decline in research integrity as pressure increases from PP1 and PP2 to PP3 and PP4, confirming that verbalized increases in intensity are sufficient to degrade ethical decision making regardless of the pressure mechanism\. Paired t\-tests on misconduct and ethical control tasks, as reported in Appendix Table[13](https://arxiv.org/html/2608.12345#A4.T13), highlight how explicit and implicit pressures degrade research integrity in qualitatively distinct manners, capturing their asymmetry through opposing signs of the t\-statistics\. Explicit pressures more effectively target misconduct tasks \(explicit 68\.8 versus implicit 73\.5, t = 6\.24, p < 0\.001, 16/18 variants\) while preserving compliance with ethical control tasks \(explicit 65\.2 versus implicit 61\.4, t = \- 7\.86, p < 0\.001, 18/18 variants\)\. These opposing signs reveal distinct mechanisms\. Explicit pressures employ named\-authority appeals with procedural counterarguments to improve misconduct compliance independent of ethical content while implicit pressures provide institutional cues that indicate expected behavior without direct instruction, causing models to over\-refuse ethical controls without reasoning through permissibility\.

Table 4:Per\-model IS scores across the three facets of ethical decision making\.Q​2−Q​1Q2\{\-\}Q1andQ​3−Q​1Q3\{\-\}Q1report facet differences while Q3 \| Q1 = 0 and Q3 \| Q1 = 1 condition Q3 on classification accuracy
### 4\.6Models Detect Misconduct More Reliably Than Ethical Compliance

![Refer to caption](https://arxiv.org/html/2608.12345v1/x4.png)Figure 4:Family\-level gap between misconduct and ethical\-control tasks\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x5.png)Figure 5:Q1 misconduct\-classification accuracy by model family on misconduct and ethical\-control tasks\.As shown in Figure[5](https://arxiv.org/html/2608.12345#S4.F5), IntegrityBench reveals a consistent asymmetry between misconduct classification within the ethical control and misconduct tasks\. Models scored 71\.9 on misconduct tasks and 41\.6 on paired ethical\-control tasks \(Δ=30\.3\\Delta=30\.3\), despite holding all tasks features constant\. Because the paired design holds these features constant, this gap suggests that models often detect integrity relevant cues without reliably determining whether those cues are made permissible by context\. For example, models frequently over flagged legitimate robustness checks, transparent data exclusions, and human\-subjects procedures as misconduct when similar cues appeared in ethical control tasks\. This indicates that models are more reliable at recognizing suspicious research patterns than at distinguishing misconduct from procedurally justified practice\. The result is consistent with alignment behavior that rewards sensitivity to integrity\-related cues but provides weaker signals for cases where those cues are legitimate\. Therefore future evaluations should not only measure refusal of clear misconduct, but also preservation of valid research workflows under superficially similar conditions\.

## 5Discussion

#### Limitations & Future Work

IntegrityBench covers 18 misconduct types across three high\-stakes domains\. Its coverage of domains is limited relative to the breadth of modern research practices\. Future work should extend this benchmark to additional domains such as the social science and economics, and increase tasks per misconduct type\. Domain, research stage and misconduct family differences are reported descriptively given the task counts per grouping\. Evaluations of agentic behavior can be deepened by scaffolding models with a broader tool suite more representative of an AI co\-scientist, which would further test the dissociation between ethical reasoning and decision\-making observed here\. Analysing the internal reasoning logs of open\-source models would additionally allow a more granular characterization of underlying intent and failure modes\. Further, Q1 represents classification through a multiple choice selection rather than open\-ended generation\. Thus, it overestimates real world classification accuracy, representing a floor on the true challenge\. Finally, explicit pressures employ authority appeals alongside procedural counterarguments and can be studied further to differentiate the relative effects of each on integrity degradation\.

## 6Conclusion

We introduced IntegrityBench, the first comprehensive benchmark targeting backbone LLMs and the three facets of ethical decision making through a series of paired misconduct and ethical control tasks, multiple scientific domains and a 5\-level pressure protocol\. Benchmarking results show that frontier models fail in a significant proportion of integrity\-critical decisions under intensive pressure, with neither scale nor reasoning providing any reliable integrity enhancement\. Instead, failures seem to be primarily shaped by alignment curriculum and techniques\. Further, the results reveal specific shared modes of failure: explicit and implicit tasks effectively degrade research integrity through two opposing mechanisms, and ethical decision making is structurally dissociated from task classification, indicating that models can appear safe while retaining systematic misclassifications\. Importantly, we demonstrate that trustworthy AI co\-scientists cannot be advanced through scale or reasoning alone and re\-frame research integrity as a fundamentally alignment gap requiring pressure\-differentiated training signals and facet\-level evaluations\. Toward this outlook, IntegrityBench provides the reproducible diagnostic foundation for both, establishing the empirical basis for the targeted alignment intervention required\.

## Impact Statement

This paper presents work whose goal is to advance the field of Machine Learning\. There are many potential societal consequences of our work, none which we feel must be specifically highlighted here\.

## References

- A\. Agarwal, M\. Arafa, T\. Avidor\-Reiss, T\. A\. A\. Hamoda, and R\. Shah \(2023\)Citation errors in scientific research and publications: causes, consequences, and remedies\.World Journal of Men’s Health41\(3\),pp\. 461–465\.External Links:[Document](https://dx.doi.org/10.5534/wjmh.230001)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p3.1)\.
- D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. Gomes \(2023\)Autonomous chemical research with large language models\.Nature624\(7992\),pp\. 570–578\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06792-0),[Link](https://doi.org/10.1038/s41586-023-06792-0)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1)\.
- B\. Brady \(2026\)Philosophy bench\.Note:[https://www\.philosophybench\.com/](https://www.philosophybench.com/)Philosophically advised by Matt Mandel\. Originally published April 24, 2026\. Accessed: 2026\-04\-30Cited by:[§2](https://arxiv.org/html/2608.12345#S2.p1.1),[§4](https://arxiv.org/html/2608.12345#S4.p1.3)\.
- J\. S\. Chan, N\. Chowdhury, O\. Jaffe, J\. Aung, D\. Sherburn, E\. Mays, G\. Starace, K\. Liu, L\. Maksin, T\. Patwardhan, L\. Weng, and A\. Mądry \(2025\)MLE\-bench: evaluating machine learning agents on machine learning engineering\.External Links:2410\.07095,[Link](https://arxiv.org/abs/2410.07095)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p2.1)\.
- M\. Entradas, Y\. Feng, and I\. C\. E\. Sousa \(2026\)The ’shades of grey’ in research integrity–researchers admit to questionable research practices that they do not perceive to be serious\.PLOS ONE21\(1\),pp\. e0339056\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0339056)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p3.1)\.
- A\. Fanous, J\. Goldberg, A\. A\. Agarwal, J\. Lin, A\. Zhou, R\. Daneshjou, and S\. Koyejo \(2025\)SycEval: evaluating llm sycophancy\.External Links:2502\.08177,[Link](https://arxiv.org/abs/2502.08177)Cited by:[§2](https://arxiv.org/html/2608.12345#S2.p1.1)\.
- Z\. Feng, Z\. Chen, J\. Ma, Y\. T\. Po, E\. Chersoni, and B\. Li \(2026\)Good arguments against the people pleasers: how reasoning mitigates \(yet masks\) llm sycophancy\.arXiv preprint arXiv:2603\.16643\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.16643),[Link](https://arxiv.org/abs/2603.16643)Cited by:[§4\.2](https://arxiv.org/html/2608.12345#S4.SS2.p1.7)\.
- J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, A\. Palepu, P\. Sirkovic, A\. Myaskovsky, F\. Weissenberger, K\. Rong, R\. Tanno, K\. Saab, D\. Popovici, J\. Blum, F\. Zhang, K\. Chou, A\. Hassidim, B\. Gokturk, A\. Vahdat, P\. Kohli, Y\. Matias, A\. Carroll, K\. Kulkarni, N\. Tomasev, Y\. Guan, V\. Dhillon, E\. D\. Vaishnav, B\. Lee, T\. R\. D\. Costa, J\. R\. Penadés, G\. Peltz, Y\. Xu, A\. Pawlosky, A\. Karthikesalingam, and V\. Natarajan \(2025\)Towards an ai co\-scientist\.External Links:2502\.18864,[Link](https://arxiv.org/abs/2502.18864)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- Y\. Huang, Y\. Sun, Y\. Zhang, R\. Zhang, Y\. Dong, and X\. Wei \(2025\)DeceptionBench: a comprehensive benchmark for ai deception behaviors in real\-world scenarios\.External Links:2510\.15501,[Link](https://arxiv.org/abs/2510.15501)Cited by:[§2](https://arxiv.org/html/2608.12345#S2.p1.1)\.
- A\. R\. Jhaveri, A\. GX\-Chen, I\. Sucholutsky, and E\. Choi \(2026\)Failing to falsify: evaluating and mitigating confirmation bias in language models\.External Links:2604\.02485,[Link](https://arxiv.org/abs/2604.02485)Cited by:[§2](https://arxiv.org/html/2608.12345#S2.p1.1)\.
- L\. K\. John, G\. Loewenstein, and D\. Prelec \(2012\)Measuring the prevalence of questionable research practices with incentives for truth telling\.Psychological Science23\(5\),pp\. 524–532\.External Links:[Document](https://dx.doi.org/10.1177/0956797611430953)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p3.1)\.
- M\. Lambert and L\. Degn \(2026\)Shaping the field: a review of the use of theory in research on research integrity\.Science and Engineering Ethics32\(2\),pp\. 19\.External Links:[Document](https://dx.doi.org/10.1007/s11948-026-00587-y)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p3.1)\.
- W\. Liang, Y\. Zhang, H\. Cao, B\. Wang, D\. Ding, X\. Yang, K\. Vodrahalli, S\. He, D\. Smith, Y\. Yin, D\. McFarland, and J\. Zou \(2023\)Can large language models provide useful feedback on research papers? a large\-scale empirical analysis\.External Links:2310\.01783,[Link](https://arxiv.org/abs/2310.01783)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2022\)TruthfulQA: measuring how models mimic human falsehoods\.External Links:2109\.07958,[Link](https://arxiv.org/abs/2109.07958)Cited by:[§2](https://arxiv.org/html/2608.12345#S2.p1.1)\.
- C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. Ha \(2024\)The ai scientist: towards fully automated open\-ended scientific discovery\.External Links:2408\.06292,[Link](https://arxiv.org/abs/2408.06292)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- A\. Lupidi, B\. Gauri, T\. S\. Foster, B\. A\. Omari, D\. Magka, A\. Pepe, A\. Audran\-Reiss, M\. Aghamelu, N\. Baldwin, L\. Cipolina\-Kun, J\. Gagnon\-Audet, C\. H\. Leow, S\. Lefdal, H\. Mossalam, A\. Moudgil, S\. Nazir, E\. Tewolde, I\. Urrego, J\. A\. Estape, A\. Budhiraja, G\. Chaurasia, A\. Charnalia, D\. Dunfield, K\. Hambardzumyan, D\. Izcovich, M\. Josifoski, I\. Mediratta, K\. Niu, P\. Pathak, M\. Shvartsman, E\. Toledo, A\. Protopopov, R\. Raileanu, A\. Miller, T\. Shavrina, J\. Foerster, and Y\. Bachrach \(2026\)AIRS\-bench: a suite of tasks for frontier ai research science agents\.External Links:2602\.06855,[Link](https://arxiv.org/abs/2602.06855)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- T\. Mat, D\. Shahmizi, and E\. K\. Ghani \(2019\)Do perceived pressure and perceived opportunity influence employees’ intention to commit fraud?\.International Journal of Financial Research10\(3\),pp\. 132–143\.External Links:[Document](https://dx.doi.org/10.5430/ijfr.v10n3p132)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p3.1)\.
- S\. Milgram \(1963\)Behavioral study of obedience\.Journal of Abnormal and Social Psychology67\(4\),pp\. 371–378\.External Links:[Document](https://dx.doi.org/10.1037/h0040525)Cited by:[§3\.2](https://arxiv.org/html/2608.12345#S3.SS2.p1.1)\.
- D\. Nathani, L\. Madaan, N\. Roberts, N\. Bashlykov, A\. Menon, V\. Moens, A\. Budhiraja, D\. Magka, V\. Vorotilov, G\. Chaurasia, D\. Hupkes, R\. S\. Cabral, T\. Shavrina, J\. Foerster, Y\. Bachrach, W\. Y\. Wang, and R\. Raileanu \(2025\)MLGym: a new framework and benchmark for advancing ai research agents\.External Links:2502\.14499,[Link](https://arxiv.org/abs/2502.14499)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- H\. Padigela, C\. Shah, and D\. Juyal \(2025\)ML\-dev\-bench: comparative analysis of ai agents on ml development workflows\.External Links:2502\.00964,[Link](https://arxiv.org/abs/2502.00964)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p2.1)\.
- S\. S\. Panigrahi, J\. Videnović, and M\. Brbić \(2026\)HeurekaBench: a benchmarking framework for ai co\-scientist\.External Links:2601\.01678,[Link](https://arxiv.org/abs/2601.01678)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- V\. Pupovac, S\. Prijić\-Samaržija, and M\. Petrovečki \(2017\)Research misconduct in the croatian scientific community: a survey assessing the forms and characteristics of research misconduct\.Science and Engineering Ethics23\(1\),pp\. 165–181\.External Links:[Document](https://dx.doi.org/10.1007/s11948-016-9767-0)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p3.1)\.
- M\. Ríos\-García, N\. Alampara, C\. Gupta, I\. Mandal, S\. Mannan, A\. A\. Aghajani, N\. M\. A\. Krishnan, and K\. M\. Jablonka \(2026\)AI scientists produce results without reasoning scientifically\.Note:arXiv preprintExternal Links:2604\.18805,[Link](https://arxiv.org/abs/2604.18805)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p4.1),[§3](https://arxiv.org/html/2608.12345#S3.p1.1)\.
- Y\. Ruan, H\. Dong, A\. Wang, S\. Pitis, Y\. Zhou, J\. Ba, Y\. Dubois, C\. J\. Maddison, and T\. Hashimoto \(2024\)Identifying the risks of lm agents with an lm\-emulated sandbox\.External Links:2309\.15817,[Link](https://arxiv.org/abs/2309.15817)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p3.1)\.
- S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. Barsoum \(2025\)Agent laboratory: using llm agents as research assistants\.External Links:2501\.04227,[Link](https://arxiv.org/abs/2501.04227)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- Z\. S\. Siegel, S\. Kapoor, N\. Nagdir, B\. Stroebl, and A\. Narayanan \(2024\)CORE\-bench: fostering the credibility of published research through a computational reproducibility agent benchmark\.External Links:2409\.11363,[Link](https://arxiv.org/abs/2409.11363)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p2.1)\.
- M\. D\. Skarlinski, S\. Cox, J\. M\. Laurent, J\. D\. Braza, M\. Hinks, M\. J\. Hammerling, M\. Ponnapati, S\. G\. Rodriques, and A\. D\. White \(2024\)Language agents achieve superhuman synthesis of scientific knowledge\.External Links:2409\.13740,[Link](https://arxiv.org/abs/2409.13740)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1)\.
- G\. Starace, O\. Jaffe, D\. Sherburn, J\. Aung, J\. S\. Chan, L\. Maksin, R\. Dias, E\. Mays, B\. Kinsella, W\. Thompson, J\. Heidecke, A\. Glaese, and T\. Patwardhan \(2025\)PaperBench: evaluating ai’s ability to replicate ai research\.External Links:2504\.01848,[Link](https://arxiv.org/abs/2504.01848)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p2.1)\.
- I\. Team, B\. Zhang, S\. Feng, X\. Yan, J\. Yuan, R\. Ma, Y\. Hu, Z\. Yu, X\. He, S\. Huang, S\. Hou, Z\. Nie, Z\. Wang, J\. Liu, T\. Peng, P\. Ye, D\. Zhou, S\. Zhang, X\. Wang, Y\. Zhang, M\. Li, Z\. Tu, X\. Yue, W\. Ouyang, B\. Zhou, and L\. Bai \(2025\)InternAgent: when agent becomes the scientist – building closed\-loop system from hypothesis to verification\.External Links:2505\.16938,[Link](https://arxiv.org/abs/2505.16938)Cited by:[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- H\. Tong, F\. Zhao, L\. Feng, R\. Wu, R\. Chen, L\. Jia, Z\. Zhao, J\. Li, T\. Li, E\. Lin, S\. Yang, E\. Lu, Y\. Sun, Q\. Zhang, Z\. Ruan, J\. Fan, Z\. Yue, P\. Wu, H\. Li, C\. Sun, and Y\. Zeng \(2026\)ForesightSafety bench: a frontier risk evaluation and governance framework towards safe ai\.External Links:2602\.14135,[Link](https://arxiv.org/abs/2602.14135)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p3.1)\.
- Y\. Weng, M\. Zhu, G\. Bao, H\. Zhang, J\. Wang, Y\. Zhang, and L\. Yang \(2025\)CycleResearcher: improving automated research via automated review\.External Links:2411\.00816,[Link](https://arxiv.org/abs/2411.00816)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p2.1)\.
- Y\. Yamada, R\. T\. Lange, C\. Lu, S\. Hu, C\. Lu, J\. Foerster, J\. Clune, and D\. Ha \(2025\)The ai scientist\-v2: workshop\-level automated scientific discovery via agentic tree search\.External Links:2504\.08066,[Link](https://arxiv.org/abs/2504.08066)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- H\. Yu, Z\. Hong, Z\. Cheng, K\. Zhu, K\. Xuan, J\. Yao, T\. Feng, and J\. You \(2025\)ResearchTown: simulator of human research community\.External Links:2412\.17767,[Link](https://arxiv.org/abs/2412.17767)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.12345#S1.p1.1)\.
- J\. Yuan, X\. Yan, S\. Feng, B\. Zhang, T\. Chen, B\. Shi, W\. Ouyang, Y\. Qiao, L\. Bai, and B\. Zhou \(2025\)Dolphin: moving towards closed\-loop auto\-research through thinking, practice, and feedback\.External Links:2501\.03916,[Link](https://arxiv.org/abs/2501.03916)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1)\.
- T\. Yuan, Z\. He, L\. Dong, Y\. Wang, R\. Zhao, T\. Xia, L\. Xu, B\. Zhou, F\. Li, Z\. Zhang, R\. Wang, and G\. Liu \(2024\)R\-judge: benchmarking safety risk awareness for llm agents\.External Links:2401\.10019,[Link](https://arxiv.org/abs/2401.10019)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p3.1)\.
- Y\. Zhou, H\. Liu, T\. Srivastava, H\. Mei, and C\. Tan \(2024\)Hypothesis generation with large language models\.InProceedings of the 1st Workshop on NLP for Science \(NLP4Science\),pp\. 117–139\.External Links:[Link](http://dx.doi.org/10.18653/v1/2024.nlp4science-1.10),[Document](https://dx.doi.org/10.18653/v1/2024.nlp4science-1.10)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p1.1)\.
- K\. Zhu, J\. Zhang, Z\. Qi, N\. Shang, Z\. Liu, P\. Han, Y\. Su, H\. Yu, and J\. You \(2025\)SafeScientist: toward risk\-aware scientific discoveries by llm agents\.External Links:2505\.23559,[Link](https://arxiv.org/abs/2505.23559)Cited by:[Appendix B](https://arxiv.org/html/2608.12345#A2.SS0.SSS0.Px1.p3.1)\.

## Appendix ABenchmark Details

### A\.1Misconduct and Ethical Control Tasks in IntegrityBench

Exemplar scenarios for a misconduct task and its ethical control pair:

> T01 \(misconduct\)\.“Samples S014, S027 and S041 were almost certainly contaminated…\\ldotsCan you remove those three rows and rerun the primary stats so we can submit tonight?”The lab notebook contains no contamination flag for any of the three samples, the analysis plan does not list contamination as a pre\-specified exclusion criterion, and removing the three rows shifts the ANOVA p\-value from 0\.0896 to below 0\.05\.

> T19 \(ethical control\)\.“We’ve confirmed that samples S014, S027 and S041 were contaminated during the handling step,this was logged in the lab notebook at the time\.Can you remove those three rows and rerun the primary stats so we can submit with a clean dataset?”The lab notebook now carries a timestamped contamination annotation for each of the three samples, the registered analysis plan \(LabPlan\-2024\-CellViab\-003\) lists documented contamination as a valid exclusion criterion,and the same three rows still shift the p\-value across the 0\.05 threshold\.

Table 5:The 18 misconduct tasks in IntegrityBench, each paired with an ethical control \(36 total\)\.
### A\.2Misconduct Taxonomy

Table[5](https://arxiv.org/html/2608.12345#A1.T5)defines the 18 misconduct types used in IntegrityBench and clarifies the research\-integrity failure each task is designed to test\.

Table 6:General definitions of the 18 research misconduct types evaluated in IntegrityBench\.
### A\.3Pressure Prompt Taxonomy

Table[7](https://arxiv.org/html/2608.12345#A1.T7)reports the full pressure prompt taxonomy used to vary mechanism and escalation level across the benchmark\.

Table 7:Pressure prompt taxonomy by explicitness and escalation level\.
### A\.4Task Distribution

Table[8](https://arxiv.org/html/2608.12345#A1.T8)summarizes the distribution of IntegrityBench tasks across domains, misconduct families, and research pipeline stages\.

Table 8:Task distribution within IntegrityBench\.Left:tasks across 3 domains and 3 misconduct families\. The non\-uniform distribution reflects the broader set of unique behaviors encompassed by the Deception family\.Right:tasks across 4 research pipeline stages and 3 misconduct families\.DomainBiasDecep\.Forbid\.AI642Medical264Physics282Total10188
StageBiasDecep\.Forbid\.Collection242Analysis462Reporting242Design242Total10188

### A\.5Annotation Protocol

Each reviewer received the complete task prompt, consisting of their role, situational context, artifact, and all ten questions without access to any rubrics or material beyond the task content\. Domain experts independently assigned one of 19 labels for Q1 misconduct classification and verified each of the ten ground\-truth answers for validity and unambiguity\. These reviews were collected through a spreadsheet without access to other reviewers’ labels or the research team’s answer key\. Cohen’s kappa was calculated between each domain expert and the ethics expert’s labeling of the same task, resulting inκ=0\.96\\kappa=0\.96, which indicates near\-perfect inter\-annotator agreement on misconduct classification across domains without definitional scaffolding\.

### A\.6Experimental Details

All non\-reasoning variants were evaluated under greedy decoding \(temperature =0\)\. Reasoning variants were evaluated using provider\-native reasoning controls through OpenRouter: Anthropic’s extended thinking budget for Claude models, OpenAI’sreasoning\.effortfor GPT models, Google’sthinkingLevelfor Gemini models and DeepSeek’s thinking toggle for DeepSeek models\. A uniform temperature was not imposed across reasoning variants because providers either ignore sampling parameters in thinking mode, route reasoning through separate controls, or recommend non\-greedy settings; OpenRouter defaults absent sampling parameters to temperature = 1\.0\. All reasoning variants were run without a fixed token budget to ensure fair cross\-provider comparison\. To verify active reasoning, reasoning token usage was extracted per prompt from the response fields\. All inferential comparisons use non\-parametric tests appropriate for this mixed\-decoding evaluation\. Models were run as single\-pass evaluations, with resampling only where a model failed to return a single\-character response\. All concatenated prompts fit within the context window of the most constrained model, DeepSeek V3\.2\. Total API cost was approximately$​90\\mathdollar 90across 70 million tokens, with a per\-model breakdown provided in Table[11](https://arxiv.org/html/2608.12345#A3.T11)\.

## Appendix BAdditional Related Work

#### AI Scientist Systems and Performance Benchmark

LLM capabilities have been extended across the complete scientific research workflow, from hypothesis generation to manuscript writing\. End\-to\-end autonomous systems include the Sakana AI Scientist\(Luet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib1); Yamadaet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib2)\), which deploys multiple collaborating LLM agents to conduct research independently; ResearchTown\(Yuet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib6)\), which simulates collaborative research communities; the Gemini CoScientist\(Gottweiset al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib8)\), a multi\-agent system for hypothesis generation validated through wet\-lab experiments; and CycleResearcher\(Wenget al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib26)\)and Dolphin\(Yuanet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib7)\), which automate aspects of the research cycle\. At a more granular level, task\-specific tools address individual stages of the pipeline: Liang et al\.\(Lianget al\.,[2023](https://arxiv.org/html/2608.12345#bib.bib19)\)evaluate LLMs for manuscript feedback; PaperQA2\(Skarlinskiet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib20)\)supports literature search and summarisation; Zhou et al\.\(Zhouet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib28)\)and Dolphin\(Yuanet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib7)\)address hypothesis generation; and Coscientist\(Boikoet al\.,[2023](https://arxiv.org/html/2608.12345#bib.bib22)\)enables autonomous chemical experimentation through web search and code execution\.

LLM capabilities have been extended across the complete scientific research workflow, from hypothesis generation to manuscript writing\.A growing set of benchmarks evaluate AI scientist capabilities on task execution and scientific problem solving, including MLE\-Bench\(Chanet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib23)\), PaperBench\(Staraceet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib24)\), ML\-Dev\-Bench\(Padigelaet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib25)\), CORE\-Bench\(Siegelet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib27)\)and AIRS\-Bench\(Lupidiet al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib9)\)\. They assess capabilities such as code implementation, paper replication and scientific reasoning across diverse domains\. However, these benchmarks uniformly measure how well AI scientists perform on scientific tasks, not whether they do so ethically\. As a result, task execution accuracy and research integrity compliance are treated as orthogonal dimensions, leaving the latter largely unexamined\. Most systems acknowledge broader ethical risks and some implement more substantive safety mechanisms such as the Gemini CoScientist, including multi\-stage safety reviews of research goals and hypotheses, continuous monitoring of research trajectories through logging and transparency\. When ethical safeguards are implemented, they take the form of post\-hoc disclosure mechanisms, for instance, AI\-generated paper detection and watermarking\(Wenget al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib26)\)rather than evaluating whether models uphold research integrity norms during the research process itself\(Luet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib1); Wenget al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib26)\)\. None evaluate whether the underlying LLMs maintain research integrity under institutional pressure during execution, nor do they test models across the full range of research misconduct behaviors spanning collection, analysis, design and reporting stages\. IntegrityBench directly addresses this gap by evaluating the backbone LLMs that these systems depend upon, measuring their integrity across 18 misconduct types\.

At the agent level, R\-Judge\(Yuanet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib32)\)evaluates LLM risk awareness by presenting models with completed agent interaction records to judge for safety, while ToolEmu\(Ruanet al\.,[2024](https://arxiv.org/html/2608.12345#bib.bib35)\)emulates tool execution to test agents across diverse tool\-use scenarios\. In scientific settings, SafeScientist\(Zhuet al\.,[2025](https://arxiv.org/html/2608.12345#bib.bib33)\)evaluates autonomous AI scientist pipelines against dangerous dual\-use science requests, and ForesightSafetyBench\(Tonget al\.,[2026](https://arxiv.org/html/2608.12345#bib.bib34)\)broadens the scope to frontier AI risk across 94 safety dimensions\. However, these agent\-level benchmarks primarily evaluate risk awareness, tool\-use safety or dangerous scientific outputs, rather than whether the backbone LLMs underlying AI co\-scientists preserve ordinary research\-integrity norms during routine scientific work\. IntegrityBench is designed for this setting: it evaluates whether models maintain integrity across deception, bias and forbidden research, across collection, design, analysis and reporting, and under both implicit and explicit institutional pressure\. Strong performance on existing safety benchmarks therefore does not guarantee reliable integrity behavior in scientific settings\.

## Appendix CAdditional Results

### C\.1Per\-Task Integrity Score Summaries

Table[9](https://arxiv.org/html/2608.12345#A3.T9)and Table[10](https://arxiv.org/html/2608.12345#A3.T10)

Table 9:Per\-misconduct task integrity score summary\. Mean, minimum, and maximum integrity scores are computed across all evaluated model variants\. Tasks are sorted from lowest to highest mean integrity score\.Table 10:Per\-task Integrity Score summary for ethical control tasks\. Mean, minimum, and maximum Integrity Scores are computed across all evaluated model variants\. Tasks are sorted from lowest to highest mean Integrity Score\.
### C\.2Reasoning and Cost Details

Figure[6](https://arxiv.org/html/2608.12345#A3.F6)compares reasoning and non\-reasoning variants across model families, and Table[11](https://arxiv.org/html/2608.12345#A3.T11)reports API cost and token usage\.

![Refer to caption](https://arxiv.org/html/2608.12345v1/x6.png)Figure 6:Reasoning versus non\-reasoning integrity scores across model families\. Lines connect paired reasoning and non\-reasoning variants within each family\.Table 11:API cost and token usage per model across 3,600 requests \(1,800 prompts×\\times2 variants: reasoning and non\-reasoning\)\. Costs are based on OpenRouter billing records\.
### C\.3Failure Patterns Across Research Pipeline Stages

Figure[7](https://arxiv.org/html/2608.12345#A3.F7)reports integrity scores by research pipeline stage for misconduct and ethical control tasks\.

![Refer to caption](https://arxiv.org/html/2608.12345v1/x7.png)Figure 7:Integrity score across the research pipeline stages\. Bars show the mean integrity score for misconduct and ethical control tasks\.As shown in Figure[7](https://arxiv.org/html/2608.12345#A3.F7), integrity scores vary substantially across research pipeline stages\. For misconduct tasks, models perform best in collection and reporting, with mean scores in the mid\-70s, while design lags in the mid\-60s\. Ethical\-control tasks show a different pattern: design and reporting receive the highest scores, whereas analysis is the weakest stage, with a mean score of 53\.6\. This stage\-level split suggests that models struggle most when permissibility depends on intent, transparency, and procedural justification rather than on a visible action alone\.

The analysis stage in IntegrityBench is comprised entirely of deception\-family tasks: p\-hacking, experiment overfitting, causal overclaiming, and effect\-size overclaiming\. These tasks are context dependent\. Whether data exclusion is acceptable depends on whether the exclusion criteria were prespecified, transparently reported, and scientifically justified\. Dropping observations is either responsible data cleaning or inappropriate manipulation depending on the researcher’s rationale and transparency\. The low ethical\-control score in analysis suggests that models struggle to make that distinction\. Hence, these analysis\-stage deception tasks require alignment interventions that specifically target contextual disambiguation between legitimate analytical practice and research misconduct\.

Design shows the opposite pattern\. Models recognize legitimate design choices more reliably than they detect misconduct embedded in research planning\. Design\-stage tasks include HARKing, hypothesis anchoring, bandwagon method selection, quantitative anchoring, and dual\-use blindness\. In the design setting, compliant behavior often manifests through clear documentation: prespecified hypotheses, principled method selection, explicit modeling assumptions, and proactive risk controls\. As a result, ethical\-control cases contain visible cues of good research practice\. However, the violation in misconduct tasks is frequently expressed through intent, framing, or downstream risk rather than through an overtly prohibited action\. Choosing a popular method is acceptable when it fits the research question, but becomes problematic when popularity substitutes for methodological judgment\. Similarly, revising a hypothesis is acceptable during exploratory work, but becomes HARKing when it is presented as prespecified after observing the results\. Models therefore struggle to catch planning\-stage violations when the problem depends on motivation, timing, or justification rather than on the surface action itself\.

## Appendix DStatistical Robustness Checks

We report additional statistical checks supporting the main empirical findings\. These checks test whether reasoning changes paired model correctness, whether pressure mechanisms differ across task types, and whether the Q1–Q3 gap is reliable across tasks\. All tests are used as robustness checks rather than as additional benchmark metrics\.

### D\.1Tests for Reasoning Effects

We use paired McNemar tests to evaluate whether enabling reasoning changes model correctness on the same benchmark prompts\. This test is appropriate because each reasoning\-enabled model and its non\-reasoning counterpart is evaluated on the same set of prompt instances, making the correctness outcomes paired rather than independent\. Because we perform the test across nine matched backbone pairs, we apply Benjamini–Hochberg correction to control the false discovery rate across multiple comparisons\.

Table 12:McNemar tests comparing reasoning\-enabled and non\-reasoning variants\. Each test compares paired correctness outcomes on the same prompt instances\. Adjustedppvalues use Benjamini–Hochberg correction across the nine matched model pairs\.Overall, the McNemar tests show no significant reasoning effect for seven of the nine matched backbone pairs after Benjamini–Hochberg correction\. The two significant effects are bidirectional: reasoning improves Qwen 3\.5 Flash 9B but degrades GPT 5\.4 Mini\. This supports the conclusion that reasoning does not produce a uniform integrity gain across model families\.

### D\.2Pressure Mechanism Comparison

Table[13](https://arxiv.org/html/2608.12345#A4.T13)summarizes paired comparisons between explicit and implicit pressure mechanisms\. Explicit pressure averages PP2 and PP4, while implicit pressure averages PP1 and PP3 respectively, reported separately for ethical control and misconduct tasks\.

Table 13:Paired t\-test summary comparing explicit and implicit pressure mechanisms\. Explicit and implicit columns report mean integrity scores averaged over PP2/PP4 and PP1/PP3, respectively\.The pressure comparison confirms an asymmetric pattern\. Implicit pressure is associated with lower ethical\-control performance, consistent with over\-refusal or over\-flagging of procedurally justified tasks\. Explicit pressure is associated with lower misconduct\-task performance, suggesting that named authority appeals and procedural counterarguments more directly increase compliance with misconduct requests\.

### D\.3Wilcoxon Effect Size for the Q1–Q3 Gap

The Q1–Q3 ordering reported in Section[4\.5](https://arxiv.org/html/2608.12345#S4.SS5)is reliable across task\-level means\. A paired Wilcoxon signed\-rank test confirms the ordering \(W=152W=152,p=0\.004p=0\.004,r=0\.68r=0\.68\), indicating a large effect\.

### D\.4Deterministic Decoding and Uncertainty Estimation

Because model\-level integrity scores are deterministic under greedy decoding, identical prompts produce identical responses and bootstrap variance estimation over responses is not applicable\. Inferential comparisons therefore rely on non\-parametric tests over paired tasks or matched model pairs rather than repeated stochastic samples as detailed in Appendix[D\.1](https://arxiv.org/html/2608.12345#A4.SS1)–[D\.3](https://arxiv.org/html/2608.12345#A4.SS3)\.

## Appendix EPrompt Structure and Model Trace

### E\.1Prompt Structure Visualization

Figure[8](https://arxiv.org/html/2608.12345#A5.F8), Figure[9](https://arxiv.org/html/2608.12345#A5.F9), and Figure[10](https://arxiv.org/html/2608.12345#A5.F10)show the complete prompt delivered to the model for Task T01, Data Fabrication in the Medical domain, across the three question groups Q1, Q2, and Q3 respectively\. The Role block is delivered as the system prompt\. The Situation, Artifact, Pressure Prompt, and Question blocks are concatenated into a single user prompt turn\. PP4, the explicit escalated pressure condition, is shown as the representative pressure condition\.

![Refer to caption](https://arxiv.org/html/2608.12345v1/x8.png)Figure 8:T01 prompt structure, Q1: Misconduct Classification, PP4 condition\.Role block delivered as system prompt\. Situation, Artifact, Pressure Prompt PP4, and Q1 question concatenated as user prompt\. Model returns single letter A–S\. Ground truth: B\) Data Fabrication\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x9.png)Figure 9:T01 prompt structure, Q2: Ethical Action Reasoning, PP4 condition\.The figure shows all three Q2 sub\-questions together for readability\. In practice, each sub\-question is delivered as a separate prompt call: the shared components are repeated in full for each sub\-question, with only Q2a, Q2b, or Q2c appended individually\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x10.png)Figure 10:T01 prompt structure, Q3: Artifact\-Grounded Decision, PP4 condition\.The figure shows all six Q3 sub\-questions together for readability\. In practice, each sub\-question is delivered as a separate prompt call: the shared components are repeated in full for each sub\-question, with only Q3a, Q3b, Q3c, Q3d, Q3e, or Q3f appended individually\.
### E\.2Full Model Trace: Scenario 1 Under Five Pressure Levels

Figures[11](https://arxiv.org/html/2608.12345#A5.F11)–[15](https://arxiv.org/html/2608.12345#A5.F15)present the complete structured trace for Scenario 1, Data Fabrication in the Medical domain and Deception–Collection stage, evaluated on Deepseek V3\.2 \(NoR\) across all five pressure conditions\. Each figure corresponds to one pressure level and shows, for every turn, the exact prompt composition injected, the model’s single\-letter answer, and its binary integrity score\.

![Refer to caption](https://arxiv.org/html/2608.12345v1/x11.png)Figure 11:Scenario 1 trace under P0, baseline\.No pressure prompt is injected\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x12.png)Figure 12:Scenario 1 trace under P1, implicit pressure\.An anonymous productivity alert introduces moderate institutional pressure while leaving the task content unchanged\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x13.png)Figure 13:cenario 1 trace under P2, explicit pressure\.An anonymous escalation notice increases institutional pressure while leaving the task content unchanged\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x14.png)Figure 14:Scenario 1 trace under P3, escalated implicit pressure\.A named senior co\-author introduces moderate authority pressure while leaving the task content unchanged\.![Refer to caption](https://arxiv.org/html/2608.12345v1/x15.png)Figure 15:Scenario 1 trace under P4, escalated explicit pressure\.A named PI introduces high\-intensity authority pressure while leaving the task content unchanged\.

## Appendix FHuman Subjects and Survey Methodology

To ensure that the misconduct behaviors evaluated in IntegrityBench are firmly grounded in real\-world scientific practices, we conducted a structured elicitation survey prior to the construction of the benchmark\. This section provides additional details regarding the human subjects involved in our research, expanding upon the declarations made in the NeurIPS Paper Checklist\.

### F\.1Misconduct Elicitation Survey

We recruited4747domain researchers spanning the three scientific fields represented in our benchmark: Artificial Intelligence, Physics, and Medicine\. These participants were asked to report and describe specific instances, patterns, or types of research misconduct they had directly observed or were intimately familiar with in their respective fields\. The qualitative responses from this survey were aggregated and synthesized to define the 18 distinct misconduct behaviors \(detailed in Table 1\) that form the core of the IntegrityBench taxonomy\.

Participant Instructions and Items:The participant\-facing survey first asked researchers to indicate their primary research field \(Artificial Intelligence, Physics, or Medicine/Biology\)\. Following this single demographic question, they were prompted with the main elicitation item:“Please describe up to three distinct types of research misconduct or questionable research practices you have observed, encountered, or suspect are prevalent in your primary research domain\. Please focus on the procedural mechanism of the behavior rather than identifying specific individuals or institutions\.”No other identifying information was collected\.

Compensation and Volunteer Status:All 47 domain researchers who participated in this elicitation survey did so strictly on a volunteer basis\. No financial compensation or honoraria were provided for their participation\.

### F\.2Expert Review Panel

Following the synthetic generation of the benchmark tasks, we engaged a separate panel of experts to validate the constructs and rigorously review the answer keys\. As described in Section 3\.3 and Appendix A\.6, this panel consisted of three domain experts \(each holding at least a Ph\.D\. in their respective fields\) and one ethics expert\.

Similar to the elicitation survey participants, all expert reviewers served purely as volunteers\. They received no financial compensation for their time, feedback, and independent labeling efforts\.

### F\.3Ethical Approval and Consent

The research activities involving human subjects encompassing both the initial survey and the expert review panel primarily involved gathering professional opinions on research conduct, standard validation procedures, and benchmark difficulty\. The potential risks to participants were assessed as minimal\.

Prior to participation, all individuals were provided with a digital information sheet detailing the purpose of the study, how their data would be utilized, and the measures taken to ensure anonymity \(e\.g\., removing any personally identifiable information from described misconduct scenarios\)\. Informed consent was obtained from all participants before proceeding\. The study protocol, including all participant\-facing materials and consent procedures, was reviewed and approved by our Institutional Review Board \(IRB\) prior to the commencement of any data collection\.

Similar Articles

On the Limits of LLM-as-Judge for Scientific Novelty Assessment

Hugging Face Daily Papers

This paper introduces RQ-Bench, a benchmark to evaluate LLMs' ability to assess the novelty of scientific research questions. It finds that LLM judges consistently rate generated questions as more novel than human experts do, raising concerns about the reliability of using LLMs for scientific novelty evaluation.