Position: Let's Strengthen Verifiability If We Can't Enforce Reproducibility
Summary
This position paper analyzes the reproducibility crisis in machine learning—quantifying unavailable code and unreproducible results—and proposes concrete measures to strengthen result checkability and verifiability in ML research.
View Cached Full Text
Cached at: 09/30/26, 09:38 AM
# Let’s Strengthen Verifiability If We Can’t Enforce Reproducibility
Source: [https://arxiv.org/html/2609.35854](https://arxiv.org/html/2609.35854)
\\DeclareCaptionType
\[within=none\]promptbox\[Prompt\]\[List of prompts\]
Nermin SametAffiliation:Valeo\.aiAffiliation:Paris, FranceRenaud MarletAffiliation:Valeo\.aiAffiliation:Paris, France
###### Abstract
In the field of Machine Learning, many papers contain empirical results supporting claimed statements or illustrating the performance of a proposed method\. However, most practitioners know that \(1\) results are generally hard to reproduce, and increasingly so, \(2\) code is not often available to do so, and \(3\) it hinders the development of research\. In this position paper, we analyze and quantify these issues, and make concrete proposals to improve result checkability, if not reproducibility\. Code and supporting materials are available at https://github\.com/giddyyupp/position\-enforce\-verifiability\.
## 1Introduction
Reproducibility is the cornerstone of scientific progress\. This foundational statement\[[1](https://arxiv.org/html/2609.35854#bib.bib1)\], likely as old as epistemology, with figures as Plato and Aristotle, has been recalled many times in the last decades\[[2](https://arxiv.org/html/2609.35854#bib.bib2),[3](https://arxiv.org/html/2609.35854#bib.bib3)\], especially since the replication crisis gained momentum about ten years ago\[[4](https://arxiv.org/html/2609.35854#bib.bib4),[5](https://arxiv.org/html/2609.35854#bib.bib5),[6](https://arxiv.org/html/2609.35854#bib.bib6),[7](https://arxiv.org/html/2609.35854#bib.bib7)\]\.
The fact is that the figures on failures to reproduce an experiment, whether one’s own or somebody else’s, are compelling: according to a study involving about 1,600 researchers in various fields of science\[[5](https://arxiv.org/html/2609.35854#bib.bib5)\], more than 70% of them have tried and failed to reproduce another scientist’s experiments, and more than half have failed to reproduce one of their own experiments\. In contrast, a recent study on AI/ML publications\[[8](https://arxiv.org/html/2609.35854#bib.bib8)\]found a low 0\.05% retraction rate, with papers that often continue to be cited, sometimes receiving up to 8 times more citations after retraction\.
Several factors amplify this crisis: each year there are more researchers, more papers per researcher, and papers are more complex, harder to write and review\. Studies show indeed that, between 2014 and 2018, the number of researchers worldwide \(approaching 9 M\) has increased by 3\.3% per year, which is 3 times faster than the global population\[[9](https://arxiv.org/html/2609.35854#bib.bib9)\]\. Also, from 2016 to 2022, the number of scientific publications worldwide \(currently over 3 M/year\) has grown by 5\.6% per year\[[10](https://arxiv.org/html/2609.35854#bib.bib10)\]\. Because of the*burden of knowledge*\[[11](https://arxiv.org/html/2609.35854#bib.bib11)\], as papers require more expertise to write and review, reproducibility may also be hindered\. Our experience is that it is increasingly difficult to publish a simple method, even if it outperforms the state of the art \(SOTA\), because then of an alleged “lack of technical novelty”; this encourages authors to propose artificially complex methods, which are then also harder to reproduce\.
Last, Generative AI \(GenAI\) now automates the production of scientific contents, with possible errors or misuses\. There are numerous reports of dishonest papers, partly or totally generated by AI, e\.g\., with hallucinated references\[[12](https://arxiv.org/html/2609.35854#bib.bib12)\], copycats of genuine papers\[[13](https://arxiv.org/html/2609.35854#bib.bib13)\], or paper mills\[[14](https://arxiv.org/html/2609.35854#bib.bib14)\]with fake papers to boost citations\[[15](https://arxiv.org/html/2609.35854#bib.bib15)\], pushing arXiv to now require first\-time posters to be endorsed\[[16](https://arxiv.org/html/2609.35854#bib.bib16)\]\.
Compared to social, life, and physical sciences, computer science is fortunate to have straightforward assets to ensure reproducibility: code and data\. Putting aside variations due to stochastic algorithms, fixed\-precision computation, and residual randomness in modern hardware, which however are prevalent in Machine Learning \(ML\), computer science is assumed to be an exact science involving deterministic machines\. In theory, reproduction can be confirmed by a simple keystroke\. In practice, there are many barriers, and the reproducibility crisis affects computer science too, including ML\[[7](https://arxiv.org/html/2609.35854#bib.bib7)\]\.
In this position paper, we study code availability \([Section3](https://arxiv.org/html/2609.35854#S3)\) and its status regarding replication \([Section4](https://arxiv.org/html/2609.35854#S4)\)\. Considering top\-tier ML and Computer Vision \(CV\) conferences, we observe a recent decrease in the ratio of accepted papers with accessible code, though papers with code have twice as many citations on average\.Considering that reproducibility and code availability will remain hard to generalize, our position is that the community should at least improve the verifiability of experiments\.We therefore propose two verifications \([Section5](https://arxiv.org/html/2609.35854#S5)\) and an associated publication process \([Section6](https://arxiv.org/html/2609.35854#S6)\) that can contribute to improving the confidence in reported experiments\.
## 2Related analyses and proposals
A broad consensus has emerged regardingreproducibilityin ML\[[17](https://arxiv.org/html/2609.35854#bib.bib17),[18](https://arxiv.org/html/2609.35854#bib.bib18)\]\. Following ACM terminology,reproducibilityasks: “Can an independent team obtain the reported results using the code and data provided by the original authors?” Under this definition,reproducibilityassumes \(i\) public access to the relevant artifacts, and \(ii\) third\-party verification of the reported outcomes\. While access to code and data is necessary, some argue that this artifact release alone does not guaranteereproducibility\[[19](https://arxiv.org/html/2609.35854#bib.bib19),[20](https://arxiv.org/html/2609.35854#bib.bib20)\]\. The more prominent obstacle however remains that code or data are often not released, motivating efforts to identify what makes a paper reproducible beyond code availability\.
[Raff \[21\]](https://arxiv.org/html/2609.35854#bib.bib21)provides a concrete illustration: he attempted to reimplement 255 papers without consulting official code, even when available\. Only 63\.5% of the papers could be replicated under the reproducibility criterion, requiring that at least 75% of the claims be validated\. A key finding is that readability, meaning clear implementation details and informative pseudocode, is the strongest predictor ofreproducibility\. Furthermore, when the authors of the original paper help, thereproducibilityincreases to 85%\. More recently, Paper2Code\[[22](https://arxiv.org/html/2609.35854#bib.bib22)\]explores whetherreproducibilitycan be automated by generating code from the paper alone\. This LLM\-based multi\-agent system achieves an overall replication rate of roughly 45% on a public benchmark\[[23](https://arxiv.org/html/2609.35854#bib.bib23)\]\.
Several other works\[[24](https://arxiv.org/html/2609.35854#bib.bib24),[25](https://arxiv.org/html/2609.35854#bib.bib25),[26](https://arxiv.org/html/2609.35854#bib.bib26),[27](https://arxiv.org/html/2609.35854#bib.bib27)\]investigated the prominent issues hinderingreproducibilityin broader ML\-based methods across various fields\.
This ever growingreproducibilityconcern in different domains urged the community to take action\. NeurIPS 2019 has introduced a “ReproducibilityChallenge” and an “ML reproducibility checklist” to encourage and improve thereproducibilityof accepted papers\[[17](https://arxiv.org/html/2609.35854#bib.bib17)\]\. The former aims to provideindependentverification and validation of the empirical claims in accepted papers\. \(Only 173 papers were submitted for this challenge out of 1,427 accepted papers\.\) The latter involves responding to a questionnaire assessing whether the paper includes various essential information to ensurereproducibility; it has to be completed during the initial submission phase\.
The Journal of Artificial Intelligence Research \(JAIR\)\[[28](https://arxiv.org/html/2609.35854#bib.bib28)\]introduces four novel mechanisms to make AI research more reproducible: \(1\) “reproducibilitychecklists”, similar to the one mentioned above, \(2\) “structured abstracts” to present key components of the paper in a more structured way, \(3\) “reproducibilitybadges” to incentivize transparency and reproducibility by making data, code, both, or independent reproductions available, and finally \(4\) “reproducibilityreports” to encourage independentreproducibilityof JAIR papers\. More specifically, JAIR created an additional track for papers that present a public implementation and validation of a previously accepted JAIR paper\.
EuroSys\[[27](https://arxiv.org/html/2609.35854#bib.bib27)\]proposes various short\- and long\-term directions to overcome thereproducibilitychallenges after carefully analyzing artifact evaluation processes in previous editions\. These recommendations include artifact submission at various stages of the reviewing process\.
## 3Code availability study
Software \(code and data\) is central to reproducibility for computer science papers\. In this section, we first consider general issues related to software availability \([Section3\.1](https://arxiv.org/html/2609.35854#S3.SS1)\), then study access to code of published papers in five top\-tier ML and CV conferences over the last five years \([Section3\.2](https://arxiv.org/html/2609.35854#S3.SS2)\)\.
### 3\.1Issues with software availability
#### Various degrees of code accessibility\.
The specificity of ML is that there are two computing stages: model learning form training data, and model inference, i\.e\., testing\. It results in three common cases of code availability: \(1\) no code available; \(2\) trained model available, i\.e\., inference code and learned parameters \(e\.g\., network weights\) but not training code; \(3\) both training and inference code available, possibly with some learned parameters\. Still, it is not uncommon that, due to partial code availability, only a fraction of the experiments can be reproduced, e\.g\., not on all datasets, not with test\-time augmentations, or not with ensembling\.
20212022202320242025303040405050Ratio \(%\)\(a\) Proportion of accepted papers with train or test code20212022202320242025101020203030Ratio \(%\)\(b\) Proportion of accepted papers with available model weights2021202220232024202510102020YearRatio \(%\)\(c\) Proportion of accepted papers acknowledging other repositoriesCVPRICCVICLRICMLNeurIPSAll
Figure 1:Proportion of accepted papers at CVPR, ICCV, ICLR, ICML and NeurIPS 2021–2025 that publicly \(a\) share train or test code \[reducing\], \(b\) provide pretrained model weights \[reducing\], or \(c\) acknowledge previous repositories \[growing\]\.
#### Common reasons why software is unavailable\.
In our experience, there are three main reasons explaining \(rather than justifying\) why some code or data is not available\.
- •*No time to polish*is not a legit argument\. If the software is not clean, chances are that the results are not either, hampering reliability\.
- •*The developer left*is not a valid argument either\. Software polishing should be an integral part of the developer’s work, including if s/he is an intern bound to leave after a few months\.
- •The company does not want to disclose software, for legal or security reasons, or to make it more difficult for competitors to reproduce the work\. The scientific implication is debatable: either the paper contains enough information to be reimplemented and the company only gains a little time over competitors, or it does not and the paper has little scientific value as it is not reproducible\. A usually accepted compromise is to only provide learned parameters and inference code, possibly with a binding license\. In contrast, academic papers are expected to give full access to artifacts for the sake of open science\[[27](https://arxiv.org/html/2609.35854#bib.bib27)\]\.
### 3\.2Code availability in ML/CV conferences
To quantify code availability, we conducted a statistical study on the accepted papers of recent \(2021\-2025\) top\-tier conferences in ML/CV, including CVPR, ICCV, ICLR, ICML, and NeurIPS\. The main collected data are represented in[Figures1](https://arxiv.org/html/2609.35854#S3.F1),[2](https://arxiv.org/html/2609.35854#S3.F2)and[3](https://arxiv.org/html/2609.35854#S3.F3)\. More details are available in App\.[A](https://arxiv.org/html/2609.35854#A1)\.
202120222023202420250010102020Topic share \(%\)\(a\) Evolution of topics in years20212022202320242025YearCode availability \(%\)\(b\) Code availability ratio in each topic3D Vision and Neur\. Render\.Image and Video Synt\.Multimodal LearningLLMs and Found\. ModelsGen\. AI and Diff\. ModelsReinforcement LearningTrust\., Fair and Robust MLOptim\. and Theory of Learn\.Scal\., Syst\. & Distrib\. MLApplication\-Driven ML
Figure 2:Proportion of papers per topic in accepted papers to CVPR, ICCV, ICLR, ICML and NeurIPS 2021–2025 \(left\), and code availability in each topic \(right\)\. Except for a few topics, the general trend shows a proportional decrease in open\-sourcing official codes\.#### Code sharing has recently started to decrease\.
After several years of noticeable increases in the proportion of accepted papers with \(train or test\) code publicly available, a marked decline is observed in 2025 for all conferences \([Figure1](https://arxiv.org/html/2609.35854#S3.F1)\(a\)\)\. The deterioration even dates back to 2023 for CVPR and ICLR\. Besides, this drop occurs while the proportion of papers with topics calling for empirical validation is significantly rising, as opposed, for instance, to more theoretical topics \([Figure2](https://arxiv.org/html/2609.35854#S3.F2)\)\. A similar decline is visible for accepted papers with model weights, although it is not as marked \([Figure1](https://arxiv.org/html/2609.35854#S3.F1)\(b\)\)\.
Additionally, the 20% most acknowledged repositories in our study \([Figure1](https://arxiv.org/html/2609.35854#S3.F1)\(c\)\) account for 81\.5% of all acknowledgements, following the Pareto principle\. Looking at the most acknowledged repos, it appears that a significant part of the recent progress in empirical machine learning originates from foundation models that make weights available, although not always training code and data\.
2025202420232022202100100100200200YearCitations per paperPapers with code \(median\)Papers without code \(median\)Papers with code \(mean\)Papers without code \(mean\)
Figure 3:Mean & median citation counts for publications at CVPR, ICCV, ICLR, ICML and NeurIPS \(2021–2025\), grouped by code availability\.
#### Papers with code are cited twice as often\.
We also collected citation counts from[Semantic Scholar \[29\]](https://arxiv.org/html/2609.35854#bib.bib29)for all the accepted papers and calculated the mean and median citation values for the papers with and without public code \([Figure3](https://arxiv.org/html/2609.35854#S3.F3)\)\. We observe that papers with official public code implementations have more than twice as many citations on average as those without any code\. While we understand that the community is not ready to require code availability for publications, and that our analysis is correlational rather than causal, this observation should nevertheless encourage authors to share their code\.
#### Code reuse stagnates but still slightly boosts research\.
Finally, we counted the number of repositories of accepted papers that acknowledge other repositories \([Figure1](https://arxiv.org/html/2609.35854#S3.F1)\(c\)\)\. We consider that it typically corresponds to implementations that borrow or modify some existing code to build their own\. We observe that the number of acknowledging repositories has not changed significantly over the past three years, although there are differences between conferences that balance each other out\. Relative to the drop in accepted papers with code \([Figure1](https://arxiv.org/html/2609.35854#S3.F1)\(a\)\), this stagnation means that public implementations have nevertheless slightly increased their capacity to boost research\.
## 4Code availability does not ensure paper reproducibility
Code is thus not always available\. But even if so, it does not ensure reproducibility\. In this section, we study inconsistencies between code and paper \([Section4\.1](https://arxiv.org/html/2609.35854#S4.SS1)\), and possible execution issues \([Section4\.2](https://arxiv.org/html/2609.35854#S4.SS2)\)\.
### 4\.1Discrepancies between paper and code
Getting a hold on the code is not just about obtaining the same numbers as those in the paper\. It is also about making sure the code does what the paper says\. It is the paper that conveys the ideas and that provides a rationale to evaluate them; the code serves as support, and it must be a faithful one\.
Yet, due to page limitations, some implementation details are commonly found only in the code, rather than in the paper\. As newly proposed architectures are increasingly complex, reimplementing a paper has thus become a real challenge\. Besides, it is not uncommon to find differences between the paper description and what the code actually does\. It impairs reproducibility, making it difficult for reimplementations to match the original results\. It misleads the reader and thus also impairs progress\.
Conversely, some processing can be missing from the code that is made public, although related results are reported in the paper\. It can be the case, e\.g\., of test\-time augmentations and ensembling, which are sometimes used to boost a metric to reach the SOTA on a specific benchmark, while being barely mentioned in the paper\. In our experience, e\.g\., for semantic segmentation, such precious processing tricks can remain undisclosed and thus unreproducible\. It boosts the performance on a benchmark with server and hidden ground truth, while the basic version of the method, showcased in the paper on another dataset with public ground truth, is on the contrary fully reproducible\.
There may also be methodological mistakes in the code, which are not visible in the paper but that could invalidate the reported experiment\. For instance, a model can be trained by introducing more data or information than intended and indicated\. Concrete bad practices include to pretrain unsupervisedly on test data, or to train on test data using a semi\-supervised approach\.
### 4\.2Issues when running code
Even if the code is sensible and consistent with the paper, some issues may occur at execution time\.
Execution is often stochastic: during training \(model initialization, batch randomness, asynchronous operations…\) but also at test time \(e\.g\., GenAI\)\. Yet, not all authors present metrics averaged over several runs, with an indication of variance\. This may lead to reproduction differences, which are however often accepted if the variations stay moderate, depending on the community, task and dataset\.
There are also bad practices related to training monitoring\. A classical one consists of peeking at the test set, quantitatively if the ground truth is known, or qualitatively otherwise\. Training can then be stopped early to prevent overfitting, as often done, e\.g\., on target sets for supposedly unsupervised domain adaptation, instead of using validators\[[30](https://arxiv.org/html/2609.35854#bib.bib30)\]\. Another bad practice is cherry\-picking a best checkpoint among several training attempts, possibly explaining results that nobody succeeds in reproducing\. Besides, when only model parameters are given, it is hardly possible to know what learning scheme and actual data were used for training\. As for inference, difficult samples can be excluded, as we already witnessed\. Test\-time augmentation or ensembling can be used but unreported\.
Last, a practical issue, besides installation problems, is that the \(train or test\) code might be too long to run, or require specific hardware \(e\.g\., lots of high\-end GPUs\)\.
## 5Verifiable indicators of reproducibility
To err is human\. But in the era of \(M\)LLMs, ensuring that papers present genuine experiment outputs rather than erroneous or fake results is more critical than ever\. As code, even if accessible, is in any case difficult to assess, we propose instead to focus on the verification of easily\-available indicators\.
The problem is that any execution that can only be performed by the authors is a source of vulnerability regarding reproducibility and verifiability\. The most reliable way to enable the validation of an experimental result remains to open\-source the training and inference code, the data, and the trained models\. However, we recognize that open\-sourcing is sometimes impossible due to privacy constraints, licensing, commercial restrictions, or security concerns\. To maintain a healthy, trustworthy research environment and prevent bad research practices, alternative verification methods are required\.
In this section, we describe two verifications that can contribute to improving confidence in reported experiments: checking experiment logs \([Section5\.1](https://arxiv.org/html/2609.35854#S5.SS1)\) and checking metric evaluations \([Section5\.2](https://arxiv.org/html/2609.35854#S5.SS2)\)\. The complete associated workflow is presented in[Section6](https://arxiv.org/html/2609.35854#S6)\. We also discuss code release \([Section5\.3](https://arxiv.org/html/2609.35854#S5.SS3)\)\.
### 5\.1Verification of experiment logs
#### Log verification material\.
Most, if not all experimental environments used in the ML community support the generation of logs for monitoring the experiments, including training and inference\. Such environments include TensorBoard\[[31](https://arxiv.org/html/2609.35854#bib.bib31)\], MLflow\[[32](https://arxiv.org/html/2609.35854#bib.bib32)\], and Weights & Biases \(W&B\)\[[33](https://arxiv.org/html/2609.35854#bib.bib33)\]\. Logs to provide for verification purpose concern each experiment related to empirical claims in the paper, in particular state\-of\-the art \(SOTA\) claims, including via tables and graphs\. Ablation experiment logs are optional\. Provided log data should include a list of basic information on the experiment, as described in App\.[B\.1](https://arxiv.org/html/2609.35854#A2.SS1), and learning curves of trained models \(examplified also in App\.[B\.1](https://arxiv.org/html/2609.35854#A2.SS1)\)\.
#### Providing log verification material\.
Experiment logs are to be provided by the authors at paper submission time, in the supplementary material\. To ease verification, authors should also include any relevant plots, exported from the logs\. Contrary to the code, experiment logs hardly raise any issues regarding privacy, licensing or security\. Their disclosure should thus be easily accepted\. Since logging useful training and inference metrics with free tools is already common practice for model analysis, the additional burden on authors is minimal, e\.g\., similar to filling the NeurIPS checklist\.
#### Checking log verification material\.
The consistency of the experiment logs with the paper is checked by the reviewers during the review period\. If the conference features a rebuttal, the reviewer may ask the authors for clarification or additional log information\. The report on log consistency is part of the review\. \(More details are in[Section6](https://arxiv.org/html/2609.35854#S6)\.\) While CVPR 2026 has introduced optional W&B log uploads, the focus is on compute consumption\. Our proposal is to use it for experiment verification\.
#### Goal and limitations of log verification\.
Log\-based verification ensures that reported numbers are backed by real training and evaluation traces\. But it does not fully prevent data leakage, dataset manipulation, or “lookup cheating”, where predictions are recovered using sample identifiers\.
### 5\.2Verification of metric computation
#### Metric verification material\.
Most experiments in ML try to quantify the quality of predictions, including stochastic generations\. It is performed by running some code that computes metrics over result files, typically by measuring a form of distance to some ground truth, including distribution models\. To make sure no error is introduced when computing these metrics and when reporting them, each experiment related to a claim in the paper should come withresults files,evaluation code, and possibly ground truth information\. \(See details in App\.[B\.2](https://arxiv.org/html/2609.35854#A2.SS2)\.\) No license agreement should be needed to access proprietary data because it could break the anonymity of the authors or of the reviewer\.
#### Providing metric verification material\.
The metric verification material is to be provided by authors at paper submission time, in the supplementary material or via links to anonymized repositories\. Even though there may be cases where the output of a model is sensitive and cannot be disclosed, results most often concern publicly available datasets and do not have disclosure issues\. Reviewers should be able to inspect the evaluation material and rerun it easily, avoiding installation issues\. We therefore recommend that authors actually provide an anonymized notebook, as available in free shared execution environment such as Google Colab, Kaggle Notebooks, Paperspace Gradient, or AWS SageMaker Studio Lab\. We provide two examples of such notebooks111https://github\.com/giddyyupp/position\-enforce\-verifiability: a short script using only public libraries and a longer script containing a custom metric implementation\. Preparing such metric evaluation code and data for review introduces only minimal overhead for authors, as it mostly involves reusing code they already have written for their own evaluation purpose\.
Table 1:Provided indicators of reproducibility\.
#### Checking metric verification material\.
To reduce the overall workload and depending on the venue policy, one or several of the reviewers are to be appointed as*metric evaluators*for the same paper\. The responsibility of a metric evaluator is to scrutinize the metric evaluation code, to run it on provided result files, to check metric consistency with the paper, and to write a comment about it as part of the paper review\. We conducted a small user study with five experienced reviewers to measure the average time required to apply the verification instructions\. The short script required approximately 5 minutes on average, while the longer script required between 45 minutes and 1 hour\. Overall, this workload is substantially lighter than reviewing a full paper and providing detailed feedback\. We consider this an acceptable cost to gain trust in empirical results\.
#### Goal and limitations of metric verification\.
This verification ensures that the metrics reported in the paper are identical \(up to acceptable stochastic variations\) to those obtained by an independent evaluator\. While the result files may be forged, this verification, and the care required for authors to prepare it, should help reduce mistakes made when quantifying a result and reporting it\.
### 5\.3Verification of code release
After acceptance, authors may provide code, typically by inserting in their paper a link to a repository\. There is currently no verification that such links actually point to sensible code\. Still, reviewers sometimes treat code promises positively, even though such commitments can be fragile\. \(It is easy to find GitHub repositories of papers from our five target conferences, that contain the \(in\)famous “code coming soon” and that are at least one year old, if not much more\.\)
To promote timely code release, we propose that, by the camera\-ready deadline, the authors provide a software link if they wish, and that, by the time the conference opens, an assigned reviewer checks the repository to verify it does contain the expected software\. As the goal is not to scrutinize the software but just to check that it exists, the task is very lightweight\. It can actually by largely automated, as we did in[Section3\.2](https://arxiv.org/html/2609.35854#S3.SS2), with a possible human control for failing repos\. This official software link, if any and if confirmed by the reviewer, would appear on the proceedings web site, as done, e\.g\., for ICML in[PMLR \[34\]](https://arxiv.org/html/2609.35854#bib.bib34)\. Although, software completeness will not be checked, we believe it would nonetheless push authors not to delay the code release, and operate as an incentive to give value to their paper\.
Submission PackageMain PaperSupplementaryTraining and Inference LogsOutputJSONEvaluationScriptReview ProcessReview Consensus onResult ValidationMetric Validator ConfirmationACValidationYESYESIf Accepted,Update OpenReviewYESYES
Figure 4:Overview of the proposed verification workflow\. At submission time,training and inference logsare attached to the submission package and verified by one or several reviewers \(Level 1\)\. The metric verification step \(Level 2\), indicated with red background, requires the submission of JSON file\(s\) that contain model predictions on target dataset\(s\), along with an evaluation script\. The assigned metric validator runs the evaluation script and compares with the numbers in the main paper\. Finally, the AC updates the corresponding statuses \([Table1](https://arxiv.org/html/2609.35854#S5.T1)\) in the submission site\.
## 6Call to action: a reinforced review process
To improve trust in accepted papers, we propose a verification workflow \([Figure4](https://arxiv.org/html/2609.35854#S5.F4)\) that is realistic, implementable immediately, and low\-overhead for conferences, authors and reviewers\.
When a paper is submitted:
- P1\.The authors mustmake available the experimental material \(logs, results and evaluation code,[Section5](https://arxiv.org/html/2609.35854#S5)\) of their main experiments, i\.e\., related to claimed contributions\. It can be given via anonymized links or provided in supplementary material on the submission web site\.
When a paper is reviewed and discussed:
- P2\.The reviewers mustcomment on the consistency of the providedexperimental materialw\.r\.t\. the empirical results presented in the submitted paper\. The reviewers must also give a formal rating on this experimental consistency, as defined in[Table1](https://arxiv.org/html/2609.35854#S5.T1)\. This is an integral part of the review, which contributes to the final recommendation\. To lighten the overall workload, a venue policy can be to assign a singleexperiment reviewerper paper\.
- P3\.The area chairs mustsummarize the assessment of the experimental consistency and provide a rationale for a final consistency status\. It is an integral part of the meta\-review\.
- P4\.The reviewers and area chairs mayvalue more, when reviewing a paper, a comparison with another paper that has been reported to have experimental consistency\.
- P5\.The program chairs mustgive each submitted paper an experimental consistency status, based on the recommendations of the area chairs\. It is an integral part of the final decision\.
When a paper is accepted:
- P6\.The authors mustprovide, before the camera\-ready deadline, a status as defined in[Table1](https://arxiv.org/html/2609.35854#S5.T1)regarding the software \(code and/or data, including model weights\) that supports the empirical results in their paper\. This software status will be visible on the proceedings web site, as well as the experimental consistency rating\. If a link to a repository is provided, the software must be readily available, although possibly after signing a license agreement\.
When a paper is published:
- P7\.The program chairs mustmake visible on the submission or proceedings web sites: the experimental consistency rating, the software availability status, the experimental material\.
When a paper is written:
- P8\.The authors of a new paper maystructure their arguments, tables and graphs to highlight results and comparisons to papers that are stamped as experimentally consistent or that provide software reproducing the experiments, separating them from unlabeled other papers\. \(Papers published before this policy have an ‘unavailable’ experimental consistency status\.\)
We note that the experimental consistency is limited to submitted logs and/or metric\-evaluation material with the reported results\. This status does not certify full reproducibility, the correctness of the training or data pipeline, the authenticity of the outputs, or overall reliability\.
Exemption\.
Not providing verification material is possible, but a justification must then be provided, which reviewers and area chairs will assess\. It may concern theoretical papers \(which may however have motivating experiments\), experiments with hidden\-test benchmark servers, experiments on private data \(which reviewer may however not value as highly\), or excessively large output files\. In contrast, copyright \(vs private\) data is not an excuse for not providing evaluation material as reviewers already commit to treat submissions as confidential\.
Implementation\.
This policy can easily be implemented in OpenReview: \(i\) It already hosts many major ML and CV conferences; \(ii\) More conferences continue to adopt it due to its flexibility and open\-source nature; \(iii\) It is actively supported by the research community\. Moreover, OpenReview is easily and commonly tailored according to the specific requirements of different venues\. Final statuses \([Table1](https://arxiv.org/html/2609.35854#S5.T1)\) and code links may also appear on the web sites of NeurIPS, PMLR and The CVF\.
This policy can be implemented incrementally, first adopting Level 1 \(log verification\), then when Level 1 is established, introducing Level 2 \(metric checking\) towards a stronger validation\.
## 7Alternative views
Imposing code submission\.Some conferences, such as VLDB, require code to be submitted\[[35](https://arxiv.org/html/2609.35854#bib.bib35)\]\. So does ASIACCS, although there was no code evaluation process this year\[[36](https://arxiv.org/html/2609.35854#bib.bib36)\]and a valid reason for not doing so could be provided\. Journals such as IPOL\[[37](https://arxiv.org/html/2609.35854#bib.bib37)\]also tightly couple papers and code\.
Incentivizing code submission\.ICPR prefers an incentive to promote code submission, with a Reproducible Research in Pattern Recognition \(RRPR\) Badge\[[38](https://arxiv.org/html/2609.35854#bib.bib38)\], which has been introduced since 2016\. The evaluation criteria are at the discretion of specially\-designated RRPR reviewers, and Reproducibility Chairs add their own meta\-review to the reproducibility reviews\. This review is performed only on accepted papers, thus having no impact on acceptance decisions\.
Encouraging code submission\.NeurIPS, ICLR, CVPR and ICCV encourage code submission, which reviewers are “welcome” to read, but “not required” to\. ICML allows code submission too, but is somewhat ambiguous: while “reproducibility of results and easy availability of code will be taken into account in the decision\-making process”, “it is entirely up to the reviewers to decide whether they wish to consult any of the appendices in the submitted paper or the supplementary material”\. However, the current practice is not full\-fledged code evaluation but code browsing for missing details\.
The code virtue hypothesis\.A previous reader of this position paper thinks that authors “should provide code, and that should be its own incentive” because “it better advances science”, as also emphasized in\[[39](https://arxiv.org/html/2609.35854#bib.bib39)\]\. However, while we note that publications with code have twice as many citations \([Section3\.2](https://arxiv.org/html/2609.35854#S3.SS2)\), we also observe a stagnation in code availability, between 40% and 50% of published papers, depending on venues \([Section3\.2](https://arxiv.org/html/2609.35854#S3.SS2)\)\. Besides, code success as an a posteriori assessment does not prevent some “toxic” SOTA but codeless papers from hindering the publication of new methods which would not reach this SOTA\. In contrast, we propose an a priori verification\.
No code, no cite\.Still more radical than our proposal, we know of colleagues who have decided in their own publications to never compare to any paper without code\. Our position is not quite so clear\-cut: we do not ban codeless papers, but prefer to highlight methods coming with code\.
De\-emphasizing reproducibility\.Contrary to the above,[Drummond \[40\]](https://arxiv.org/html/2609.35854#bib.bib40)argues that reproducibility is not essential to science, that requiring code to be submitted is unnecessary and would even reduce paper acceptance to narrow technical criteria, and that misconducts actually have little impact\.
De\-emphasizing experiments\.Not exactly contradictory, result\-blind peer review focuses on the ideas presented in the paper regardless of any empirical evidence, with a submitted paper deprived from results and conclusion\[[41](https://arxiv.org/html/2609.35854#bib.bib41)\]\. The full paper is to be provided in a second reviewing stage\[[42](https://arxiv.org/html/2609.35854#bib.bib42)\]\.
Code review burden\.Reviewing code, that took weeks or months to be written, is a heavy burden for reviewers, which probably explains why the RRPR badge is awarded via a specific committee\. One might thus consider that the workload is too high a price to pay for the added guarantee\. Our proposal however only concerns the evaluation code, which is usually small, if not already in libraries\.
## 8Perspectives
Our proposal is not bullet\-proof, but it should still make the lives of fraudsters a bit harder\. It would also make their infractions more explicit and less forgivable if they were discovered\.
To go beyond this log and metric checking would require code available at review time\. However, the decision for a venue to ignore possible reasons not to disclose code and impose code submission seems to be hard to make\.
We believe that checking the consistency of code with respect to a paper is actually a long shot, if only because of the significant reviewer workload and compute it would require\.
Still, Paper2Code\[[22](https://arxiv.org/html/2609.35854#bib.bib22)\], which proposes a first approach to automatically generate code from papers, interestingly also proposes an LLM\-based method to evaluating how well a code repository reflects the contents of a paper, with a level of performance close to that of a human\. This verification task is indeed simpler than code generation, and we may expect in the near future automated systems to provide reliable reports on the consistency between code and paper\. This should alleviate the load of*reproducibility reviewers*, who could concentrate on possible issues mentioned in generated consistency reports\. Security would, however, have to be guaranteed so that proprietary code cannot leak via LLM requests\. It would be up to program chairs to provide reviewers with reports produced by a secure LLM, keeping a human in the loop for interpretation\.
Another perspective is for the authors to upload code, models, and data to an independent third\-party server, secured to allow uploads of proprietary software\. The server would run the code and automatically generate a report, while an LLM\-based auditor would inspect the code for suspicious behavior, such as hard\-coded outputs, lookup tables, or other shortcuts that could artificially reproduce the claimed numbers\. Unfortunately, this framework is unlikely to be feasible in the near future because hosting, securing, and maintaining such a platform is expensive and operationally complex\.
## 9Conclusion
Our study shows that, after a period of marked improvement, the ability to reproduce empirical results of publications in top\-tier ML and CV conferences has recently deteriorated\. This observation comes at a time when the capacity of \(M\)LLMs is increasing dramatically, changing the way code is developed and papers are written and reviewed, while casting doubt on the reliability of publications\.
Although code and data availability remains the key to evaluating empirical results in computer science publications, many authors are still reluctant to make their software accessible\. Considering that providing enough details to make a paper really reproducible and assessable by reviewers is not very far from providing actual code should defeat arguments against code availability\. Still, the current consensus of the scientific community is to accept papers coming without supporting software, hence papers that could be hard to reproduce and thus to fully evaluate in the first place\.
In this context, while still promoting code availability for reproducibility and research acceleration, we approach the problem from theverifiabilityperspective, which is much lighter than actualreproducibility, although it remains efficient\. We propose a new policy to improve the ability to evaluate empirical results, based on the verification of experiment logs and metric computation, without the need for actual code, which does not guarantee replication anyway\. And this policy seamlessly integrates into the existing reviewing and publication process\.
There is no free lunch tough; our proposal requires some additional effort from authors and reviewers\. We however quantify it and argue it is minimal\. We believe this light extra workload is worth it, because it can significantly improve the quality and reproducibility of published papers\.
This is an initial proposal for establishing a new practice for evaluating empirical results and citing them\. We hope the community can discuss it, improve it, and start implementing it in upcoming conferences, at least as an experiment\.
## 10Acknowledgment
We acknowledge the EuroHPC Joint Undertaking for awarding the project ID EHPC\-REG\-2024R02\-234 access to Karolina, Czech Republic\.
## References
- \[1\]Karl Popper\.*Logik der Forschung*\.1934\.Translated to English as*The Logic of Scientific Discovery*in 1959\. Republished by Routledge, London, 1992\.
- \[2\]Ramal Moonesinghe, Muin J Khoury, and A\. Cecile J\.W\. Janssens\.Most published research findings are false\-but a little replication goes a long way\.*PLOS Medicine*, 4\(2\), 2007\.
- \[3\]Daniel J\. Simons\.The value of direct replication\.*Perspectives on Psychological Science*, 9\(1\):76–80, 2014\.
- \[4\]Harold Pashler and Christine R\. Harris\.Is the replicability crisis overblown? three arguments examined\.*Perspectives on Psychological Science*, 7\(6\):531–536, 2012\.
- \[5\]Monya Baker\.1,500 scientists lift the lid on reproducibility\.*Nature*, 533:452–454, 2016\.
- \[6\]Daniele Fanelli\.Is science really facing a reproducibility crisis, and do we need it to?*Proceedings of the National Academy of Sciences*, 115\(11\):2628–2631, 2018\.
- \[7\]Benjamin Antunes and David R\.C\. Hill\.Reproducibility, replicability and repeatability: A survey of reproducible research with a focus on high performance computing\.*Computer Science Review*, 53:100655, 2024\.
- \[8\]Sushree Namita Nag, Abhijit Roy, and KG Sudhier\.Global perspectives on retracted papers in artificial intelligence and machine learning: a bibliometric study\.*Global Knowledge, Memory and Communication*, 2025\.
- \[9\]UNESCO\.*UNESCO Science Report: the race against time for smarter development*\.UNESCO, 2021\.
- \[10\]Mark A\. Hanson, Pablo Gómez Barreiro, Paolo Crosetto, and Dan Brockington\.The strain on scientific publishing\.*Quantitative Science Studies*, 5\(4\):823–843, 2024\.
- \[11\]Benjamin F\. Jones\.The burden of knowledge and the “Death of the Renaissance Man”: Is innovation getting harder?*The Review of Economic Studies*, 76\(1\):283–317, January 2009\.ISSN 0034\-6527\.
- \[12\]Nazar Shmatko, Alex Adam, and Paul Esau\.GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers, Jan 2026\.[https://gptzero\.me/news/neurips/](https://gptzero.me/news/neurips/)\. Accessed Jan 27th, 2026\.
- \[13\]Miryam Naddaf\.Journals infiltrated with ‘copycat’ papers that can be written by AI\.*Nature*, September 2025\.
- \[14\]João Phillipe Cardenuto, Daniel Moreira, and Anderson Rocha\.Unveiling scientific articles from paper mills with provenance analysis\.*PLOS ONE*, 19\(10\):1–28, 2024\.doi:10\.1371/journal\.pone\.0312666\.
- \[15\]Haitham S\. Al\-Sinani and Chris J\. Mitchell\.From content creation to citation inflation: A GenAI case study, 2025\.arXiv preprint arXiv:2503\.23414\.
- \[16\]Kat Boboris\.Attention authors: updated endorsement policy, 2026\.arXiv blog,[https://blog\.arxiv\.org/2026/01/21/attention\-authors\-updated\-endorsement\-policy/](https://blog.arxiv.org/2026/01/21/attention-authors-updated-endorsement-policy/)\. Accessed Jan 27th, 2026\.
- \[17\]Joelle Pineau, Philippe Vincent\-Lamarre, Koustuv Sinha, Vincent Larivière, Alina Beygelzimer, Florence d’Alché Buc, Emily Fox, and Hugo Larochelle\.Improving reproducibility in machine learning research \(a report from the NeurIPS 2019 reproducibility program\)\.*Journal of machine learning research*, 22\(164\), 2021\.
- \[18\]Edward Raff, Michel Benaroch, Sagar Samtani, and Andrew L Farris\.What do machine learning researchers mean by “reproducible”?In*AAAI Conference on Artificial Intelligence*, volume 39, pages 28671–28683, 2025\.
- \[19\]Chris Drummond\.Replicability is not reproducibility: nor is it good science\.In*Proceedings of the Evaluation Methods for Machine Learning Workshop at the 26th ICML*, 2009\.
- \[20\]Faisal Shehzad, Timo Breuer, Maria Maistro, and Dietmar Jannach\.“we share our code online”: Why this is not enough to ensure reproducibility and progress in recommender systems research\.In*ACM Conference on Recommender Systems*, 2025\.
- \[21\]Edward Raff\.A step toward quantifying independently reproducible machine learning research\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, volume 32, 2019\.
- \[22\]Minju Seo, Jinheon Baek, Seongyun Lee, and Sung Ju Hwang\.Paper2Code: Automating code generation from scientific papers in machine learning, 2025\.arXiv preprint arxiv:2504\.17192\.
- \[23\]Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al\.PaperBench: Evaluating AI’s ability to replicate AI research\.In*International Conference on Machine Learning \(ICML\)*, 2025\.
- \[24\]Benjamin Haibe\-Kains, George Alexandru Adam, Ahmed Hosny, Farnoosh Khodakarami, Thakkar Shraddha, Rebecca Kusko, Susanna\-Assunta Sansone, Weida Tong, Russ D\. Wolfinger, Christopher E\. Mason, Wendell Jones, Joaquin Dopazo, Cesare Furlanello, Levi Waldron, Bo Wang, Chris McIntosh, Anna Goldenberg, Anshul Kundaje, Casey S\. Greene, Tamara Broderick, Michael M\. Hoffman, Jeffrey T\. Leek, Keegan Korthauer, Wolfgang Huber, Alvis Brazma, Joelle Pineau, Robert Tibshirani, Trevor Hastie, John P\. A\. Ioannidis, John Quackenbush, and Hugo J\. W\. L\. Aerts\.Transparency and reproducibility in artificial intelligence\.*Nature*, 586\(7829\):E14–E16, October 2020\.
- \[25\]Sayash Kapoor and Arvind Narayanan\.Leakage and the reproducibility crisis in machine\-learning\-based science\.*Patterns*, 4\(9\), 2023\.
- \[26\]Nikita Ravi, Abhinav Goel, James C Davis, and George K Thiruvathukal\.Improving the reproducibility of deep learning software: An initial investigation through a case study analysis, 2025\.arXiv preprint arXiv:2505\.03165\.
- \[27\]Daniele Cono D’Elia, Thaleia Dimitra Doudali, Cristiano Giuffrida, Miguel Matos, Mathias Payer, Solal Pirelli, Georgios Portokalidis, Valerio Schiavoni, Salvatore Signorello, and Anjo Vahldiek\-Oberwagner\.Lessons learned from five years of artifact evaluations at EuroSys\.In*3rd ACM Conference on Reproducibility and Replicability*, pages 108–120, 2025\.
- \[28\]Odd Erik Gundersen, Malte Helmert, and Holger Hoos\.Improving reproducibility in AI research: Four mechanisms adopted by JAIR\.*Journal of Artificial Intelligence Research*, 81, 2024\.
- \[29\]Semantic Scholar\.Semantic Scholar API, 2026\.[https://api\.semanticscholar\.org/graph/v1/paper/search](https://api.semanticscholar.org/graph/v1/paper/search)\. Accessed Jan 18th, 2026\.
- \[30\]Kevin Musgrave, Serge Belongie, and Ser\-Nam Lim\.Three new validators and a large\-scale benchmark ranking for unsupervised domain adaptation, 2022\.arXiv preprint arXiv:2208\.07360\.
- \[31\]TensorFlow Team\.TensorBoard: TensorFlow’s visualization toolkit\.[https://www\.tensorflow\.org/tensorboard](https://www.tensorflow.org/tensorboard), 2026\.Accessed Jan 25th, 2026\.
- \[32\]Matei Zaharia, Andrew Chen, Aaron Davidson, Ali Ghodsi, Sue Ann Hong, Andy Konwinski, Siddharth Murching, Tomas Nykodym, Paul Ogilvie, Mani Parkhe, Fen Xie, and Corey Zumar\.Accelerating the machine learning lifecycle with MLflow\.*Bulletin of the IEEE Computer Society Technical Committee on Data Engineering*, 41\(4\):39–45, December 2018\.
- \[33\]Lukas Biewald\.Experiment tracking with weights and biases, 2020\.URL[https://www\.wandb\.com/](https://www.wandb.com/)\.Software available from wandb\.com\.
- \[34\]PMLR\.Proceedings of Machine Learning Research \(PMLR\), 2026\.Proceedings of ICML 2025\.[https://proceedings\.mlr\.press/v267/](https://proceedings.mlr.press/v267/)\. Accessed Jan 25th, 2026\.
- \[35\]PVLDB\.Proceedings of the Very Large Data Bases \(VLDB\) Endowment, 2026\.Transparency and reproducibility\.[https://www\.vldb\.org/pvldb/volumes/19/submission](https://www.vldb.org/pvldb/volumes/19/submission)\. Accessed Jan 25th, 2026\.
- \[36\]ASIACCS\.ACM ASIA Conference on Computer and Communications Security \(ASIACCS\), Call for Papers, 2026\.[https://asiaccs2026\.cse\.iitkgp\.ac\.in/call\-for\-papers/](https://asiaccs2026.cse.iitkgp.ac.in/call-for-papers/)\. Accessed Jan 25th, 2026\.
- \[37\]IPOL\.Image Processing On Line \(IPOL\) Journal, Editorial policy, 2026\.[https://www\.ipol\.im/meta/policy/](https://www.ipol.im/meta/policy/)\. Accessed Jan 25th, 2026\.
- \[38\]ICPR\.International Conference on Pattern Recognition \(ICPR\), Reproducible Research in Pattern Recognition \(RRPR\) Badge, 2026\.[https://icpr2026\.org/rrprBadges\.html](https://icpr2026.org/rrprBadges.html)\. Accessed Jan 25th, 2026\.
- \[39\]David Donoho\.Data science at the singularity\.*Harvard Data Science Review*, 6\(1\), 2024\.
- \[40\]Chris Drummond\.Reproducible research: a minority opinion\.*Journal of Experimental & Theoretical Artificial Intelligence*, 30:1–11, 2018\.
- \[41\]Robert Rosenthal\.*Experimenter Effects in Behavioral Research*\.John Wiley & Sons, 1966\.
- \[42\]Michael J\. Mahoney\.Publication prejudices: An experimental study of confirmatory bias in the peer review system\.*Cognitive Therapy and Research*, 1:161–175, 1977\.URL[https://api\.semanticscholar\.org/CorpusID:7350256](https://api.semanticscholar.org/CorpusID:7350256)\.
- \[43\]Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang, Penglong Jiao, Zhenjiang Jin, Zhikai Lei, Jiaxing Li, Jingwen Li, Linyang Li, Shuaibin Li, Wei Li, Yining Li, Hongwei Liu, Jiangning Liu, Jiawei Hong, Kaiwen Liu, Kuikun Liu, Xiaoran Liu, Chengqi Lv, Haijun Lv, Kai Lv, Li Ma, Runyuan Ma, Zerun Ma, Wenchang Ning, Linke Ouyang, Jiantao Qiu, Yuan Qu, Fukai Shang, Yunfan Shao, Demin Song, Zifan Song, Zhihao Sui, Peng Sun, Yu Sun, Huanze Tang, Bin Wang, Guoteng Wang, Jiaqi Wang, Jiayu Wang, Rui Wang, Yudong Wang, Ziyi Wang, Xingjian Wei, Qizhen Weng, Fan Wu, Yingtong Xiong, Chao Xu, Ruiliang Xu, Hang Yan, Yirong Yan, Xiaogui Yang, Haochen Ye, Huaiyuan Ying, Jia Yu, Jing Yu, Yuhang Zang, Chuyu Zhang, Li Zhang, Pan Zhang, Peng Zhang, Ruijie Zhang, Shuo Zhang, Songyang Zhang, Wenjian Zhang, Wenwei Zhang, Xingcheng Zhang, Xinyue Zhang, Hui Zhao, Qian Zhao, Xiaomeng Zhao, Fengzhe Zhou, Zaida Zhou, Jingming Zhuo, Yicheng Zou, Xipeng Qiu, Yu Qiao, and Dahua Lin\.Internlm2 technical report, 2024\.arXiv preprint arXiv:2403\.17297\.
- \[44\]Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter\.GANs trained by a two time\-scale update rule converge to a local Nash equilibrium\.In*NeurIPS*, 2017\.
## Appendix
In this appendix, we provide:
- •details on our study regarding code availability and citations \(App\.[A](https://arxiv.org/html/2609.35854#A1)\),
- •details and examples regarding the verification framework \(App\.[B](https://arxiv.org/html/2609.35854#A2)\),
- •an additional discussion regarding code, before and after acceptance \(App\.[C](https://arxiv.org/html/2609.35854#A3)\)\.
We also wish to recall that our proposal is a starting point, not a set\-in\-stone procedure\. Thanks to this position paper, we hope the community will be able to debate the question, improve the proposition into a clear Instruction Guide and FAQ, and implement it at an upcoming event \(e\.g\., workshop\) on a trial basis before it can be applied more broadly and at a larger scale\. To gain credibility and be in a position to convince the organizers of a venue to support a pilot study, we also believe our proposal will benefit from a recognition as a position paper in a major conference like NeurIPS\.
## Appendix ADetails of the study on code availability
In this section, we explain our study on code availability \([Section3\.2](https://arxiv.org/html/2609.35854#S3.SS2)\) in detail\. Our focus is on the top\-tier Computer Vision \(CV\) and Machine Learning \(ML\) conferences, i\.e\., CVPR, ICCV, ICLR, ICML and NeurIPS\. We also narrow our scope to the last 5 years,i\.e\., from 2021 to 2025 inclusive\.
We first downloaded the PDFs \(main paper and supplementary\) of all the accepted papers from their official website \([thecvf\.com](https://www.thecvf.com/),[iclr\.cc](https://iclr.cc/),[icml\.cc](https://icml.cc/),[neurips\.cc](https://neurips.cc/)\)\. Then, we analyzed their contents to detect possible code URLs\. We extracted all the URL in the paper using a generic regular expression and fed them, along with the context surrounding them, to an LLM to determine which URL is more likely to be the implementation of the method described in the paper\. We present the prompt used during this process in Prompt[A](https://arxiv.org/html/2609.35854#A1)\.
In the next stage, we tried to clone into a local directory the code URL of each paper, if available\. It is common practice in CV/ML papers to provide a project page containing a link to the code repository along with other useful information, such as visual results\. To overcome the cloning failures in such cases, we added a resolver function if the detected URL is not a “Git” repository but a project page\. This resolver looks for any repository link on the project page\.
For each cloned repository, we again prompted the LLM using the downloaded files and the code URL to tell:
- •If the repository is training based,
- •If specific training is required,
- •If the training code is available,
- •If the inference or test code is available,
- •If the training instructions are available,
- •If the model weights are available,
- •If there are open issues regarding the installation or reproducibility of results in the paper,
- •If other repositories are acknowledged, suggesting that external code may have been borrowed and adapted\.
In Prompt[A](https://arxiv.org/html/2609.35854#A1), we present the prompt used to obtain this assessment of the code repositories\.
Table 2:Detailed statistics for CVPR\.Finally, to understand which fields \(topics\) in the CV/ML domain are more active in open\-sourcing their code, we classify each paper into a predefined set of topics\. Specifically, we prompted the LLM using only the first two pages of the downloaded PDF and asked it to determine a single topic\.
The predefined topics \(which were also LLM\-generated to cover the 10 main topics of our five target conferences\) are:
1. 1\.3D Vision and Neural Rendering,
2. 2\.Image and Video Synthesis,
3. 3\.Multimodal Learning \(Vision \+ Language \+ Reasoning\),
4. 4\.Large Language Models and Foundation Models,
5. 5\.Generative AI and Diffusion Models,
6. 6\.Reinforcement Learning and Decision Making,
7. 7\.Trustworthy, Fair and Robust Machine Learning,
8. 8\.Optimization and Theory of Learning,
9. 9\.Scalability, Systems and Distributed ML,
10. 10\.Application\-Driven ML\.
In Prompt[A](https://arxiv.org/html/2609.35854#A1), we share the prompt used to classify each paper’s topic\.
We used InternLM\[[43](https://arxiv.org/html/2609.35854#bib.bib43)\]\(internlm3\-8b\-instruct\) as the LLM\. We finetuned the prompts until the provided results were identical to a human analysis based on a sample of 100 randomly selected papers from the latest editions of all venues\. In all prompts, we asked the LLM to also supply a reason and evidence for the final response\. We observed that pushing the LLM to reason significantly improved the response quality and reduced the number of hallucinated responses\.
After optimizing our prompts on the 100 sampled papers, residual errors were marginal, with a slight tendency to overestimate the count of papers with code or weights\. The actual situation of code and weight availability is thus actually slightly worse than what we report\. \(We do not have resources to get meaningful error bars for the analysis of these 55,377 papers, which would require manually creating a large pool of ground\-truth statuses\.\)
We also extracted the number of citations for each accepted paper using the Semantic Scholar API\[[29](https://arxiv.org/html/2609.35854#bib.bib29)\]to further analyze the effect of code sharing on citations\. \(Forks and stars are not good metrics: some repositories without any code and just introducing a paper may have thousands of stars\.\)
Table 3:Detailed statistics for ICCV\.We present detailed results of this study for each venue in[Tables2](https://arxiv.org/html/2609.35854#A1.T2),[3](https://arxiv.org/html/2609.35854#A1.T3),[4](https://arxiv.org/html/2609.35854#A1.T4),[5](https://arxiv.org/html/2609.35854#A1.T5)and[6](https://arxiv.org/html/2609.35854#A1.T6)for CVPR, ICCV, ICLR, ICML, and NeurIPS, respectively\. There are some discrepancies between the number of accepted and downloaded papers, which is due to either scraping errors or, in times, the official numbers being incorrect, as in the case of CVPR’2022, where the CVF open access page indeed contains 2074 papers\. While the regular expression extracts all potential URLs, the LLM typically filters these to identify only the subset likely to represent the paper’s actual implementation\. Notably, the “No URL in PDF” row reveals a significant number of instances where papers contain no URLs at all\. Unfortunately, not all the detected code URLs by the LLM, in fact, point to a valid code repository\. Sometimes, the LLM detects non\-code repositories, such as links to tools or software referred to in the paper, as possible code URL\. However, most of the time, the code URL presented in the paper is not reachable, which reduces the number of cloned repositories\.
In some cases, even though a repository is successfully cloned to a local directory, it does not include any files except for a brief Readme file\. This leads to a difference between the number of cloned repositories and the repositories with training and/or test code available\. Another important observation is that a significant number of papers present only the training code but not the pretrained model weights, which is another barrier to reproducibility\.
We see the importance of open\-sourcing the code in the “Acknowledges Previous Repos” row, which shows the number of repositories that acknowledge one or more publicly available repositories\. As expected, this number increases each year for all the venues, notably doubling for ICLR from 2024 to 2025\. Finally, the number of issues regarding reproducibility in corresponding code pages is very low for open\-sourced papers\.
Table 4:Detailed statistics for ICLR\.We observe \([Figure2](https://arxiv.org/html/2609.35854#S3.F2)that the proportion of theoretical papers \(i\.e\., with topic “Optimization and Theory of Learning”, including “Statistical learning theory, optimization methods, generalization bounds, theory insights”\) is decreasing while the proportion of theoretical papers with available code is generally increasing or stable\. We thus concluded there was no surge of code\-less theoretical papers, and thus that the observation of the decline or stagnation of papers with code was not biased by a reduction of the needs for empirical evaluation\.
A similar code availability study is conducted in Paper2Code\[[22](https://arxiv.org/html/2609.35854#bib.bib22)\]\. However, since the authors only inspected the abstracts of the papers, their numbers differ from ours\. Their inspection includes the accepted papers of the 2024 editions of ICLR, ICML, and NeurIPS, and they report that only 20% of the papers have publicly available code, which is significantly lower than ours \(around 50%\)\. Our study is more detailed, as we search for the code URLs in the entire paper \(with a regexp\)\.
Table 5:Detailed statistics for ICML\.Table 6:Detailed statistics for NeurIPS\.We attach all the code used in this study to the supplementary material\. It will be made publicly available upon acceptance\.
Code URL Identification PromptYou will be given URL candidates extracted from a paper PDF, together with page snippets\.Your task is to identify which URL most likely points to the paper’s open\-source code:•GitHub repository•GitLab repository•code release pageAssume there is at most one true code repository for the paper\. All remaining URLs should be assigned toother\_urls\. A single evidence snippet is sufficient\.Rules\.•Be strict\.•Include a URL only if the snippet suggests code or implementation, or if the URL itself is clearly a repository\.•Include at most the top\-4 most relevant entries inother\_urls\.•Return only valid JSON\.•Do not return markdown or commentary\.Output schema\.``` { "code_urls": [ { "url": "string", "confidence": 0.0, "reason": "short string", "evidence": [ { "page": 0, "snippet": "string" } ] } ], "other_urls": [ { "url": "string", "type": "dataset|project_page|paper|supplement|other", "reason": "short string", "evidence": [ { "page": 0, "snippet": "string" } ] } ] } ``` \{promptbox\}
Prompt template used for code URL identification\.
Repository Audit PromptYou are auditing a GitHub repository for training and inference availability\.You must base every answer only on the provided repository evidence bundle:•README•file tree•embedded file contents•optional issue excerptsRules\.•Do not guess\.•If evidence is missing, setvaluetonulland use:unknown \(not found in provided repo materials\)\.•Return only one valid JSON object\.•Do not return markdown or commentary\.Evidence requirements\.•Every field must include anevidencelist\.•Each evidence item must contain:–source\_type: one ofREADME,FILE\_TREE,FILE\_CONTENT, orISSUE–location: a precise pointer such asREADME \> Training, a file path, or an issue identifier–quote: an exact quote of at most 25 wordsConservative interpretation\.•training\_code\_available = trueonly if training\-related code or explicit training instructions are present\.•inference\_or\_test\_code\_available = trueonly if inference/evaluation code or explicit instructions are present\.•weights\_available\_for\_this\_method = trueonly if checkpoints are provided for the audited method\.•presents\_dataset = trueonly if the repository itself presents a dataset\.•open\_issues = trueonly if issues explicitly mention installation problems, missing steps, reproducibility, or similar concerns\.•Ignore issues authored byNielsRogge\.•Include the full repository URL\.Output schema\.``` { "repo": { "url": "", "name": "" }, "assessment": { "is_training_based": { "value": null, "rationale": "", "evidence": [] }, "requires_specific_training": { "value": null, "rationale": "", "evidence": [] }, "training_code_available": { "value": null, "rationale": "", "evidence": [] }, "inference_or_test_code_available": { "value": null, "rationale": "", "evidence": [] }, "training_instructions_available": { "value": null, "rationale": "", "evidence": [] }, "inference_or_test_instructions_available": { "value": null, "rationale": "", "evidence": [] }, "weights_available_for_this_method": { "value": null, "rationale": "", "evidence": [] }, "presents_dataset": { "value": null, "rationale": "", "evidence": [] }, "open_issues": { "value": null, "rationale": "", "evidence": [] }, "acknowledgement": { "value": null, "rationale": "", "evidence": [] } }, "open_issues_lowerbound_notes": { "value": "", "evidence": [] } } ``` \{promptbox\}
Prompt template used for repository audit \(training/inference availability\)\.
Paper Topic Classification PromptYou are classifying research paper PDFs into exactly one predefined topic\.You will be given a bundle of papers\.For each paper:•choose the single best topic•justify it using evidence from the provided PDF snippets onlyRules\.•Do not guess\.•Use only the provided PDF snippets\.•Do not use outside knowledge\.•Each paper must have exactly one selected topic, ornullif unknown\.•If the paper cannot be confidently assigned, settopic\_idtonulland use exactly:unknown \(insufficient evidence in provided PDF snippets\)•Return only valid JSON\.•Do not return markdown or commentary\.Evidence requirements\.•Every paper entry must include anevidencelist\.•Each evidence item must contain:–page: integer page index–snippet: exact snippet copied from the provided materials•Keep each snippet at most 25 words\.•Do not include double quotes in the snippet field\.•Remove or avoid unparsable characters such as math symbols or broken glyphs\.•Keep the rationale short, ideally 1–2 sentences\.Topics\.1\.3D Vision and Neural Rendering2\.Image and Video Synthesis3\.Multimodal Learning \(Vision \+ Language \+ Reasoning\)4\.Large Language Models and Foundation Models5\.Generative AI and Diffusion Models6\.Reinforcement Learning and Decision Making7\.Trustworthy, Fair and Robust Machine Learning8\.Optimization and Theory of Learning9\.Scalability, Systems and Distributed ML10\.Application\-Driven MLOutput schema\.``` { "results": [ { "paper": { "paper_id": "string", "filename": "string" }, "topic": { "topic_id": 0, "topic_name": "string" }, "confidence": 0.0, "rationale": "string", "evidence": [ { "page": 0, "snippet": "string" } ] } ] } ``` \{promptbox\}
Prompt template used to classify the topic of each paper\.
## Appendix BDetails and examples regarding the verification framework
### B\.1Log verification details and examples
If the authors trained a model, they most certainly already have logs, from which they visually monitored training and execution curves while developing their approach\. The extra work to provide log verification material for reviewers is then just to document log information\.
Besides snapshots of log curves pointing at the corresponding experiments in the paper \(e\.g\., a line in a SOTA table\), log verification material should include a common core of information for both training and test\. This information includes the following fields:
- •a clearreference to an experimentin the paper, e\.g\., a line in a SOTA table,
- •thedataset sizeused in the run,
- •GPU utilizationandCPU utilization,
- •model size\(number of parameters\),
- •FLOPs\(or a comparable compute estimate\)\.
On top of these shared fields, training logs must report:
- •thelearning schemeandtraining parameters,
- •theloss valuesover steps or epochs,
while validation logs must report
- •the finalevaluation metrics\.
including what was actually reported in the paper\. In the Appendix, we present example log snapshots taken from TensorBoard, MLflow, and W&B with the required fields\.
As examples of log verification material, we show representative visualizations produced by three widely adopted experiment\-tracking tools for custom training: MLflow \(Figure[5](https://arxiv.org/html/2609.35854#A2.F5)\), TensorBoard \(Figure[6](https://arxiv.org/html/2609.35854#A2.F6)\), and Weights & Biases \(Figure[7](https://arxiv.org/html/2609.35854#A2.F7)\)\. To facilitate comparison, we organize the plots into four thematic groups that roughly follow the life cycle of a training run\.
The first group summarisesmodel characteristics, including the total and trainable parameter counts, as well as an estimate of computational cost \(FLOPs\)\. The second group reportssystem\-level signalscaptured during training, such as GPU memory usage \(used, allocated, and reserved\)\. When supported by the logging backend, we additionally track host memory utilization and GPU utilization\. The third group focuses onoptimization procedure, recording the number of epochs completed alongside training loss and the learning\-rate schedule\. Finally, the fourth group presentsevaluation metricson the validation set and, when available, the test set, together with per\-epoch validation loss\.
For each experiment to verify, the task \(P2\) of the reviewer includes, but is not limited to:
- •checking that the training process is consistent with the paper,
- •assessing training convergence, including performance variance,
- •checking the performance consistency with the paper,
- •detecting bad practices, e\.g\., training on more data that said \(e\.g\., comparing the number of epochs vs the batch size and iteration count\) or cherry\-picking checkpoints\.




Figure 5:Example log plots from MLFlow\.



Figure 6:Example log plots from TensorBoard\.



Figure 7:Example log plots from W&B\.
### B\.2Metric verification details
As mentioned in the main paper, software \(code and data\) is central to reproducibility for computer science papers\. \(Note however that computer science papers sometimes also include experiments that are not purely virtual, such as robotic manipulations\. They may include human studies too, e\.g\., to evaluate acceptability, assess image realism, or measure the time in a computer\-assisted human labeling process\. We only consider here repeatable*virtual*experiments with mathematically\-defined metrics\.\)
If the authors reported some metrics, then they already processed generated outputs using some evaluation code and, possibly, ground\-truth data\. While training or inference code could be sensitive for the authors, metric code should not, or it should at least be accessible to reviewers, who already currently commit to treat all information related to submissions as confidential\. In fact, a metric must anyway have a public definition that enables its implementation\. We presume that, most of the time, the evaluation code and data will be standard and readily available\.
The extra task then just amounts to:
- •packaging the output files and ground\-truth data \(if any\),
- •making the evaluation code stand\-alone, if not already the case\.
A code assistant can typically see to it\.
To control if errors were introduced when computing metrics or when reporting them, the verification material of an experiment should include:
- •adefinitionorreferenceto the experiment in the paper,
- •aJSON filerepresenting the output of the experiment,
- •a link to themetric evaluation codefrom an official dataset or benchmark, or if the metric is original, the anonymized evaluation code, which may include a trained model \(e\.g\., to propose alternative features to an FID\-like evaluation\[[44](https://arxiv.org/html/2609.35854#bib.bib44)\]\),
- •a link toground\-truth dataused for the evaluation, whether public or proprietary if the test data is original,
- •aguidefor download and usage\.
No license agreement should be needed to access proprietary data because it could break the anonymity of the authors or of the reviewer\.
To prevent installation issues, as reviewers have to check the evaluation material and rerun it, we actually recommend that authors prepare an anonymized Google Colab\-like notebook containing:
- •output data, or code to download them anonymously,
- •ground\-truth dataor, preferably, code to download them from standard repositories, e\.g\., original dataset web site or Hugging Face datasets,
- •evaluation code, or code to download it, using standard procedures when applicable, e\.g\., calls to PyTorch libraries or dataset devkits\.
Examples of such notebooks are available from[https://github\.com/papersubmissions13/PositionPaper](https://github.com/papersubmissions13/PositionPaper)\.
Checking the metric evaluation software includes, but is not limited to:
- •checking the genuineness of provided or downloadable ground\-truth data, e\.g\., making sure difficult samples were not excluded,
- •checking the validity of provided or downloadable evaluation code,
- •running the evaluation code and checking the result consistency with the paper,
- •reporting on it in the review and balancing it for the recommendation \(P4\)\.
The implied discipline to provide and review such material is also expected to help the community standardize evaluation codes, contributing to more fair evaluations and improved paper quality\.
### B\.3Verification process
#### A priori verification vs a posteriori replication\.
The target of our position paper is not science in general but, more modestly, the organization of ML / CV conferences\. It is the responsibility of program chairs to select papers \(1\) that are sound and \(2\) that have potential to make a significant impact in their field\. We believe a systematic a priori verification at review time, even if it is partial and possibly sidestepped, would be a notable improvement towards more soundness and reproducibility\. In contrast, a posteriori code verification is highly fortuitous \(limited code availability, need for good\-will users, substantial effort, resource requirements\) and comes too late, after acceptance, in a publication landscape which is certainly not without errors but that chiefly ignores retraction\.
In fact, our original motivation is to address “toxic” papers, with partly hidden experimental protocols and irreproducible results\. We believe that such papers, which prevent meaningful other papers from being submitted or accepted because they are not SOTA, would have a harder time being published if some basic verifications could be done at review time to find traces of secret recipes\.
Our hope, if such a verification becomes a standard, is that it will actually contribute to filtering out papers from careless authors, thus promoting better scientific practices\. However, it will not prevent authors that deliberately forge results from also forging verification material\. Totally preventing forgery would not only require full code and data availability \(hence raising proprietary issues\), but also demand a crazy amount of verification effort from reviewers, including rerunning training and inferences \(assuming reviewers have the time and resources to do so\)\.
#### Verification burden\.
For common cases, verifying both log data and evaluation code should only be a matter of minutes \(not counting execution time\)\. Let’s considering the short evaluation script example that we provide, It is 35 lines long but less than half of them actually matter \(disregarding empty lines, comments, title printing, file loading check, and import commands\)\. As they contain little or no algorithmic content, checking the logic of these lines actually requires less than 5 min\. Arguably, as a comparison, this is less difficult and takes less time than checking a formal proof in a paper\. For authors, typesetting a proof also takes longer than writing such a code\.
Assuming that most papers have 1\-4 main empirical claims regarding 1\-3 metrics, with some sharing regarding data and metrics, and given that reviewers are often familiar with the dataset and metrics, we consider the total verification time to be less than one hour in general, and often much less\. It is a price to pay, but it is particularly small compared to a full\-fledged code verification\.
Area and program chairs only have to take into account extra comments regarding verification when assessing a paper \(P3, P5\)\. And the extra load after acceptance \(P6\-8\) is very light for all actors\. In any case, improving verifiability cannot come for free\. Our verification effort however is minor for all actors, offering a good compromise regarding enhanced guarantees\.
#### Overhead variance\.
Just like some papers are much easier than others to read and review, we expect some variance in the verification effort\. While multiple datasets further increase the overheads, limiting the verification to SOTA claims restrains the required effort, including for large benchmark papers\. As for heavy outputs, we expect most of the complexity to be hidden in JSON files and, most often, in standard evaluation procedures\. To reduce the overall effort, we also suggest a single metric reviewer \(see below\)\.
#### Single evaluator\.
To keep the workload low, we propose that there be a single metric evaluator per paper\. This substantially reduces the workload of reviewers on average\. For NeurIPS last year \(4\-6 papers per reviewer and 4\-5 reviewers per paper\), metric evaluation would have concerned 1 paper per reviewer, occasionally 2\. And it is overestimated as it assumes that all submissions provide evaluation code, whereas there is actually a number of papers for which such an evaluation either does not apply \(e\.g\., theoretical papers\) or is not possible \(see Exemptions in[Section6](https://arxiv.org/html/2609.35854#S6)\) — which is ok\. A possible lighter alternative, to be discussed with the community, is that the reviewer assigned to log and metric verifications reads the paper to understand experiments but does not review it in depth, focusing only on experiment verification\.
#### Escaping verification and forged verification material\.
Verification can be escaped with proper justification \(Exemptions in[Section6](https://arxiv.org/html/2609.35854#S6)\)\. Still, we think it is better than nothing\. We also hope the authors can catch some errors themselves while preparing verification material, which they could have missed otherwise\. Besides, forged verification material constitutes evidence of misconduct\. If offenders are caught, it will be harder for them to pretend it was a lapse in attention\.
## Appendix CCode before and after acceptance
### C\.1Code as full\-fledged part of the reviewing
#### Reviewing code\.
The safest but most radical option to try to enforce reproducibility is to impose code reviewing\. Although it would not address all possible reasons for not disclosing code, the authors providing code could get the formal guarantee that reviewers will undertake not to uncover any information from the submitted material\.
It opens the possibility for reviewers to fully check the code attached to an experiment and rerun it\. However, not all experiments could be reproduced in this way, particularly those requiring a lot of resources \(compute, memory space, time, etc\.\) Besides, it requires a lot of effort from reviewers\.
Ideally, validation would involve uploading code, models, and data to an independent third\-party server, secured to allow uploads of proprietary software\. The server would run the code and automatically generate a report, while an LLM\-based auditor would inspect the code for suspicious behavior, such as hard\-coded outputs, lookup tables, or other shortcuts that could artificially reproduce the claimed numbers\. Unfortunately, this framework is unlikely to be feasible in the near future because hosting, securing, and maintaining such a platform is expensive and operationally complex\.
#### A few authors already submit code\.
We do not have large\-scale statistics regarding authors that choose to provide code as supplementary material\. We only observe, as ACs for recent CV conferences \(CVPR 2025, 2026 and ECCV 2024, 2026\), that we only got 16/74 submissions with some code, i\.e\., a bit more than 20%\. And out of the 48 reviews for the 16 submissions with code, only two reviewers mentioned the code, and just to say they “appreciate the effort toward reproducibility”\. Besides, one reviewer of a codeless submission complained that code was not submitted “for verification”\. Last, a number of reviewers asked for the code release plan, thus implicitly trusting the authors\.
However incomplete and topically biased this experience may be, we consider that only a small minority of submissions include code, and that it is hardly ever used in the review process\. This is understandable given the effort that is required to actually review code; the best that can be expected is to allow reviewers to possibly peek at the code when something is unclear in the paper description\.
### C\.2Code after acceptance
#### The “code being polished” excuse\.
The argument that “time is needed to polish the code” \([Section3](https://arxiv.org/html/2609.35854#S3)\) does not hold much nowadays, as cleaning up and reorganizing code can largely be done by code assistants\. AI indeed makes code preparation easier, but it is not a magic wand\. Most of the time spent to make code available is not in polishing it or even documenting its use, but in making sure it can be installed, with relevant data, and rerun with results that are consistent with the paper\. Code polishing sounds more of an excuse to delay availability than a real reason\.
#### Discussing the code virtue hypothesis\.
The code virtue hypothesis is to consider that, by making code available, authors get their own reward: if they provide good code, it will be reused, their paper will be cited more \(which is good for them\) and science will advance \(which is good for humanity\)\.
Now if, after acceptance, authors provide no code, or only partial code, or code/weights that does not succeed in reproducing results in the paper, or weights that do reproduce good results but that do not generalize because they were trained on test sets, they will nevertheless be cited \(which is good for them\) but it will not advance science; on the contrary, it will prevent other papers from being published or cited because they will not be SOTA \(which is unfair and bad for humanity\)\.
In fact, if providing good code was enough of an incentive to get more citations, it would have become the standard and we would not witness a plateauing in the proportion of code made available \(40\-50%\) nor an increase in open issues related to reproducibility in code repositories of accepted papers in the last 3 years: 1\.9% in 2023, 2\.3% in 2024, 4\.0% in 2025 \(numbers computed from[Tables2](https://arxiv.org/html/2609.35854#A1.T2),[3](https://arxiv.org/html/2609.35854#A1.T3),[4](https://arxiv.org/html/2609.35854#A1.T4),[5](https://arxiv.org/html/2609.35854#A1.T5)and[6](https://arxiv.org/html/2609.35854#A1.T6)\)\. The proportion goes even up to 4\.9% in 2025 for papers that make both train and test code available\.
We believe that the increasing pace of science and the pressure to publish more undermines a behavior that otherwise should indeed be virtuous\. We are not saying that all authors of papers with reproducibility issues are intentional fraudsters; some probably are, but others are simply the victims of carelessness or inadvertent errors\.
Rather than let time decide on the soundness and significance of a paper, our proposal is to include an extra verification dimension at review time to try to filter out some of these problematic papersbefore they are published\. Because afterwards, it is too late\[[8](https://arxiv.org/html/2609.35854#bib.bib8)\]\.
Related to this hypothesis, our citation analysis \([Figure3](https://arxiv.org/html/2609.35854#S3.F3)\), which observes that papers with code have twice as many citations as paper without code, is correlational and not causal\. Nevertheless, we consider that a large number of people think that available code helps increase the number of citations\. This belief should be enough as an incentive to provide code, but it seems to be reaching its limits \([Section3\.2](https://arxiv.org/html/2609.35854#S3.SS2)\)\. In any case, this motivating assumption does not invalidate our proposition for an a priori verification\.
#### Limitations\.
While well\-captioned training curves can reveal cherry\-picked checkpoints, and training on more data than said, it cannot catch all forms of data leakage and other methodological flaws, nor lookup cheating and forged logs or output files\. In fact, the verification we propose is only partial, and can be fooled by forged verification data, which is easily done nowadays with an AI\. It is an accepted limitation\.Similar Articles
Reproducibility seems to be headed towards irrelevance in ML research. Is it too late? [D]
An opinion piece arguing that reproducibility in machine learning research is becoming a lost cause due to the rise of physical AI requiring expensive hardware, unverifiable performance claims from big tech companies, and competitive incentives that discourage authors from sharing code.
Improving verifiability in AI development
OpenAI publishes a report on mechanisms to improve verifiability in AI development, addressing how stakeholders can verify organizations' claims about AI system properties and safety practices.
@houjun_liu: new method day! Trust[ing results] in ML conferences is utterly broken. Let's fix it with an algorithm! We are excited …
Researchers introduce "Training Witnesses," a system for generating "honesty" certificates of ML training data and evaluations that conference reviewers can cheaply verify and trust, aiming to fix the broken trust model in ML conferences.
Computer Science Conferences Should Require Nonrepudiable Experimental Results
This paper argues that computer science conferences should require nonrepudiable experimental results to prevent tampering and denial, and introduces K-Veritas, a reference implementation for signed reports without accessing training data.
Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D]
The article argues that training-side decontamination in AI models cannot be verified due to inherent trust and inspection issues, and proposes an evaluation-side rule to ensure reproducibility by controlling the evaluation process.