It's time to desk reject papers that don't include code that can reproduce the results [D]

Reddit r/MachineLearning News

Summary

The author, after reviewing for three major conferences, argues that papers without code to reproduce results should be desk rejected, citing that only 1 of 12 papers reviewed provided full code and 7 provided none.

As review season for NeurIPS wraps up, I have now reviewed for 3 major conferences this year. And I'm noticing a worrying trend: Out of the 12 papers I reviewed this year, only 1 provided full code (that runs the whole training pipeline from input dataset to output AUROC). 4 provided partial code with fragments of their method, but no ability to run the experiment end to end. And 7 provided no code. This is really bad for ensuring quality and reproducibility. Of the 5 papers that provided at least some code, 3 of them contained obvious bugs that completely invalidated the results. ML is highly technical and small bugs can have huge impacts if they are in the wrong place. Who knows what was going on in the remaining 7 papers. The fundamental issue here is of incentives: there is almost no cost to hiding code during the review process. Releasing code only increases odds of rejection due to reviewers finding bugs. The only way to fix this is to change the game by imposing real penalties on hiding code.
Original Article

Similar Articles

AAAI 2027 Review: No code submission? [D]

Reddit r/MachineLearning

A reviewer for AAAI 2027 expresses surprise at the low number of paper submissions with code, despite the conference's emphasis on reproducibility, and asks for opinions on whether lack of code should affect review scores.

Failure to Reproduce Modern Paper Claims [D]

Reddit r/MachineLearning

A researcher reports failure to reproduce claims from modern papers, with 4 out of 7 checked claims being irreproducible and 2 having unresolved GitHub issues, raising concerns about research quality standards.

When I reject AI code even if it works

Hacker News Top

The author explains why they often reject AI-generated code even when it works, citing reasons like inability to explain the approach, overly large diffs, premature abstractions, and reduced system reasoning, and argues for mandatory human review.

Asking Authors About Their Own Papers

Hacker News Top

An editor at Transactions on Machine Learning Research conducted meetings with authors of papers slated for desk rejection, revealing many authors could not adequately explain their own papers, raising concerns about authorship, research integrity, and the review process in AI academia.