Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System

arXiv cs.AI Papers

Summary

This position paper argues that improving machine learning peer review requires enforceable procedural safeguards and a spendable credit system, such as OpenReview Points, to incentivize good reviewing and limit submission volume.

arXiv:2608.14571v1 Announce Type: new Abstract: With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the machine learning (ML) community has one of the largest scholarly presences among all scientific fields. And yet, \textbf{almost \textit{everyone} has \textit{many} unpleasant things to share about their review experience.} Worse, there is little public space to seriously discuss, let alone debate, what makes a review system effective or how it might be improved.\quad In this position paper, we expand our discussion from two core problems: \textit{How can we reasonably limit submission volume?} and \textit{How can we incentivize good and discourage bad reviewing?} We first assess the strengths and shortcomings of existing attempts to address such problems. Specifically, we present four takes on some popular conference mechanisms and propose two alternative designs for improvement.\quad Our general position is that meaningful improvement in ML peer review won't come from polite best-practice suggestions tucked into Calls for Papers or Reviewer Guidelines: it requires \textbf{enforceable yet fine-grained procedural safeguards} paired with \textbf{a currency-like credit system (e.g., our proposed \textit{OpenReview Points})}. ML practitioners can ``earn'' such points by contributing good review practices, and ``spend'' them across one or multiple major conferences to redeem different kinds of ``perks,'' such as complimentary registration or the right to request additional review resources.
Original Article
View Cached Full Text

Cached at: 08/18/26, 09:45 AM

# Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System
Source: [https://arxiv.org/html/2608.14571](https://arxiv.org/html/2608.14571)
###### Abstract

With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the machine learning \(ML\) community has one of the largest scholarly presences among all scientific fields\. And yet,almosteveryonehasmanyunpleasant things to share about their review experience\.Worse, there is little public space to seriously discuss, let alone debate, what makes a review system effective or how it might be improved\. In this position paper, we expand our discussion from two core problems:How can we reasonably limit submission volume?andHow can we incentivize good and discourage bad reviewing?We first assess the strengths and shortcomings of existing attempts to address such problems\. Specifically, we present four takes on some popular conference mechanisms and propose two alternative designs for improvement\. Our general position is that meaningful improvement in ML peer review won’t come from polite best\-practice suggestions tucked into Calls for Papers or Reviewer Guidelines: it requiresenforceable yet fine\-grained procedural safeguardspaired witha currency\-like credit system \(e\.g\., our proposedOpenReview Points\)\. ML practitioners can “earn” such points by contributing good review practices, and “spend” them across one or multiple major conferences to redeem different kinds of “perks,” such as complimentary registration or the right to request additional review resources\.

Machine Learning, ICML

## 1Introduction

This position paper argues that peer review in machine learning \(ML\) is unlikely to improve through polite requests or optimistic guidance tucked into Calls for Papers or Reviewer Guidelines\. Fine\-grained yet enforceable procedural guardrails, combined with a spendable, across\-conference credit system, are almost mandatory for a sustainable review ecosystem\.

Machine learning has scaled faster than nearly any other scientific field in both volume and visibility\. We now have tens of thousands of paper submissions to a single conference,111The most recent[NeurIPS 2025](https://blog.neurips.cc/2025/09/30/reflections-on-the-2025-review-process-from-the-program-committee-chairs/)had 21,575 valid submissions to the main conference alone\.open\-access platforms like OpenReview that support interactive discussions, and increasingly reciprocal reviewing obligations to match supply with demand\. On paper, the ML community has everything it needs to sustain a robust yet pleasant peer\-review pipeline: we have the largest scholarly presence and the most modern review technology, all without the typical bottlenecks of paywalls, publication fees, or expensive memberships\. However, the lived reality often feels far less functional\. From cryptic or dismissive reviews to wildly inconsistent standards, frustrations with the review process appear widely shared\(Thorn Jakobsen and Rogers,[2022](https://arxiv.org/html/2608.14571#bib.bib21)\), voiced by PhD students, seasoned professors, and industry researchers alike \(e\.g\., see examples in Section[2\.1](https://arxiv.org/html/2608.14571#S2.SS1)\)\.222In the[NeurIPS 2025 Datasets & Benchmarks author survey](https://blog.neurips.cc/2025/12/05/neurips-datasets-benchmarks-track-from-art-to-science-in-ai-evaluations/), for instance, roughly 25% of respondents flagged review quality as needing improvement\.Worse, there is little to no structured way to hold bad actors accountable, nor are there incentives to encourage good actors to go the extra mile\.

In this position paper, we expand our discussion of the two core challenges we have identified:

1. 1\.How can we reasonably limit submission volume?
2. 2\.How can we incentivize good and discourage bad reviewing?

We first lay the background on why these two issues are the root causes of much unpleasantness in ML review\. Then, we assess some existing attempts to mitigate such issues as implemented in several ML conferences\. We present our takes on such measures and, finally, propose two new mechanisms:fine\-grained procedural safeguards that could be enforced at scale; and a credit system based on something we call“OpenReview Points”, which would let researchers “earn” and “spend” their reviewing efforts in tangible ways across major conferences and review cycles\.We believe such mechanisms would have a fair chance of addressing many of the aforementioned shortcomings effectively and, more importantly, are flexible enough to allow each conference to adopt its own variants\. Beyond our proposed mechanism, we engage many alternative views, where we discuss how such views are valid \(or not\) and how our proposed mechanism shall be able to take such concerns into consideration\. We conclude our paper with aRecommended Practicessection, which outlines our vision on how the first few conferences adopting a similar credit system should proceed and what aspects should be considered cautiously\. Additional materials, such as anecdotal case studies based upon real conferences, can be found in Appendix[B](https://arxiv.org/html/2608.14571#A2)\. Due to limited immediately relevant work and page limits, we present related works in Appendix[A](https://arxiv.org/html/2608.14571#A1)\.

We emphasize that our goal is not to perfect ML peer review \(as it would be unfaithful and condescending for anyone to claim so\), but to make its failures rarer, less painful, and, most importantly, more accountable and sustainable\. We are also not here to propose a specific rulebook that all conferences must follow, but rather to advocate for a promising general direction that future conference organizers can explore and adapt to their own needs\.

## 2Root Causes

### 2\.1Overwhelming number of submissions causes all kinds of challenges\.

We believe it is common knowledge that ML conferences typically receive an overwhelming number of submissions\(Kimet al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib2); Yang,[2025](https://arxiv.org/html/2608.14571#bib.bib1)\), driven by the field’s rapid growth and increasingly accessible AI\-assisted research \(Section[3\.1](https://arxiv.org/html/2608.14571#S3.SS1)\)\. Naturally, this causes all kinds of practical challenges\. From a manpower perspective, more submissions directly mean greater demand for reviewers and Area Chairs \(ACs\), which translates to a heavier workload for Senior Area Chairs \(SACs\) and, eventually, Program Chairs \(PCs\)\. With such great pressure on every aspect of the conference review system, the results are almost predictable: thinner attention per paper, more rushed triage, and greater variance in both review quality and decision outcomes\(Suet al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib3); Shah,[2022](https://arxiv.org/html/2608.14571#bib.bib20)\)\.

Moreover, practicality\-wise, most ML conferences often guarantee the right to in\-person presentation exposure once a paper is accepted\.333That said, this tradition might be undergoing a serious update, with[ICML 2026](https://icml.cc/Conferences/2026/CallForPapers)no longer requiring in\-person attendance\.This makes physical capacity restrictions come into play, directly imposing an upper limit on how many papers can be accepted\. Many borderline or acceptance\-inclined papers might be ruled out purely based on capacity constraints, a role often delegated to SACs\. But considering the number of submissions versus the number of SACs,444In[NeurIPS 2024](https://media.neurips.cc/Conferences/NeurIPS2024/NeurIPS2024-Fact_Sheet.pdf), there were 195 main\-conference SACs against 15,671 submissions, roughly 80 papers per SAC\.this kind of assignment is, by design, unreasonable and unsustainable, as few SACs would have the bandwidth or appetite to go through the content and review record of that many papers\. In practice, this pressure incentivizes shortcutting \(e\.g\., relying more heavily on numerical scores or other similar quick heuristics\), which further amplifies randomness and weakens accountability\. In fact, we have seen many SACs publicly pushing against such “force rejection for capacity” practices, as exemplified by LinkedIn posts from NeurIPS SACs[Ahmad Beirami](https://www.linkedin.com/posts/abeirami_my-thoughts-on-the-broken-state-of-ai-conference-activity-7375191053093728256-Jxdv)and[Atlas Wang](https://www.linkedin.com/posts/atlas-wang-41b726259_neurips-activity-7374496205533466624-iDjl)\.

### 2\.2Lack of oversight, feedback loops, and incentives for good actors and consequences for bad ones\.

While reviewers, ACs, and SACs have the right to provide feedback, ML conferences lack proper oversight and feedback loops\. A reviewer can act with near\-total impunity — submitting dismissive, inconsistent, or low\-effort assessments — so long as they are not extreme enough to trigger formal intervention, and little exists to correct or even surface such behavior\.This ecosystem leaves actors with few avenues to learn how to become better, let alone much incentive to go the extra mile\.Without routinized feedback, transparent metrics, or positive incentives, the system neither rewards exemplary stewardship nor deters poor practices; we elaborate on this in Section[3\.4](https://arxiv.org/html/2608.14571#S3.SS4)\.

## 3Why Existing Fixes Fall Short

### 3\.1Soft and hard submission caps offer limited help\.

With the growing body of research in the ML community\(Yang,[2025](https://arxiv.org/html/2608.14571#bib.bib1); Kimet al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib2)\)and with AI\-assisted research becoming more accessible555This is particularly evidenced by recent automatic research tooling such as Karpathy’s[autoresearch](https://github.com/karpathy/autoresearch), and by coding\-assistant plans such as[Claude Code](https://x.com/bcherny/status/2032514807388123255)moving Opus’s 1M\-token context from extra usage into the standard subscription, both lowering the access barrier to AI\-assisted work\. We can even find entirely AI\-driven work accepted at major NLP venues, e\.g\., an AI scientist getting a paper into the main track of ACL 2025 \(see[this blog](https://www.lesswrong.com/posts/LtsgfGsXpiLTSGpaW/zochi-publishes-a-paper?utm_campaign=post_share&utm_source=link)andZhou and Arel,[2025](https://arxiv.org/html/2608.14571#bib.bib5)\)\.\(Egeret al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib6)\), the volume of submissions continues to grow at a pace that far outstrips the community’s reviewing capacity\. This escalation naturally prompts discussions around mechanisms for curbing submission rates and maintaining a manageable reviewing load, where submission caps are often proposed as one of the most direct ways to reduce such volume\.

Table 1:Most\-submitted authors at ICLR 2025, counted by author positionper PaperCopilot statistics\(Yang,[2025](https://arxiv.org/html/2608.14571#bib.bib1)\)\. Each cell reports the submission count of thenn\-th most prolific author under that counting\. Notably, the top*any author*counts far exceed the top*first*or*last author*counts\.Most submitted\#1\#25\#50\#75\#100First author74333Last author34131198Any author4221171614We argue that submission caps, whether “soft” \(e\.g\., mandatory reciprocal review beyond a certain number of submissions\) or “hard” \(e\.g\., strict per\-author quotas on how many papers can be submitted\), provide, at best, marginal relief\. The core problem is not that a small set of “hyper\-prolific” lead authors are personally flooding the system with many new submissions\(Yang,[2025](https://arxiv.org/html/2608.14571#bib.bib1)\); rather, it is distributed across the community, as there is essentially no immediate downside for any author to submit unready manuscripts or endlessly recycle previously rejected work with critical flaws, straining reviewing capacity in aggregate\.

We suspect that, in practice, per\-author caps mostly trim auxiliary authors from the byline so that teams can fit under the quota\. In other words — without penalties for low\-quality submissions —submission caps may largely change who gets listed on a paper, rather than whether the paper is submitted\.While we lack direct evidence, public statistics are at least suggestive: as shown in Table[1](https://arxiv.org/html/2608.14571#S3.T1), a non\-trivial number of authors are listed on many ICLR 2025 submissions, yet very few arefirstorlastauthors on comparably many, making it unlikely that they are core contributors to all those submissions\. Under a harsh cap, such authors would likely remove their names from lower\-priority submissions rather than withhold those papers from submission altogether\. Thus, we argue that until there is a genuine negative incentive that discourages unlimited resubmission and/or rewards restraint, submission caps can only nibble at the edges of the volume problem, instead of making a significant impact at scale\.

### 3\.2Irresponsible Reviewers Care Most About Their Own Works, So Asking Nicely is Not Helpful

Bad or irresponsible review practices appear widespread, in large part because of the lack of accountability built into current conference mechanisms\. Until very recently, most ML conferences enforced no sanction against irresponsible reviewers, leaving bad practices essentially unchecked\. For a reviewer who is recruited by force \(e\.g\., through mandatory reciprocal reviewing\) and treats the assigned duty as a mere task to be finished, the things that matter most are likely their own current or future submissions\. We argue that,to effectively discourage irresponsible reviewing, some form of sanction must be enforced at the submission end\.Otherwise, conferences have little real leverage and often can only resort to asking nicely in Calls for Papers or Reviewer Guidelines; and despite the existence of some extremely thoughtful guidelines like the[ARR Reviewer Guidelines](https://aclrollingreview.org/reviewerguidelines), their effectiveness leaves much to be desired\.

Recently, starting with CVPR 2025, several ML conferences have adopted what is essentially a retaliatory desk\-rejection policy targeting irresponsible reviewers\. At CVPR 2025, its Area Chairs \(ACs\)“identified a number of highly irresponsible reviewers, those who either abandoned the review process entirely or submitted egregiously low\-quality reviews, including some generated by large language models”and ultimately issued[desk rejections for 19 otherwise accepted papers](https://x.com/CVPR/status/1894853624200863958)involving those reviewers\.

While this act marks a meaningful start to enforcing hard procedural guardrails to protect review quality and integrity,666E\.g\., at[ICML 2026](https://blog.icml.cc/2026/03/18/on-violations-of-llm-review-policies/), watermark detection flagged roughly 1% of all reviews as LLM\-generated despite a no\-LLM policy, and about 2% of submissions were desk\-rejected over such violations — both are small numbers\. We also expect such a mechanism to grow less fruitful as awareness of these watermarks spreads\.we argue thatretaliatory procedures as harsh as desk rejection can only offer marginal benefits to the conference at large, as only a few bad actors would be extreme enough to blatantly ignore direct instructions while leaving indisputable evidence behind\.This is because a harsh penalty like desk rejection is, by nature, a blunt instrument: it is too harsh to apply broadly and can only be reasonably used for the most extreme violations with verifiable signals\. In CVPR’s case, it was mostly reserved for reviewers who outright abandoned their review duties, a clean, rule\-based, verifiable breach that leaves no room for ambiguity\. Unfortunately, most reviewer problems in ML are arguably more subtle than complete negligence\. Irresponsible reviews can manifest in many forms: from thoughtless boilerplate complaints like “no theory” or “needs more experiments” applied indiscriminately to every submission, to a gross misunderstanding of basic facts and refusal to reconsider, to a raised concern with no concrete support, or even the famous[“Who is Adam?”](https://x.com/2prime_PKU/status/1948549824594485696)\. These reviews are much harder to police, but no less damaging\.

We argue that desk rejection is too coarse a penalty to handle the long tail of poor reviewing behaviors that fall short of full abandonment\. If we want meaningful deterrents at scale, we need a system that applies graduated, proportional penalties, not just all\-or\-nothing rulings\. The fact that the CVPR 2025 procedure only resulted in 19 desk rejections suggests that the irresponsible review issue in the ML community is far from resolved by adopting this retaliatory desk rejection policy alone; more enforceable yet fine\-grained procedural safeguards are necessary to handle the wide spectrum of irresponsible review practices\. In other words,desk rejections are like felony charges, but we also need misdemeanors, infractions, and everything in between to moderate a proper review community\.

### 3\.3100% Mandatory Reciprocal Reviewer Recruitment is a Slow\-Acting Poison

To keep up with rising submission counts, many ML conferences, such as the most recent[EMNLP 2025 \(ARR May\)](https://aclrollingreview.org/incentives2025), now rely on 100% reciprocal reviewer recruitment: every eligible author777Where such eligibility is determined by having at least a few published works at certain recognized ML/AI venues\.must also be available to review, except in extreme circumstances like parental leave\. On paper, this sounds fair: if one wants to publish, one should also contribute to the review pool\. But in practice, we argue it is a slow\-acting poison that comes at the cost of review quality\.

The policy assumes that all eligible authors of every accepted paper are both capable and willing to provide thoughtful reviews\. That assumption does not hold across the ML community: many authors may have played only auxiliary roles or contributed as expert consultants on highly specialized components\. They are therefore ill\-suited to reviewing general\-purpose ML submissions\. Worse, mandatory review removes the ability for people to decline, even when they know they cannot meaningfully contribute due to sensible \(but non\-medical\) emergencies and circumstances\. Once in the reviewer pool, conferences often allow very limited flexibility for exemption\. For instance, AAAI 2026 instructed their reviewers to“do your best”even if the assigned paper is outside their area of expertise\.

We argue that such enforced cultures would likely result in a series of rushed, templatized, or often disengaged reviews/meta\-reviews\. We suspect this is a meaningful contributor to much of the unpleasantness in ML peer review:when recruited by force and with limited motivation and bandwidth, reviewers are likely to invest only the minimum effort, as their main drive becomes fulfilling their assigned duties so their own submitted work is not desk\-rejected\.Perhaps recognizing this, many later conferences appear to be on the same page with us \(in terms of realizing the negative side of mandatory reciprocal review\): starting from[ARR July 2025](https://aclrollingreview.org/exemptions2025), technically qualified authors may request an exemption from review duty on a case\-by\-case basis if they find themselves lacking the relevant expertise\. That said, broader allowance for exemption \(e\.g\., lack of bandwidth\) has yet to be widely adopted\.

We find it ironic that a mechanism designed to distribute the workload ends up degrading its quality\. Yet, if everyone could opt out with no consequences, the system would collapse under the sheer volume of submissions\. The reasonable middle ground, then, is to allow sensible opt\-outs while still holding authors accountable for their share of the reviewing load — ideally alongside a way to mobilize additional willing reviewers at will to cover any resulting gaps\. We expand on both in Section[3\.4](https://arxiv.org/html/2608.14571#S3.SS4)and Section[4](https://arxiv.org/html/2608.14571#S4)\.

### 3\.4Helpful Implicit Expectations Often Go Unmet

As we previously teased in Section[2\.2](https://arxiv.org/html/2608.14571#S2.SS2), conference processes implicitly assume that reviewers, ACs, and SACs will self\-initiate best practices, such as timely calibration, substantive internal discussions, careful revision after rebuttals, and principled follow\-ups, to ensure that informed decisions are made for each submission\. We argue that in reality,helpful practices like internal reviewer discussions rarely happen effectively at scale, because reviewers are seldom incentivized to “go the extra mile\.”

These practices bring little recognition, credit, or accountability for the extra coordination and time they require, so under deadline pressure most actors would likely default to the minimum required\. Steps like internal reviewer discussion thus tend to be perfunctory or skipped, and decision quality can degrade, not for lack of guidance, but for lack of aligned incentives to act on it\.

As[Eli Goldratt](https://www.duperrin.com/english/2014/06/23/quote-tell-how-you-measure-and-i-will-tell-you-how-i-will-behave/)put it:“Tell me how you measure me and I will tell you how I will behave”— when the only thing measured is whether one’s assigned duties are formally complete, that is all most will optimize for\. We arguethe key to better review is therefore not to compel more participation, but to find and motivate those with the bandwidth and genuine enthusiasm to engage well — including capable scholars who currently sit out the process entirely\.This is admittedly difficult, as review remains, for now, a largely thankless activity; we explore how to incentivize in Section[4](https://arxiv.org/html/2608.14571#S4)and Section[5\.2](https://arxiv.org/html/2608.14571#S5.SS2)\.

## 4Our Proposal: Fine\-Grained Procedural Guardrails with a Currency\-Like Incentive System

To make meaningful progress in peer review reform, we argue that two ingredients are essential:enforceable procedural safeguards at different granularities, andan incentive structure that rewards good\-faith participation while offering flexibility\.We propose a system based on a community\-wide, cross\-conference\-supported economy called“OpenReview Points”, mainly due to the widespread adoption of OpenReview, which is also well positioned to track such balances\. This section outlines the basic principles of such a system, discusses potential enforcement strategies, and explores the feasibility of a conference\-wide credit market that could finally provide conference organizers both the “stick” and the “carrot” they currently lack\.

### 4\.1OpenReview Points: A Currency\-Like Economy Enabling Flexible Options

The current review ecosystem operates on the honor system: a reviewer is expected to perform review duties diligently and hope that others will do so as well\. However, we argue that optimistic hoping is not a system\. To instill accountability, we propose a credit system, giving contributors to the review pipeline something to earn, spend, and track\.

Under our proposal, ML practitioners would accumulate OpenReview Points based on their contributions to the community\. For instance, in the context of reviewing, completing a standard review might earn 1 point, helping with an emergency review might earn 2, and being recognized as an “outstanding reviewer” could grant an additional 3 points\.888We emphasize that all point values mentioned in this section are intuitively assigned for hypothetical purposes\. A real OpenReview Point\-based economy would require significantly more sophisticated balancing, subject to each conference’s own preferences\. More on such specifications in Section[6](https://arxiv.org/html/2608.14571#S6)\.Once earned, OpenReview Points could be spent to gain access to certain “perks” and privileges\. For example:

- •A reviewer can spend 5 points to opt out of an assigned review duty\.
- •An author can spend 10 points to exempt a co\-author from their reciprocal reviewing obligation\.
- •An author can spend 50 points to request an additional expert reviewer in the case of a highly controversial or borderline decision\. In the meantime, a reviewer/AC can take this job and earn those 50 points\.
- •An author can spend 100 points to redeem free registration\.

This economy introduces direct incentives: if one contributes meaningfully, one gains flexibility and optionality\. If one does not, one’s publishing privileges will begin to shrink\. We emphasize that,because credits can be flexibly awarded or deducted for almost any behavior, the system gives organizers a broad design space to influence community behavior in ways never possiblewith just blunt\-force policies like universal reciprocal reviewing or desk rejections\.

On the*reward*side, for instance, much of our work argues that there is a lack of incentive for actors to “go the extra mile,” even when such effort can be immensely helpful \(e\.g\., as anecdotally demonstrated in Appendix[B](https://arxiv.org/html/2608.14571#A2)\)\. With point incentives, however, such “extra miles” can be explicitly encouraged: reviewers may become more willing to initiate and engage in internal discussion, and ACs more willing to investigate, because exemplary actions can now be potentially rewarded\. Similarly, the lack of a reviewer feedback loop, discussed in Section[2\.2](https://arxiv.org/html/2608.14571#S2.SS2), can be mitigated by awarding points to authors who are willing to provide detailed reviewer feedback, which reviewers can consult to improve their future practices\.

On the*penalty*side, as noted in Section[2\.1](https://arxiv.org/html/2608.14571#S2.SS1), many reviewing issues stem from inflated submission counts\.One way to mitigate this is to require a small and refundable “submission fee” \(e\.g\., 10 OpenReview Points\) per paper\.If the paper is accepted or meets a reasonable “fair attempt” bar, the points are refunded; otherwise, they are forfeited\. This soft deterrent discourages unready submissions by linking low\-quality or premature work to a corresponding reduction in future publication privileges\. We find this use of the credit system particularly appealing, as it directly targets the core issue of submission quantity\. It also rests on a relatively reliable signal: the[NeurIPS Consistency Experiments](https://blog.neurips.cc/2021/12/08/the-neurips-2021-consistency-experiment/)found that when two independent reviewer teams assess the same papers, their agreement is strongest on clear rejections\. In other words, while review outcomes can be quite noisy for borderline or even spotlight\-worthy work, they are far more consistent at flagging low\-quality submissions\. We envision that such cases can be reliably identified by signals like uniformly low ratings from all reviewers — exactly what a refundable submission fee is meant to discourage\.

While we have proposed several specific policies in this section,we emphasize that we do not argue for the enforcement of any specific rule\. Rather, we argue that a credit system would grant every participant in the ML community far greater flexibility in how they interact with the review process\. While we fully expect friction or disagreement regarding any particular rule or redemption policy, we believe it would be difficult to argue against the utility of having a credit system*at all*, since it makes sense for different conferences to carve out their own rules to cater to their own communities\.

### 4\.2Making the Credit System Enforceable: A Four\-Part Defense

A credit system is only as useful as it is enforceable\. Beyond awarding points for good\-faith contributions, organizers need levers to deter malicious gaming behaviors, such as fraud, point farming, and bulk low\-effort reviewing\. Granted the vast scale of potential abuse attempts, these levers should operate at a finer granularity than the blunt instrument of desk rejection\. Of course, no system is immune to abuse — even real\-world economies with actual laws and tangible consequences face persistent bad actors\. Nonetheless, we argue that, although imperfect, a credit system is far better positioned than the status quo, precisely because organizers can award and deduct points to impose countermeasures and realign incentives\. We group these countermeasures into four reusable primitives that conferences can mix and match:

- •Duty Delegation:Restrict certain actions to specific roles\. For example, only a paper’s own \(lead\) authors may spend points to request an additional expert reviewer, and submission\-fee\-like charges fall on lead authors rather than auxiliary ones\. This prevents charges and privileges from being offloaded onto, or routed through, less accountable parties\.
- •Upper Limits:Cap how often an action may be taken per cycle, e\.g\., how many papers one may review for points, or how many additional or emergency reviewer requests one may file\. This bounds low\-effort point farming and bulk abuse\.
- •Dynamic Pricing:Make repeated use of an action or perk progressively more expensive, e\.g\., beyond a threshold, each submission or each successive additional\-reviewer request would cost more points\. This reserves scarce resources for the cases that matter most\.
- •Voting\-Based Penalties and Awards:Let peer reviewers and the AC issue fine\-grained, peer\-driven judgments, deducting points for low\-quality reviews that peers and the AC confirm, and awarding points for exemplary ones \(e\.g\., an “outstanding reviewer” bonus\)\.

These primitives compose against concrete threats\. For instance, bulk outsourcing of reviews for points is curbed byupper limits\(a few reviews per actor\) andvoting\-based penalties\(low\-quality reviews lose points\), making writing good reviews the easier path forward\. Schemes that funnel points to a dedicated “point person” to be listed as a coauthor across many papers to buy perks are blunted byduty delegation\(has to be lead authors\) anddynamic pricing\(increasing costs for more perk purchases\), and, absent person\-to\-person transfers, are simply hard to pull off at scale\. Similarly, spamming the “additional expert reviewer” option is also contained by per\-cycle caps and escalating prices\.While these attack angles represent only a fraction of potential threat models, we argue that the primitives themselves are flexible enough to be adapted to a wide range of abuse attempts, and that the ability to deploy such countermeasures is a key advantage of a credit system over the status quo\.

Crucially, much of this rests on one principle: points must be earned through labor, not bought\. We therefore strongly advocate forbidding person\-to\-person point transfers and any money\-to\-points conversion\. The moment points can be purchased, the system risks degenerating into pay\-to\-play, which is plausibly worse than the status quo\. The voting\-based lever in particular raises concerns about false positives and politicization, which we address in Section[5\.4](https://arxiv.org/html/2608.14571#S5.SS4)\.

## 5Alternative Views

While we advocate for enforceable procedural safeguards and a credit system, we recognize that not everyone will agree with this approach\. Below, we discuss several alternative perspectives and respond to their concerns\.

### 5\.1“A credit system gamifies peer review and adds needless bureaucracy\.”

A common objection is that introducing a credit system risks gamifying the review process, turning what should be a scholarly, community\-driven responsibility into a transactional system; a related worry is that tracking points and adjudicating review quality would introduce too much bureaucracy\. Both concerns are understandable, but neither names anything fundamentally new: peer review is already governed by incentives, such as reviewing others’ submissions in return for having one’s own work properly reviewed, and conferences already invest massive effort coordinating thousands of reviews and rebuttals with different rules and workflows\. A credit system creates little out of thin air; it largely formalizes and aligns those incentives with the broader health of the ecosystem, and makes the existing effort fairer, more consistent, and more sustainable\. That said, we concede that the full system may be too heavy to implement at once, which is why we favor a gradual rollout of changes — a “soft landing” for existing community members, as we detail in Section[6\.2](https://arxiv.org/html/2608.14571#S6.SS2)\.

### 5\.2“A credit system favors the privileged, and review duty exemptions would lose reviewers\.”

Some may argue that a credit system will disproportionately benefit researchers with more time, institutional support, or prior connections, allowing them to “buy” their way out of responsibilities \(e\.g\., being exempt from review duties\) while leaving others to “pick up the slack\.” This is a legitimate concern, but in our design, points are earned through labor, not status\. There is no “premium tier of citizen,” only accumulated contributions through hard work\. While it is still true that researchers with strong support will likely have more opportunities to contribute \(as they are not otherwise occupied by some chores\), their “surplus contributions” are still a net gain to the community\.

While exemptions from review duties might indeed cost us some reviewers, it is worth asking whether those willing to pay a high price to opt out are producing quality reviews \(if kept by force\), and whether they have the bandwidth to stay engaged with the authors\. We tend to believe such answers lean toward the negative, and argue that a better alternative might be to just let them be exempted, then utilize the collected points to incentivize reviewers who do have the bandwidth and motivation in this particular cycle\.

We would also strongly advocatemobilizing researchers who are not main authors to participate more in the review pipeline, as they likely have better bandwidth \(since they are not under the pressure of author deadlines\), and their reviews will not be as affected by feedback on their own submitted work\. Under the current system, there is little incentive for researchers to do so, as most reviewers are recruited by mandatory reciprocity, which no longer applies without being an author\. Our proposed credit system might provide them with a strong incentive to participate, as they can earn points to enrich their publication privileges; and specifically, have the option to spend such points to be exempt from reviewer duties when they are submitting lead\-authored work, granting themselves wider bandwidth as authors when they are under rebuttal pressure\.

### 5\.3“What about early\-career researchers?”

Much like how games onboard novice players and companies onboard new employees, the point\-hosting platform could grant a baseline amount of points to first\-time contributors, perhaps along with a protection period, giving them enough time and capital for trial\-and\-error\. Concretely, such a protection period \(e\.g\., the first three to five months, or one’s first few conference cycles\) might allow a limited number of initial submissions to be fully refunded and non\-extreme penalties to be softened, so that newcomers can learn from early missteps without those missteps derailing a nascent career\. A central theme of the credit system is that a single mistaken penalty is not catastrophic; the same spirit should carry over to onboarding, where we can afford to be more forgiving\.

We further note that a credit system can reward “good acts” well beyond reviewing itself\. For instance, conference organizers might host onboarding workshops, or pair early\-career researchers with more experienced “research buddies” who mentor them on submission compliance, review writing, and community norms\. Rewarding such mentorship can be far more constructive than relying on point awards and penalties alone, yet would be hard to establish without a fine\-grained credit system in place\.

### 5\.4“Voting\-based penalties will be abused or weaponized\.”

The voting\-based lever introduced in Section[4\.2](https://arxiv.org/html/2608.14571#S4.SS2)naturally raises the concern that it could be misused, weaponized in borderline cases, or influenced by interpersonal bias\. This concern is legitimate\. However, as argued in Section[3\.2](https://arxiv.org/html/2608.14571#S3.SS2), the reviewing problems that matter most are the subtle ones that desk rejection is too blunt to police, which is exactly why a finer\-grained, voting\-based penalty is needed: if the authors report a reviewer and that reviewer’s peers, along with the area chair, agree that a review is unacceptably low in quality or that the reviewer engaged in unprofessional conduct, the reviewer could face penalties ranging from a warning to graduated point deductions\.

Such a system does introduce the possibility of false positives, but we argue this is largely acceptable\. First, with guardrails such as close\-unanimous \(if not fully unanimous\) agreement among the other reviewers on the paper, area chair confirmation, and an appeal mechanism, the practical false\-positive rate can be kept low\. Second, a point deduction is far less extreme than desk rejection or a submission ban, so even a wrong call is unlikely to cause severe or irrecoverable harm\. Most importantly, the current system has the opposite problem: a nearly 0% true\-positive rate, since no matter how badly a reviewer behaves, there are essentially zero consequences short of automatically verifiable atrociousness\. The status quo, in our view, is worse\.

No penalty system will ever be perfect, but the absence of one leaves little room for improvement\. We argue that a small risk of overcorrection is a worthwhile price for finally holding peer review to a higher standard\. An even lower\-risk alternative is to lean on the*award*side of the same lever, granting credits to ACs and reviewers who provide detailed feedback on their peers’ reviews, enabling a positive feedback loop where actors have a channel to learn and improve\. We discuss such recommended practices in Section[6](https://arxiv.org/html/2608.14571#S6)\.

### 5\.5“Restricting point transfers is undemocratic\.”

It is worth first being explicit about why unrestricted transfer would break the system\. Points are meant to certify that a specific person performed a specific service; once they can be gifted or sold, a researcher’s privileges no longer reflect their own contributions, and the “labor for perks” principle the whole design rests on \(Section[4\.2](https://arxiv.org/html/2608.14571#S4.SS2)\) loses its meaning; transferability would also neutralize much of our four\-part defense, as a well\-resourced actor could simply pool points from others to absorb the rising costs ofdynamic pricing\. A transfer channel is, moreover, the natural on\-ramp to a black market, leading to a money\-to\-points and pay\-to\-play outcome all reasonable scholars would want to avoid, and plausibly worse than what we have today\.

Given this, one might still object that strictly regulating transfers is paternalistic, much like tightly regulating the transfer of money\. We find the analogy imperfect: points are not private wealth but a record of personal community service, closer to a professional credential, a citation, or an authorship than to money\. One cannot sell a degree or a reviewing record, and few would call that undemocratic\. After all, non\-transferability is what gives the credential its meaning, not a liberty taken away\. The restriction on person\-to\-person transfer is also not a blanket ban on all sharing/pooling operations: team\-based redemptions \(e\.g\., coauthors jointly spending points\) remain sensible, and each conference can set its own boundaries\. What we resist is decoupling points from the labor that earned them\.

### 5\.6“Will points simply inflate as the system scales?”

A natural concern at scale is not the raw submission count but the points themselves: as the system runs across many cycles and venues, the total points in circulation could grow until they inflate and lose value as either a deterrent or a reward\. This is a familiar problem for real\-world point economies, and we can borrow their remedies\. Credit\-card and loyalty\-point programs routinely curb inflation and hoarding through point expiration and time\-limited, discounted redemption windows that nudge members to spend rather than stockpile, giving organizers meaningful influence over the effective point supply\. A second, related worry is that the system could be gamed at scale through cheap, bulk, or AI\-driven point farming; this is contained by the four\-part defense elaborated in Section[4\.2](https://arxiv.org/html/2608.14571#S4.SS2)\. None of this makes the system immune to abuse at scale, but it does hand organizers concrete moderation levers that current conference mechanisms simply lack\.

### 5\.7“A cross\-conference reciprocity must exist first\.”

One clear and legitimate criticism of our credit system is that, for it to work to its full potential, multiple major conferences must adopt it\. Granted, conferences like ICML, NeurIPS, and ICLR rarely collaborate explicitly, so this prerequisite is admittedly hard to meet\. However, we argue that there are ML conferences well\-positioned to adopt such practices: for example, the ARR series of conferences has long implemented cross\-conference measures \(e\.g\., submission bans from the next ARR cycle\), as experimented with in[EMNLP 2025](https://2025.emnlp.org/reviewer-policies/), making them more openminded to adopting similar measures\. Further, even if the credit system is per\-conference, it can still function better than nothing; it is just that features requiring accumulated effort may be harder to activate and experiment with\.

### 5\.8“Who pays for the cost?”

Our mechanism does not require conferences to offer more complimentary registrations than they already do\. Major ML conferences already grant free registration to top reviewers\.999NeurIPS 2025, for instance, offers this to an estimated[1,900\+ individuals](https://neurips.cc/Conferences/2025/ProgramCommittee#top-reviewer), and[ICML 2026](https://x.com/icmlconf/status/2049919247065694669)similarly awarded free registration to its top 25% of reviewers\.With conferences increasingly open to virtual attendance, the material cost per registration is further reduced\. As discussed in Section[6](https://arxiv.org/html/2608.14571#S6), point thresholds can be calibrated to match existing resource constraints\.

The key difference is simply how these perks are allocated: instead of relying on AC discretion, we can rank reviewers by credits earned across multiple papers \(or even venues\), providing a more objective way to identify top contributors — and arguably giving people more freedom in how they spend their points\.

### 5\.9“Just enforce monetary cost per submission\.”

Some ML conferences have begun experimenting with monetary submission fees to curb submission volume\.101010[IJCAI\-ECAI 2026](https://2026.ijcai.org/ijcai-ecai-2026-call-for-papers-main-track/), for instance, charges $100 per paper from the second submission onward\.We argue that points are preferable to money for one practical reason: if the monetary threshold is too low, it becomes meaningless as a deterrent; if too high, it excludes researchers without strong institutional support — and blocking the accessibility of science is of course undesirable\. Points, by contrast, can be earned by contributing to the community \(e\.g\., by reviewing papers\), making them a fairer currency that rewards effort rather than financial privilege\. The restrictiveness of point transfers also makes them much more resistant to gaming and abuse, as we discussed in Section[5\.5](https://arxiv.org/html/2608.14571#S5.SS5)\.

## 6Recommended Practices / Call to Action

As emphasized throughout this position paper, our goal is not to promote a single, prescriptive rulebook that every conference must follow, but to advocate for a flexible framework that can adapt to different conference idiosyncrasies\. Under our credit system, conference panels and authors are akin tostore owners and customers: the panels decide what goods are offered and at what price, while the customers decide where and how they wish to spend their money\.However, we recognize that without concrete discussion of how such a system might operate, adoption could invite resistance or, worse, chaos\. This section therefore offers practical guidance on how the first few conferences adopting a credit\-like system might proceed\.

### 6\.1Heavy on existing perks

We believe that early adopters should anchor their credit\-like system around perks ML conferences already offer \(e\.g\., complimentary registration, emergency reviewer invitations\), rather than immediately introducing entirely new perks \(e\.g\., exemption from review duties, invitation of extra reviewers\)\. Staying with existing perks provides two immediate benefits: 1\) Because these perks are already part of established workflows, redistributing them according to the credit system \(e\.g\., ranking reviewers by awarded points per conference cycle rather than relying on an AC’s subjective judgment\) keeps overall impact bounded\. If the new distribution turns out problematic, its effects are still confined to the known scope of these already\-tested perks; whereas brand\-new perks introduce unknown risks\. 2\) We can directly compare credit\-system conferences’ key metrics against their historical data\. Any improvement or decline is then more likely attributable to the credit system \(or its specific implementation\), rather than being confounded by the introduction of new perks\. Reusing the same perks under credit\-based allocation also yields ablated data on how the credit system behaves in practice — a baseline for testing new perks later\.

### 6\.2Gentle and gradual rollout of new perks, potentially with one\-off tests

When launching new perks, it is best to roll them out gently and gradually rather than all at once\. This approach offers a clean testbed to monitor each perk’s contribution and reduces the information load on all involved parties, who will need time to adjust\.

Observant readers may notice that some of our proposed policies \(e\.g\., free registration, the right to request additional reviewers, or refundable submission fees discussed in Section[4](https://arxiv.org/html/2608.14571#S4)\) require a relatively long\-term accumulation of points before they become useful, stretching the evaluation horizon\.

A simple way to expedite early evaluation is to introduce one\-off tests\. For example, in addition to point awards, conferences might grant top point\-earners a one\-off right to request an additional reviewer to help resolve borderline cases, with this right expiring at the end of the conference cycle\. This allows organizers to directly observe whether such redeemable incentives meaningfully help, and to what extent these improvements propagate through the reviewer–AC pipeline\. This kind of fast feedback might help conference organizers trim unhelpful policies quickly and enable faster iterations of rule sets\.

### 6\.3Determining point values for contributions and perks

One reason we did not specify exact point values for different contributions \(e\.g\., how many points an emergency review should yield\) is that we currently lack the empirical data needed to set these responsibly\. As a general guideline, we believe it is reasonable to treat the completion of one regular review duty as the base “unit price” of this ecosystem\.

Rather than arguing directly about how many points each contribution “should” receive, we propose working backwards from the perks: estimate how many free registrations or similar rewards a conference can offer, determine what percentile of contributors this represents, and calibrate point values accordingly\.

### 6\.4Track key metrics and publicize such statistics

Finally, for a credit system to have a lasting impact, conferences must make informed decisions about which rules to adopt and at what point\-values\. Such decisions require cross\-conference consistency\. If, for example, NeurIPS values its perks at 10x ICLR’s level for no meaningful reason, the ecosystem loses the interoperability we envision\. Thus, each conference should monitor key metrics and publish these statistics as part of their post\-conference fact sheets\.

For example: If a new rule is implemented, do we observe increased interaction among reviewers and ACs? Do ACs report that these additional exchanges help them make more confident decisions? These statistics and reports can form the foundation for iterating toward a better implementation of the credit system, and can serve as a strong signal, even an advertisement, encouraging more conferences to adopt a shared credit currency\.

## 7Limitations

A proposal is only faithful if it also highlights its main weaknesses\. In our case,the lack of numerical experiment resultsis one\. However, we believe discussions about review mechanisms are most meaningful in the hypothetical space, since there is no way to rewind history and A/B\-test different conference measures\. LLM\-powered simulation is a natural alternative, but at the scale of multiple conferences, we find it adds little value: too many moving factors compound, with too few publicly available metrics to ground them, making even crude simulations hand\-wavy and easy to rig absent real baselines\. That said, recognizing the desire for some anchoring to real conferences, we share three case studies \(from top ML venues where we served as reviewers\) in Appendix[B](https://arxiv.org/html/2608.14571#A2)as a compromise between realism and faithfulness\.

A second limitation is in terms of scope, asour proposal is largely*corrective*: it polices and rationalizes an already\-large pool of submissions,rather than addressing the root cause behind that volume\.The explosion in ML submissions has turned review by a committee of experts into reliance on a committee of authors \(often recruited by force\), and we don’t see this fundamental scaling challenge being addressed even with a credit system\. This work thus leans toward the*reactive*side of the problem — making review fairer and more sustainable given the volume we already face — rather than*proactively*rethinking the publication system itself\. Some of our measures do curb submission volume, but they remain closer to policing than to reducing the innate drive behind it; we regard the latter as a deeper question deserving dedicated study, and encourage future work in that direction\.

## Acknowledgements

I thank thereviewers and area chairfor their concrete and constructive feedback, and in particular for the many interesting threat models they raised, which shaped the four\-part defense of the credit system\. Many of the added alternative views were also inspired by their comments\. Due to page limitations, I unfortunately could not include a diagram of the proposed pipeline in the main text as promised, but I will make sure one is available on my poster\.

I would also like to thankMoshe Y\. Vardifor his insightful comments on this paper, many of which I incorporated into the final version\. Beyond individual fixes, Moshe pointed me to several relevant works I had overlooked, letting me attribute credit more properly — fitting, perhaps, for a paper that itself advocates crediting contributions\. His observation that this proposal is more corrective than root\-solving also prompted the scope discussion now in Section[7](https://arxiv.org/html/2608.14571#S7)\.

I am also grateful toBuxin Su— one of the core authors of the Isotonic Mechanism series of works\(Suet al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib3),[2026](https://arxiv.org/html/2608.14571#bib.bib22)\)pioneered by Weijie Su\(Su,[2021](https://arxiv.org/html/2608.14571#bib.bib13)\)— for an online discussion in which he shared feedback on my initial ideas\. Many of this work’s arguments were also challenged and polished byMinghao YanandZhaozhuo Xuwhen we drove back from NeurIPS 2025\. I thank them for their company and sharp comments, and for letting me stay a passenger princess throughout the ride\.

This work is written for the love of the game, out of a stubborn hope that ML review can be better: that one day a system — this one, something similar, or a different but better one — might provide an improved review experience for all actors in the pipeline\. The nature of this work is not technical, and its development was not funded by any organization or grant\. That said, I would like to thank theDepartment of Computer Science at Rice Universityfor providing a supportive research environment\. On a similar note, much of the writing was done while I was visitingMATS ResearchandConstellation \(with Anthropic Fellows Program\), and I thank these organizations for their hospitality and for the chance to bounce ideas off my peers\. In particular, I thankJonathan Michalaat MATS for being a wonderful research manager and for enduring much of my ranting and rambling about peer review\.

The views expressed in this paper are solely my own and do not reflect the views of any organization or any individual mentioned\.

## References

- A\. Beygelzimer, Y\. N\. Dauphin, P\. Liang, and J\. W\. Vaughan \(2023\)Has the machine learning review process become more arbitrary as the field has grown? the neurips 2021 consistency experiment\.arXiv preprint arXiv:2306\.03262\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px3.p1.1)\.
- C\. Cortes and N\. D\. Lawrence \(2021\)Inconsistency in conference peer review: revisiting the 2014 neurips experiment\.arXiv preprint arXiv:2109\.09774\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px3.p1.1)\.
- S\. Eger, Y\. Cao, J\. D’Souza, A\. Geiger, C\. Greisinger, S\. Gross, Y\. Hou, B\. Krenn, A\. Lauscher, Y\. Li,et al\.\(2025\)Transforming science with large language models: a survey on ai\-assisted scientific discovery, experimentation, content generation, and evaluation\.arXiv preprint arXiv:2502\.05151\.Cited by:[§3\.1](https://arxiv.org/html/2608.14571#S3.SS1.p1.1)\.
- M\. Francia, E\. Gallinucci, and M\. Golfarelli \(2026\)From volunteerism to duty: reforming peer review with tokens\.Commun\. ACM69\(6\),pp\. 78–85\.External Links:ISSN 0001\-0782,[Document](https://dx.doi.org/10.1145/3770921)Cited by:[1st item](https://arxiv.org/html/2608.14571#A1.I1.i1.p1.1),[4th item](https://arxiv.org/html/2608.14571#A1.I1.i4.p1.1),[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px2.p1.1),[footnote 13](https://arxiv.org/html/2608.14571#footnote13)\.
- S\. Fund and G\. Dyke \(2024\)Recognition and reward in peer review: the reviewercredits vision\.Minerva Cardiology and Angiology\.External Links:[Document](https://dx.doi.org/10.23736/S2724-5683.23.06487-6)Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px2.p4.1)\.
- A\. Y\. Gasparyan, A\. N\. Gerasimov, A\. A\. Voronov, and G\. D\. Kitas \(2015\)Rewarding peer reviewers: maintaining the integrity of science communication\.Journal of Korean medical science30\(4\),pp\. 360–364\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px2.p2.1),[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px2.p3.1)\.
- A\. Goldberg, I\. Stelmakh, K\. Cho, A\. Oh, A\. Agarwal, D\. Belgrave, and N\. B\. Shah \(2025\)Peer reviews of peer reviews: a randomized controlled trial and other experiments\.PloS one20\(4\),pp\. e0320444\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px3.p1.1)\.
- S\. Jecmen, N\. B\. Shah, F\. Fang, and L\. Akoglu \(2025\)On the detection of reviewer\-author collusion rings from paper bidding\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px4.p1.1)\.
- J\. Kim, Y\. Lee, and S\. Lee \(2025\)Position: the AI conference peer review crisis demands author feedback and reviewer rewards\.InForty\-second International Conference on Machine Learning Position Paper Track,Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px4.p1.1),[§2\.1](https://arxiv.org/html/2608.14571#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.14571#S3.SS1.p1.1)\.
- R\. Liu and N\. B\. Shah \(2023\)Reviewergpt? an exploratory study on using large language models for paper reviewing\.arXiv preprint arXiv:2306\.00622\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px4.p1.1)\.
- D\. Mimno and A\. McCallum \(2007\)Expertise modeling for matching papers with reviewers\.InProceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,KDD ’07,New York, NY, USA,pp\. 500–509\.External Links:ISBN 9781595936097,[Document](https://dx.doi.org/10.1145/1281192.1281247)Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px4.p1.1)\.
- Publons \(2018\)Global state of peer review\.Note:Clarivate AnalyticsCited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px2.p4.1)\.
- V\. Rao, J\. Payan, A\. McCallum, and N\. B\. Shah \(2025\)ML researchers support openness in peer review but are concerned about resubmission bias\.External Links:2511\.23439Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px4.p1.1)\.
- A\. Rogers and I\. Augenstein \(2020\)What can we do to improve peer review in NLP?\.InFindings of the Association for Computational Linguistics: EMNLP 2020,T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 1256–1262\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.112)Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p4.1)\.
- N\. B\. Shah \(2022\)Challenges, experiments, and computational solutions in peer review\.Communications of the ACM65\(6\),pp\. 76–87\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p3.1),[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p4.1),[§2\.1](https://arxiv.org/html/2608.14571#S2.SS1.p1.1)\.
- B\. Su, J\. Zhang, N\. Collina, Y\. Yan, D\. Li, K\. Cho, J\. Fan, A\. Roth, and W\. Su \(2025\)The icml 2023 ranking experiment: examining author self\-assessment in ml/ai peer review\.Journal of the American Statistical Association0\(0\),pp\. 1–12\.External Links:[Document](https://dx.doi.org/10.1080/01621459.2025.2510006),https://doi\.org/10\.1080/01621459\.2025\.2510006Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p2.1),[§2\.1](https://arxiv.org/html/2608.14571#S2.SS1.p1.1),[Acknowledgements](https://arxiv.org/html/2608.14571#Sx1.p3.1)\.
- B\. Su, J\. Zhang, N\. Collina, Y\. Yan, D\. Li, K\. Cho, J\. Fan, A\. Roth, and W\. Su \(2026\)Rejoinder: the icml 2023 ranking experiment: examining author self\-assessment in ml/ai peer review\.arXiv preprint arXiv:2605\.25172\.Cited by:[Acknowledgements](https://arxiv.org/html/2608.14571#Sx1.p3.1)\.
- W\. J\. Su \(2021\)You are the best reviewer of your own papers: an owner\-assisted scoring mechanism\.InAdvances in Neural Information Processing Systems,A\. Beygelzimer, Y\. Dauphin, P\. Liang, and J\. W\. Vaughan \(Eds\.\),Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p2.1),[Acknowledgements](https://arxiv.org/html/2608.14571#Sx1.p3.1)\.
- T\. Thorn Jakobsen and A\. Rogers \(2022\)What factors should paper\-reviewer assignments rely on? community perspectives on issues and ideals in conference peer\-review\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 4810–4823\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.354)Cited by:[§1](https://arxiv.org/html/2608.14571#S1.p2.1)\.
- J\. Wu, H\. Xu, Y\. Guo, and W\. Su \(2023\)A truth serum for eliciting self\-evaluations in scientific reviews\.arXiv preprint arXiv:2306\.11154\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p2.1)\.
- J\. Yang \(2025\)Position: the artificial intelligence and machine learning community should adopt a more transparent and regulated peer review process\.InForty\-second International Conference on Machine Learning Position Paper Track,Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px3.p1.1),[§2\.1](https://arxiv.org/html/2608.14571#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.14571#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.14571#S3.SS1.p2.1),[Table 1](https://arxiv.org/html/2608.14571#S3.T1)\.
- Y\. Zhang, F\. Yu, G\. Schoenebeck, and D\. Kempe \(2022\)A system\-level analysis of conference peer review\.InProceedings of the 23rd ACM Conference on Economics and Computation,pp\. 1041–1080\.Cited by:[Appendix A](https://arxiv.org/html/2608.14571#A1.SS0.SSS0.Px1.p4.1)\.
- A\. Zhou and R\. Arel \(2025\)Tempest: autonomous multi\-turn jailbreaking of large language models with tree search\.arXiv preprint arXiv:2503\.10619\.Cited by:[footnote 5](https://arxiv.org/html/2608.14571#footnote5)\.

## Appendix ARelated Works

#### Proposal of conference mechanisms for better ML review quality

To the best of our knowledge, few published works have addressedhow to improve ML review quality by proposing new conference mechanisms\.The closest work to ours might beKimet al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib2)\), where the authors advocate establishing a feedback loop to reviewers and promoting reviewer rewards, two points that we also share\. Specifically,Kimet al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib2)\)highlights that if we allow authors to rate reviewers, those ratings will almost always be heavily influenced by the specific strengths and weaknesses \(and, by extension, the scores\) listed by the reviewer\(Goldberget al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib9)\)\. Reviewers who fairly rate papers negatively may be subjected to unfair retaliatory ratings from authors\. To address this,Kimet al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib2)\)suggests a two\-stage reveal: authors first read only the reviewer\-written summary and strengths and provide a rating, after which the weaknesses are revealed\. We argue that this system might work to some degree, but in reality, much of a review’s quality is determined by whether the highlighted weaknesses are sound and well supported\. Rating reviewers without seeing these details would likely produce noisy signals, and reviewers would be incentivized to write vague summaries that lack sharp substance; even worse, reviewers could inflate their strengths sections to manipulate the rating\.Kimet al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib2)\)themselves acknowledge that their proposed mechanism cannot address low\-quality negative reviews,111111[https://openreview\.net/forum?id=l8QemUZaIA&noteId=SXdGgGs6SV](https://openreview.net/forum?id=l8QemUZaIA&noteId=SXdGgGs6SV)which are likely the main complaints of most authors\. As for reviewer rewards,Kimet al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib2)\)mostly argues for vanity perks like digital badges \(e\.g\., ones similar to the “Pull Shark” badge on GitHub\)\. We argue that such vanity\-only perks would have much less influence than the review, submission, and cost\-influencing perks we propose here\.

Another piece of related work is the Isotonic Mechanism score pioneered bySu \([2021](https://arxiv.org/html/2608.14571#bib.bib13)\)and its follow\-up works likeWuet al\.\([2023](https://arxiv.org/html/2608.14571#bib.bib14)\); Suet al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib3)\), which survey authors \(with multiple submissions\) and ask them to rank their submitted papers\. A score is thereby calculated and compared with the mean of reviewers’ raw scores\. Should these two scores exhibit too drastic a gap, it may trigger AC intervention with an additional reviewer request, etc\. We note that this work is, by and large, orthogonal to ours, as it is yet another safeguard one can implement under our proposed credit system\. The two works overlap in the sense that some countermeasures do show resemblance \(e\.g\., requesting an additional reviewer\)\.

A broader but still relevant work is the survey byShah \([2022](https://arxiv.org/html/2608.14571#bib.bib20)\), which systematically examines peer review challenges across several dimensions: mismatched reviewer expertise, dishonest behavior, miscalibration, etc\. For each dimension, Shah discusses computational solutions, such as randomized assignment algorithms to mitigate dishonest bidding, or machine learning approaches to address commensuration bias\. While comprehensive, the survey proposes separate solutions tailored to each problem rather than a unified mechanism\. Our credit system, by contrast, offers a single flexible framework that can incorporate many of these solutions as special cases\.

Two other pieces of related work areRogers and Augenstein \([2020](https://arxiv.org/html/2608.14571#bib.bib12)\)andZhanget al\.\([2022](https://arxiv.org/html/2608.14571#bib.bib11)\)\.Rogers and Augenstein \([2020](https://arxiv.org/html/2608.14571#bib.bib12)\)discusses why incentives clash under a peer\-review context and outlines many potential proposals \(e\.g\., better review–paper matching, more tracks, abolishing “score\-based” feedback, track\-specific review formats, etc\.\)\. Similarly,Zhanget al\.\([2022](https://arxiv.org/html/2608.14571#bib.bib11)\)investigates various policies but focuses on modeling why resubmission is so prevalent despite many works eventually getting accepted at a top venue\. LikeShah \([2022](https://arxiv.org/html/2608.14571#bib.bib20)\), these works propose custom solutions for each problem or edge case, rather than a unified recipe\.

#### Credit\-based review systems

A small but growing set of works, like ours, frames the remedy as a credit\-, token\-, or general reward\-based currency for review\. Most directly comparable in mechanism, though set outside ML, is the concurrent work ofFranciaet al\.\([2026](https://arxiv.org/html/2608.14571#bib.bib23)\)\.121212Which first appeared publicly after ours, despite an earlier initial submission\.Like ours, it holds that voluntary fixes fall short and builds the remedy on a credit\-like currency for review; our differing settings, however, lead to substantially different designs along several axes:

- •Venue Setting\.The proposed token system underFranciaet al\.\([2026](https://arxiv.org/html/2608.14571#bib.bib23)\)is*primarily*aimed at reducing review delay \(shrinking editorial queues and submission\-to\-decision times\) under constraints especially salient for scientific journals: reviewer scarcity, hard\-to\-place “unattractive” papers, and slow rolling turnaround\. These are far less pressing at ML conferences, which draw on a comparatively ample reviewer pool and operate in fixed, batched cycles; our work’s central problems are instead submission*volume*and review*quality and accountability*\.131313We want to be clear thatFranciaet al\.\([2026](https://arxiv.org/html/2608.14571#bib.bib23)\)also employ levers regarding review quality, though more at a conceptual level, with fewer operational proposals\.One of its core remedies,*compulsory*reciprocity \(“from volunteerism to duty”\), is moreover already standard practice in ML and is, as argued in Section[3\.4](https://arxiv.org/html/2608.14571#S3.SS4), regarded as problematic, plausibly because ML’s far larger scale can turn compelled reviewing into many unwilling, ill\-suited reviews that erode quality and are much harder to police\.
- •Penalty\.The scheme is almost entirely reward\-*reduction*rather than punishment: a late or weak review merely earns a fraction of a token, and a balance typically turns negative only for an*undelivered*review\. In ML, such a case is already egregious enough to warrant desk rejection, precisely the blunt instrument we argue is too coarse; it thus offers comparatively little for the subtle, long\-tail misconduct that most needs policing, whereas we provide graduated, peer\- and AC\-adjudicated point*deductions*spanning minor to severe\.
- •Reward\.The currency is largely single\-purpose: one earns primarily by reviewing and spends primarily to submit, so it exists mainly to balance the review\-to\-submission loop\. We instead treat reward as a general lever for shaping behavior across the pipeline, attaching points to many actions \(e\.g\., emergency and “outstanding” reviews, initiating internal discussion, AC investigation, authors who return useful feedback to reviewers, and mentoring early\-career researchers\) and to many redemptions\.
- •Transferability\.The two designs largely agree: both let credits travel with the holder across venues, and both permit coauthors to pool credits toward a shared submission\. They diverge on one point:Franciaet al\.\([2026](https://arxiv.org/html/2608.14571#bib.bib23)\)additionally allow transfers between unrelated parties \(e\.g\., a non\-coauthor “mentor” may supply a rookie’s tokens\), which the balancing mechanism relies on but which, as the authors acknowledge, invites “fake coauthorship\.” We forbid such person\-to\-person transfer, along with any money\-to\-points conversion, keeping points tied to the labor that earned them \(Section[5\.5](https://arxiv.org/html/2608.14571#S5.SS5)\)\.

Both works engage a range of threat models and operational concerns; ours places somewhat more emphasis on defending against malicious gaming, through the four\-part defense of Section[4\.2](https://arxiv.org/html/2608.14571#S4.SS2)\.

Outside the machine learning community, we also haveGasparyanet al\.\([2015](https://arxiv.org/html/2608.14571#bib.bib7)\)analyzing peer review incentives under a mainly medical\-focused context\. The main argument of this work is“none of these \(financial or nonfinancial incentives\) is proven effectiveon its own”; however, the authors envision“a strategy of combined rewards and credits for the reviewers’ creative contributions seems a workable solution\.”Our work makes essentially the same argument, though framed under a credit system and tailored to the ML setting\. Our system supports many of the typical financial and non\-financial incentives mentioned inGasparyanet al\.\([2015](https://arxiv.org/html/2608.14571#bib.bib7)\)\. For instance, several exemplary policies we discuss in Section[4\.1](https://arxiv.org/html/2608.14571#S4.SS1)range from incentives for reviewers to write higher\-quality and more timely reviews, to deterrents for authors submitting unready work, to non\-financial privileges such as the right to be exempt from review assignments, and even financial compensation like free registration\. We do not advocate for any particular type of incentive, but rather a collection of them, unified under a credit system so that their impact can last beyond a single conference\.

Another point specific toGasparyanet al\.\([2015](https://arxiv.org/html/2608.14571#bib.bib7)\)is that many reviewers dismiss incentives such as paper purchase discounts or free publication access, since their institutions already pay for publisher subscriptions, making such incentives mostly relevant to members with non\-academic backgrounds\. In our proposal, however, conference hosts can offer services that no institution would purchase in bulk \(e\.g\., registration\), or even privileges that money cannot buy \(e\.g\., the right to request additional review resources\)\. While we do not claim these incentives are inherently better, we do believe our framework offers a flexible way to combine different incentives to suit diverse needs, effectively pushing forward a system thatGasparyanet al\.\([2015](https://arxiv.org/html/2608.14571#bib.bib7)\)envisions\.

Real\-world platforms such asReviewerCredits\(Fund and Dyke,[2024](https://arxiv.org/html/2608.14571#bib.bib24)\)andPublons\(Publons,[2018](https://arxiv.org/html/2608.14571#bib.bib25)\)have also explored giving reviewers some form of recognition or credit for their review activity\. Such efforts broadly share our premise that review labor merits durable, portable credit, though they generally operate as voluntary, third\-party services rather than mechanisms tied to a venue’s own decisions\. This is part of why we instead tie credit to concrete stakes within the conference pipeline\. Related ideas also surface outside academia, in the opensource community\. For instance,vouch141414[https://github\.com/mitchellh/vouch](https://github.com/mitchellh/vouch)implements a “web of trust” in which members vouch for \(or denounce\) one another to manage participation\. While conceptually adjacent to our credit system, the resemblance is loose:vouchis largely focused on access control, gating who is admitted in the first place, whereas our proposal is aimed at shaping behavior after admission \(for obvious fair\-science reasons\), with greater flexibility in how earned points can be spent\. We thus see it as related in spirit but distinct in goal\.

#### Conference Review Statistics

There are a few works likeYang \([2025](https://arxiv.org/html/2608.14571#bib.bib1)\); Cortes and Lawrence \([2021](https://arxiv.org/html/2608.14571#bib.bib15)\); Beygelzimeret al\.\([2023](https://arxiv.org/html/2608.14571#bib.bib16)\); Goldberget al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib9)\)that collect real statistics and conduct controlled experiments from past conferences\. While such works typically do not propose mechanism\-based solutions, their numerical presentations help illustrate the scale and disorder of current conferences; nonetheless, such controlled experiments shall offer us insight into the practical dynamics of a particular mechanism design\.

#### Broader Relevant Art

Beyond the effectiveness of any single incentive system, much of a conference’s experience also depends on many other aspects, such as the reduction of collusion rings\(Jecmenet al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib10)\), finding better reviewer–paper mappings\(Mimno and McCallum,[2007](https://arxiv.org/html/2608.14571#bib.bib17)\), determining the level of openness\(Raoet al\.,[2025](https://arxiv.org/html/2608.14571#bib.bib19)\), or the use of LLMs for reviewers\(Liu and Shah,[2023](https://arxiv.org/html/2608.14571#bib.bib18)\)\. We refer readers to these works to build a deeper understanding of these important topics, and recommendKimet al\.\([2025](https://arxiv.org/html/2608.14571#bib.bib2)\)for an overview of such efforts\.

## Appendix BThree Case Studies Where We Served as Reviewers

While our position paper, unfortunately, lacks real conference data to support why our proposed framework would be helpful, we believe there is even less reason to run an LLM\-roleplaying simulation\. We understand the perspective that having some anchors to real conferences is preferred\. Here, we share three case studies, all from similar top ML conferences, where we served as reviewers\. We share these not to highlight our own conduct, but to illustrate how small reviewer initiatives can help\.

### B\.1Case 1: Reviewers criticizing matters outside the paper’s scope\.

In this case, we observed that another reviewer criticized the submitted work for reasons clearly outside its intended scope\. We therefore raised our concern to the AC:

Internal comment to PC/SAC/ACThis message is set to be only visible to PC, SAC, AC, and the authors\.I want to disclose that I find reviewerA’s evaluation of this paper quite unreasonable\. This reviewer writes:•Certain methods, such asmethod type, demonstrate limited effectiveness insetting, which may restrict their practical deployment\.•The paper points out that many of thean important taskmethods tested are essentially extensions of existing models adapted foranother important task, such asa famous method\.It doesn’t make much sense to cite the low performance of certain featured methods as weaknesses of a dataset\-proposing/benchmark paper\. It is not the authors’ problem if an established method underperforms\. Instead, the point of benchmarking is precisely to show when a method would fail\. Many benchmark works have done this —some examples— and it is incomprehensible why this is considered a weakness\.Another criticism from reviewerAis:•The tasks within the benchmark may not capture all possible real\-world application scenarios, possibly overlooking specific needs within certain domains\.This, in my opinion, is a boilerplate concern that can be said for literally*any*dataset\. While I do agree that the proposed dataset does not capture some importanttaskscenarios —some examples— criticizing it for “not capturing all possible real\-world applications” crosses the line and feels borderline hostile\. This is akin to criticizing a method paper for not evaluating on every possible dataset\.I recommend the AC to either disregardA’s review or consider encouraging the reviewer to revisit the evaluation\.

This paper was ultimately accepted\. This anecdote shows that without internal reviewer analysis, simple rule\-based policies such as “inactive→\\rightarrowdesk rejection” fail to capture cases of severely low\-quality reviews\. Finer\-grained measures must be in place to ensure that positive impact scales broadly, rather than being limited to a few desk\-rejected papers\.

### B\.2Case 2: Reviewers asking for particular experiments after the rebuttal deadline\.

In a top ML conference where the exchange between authors and reviewers is limited to a certain time window, we had a split decision situation where the non/late\-responding reviewers were not supportive of the submission\. As the AC was calling for consensus, we stepped in and asked:

Internal reviewer discussionI skimmed over the two negative reviews of this work and found merits in many of the reviewer\-raised points\. However, I also find the authors’ rebuttal to be proper in many regards — especially when the raised concern demands a clarification\-like answer\.It looks like the two negative reviewers have yet to address the authors’ rebuttal in a meaningful way \(only acks are issued, cmiiw\)\. So, to reach a consensus, I believe it would be helpful if the two reviewers could elaborate a bit on their leftover concerns\. I am happy to set aside some time to discuss such leftover issues from my perspective\.

Essentially, the two negative reviewers believed that certain experiments were missing — one of which can be seen as a combination of two existing methods, and another as a specific investigatory study of the author\-proposed method\. While we find such suggestions to have merit, we believe they were not raised appropriately from a procedural standpoint:

Internal reviewer discussionI appreciateA’s detailed response and updated review\. I believeA\(as well asB\)’s main concerns regardingProposedMethodvs\.PriorWork1\+PriorWork2are legitimate and sound\.That being said, I am always the kind of reviewer who is “more in the authors’ shoes”—for lack of better words—and I would like to present two alternative arguments regarding this concern\.First, I believe experiment\-comparison requests that touch on*combinations of existing works*should be cautiously brought up\. Many methods can be combined, but their combinations typically require a number of discretionary design decisions, and it is often unlikely for authors to feature the exact combination a reviewer has in mind\. In this case,ProposedMethodproposes a paradigm of\[redacted\], where the scope of eligible combinations is wide\. Thus, in my opinion,if reviewers are specifically interested in the comparisonProposedMethodvs\.PriorWork1\+PriorWork2, such a request should be made*explicitly*before the rebuttal deadline, rather than mentioned in hindsight when the authors have no channel to address it\.From the look of it, theProposedMethodauthors submitted their initial rebuttal onan early date, which is\[redacted\]days after the review post\. However, onlyAengaged substantively ona late date\. I must note that this year’sBigConferenceNamerequires only two rounds of exchange, yet only 2/4 reviewers provided those to the authors, with all engaged reviewers leaning positive\.For such reasons, while I am also interested in this comparative result and agree withA’s analysis, I do not believe we can use it against the authors \(at least not as a singular veto reason\), as the request was not properly raised from a procedural standpoint\.Imagine we were submitting a paper where reviewers were largely non\-responsive, and the paper was then rejected for missing an experiment that was never explicitly asked for—it would be hard not to feel that is unfair\. While I understand we all have different priorities and may have limited bandwidth for various reasons, we need to properly compensate authors when such situations occur\.\(I know it is uncommon for a reviewer to argue on behalf of authors, but I always do so when it is warranted\. The AC is welcome to confirm that this advocacy is from me for good reasons and not the result of any collusion\.\)

Later, we and the two reviewers exchanged more than five comments in total, which contributed to the AC’s eventual decision\. This anecdote shows that many reviewers are willing to engage in internal discussions, and such exchanges are profoundly helpful in deciding borderline papers — they simply never had the “push” to initiate such discussion voluntarily\. We argue that our credit system could help elicit more of these productive discussions\.

Another takeaway from this exchange is the importance of having enforceable safeguards \(e\.g\., reply deadlines\), as otherwise authors may have no channel to meaningfully rebut at all\. Attaching rewards and penalties to such actions is also crucial to ensure procedural fairness\. While we respect and appreciateAandBfor engaging in our discussion, they were fine with leaving authors unanswered or replying late only because there is virtually no penalty to them \(as their “misconduct” would be too minute for desk rejection, yet the conference organizer has no other finer\-grained tools\) — something our credit system would help address\.

### B\.3Case 3: Authors presenting unsupportive results while claiming otherwise\.

In this case, we found that the authors were presenting unsupportive results to one reviewer while verbally claiming the opposite — so we stepped in:

Internal reviewer discussionAs another reviewer, I find the reading onProblemXtricky\. The goal ofTaskis to have the\[redacted\]\.InProblemX, when thecomponentis trained only ondataset1, thedataset2accuracy drops quite significantly — a disadvantaged result\. Once thecomponentis trained on bothdataset1anddataset2:•Thedataset1performance improvesa small numberover thea baseline, but this also comes at the cost of generating many more tokens than thea baseline\.•\[redacted as it is too technical\]•In the end, the system shows only asmall numberaccuracy gain on the task it was specifically trained on, whiledoing something at a higher cost\. I think the added experiment makes the work more negative \(though I appreciate the transparency\), as it feels like the proposedcomponentis very task\-specific and does not generalize well\.\(This message is set to be not visible to authors so that we can have a discussion with supposedly no bias, but I am happy to adjust the visibility should you want a more public discussion\.\)

The reviewer exchanged views with us and we reached an agreement\. The paper was ultimately rejected, with the AC citing concerns raised during the discussion\. Specifically, this other reviewer noted:

Internal reviewer discussion\(BTW, you’re among the very few reviewers who respond to other review comments\. I truly admire this level of dedication\.\)

This suggests that internal reviewer discussion is indeed rare, even though we lack direct statistical evidence\.

We hope these three anecdotal case studies illustrate how even a small component of our proposed framework \(encouraging more internal reviewer discussion\) can meaningfully facilitate paper evaluation\. While we acknowledge our bias, we believe that in all three cases, our initiative in starting these discussions clearly shaped the outcome\. This suggests that most reviewers*want*to contribute and most ACs will take such discussions seriously; they simply need a small push to take the initiative — a push that our credit system could very well provide\.

Similar Articles

NeurIPS 2026 AI-generated reviews [D]

Reddit r/MachineLearning

Discussion about the use of AI-generated reviews at NeurIPS 2026, including concerns over prompt injection and lack of consequences for reviewers using LLMs without proper oversight.