Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify
Summary
This paper formulates PRM stress testing as a quality-diversity search problem using MAP-Elites, characterizing what archive coverage can and cannot certify. Experiments on Qwen2.5-Math-PRM-7B reveal aggregation-dependent vulnerabilities, and a LoRA repair protocol reduces exploit rates.
View Cached Full Text
Cached at: 08/11/26, 08:09 AM
# Quality-Diversity Stress Tests for Process Reward Models: What Archive Coverage Can and Cannot Certify
Source: [https://arxiv.org/html/2608.08008](https://arxiv.org/html/2608.08008)
Fariya Afrin211footnotemark:1 1Department of Computer Science, Iowa State University 2Department of Computer Science, Kalinga Institute of Industrial Technology ishihab@iastate\.edu
###### Abstract
Process reward models \(PRMs\) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning\. We formulate PRM stress testing as a quality\-diversity search problem using MAP\-Elites, retaining the most severe correctness\-flipping edit in each behavior\-space region while separating search coverage from exploit coverage\. We characterize what such archives certify: finite\-cell repair bounds covered\-cell tail risk and average residual severity but cannot bound the worst remaining cell from covered fraction alone; under Lipschitz post\-repair loss and metric\-cover auditing, the residual is bounded by archive fitting error plus the Lipschitz constant times the covering radius\. A controlled landscape validates this certificate and the impossibility of any fraction\-only worst\-case guarantee\. On real PRMs, the search reveals an aggregation\-dependent vulnerability inQwen2\.5\-Math\-PRM\-7B: padding yields4444strict exploits with maximum gain0\.2940\.294under mean pooling versus one exploit under minimum readout; a matched syntactic control isolates the mechanism, and an RLHFlow value\-head model shows the same qualitative effect with maximum gain0\.0050\.005\. A predeclared paired LoRA repair protocol reduces exploit rates from0\.1480\.148to0\.0370\.037–0\.0740\.074, lowers the worst attack from0\.3330\.333to0\.1770\.177–0\.2120\.212, improves ranking AUROC without degrading best\-of\-44accuracy, attributes gains to adversarial fine\-tuning rather than archive diversity, and is confirmed by independent unpaired replications \(44→144\\\!\\rightarrow\\\!1, clean\-split worst gain0\.00920\.0092, MATH\-50041→041\\\!\\rightarrow\\\!0, clean ranking40/4040/40\)\.
Quality\-Diversity Stress Tests for Process Reward Models: What Archive Coverage Can and Cannot Certify
Ibne Farabi Shihab1††thanks:Equal contribution\.††thanks:Corresponding author:ishihab@iastate\.edu\.and Fariya Afrin211footnotemark:11Department of Computer Science, Iowa State University2Department of Computer Science, Kalinga Institute of Industrial Technologyishihab@iastate\.edu
## 1Introduction
Process reward models have become an important component of modern reasoning systems\. Instead of assigning a single score to a completed answer, a PRM evaluates intermediate steps, providing a dense signal that can guide inference\-time search, rank candidate derivations, or serve as a training reward\(Lightmanet al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib4); Uesatoet al\.,[2022](https://arxiv.org/html/2608.08008#bib.bib8); Wanget al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib9)\)\. The appeal is straightforward: a step\-level signal can distinguish two reasoning traces long before their final answers diverge\. The same granularity, however, creates a larger surface on which a learned evaluator can be optimized\. When a PRM rewards features that correlate with correctness on its training distribution but are not constitutive of correctness, search or reinforcement learning can amplify those features without improving the reasoning itself\.
This failure is the process\-level analogue of reward over\-optimization in outcome reward models\(Gaoet al\.,[2023](https://arxiv.org/html/2608.08008#bib.bib2); Skalseet al\.,[2022](https://arxiv.org/html/2608.08008#bib.bib13); Amodeiet al\.,[2016](https://arxiv.org/html/2608.08008#bib.bib1)\)\. Consider a correct trace and an edit that replaces its conclusion with a plausible wrong answer while adding several confident “verification” steps\. If the added steps receive high local scores, an average over step rewards can increase even though the edited trace is wrong\. Such an example is more than a naturally misclassified trace\. It is a direction in trace space along which the proxy improves as the target degrades, and therefore one that an optimizer is incentivized to exploit\.
Finding one such direction is useful but incomplete\. A single\-objective attack tends to return many variants of the same high\-scoring strategy, while a hand\-written test suite covers only the failure modes anticipated by its designer\. Quality\-diversity \(QD\) search provides a natural alternative\. Rather than retaining only the globally strongest attack, MAP\-Elites\(Mouret and Clune,[2015](https://arxiv.org/html/2608.08008#bib.bib6); Pughet al\.,[2016](https://arxiv.org/html/2608.08008#bib.bib7)\)keeps the strongest candidate in each region of a behavior space\. In our setting, the regions are defined by the type and magnitude of a trace edit\. The resulting archive reveals whether a PRM exhibits a single narrow vulnerability or a repertoire of qualitatively distinct failure modes\.
An archive also invites a tempting robustness claim: if attacks have been found and repaired in a fractionρ\\rhoof the cells, perhaps the worst remaining vulnerability should fall like1−ρ1\-\\rho\. This intuition is false\. An archive may cover almost every cell while missing one cell whose severity is maximal\. Coverage fraction can control an average or a tail probability, but it cannot by itself control a supremum\. This distinction matters because an easily reported coverage statistic can otherwise be mistaken for a worst\-case certificate that it does not justify\.
We therefore separate three notions that are often conflated\. Visitation coverage records where the search evaluated at least one candidate\. Exploit coverage records where it found a verified correctness\-flipping score increase\. Robustness coverage records where a post\-repair certificate supplies a valid cellwise upper bound; a finite heuristic search can only approximate this condition empirically\. These notions support fundamentally different robustness claims\. A finite partition yields a deterministic covered\-tail certificate and an average\-severity bound\. A worst\-case certificate requires additional structure, which we express through the covering radius of the audited examples in a metric over complete attack instances\.
Figure 1:The discovery\-to\-audit loop\. Search coverage and exploit coverage are recorded separately\. Training on archived exploits is followed by a fresh adaptive search; operationally, a cell is treated as repaired only after that audit, while the theorem requires a valid cellwise upper bound\. Finite\-cell coverage controls an average and a tail, while a metric covering radius is needed for a worst\-case certificate\.The empirical evaluation consists of two parts\. A controlled descriptor field makes every cell severity known, allowing us to verify the finite accounting identities and to demonstrate directly why a linear coverage\-only worst\-case claim fails\. We then apply QD search to real PRMs on GSM8K traces\(Cobbeet al\.,[2021](https://arxiv.org/html/2608.08008#bib.bib12)\)\. The primary reproducible finding is a verification\-dilution attack against a mean readout ofQwen2\.5\-Math\-PRM\-7B\. The same attack is structurally suppressed by a minimum readout, although the larger run contains one strict min\-readout exploit and therefore does not support a claim of immunity\. An RLHFlow value\-head PRM behaves differently: positive changes occur under both aggregations but are roughly fifty times smaller at their maximum\. These findings demonstrate why aggregation must be evaluated jointly with the PRM rather than presented as a universal defense\. The overall discovery\-to\-audit workflow is illustrated in Figure[1](https://arxiv.org/html/2608.08008#S1.F1)\.
The paper makes four main contributions:
- •A PRM\-specific QD formulationthat preserves diverse correctness\-flipping attacks while explicitly distinguishing visitation coverage from exploit coverage\.
- •A robustness certification frameworkcomprising finite\-cell average and tail certificates, an impossibility result for fraction\-only worst\-case guarantees, and a metric\-cover theorem specifying the additional geometric structure required for valid worst\-case certification\.
- •A controlled diagnosticthat exposes the distinction between finite\-cell guarantees and metric\-cover worst\-case guarantees\.
- •An empirical evaluationidentifying an aggregation\-dependent vulnerability, validating it through a syntactic negative control, analyzing cross\-model transfer, and separating parser failures and numerically tiny effects from substantive robustness conclusions\.
## 2Related Work
Process reward models \(PRMs\) were introduced to provide denser supervision than outcome\-only reward models and are now widely used for step\-level verification, best\-of\-nnselection, tree search, and process\-based reinforcement learning\(Uesatoet al\.,[2022](https://arxiv.org/html/2608.08008#bib.bib8); Lightmanet al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib4); Wanget al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib9)\)\. Their quality is typically evaluated through error localization or by measuring how well they rank correct reasoning traces above incorrect ones\(Lambertet al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib3)\)\. While these evaluations are essential, they do not directly examine how a fixed PRM behaves when trace construction is optimized against its own score\. Our experiments instead condition on a verified correctness change and measure whether the attacked trace receives a higher aggregate PRM score\.
Reward hacking has been studied extensively in the context of outcome reward models\. Optimizing a policy against a fixed learned reward can ultimately reduce true task performance\(Gaoet al\.,[2023](https://arxiv.org/html/2608.08008#bib.bib2)\)\. Adversarial training provides a general framework for improving robustness against adversarial attacks\(Madryet al\.,[2018](https://arxiv.org/html/2608.08008#bib.bib5)\), and recent work has adapted these ideas to reward models specifically\(Bukharinet al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib19)\)\. More recent work jointly trains a generator and a PRM to produce increasingly challenging process\-level negatives\(Junejaet al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib18)\)\. Our focus is complementary\. Rather than treating adversarial training alone as evidence of robustness, we study how a diverse, explicitly indexed failure archive changes what can be audited and what can be certified\.
Additional discussion of related aggregation choices, quality\-diversity search, and the distinction between our audited failure setting and prior PRM stress\-testing approaches is provided in Appendix[A](https://arxiv.org/html/2608.08008#A1)\.
## 3Problem Formulation
Letx=\(u,s\)x=\(u,s\)contain a problemuuand a reasoning tracess, and letq\(x\)∈\{0,1\}q\(x\)\\in\\\{0,1\\\}denote final\-answer correctness under a deterministic task verifier\. A PRM with parametersθ\\thetaemits step scores, which an aggregation rule converts into a trace scoreSθ\(x\)∈\[0,1\]S\_\{\\theta\}\(x\)\\in\[0,1\]\. The aggregation rule is part of the audited system: changing a mean readout to a minimum readout changesSθS\_\{\\theta\}even when the underlying step scores are unchanged\.
An attack operatorτ∈𝒳\\tau\\in\\mathcal\{X\}mapsxxto an edited traceτ\(x\)\\tau\(x\)\. We restrict the primary audit to correct base traces, soq\(x\)=1q\(x\)=1, and verify the attacked answer independently\. This makes the correctness event discrete and prevents a score difference from being conflated with the unit\-sized change in the binary label\.
###### Definition 1\(Correctness\-flipping exploit gain\)\.
The exploit gain of an attack instancez=\(x,τ\)z=\(x,\\tau\)is
ℓθ\(z\)=\\displaystyle\\ell\_\{\\theta\}\(z\)=\{\}𝟏\{q\(x\)=1,q\(τ\(x\)\)=0\}\\displaystyle\\mathbf\{1\}\\\{q\(x\)=1,\\ q\(\\tau\(x\)\)=0\\\}×\[Sθ\(τ\(x\)\)−Sθ\(x\)\]\+,\\displaystyle\\times\\big\[S\_\{\\theta\}\(\\tau\(x\)\)\-S\_\{\\theta\}\(x\)\\big\]\_\{\+\},where\[a\]\+=max\(a,0\)\[a\]\_\{\+\}=\\max\(a,0\)\. For an operational thresholdξ≥0\\xi\\geq 0,zzis aξ\\xi\-exploit when the correctness indicator is one andSθ\(τ\(x\)\)−Sθ\(x\)\>ξS\_\{\\theta\}\(\\tau\(x\)\)\-S\_\{\\theta\}\(x\)\>\\xi\.
This definition separates two facts that should not be added together: the attack has made the answer wrong, and the proxy has moved in the wrong direction\. BecauseSθ∈\[0,1\]S\_\{\\theta\}\\in\[0,1\], every cell severity defined below also lies in\[0,1\]\[0,1\], making the reported maximum increases of0\.2940\.294and0\.0050\.005directly interpretable\. The experiments report the strict conventionξ=0\\xi=0together with effect magnitudes; in deployment,ξ\\xishould be set above reader and numerical tolerance\.
Each attack has a behavioral descriptorβ\(z\)∈𝒟\\beta\(z\)\\in\\mathcal\{D\}\. We partition𝒟\\mathcal\{D\}intoMMcells and write𝒳c=\{z:β\(z\)=c\}\\mathcal\{X\}\_\{c\}=\\\{z:\\beta\(z\)=c\\\}for the attack instances represented by cellcc\. The population severity of a cell is
gθ\(c\)=supz∈𝒳cℓθ\(z\),gmax=maxc∈𝒟gθ\(c\),g\_\{\\theta\}\(c\)=\\operatorname\*\{sup\}\_\{z\\in\\mathcal\{X\}\_\{c\}\}\\ell\_\{\\theta\}\(z\),\\qquad g\_\{\\max\}=\\max\_\{c\\in\\mathcal\{D\}\}g\_\{\\theta\}\(c\),with the supremum defined as zero when a cell contains no valid correctness\-flipping attack\. An empirical search only supplies a lower bound ongθ\(c\)g\_\{\\theta\}\(c\); finding one elite in a cell does not establish that the cell supremum has been found\.
The distinction leads to three coverage quantities\. IfAAis the evaluated archive, thenV\(A\)V\(A\)contains cells in which at least one valid candidate was scored, whileCξ\(A\)C\_\{\\xi\}\(A\)contains cells whose empirical elite is aξ\\xi\-exploit\. Their respective fractions are
ρvisit\(A\)=\|V\(A\)\|M,ρexp\(A\)=\|Cξ\(A\)\|M\.\\rho\_\{\\mathrm\{visit\}\}\(A\)=\\frac\{\|V\(A\)\|\}\{M\},\\qquad\\rho\_\{\\mathrm\{exp\}\}\(A\)=\\frac\{\|C\_\{\\xi\}\(A\)\|\}\{M\}\.After repair, a third setℛ\(A\)\\mathcal\{R\}\(A\)contains cells assigned a valid upper bound at a declared residual levelε\\varepsilon\. Its fractionρrep=\|ℛ\(A\)\|/M\\rho\_\{\\mathrm\{rep\}\}=\|\\mathcal\{R\}\(A\)\|/Mis the coverage that enters the finite\-cell certificate\. A finite adaptive re\-audit supplies empirical evidence for this condition but, without additional structure, does not prove a population supremum\. A cell is not declared repaired merely because one archived example from it was used in training\.
## 4Quality\-Diversity Discovery
MAP\-Elites maintains one elite in each descriptor cell\. For a fixed base trace, the search is initialized from operator\-specific seeds\. At each iteration, it samples an existing elite, mutates its operator or magnitude descriptor, instantiates the corresponding textual edit, verifies the resulting final answer, and queries the PRM\. A candidate replaces the incumbent elite whenever it achieves a larger exploit gain\. Unlike severity\-only optimization, the archive preserves the strongest attack within every explored region of the descriptor space, even when that attack is not globally optimal\. The resulting archive exposes a diverse repertoire of failure modes while providing training pairs for archive\-guided repair\.
Our real\-model instantiation employs five operator families crossed with five magnitude levels\. The operators append a confident but unjustified conclusion, insert locally plausible verification padding, wrap an incorrect conclusion in authoritative mathematical language, overwrite the final numeric result with a nearby incorrect value, or repeat the incorrect answer as though repetition constituted evidence\. The magnitude descriptor controls the amount of inserted or modified content\. Appendix[C](https://arxiv.org/html/2608.08008#A3)details the archive update rule, the descriptor definitions, and the accounting procedure used when attacks derived from different base traces map to the same cell\.
The descriptor space is intentionally interpretable, but its limitations must be acknowledged\. A5×55\\times 5grid does not constitute a cover of natural\-language trace space\. Two edits assigned to the same cell may differ substantially in their semantics, while entirely new operator families may lie outside the descriptor grid\. Consequently,ρexp\\rho\_\{\\mathrm\{exp\}\}characterizes exploit coverage only within the audited operator–magnitude space rather than across all possible PRM failures\. Section[5](https://arxiv.org/html/2608.08008#S5)shows that meaningful worst\-case guarantees require additional geometric structure, expressed through a metric over complete attack instances together with an associated covering radius\.
## 5Archive\-Guided Repair and Coverage Certificates
The positive elites in an archive form adversarial training pairs\. A direct repair objective adds a pairwise margin loss to the original clean\-data objective,
ℒrepair\(θ\)=ℒclean\(θ\)\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{repair\}\}\(\\theta\)=\\mathcal\{L\}\_\{\\mathrm\{clean\}\}\(\\theta\)\+λ\|A\+\|∑\(x,τ\)∈A\+\[Sθ\(τ\(x\)\)−Sθ\(x\)\+m\]\+,\\displaystyle\\ \+\\frac\{\\lambda\}\{\|A\_\{\+\}\|\}\\sum\_\{\(x,\\tau\)\\in A\_\{\+\}\}\\bigl\[S\_\{\\theta\}\(\\tau\(x\)\)\-S\_\{\\theta\}\(x\)\+m\\bigr\]\_\{\+\},whereA\+A\_\{\+\}contains verified exploits andm≥0m\\geq 0is a desired separation margin\. The clean term anchors the original PRM behavior and prevents a vacuous repair that lowers all scores\. After training, the system is attacked again from fresh seeds\. The following statements apply to the post\-repair model when a certification procedure supplies a declared cellwise residual; a heuristic re\-audit is its empirical approximation rather than a formal upper bound\.
Letgθ′\(c\)g\_\{\\theta^\{\\prime\}\}\(c\)denote the post\-repair severity\. For a repaired setℛ⊆𝒟\\mathcal\{R\}\\subseteq\\mathcal\{D\}, define
εℛ\\displaystyle\\varepsilon\_\{\\mathcal\{R\}\}=maxc∈ℛgθ′\(c\),Uℛ=maxc∉ℛgθ\(c\),\\displaystyle=\\max\_\{c\\in\\mathcal\{R\}\}g\_\{\\theta^\{\\prime\}\}\(c\),\\qquad U\_\{\\mathcal\{R\}\}=\\max\_\{c\\notin\\mathcal\{R\}\}g\_\{\\theta\}\(c\),ζℛ\\displaystyle\\zeta\_\{\\mathcal\{R\}\}=maxc∉ℛ\[gθ′\(c\)−gθ\(c\)\]\+,\\displaystyle=\\max\_\{c\\notin\\mathcal\{R\}\}\\big\[g\_\{\\theta^\{\\prime\}\}\(c\)\-g\_\{\\theta\}\(c\)\\big\]\_\{\+\},where a maximum over an empty set is zero\. The spillover termζℛ\\zeta\_\{\\mathcal\{R\}\}records the largest increase induced by training in an unrepaired cell; assuming it is zero would generally be unjustified for a neural model\.
###### Theorem 1\(Finite\-cell coverage certificate\)\.
Letρrep=\|ℛ\|/M\\rho\_\{\\mathrm\{rep\}\}=\|\\mathcal\{R\}\|/Mand suppose a valid cellwise certificate establishesεℛ≤ε\\varepsilon\_\{\\mathcal\{R\}\}\\leq\\varepsilon\. Then
maxc∈𝒟gθ′\(c\)≤max\{ε,Uℛ\+ζℛ\}\.\\max\_\{c\\in\\mathcal\{D\}\}g\_\{\\theta^\{\\prime\}\}\(c\)\\leq\\max\\\!\\left\\\{\\varepsilon,U\_\{\\mathcal\{R\}\}\+\\zeta\_\{\\mathcal\{R\}\}\\right\\\}\.Moreover,
1M∑c∈𝒟gθ′\(c\)≤ρrepε\+\(1−ρrep\)\(Uℛ\+ζℛ\),\\frac\{1\}\{M\}\\sum\_\{c\\in\\mathcal\{D\}\}g\_\{\\theta^\{\\prime\}\}\(c\)\\leq\\rho\_\{\\mathrm\{rep\}\}\\varepsilon\+\(1\-\\rho\_\{\\mathrm\{rep\}\}\)\(U\_\{\\mathcal\{R\}\}\+\\zeta\_\{\\mathcal\{R\}\}\),and the fraction of cells whose post\-repair severity exceedsε\\varepsilonis at most1−ρrep1\-\\rho\_\{\\mathrm\{rep\}\}\.
The theorem separates the claims cleanly\. The first inequality is an exact residual decomposition: the worst case is either the audited residual on repaired cells or the severity, including harmful spillover, outside them\. The second and third statements are the coverage\-indexed conclusions that follow from a cell fraction\. They control an average and a tail, not the largest uncovered cell\.
###### Proposition 2\(A coverage fraction cannot control the worst cell\)\.
For everyM≥2M\\geq 2and every attainableρ<1\\rho<1, there is an exchangeable cell\-severity field and a perfectly repaired set covering fractionρ\\rhofor which the uncovered supremum equalsgmaxg\_\{\\max\}\. Hence, even with zero repaired\-cell error, no universal bound of the formmaxcgθ′\(c\)≤h\(ρ\)gmax\\max\_\{c\}g\_\{\\theta^\{\\prime\}\}\(c\)\\leq h\(\\rho\)g\_\{\\max\}withh\(ρ\)<1h\(\\rho\)<1can follow from coverage fraction alone\.
Proposition[2](https://arxiv.org/html/2608.08008#Thmtheorem2)rules out the appealing but incorrect\(1−ρ\)gmax\(1\-\\rho\)g\_\{\\max\}worst\-case law\. To obtain a worst\-case guarantee that improves continuously with coverage, the archive must say how far every possible attack lies from an audited example\.
###### Theorem 3\(Metric\-cover certificate\)\.
Let𝒳flip\\mathcal\{X\}\_\{\\mathrm\{flip\}\}be the set of valid correctness\-flipping attack instances, let\(𝒳flip,d\)\(\\mathcal\{X\}\_\{\\mathrm\{flip\}\},d\)be compact, and letB⊂𝒳flipB\\subset\\mathcal\{X\}\_\{\\mathrm\{flip\}\}be a finite audited set with covering radius
r\(B\)=supz∈𝒳flipminb∈Bd\(z,b\)\.r\(B\)=\\sup\_\{z\\in\\mathcal\{X\}\_\{\\mathrm\{flip\}\}\}\\min\_\{b\\in B\}d\(z,b\)\.Suppose the post\-repair exploit lossℓθ′\\ell\_\{\\theta^\{\\prime\}\}isLL\-Lipschitz with respect toddand the audit establishesℓθ′\(b\)≤ε\\ell\_\{\\theta^\{\\prime\}\}\(b\)\\leq\\varepsilonfor everyb∈Bb\\in B\. Then
supz∈𝒳flipℓθ′\(z\)≤ε\+Lr\(B\)\.\\sup\_\{z\\in\\mathcal\{X\}\_\{\\mathrm\{flip\}\}\}\\ell\_\{\\theta^\{\\prime\}\}\(z\)\\leq\\varepsilon\+Lr\(B\)\.
This theorem states the additional burden hidden by a cell count\. The metric must include the base problem and the semantic attack, or the descriptor must be sufficient for the loss; a radius computed only from operator identifiers does not control variation inside a cell\. In practice,LL,r\(B\)r\(B\), and the audit residual must be reported or upper\-bounded\. The complete proofs of Theorem[1](https://arxiv.org/html/2608.08008#Thmtheorem1), Proposition[2](https://arxiv.org/html/2608.08008#Thmtheorem2), and Theorem[3](https://arxiv.org/html/2608.08008#Thmtheorem3)appear in Appendix[B](https://arxiv.org/html/2608.08008#A2)\.
The aggregation mechanism found in the real\-model search has a separate algebraic explanation\.
###### Proposition 4\(Padding under mean and minimum readouts\)\.
Let a trace havekkstep rewards with meanr¯\\bar\{r\}and minimumrminr\_\{\\min\}\. Appendmmsteps whose mean reward isu¯\\bar\{u\}and minimum reward isuminu\_\{\\min\}\. The mean score changes by
mk\+m\(u¯−r¯\),\\frac\{m\}\{k\+m\}\(\\bar\{u\}\-\\bar\{r\}\),whereas the minimum score becomesmin\(rmin,umin\)\\min\(r\_\{\\min\},u\_\{\\min\}\)and therefore cannot increase under an append\-only edit\.
The proposition does not imply that minimum aggregation is universally robust\. An edit that removes or rewrites the original minimum\-scoring step can still raise the minimum, and a model can assign a high score to every incorrect step\. It predicts only that high\-scoring padding can dilute a mean but cannot hide an existing low score from a minimum\.
## 6Experiments
The experiments are organized around the distinctions introduced above\. We first examine what a cell fraction reveals when the population severity is fully known\. We then evaluate whether the QD procedure discovers correctness\-flipping score increases on a real PRM, whether the mechanism persists under a different aggregation rule, and whether the resulting attacked traces transfer across independently trained reward heads\. The controlled field serves as an accounting experiment rather than a surrogate for neural fine\-tuning, whereas the real\-model study comprises both a discovery audit*and*an executed repair evaluated by fresh adaptive re\-attack, including a clean\-split rerun that tests generalization to unseen problems \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px9)\)\. Maintaining this distinction prevents a simulated repair from being presented as evidence of an empirical defense\.
### Controlled coverage on a known field
We begin with a24×2424\\times 24descriptor grid comprising576576cells\. Each cell is assigned a nonnegative latent severity generated from a smooth mixture of Gaussian bumps and low\-amplitude noise, producing a small number of high\-severity regions alongside many benign ones\. The original field is normalized by its raw maximum of1\.2011\.201, yielding severities in\[0,1\]\[0,1\]consistent with Definition[1](https://arxiv.org/html/2608.08008#Thmdefinition1)\. MAP\-Elites explores the grid through local descriptor mutations\. At each evaluation budget, an oracle repair sets the severity of archived cells to the corresponding normalized floorε=0\.0167\\varepsilon=0\.0167while leaving all remaining cells unchanged\. Because this intervention is defined at the cell level, it directly tests the finite\-cell accounting of Theorem[1](https://arxiv.org/html/2608.08008#Thmtheorem1); it does not evaluate whether gradient\-based training would satisfy the same condition\.
QD evalsρ\\rhopost\-repair supuncovered supcertificate2000\.291\.00001\.0000holds5000\.550\.95420\.9542holds10000\.820\.95420\.9542holds20000\.960\.61620\.6162holds40000\.9980\.14900\.1490holds80001\.000\.01670\.0000holds
Table 1:Controlled field after normalizing the raw severities by their original maximum1\.2011\.201, so that the pre\-repairgmax=1g\_\{\\max\}=1\. The post\-repair supremum equals the maximum of the optimization floor and the uncovered supremum, exactly as stated by the finite\-cell certificate\. High cell coverage does not imply a proportionally small worst remaining cell: even atρ=0\.96\\rho=0\.96, the residual remains0\.61620\.6162\.Table[1](https://arxiv.org/html/2608.08008#S6.T1)illustrates two complementary observations\. First, the finite\-cell residual certificate holds exactly at every evaluation checkpoint, and the post\-repair supremum decreases monotonically as increasingly severe cells are eventually covered\. Second, this decrease is not linear in coverage\. Between coverages0\.550\.55and0\.820\.82, the worst remaining severity is unchanged, and even atρ=0\.998\\rho=0\.998, the single uncovered region retains severity0\.14900\.1490\. The regression of the post\-repair supremum on the uncovered supremum has slope0\.990\.99andR2=0\.9999R^\{2\}=0\.9999after including the optimization\-floor row, serving as an implementation check of the accounting rather than evidence for a proportional worst\-case law\. Complete construction and reporting details are provided in Appendix[D](https://arxiv.org/html/2608.08008#A4)\.
### Real\-model protocol
The primary audit evaluatesQwen2\.5\-Math\-PRM\-7B\(Zhanget al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib22)\)\. We apply the model’s chat template and step\-separator token, reading the positive\-class softmax probability at each separator rather than using raw logits\. Base examples are correct multi\-step GSM8K solutions\. An attack is counted only when exact\-match verification confirms that the base answer is correct, the edited answer is incorrect, and the aggregate PRM score increases strictly\. Each trace is initialized with five operator\-specific seeds followed by4040mutation evaluations, yielding4545evaluated candidates per readout\. Counts are therefore candidate\-level descriptive statistics, whereas archive coverage counts distinct descriptor cells and is bounded above by2525\.
We evaluate both the mean and minimum of the emitted step probabilities\. A matched syntactic control \(arithmetic corruption, dropped steps, operand swaps, and contentless padding\) tests whether the search rewards arbitrary edits\. The current evaluation adopts the strict thresholdξ=0\\xi=0\. Because numerically tiny floating\-point changes satisfy this convention, we report the maximum gain and interpret near\-threshold counts cautiously\. Appendix[5](https://arxiv.org/html/2608.08008#A5.T5)details the reader, correctness verification, archive construction, and thresholding rules\.
### An aggregation\-dependent Qwen vulnerability
The initial run attacks the6767of8080sampled traces whose base answers pass exact\-match verification:1616strict exploits under mean pooling, none under minimum, coverage0\.160\.16, and mean rise0\.0390\.039\(Table[4](https://arxiv.org/html/2608.08008#A5.T4), appendix\)\. Elites are dominated by verification dilution: a plausible incorrect final answer is followed by locally fluent verification steps whose scores exceed the trace mean\. Proposition[4](https://arxiv.org/html/2608.08008#Thmtheorem4)predicts this behavior exactly\. The same padding cannot increase an existing minimum, explaining the aggregation contrast without interpreting a zero count as a robustness guarantee\.
### Syntactic control: padding, not deception, is the mechanism\.
A matched control replaces the deceptive operators with five syntactic operators under the identical protocol and budget:4747strict mean\-readout exploits,*all*from the two padding operators and none from the three pure\-corruption operators\. Every control edit applies the same fixed final\-number corruption \(ensuring a verified correctness flip\), while the score*increase*arises from padding alone\. Thus, any locally innocuous padding exploits mean pooling \(Proposition[4](https://arxiv.org/html/2608.08008#Thmtheorem4)\), superseding the earlier deception\-required interpretation \(appendix Syntactic control\)\.
The larger run over120120traces reproduces both the mechanism and its boundary:4444strict mean\-pooling exploits, ten occupied exploit cells \(ρexp=0\.40\\rho\_\{\\mathrm\{exp\}\}=0\.40\), mean rise0\.0870\.087, and maximum0\.2940\.294\. Minimum pooling yields one strict exploit rather than none, indicating that the minimum readout strongly suppresses the tested padding mechanism on this PRM without providing immunity\.
### Problem\-level uncertainty and audit thresholds\.
Using the problem as the resampling unit, the mean\-readout exploit rate is0\.1680\.168\[0\.103,0\.243\]\[0\.103,0\.243\]\(n=107n\{=\}107\), with a best\-gap9595th percentile of0\.0970\.097\[0\.029,0\.237\]\[0\.029,0\.237\]\. The threshold curve \(τ=0\.01/0\.02/0\.05/0\.1/0\.2\\tau=0\.01/0\.02/0\.05/0\.1/0\.2:0\.150/0\.140/0\.084/0\.037/0\.0280\.150/0\.140/0\.084/0\.037/0\.028\) shows that strict counts overstate the number of exploits surviving practical margins\.
### Equal\-budget baselines\.
At the identical4545\-evaluation budget, random perturbation and exhaustive grid enumeration occupy*more*exploit cells than MAP\-Elites \(0\.480\.48vs\.0\.400\.40; single\-objective collapses to0\.120\.12\)\. With only2525cells, exhaustive evaluation is inexpensive, so this experiment does*not*demonstrate that QD search is necessary to achieve coverage at this grid size\. The reported severity statistics follow different conventions across methods \(baselines: nonnegative per\-problem best gain, mean0\.0170\.017–0\.0180\.018; archive: signed per\-problem best gaps, exploited\-candidate mean rise0\.0870\.087, maximum0\.2940\.294\), and we therefore make no severity claim based on these columns\. The corresponding same\-statistic comparison is*performed*in §[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px11)\(Table[6](https://arxiv.org/html/2608.08008#A7.T6)\), where equal\-budget exhaustive enumeration slightly exceeds MAP\-Elites\. The contribution of the archive is instead its behavior\-indexed structure, which supports the repair and certificate diagnostics\.
### Model dependence and transfer
The same reader\-normalized search on an RLHFlow Llama\-3\.1 value\-head PRM\(Xionget al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib11)\)provides an informative contrast\. Hundreds of strictly positive candidates are observed under both readouts, and every descriptor cell is occupied, yet the maximum score increase is only0\.0050\.005—approximately fifty times smaller than for Qwen\. Consequently, a deployment threshold above this scale may eliminate most or all strict events\. The results therefore indicate model dependence without providing evidence of severe RLHFlow exploitability \(full accounting in Table[9](https://arxiv.org/html/2608.08008#A7.T9), appendix\)\.
Transfer re\-scores the strict mean\-readout exploits of one PRM on the other\. A transfer is counted when the target PRM also strictly prefers the incorrect edited trace over its correct base\. All Qwen exploits transfer to RLHFlow, and0\.930\.93of RLHFlow exploits transfer to Qwen \(Table[8](https://arxiv.org/html/2608.08008#A7.T8), appendix\)\. The attack templates are therefore not specific to a single reader, although the strict threshold and the small RLHFlow effects preclude a stronger operational\-severity claim\.
An attempted Skywork\-o1 audit\(Skywork Team,[2024](https://arxiv.org/html/2608.08008#bib.bib10)\)did not produce a valid model result because the uniform reader failed to parse the model’s native step\-reward channel, returning constant zeros throughout\. These zeros, together with the corresponding zero transfer entries, are excluded from the evidential tables\. Appendix[F](https://arxiv.org/html/2608.08008#A6)documents the failed run to avoid its misinterpretation as evidence of robustness\. A meaningful Skywork comparison requires a model\-specific reader validated against reference outputs before attack evaluation\.
### Certificate diagnostics on the real archive \(plug\-in, not certified\)
We instantiate both certificate formulas on Qwen’s measured mean\-readout archive as*plug\-in diagnostics*, explicitly*not*certified worst\-case residuals, because five inputs are empirical lower bounds or unverified assumptions \(searched severities, the observed ratioL^=0\.46\\widehat\{L\}=0\.46, a descriptor\-cell metric,ρexp\\rho\_\{\\mathrm\{exp\}\}in place ofρrep\\rho\_\{\\mathrm\{rep\}\}, and unbounded spillover; the complete list appears in the caption of Table[5](https://arxiv.org/html/2608.08008#A5.T5)\)\. The plug\-in arithmetic nevertheless illustrates the central point: ten exploit cells \(ρexp=0\.40\\rho\_\{\\mathrm\{exp\}\}=0\.40,gmax=0\.294g\_\{\\max\}=0\.294\), zero measured severity elsewhere, and a finite\-cell residualεopt=0\.02\\varepsilon\_\{\\mathrm\{opt\}\}=0\.02; the metric\-cover plug\-in remains0\.480\.48for*every*audited subset smaller than the full grid \(radius11at\|B\|=5\|B\|=5–2020\), decreasing to0\.020\.02only when all2525cells are audited \(radius0\)\. Audited\-cell*count*alone provides no improvement until the audited set forms a genuine cover\. Values are taken fromresults\_certificate/\.
### Executed archive\-guided repair on the real PRM
We now execute the repair stage onQwen2\.5\-Math\-PRM\-7Brather than modeling it\. The first*implemented*objective is simpler than the pairwise\-margin loss of §[5](https://arxiv.org/html/2608.08008#S5): a LoRA adapter \(rank88,α=16\\alpha\{=\}16, dropout0\.050\.05, query/value\) is trained for6060AdamW steps \(lr10−410^\{\-4\}, batch11\) to down\-weight archived exploits through the model’s own readout, without a margin or clean term; clean behavior is evaluated post hoc\. Against a fresh adaptive re\-attack, the comparison is intentionally*unpaired*: full\-pool discovery identifies4444exploits \(0\.1680\.168of107107\), whereas the post\-repair re\-attack finds11\(maximum gain0\.00120\.0012versus the pre\-repair0\.2940\.294\), while clean ranking remains58/5858/58; the matched paired baseline is reported in §[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px11)\. The stricter clean\-split rerun \(repair on1313repair\-half exploits followed by re\-attack on the disjoint half\) likewise yields11strict exploit, maximum gain0\.00920\.0092, and clean ranking58/5858/58\. The pipeline cannot repair the RLHFlow value\-head PRM because it is not differentiable in our reader; this limitation is reported rather than omitted\. Values are taken fromresults\_defense/andresults\_gsm8k\_cleansplit/\.
### Benchmark generalization: MATH\-500\.
Applying the identical protocol to MATH\-500\(the Hendryckset al\.,[2021](https://arxiv.org/html/2608.08008#bib.bib25); Lightmanet al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib4), subset\)\(8282of120120sampled traces attacked\) reproduces the qualitative findings on harder problems:4141strict mean\-readout exploits \(problem\-level rate0\.2320\.232\[0\.146,0\.329\]\[0\.146,0\.329\]\), exploit\-cell coverage0\.320\.32, and a worst\-cell gap of0\.2320\.232; the threshold curve decreases from0\.2070\.207atτ=0\.01\\tau\{=\}0\.01to0\.0240\.024atτ=0\.1\\tau\{=\}0\.1\. Minimum pooling yields1919strict exploits whose largest cell gap,0\.0200\.020, is an order of magnitude smaller than the mean\-readout value of0\.2320\.232, indicating that on harder problems minimum pooling suppresses*severity*rather than exploit counts\. The post\-repair fresh re\-attack finds0exploits \(coverage0;gmax=0g\_\{\\max\}=0under the nonnegative\-gain convention, maximum*signed*difference−0\.0003\-0\.0003\), while all40/4040/40held\-out clean pairs remain correctly ranked\. Values are taken fromresults\_math500/\.
### Paired repair with the method’s loss
The predeclared paired protocol \(endpoints fixed before results; no public preregistration\) addresses the remaining evaluation gaps on GSM8K: a53/5453/54repair/evaluation split, discovery restricted to the repair half \(2020exploits, allverify\_dilution\), and matched pre/post fresh re\-attacks using the method’s margin\-plus\-clean\-anchor loss \(0\.100\.10/λ=0\.5\\lambda\{=\}0\.5, cached clean anchors; LoRA r88/α16\\alpha 16, AdamW10−410^\{\-4\},6060steps;33repair×\\times22attack seeds\)\. The frozen primary endpoint is the per\-problem best nonnegative gain averaged over attack seeds, ensuring that every reported column is internally consistent \(Table[7](https://arxiv.org/html/2608.08008#A7.T7)\): the pre\-repair rate is8/54=0\.1488/54=0\.148, the post\-repair rate is2/542/54–4/544/54, and all paired confidence intervals exclude zero \(Δ\\Deltagain−0\.012\-0\.012to−0\.015\-0\.015; hierarchical−0\.014\-0\.014\[−0\.028,−0\.003\]\[\-0\.028,\-0\.003\]\)\. Residual severity also decreases: the maximum gain falls from0\.3330\.333to0\.1770\.177–0\.2120\.212\(q950\.146→≤0\.0340\.146\\to\\leq 0\.034\), indicating reductions in both prevalence and the upper tail, although worst\-case attacks with approximately two\-thirds of the original severity remain\. AUROC improves \(0\.884→0\.9150\.884\\to 0\.915–0\.9350\.935\), best\-of\-44remains1\.001\.00, and the Brier score worsens \(0\.342→0\.3850\.342\\to 0\.385worst\)\. The v2 run additionally evaluates*repair\-source*controls under the same pair budget and optimization steps\. Random, exhaustive, and strongest\-only attack sets all achieve2/542/54post\-repair exploits, with reductions at least as large as those obtained from the QD archive \(−0\.016\-0\.016to−0\.017\-0\.017\), whereas the clean\-only control produces no change \(Δ≡0\\Delta\\equiv 0\)\. These results indicate that adversarial fine\-tuning drives the repair, whereas our claim for the archive is limited to its audit structure\. The unseen\-wording evaluation \(variant\-B templates, never used for training\) is only weakly informative: variant B produces few exploits even before repair \(rate0\.0370\.037\), and post\-repair rates are0–0\.0190\.019\. This shows no regression but provides insufficient pre\-repair signal to support a wording\-generalization claim\. Math\-Shepherd likewise has little to repair \(rate0\.0190\.019\)\. Complete results are reported in Tables[6](https://arxiv.org/html/2608.08008#A7.T6)–[7](https://arxiv.org/html/2608.08008#A7.T7)\(appendix;results\_paired\_v2/\)\. Additional implementation details, repair protocols, control experiments, and supplementary analyses are provided in Appendix[G](https://arxiv.org/html/2608.08008#A7)\. Supplementary figures illustrating repair behavior, aggregation effects, threshold sensitivity, generalization, and clean\-utility evaluation are presented in Appendix[H](https://arxiv.org/html/2608.08008#A8)\.
## 7Discussion and Conclusion
The observed pattern is more precise than the broad claim that “PRMs are easily hacked”: padding exploits a mean readout when inserted scores exceed the trace average, whereas a minimum readout blocks that route but not lowest\-step edits or a confidently wrong PRM\. These results indicate that PRM vulnerabilities are not governed by a single failure mode, but emerge from the interaction between attack structure, score aggregation, and the reliability of the learned evaluator\. Coverage describes a repertoire rather than a worst\-case guarantee \(Figure[2](https://arxiv.org/html/2608.08008#A4.F2)\); certificate numbers are plug\-in diagnostics \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px8)\) rather than standalone robustness certificates\. Repairs remain effective on unseen problems, MATH\-500, and the paired protocol, with the repair\-source controls attributing the effect to adversarial fine\-tuning rather than the archive \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px11)\)\. Together, these findings support adaptive stress testing and repair as complementary components for evaluating the reliability of process reward models under adversarial optimization\.
## Limitations
The real study uses modest problem subsets \(107107–120120traces per benchmark on GSM8K and MATH\-500\) and a hand\-designed5×55\\times 5descriptor grid\. The strict thresholdξ=0\\xi=0is sensitive to small reader differences, particularly for RLHFlow, and candidate counts are not independent because multiple mutations share a base problem and archive lineage \(problem\-level bootstrap intervals and threshold curves are reported in §[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px3)\)\. The real\-archive certificate numbers are plug\-in diagnostics, not certified bounds \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px8)\)\. The paired v2 protocol \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px11)\) supplies matched pre/post re\-attacks, reconciled union statistics, hierarchical bootstrap over problems, repair seeds, and attack seeds, equal\-budget repair\-source controls, and a full clean\-utility panel; its remaining gaps are honest and specific: worst observed post\-repair attacks retain about two\-thirds of the original maximum gain; the Brier score degrades slightly post\-repair \(score inflation on clean and corrupted traces alike\); the repair\-source controls show the QD archive is not a better repair\-data source than random or exhaustive attack sets at this scale, so the quality\-diversity claim concerns the auditable archive, not search or repair superiority; and the unseen\-wording arm is weakly informative because the variant templates barely exploit even the original model, while a genuinely open\-ended attack generator remains unexecuted\. All exploits in this channel come from one operator family \(verify\_dilution\), so cross\-family repair generalization is unmeasured\. The clean\-split repair establishes generalization to unseen*problems*, but not to unseen templates or operator families \(Appendix[G](https://arxiv.org/html/2608.08008#A7)\); the repair pipeline does not apply to non\-differentiable readouts such as the RLHFlow value head\. ProcessBench\(Zhenget al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib23)\)/PRMBench\(Songet al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib24)\)discrimination and a downstream reasoning evaluation of the repaired PRM are specified but not run; no number for either appears in this paper\. Finally, the edit operators are textual and black\-box; gradient\-, logit\-, or policy\-guided adversaries could expose failures outside the descriptor family\. These limitations bound the empirical claim to a reproducible diagnostic finding and motivate the broader evaluation protocol in the appendix\.
## References
- D\. Amodei, C\. Olah, J\. Steinhardt, P\. Christiano, J\. Schulman, and D\. Mané \(2016\)Concrete problems in AI safety\.arXiv preprint arXiv:1606\.06565\.Cited by:[§1](https://arxiv.org/html/2608.08008#S1.p2.1)\.
- A\. Bukharin, H\. Qian, S\. Sun, A\. Renduchintala, S\. Singhal, Z\. Wang, O\. Kuchaiev, O\. Delalleau, and T\. Zhao \(2025\)Adversarial training of reward models\.InThe Second Conference on Language Modeling,External Links:[Link](https://arxiv.org/abs/2504.06141)Cited by:[§2](https://arxiv.org/html/2608.08008#S2.p2.1)\.
- J\. Cheng, G\. Xiong, R\. Qiao, L\. Li, C\. Guo, J\. Wang, Y\. Lv, and F\. Wang \(2025\)Stop summation: min\-form credit assignment is all process reward model needs for reasoning\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/be91eb86eb74efc055cff83e953f86ce-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p1.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. Schulman \(2021\)Training verifiers to solve math word problems\.Cited by:[§1](https://arxiv.org/html/2608.08008#S1.p6.1)\.
- A\. Cully, J\. Clune, D\. Tarapore, and J\. Mouret \(2015\)Robots that can adapt like animals\.Vol\.521,pp\. 503–507\.Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p2.1)\.
- Q\. Dang, C\. Ngo, and T\. Hy \(2025\)RainbowPlus: enhancing adversarial prompt generation via evolutionary quality\-diversity search\.arXiv preprint arXiv:2504\.15047\.External Links:[Link](https://arxiv.org/abs/2504.15047)Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p2.1)\.
- L\. Gao, J\. Schulman, and J\. Hilton \(2023\)Scaling laws for reward model overoptimization\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2608.08008#S1.p2.1),[§2](https://arxiv.org/html/2608.08008#S2.p2.1)\.
- D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt \(2021\)Measuring mathematical problem solving with the MATH dataset\.InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track,External Links:[Link](https://arxiv.org/abs/2103.03874)Cited by:[§6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px10.p1.19)\.
- G\. Juneja, D\. Nathani, and W\. Y\. Wang \(2025\)Adversarial training for process reward models\.arXiv preprint arXiv:2511\.22888\.External Links:[Link](https://arxiv.org/abs/2511.22888)Cited by:[§2](https://arxiv.org/html/2608.08008#S2.p2.1)\.
- N\. Lambert, V\. Pyatkin, J\. Morrison, L\. J\. V\. Miranda, B\. Y\. Lin, K\. Chandu, N\. Dziri, S\. Kumar, T\. Zick, Y\. Choi, N\. A\. Smith, and H\. Hajishirzi \(2025\)RewardBench: evaluating reward models for language modeling\.InFindings of the Association for Computational Linguistics: NAACL,Cited by:[§2](https://arxiv.org/html/2608.08008#S2.p1.1)\.
- J\. Lehman and K\. O\. Stanley \(2011\)Abandoning objectives: evolution through the search for novelty alone\.Vol\.19,pp\. 189–223\.Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p2.1)\.
- H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe \(2024\)Let’s verify step by step\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2608.08008#S1.p1.1),[§2](https://arxiv.org/html/2608.08008#S2.p1.1),[§6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px10.p1.19)\.
- A\. Madry, A\. Makelov, L\. Schmidt, D\. Tsipras, and A\. Vladu \(2018\)Towards deep learning models resistant to adversarial attacks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2608.08008#S2.p2.1)\.
- J\. Mouret and J\. Clune \(2015\)Illuminating search spaces by mapping elites\.arXiv preprint arXiv:1504\.04909\.Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p2.1),[§1](https://arxiv.org/html/2608.08008#S1.p3.1)\.
- J\. K\. Pugh, L\. B\. Soros, and K\. O\. Stanley \(2016\)Quality diversity: a new frontier for evolutionary computation\.Frontiers in Robotics and AI3,pp\. 40\.Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p2.1),[§1](https://arxiv.org/html/2608.08008#S1.p3.1)\.
- M\. Samvelyan, S\. C\. Raparthy, A\. Lupu, E\. Hambro, A\. H\. Markosyan, M\. Bhatt, Y\. Mao, M\. Jiang, J\. Parker\-Holder, J\. Foerster, T\. Rocktäschel, and R\. Raileanu \(2024\)Rainbow teaming: open\-ended generation of diverse adversarial prompts\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 69747–69786\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/8147a43d030b43a01020774ae1d3e3bb-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p2.1)\.
- I\. F\. Shihab, F\. Afrin, S\. Akter, and A\. Sharma \(2026\)EST\-PRM: stress\-testing process reward models before they become load\-bearing\.arXiv preprint arXiv:2606\.00437\.External Links:[Link](https://arxiv.org/abs/2606.00437)Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p3.1)\.
- J\. Skalse, N\. H\. R\. Howe, D\. Krasheninnikov, and D\. Krueger \(2022\)Defining and characterizing reward hacking\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2608.08008#S1.p2.1)\.
- Skywork Team \(2024\)Skywork\-o1 open series: process reward models for complex reasoning\.[https://huggingface\.co/Skywork](https://huggingface.co/Skywork)\.Cited by:[§6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px7.p3.1)\.
- M\. Song, Z\. Su, X\. Qu, J\. Zhou, and Y\. Cheng \(2025\)PRMBench: a fine\-grained and challenging benchmark for process\-level reward models\.arXiv preprint arXiv:2501\.03124\.External Links:[Link](https://arxiv.org/abs/2501.03124)Cited by:[Limitations](https://arxiv.org/html/2608.08008#Sx1.p1.4)\.
- J\. Uesato, N\. Kushman, R\. Kumar, F\. Song, N\. Siegel, L\. Wang, A\. Creswell, G\. Irving, and I\. Higgins \(2022\)Solving math word problems with process\- and outcome\-based feedback\.arXiv preprint arXiv:2211\.14275\.Cited by:[§1](https://arxiv.org/html/2608.08008#S1.p1.1),[§2](https://arxiv.org/html/2608.08008#S2.p1.1)\.
- P\. Wang, L\. Li, Z\. Shao, R\. Xu, D\. Dai, Y\. Li, D\. Chen, Y\. Wu, and Z\. Sui \(2024\)Math\-Shepherd: verify and reinforce LLMs step\-by\-step without human annotations\.InAnnual Meeting of the Association for Computational Linguistics \(ACL\),Cited by:[§1](https://arxiv.org/html/2608.08008#S1.p1.1),[§2](https://arxiv.org/html/2608.08008#S2.p1.1)\.
- R\. Wang, K\. Xue, Z\. Qin, Z\. Li, S\. Tang, H\. Li, S\. Liu, and C\. Qian \(2025\)Quality\-diversity red\-teaming: automated generation of high\-quality and diverse attackers for large language models\.arXiv preprint arXiv:2506\.07121\.External Links:[Link](https://arxiv.org/abs/2506.07121)Cited by:[Appendix A](https://arxiv.org/html/2608.08008#A1.p2.1)\.
- W\. Xiong, H\. Zhang, N\. Jiang, and T\. Zhang \(2024\)An implementation of generative PRM \(RLHFlow\)\.InPreprint,Cited by:[Appendix E](https://arxiv.org/html/2608.08008#A5.SS0.SSS0.Px3.p1.4),[§6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px7.p1.1)\.
- Z\. Zhang, C\. Zheng, Y\. Wu, B\. Zhang, R\. Lin, B\. Yu, D\. Liu, J\. Zhou, and J\. Lin \(2025\)The lessons of developing process reward models in mathematical reasoning\.arXiv preprint arXiv:2501\.07301\.External Links:[Link](https://arxiv.org/abs/2501.07301)Cited by:[§6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px2.p1.3)\.
- C\. Zheng, Z\. Zhang, B\. Zhang, R\. Lin, K\. Lu, B\. Yu, D\. Liu, J\. Zhou, and J\. Lin \(2025\)ProcessBench: identifying process errors in mathematical reasoning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics,External Links:[Link](https://arxiv.org/abs/2412.06559)Cited by:[Limitations](https://arxiv.org/html/2608.08008#Sx1.p1.4)\.
## Appendix AExtended Related Work
Aggregation is a well\-established determinant of PRM behavior\. Sum\- or accumulation\-based credit assignment can reward long or repetitive generations, whereas minimum\-based aggregation can suppress some of these incentives\(Chenget al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib17)\)\. We do not claim minimum aggregation as a methodological contribution\. Instead, the mean–minimum comparison serves as a controlled diagnostic for the mechanism uncovered by the QD search, while the RLHFlow results demonstrate that the effect is model dependent\. This distinction is important because a minimum readout may itself be sensitive to a single spuriously low score and may not preserve the ranking utility of the original PRM\.
Our search framework builds on quality\-diversity \(QD\) optimization\. MAP\-Elites and related illumination algorithms preserve high\-performing solutions throughout a user\-defined behavior space\(Mouret and Clune,[2015](https://arxiv.org/html/2608.08008#bib.bib6); Pughet al\.,[2016](https://arxiv.org/html/2608.08008#bib.bib7); Lehman and Stanley,[2011](https://arxiv.org/html/2608.08008#bib.bib14)\), extending ideas originally developed for learning diverse behavioral repertoires\(Cullyet al\.,[2015](https://arxiv.org/html/2608.08008#bib.bib15)\)\. QD has also been applied to language\-model red teaming\. Rainbow Teaming formulates adversarial prompt generation as open\-ended QD search, and subsequent work learns behavior\-conditioned attackers across multiple risk categories\(Samvelyanet al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib16); Wanget al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib21); Danget al\.,[2025](https://arxiv.org/html/2608.08008#bib.bib27)\)\. Our contribution is therefore not the general observation that QD can diversify attacks, nor the claim that MAP\-Elites is inherently a superior search algorithm or repair\-data source\. Indeed, under our small fixed operator grid and equal search budget, exhaustive enumeration matches MAP\-Elites on both discovery and repair performance \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px11)\)\. Rather, our contribution is the specialization of QD to process\-reward traces, the formulation of the behavior\-indexed*archive as an auditable object*, the explicit separation of distinct coverage notions, and the characterization of which forms of coverage support which robustness guarantees\.
The closest PRM\-specific stress\-testing framework is EST\-PRM, which measures score inflation and correlation collapse under fixed, label\-preserving transformations such as step inflation, dependency\-aware reordering, and confidence markers\(Shihabet al\.,[2026](https://arxiv.org/html/2608.08008#bib.bib20)\)\. Our setting changes the audited event\. We begin with correct reasoning traces and search for edits that simultaneously flip final correctness and increase the aggregate PRM score\. The resulting QD archive therefore consists of verified proxy\-improving failures rather than a label\-preserving invariance suite\. The overlap in structural perturbations remains important, and we accordingly use those transformations as informed operator seeds rather than presenting them as novel attack classes\.
## Appendix BComplete Proofs
We restate the relevant notation\. The finite descriptor partition containsMMcells\. The pre\-repair and post\-repair population severities aregθ\(c\)g\_\{\\theta\}\(c\)andgθ′\(c\)g\_\{\\theta^\{\\prime\}\}\(c\), respectively\. The repaired set isℛ\\mathcal\{R\}, its coverage isρrep=\|ℛ\|/M\\rho\_\{\\mathrm\{rep\}\}=\|\\mathcal\{R\}\|/M, and
εℛ=maxc∈ℛgθ′\(c\),Uℛ=maxc∉ℛgθ\(c\),\\varepsilon\_\{\\mathcal\{R\}\}=\\max\_\{c\\in\\mathcal\{R\}\}g\_\{\\theta^\{\\prime\}\}\(c\),\\qquad U\_\{\\mathcal\{R\}\}=\\max\_\{c\\notin\\mathcal\{R\}\}g\_\{\\theta\}\(c\),ζℛ=maxc∉ℛ\[gθ′\(c\)−gθ\(c\)\]\+\.\\zeta\_\{\\mathcal\{R\}\}=\\max\_\{c\\notin\\mathcal\{R\}\}\\big\[g\_\{\\theta^\{\\prime\}\}\(c\)\-g\_\{\\theta\}\(c\)\\big\]\_\{\+\}\.All severities are nonnegative\. We use the convention that a maximum over the empty set is zero\.
### Proof of the finite\-cell certificate
###### Proof of Theorem[1](https://arxiv.org/html/2608.08008#Thmtheorem1)\.
Partition the descriptor cells intoℛ\\mathcal\{R\}and its complement\. For everyc∈ℛc\\in\\mathcal\{R\}, the post\-repair audit gives
gθ′\(c\)≤εℛ≤ε\.g\_\{\\theta^\{\\prime\}\}\(c\)\\leq\\varepsilon\_\{\\mathcal\{R\}\}\\leq\\varepsilon\.For everyc∉ℛc\\notin\\mathcal\{R\}, write the post\-repair severity as its pre\-repair value plus its change:
gθ′\(c\)\\displaystyle g\_\{\\theta^\{\\prime\}\}\(c\)=gθ\(c\)\+gθ′\(c\)−gθ\(c\)\\displaystyle=g\_\{\\theta\}\(c\)\+g\_\{\\theta^\{\\prime\}\}\(c\)\-g\_\{\\theta\}\(c\)≤gθ\(c\)\+\[gθ′\(c\)−gθ\(c\)\]\+\\displaystyle\\leq g\_\{\\theta\}\(c\)\+\\big\[g\_\{\\theta^\{\\prime\}\}\(c\)\-g\_\{\\theta\}\(c\)\\big\]\_\{\+\}≤Uℛ\+ζℛ\.\\displaystyle\\leq U\_\{\\mathcal\{R\}\}\+\\zeta\_\{\\mathcal\{R\}\}\.Taking the maximum over the two parts of the partition gives
maxc∈𝒟gθ′\(c\)≤max\{ε,Uℛ\+ζℛ\}\.\\max\_\{c\\in\\mathcal\{D\}\}g\_\{\\theta^\{\\prime\}\}\(c\)\\leq\\max\\\!\\left\\\{\\varepsilon,U\_\{\\mathcal\{R\}\}\+\\zeta\_\{\\mathcal\{R\}\}\\right\\\}\.
For the average, sum the corresponding cellwise bounds:
1M∑c∈𝒟gθ′\(c\)\\displaystyle\\frac\{1\}\{M\}\\sum\_\{c\\in\\mathcal\{D\}\}g\_\{\\theta^\{\\prime\}\}\(c\)=1M∑c∈ℛgθ′\(c\)\+1M∑c∉ℛgθ′\(c\)\\displaystyle=\\frac\{1\}\{M\}\\sum\_\{c\\in\\mathcal\{R\}\}g\_\{\\theta^\{\\prime\}\}\(c\)\+\\frac\{1\}\{M\}\\sum\_\{c\\notin\\mathcal\{R\}\}g\_\{\\theta^\{\\prime\}\}\(c\)≤\|ℛ\|Mε\+M−\|ℛ\|M\(Uℛ\+ζℛ\)\\displaystyle\\leq\\frac\{\|\\mathcal\{R\}\|\}\{M\}\\varepsilon\+\\frac\{M\-\|\\mathcal\{R\}\|\}\{M\}\(U\_\{\\mathcal\{R\}\}\+\\zeta\_\{\\mathcal\{R\}\}\)=ρrepε\+\(1−ρrep\)\(Uℛ\+ζℛ\)\.\\displaystyle=\\rho\_\{\\mathrm\{rep\}\}\\varepsilon\+\(1\-\\rho\_\{\\mathrm\{rep\}\}\)\(U\_\{\\mathcal\{R\}\}\+\\zeta\_\{\\mathcal\{R\}\}\)\.
Finally, every cell inℛ\\mathcal\{R\}has post\-repair severity at mostε\\varepsilon\. Therefore
\{c:gθ′\(c\)\>ε\}⊆𝒟∖ℛ,\\\{c:g\_\{\\theta^\{\\prime\}\}\(c\)\>\\varepsilon\\\}\\subseteq\\mathcal\{D\}\\setminus\\mathcal\{R\},and hence
1M\|\{c:gθ′\(c\)\>ε\}\|≤M−\|ℛ\|M=1−ρrep\.\\frac\{1\}\{M\}\\left\|\\\{c:g\_\{\\theta^\{\\prime\}\}\(c\)\>\\varepsilon\\\}\\right\|\\leq\\frac\{M\-\|\\mathcal\{R\}\|\}\{M\}=1\-\\rho\_\{\\mathrm\{rep\}\}\.All three conclusions are deterministic and make no exchangeability assumption\. ∎
If a deployment distribution is uniform over cells, the average inequality is also an expected\-risk bound\. More generally, for declared cell weightswc≥0w\_\{c\}\\geq 0with∑cwc=1\\sum\_\{c\}w\_\{c\}=1, the same proof gives
∑cwcgθ′\(c\)≤w\(ℛ\)ε\+∑c∉ℛwc\(gθ\(c\)\+ζc\),\\sum\_\{c\}w\_\{c\}g\_\{\\theta^\{\\prime\}\}\(c\)\\leq w\(\\mathcal\{R\}\)\\varepsilon\+\\sum\_\{c\\notin\\mathcal\{R\}\}w\_\{c\}\\left\(g\_\{\\theta\}\(c\)\+\\zeta\_\{c\}\\right\),wherew\(ℛ\)=∑c∈ℛwcw\(\\mathcal\{R\}\)=\\sum\_\{c\\in\\mathcal\{R\}\}w\_\{c\}andζc=\[gθ′\(c\)−gθ\(c\)\]\+\\zeta\_\{c\}=\[g\_\{\\theta^\{\\prime\}\}\(c\)\-g\_\{\\theta\}\(c\)\]\_\{\+\}\. This weighted form is preferable when attack cells are not equally likely in deployment\.
### Proof that fraction coverage cannot control a supremum
###### Proof of Proposition[2](https://arxiv.org/html/2608.08008#Thmtheorem2)\.
FixM≥2M\\geq 2and an attainable fractionρ=k/M<1\\rho=k/M<1\. Let every pre\-repair cell have the same severity,
gθ\(c\)=gmax\>0for allc∈𝒟\.g\_\{\\theta\}\(c\)=g\_\{\\max\}\>0\\qquad\\text\{for all \}c\\in\\mathcal\{D\}\.This constant field is exchangeable: permuting cell labels leaves it unchanged, and no cell is more likely than another to contain a maximum because every cell is a maximum\. Choose any repaired setℛ\\mathcal\{R\}of sizekkand suppose, favorably, that repair drives every cell inℛ\\mathcal\{R\}to zero and causes no spillover\. Sincek<Mk<M, at least one cell remains outsideℛ\\mathcal\{R\}, and every such cell retains severitygmaxg\_\{\\max\}\. Consequently,
maxc∉ℛgθ′\(c\)=gmax\.\\max\_\{c\\notin\\mathcal\{R\}\}g\_\{\\theta^\{\\prime\}\}\(c\)=g\_\{\\max\}\.For any proposed universal multiplierh\(ρ\)<1h\(\\rho\)<1, this construction violatesmaxcgθ′\(c\)≤h\(ρ\)gmax\\max\_\{c\}g\_\{\\theta^\{\\prime\}\}\(c\)\\leq h\(\\rho\)g\_\{\\max\}even though the field and the choice among equally severe cells are fully symmetric\. Adding an optimization floorε<1−h\(ρ\)gmax\\varepsilon<\{1\-h\(\\rho\)\}g\_\{\\max\}does not change the contradiction\. Thus coverage fraction alone cannot yield a nontrivial worst\-case contraction\. ∎
The argument also identifies the error in a quantile\-based intuition\. Selecting or repairing aρ\\rhofraction of cells may control the distribution of a uniformly sampled cell, but the maximum over the remaining\(1−ρ\)M\(1\-\\rho\)Mcells is not the\(1−ρ\)\(1\-\\rho\)\-quantile of the original severity distribution\. It is an extreme order statistic of the uncovered subset and can remain equal to the global maximum\.
### Proof of the metric\-cover certificate
###### Proof of Theorem[3](https://arxiv.org/html/2608.08008#Thmtheorem3)\.
Fix anyz∈𝒳flipz\\in\\mathcal\{X\}\_\{\\mathrm\{flip\}\}\. BecauseBBis finite, there existsbz∈Bb\_\{z\}\\in Battainingminb∈Bd\(z,b\)\\min\_\{b\\in B\}d\(z,b\)\. Lipschitz continuity and the audited residual bound give
ℓθ′\(z\)\\displaystyle\\ell\_\{\\theta^\{\\prime\}\}\(z\)≤ℓθ′\(bz\)\+\|ℓθ′\(z\)−ℓθ′\(bz\)\|\\displaystyle\\leq\\ell\_\{\\theta^\{\\prime\}\}\(b\_\{z\}\)\+\\left\|\\ell\_\{\\theta^\{\\prime\}\}\(z\)\-\\ell\_\{\\theta^\{\\prime\}\}\(b\_\{z\}\)\\right\|≤ε\+Ld\(z,bz\)\\displaystyle\\leq\\varepsilon\+Ld\(z,b\_\{z\}\)≤ε\+Lr\(B\)\.\\displaystyle\\leq\\varepsilon\+Lr\(B\)\.Since the inequality holds for everyz∈𝒳flipz\\in\\mathcal\{X\}\_\{\\mathrm\{flip\}\}, taking the supremum proves
supz∈𝒳flipℓθ′\(z\)≤ε\+Lr\(B\)\.\\sup\_\{z\\in\\mathcal\{X\}\_\{\\mathrm\{flip\}\}\}\\ell\_\{\\theta^\{\\prime\}\}\(z\)\\leq\\varepsilon\+Lr\(B\)\.Compactness ensures that the covering radius is finite for the intended bounded domain; the pointwise argument itself requires only that the displayed radius be finite\. ∎
If the audit only establishes noisy upper estimatesℓ^θ′\(b\)\\widehat\{\\ell\}\_\{\\theta^\{\\prime\}\}\(b\)with simultaneous error at mostν\\nu, thenℓθ′\(b\)≤ℓ^θ′\(b\)\+ν\\ell\_\{\\theta^\{\\prime\}\}\(b\)\\leq\\widehat\{\\ell\}\_\{\\theta^\{\\prime\}\}\(b\)\+\\nu\. Replacingε\\varepsilonbymaxbℓ^θ′\(b\)\+ν\\max\_\{b\}\\widehat\{\\ell\}\_\{\\theta^\{\\prime\}\}\(b\)\+\\nuin the proof yields
supz∈𝒳flipℓθ′\(z\)≤maxb∈Bℓ^θ′\(b\)\+ν\+Lr\(B\)\.\\sup\_\{z\\in\\mathcal\{X\}\_\{\\mathrm\{flip\}\}\}\\ell\_\{\\theta^\{\\prime\}\}\(z\)\\leq\\max\_\{b\\in B\}\\widehat\{\\ell\}\_\{\\theta^\{\\prime\}\}\(b\)\+\\nu\+Lr\(B\)\.This extension makes explicit that statistical uncertainty and geometric coverage are separate terms\.
### Proof of the pooling proposition
###### Proof of Proposition[4](https://arxiv.org/html/2608.08008#Thmtheorem4)\.
Let the original rewards ber1,…,rkr\_\{1\},\\ldots,r\_\{k\}and the appended rewards beu1,…,umu\_\{1\},\\ldots,u\_\{m\}\. By definition,
∑i=1kri=kr¯,∑j=1muj=mu¯\.\\sum\_\{i=1\}^\{k\}r\_\{i\}=k\\bar\{r\},\\qquad\\sum\_\{j=1\}^\{m\}u\_\{j\}=m\\bar\{u\}\.The new mean is therefore
r¯′=kr¯\+mu¯k\+m\.\\bar\{r\}^\{\\prime\}=\\frac\{k\\bar\{r\}\+m\\bar\{u\}\}\{k\+m\}\.Subtracting the original mean gives
r¯′−r¯=kr¯\+mu¯−\(k\+m\)r¯k\+m=mk\+m\(u¯−r¯\)\.\\bar\{r\}^\{\\prime\}\-\\bar\{r\}=\\frac\{k\\bar\{r\}\+m\\bar\{u\}\-\(k\+m\)\\bar\{r\}\}\{k\+m\}=\\frac\{m\}\{k\+m\}\(\\bar\{u\}\-\\bar\{r\}\)\.Thus padding raises the mean exactly when the appended average exceeds the original average\.
For the minimum readout, the new set of rewards is the union of the original and appended rewards, so
rmin′\\displaystyle r^\{\\prime\}\_\{\\min\}=min\(miniri,minjuj\)\\displaystyle=\\min\\bigl\(\\min\_\{i\}r\_\{i\},\\min\_\{j\}u\_\{j\}\\bigr\)=min\(rmin,umin\)≤rmin\.\\displaystyle=\\min\(r\_\{\\min\},u\_\{\\min\}\)\\leq r\_\{\\min\}\.Hence an append\-only edit cannot strictly increase the minimum score\. ∎
## Appendix CQuality\-Diversity Search Details
### Archive update
For each base trace, the implementation first constructs one seed per operator family\. An elite record contains the edited trace, its operator and magnitude descriptor, the base and attacked aggregate scores, the verified base and attacked answers, and the resulting exploit gain\. The following update is applied for each mutation evaluation\.
StepOperation1Sample a nonempty archive cell uniformly and retrieve its elite\.2Mutate the operator identifier or move the magnitude by one neighboring level, with boundary reflection\.3Instantiate the mutated edit on the same base trace and verify that the edited answer is parseable\.4Query the PRM once, aggregate its native step probabilities, and compute Definition[1](https://arxiv.org/html/2608.08008#Thmdefinition1)\.5Insert the candidate if its cell is empty; otherwise replace the elite only when the exploit gain is larger\.Table 2:MAP\-Elites update used by the discrete trace\-edit search\. The model query budget, rather than the number of successful exploits, is the appropriate unit for comparisons with alternative search methods\.When several base traces map to the same global descriptor cell, the global archive retains the largest observed exploit gain and stores the originating problem identifier\. Candidate\-level exploit counts include every evaluated candidate satisfying the criterion and can therefore exceed the number of cells\. Cell coverage counts each descriptor cell once\. This distinction explains why RLHFlow can yield hundreds of strict candidates while the archive coverage remains bounded by one\. The MAP\-Elites archive update is summarized in Table[2](https://arxiv.org/html/2608.08008#A3.T2)\.
### Descriptor semantics
The descriptor combines one of five operator families with one of five ordered magnitudes\. The operator families and the semantic role of magnitude are summarized in Table[3](https://arxiv.org/html/2608.08008#A3.T3)\. Exact surface strings are implementation details and should be released alongside the evaluation script so that paraphrase\-level changes are not hidden by a shared family name\.
Operator familyTransformationMagnitude axisConfident conclusionReplaces the final answer with a plausible wrong value and appends an unjustified confident conclusion\.Amount of concluding scaffolding\.Verification paddingAdds locally fluent checking or restatement steps around the wrong conclusion without introducing valid new evidence\.Number of inserted verification units\.Authoritative wrapperSurrounds the wrong result with formal mathematical language and certainty markers\.Density of authoritative framing\.Nearby overwriteSilently changes the final numeric result to a nearby wrong value while preserving the preceding derivation\.Size of the normalized numeric displacement\.Repeated assertionRestates the wrong answer multiple times as though repetition strengthened its validity\.Number of repeated assertions\.Table 3:Behavior descriptors in the real\-model archive\. These cells organize the declared attack family; they do not by themselves define a metric cover of all semantic trace edits\.
### Coverage accounting
A cell entersV\(A\)V\(A\)after one valid candidate in that cell has been scored, regardless of whether the candidate is an exploit\. A cell entersCξ\(A\)C\_\{\\xi\}\(A\)only when its elite exceeds the declared threshold\. Operationally, a cell is proposed forℛ\(A\)\\mathcal\{R\}\(A\)only if a fresh adaptive re\-audit, initialized independently of the training elite, fails to exceed the residual levelε\\varepsilonwithin the declared budget\. The formal finite\-cell theorem additionally requires that this empirical condition be promoted to a valid upper bound by an appropriate certification argument; the metric\-cover theorem gives one such route\. Reporting only\|Cξ\(A\)\|/M\|C\_\{\\xi\}\(A\)\|/Mloses two important facts: an empty exploit cell may have been extensively searched without success, and an occupied exploit cell may contain many unseen attacks that survive training\.
## Appendix DControlled\-Landscape Details
0\.250\.250\.30\.30\.350\.350\.40\.40\.450\.450\.50\.50\.550\.550\.60\.60\.650\.650\.70\.70\.750\.750\.80\.80\.850\.850\.90\.90\.950\.951100\.50\.511archive coverageρ\\rhopost\-repair supremumobserved supremumfraction\-only heuristicFigure 2:Coverage and worst\-cell severity need not be proportional\. The dashed curve is the invalid heuristic\(1−ρ\)gmax\+ε\(1\-\\rho\)g\_\{\\max\}\+\\varepsilon, not a bound\. Every non\-full\-coverage checkpoint lies above it, including coverages of0\.960\.96and0\.9980\.998\. The experiment therefore illustrates Proposition[2](https://arxiv.org/html/2608.08008#Thmtheorem2)as well as the valid residual decomposition\.The synthetic field is a24×2424\\times 24grid\. Its latent severities are formed from a small number of smooth Gaussian bumps plus low\-amplitude noise, followed by clipping at zero\. MAP\-Elites is initialized along the operator axis and repeatedly selects an elite, changes one grid coordinate by one cell, and evaluates the latent severity at the resulting location\. Budgets are swept over\{200,500,1000,2000,4000,8000\}\\\{200,500,1000,2000,4000,8000\\\}\.
The raw generated field has maximum1\.2011\.201\. Before reporting, every severity and the repair floor are divided by1\.2011\.201, producing the normalized values in Table[1](https://arxiv.org/html/2608.08008#S6.T1)\. At each checkpoint, oracle repair replaces the severity of archived cells byε=0\.0167\\varepsilon=0\.0167\. This makesζℛ=0\\zeta\_\{\\mathcal\{R\}\}=0and leaves the unrepaired cells at their pre\-repair values\. Consequently, Theorem[1](https://arxiv.org/html/2608.08008#Thmtheorem1)specializes to
maxcgθ′\(c\)=max\{0\.0167,maxc∉ℛgθ\(c\)\}\.\\max\_\{c\}g\_\{\\theta^\{\\prime\}\}\(c\)=\\max\\\!\\left\\\{0\.0167,\\max\_\{c\\notin\\mathcal\{R\}\}g\_\{\\theta\}\(c\)\\right\\\}\.The equality is expected because the experiment constructs the two terms directly\. Its useful role is to confirm the archive bookkeeping and to display the nonlinearity of the uncovered maximum\. All five non\-full\-coverage checkpoints in Table[1](https://arxiv.org/html/2608.08008#S6.T1)violate the fraction\-only heuristic\(1−ρ\)gmax\+0\.0167\(1\-\\rho\)g\_\{\\max\}\+0\.0167, while the finite\-cell certificate holds at every checkpoint\.
The released diagnostic files areverify\_summary\.jsonandverify\_coverage\.csv\. A complete release should additionally store the archive cell identifiers at every budget, allowingUℛU\_\{\\mathcal\{R\}\}and any metric covering radius to be recomputed rather than inferred from plotted values\.
## Appendix EReal\-PRM Discovery Details
Settingtracesevalselitesexploitsρexp\\rho\_\{\\mathrm\{exp\}\}Deceptive, mean6730151213160\.16Deceptive, min673015124100\.00
Table 4:Initial discovery audit onQwen2\.5\-Math\-PRM\-7B\. The evaluation budget is five operator seeds plus4040mutation evaluations per attacked trace \(45×67=301545\\times 67=3015\); “elites” counts the distinct occupied descriptor cells summed over traces, which is smaller because the search revisits cells\. Coverage is the fraction of the2525descriptor cells containing at least one strict exploit; it is not the fraction of natural\-language attack space covered\.Audited setBBr\(B\)r\(B\)L^\\widehat\{L\}εopt\+L^r\(B\)\\varepsilon\_\{\\mathrm\{opt\}\}\+\\widehat\{L\}\\,r\(B\)Top\-5 cells1\.000\.460\.48Top\-10 cells1\.000\.460\.48Top\-15 cells1\.000\.460\.48Top\-20 cells1\.000\.460\.48All 25 \(every op family\)0\.000\.460\.02Table 5:Metric\-cover*plug\-in diagnostic*on the real Qwen archive \(not a certified bound; §[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px8)\)\. Five substitutions prevent a certified reading: \(i\) search assigns non\-exploit cells severity zero, but a failed search is only a*lower*bound on cell severity; \(ii\)L^=0\.46\\widehat\{L\}=0\.46is the largest*observed*loss ratio, a lower bound on the true Lipschitz constant; \(iii\) the metricd=𝟏\{operator differs\}\+\|Δmag\|/4d=\\mathbf\{1\}\\\{\\text\{operator differs\}\\\}\+\|\\Delta\\text\{mag\}\|/4covers descriptor cells, whereas the theorem requires coverage over complete problem–attack instances, sor\(B\)=0r\(B\)=0over the2525cells does not cover natural\-language attack space; \(iv\)ρexp\\rho\_\{\\mathrm\{exp\}\}is substituted forρrep\\rho\_\{\\mathrm\{rep\}\}; \(v\) post\-training spilloverζℛ\\zeta\_\{\\mathcal\{R\}\}is not bounded\. The plug\-in residualεopt\+L^r\(B\)\\varepsilon\_\{\\mathrm\{opt\}\}\+\\widehat\{L\}\\,r\(B\)is additionally conditional on the repair achievingεopt=0\.02\\varepsilon\_\{\\mathrm\{opt\}\}=0\.02onBB\. Within those assumptions it shrinks only when the covering radius does, i\.e\. when every operator family is audited, not merely when many cells are; the finite\-cell plug\-in equalsεopt\\varepsilon\_\{\\mathrm\{opt\}\}because no*searched*uncovered cell carries an exploit\.### Base traces and correctness
The current runs use the first correct GSM8K training solutions with at least two identifiable reasoning steps\. Before an attack is scored, the base final answer is parsed and matched against the dataset answer\. The edited final answer is parsed through the same path and must be wrong\. An unparseable edit is a failed candidate, not an exploit\. Because these examples come from a training split and may have appeared in model pretraining, the experiment measures attack discovery on a controlled trace source rather than out\-of\-distribution generalization\.
Final\-answer verification is sufficient for the correctness\-flip event in Definition[1](https://arxiv.org/html/2608.08008#Thmdefinition1)but does not establish that every inserted sentence is locally coherent\. A stronger release should include an independent audit of fluency, semantic plausibility, and whether the first newly incorrect step is located where the operator intended\. Such validation can be performed on a stratified sample without changing the deterministic final\-answer label\.
### Qwen reader
ForQwen2\.5\-Math\-PRM\-7B, the trace is formatted with the model’s native chat template and step separator\. At each separator position, the implementation extracts the two\-class reward\-head logits, applies a softmax, and retains the positive\-class probability\. Mean and minimum readouts are computed over exactly the same extracted step sequence\. The base and attacked traces must contain at least one valid separator score\. The reader should be validated against a small set of reference examples from the model implementation before attack statistics are produced\.
### RLHFlow reader and normalization
The RLHFlow Llama\-3\.1 PRM\(Xionget al\.,[2024](https://arxiv.org/html/2608.08008#bib.bib11)\)loads as a causal LM \(causal\_steppath\): for each stepiithe reader scores the prefix through stepiifollowed by the literal query “Is this step correct? ” and takes the next\-token probabilityP\(\+\)/\(P\(\+\)\+P\(\-\)\)P\(\\texttt\{\+\}\)/\(P\(\\texttt\{\+\}\)\+P\(\\texttt\{\-\}\)\), so every RLHFlow step score is already a probability in\[0,1\]\[0,1\]—the same scale as the Qwen two\-class softmax readout; no further normalization is applied to either model, and mean/min readouts aggregate these per\-step probabilities directly\.
### Thresholds and dependence
The existing tables useξ=0\\xi=0, so any strictly positive aggregate difference is counted\. This choice is transparent but is not sufficient for a deployment claim\. Recommended reporting includes a threshold curve over absolute score increases, at leastξ∈\{0,10−3,5×10−3,10−2,5×10−2\}\\xi\\in\\\{0,10^\{\-3\},5\\\!\\times\\\!10^\{\-3\},10^\{\-2\},5\\\!\\times\\\!10^\{\-2\}\\\}, together with the maximum and median positive gain\. Thresholds should be applied before transfer is computed\.
Mutations descending from one base trace are correlated\. Confidence intervals should therefore resample base problems, not individual candidates\. A problem\-level bootstrap can recompute attack success, exploit\-cell coverage, maximum gain, and transfer for each replicate\. With only120120problems, maximum\-gain intervals may remain wide; reporting a high quantile such as the9595th\-percentile gain alongside the maximum reduces sensitivity to a single candidate\.
### Syntactic control
The five negative\-control operators split into two groups: three*pure\-corruption*edits \(add a constant to an intermediate number, drop a reasoning step, swap arithmetic operands\) and two*padding*edits \(contentless filler, goal restatement\), none adding deceptive justification\. Every control operator, including the padding pair,*additionally corrupts the final numeric value*\(a fixed\+2\+2shift of the last number\), because Definition 1 requires a correctness flip: “contentless” refers to the inserted text, and the score*rise*comes from the padding while the correctness flip comes from the mechanical final\-number shift—the pure\-corruption operators apply the same shift without padding and never gain score\. In the recorded run \(results\_syntactic/prm\_syntactic\_control\_summary\.json\), the mean readout yields4747strict exploits,*all*from the two padding operators; the three pure\-corruption operators produce none—consistent with the main\-text finding \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px3)\) that any locally innocuous padding, not deception, drives the mean\-pooling failure\. This does not establish that corruption attacks are universally ineffective: the control set is small and no gradient\- or model\-guided corruption is included\.
## Appendix FExcluded Reader Failure
The attempted Skywork\-o1\-Open\-PRM run produced an unparsed native reward channel under the uniform reader\. All extracted per\-cell differences were therefore exactly zero\. The raw diagnostic row was
PRMmean countmin countcoveragereported gainSkywork\-o1\-PRM000\.000\.000
This row is preserved only as a failed\-pipeline record\. It is excluded from Tables[9](https://arxiv.org/html/2608.08008#A7.T9)and[8](https://arxiv.org/html/2608.08008#A7.T8)because a constant reader cannot distinguish robustness from extraction failure\. The corresponding Qwen\-to\-Skywork and RLHFlow\-to\-Skywork transfer values of zero are likewise non\-results\. A corrected study must implement the model\-specific tokenization and reward extraction path, verify nonconstant outputs on documented examples, and rerun every base and attacked trace from the beginning\.
## Appendix GArchive\-Guided Repair Protocol
The following protocol turns the theoretical repair objective into a falsifiable experiment without reusing attack examples for evaluation\. Execution status, stated exactly: the repair stage, fresh adaptive re\-attack, clean\-split problem generalization, MATH\-500 replication \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px9)\), the*paired*pre/post re\-attack with the pairwise\-margin loss plus explicit clean anchor \(three repair seeds, two attack seeds, reconciled union statistics, hierarchical bootstrap\), the equal\-budget*repair\-source*controls \(random, exhaustive, strongest\-only, clean\-only\), the unseen\-wording variant\-B template arm, and the full clean\-utility panel \(§[6](https://arxiv.org/html/2608.08008#S6.SS0.SSS0.Px11)\) are*executed*\. The held\-out operator family \(v1\) returned*no repair pairs*\(every repair\-half exploit belonged to the held\-out family\) and contributes no generalization evidence; an open\-ended attack generator, readout\-side defenses, downstream solver evaluation, and ProcessBench/PRMBench discrimination remain*unexecuted*, are written in the conditional below, and contribute no number to any table\.
StateAUROCBrierBest\-of\-4MarginPre\.884\.3421\.00\.192QD s0/s1/s2\.935/\.915/\.921\.385/\.365/\.3721\.00\.198–\.201Random s0/s1\.938/\.937\.392/\.3901\.00\.196Exhaustive s0/s1\.938/\.939\.392/\.3911\.00\.196–\.197Strongest s0/s1\.941/\.949\.393/\.4021\.00\.194–\.196Clean\-only\.884\.3421\.00\.192*Equal\-budget search baselines \(original model, eval half,**budget4545/problem\): mean best gain / exploit\-cell coverage:*MAP\-Elites\.021\.021/\.40\.40; random\.022\.022/\.40\.40;single\-obj\.\.009\.009/\.18\.18; exhaustive\.024\.024/\.48\.48
Table 6:Top: full clean\-utility panel on the evaluation half \(ranking AUROC of correct vs\. corrupted, Brier of mean\-step score against correctness, mean ranking margin\), per repair condition and seed\. Best\-of\-44is defined exactly: per problem, the clean correct trace versus three corruption variants \(confident\_wrong,plausible\_number,authoritative\_frameat random magnitude\); success iff the clean trace scores strictly highest of the four—a clean\-vs\-corrupted ranking ceiling, not robustness in an attacked candidate pool\. Every adversarial repair improves AUROC and worsens Brier \(score inflation\); clean\-only changes nothing\. Bottom: same\-statistic equal\-budget*search*baselines; the budget exceeds the2525\-cell grid, so exhaustive enumeration is the strongest per\-problem searcher at this scale\.ConditionRate \(k/54k/54\)Mean gainMaxq95Δ\\Deltagain \[CI\]Δ\\DeltaratePre \(original\)\.148 \(8\)\.0212\.333\.146——QD, s0\.037 \(2\)\.0060\.194\.000−\.015\[−\.030,−\.004\]\-\.015\\,\[\-\.030,\-\.004\]−6/54\-6/54QD, s1\.056 \(3\)\.0068\.177\.011−\.014\[−\.028,−\.004\]\-\.014\\,\[\-\.028,\-\.004\]−5/54\-5/54QD, s2\.074 \(4\)\.0092\.212\.034−\.012\[−\.023,−\.003\]\-\.012\\,\[\-\.023,\-\.003\]−4/54\-4/54Random, s0\.037 \(2\)\.0043\.168\.000−\.017\[−\.033,−\.004\]\-\.017\\,\[\-\.033,\-\.004\]−6/54\-6/54Random, s1\.037 \(2\)\.0048\.172\.000−\.016\[−\.032,−\.004\]\-\.016\\,\[\-\.032,\-\.004\]−6/54\-6/54Exhaustive, s0\.037 \(2\)\.0042\.166\.000−\.017\[−\.034,−\.004\]\-\.017\\,\[\-\.034,\-\.004\]−6/54\-6/54Exhaustive, s1\.037 \(2\)\.0048\.173\.000−\.016\[−\.032,−\.004\]\-\.016\\,\[\-\.032,\-\.004\]−6/54\-6/54Strongest, s0\.037 \(2\)\.0053\.191\.000−\.016\[−\.031,−\.004\]\-\.016\\,\[\-\.031,\-\.004\]−6/54\-6/54Strongest, s1\.037 \(2\)\.0048\.187\.000−\.016\[−\.032,−\.004\]\-\.016\\,\[\-\.032,\-\.004\]−6/54\-6/54Clean\-only, s0/s1\.148 \(8\)\.0212\.333\.1460\[0,0\]0\\,\[0,0\]0
Table 7:Paired repair v2 onQwen2\.5\-Math\-PRM\-7B\(GSM8K, evaluation half of5454problems; identical budgets and denominators pre/post\)\. Every statistic, including the rate column, is computed on the*same*per\-problem best\-nonnegative\-gain vector averaged over the two attack seeds, so rate columns subtract exactly toΔ\\Deltarate\.Δ\\Deltacolumns are per\-problem paired bootstrap95%95\\%intervals; all adversarially trained conditions exclude zero\. Repair\-source conditions use an equal pair budget \(cap2020\) and equal steps; QD used2020pairs, random/exhaustive/strongest found1313\. Threshold curves per condition are in the released JSON \(results\_paired\_v2/prm\_paired\_v2\_summary\.json\); the pre\-repair curve falls from0\.1480\.148atτ=0\\tau\{=\}0to0\.0560\.056atτ=0\.1\\tau\{=\}0\.1\.Source PRMTarget PRMstrict transferQwenRLHFlow1\.00RLHFlowQwen0\.93Table 8:Cross\-PRM transfer for mean\-readout exploits\. Transfer records preservation of a strict positive score change, not preservation of the source effect size\. Thresholded transfer curves and problem\-level confidence intervals are required before deployment\-level conclusions\.PRMevals/readoutmeanminmean cov\.gmaxg\_\{\\max\}Qwen2\.5\-Math\-PRM\-7B54004410\.400\.294RLHFlow\-Llama3\.1\-PRM54006456751\.000\.005
Table 9:Larger discovery audit over120120sampled base traces\. The evals/readout column is the*allocated*budget \(120×45=5,400120\\times 45=5\{,\}400\); only the107107traces whose base answer passes exact\-match verification are attacked, consuming107×45=4,815107\\times 45=4\{,\}815evaluations, and all counts, coverage, and problem\-level statistics usen=107n=107as denominator\. “mean” and “min” are candidate\-level strict exploit counts\. The coverage and maximum\-gain columns summarize the mean\-readout archive, matching the archive used for the transfer audit\. Counts should be interpreted together withgmaxg\_\{\\max\}: a strict positive event can be numerically small\.### Problem and attack splits
Problems should be divided into disjoint archive\-training, calibration, and final\-test partitions\. Operator surface forms should be split independently so that held\-out paraphrases instantiate known families, while at least one entire operator family is reserved for an unseen\-mechanism test\. The QD search runs only on the archive\-training problems when constructingA\+A\_\{\+\}\. Hyperparameters, including the marginmm, repair weightλ\\lambda, and operational thresholdξ\\xi, are selected on the calibration problems\.
### Repair and clean anchoring
The PRM is fine\-tuned with the pairwise margin objective from Section[5](https://arxiv.org/html/2608.08008#S5)\. Each batch should mix clean process examples, ordinary incorrect traces, and archived exploits\. A frozen\-copy regularizer can penalize movement of clean step probabilities, while the original clean objective preserves error\-detection ability\. Merely lowering every attacked score is insufficient: the repair must retain or improve the separation of correct and naturally incorrect traces\.
### Fresh adaptive re\-audit
After training, MAP\-Elites is restarted from independent seeds on final\-test problems\. The re\-audit includes known operators with held\-out wording, the reserved operator family, and an open\-ended LLM mutation channel\. The attack budget and PRM\-query count are fixed before comparison\. Cells are added toℛ\\mathcal\{R\}only from this post\-training re\-audit; training\-set attack fit is reported separately\. To measure harmful spillover, the same search is run in cells outsideℛ\\mathcal\{R\}before and after repair, producing an empirical estimate ofζℛ\\zeta\_\{\\mathcal\{R\}\}\.
### Baselines and outcomes
Archive\-guided repair should be compared under equal training\-example and PRM\-query budgets with random attack augmentation, training on globally strongest attacks only, exhaustive evaluation of the discrete2525\-cell grid, and a no\-repair control\. Readout baselines should include mean, minimum, product, last\-step, a lower quantile, and a differentiable soft minimum\. Discovery is reported through attack success, exploit\-cell coverage, QD score, maximum and high\-quantile gain, and semantic uniqueness\. Repair is reported on held\-out and adaptive attacks\. Clean utility is reported through step\-error detection, correct\-versus\-incorrect ranking, calibration, and best\-of\-nnanswer selection\. All intervals should resample problems, and all model\-reader failures should be excluded before aggregation\.
This design separates four outcomes that a single exploit count cannot distinguish: memorization of archived strings, generalization within a descriptor cell, transfer to an unseen attack family, and preservation of clean PRM utility\. Only the latter three support a substantive defense claim\.
## Appendix HSupplementary Figures
We next evaluate the discovered vulnerabilities and the effectiveness of the repair procedure from several complementary perspectives\. Beyond measuring whether exploits can be found, we analyze their prevalence, severity, threshold sensitivity, generalization across reasoning benchmarks, and impact on clean\-task behavior\. Since exploit frequency and exploit magnitude capture different failure modes, we report both problem\-level prevalence and reward gain statistics throughout the evaluation\.
The repair procedure reduces both attack prevalence and attack severity\. Figure[3](https://arxiv.org/html/2608.08008#A8.F3)\(a\) demonstrates a substantial reduction in exploit frequency across repair conditions, while Figure[3](https://arxiv.org/html/2608.08008#A8.F3)\(b\) shows a corresponding suppression of high\-gain attack tails\. Residual maximum gains after fresh adaptive re\-attack indicate that repair reduces, but does not completely eliminate, worst\-case vulnerability\.
Aggregation choice and PRM architecture also affect the observed exploit landscape\. Figure[4](https://arxiv.org/html/2608.08008#A8.F4)\(a\) shows that mean and minimum aggregation produce substantially different strict exploit counts for Qwen2\.5\-Math\-PRM\-7B and RLHFlow\-Llama3\.1\-PRM\. However, exploit count alone does not determine practical risk: Figure[4](https://arxiv.org/html/2608.08008#A8.F4)\(b\) shows that the two PRMs exhibit markedly different maximum reward gains, highlighting that vulnerability depends on both aggregation strategy and the underlying reward model\.
To distinguish strict detections from practically meaningful attacks, we perform a threshold sensitivity analysis\. As shown in Figure[5](https://arxiv.org/html/2608.08008#A8.F5), the exploit rate decreases monotonically from0\.1680\.168atτ=0\\tau=0to0\.0280\.028atτ=0\.2\\tau=0\.2, demonstrating that zero\-threshold exploit counts include a substantial number of low\-margin events\.
We further examine whether the discovered phenomena generalize beyond GSM8K\. Figure[6](https://arxiv.org/html/2608.08008#A8.F6)\(a\) shows that mean readout identifies 41 strict exploits on MATH\-500, while minimum pooling reduces this count to 19 and post\-repair evaluation eliminates all detected exploits\. Figure[6](https://arxiv.org/html/2608.08008#A8.F6)\(b\) shows that minimum pooling primarily reduces severity, decreasing the worst\-case cell gap from0\.2320\.232to0\.0200\.020, whereas repair removes the observed nonnegative\-gain exploits\.
Finally, we evaluate whether repair preserves clean\-task utility\. Figure[7](https://arxiv.org/html/2608.08008#A8.F7)\(a\) shows improved ranking AUROC across repair conditions compared with the original model, while Figure[7](https://arxiv.org/html/2608.08008#A8.F7)\(b\) reveals the corresponding increase in Brier score, consistent with score inflation\. Despite this calibration shift, Figure[7](https://arxiv.org/html/2608.08008#A8.F7)\(c\) confirms that Best\-of\-4 clean ranking remains perfect at1\.001\.00across all conditions, indicating no degradation in the evaluated clean ranking behavior\.
Figure 3:Paired repair analysis across repair conditions on the GSM8K evaluation set\. \(a\) Exploit prevalence, measured as the fraction of problems exhibiting positive\-gain exploits\. \(b\) Attack severity, measured by the maximum observed reward gain and the 95th\-percentile reward gain\. Repair substantially reduces exploit prevalence and tail severity, but residual maximum gains remain after fresh adaptive re\-attack\.Figure 4:Aggregation and PRM dependence of exploit discovery\. \(a\) Candidate\-level strict exploit counts under mean and minimum readouts for Qwen2\.5\-Math\-PRM\-7B and RLHFlow\-Llama3\.1\-PRM\. \(b\) Maximum gain \(gmaxg\_\{\\max\}\) from the mean\-readout archive, showing substantially different attack severity across the two PRMs\. The results show that exploit counts and severity vary with both the aggregation rule and the PRM\.Figure 5:Exploit rate as a function of the gain threshold \(τ\\tau\) on GSM8K\. Increasingτ\\taureduces the exploit rate, showing that strict zero\-threshold counts include many low\-margin events\.Figure 6:Generalization of exploit discovery and repair on MATH\-500\. \(a\) Strict exploit counts decrease from mean to minimum pooling and fall to zero after repair\. \(b\) Maximum cell\-gap severity decreases substantially under minimum pooling and reaches zero under the nonnegative\-gain convention after repair\. The results indicate that minimum pooling primarily suppresses exploit severity, whereas repair eliminates the observed nonnegative\-gain exploits\.Figure 7:Clean\-task ranking and scoring evaluation after repair\. \(a\) Ranking AUROC improves across adversarial repair conditions relative to the original model, indicating improved discrimination between correct and corrupted traces\. \(b\) Brier scores increase after repair, reflecting score inflation despite improved ranking quality\. \(c\) Best\-of\-4 clean\-vs\-corruption ranking remains perfect across all conditions, showing no degradation on this evaluation\.Similar Articles
PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
PRM-as-a-Judge 1.5 is a toolkit that provides fine-grained metrics and reliability tools for evaluating embodied robotic models, moving beyond binary success rates to assess process progress and execution quality.
Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
The paper shows that deterministic PRM guidance in discrete diffusion language models underperforms simpler baselines when compute is matched, identifying two key failures: weak pruning signals and poor final judgment, and releases a corpus and toolkit for reproducible comparisons.
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
KV-PRM introduces a process reward model that leverages KV-cache transfer to avoid re-encoding, achieving up to 5000x FLOP reduction while maintaining or improving performance on reasoning benchmarks.
DEI: Diversity in Evolutionary Inference for Quality-Diversity Search
DEI introduces a distributed Quality-Diversity search framework using heterogeneous LLMs as mutation operators, showing that model diversity improves performance over homogeneous parallel approaches. Evaluated on the Core War domain, a four-node heterogeneous ensemble achieves significant gains in QD-Score and coverage.
Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA
This paper proposes Archive, a framework for ambiguity detection in open-domain QA that distinguishes ambiguity from answer diversity using logical conflict, and introduces QuireQA, a 4,703-query benchmark. Experiments show Archive improves F1 by up to 21.6% while being 16x faster than competitors.