Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Summary
This paper introduces Reasoning Jury, a system that uses a jury of open-weight LLMs with a moderated consensus mechanism to evaluate long reasoning traces, significantly outperforming frontier models at identifying reasoning defects while costing a fraction of the price.
View Cached Full Text
Cached at: 08/14/26, 09:26 AM
# Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Source: [https://arxiv.org/html/2608.12585](https://arxiv.org/html/2608.12585)
Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu,\[4pt\] Mahdi Namazifar\[9pt\]Amazon AGI
###### Abstract
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in\-depth understanding of reasoning behaviors during model performance evaluation\. Additionally, surfacing reasoning mistakes that the model makes would enable improving the model’s performance at runtime through providing feedback\. Due to the difficulty of this complex task on long reasoning traces, single\-model judges \(even frontier models\) do not do well at identifying reasoning defects\. Additionally, leveraging frontier models during online training of reasoning LLMs is generally prohibited due to guardrails in terms of use\. In this work, we introduce*Reasoning Jury*, a system that replaces the single judge with a jury of LLMs and a moderated consensus mechanism, to improve the fidelity of judgments for identifying reasoning defects\. The moderator derives a consensus through deliberation amongst jurors or consolidation of judgements\. We show that Reasoning Jury with a jury of open\-weight models is able to significantly outperform frontier models \(Opus, Sonnet, and Gemini\) at correctly identifying reasoning defects\. Besides accuracy performance improvements, the aggregated cost of the jury \(initial verdicts, deliberations, consolidation, etc\.\) is a fraction \(88to16%16\\%\) of the cost of running frontier models in LLM\-as\-a\-judge setup\. We also show how these judgements can be leveraged to understand failure modes of reasoning LLMs on benchmarks, which allows much deeper understanding of a model’s performance\.
## 1Introduction
Reliable evaluation of reasoning traces supports several stages of development of reasoning LLMs\. Supervised fine\-tuning depends on high\-quality reasoning traces and filtering out low quality data; reinforcement learning depends on reward signals that separate good reasoning from bad; and performance evaluation of reasoning LLMs and error analysis depend on diagnosing reasoning outputs\. All of these depend on the ability to decide how good a reasoning trace actually is\. Unlike non\-reasoning models where evaluation concerns short and self\-contained answers, in reasoning LLMs the content being judged has shifted to long chain\-of\-thought traces that include tens to hundreds of interdependent steps, in which a single bad step could propagate through everything that follows, and in which spotting the defect at all is often highly challenging\. In this setting the evaluation is no longer only “is the final answer right” but “which reasoning step went wrong, and how badly”\. Such fine\-grained evaluations can provide actionable signals for improving other dimensions of reasoning quality, including token efficiency, self\-checking behavior, and confidence calibration\.
Figure 1:Comparing Balanced F1 scores of reasoning defect detection on Hard2Verify benchmark with frontier models \(opus\-4\.6, sonnet\-4\.6, gemini\-3\.1\-pro\) jury vs an open\-weight jury \(3 instances of gpt\-oss\-120b\)\. Dots are individual jurors and stars are jury results \(green for deliberation and orange for consolidation\)\. The frontier models outperform gpt\-oss\-120b individually\. However Reasoning Jury of gpt\-oss\-120b outperforms these frontier models by up to 12 points\. The dollar cost of this Reasoning Jury in deliberation mode is 15\.6% and in consolidation mode is 8% of the cost of opus\-4\.6 at this task \(see[Table2](https://arxiv.org/html/2608.12585#S4.T2)\)\.Identifying defective reasoning traces, especially those with higher severity enables effective training data filtering for mid\-training and SFT stages\. Additionally identifying where exactly such defects occur enables mechanisms to rewrite the defective parts of a reasoning trace either leveraging expert human annotators or using LLMs\. Moreover, such high\-fidelity defect detection could be leveraged during RL to steer reasoning behaviors of the models\. Approaches such as NuRL\[[5](https://arxiv.org/html/2608.12585#bib.bib13)\]\(uses offline\-generated hints during generation\), scaf\-GRPO\[[30](https://arxiv.org/html/2608.12585#bib.bib14)\]\(uses in\-prompt hints when training stagnation is detected\), SDPO\[[16](https://arxiv.org/html/2608.12585#bib.bib15)\]\(uses natural language feedback as hints\), etc\. are RL approaches that rely on rich signals beyond an outcome score\. These signals could include a binary defect based score, natural language hints to avoid defects, or specific tokens where a defect occurred; and RL could strongly benefit from high\-fidelity and rich reasoning defect detection\. Besides model training, new inference time scaling approaches such as\[[1](https://arxiv.org/html/2608.12585#bib.bib16)\]leverage evaluation of the model’s outputs in an iterative loop as input back to the model to improve the model’s performance at run time\. Following this approach, identified reasoning defects could be fed back to the model to correct its reasoning mistakes and improve its performance\. Also in model evaluations, instead of only reporting accuracy scores on benchmarks, reasoning defects, their distributions, and major failure modes could also be reported that would guide the training process and model improvements\.
The dominant solution for such judgements is LLM\-as\-a\-judge with a single strong model\[[32](https://arxiv.org/html/2608.12585#bib.bib2),[14](https://arxiv.org/html/2608.12585#bib.bib3)\]\. This is adequate on short outputs, but breaks down precisely on long reasoning due to the long context of reasoning traces and the complexity of reasoning defect detection task \(results in[Section4\.1](https://arxiv.org/html/2608.12585#S4.SS1)\)\. For such challenging tasks, a single model brings a single perspective to the judging of a reasoning trace, and whatever it misses or misjudges stays incorrect\. Here a natural remedy mirrors how human institutions handle high\-stakes judgements, namely convene a*jury*, where several independent judges drawn from diverse model families have decorrelated blind spots, so a misjudgement by one is often caught by another\. This is a growing line of work where panels of diverse LLMs \(and debate among them\) improve evaluation over any single judge\[[28](https://arxiv.org/html/2608.12585#bib.bib1),[10](https://arxiv.org/html/2608.12585#bib.bib4),[4](https://arxiv.org/html/2608.12585#bib.bib8)\]\. We build on this line and specialize it to the problem of detecting step\-grounded defects in long reasoning traces\.
To this end, we introduce*Reasoning Jury*\. The main assumption of the approach is that every reasoning trace that is input to Reasoning Jury is broken into reasoning “steps” demarcated by\[STEP\-x\]wherexis a step counter\. In turn, every verdict highlighted by Reasoning Jury is grounded to a specific\[STEP\-x\]marker in the trace, which makes each judgement auditable \(a human can check it\) and actionable \(it points at the exact step where a defect occurs\)\. To identify reasoning defects, Reasoning Jury runs a two\-phase process over a panel, where in Phase 1 independent judgements are obtained from the jurors and in Phase 2 a consensus is reached based on these verdicts\.
In Phase 1, each juror reads the problem and the step\-segmented reasoning trace and,*independently*, emits a structured list of defects\. Each defect comes with 1\) a self\-contained description of the defect, 2\) a list of\[STEP\-x\]in which the defect is, 3\) severity of the defect, and 4\) detailed evidence of the defect in the reasoning trace\. A representative defect example for a reasoning trace for solving a math Olympiad problem is shown in[Figure2](https://arxiv.org/html/2608.12585#S1.F2)\. As a result, the jurors commit decorrelated views before any of them can anchor on the others\.
In Phase 2, a consensus from these verdicts is derived\. This phase has two modes, namely Consolidation and Deliberation\. In Consolidation mode the original content along with Phase 1 verdicts are passed to a judge, and the judge is asked to consolidate them\. In Deliberation mode, a moderator runs a multi\-turn debate in which jurors argue over these defects until a consensus is reached or the deliberation is stalled\. The moderator decides which juror speaks on each turn, what the juror should address, and when the discussion has converged enough to terminate\. The moderator then synthesizes the panel’s consensus from the full transcript of deliberations\. It is worth note that the moderator is not given the original problem, reasoning trace, or candidate solution\. It nevertheless sees content\-rich juror arguments, maintains a claim\-level consensus state, selects speakers, and determines termination\. Thus, the design separates direct access to the source material from control of the discussion\. We treat trace withholding as a design heuristic intended to limit direct re\-adjudication, not as a structural guarantee of unbiased moderation\. Its causal effect is not isolated in our experiments\. By design, Phase 1 and Phase 2 preserve a compatible core output schema for individual defect findings, while Phase 2 adds consensus and deliberation metadata where applicable\.
Fatal defectProblem:Letk≥2k\\geq 2be an integer\. Determine all sequences of positive integersa1,a2,…a\_\{1\},a\_\{2\},\\ldotsfor which there exists a monic polynomialPPof degreekkwith non\-negative integer coefficients such thatP\(an\)=an\+1an\+2⋯an\+kP\(a\_\{n\}\)=a\_\{n\+1\}a\_\{n\+2\}\\cdots a\_\{n\+k\}for every integern≥1n\\geq 1\.What went wrong:In\[STEP\-3\]the trace claims that, for the characteristic polynomialrk\+rk−1\+⋯\+r−k=0r^\{k\}\+r^\{k\-1\}\+\\cdots\+r\-k=0, “all roots other thanr=1r=1must have absolute value strictly less than11\.” This is false: fork=2k\{=\}2the polynomial factors as\(r−1\)\(r\+2\)\(r\-1\)\(r\+2\), giving a rootr=−2r=\-2with\|r\|=2\>1\|r\|=2\>1\. The trace’s triangle\-inequality argument only rules out roots on the unit circle other than11; it never excludes\|r\|\>1\|r\|\>1\. This false claim is exactly what the derivation needs to conclude thatdn=an\+1−and\_\{n\}=a\_\{n\+1\}\-a\_\{n\}is eventually constant, so the rest of the case rests on a false foundation\.Statement references:\["STEP\-3"\]Impact:FatalEvidence:“ So all roots other thanr=1r=1must have absolute value strictly less than11\. The general solution to the recurrence isdn=A⋅1n\+∑iPi\(n\)rind\_\{n\}=A\\cdot 1^\{n\}\+\\sum\_\{i\}P\_\{i\}\(n\)r\_\{i\}^\{n\}where\|ri\|<1\|r\_\{i\}\|<1\. Asn→∞n\\to\\infty,dn→Ad\_\{n\}\\to A\. ”
Figure 2:A representative defect emitted by a juror in Phase 1, on an olympiad problem from Hard2Verify\. The juror grounds its judgement to a specific\[STEP\-x\], gives a self\-contained explanation of the mistake, tags a severity, and quotes the trace as evidence\.To evaluate Reasoning Jury we use reasoning defect benchmarks Hard2Verify\[[25](https://arxiv.org/html/2608.12585#bib.bib6)\]and DeltaBench\[[15](https://arxiv.org/html/2608.12585#bib.bib18)\]\. On these benchmarks we show that a jury of open\-weight models could easily outperform frontier models\. Additionally, although jury consolidation or deliberation consumes significantly more tokens than a single model judge, we show that Reasoning Jury costs a fraction compared to running a frontier LLM as a judge\. As there are legal and practical limitations on using closed\-weight frontier models for training LLMs, the jury of open\-weight models with high fidelity would provide a flexible path to use such juries for LLM training\.
Our benchmark evaluations measure whether the jury localizes defective steps, but they do not directly validate the factual accuracy ofwhat\_went\_wrong, the calibration of severity labels, or the general downstream usefulness of these fields\. A comprehensive assessment of those properties across models and tasks remains future work\. We nevertheless provide an initial task\-based test of whether richer defect descriptions help a model repair its reasoning\. For each AIME2026 problem, we generate6464solutions with nemotron\-3\-super \(1,9201\{,\}920traces\), evaluate every trace with the jury, and ask the same model to retry each of the877877traces flagged with at least one non\-neutral defect\. Supplying thewhat\_went\_wrongdiagnosis and severity raises retry accuracy to76\.2%76\.2\\%from71\.2%71\.2\\%with step locations alone\.[Section6](https://arxiv.org/html/2608.12585#S6)summarizes the experiment, and[AppendixN](https://arxiv.org/html/2608.12585#A14)provides the full design and results\.
## 2Reasoning Jury
### 2\.1Segmenting the trace into steps
Grounding of detected defects of a reasoning trace would require some way to reference where exactly the defect occurs within a long reasoning trace\. Splitting a long reasoning trace into*reasoning steps*would enable this grounding where a juror can cite for example\[STEP\-14\]when claiming a defect\. In order to achieve this reasoning step segmentation the vast majority of work in the literature insert a step boundary at every double newline \(the pattern\\n\\n\)\[[17](https://arxiv.org/html/2608.12585#bib.bib5),[31](https://arxiv.org/html/2608.12585#bib.bib9),[33](https://arxiv.org/html/2608.12585#bib.bib10),[6](https://arxiv.org/html/2608.12585#bib.bib11),[27](https://arxiv.org/html/2608.12585#bib.bib12)\]\. Although simple, this approach has its drawbacks, including potentially placing step boundaries in the middle of a code or pseudo\-code snippet, a multi\-line algebraic derivation, or a simple enumeration of different cases\. Additionally if a model simply uses a single newline instead of double, or in general uses line breaks less frequently this approach becomes less robust\. Another approach for this could be using a strong LLM as a one\-shot segmenter, where an LLM is prompted to add step markers, and produce semantically coherent steps for a long reasoning trace\. But the same long\-content self\-instability that motivates this paper applies here as well where segmentation is not stable across re\-samples, and most importantly the large LLM may not remain fully faithful to the main reasoning text and makes modifications to it, and as a result, the segmented reasoning trace is different from the original reasoning trace\. Ideally a dedicated, specialized reasoning step segmentation model would address this need; but absent that, we use the pattern of an end\-of\-sentence character followed by\\n\\n\(regex pattern\(?<=\[\.\!?\]\)\\s\*\\n\\s\*\\n\+\) to segment a reasoning trace to reasoning steps\. The addition of end\-of\-sentence characters improves the robustness of the pattern for traces including code and multi\-step algebraic derivations\.
### 2\.2Reasoning Jury Pipeline
#### 2\.2\.1Phase 1: independent judgement
All jurors receive the same prompt that includes the original problem, the step\-annotated trace, and the final solution\. Each independently returns a JSON object of verdicts, where each verdict carrieswhat\_went\_wrong, a self\-contained description of the defect;statement\_refs, a list of steps that the defect refers to;impact, the severity of the defect \(𝚗𝚎𝚞𝚝𝚛𝚊𝚕\|𝚖𝚒𝚗𝚘𝚛∣𝚖𝚊𝚓𝚘𝚛∣𝚏𝚊𝚝𝚊𝚕\\mathtt\{neutral\}\\\!\\mid\\\!\\mathtt\{minor\}\\\!\\mid\\\!\\mathtt\{major\}\\\!\\mid\\\!\\mathtt\{fatal\}\); andevidence, the supporting evidence of the identified defect\. The full prompt can be found in[AppendixB](https://arxiv.org/html/2608.12585#A2)\. It enforces a*genericness test*\(“could this comment apply to a different problem with no edits? if so, rewrite it”\) and a*specificity self\-check*\(1–5; include only issues scoring≥4\\geq 4\) to suppress vague, non\-grounded criticism\.what\_went\_wrongtries to capture weaknesses in the reasoning trace by pointing to bad reasoning moves\. It is intentionally kept high\-level and open\-ended with requirements on specificity to not limit the jurors in identifying different kinds of defects\. Phase 1 calls fan out in parallel, and the pipeline proceeds as long as a configurable minimum number of jurors succeed\.
#### 2\.2\.2Phase 2: Consensus
Consensus from Phase 1 independent judgements is reached in two different modes, namely Consolidation and Deliberation\.
##### Consolidation\.
In this mode a judge is asked to perform the task as Phase 1, except it is also given all the verdicts from Phase 1\. The judge \(moderator\) is given instructions on how to verify, merge, and fill in the gaps in the Phase 1 judgements\. The final output of Consolidation is in the format of Phase 1 outputs\. The full prompt is provided in[SectionC\.1](https://arxiv.org/html/2608.12585#A3.SS1)\.
##### Deliberation\.
The jury deliberations start based on Phase 1 verdicts\. The moderator \(see full prompt in[SectionC\.2](https://arxiv.org/html/2608.12585#A3.SS2)\) is intentionally designed to be blind to the problem, the reasoning trace, and the candidate solution\. It sees only the deliberation transcript, and its role is purely procedural \(manage deliberation turns, surface disagreements, detect convergence\), never speculating about reasoning trace content or hinting at the right answer\. Deliberation is organized into logical rounds, within each of which every juror must speak at least once\. Each turn proceeds in three parts\.
First, the moderator selects the next speaker and issues a process\-oriented instruction, prioritising jurors involved in unresolved disagreements, recalling any juror silent for two or more turns, and asking the chosen juror to clarify a specific point and defend or concede it\. Second, the selected juror \(see full prompt in[SectionC\.3](https://arxiv.org/html/2608.12585#A3.SS3)\) re\-reads the problem and the reasoning trace, as well as the moderator’s instructions, and contributes a natural\-language argument \(agreeing, disagreeing with evidence, raising a new defect, or conceding\) addressing the moderator’s asks\. Third, the moderator folds the contribution into a running consensus state and checks for deliberation termination conditions\. The loop terminates on a full logical round with no new substantive argument, on universal agreement, or on detected cycling, with a hard cap on total turns as a backstop regardless\. A final extraction call to the moderator receives the Phase 1 judgements, the full transcript, and the moderator’s running deliberations\. The moderator then synthesizes the*consensus defects*\(a defect reaches consensus if a dynamically computed majority endorsed it, or if it was raised and never contested\), a*confidence*in\[0,1\]\[0,1\]reflecting the degree of agreement,*dissenting views*that did not reach consensus, and the*termination reason*\. The majority threshold is computed from the number of jurors that actually succeeded, so partial failures do not silently change the voting rule\.
The integrity of the deliberation rests on a deliberate division of labor among non\-juror roles\. In deliberation mode, if the moderator could see the trace, it would inevitably form opinions about the defects, and those opinions would leak into*whom it calls on*and*what it tells them to address*, making it a covert extra juror with the unique power to suppress dissent by never giving it the floor\. By restricting the moderator to the*conversation only*\(speaker names and message content, never the problem or trace\), its decisions are necessarily about*argument dynamics*\(who has not spoken, which disagreement is unresolved, whether the round produced anything new\) not about content\.
## 3Experimental Setup
We now describe how we evaluate Reasoning Jury, the benchmarks and the evaluation procedures\.
### 3\.1Benchmarking Accuracy of Reasoning Jury
Since reasoning defect detection is an under\-studied task in the published literature, there are very few available datasets to leverage for benchmarking this task\. Our evaluations mainly use Hard2Verify\[[25](https://arxiv.org/html/2608.12585#bib.bib6)\]which is a human\-annotated, step\-level verification benchmark for open\-ended advanced mathematics competitions with 200 records\. These records have approximately 780 gold error steps \(on average around 9\.3 steps per record, of which around 3\.9 are labeled erroneous\)\. Additionally we also use DeltaBench\[[15](https://arxiv.org/html/2608.12585#bib.bib18)\], which is a benchmark for this task covering STEM, coding, and general reasoning with 1,236 records, and includes defect localization over reasoning traces\. The reasoning traces in these benchmarks are substantially longer than those of prior step\-level benchmarks such as ProcessBench\[[31](https://arxiv.org/html/2608.12585#bib.bib9)\]\. For that reason we do not consider using ProcessBench in this work \(discussed further in[AppendixI](https://arxiv.org/html/2608.12585#A9)\)\. For cost related concerns, the vast majority of our evaluations are done on the smaller benchmark Hard2Verify, but we also provide detailed results for a subset of evaluations for DeltaBench\.
In order to evaluate reasoning jury on these benchmarks we take the rich and detailed step\-level detected defects and turn them into step level binary signals\. If a step was highlighted in a detected defect \(in thestatement\_refsfield of the defect\) with severity*minor*,*major*, or*fatal*, that step is labeled as 1; otherwise step labels are 0\. Following Hard2Verify paper we mainly look at Balanced F1, the harmonic mean of error\-recall and specificity, micro\-averaged over steps\. Alongside Balanced F1 we report Balanced Accuracy, Accuracy, Precision, Recall, and F1\. Note that we evaluate only step\-level defect localization\. This scoring discards what went wrong, evidence, confidence, dissent, and the distinction among minor, major, and fatal severity\. Consequently, these experiments do not evaluate explanation factuality, evidential support, severity calibration, semantic deduplication, or downstream usefulness\.
The Hard2Verify paper uses a simple prompt to output a list of verdicts for each step \(correct or incorrect\) as well as a list of reasoning for each verdict\. The prompt that we leverage in this paper is much more elaborate and produces verdicts with severity tags \(which helps with reasoning data curation and filtering and provides reliable and detailed signals for RL\), as well as mechanisms to ensure defect specificity that, intuitively speaking, helps in root causing issues in other domains\. The Hard2Verify paper reports Balanced F1 score of 85\.8 with their prompt using gpt\-5\. We replicate that experiment with gpt\-5\.4 \(which is the version we use in this paper\) and we get the Balanced F1 score of 85\.3\. Using our prompt \([AppendixB](https://arxiv.org/html/2608.12585#A2)\) with gpt\-5\.4 we get the Balanced F1 score of 83\.9\. Based on these numbers and the additional utilities that our prompt provides, for the rest of the experiments we leverage our prompt for reasoning defect detection\.
We score against the benchmarks’ own gold step boundaries and gold error labels\. We do*not*re\-segment solutions for scoring purposes\. This keeps the step\-level comparison apples\-to\-apples against the human annotation\. As an additional point, the Hard2Verify authors evaluate 29 generative critics and process reward models, and the strongest process reward models score in the range 0\.35–0\.60 Balanced F1 at the step level\. We cite this range as context for what a strong model achieves on this task, without importing any per\-model figure\.
### 3\.2Evaluation Configuration
In all of our experiments we set the reasoning effort of all LLM calls to “High”\. Temperature 1\.0 is used across all LLM calls\. For gpt\-5\.4 calls \(except gpt\-oss\-120b\) we use Amazon Mantle API, for opus\-4\.6 and sonnet\-4\.6 we use Amazon Bedrock API, for gemini\-3\.1\-pro we use OpenRouter API, and for all open\-weight models we serve them on a local cluster using vLLM\.
## 4Results
Most of our results in this section are on the Hard2Verify benchmark, and we also report results on DeltaBench for some of the key experiments in[Section4\.6](https://arxiv.org/html/2608.12585#S4.SS6)\. Throughout this section we mostly report on Balanced F1 score at step level, and we provide full tables in[AppendixG](https://arxiv.org/html/2608.12585#A7)\. For Hard2Verify, because each score is measured on roughly 200 records, every number carries sampling uncertainty; we report 95% bootstrap confidence intervals \(10,000 record\-level resamples\) on the Balanced\-F1\.
### 4\.1Single\-model as a judge
We first establish how well individual models detect reasoning defects\. From[Figure3](https://arxiv.org/html/2608.12585#S4.F3), it is clear that gpt\-5\.4 at83\.983\.9Balanced F1 is a saturated outlier that sits roughly ten points clear of the next best single model, opus\-4\.6 at73\.773\.7\. The remainder of the frontier tier \(sonnet\-4\.6 at71\.271\.2, gemini\-3\.1\-pro at69\.569\.5\) and the strongest open model glm\-5\.2 at69\.469\.4performs lower than opus\-4\.6\. The open\-weight models as a whole span a wide band, from roughly5050to7070\.
The contrast between the two leading models is notable\. gpt\-5\.4’s lead is through a balance between precision and recall in the high7070s to low8080s \([Table4](https://arxiv.org/html/2608.12585#A7.T4)\), whereas opus\-4\.6 has a high precision \(81\.781\.7\) but a low recall \(62\.562\.5\), and it misses many genuine defects\. This asymmetry foreshadows the aggregation results below, where combining jurors mostly recovers recall\.
### 4\.2Jury versus a single judge
Next we move on from a single judge to a jury with a consensus verdict\. For these experiments we use Hard2Verify benchmark and we create the following juries\. \(1\) Frontier which includes frontier proprietary models gpt\-5\.4\[[24](https://arxiv.org/html/2608.12585#bib.bib19),[23](https://arxiv.org/html/2608.12585#bib.bib20)\], opus\-4\.6\[[2](https://arxiv.org/html/2608.12585#bib.bib21)\], sonnet\-4\.6\[[3](https://arxiv.org/html/2608.12585#bib.bib22)\], gemini\-3\.1\-pro\[[13](https://arxiv.org/html/2608.12585#bib.bib23)\]\. We also include glm\-5\.2\[[12](https://arxiv.org/html/2608.12585#bib.bib24),[29](https://arxiv.org/html/2608.12585#bib.bib25)\]which is the largest open\-weight model in our mix of models to the Frontier jury\. \(2\) Frontier−\-gpt\-5\.4 which is the Frontier jury excluding gpt\-5\.4\. Through this jury we study the removal of the outsized performance of gpt\-5\.4 from the jury\. \(3\) Large OSS, which includes the largest open\-weight models, namely glm\-5\.2, deepseek\-v4\-pro\[[8](https://arxiv.org/html/2608.12585#bib.bib26)\], kimi\-k2\.6\[[20](https://arxiv.org/html/2608.12585#bib.bib28)\], nemotron\-3\-super\[[21](https://arxiv.org/html/2608.12585#bib.bib27)\], and minimax\-m3\[[19](https://arxiv.org/html/2608.12585#bib.bib29)\]\. \(4\) Small/Medium OSS models, which include qwen3\.6\-27b\[[26](https://arxiv.org/html/2608.12585#bib.bib32)\], nemotron\-3\-super, gpt\-oss\-120b\[[22](https://arxiv.org/html/2608.12585#bib.bib30)\], gemma\-4\-31b\[[11](https://arxiv.org/html/2608.12585#bib.bib33)\], minimax\-m2\.7\[[18](https://arxiv.org/html/2608.12585#bib.bib31)\]\.[Table1](https://arxiv.org/html/2608.12585#S4.T1)reports the consolidation and deliberated final verdict for the four juries across 6 metrics, and[Figure3](https://arxiv.org/html/2608.12585#S4.F3)places each jury’s consensus against its constituent jurors’ Balanced F1 scores\. From the figure it is clear that at84\.484\.4Balanced F1 the Frontier jury is near\-saturated by gpt\-5\.4’s solo score of83\.983\.9, so the jury adds almost nothing on top of its dominant member\. In Frontier−\-gpt\-5\.4 we remove the dominant gpt\-5\.4 from the jury while keeping the remaining four jurors’ Phase 1 judgements fixed, and a Balanced F1 of81\.781\.7is achieved, which is 8 points above its best juror opus\-4\.6, and it lands within 3 points of the gpt\-5\.4 anchored jury\. Here a conclusion is that when one juror already saturates the task, consensus tracks that juror, and when that is not the case, aggregation across the merely\-strong jurors does improve the performance\.
Among open\-weight models, the deliberating OSS juries substantially outperform every constituent juror\. The strongest individual OSS jurors achieve only65\.465\.4–69\.469\.4Balanced F1, whereas the Small/Medium and Large OSS juries reach80\.880\.8and80\.280\.2, gains of15\.415\.4and10\.810\.8points, respectively\. These gains are driven primarily by improved recall \([Table4](https://arxiv.org/html/2608.12585#A7.T4)\), while maintaining precision near8080\. Notably, the Small/Medium OSS jury outperforms opus\-4\.6, sonnet\-4\.6, and gemini\-3\.1\-pro by7\.17\.1–11\.311\.3points and performs on par with the Large OSS jury\.
Across all four juries, deliberation yields numerically higher Balanced F1 than consolidation, with differences ranging from0\.40\.4to1\.51\.5points \([Table1](https://arxiv.org/html/2608.12585#S4.T1)\)\. For the OSS juries, the distinction is primarily a precision–recall trade\-off: deliberation raises recall from71\.371\.3to75\.075\.0for Large OSS and from71\.371\.3to77\.477\.4for Small/Medium OSS, whereas consolidation raises precision from79\.679\.6to81\.181\.1and from78\.378\.3to82\.982\.9, respectively\. Thus, deliberation recovers more true defects, while consolidation produces more conservative verdicts\.
The Balanced\-F1 confidence intervals in[Table4](https://arxiv.org/html/2608.12585#A7.T4)reflect sampling uncertainty over the approximately200200Hard2Verify records\. Single\-model estimates have wider intervals \(±3\\pm 3to±7\\pm 7points\), while the deliberated consensus intervals range from±2\.5\\pm 2\.5to±3\.1\\pm 3\.1points\. The exception is gpt\-5\.4, whose solo interval \(83\.9±2\.883\.9\\pm 2\.8\) is comparable to the Frontier consensus\. Full per\-juror intervals are provided in[Table4](https://arxiv.org/html/2608.12585#A7.T4)\([AppendixG](https://arxiv.org/html/2608.12585#A7)\)\.
Table 1:Jury verdict for four juries under both Phase 2 modes: deliberation and consolidation\. Both modes share identical Phase 1 judgements per jury; for Frontier−\-gpt\-5\.4 the Phase 1 judgements are inherited from the Frontier run \(dropping only the gpt\-5\.4 seat\), so that row isolates exactly the effect of removing the dominant juror\. All six metrics are step\-level, with severity minor and above\. OSS juries lift Balanced F110\.810\.8to15\.415\.4over their best solo juror under deliberation\. Consolidation retains most of that lift giving up0\.40\.4–1\.51\.5Balanced F1 by trading recall for precision\.Figure 3:Solo jurors \(dots\) versus deliberated consensus \(star\) for each jury, with the lift over best juror on the right\. The open\-source juries gain the most; the full Frontier jury is near\-saturated by gpt\-5\.4, so its lift is small, whereas removing gpt\-5\.4 restores a large lift\. Exact numbers, with95%95\\%confidence intervals, in[Table4](https://arxiv.org/html/2608.12585#A7.T4)\.
### 4\.3How many jurors are needed?
To understand the impact of the size of the jury on its performance, for Hard2Verify we drop jurors from the jury one at a time\. In each run, we drop the juror contributing the fewest*unique*gold defective steps \(steps that no other juror in the jury also caught\), computed from the preceding run’s Phase 1 outputs\. To isolate jury size from sampling noise, we hold Phase 1 fixed: each smaller jury reuses the exact independent Phase 1 judgements from the full jury’s run, and only Phase 2 with both deliberation and consolidation are run over the surviving subset\.[Figure4](https://arxiv.org/html/2608.12585#S4.F4)depicts the results\. The full Frontier jury performs the same from five to two jurors \(84\.4→84\.5→84\.5→84\.784\.4\\to 84\.5\\to 84\.5\\to 84\.7for deliberation\) because gpt\-5\.4 carries the verdict regardless of what other models are in the jury\. The Frontier−\-gpt\-5\.4 and Large OSS juries decline gently as jurors are removed, to80\.180\.1and77\.177\.1respectively with deliberation at two jurors\. The Small/Medium OSS jury holds well through three jurors \(80\.8→79\.7→76\.880\.8\\to 79\.7\\to 76\.8with deliberation\) and then drops more sharply to73\.573\.5with two jurors\.
Taken together, three jurors provide a reasonable operating point for the Frontier, Frontier−\-gpt\-5\.4, and Large OSS panels, remaining within1\.31\.3points of the largest jury\. For Small/Medium OSS, four jurors appear preferable: reducing from five to four costs only1\.11\.1points, whereas reducing to three costs4\.04\.0points\.
Figure 4:Consensus versus solo jurors as each jury is reduced from five jurors down to two\. Exact numbers in[Table5](https://arxiv.org/html/2608.12585#A7.T5)\.
### 4\.4Which model should moderate?
The moderator runs the jurors’ arguments and issues the final verdict, so a question is how much the moderator choice matters when the jury itself is held fixed\.[Figure5](https://arxiv.org/html/2608.12585#S4.F5)sweeps the moderator over each jury while leaving the jurors unchanged\. Moderator choice with deliberation moves consensus by3\.73\.7and2\.62\.6points for the two OSS juries, namely Large OSS and Small/Medium OSS\. With consolidation, the corresponding ranges are6\.46\.4and6\.16\.1points, respectively, which shows that consolidation is more sensitive to the choice of moderator\. In consolidation mode we see that a weak moderator \(e\.g\., qwen3\.6\-27b\) could cause a77point drop in the performance of the jury compared to deliberation\. The results show that the moderator plays a key role in the performance of the jury, and it should therefore be chosen carefully and with caution\.
Figure 5:Moderator choice changes deliberation by2\.62\.6–3\.73\.7points and consolidation by6\.16\.1–6\.46\.4points\. Deliberation values are shown in[Table6](https://arxiv.org/html/2608.12585#A7.T6); consolidation values are shown in the figure\.
### 4\.5Does jury diversity matter?
We ask whether the jury’s gains come from model diversity or simply from aggregating multiple opinions\. To isolate this we build homogeneous juries: three independent samples of a single model, moderated by that same model, so that all diversity comes from sampling rather than from mixing architectures \([Figure6](https://arxiv.org/html/2608.12585#S4.F6)\)\. A homogeneous gpt\-oss\-120b jury with deliberation reaches82\.382\.3Balanced F1 consensus, a lift of\+13\.0\+13\.0against the best solo score and significantly higher than opus\-4\.6, sonnet\-4\.6, and gemini\-3\.1\-pro\. A homogeneous qwen3\.6\-27b jury reaches72\.572\.5against60\.760\.7solo, a lift of\+11\.8\+11\.8\. In both cases precision stays near\-constant and the lift is driven by recall, consistent with the pattern in[Section4\.2](https://arxiv.org/html/2608.12585#S4.SS2)\.
These results indicate that model diversity is not necessary for large deliberation gains and, in this setting, matters less than model capability\. The homogeneous gpt\-oss jury reaches82\.382\.3Balanced F1, within the roughly8080to8484range of the diverse juries\. In contrast, the homogeneous qwen3\.6\-27b jury reaches only72\.572\.5, despite achieving a comparable improvement over its solo baseline\. Thus, deliberation across independent samples can be effective without mixing models, but the final performance remains strongly constrained by the capability of the underlying model\. These results do not rule out additional benefits from model diversity\.
Figure 6:Diversity control: homogeneous juries of three independent temperature\-1\.01\.0samples of one model with a same\-model moderator\. Jury deliberation lifts consensus over the best single solo even with zero diversity \(\+13\.0\+13\.0for gpt\-oss\-120b,\+11\.8\+11\.8for qwen3\.6\-27b\), but the ceiling tracks base\-model strength\. Exact numbers in[Table7](https://arxiv.org/html/2608.12585#A7.T7)\.
### 4\.6Generalization to DeltaBench
In this section we evaluate Reasoning Jury on DeltaBench\[[15](https://arxiv.org/html/2608.12585#bib.bib18)\], which includes long reasoning traces with human\-annotated erroneous steps\. We compare a homogeneous3×3\\timesgpt\-oss\-120b jury against a single opus\-4\.6 Phase 1 pass, on the1,2261\{,\}226records\. For evaluation details refer to[AppendixM](https://arxiv.org/html/2608.12585#A13)\.[Figure7](https://arxiv.org/html/2608.12585#S4.F7)shows that the jury’s deliberated consensus reaches61\.561\.5Balanced F1 against58\.458\.4for opus\-4\.6, with each individual gpt\-oss juror below both \(53\.353\.3–55\.755\.7\), and we see that deliberation lifts the panel\+5\.8\+5\.8over its best member \(full table in[AppendixM](https://arxiv.org/html/2608.12585#A13)\)\.
Figure 7:DeltaBench generalization: Balanced F1 for a single opus\-4\.6 Phase 1 pass versus a homogeneous3×3\\timesgpt\-oss\-120b jury with deliberation\. Every individual gpt\-oss\-120b juror trails opus by around33–55points, but deliberation lifts the panel\+5\.8\+5\.8over its best member and past opus\-4\.6 and at the same level as gpt\-5\.4\. Under DeltaBench’s own macro\-F1 protocol \(see[AppendixM](https://arxiv.org/html/2608.12585#A13)\) the jury deliberation scores47\.947\.9vs\.43\.443\.4for opus\-4\.6 \([AppendixM](https://arxiv.org/html/2608.12585#A13)\)\.
### 4\.7Jury deliberation vs majority vote or union of defects
One natural question is what if jury deliberation or consolidation is replaced with something simple such as majority vote or a simple union of defects\.[Tables8](https://arxiv.org/html/2608.12585#A8.T8)and[9](https://arxiv.org/html/2608.12585#A8.T9)report the full metric breakdown for both baselines against deliberation across all six panels\.
Majority voting consistently underperforms because exact\-step agreement is sparse\. This collapses recall to 38–63% despite precision holding at 86–93%, dragging Balanced F1 roughly 9–26 points below consensus\. Union aggregation captures most of the benefit of multiple independent judgements\. Deliberation outperforms union in four of six panels, but underperforms it in two\. Thus, the principal accuracy gain comes from pooling independent judgements, but it is just a bag of possibly\-duplicate, possibly\-contradictory raw claims with no resolution of severity or evidence quality\. Consolidation or deliberation, on the other hand, primarily adjudicates, deduplicates, and structures those judgements rather than uniformly improving step\-level Balanced F1\. These modes produce a single deduplicated, evidence\-backed verdict, which has real downstream applications \([Section1](https://arxiv.org/html/2608.12585#S1)\) independent of any metric\. However, since benchmark scoring ignores explanation text, severity distinctions, evidence, and dissent, these results do not establish that consolidation or deliberation improves those properties\. A more detailed analysis of this is available in[AppendixH](https://arxiv.org/html/2608.12585#A8)\.
### 4\.8Jury cost analysis
How much does a deliberating jury cost compared to a single strong model? We compare two configurations that ran on the same200200Hard2Verify records: a3×3\\timesgpt\-oss\-120b homogeneous jury against a single opus\-4\.6 Phase 1 call \(one model judging alone\)\. Token counts are taken from each run’s logged LLM calls, and pricing is per\-million tokens at list rates\.111Pricing used: opus\-4\.6 $5\.00 / 1M input, $25\.00 / 1M output\. gpt\-oss\-120b $0\.15 / 1M input, $0\.62 / 1M output\. Source: https://aws\.amazon\.com/bedrock/pricing/ as of July 2026
3×3\\timesgpt\-oss, consolidation3×3\\timesgpt\-oss, deliberationMetricopus \(single\)ValueRatioValueRatioOutput tokens2,905,7192\{,\}905\{,\}7199,276,7939\{,\}276\{,\}7933\.2×3\.2\\times16,106,48916\{,\}106\{,\}4895\.5×5\.5\\timesInput tokens1,297,8661\{,\}297\{,\}8664,975,3544\{,\}975\{,\}3543\.8×3\.8\\times15,765,86715\{,\}765\{,\}86712\.1×12\.1\\timesTotal tokens4,203,5854\{,\}203\{,\}58514,252,14714\{,\}252\{,\}1473\.4×3\.4\\times31,872,35631\{,\}872\{,\}3567\.6×7\.6\\timesLLM calls2002008008004\.0×4\.0\\times2,1732\{,\}17310\.9×10\.9\\timesInput cost$6\.496\.49$0\.750\.750\.12×0\.12\\times$2\.362\.360\.36×0\.36\\timesOutput cost$72\.6472\.64$5\.755\.750\.08×0\.08\\times$9\.999\.990\.14×0\.14\\timesTotal cost \(200 rec\)$79\.13$6\.500\.08×\\mathbf\{0\.08\\times\}$12\.350\.16×\\mathbf\{0\.16\\times\}Cost per record$0\.3960\.396$0\.0330\.0330\.08×0\.08\\times$0\.0620\.0620\.16×0\.16\\timesBalanced F173\.773\.780\.380\.3—82\.382\.3—Table 2:Cost comparison on Hard2Verify: a single opus\-4\.6 Phase 1 call versus the3×3\\timesgpt\-oss\-120b panel under both Phase 2 modes, consolidation and deliberation\. Ratios are relative to the opus column\. Deliberation uses7\.6×7\.6\\timesmore tokens than opus yet costs6\.4×6\.4\\timesless; consolidation cuts the jury’s cost roughly in half again \($6\.506\.50,12×12\\timescheaper than opus\) while giving up22Balanced F1 points of the deliberation lift\. Both jury modes are more accurate than the frontier single judge\.[Table2](https://arxiv.org/html/2608.12585#S4.T2)reports the comparison\. In deliberation mode, despite using7\.6×7\.6\\timesmore tokens in total and making10\.9×10\.9\\timesmore LLM calls, the three\-juror gpt\-oss jury costs$12\.35for the full200200\-record run, versus$79\.13for a single opus pass, which is roughly6\.4×6\.4\\timescheaper\. The per\-token price difference \(opus\-4\.6 $25/M output versus gpt\-oss $0\.62/M\) more than compensates for the higher token volume\. In consolidation mode we see that the cost drops to$6\.50\\$6\.50\. On accuracy the jury with deliberation scores82\.382\.3Balanced F1 and with consolidation scores80\.380\.3against opus\-4\.6’s73\.773\.7, and gpt\-5\.4’s83\.983\.9\. Compared to opus\-4\.6 the jury is both more accurate and substantially cheaper \(up to 12×\\times\)\. Compared to gpt\-5\.4 the performance of the jury is similar while still at a fraction of its cost\. The practical implication is that when the jurors are cheap open\-weight models, the deliberation overhead is dwarfed by the per\-token savings relative to a single expensive frontier judge\. Multi\-turn deliberation comes with additional latency cost which should also be considered\.
## 5Exploratory jury output profiling
The results so far use the jury as an evaluator scored against a benchmark\. Because its output is a structured, step\-grounded, severity\-tagged list of defects, the jury can also be used as an*instrument*to run over many reasoning traces from a given model and aggregate the defects into a distribution over defect kinds\. This especially would add a rich layer of analysis and insights on top of performance benchmarks, and would go beyond accuracy numbers to surface detailed failure modes\. This section applies Reasoning Jury this way to reasoning traces for AIME2026\[[9](https://arxiv.org/html/2608.12585#bib.bib17)\]generated by nemotron\-3\-super as a case study, producing a reasoning defect fingerprints over a shared taxonomy of defect types for math problems\. Such reasoning defect analysis does require a taxonomy as a pre\-requisite\. For demonstration of this use case, to achieve this taxonomy, we categorize reasoning defects from different models on this benchmark\. Next we map defects from nemotron\-3\-super to this taxonomy and analyze the results\. We should emphasize that this taxonomy is not a comprehensive and validated taxonomy of defects for math reasoning, and it’s created simply for demonstration purposes, and it does not establish a standard\. Also since the jury natural language verdicts are not validated in our experiments the identified defects cannot be taken at face value\.
### 5\.1Setup
We generate reasoning traces from gpt\-oss\-120b, qwen3\.6\-27b, and gemma\-4\-31b on the AIME2026 benchmark\. For each model we draw 32 samples per problem\. Every reasoning trace is segmented into\[STEP\-x\]units with the regular\-expression segmenter of[Section2\.1](https://arxiv.org/html/2608.12585#S2.SS1), and Reasoning Jury is run over each trace; the final verdict’s defects are the signal we aggregate\. The jury \(glm\-5\.2, minimax\-m3, deepseek\-v4\-pro, kimi\-k2\.6\) is*disjoint*from the three generators under study, so no model grades its own traces\. This is a use of the jury as a defect detector, not an accuracy measurement against human defect labels\. For the purpose of the rest of this section, we then induce a single*emergent, shared*taxonomy over the*pooled*defects of all three models\. For details please refer to[AppendixJ](https://arxiv.org/html/2608.12585#A10)\.
### 5\.2Defect fingerprints
The taxonomy makes it possible to ask, for a given model,*what kind*of reasoning failure it exhibits and how severe each kind tends to be\. We illustrate this with a close read of one generator, nemotron\-3\-super, on AIME2026\. Over its1,9201\{,\}920traces \(295,453295\{,\}453steps, an average of84\.684\.6tokens per step and13,02513\{,\}025tokens per trace\), the jury flags8,7118\{,\}711distinct steps, counting each step once at the highest severity assigned to it:3,6813\{,\}681at*fatal*severity,3,2703\{,\}270*major*,1,7031\{,\}703*minor*, and5757*neutral*\.
##### Defective steps by category and severity
[Figure10](https://arxiv.org/html/2608.12585#A11.F10)in the Appendix breaks the flagged steps down by the eight top\-level taxonomy categories, each shown as a bar stacked by severity \(within a category a step is counted once at its highest severity; a step flagged under two different categories appears in both bars\)\.[Figure8](https://arxiv.org/html/2608.12585#S5.F8)drills one level deeper, restricting each category’s bar to its top three taxonomy leaves \(still stacked by severity\), which makes visible which specific defect type is doing the work behind each category’s total in[Figure10](https://arxiv.org/html/2608.12585#A11.F10)\. Within self\-monitoring and error\-handling,*extended reasoning on an uncorrected false premise*is both the largest single leaf and almost entirely fatal, confirming it as the dominant driver of that category’s fatal mass\. Within proof methodology and rigor, by contrast, the leading leaves are*incomplete case analysis or verification*and*unproven assumptions used as established facts*, carrying much less fatal mass than the logical\-reasoning and self\-monitoring leaves, consistent with that category’s lapses\-in\-rigor profile\. The full subcategory \(leaf\) distribution behind this figure, together with a worked example of a trace carrying defects at every severity level, is given in[SectionK\.1](https://arxiv.org/html/2608.12585#A11.SS1)\.
Figure 8:nemotron\-3\-super: flagged steps for the top three taxonomy leaves within each category, stacked by severity \(regex segmentation, AIME2026\)\. Within self\-monitoring and error\-handling,*extended reasoning on an uncorrected false premise*is by far the largest leaf and heavily fatal; within proof methodology and rigor, the dominant leaves \(*incomplete case analysis*and*unproven assumptions*\) carry far less fatal mass\.
### 5\.3Correlation between defects and final answer correctness
We split nemotron\-3\-super’s traces by whether its final integer answer was correct, and ask whether the jury’s defect signal tracks this independent notion of solution quality\. If it does, traces with wrong answers should carry both more defects and more*fatal*defects than correct\-answer traces from the same model\. This is exactly what we observe \([Figure9](https://arxiv.org/html/2608.12585#S5.F9.fig1)\)\. nemotron\-3\-super answers80%80\\%\(mean@64\) of AIME2026 problems correctly; on its correct\-answer traces the jury flags0\.990\.99defects per trace on average, of which0\.170\.17are fatal, while on its wrong\-answer traces this rises to4\.864\.86defects per trace, of which2\.312\.31are fatal, roughly a5×5\\timesincrease in overall defect rate and a14×14\\timesincrease in the fatal rate specifically\.
Figure 9:Accuracy vs defect severity for nemotron\-3\-super on AIME2026 \(on which it achieves80%80\\%accuracy\)\. Splitting its traces by correctness of the final answer, the jury flags0\.990\.99defects for traces with correct answer, versus4\.864\.86for traces with wrong answer, with the share of fatal defects rising from0\.170\.17to2\.312\.31per trace\. The defect signal tracks an external notion of solution quality, and reasoning defects for traces with wrong answer are disproportionately fatal\.The correct\-answer figures are informative in their own right: they are not zero\. Even when nemotron\-3\-super lands on the right final integer, the jury still flags on average just over one defect per trace, exposing “right answer, flawed reasoning” traces in which a genuine reasoning defect exists without changing the outcome\. The fatal rate on correct traces \(0\.170\.17per trace\) being far from zero but far below the wrong\-answer rate \(2\.312\.31\) is the expected signature of a defect signal\.
### 5\.4Caveats
Three caveats bound these fingerprints\. First, they are the jury’s judgements, not ground truth\. Second, the taxonomy is a single induction run at non\-zero temperature, so a re\-run can produce slightly different labels and counts — the fingerprints are over*this*taxonomy snapshot\. Third, per\-step rates are confounded by how finely the segmenter splits each reasoning trace, and as we discussed earlier the approach used for reasoning trace segmentation requires further study\.
## 6Exploratory downstream use: defect\-guided retry
The final answer correctness correlation above shows that the jury’s defect signal tracks solution correctness, but it does not establish whether information beyond the flagged step locations helps a model correct its reasoning\. We therefore run a larger, task\-based evaluation in which the generating model receives the jury’s feedback and attempts the problem again\.
##### Design\.
We draw6464nemotron\-3\-super samples for each of the3030AIME2026 problems, giving1,9201\{,\}920traces under the same generation settings as[Section5\.1](https://arxiv.org/html/2608.12585#S5.SS1)\. The original traces achieve80\.0%80\.0\\%pass@1\. We evaluate every trace with the same jury, which flags877877traces \(45\.7%45\.7\\%\) with at least one non\-neutral defect\. Among all of the traces, there are384384that have wrong final answers\. Of these384384traces372372of them include at least one non\-neutral defect, which corresponds to96\.9%96\.9\\%recall\.
For each flagged trace, the same generator retries the problem under two matched conditions\. In the*step\-anchors\-only*condition, it receives the cited\[STEP\-k\]locations but no description of the defects\. In the*rich\-feedback*condition, it receives those same locations together with thewhat\_went\_wrongdiagnosis and severity of each defect\. Both conditions include the original problem and trace and instruct the model to re\-solve the problem from scratch\. The comparison therefore holds localization fixed and tests the joint incremental value of diagnosis and severity\.
Table 3:Defect\-guided retry on nemotron\-3\-super AIME2026 traces \(6464samples per problem\)\. Retry accuracy is measured on the877877traces that have one or more defects with severity more than neutral\. Before retry, accuracy is505/877505/877\(57\.6%57\.6\\%\) on the flagged subset and1,536/1,9201\{,\}536/1\{,\}920\(80\.0%80\.0\\%\) over the full corpus\. Steps only is when in a retry the model is given the step indices that include a defect\. In Rich defect feedback the model also receives what went wrong and the severity of each defect for the defective steps\.
##### Results\.
Out of initial877877traces with non\-neutral defects,505505of them \(57\.6%57\.6\\%\) initially have a correct answer\. Providing rich feedback back to the model for these877877problems in a retry produces668668correct answers \(76\.2%76\.2\\%\), compared with624624\(71\.2%71\.2\\%\) from step anchors alone\. Compared to defective steps only information, rich feedback both repairs more of the372372flagged wrong answers \(181181vs\.146146\) and causes fewer regressions among the505505flagged correct answers \(1818vs\.2727\)\. In the paired comparison,7878traces are correct only under rich feedback and3434only under step anchors \(exact McNemarp=4×10−5p=4\\times 10^\{\-5\}\)\. Because traces from the same problem are correlated, we also perform a paired analysis at the problem level\. The mean advantage is\+4\.1\+4\.1percentage points, with a bootstrap95%95\\%confidence interval of\[\+2\.3,\+6\.1\]\[\+2\.3,\+6\.1\]; rich feedback outperforms step anchors on1515of the2929problems with flagged traces, underperforms on one, and ties on1313\.
These results provide task\-based evidence for the*joint*value of thewhat\_went\_wrongdiagnosis and severity beyond defective step localization alone\. The complete prompts, results, and limitations are in[AppendixN](https://arxiv.org/html/2608.12585#A14)\.
## 7Limitations
Reasoning Jury gets its accuracy gains with additional latency and token consumption\. Because the jury deliberates, it is inherently slower than a single judge\. Phase 1 is fast due to judgements produced in parallel\. The consolidation mode also is relatively fast since it is a single LLM call on top of Phase 1\. The deliberation mode, however is sequential by design and the number of turns is not known in advance\. For offline uses such as data curation or evaluation this added wall\-clock time might not introduce a risk\. However for online use cases the current deliberation pipeline might not be directly usable for on\-policy RL\. For off\-policy RL where the actor could be a few optimization step behind the policy being optimized, this latency is still tolerable\. Additionally, deliberation also consumes far more tokens than a single pass\. Running many jurors, and then having them argue over several turns, means the same trace is read and reasoned about repeatedly\. In our cost comparison \([Section4\.8](https://arxiv.org/html/2608.12585#S4.SS8)\), the jury used roughly7\.6×7\.6\\timesmore tokens and made about10\.9×10\.9\\timesmore LLM calls than a single\-model judge on the same records\. As we show there, this does not necessarily translate into higher dollar cost, but still, the raw token and call counts are substantially higher, which matters wherever throughput and rate limits are the binding constraint\.
## 8Conclusion
We introduced*Reasoning Jury*, a framework for identifying defects in long reasoning traces with a panel of LLM judges\. Jurors first inspect a step\-segmented trace independently; their findings are then combined through either consolidation or moderated deliberation\. The resulting verdict grounds each defect in one or more\[STEP\-x\]locations and records what went wrong, its severity, and supporting evidence\. This moves reasoning evaluation beyond a single score toward feedback that can be inspected, aggregated, and returned to a model\. The benchmark evaluations in this paper directly measure defect localization; they do not by themselves validate every field in the structured verdict\.
The experiments show that the value of a jury depends on the strength and composition of its members\. On Hard2Verify, the full Frontier jury is already saturated by gpt\-5\.4, reaching84\.484\.4Balanced F1 compared with83\.983\.9for gpt\-5\.4 alone\. When no single juror dominates, however, pooling independent judgements substantially improves recall: the two open\-weight juries reach80\.280\.2–80\.880\.8Balanced F1, roughly1111–1616points above their best individual members\. Homogeneous juries also improve over repeated samples of the same model, showing that independent sampling contributes even without model diversity\. The DeltaBench results exhibit the same pattern, with a three\-member gpt\-oss\-120b jury reaching61\.561\.5Balanced F1, compared with58\.458\.4for opus\-4\.6 and61\.861\.8for gpt\-5\.4\.
The aggregation controls refine the interpretation of these gains\. Majority voting performs poorly because jurors rarely identify exactly the same steps, whereas the union of their findings captures much of the improvement\. Deliberation does not uniformly outperform union on step\-level Balanced F1; its role is instead to adjudicate conflicting claims, remove duplicates, and produce one structured verdict\. This distinction matters operationally\. On Hard2Verify, a homogeneous gpt\-oss\-120b jury scores82\.382\.3with deliberation and80\.380\.3with consolidation, compared with73\.773\.7for a single opus\-4\.6 judge\. Despite using more calls and tokens, the two jury configurations cost $12\.35 and $6\.50, respectively, versus $79\.13 for opus\-4\.6 at the list prices used in our analysis\.
The structured findings also support uses that step\-level benchmark scores do not capture\. In the AIME2026 case study, the number and severity of detected defects track final\-answer correctness and expose flawed reasoning even in some correct\-answer traces\. More directly, when the generating model retries flagged traces, providing the jury’s diagnosis and severity raises accuracy from71\.2%71\.2\\%with step locations alone to76\.2%76\.2\\%\. This result provides initial task\-based evidence that information beyond localization is useful, although it evaluates diagnosis and severity jointly and does not establish the accuracy of every individual finding\.
Taken together, the results position Reasoning Jury as a practical approach for offline reasoning evaluation when a single judge is insufficient and a structured verdict is more useful than a scalar score\. Independent sampling provides the main accuracy benefit, while consolidation and deliberation turn the pooled findings into a usable output with different cost–latency tradeoffs\. Future work should directly evaluate the factuality and usefulness of the defect descriptions, calibrate severity labels, isolate the effect of deliberation under matched compute, and test downstream use across additional models and domains\.
## References
- \[1\]L\. Agrawal, J\. Lee, S\. Tan, S\. A\. Seshia, K\. Sen, D\. Klein, I\. Stoica, and J\. E\. Gonzalez\(2026\)Optimize\_anything: a universal api for optimizing any text parameter\.InProceedings of the ACM Conference on AI Engineering: Software Engineering for AI \(CAIS ’26\),External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.19633)Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p2.1)\.
- \[2\]Anthropic\(2026\)Claude opus 4\.6 system card\.Note:AnthropicPublished February 5, 2026External Links:[Link](https://www.anthropic.com/claude-opus-4-6-system-card)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[3\]Anthropic\(2026\)Claude sonnet 4\.6 system card\.Note:AnthropicPublished February 17, 2026External Links:[Link](https://www.anthropic.com/claude-sonnet-4-6-system-card)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[4\]J\. Chen, Y\. Lu, X\. Wang, H\. Zeng, J\. Huang, J\. Gesi, Y\. Xu, B\. Yao, and D\. Wang\(2025\)Multi\-agent\-as\-judge: aligning LLM\-agent\-based automated evaluation with multi\-dimensional human evaluation\.arXiv preprint arXiv:2507\.21028\.Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p3.1)\.
- \[5\]J\. C\. Chen, X\. Peng, P\. K\. Choubey, K\. Huang, J\. Zhang, M\. Bansal, and C\. Wu\(2026\)Nudging the boundaries of llm reasoning\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p2.1)\.
- \[6\]J\. Cheng, G\. Xiong, R\. Qiao, L\. Li, C\. Guo, J\. Wang, Y\. Lv, and F\. Wang\(2025\)Stop summation: min\-form credit assignment is all process reward model needs for reasoning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2504\.15275Cited by:[§2\.1](https://arxiv.org/html/2608.12585#S2.SS1.p1.1)\.
- \[7\]K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. Schulman\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[Appendix I](https://arxiv.org/html/2608.12585#A9.SS0.SSS0.Px1.p1.1)\.
- \[8\]DeepSeek\-AI\(2026\)DeepSeek\-V4\-Pro model card\.Note:Hugging FaceAccessed July 28, 2026External Links:[Link](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[9\]J\. Dekoninck, N\. Jovanović, T\. Gehrunger, K\. Rögnvaldsson, I\. Petrov, C\. Sun, and M\. Vechev\(2026\)Beyond benchmarks: matharena as an evaluation platform for mathematics with llms\.External Links:2605\.00674,[Link](https://arxiv.org/abs/2605.00674)Cited by:[§5](https://arxiv.org/html/2608.12585#S5.p1.1)\.
- \[10\]Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. Mordatch\(2024\)Improving factuality and reasoning in language models through multiagent debate\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p3.1)\.
- \[11\]Gemma Team\(2026\)Gemma 4 technical report\.arXiv preprint arXiv:2607\.02770\.External Links:[Link](https://arxiv.org/abs/2607.02770)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[12\]GLM\-5 Team\(2026\)GLM\-5: from vibe coding to agentic engineering\.arXiv preprint arXiv:2602\.15763\.External Links:[Link](https://arxiv.org/abs/2602.15763)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[13\]Google DeepMind\(2026\)Gemini 3\.1 pro model card\.Note:Google DeepMindPublished February 19, 2026External Links:[Link](https://deepmind.google/models/model-cards/gemini-3-1-pro/)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[14\]J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu,et al\.\(2024\)A survey on LLM\-as\-a\-judge\.arXiv preprint arXiv:2411\.15594\.Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p3.1)\.
- \[15\]Y\. He, S\. Li, J\. Liu, W\. Wang, X\. Bu, G\. Zhang, Z\. Peng, Z\. Zhang, Z\. Zheng, and W\. Su\(2025\)Can large language models detect errors in long chain\-of\-thought reasoning?\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Sunnyvale, California,pp\. 18468–18489\.External Links:[Link](https://aclanthology.org/2025.acl-long.905/)Cited by:[Appendix M](https://arxiv.org/html/2608.12585#A13.SS0.SSS0.Px1.p1.1),[Appendix M](https://arxiv.org/html/2608.12585#A13.SS0.SSS0.Px2.p1.1),[Table 12](https://arxiv.org/html/2608.12585#A13.T12),[§1](https://arxiv.org/html/2608.12585#S1.p7.1),[§3\.1](https://arxiv.org/html/2608.12585#S3.SS1.p1.1),[§4\.6](https://arxiv.org/html/2608.12585#S4.SS6.p1.1)\.
- \[16\]J\. Hübotter, F\. Lübeck, L\. D\. Behric, A\. Baumann, M\. Bagatella, D\. Marta, I\. Hakimi, I\. Shenfeld, T\. K\. Buening, C\. Guestrin, and A\. Krause\(2026\)Reinforcement learning via self\-distillation\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=k8DcHShsrJ)Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p2.1)\.
- \[17\]H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe\(2024\)Let’s verify step by step\.International Conference on Learning Representations \(ICLR\)\.Cited by:[§2\.1](https://arxiv.org/html/2608.12585#S2.SS1.p1.1)\.
- \[18\]MiniMax AI\(2026\)MiniMax\-M2\.7 model card\.Note:Hugging FaceAccessed July 28, 2026External Links:[Link](https://huggingface.co/MiniMaxAI/MiniMax-M2.7)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[19\]MiniMax AI\(2026\)MiniMax\-M3 model card\.Note:Hugging FaceAccessed July 28, 2026External Links:[Link](https://huggingface.co/MiniMaxAI/MiniMax-M3)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[20\]Moonshot AI\(2026\)Kimi K2\.6 model card\.Note:Hugging FaceAccessed July 28, 2026External Links:[Link](https://huggingface.co/moonshotai/Kimi-K2.6)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[21\]NVIDIA\(2025\)NVIDIA Nemotron 3: efficient and open intelligence\.Note:White paperExternal Links:[Link](https://arxiv.org/abs/2512.20856)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[22\]OpenAI\(2025\)gpt\-oss\-120b & gpt\-oss\-20b model card\.arXiv preprint arXiv:2508\.10925\.External Links:[Link](https://arxiv.org/abs/2508.10925)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[23\]OpenAI\(2026\)gpt\-5\.4 model\.Note:OpenAI API documentationAccessed July 28, 2026External Links:[Link](https://developers.openai.com/api/docs/models/gpt-5.4)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[24\]OpenAI\(2026\)Introducing gpt\-5\.4\.Note:OpenAIPublished March 5, 2026External Links:[Link](https://openai.com/index/introducing-gpt-5-4/)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[25\]S\. Pandit, A\. Xu, X\. Nguyen, Y\. Ming, C\. Xiong, and S\. Joty\(2026\)Hard2Verify: a step\-level verification benchmark for open\-ended frontier math\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 22502–22517\.Cited by:[Appendix I](https://arxiv.org/html/2608.12585#A9.p1.1),[§1](https://arxiv.org/html/2608.12585#S1.p7.1),[§3\.1](https://arxiv.org/html/2608.12585#S3.SS1.p1.1)\.
- \[26\]Qwen Team\(2026\)Qwen3\.6\-27B: flagship\-level coding in a 27B dense model\.External Links:[Link](https://qwen.ai/blog?id=qwen3.6-27b)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[27\]R\. Sharma, W\. Chen, N\. Provenzano, and T\. Vu\(2026\)PRISM: pushing the frontier of deep think via process reward model\-guided inference\.arXiv preprint arXiv:2603\.02479\.Cited by:[§2\.1](https://arxiv.org/html/2608.12585#S2.SS1.p1.1)\.
- \[28\]P\. Verga, S\. Hofstatter, S\. Althammer, Y\. Su, A\. Piktus, A\. Arkhangorodsky, M\. Xu, N\. White, and P\. Lewis\(2024\)Replacing judges with juries: evaluating LLM generations with a panel of diverse models\.arXiv preprint arXiv:2404\.18796\.Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p3.1)\.
- \[29\]Z\.ai\(2026\)GLM\-5\.2\-FP8 model card\.Note:Hugging FaceAccessed July 28, 2026External Links:[Link](https://huggingface.co/zai-org/GLM-5.2-FP8)Cited by:[§4\.2](https://arxiv.org/html/2608.12585#S4.SS2.p1.1)\.
- \[30\]X\. Zhang, S\. Wu, Y\. Zhu, H\. Tan, S\. Yu, Z\. He, and J\. Jia\(2026\)Scaf\-grpo: scaffolded group relative policy optimization for enhancing llm reasoning\.External Links:2510\.19807,[Link](https://arxiv.org/abs/2510.19807)Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p2.1)\.
- \[31\]C\. Zheng, Z\. Zhang, B\. Zhang, R\. Lin, K\. Lu, B\. Yu, D\. Liu, J\. Zhou, and J\. Lin\(2025\)ProcessBench: identifying process errors in mathematical reasoning\.InProceedings of the Association for Computational Linguistics \(ACL\),Note:arXiv:2412\.06559Cited by:[Appendix I](https://arxiv.org/html/2608.12585#A9.p1.1),[§2\.1](https://arxiv.org/html/2608.12585#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.12585#S3.SS1.p1.1)\.
- \[32\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\(2023\)Judging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.Advances in Neural Information Processing Systems \(NeurIPS\)\.Cited by:[§1](https://arxiv.org/html/2608.12585#S1.p3.1)\.
- \[33\]J\. Zou, L\. Yang, J\. Gu, J\. Qiu, K\. Shen, J\. He, and M\. Wang\(2025\)ReasonFlux\-PRM: trajectory\-aware PRMs for long chain\-of\-thought reasoning in LLMs\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2506\.18896Cited by:[§2\.1](https://arxiv.org/html/2608.12585#S2.SS1.p1.1)\.
## Appendix ADefect schema and an illustrative example
Each Phase 1*negative*is a JSON object grounded to specific steps\. The prompt enforces a genericness test and a 1–5 specificity self\-check, keeping only issues scoring≥4\\geq 4\. The contrast below \(both abridged from the prompt templates\) shows the bar\.
\{
"statement\_refs":\["STEP\-6","STEP\-14"\],
"what\_went\_wrong":"Thereasoningmakesunsupportedassumptionsaboutthe
connectionbetweenthetwodomainsandappliesformulasincorrectly\.",
"impact":"major",
"evidence":"Thetraceassumesoverlapwithoutjustification\."
\}
Listing 1:Insufficient \(specificity score 2 — rejected\)\.\{
"statement\_refs":\["STEP\-14","STEP\-17"\],
"what\_went\_wrong":"In\[STEP\-14\],thetracestates’Energyrequiredis
proportionalto1/sqrt\(finalparticlesize\)’citingRittinger’sLaw,but
Rittinger’sLawisEproportionalto\(1/D2\-1/D1\),i\.e\.1/Dnot1/sqrt\(D\)\.
Thismisformulationpropagatesto\[STEP\-17\],yieldinganenergyestimate
~4\.5xtoolow\.",
"impact":"fatal",
"evidence":"\[STEP\-14\]’Energy\.\.\.\(proportionalto1/sqrt\(finalparticle
size\)\)’;\[STEP\-17\]’soenergyscalesas1/sqrt\(50\)\.\.\.’",
"correct\_value":"Eproportionalto\(1/D2\-1/D1\)\.ForD1=1mm,D2=50um:
~19xbaseline,not4\.5x\."
\}
Listing 2:Sufficient \(specificity score 5 — accepted\)\.
## Appendix BPhase 1 prompt
The complete Phase 1 independent\-judgement prompt is reproduced below: a short system prompt followed by the user prompt template, whoseproblem,reasoning\_trace, andfinal\_solutionplaceholders are filled per trace\. Mathematical symbols are transcribed to ASCII \(e\.g\. “proportional to” for∝\\propto, “1/sqrt\(D\)” for1/D1/\\sqrt\{D\}\), matching[AppendixA](https://arxiv.org/html/2608.12585#A1)\.
\[SYSTEMPROMPT\]
Youareareasoning\-traceauditor\.Youidentifyconcreteweaknessesinreasoningtracesbypointingtospecificsteps\.YoualwaysreturnvalidJSONandneverwrapitinmarkdownfences\.
\[USERPROMPT\]
Youarea\*\*reasoning\-traceauditor\*\*\.Yourjobisto\*\*evaluateaspecificreasoningtrace\*\*foraspecificproblembyidentifyingconcreteweaknesses\.
Youwillbegiven:
1\)\*\*PROBLEM\*\*:thetaskstatement
2\)\*\*REASONINGTRACE\*\*:thereasoningproducedbyamodelwhileattemptingtheproblem\.Thetraceisdividedintosegments,eachprefixedwith\[STEP\-x\]wherexisanintegerindicatingtheordinalpositionofeachsegmentinthetrace\.
3\)\*\*FINALResponse\*\*:thefinalresponsetotheproblem,basedonthereasoningtrace\.
\#\#Primaryobjective
Judgethe\*weaknesses\*oftheprovidedreasoningtracebypointingto\*\*specificbadreasoningmoves\*\*\.
\#\#Criticalconstraints
\-\*\*DoNOTwriteafreshfullresponse\*\*totheproblem\.
\-\*\*Avoidgenericfeedback\.\*\*Everycritiquemustbe\*\*trace\-grounded\*\*:
\-nametheexactquantity/expression/claiminvolved,and
\-referencewhich\[STEP\-x\]itoccursin\.
\#\#Groundingrequirement\(veryimportant\)
Alljudgementsmustreferencespecific\[STEP\-x\]markersinthereasoningtrace\.Inyour‘what\_went\_wrong‘description,youMUST:
1\.Namethespecificstep\(s\):"In\[STEP\-14\],thetracestates\.\.\."
2\.Quotetheexactproblematicclaimfromthatstep
3\.ExplainWHYitiswrong,includingwhatthecorrectvalue/reasoningshouldbe
\#\#Genericnesstest\(mustpass\)
Foreverypointyoumake,ask:
\>"Couldthiscommentapplytoatotallydifferentproblemwithnoedits?"
Ifyes,rewriteitsoitmentionsatleastone\*\*problem\-specificartifact\*\*\(variablename,expression,unit,constraint,theorem/rule,diagramtype,etc\.\)ANDa\*\*specificoperation/transition\*\*fromaspecific\[STEP\-x\]\.
\#\#Specificityself\-check\(mustpass\)
Beforeincludingeachissue,scoreit1\-5:
5=citesexactformula/number/quotefromaspecificSTEPANDexplainsthecorrectvalue
4=citesaspecificclaimfromaspecificSTEPANDnamestheerrortype
3=referencesastepbutdescribestheissuegenerically
2=couldapplytoadifferentproblemwithminoredits
1=completelygeneric
\*\*Onlyincludeissuesscoring\>=4\.\*\*Discardissuesscoring3orbelow\.
\#\#Examples
INSUFFICIENT\(score2\-\-doNOTwritelikethis\):
‘‘‘json
\{
"statement\_refs":\["STEP\-6","STEP\-14"\],
"what\_went\_wrong":"Thereasoningmakesunsupportedassumptionsabouttheconnectionbetweenthetwodomainsandappliesformulasincorrectly\.",
"impact":"major",
"evidence":"Thetraceassumesoverlapwithoutjustification\."
\}
‘‘‘
SUFFICIENT\(score5\-\-writelikethis\):
‘‘‘json
\{
"statement\_refs":\["STEP\-14","STEP\-17"\],
"what\_went\_wrong":"In\[STEP\-14\],thetracestates’Energyrequiredproportionalto1/sqrt\(finalparticlesize\)’citingRittinger’sLaw,butRittinger’sLawisEproportionalto\(1/D2\-1/D1\),proportionalto1/Dnot1/sqrt\(D\)\.Thismisformulationpropagatesto\[STEP\-17\]wheretheenergyestimateiscomputedusingthewrongexponent,yieldingavalueroughlysqrt\(20\)~4\.5xtoolow\.",
"impact":"fatal",
"evidence":"\[STEP\-14\]’Energyrequiredproportionaltosurfaceareagenerated\(proportionalto1/sqrt\(finalparticlesize\)\)’;\[STEP\-17\]’soenergyscalesas1/sqrt\(50\)relativeto1mmbaseline’",
"correct\_value":"Rittinger’sLaw:Eproportionalto\(1/D2\-1/D1\)\.ForD1=1mm,D2=50um:Eproportionalto\(1/0\.05\-1/1\)=19,so~19xbaseline,not4\.5x\."
\}
‘‘‘
\#\#Severityguidelines\(preliminary\)
Yourseverityratingisa\*\*preliminaryestimate\*\*\.Itwillbeindependentlyrecalibratedbyalaterverificationstage\.Focusyourenergyon\*\*accuratedetectionandspecificdescription\*\*ratherthanagonizingovertheexactseveritylevel\.Thatsaid,usethisguidance:
\-\*\*fatal\*\*:Thetraceteachesafalsegeneralizableruleorfact\-\-alargelanguagemodeltrainedonthiswouldlearnfundamentallybrokenreasoningthattransferstofuturetasks\.Examples:wrongformula,fabricatedevidence,factualinversion,categoryerror\.Domain\-knowledgeerrors\(wrongfacts,wrongformulas\)areparticularlystrongfatalsignals\.
\-\*\*major\*\*:Asignificanterrorthatweakensthereasoningbutislocalized\-\-itdoesnotencodeatransferablefalserule\.Examples:arithmeticmistakeinonestep,omittingarelevantconstraint,misreadingonedatapoint\.
\-\*\*minor\*\*:Asmallimprecisionorstylisticissuethatdoesnotaffectthereasoningoutcome\.Examples:roundingdifferences,informallanguage,minornotationinconsistency\.
\-\*\*neutral\*\*:Notactuallyaflaw\-\-justadifferentapproachorstylechoice\.
\*\*Important\*\*:DoNOTflag\*\*exposition\-only\*\*issues\(confusingexplanationofcorrectreasoning,poorformatting,verbosepresentation\)asmajororfatal\.Theseaffectcommunicationquality,notreasoningquality,andarealmostnevertrueflawsfortrainingpurposes\.
\#\#Input
\#\#\#PROBLEM
====beginproblem====
\{problem\}
====endproblem====
\#\#\#REASONINGTRACE\(thetracetoaudit\)
====beginreasoningtrace====
\{reasoning\_trace\}
====endreasoningtrace====
\#\#\#Finalresponse
====beginfinalresponse====
\{final\_solution\}
====endfinalresponse====
\#\#Outputformat\(JSONONLY\)
Return\*\*onlyvalidJSON\*\*\.TheJSONmustinclude:
\-‘negatives‘:arrayoftrace\-groundedweaknesses,eachwith:
\-‘statement\_refs‘:arrayofstepIDswheretheissuemanifests\(e\.g\.,\["STEP\-14","STEP\-17"\]\)
\-‘what\_went\_wrong‘:specificdescriptionthatembeds\[STEP\-x\]referencesanddirectquotesfromthetrace\.Mustnametheexactquantity/claimandexplainwhyitiswrong\.
\-‘impact‘:oneof"neutral"\|"minor"\|"major"\|"fatal"
\-‘evidence‘:directquotesfromthetrace,prefixedwiththeirstepIDs\(e\.g\.,"\[STEP\-14\]’Energyproportionalto1/sqrt\(D\)’"\)
\-‘correct\_value‘:whatthecorrectreasoning/value/conclusionshouldbe\(omitonlyiftheissueisaboutmissinganalysisratherthanawrongclaim\)
\#\#Additionalrules
\-Be\*\*surgicallyspecific\*\*:statetheexactissueandexacterrorpattern
\-Keepeverythinggroundedinthereasoningtraceandproblemstatement
\-Thereasoningtraceisnotthefinalanswer\-\-it’stheinternalreasoningprocessthatisusedtocreatethefinalresponse\.Sothereasoningtracedoesnothavetofollowallthedetailedinstructionsintheproblemsuchasputtingthefinalanswerin‘\\boxed‘,\.\.\.
\-Each‘what\_went\_wrong‘mustbeself\-contained:areadershouldunderstandtheissuefromthatfieldalone,withoutneedingtoread‘statement\_refs‘or‘evidence‘separately
\#\#Defectpropagation
\-\*\*Important\*\*:Anystepthatcontainsorisbasedonanerrorisconsideredincorrectandneedstobemarkedashavingadefect\.Thatis,iftheerroriscarriedforwardfromaprevioussteporisbasedonanerrorinthepreviousstep,considerthestepincorrectandgiveitadefectintheoutput\.Apropagatedstepinheritsthe\*\*sameseverity\(‘impact‘\)\*\*astheupstreamdefectitdependson\-\-ifitbuildsona‘fatal‘error,markit‘fatal‘\.
Listing 3:The Phase 1 independent\-judgement prompt \(system prompt and user template\)\.
## Appendix CPhase 2 prompts
Phase 2 uses two prompts: the moderator prompt that runs the deliberation loop \([SectionC\.2](https://arxiv.org/html/2608.12585#A3.SS2)\) and the juror contribution prompt that each selected juror answers on its turn \([SectionC\.3](https://arxiv.org/html/2608.12585#A3.SS3)\)\.
### C\.1Consolidation addendum
In consolidation mode \([Section4\.8](https://arxiv.org/html/2608.12585#S4.SS8)\) the multi\-turn deliberation and the separate verdict extraction are replaced by a single call: the moderator model re\-performs the Phase 1 task with every juror’s independent judgement attached as unverified candidate findings\. The consolidator’s user prompt is the Phase 1 prompt of[AppendixB](https://arxiv.org/html/2608.12585#A2)verbatim \(including the defect\-propagation addendum when propagation is enabled\) with the addendum below appended — thejury\_countandpanel\_judgementsplaceholders are filled with the panel size and the formatted Phase 1 judgements, attributed by juror id\. Unlike the deliberation moderator \([SectionC\.2](https://arxiv.org/html/2608.12585#A3.SS2)\), which is deliberately blind to the trace, the consolidator sees the full problem and reasoning trace and rules on every candidate directly; its JSON output is the final verdict, with two extra per\-finding fields \(vote\_count,voters\) recording panel support\. Mathematical symbols are transcribed to ASCII, matching[AppendixA](https://arxiv.org/html/2608.12585#A1)\.
\#\#PANELCANDIDATEFINDINGS\(unverified\-\-leads,notconclusions\)
Beforeyou,\{jury\_count\}independentauditorsexaminedthissametraceunderthe
exactinstructionsabove\.Theirrawfindingsarereproducedbelow,attributedby
auditorid\.TreatthemasCANDIDATES,notestablishedfacts:auditorscanbe
wrong,canduplicateoneanother,andcanmissdefectsentirely\.
====beginpanelfindings====
\{panel\_judgements\}
====endpanelfindings====
\#\#Howtousethepanelfindings
1\.\*\*Verifybeforeadopting\.\*\*Foreachcandidate,re\-readthecited\[STEP\-x\]
inthereasoningtrace\.Adoptitonlyifthetraceitselfsupportsthe
claim\.Rejectcandidatesthetracedoesnotsupport\-\-neverincludea
findingmerelybecauseoneorseveralauditorsraisedit\.Head\-countis
notevidence;onlythetraceis\.
2\.\*\*Mergeduplicatesatmaximumspecificity\.\*\*Whenseveralcandidates
describethesameunderlyingdefect,outputONEfinding,usingtheMOST
SPECIFICdescriptionavailableamongthem\(exactquotes,numbers,formulas,
\[STEP\-x\]refs\)\.Supplement,don’taverage:foldinsupportingevidencefrom
theotherauditors,butneverreplaceaspecificclaimwithavaguer
paraphrase\.
3\.\*\*Fillthegaps\.\*\*Youaresimultaneouslyauditingthetraceyourselfunder
theinstructionsabove\.IncludegenuinedefectsthatNOauditorraised\.
Yourownfindingsareheldtothesamegrounding,genericness,and
specificityrequirementsaseverythingelse\.
4\.\*\*Recalibrateseverityyourself\.\*\*Acandidate’sseverityratingisa
suggestion,notaconstraint:assignseverityfromyourownreadingofthe
severityguidelinesabove\.
5\.\*\*Samebarforeverything\.\*\*Everyfindinginyouroutput\-\-adopted,
merged,ornewlyadded\-\-mustpassthegenericnesstestandscore\>=4on
thespecificityself\-check\.
\#\#Outputformatreminder\(JSONONLY\-\-sameschemaasspecifiedabove\)
Yourfinaloutputremains\*\*onlyvalidJSON\*\*,exactlyasspecifiedinthe
"Outputformat"sectionabove\-\-consolidatingthepanelfindingsdoesNOT
changetheoutputschema\.Donotoutputanyprose,commentaryonthepanel,
ormarkdownfences\.TheJSONmustinclude:
\-‘negatives‘:arrayoftrace\-groundedweaknesses,eachwith:
\-‘statement\_refs‘:arrayofstepIDswheretheissuemanifests\(e\.g\.,\["STEP\-14","STEP\-17"\]\)
\-‘what\_went\_wrong‘:specificdescriptionthatembeds\[STEP\-x\]referencesanddirectquotesfromthetrace\.Mustnametheexactquantity/claimandexplainwhyitiswrong\.
\-‘impact‘:oneof"neutral"\|"minor"\|"major"\|"fatal"
\-‘evidence‘:directquotesfromthetrace,prefixedwiththeirstepIDs\(e\.g\.,"\[STEP\-14\]’Energyproportionalto1/sqrt\(D\)’"\)
\-‘correct\_value‘:whatthecorrectreasoning/value/conclusionshouldbe\(omitonlyiftheissueisaboutmissinganalysisratherthanawrongclaim\)
\-‘vote\_count‘:howmanypanelauditors’candidatessupportthisfinding\(0ifitisyoursalone\)
\-‘voters‘:arrayofthesupportingauditorids\(e\.g\.,\["jury\-0","jury\-2"\];emptyarrayforyourownfindings\)
Listing 4:The consolidation addendum, appended verbatim to the Phase 1 prompt of[AppendixB](https://arxiv.org/html/2608.12585#A2)\.
### C\.2Moderator prompt
The moderator’s prompt is reproduced below: a system prompt establishing its procedural role, followed by the per\-turn template whose placeholders \(total\_turns,recent\_transcript,previous\_consensus\_state, etc\.\) are filled each turn\. The moderator never receives the problem, the reasoning trace, or the solution; it sees only the deliberation transcript\. Mathematical symbols are transcribed to ASCII, matching[AppendixA](https://arxiv.org/html/2608.12585#A1)\.
\[SYSTEMPROMPT\]
Youareadeliberationmoderatorforapanelofreasoning\-traceauditors\.
CRITICALCONSTRAINTS:
\-YouhaveNOaccesstotheproblem,reasoningtrace,orsolutionunderreview\.
\-YoucanONLYseetheconversationbetweenpanelmembers\.
\-YoumustNEVERspeculateaboutthecontentofthereasoningtrace\.
\-Yourroleispurelyprocedural:manageturns,identifydisagreements,anddetectconvergence\.
YOURRESPONSIBILITIES:
0\.CONSENSUSTRACKING:Oneofyourjobsistoproduce\*\*onehigh\-fidelityconsolidatedjudgement\*\*by:
\-clusteringoverlappingfindingsintocanonical,problem\-specificitems
\-\*\*preservingmaximumspecificity\*\*fromthemostdetailedsourcejudgement
\#\#Criticalconstraints
\-\*\*DoNOTsolvetheproblem\.\*\*
\-\*\*DoNOTinventnewissues\*\*notsupportedbytheprovidedjudgements\.
\-\*\*DoNOTabstractawayspecifics\.\*\*Whenmultiplejudgesdescribethesameissueatdifferentlevelsofdetail,usetheMOSTSPECIFICdescriptionasthebase\.
\#\#Specificitypreservationrules\(veryimportant\)
Whenconsolidatingoverlappingissuesintoacanonicaldescription:
1\.\*\*Selectthebestsource\*\*:Identifywhichjudgegavethemostspecificdescription\(moststepreferences,directquotes,exactnumbers/formulas\)\.Usethatasthecanonicalwording\.
2\.\*\*Preservefromthebestsource\*\*:
\-Exactnumbers,formulas,orcalculations
\-Directquotesfromthetrace\(keepinquotationmarks\)
\-Specific\[STEP\-x\]references
\-The‘correct\_value‘ifanyjudgeprovidesit
3\.\*\*Supplement,don’taverage\*\*:Addsupportingevidencefromotherjudges,butneverreplaceaspecificclaimwithavaguerparaphrase\.
4\.\*\*Genericnesstest\*\*:Eachconsolidatedissuemustpass:"Couldthisdescriptionapplytoatotallydifferentproblemwithnoedits?"Ifyes,rewriteusingthemostspecificjudge’slanguage\.
Example:
\-jury\-2says:"In\[STEP\-14\],thetracestates’Eproportionalto1/sqrt\(D\)’butRittinger’sLawisEproportionalto1/D"
\-jury\-4says:"Theenergyformulaisappliedincorrectly"
\-\*\*CORRECTconsolidation\*\*:Usejury\-2’swordingverbatim\.jury\-4addssupportbutnotspecificity\.
\-\*\*WRONGconsolidation\*\*:"Theenergyformulaismisapplied"\(lostjury\-2’sdetail\)
\#\#Output\(JSON\)
Forthisresponsibility,return\*\*onlyvalidJSON\*\*,underkey"consensus\_state"inthefinaloutput\.TheJSONmustinclude:
\-‘consensus\_map‘:arrayofconsensusobjects,eachwith:
\-‘what\_went\_wrong‘:consolidatedclaim\-\-mustbetheMOSTSPECIFICversion,embedding\[STEP\-x\]refsanddirectquoteswhereavailable
\-‘source\_ref‘:\["jury\-1","jury\-3",\.\.\.\]\-\-whichjudgessupportthisclaim
\-‘best\_source‘:"jury\-2"\-\-whichjudgeprovidedthemostspecificdescriptionusedasbase
\-‘statement\_refs‘:\["STEP\-4","STEP\-7"\]\-\-all\[STEP\-x\]markersinvolved,collectedfromthesupportingjudges’descriptions
\-‘consensus\_id‘:"C1","C2",etc\.
\-‘judge\_count‘:numberofjudgessupportingthisissue
\-‘impact‘:"fatal"\|"major"\|"minor"\|"neutral"\-\-preliminaryseverityfromthebestsourcejudge
\-‘evidence‘:directquotesfromthetrace,prefixedwith\[STEP\-x\]\-\-takenfromthemostspecificjudge
\-‘correct\_value‘:whatthecorrectreasoning/valueshouldbe\(carryfrombestsourceifavailable,omitifnojudgeprovidedit\)
1\.TURNSELECTION:Choosewhichpanelistspeaksnextbasedon:
\-Ensureeverypanelistspeaksatleastonceperlogicalround\.
\-Prioritisepanelistsinvolvedinunresolveddisagreements\.
\-Recallsilentpanelistswhohavenotspokenin2\+turns\.
2\.INSTRUCTIONGENERATION:Givetheselectedpanelistaspecific,process\-orientedinstruction:
\-Pointthemtospecificdisagreementstheyshouldaddress\.
\-Askthemtoclarify,defend,orconcedespecificpointsraisedbyothers\.
\-NEVERsuggestwhatthe"right"answerisaboutthetracecontent\.
3\.TERMINATIONDETECTION:Signalthatdeliberationshouldendwhen:
\-Afulllogicalroundpasseswithnonewsubstantivearguments\.
\-Allpanelistshaveexplicitlysignalledagreementonallpoints\.
\-Argumentsarecyclingwithoutresolution\.
YouMUSTrespondwithONLYvalidJSON\(nomarkdownfences\)inthisexactformat:
\{
"should\_terminate":false,
"termination\_reason":null,
"next\_speaker":"<agent\-id\>",
"instruction":"<whatyouwantthemtodo\>",
"reasoning":"<yourinternalreasoningaboutconversationdynamics\>",
"consensus\_state":<jsonfromResponsibility0:CONSENSUSTRACKING\>
\}
\[PER\-TURNTEMPLATE\]
DELIBERATIONSTATUS:
\-Totalturnssofar:\{total\_turns\}
\-Currentlogicalround:\{logical\_round\}
\-Panelistswhohavespokenthislogicalround:\{spoken\_this\_round\}
\-PanelistswhohaveNOTspokenthislogicalround:\{silent\_this\_round\}
\-AllpanelistIDs:\{all\_agent\_ids\}
TRANSCRIPT\(last\{recent\_window\}turns\):
====begintranscript====
\{recent\_transcript\}
====endtranscript====
YOURPREVIOUSCONSENSUSSTATE:
Thisistheconsensus\_stateyouproducedonyourlastturn\.Useitasyourstartingpoint\-\-updateitbasedonwhathaschangedinthetranscriptsincethen\(newendorsements,concessions,withdrawnclaims,etc\.\)\.Onthefirstturnthiswillbeempty\.KeepinmindthatthepanelistsdoNOThaveaccesstoconsensusstate,sodonotrefertoitscontentswhenyouaddressthepanelists\.
====beginpreviousconsensusstate====
\{previous\_consensus\_state\}
====endpreviousconsensusstate====
Basedonthedeliberationdynamics,decide:
0\.Update"consensus\_state":Payspecialattentiontothelastjuror’sresponsetoyourquestionandcheckwhetherthecurrentspeakerendorsed,contested,orwassilentoneachconsensusitem\.Update‘source\_ref‘onconsensus\_stateaccordingly\.
1\.Shoulddeliberationterminate?\(Hasafullroundpassedwithnonewsubstantivearguments?Arepanelistsrepeatingthemselves?\)
2\.Ifnot,whoshouldspeaknextandwhatshouldtheyaddress?DoNOTrefertoanythingfromconsensusstatewhenyouaddressthepanelists,astheydon’thaveaccesstoit\.
Listing 5:The moderator prompt \(system prompt and per\-turn template\)\.
### C\.3Juror contribution prompt
When a juror is selected to speak, it receives the Phase 2 contribution prompt below\. The prompt opens by identifying the juror, then re\-includes the same auditor preamble, problem/trace/solution input, and additional rules as the Phase 1 prompt \([AppendixB](https://arxiv.org/html/2608.12585#A2)\), and closes with the deliberation transcript so far, the moderator’s instruction for this turn, and the critical rules governing the contribution\.
Youare\{agent\_id\}\.<<<thentheauditorpreamble,thePROBLEM/REASONING
TRACE/FINALRESPONSEinputblock,andthe"Additionalrules"block,all
identicaltothePhase1promptinAppendixB\(theJSONoutput\-formatblockis
omitted,sinceaPhase2turnisanatural\-languageargument,notaJSONlist\)\>\>\>
\-\-\-
Rememberthatyouare\{agent\_id\}\.Yourpriorcontributionsinthedeliberationssofarareinthetranscriptbelow\(markedby\{agent\_id\}\)\.
DELIBERATIONSOFAR:
====begindeliberationsofar====
\{transcript\}
====enddeliberationsofar====
\-\-\-
Nowthemoderatorinthepanelhasthefollowinginstructionforyou:
MODERATORINSTRUCTIONFORYOU:
====beginmoderatorinstructionsforyou====
\{instruction\}
====endmoderatorinstructionsforyou====
\-\-\-
Followtheinstructionsfromthemoderatorwiththefollowingrulesinmind:
CRITICALRULES:
\-EveryclaimyoumakeMUSTreferencespecific\[STEP\-x\]markersfromthereasoningtraceprovidedaboveunder====beginreasoningtrace====\.Carefullyexamineitbeforemakingdecisions\.
\-Ifyoucannotgroundaclaiminaspecificstep,donotmakeit\.
\-YoumayAGREEwithotherpanelists,DISAGREEwithevidence,orRAISEnewissues\.
\-Beconcise\.Focusonthestrongestarguments\.
\-Ifyouhavenothingnewtoadd,say"Ihavenonewargumentstopresent\."
Respondwithyourcontributiontothedeliberation\.
Listing 6:The Phase 2 juror contribution prompt\. The middle blocks are identical to the Phase 1 prompt \([AppendixB](https://arxiv.org/html/2608.12585#A2)\) and elided here\.
## Appendix DSeverity rubric
Severity is assigned on a four\-level ordinal scale \(neutral=0=0, minor=1=1, major=2=2, fatal=3=3\), with the training\-impact framing used in the prompts:
fatalThe trace teaches a*false generalizable rule or fact*—a model trained on it would learn broken reasoning that transfers \(wrong formula, fabricated evidence, factual inversion, category error\)\.
majorA significant but*localized*error that does not encode a transferable false rule \(a single arithmetic slip, an omitted constraint, a misread data point\)\.
minorA small imprecision not affecting the outcome \(rounding, notation\)\.
neutralNot a flaw—a different valid approach or style choice\.
Exposition\-only issues \(confusing but correct explanations, formatting\) are explicitly*not*to be rated major/fatal: they affect communication, not reasoning quality\.
## Appendix EVerdict object
The pipeline emits aVerdictcontaining: the consensusnegatives\(each withstatement\_refs,what\_went\_wrong,impact,evidence,vote\_count,voters\); a scalarconfidence∈\[0,1\]\\in\[0,1\];dissenting\_views;deliberation\_rounds;termination\_reason; the fulltranscriptandphase1\_judgements; and, when the corresponding mechanisms are enabled,dropped\_negatives\(removal pass\) andnegative\_adjudications\(evidence arbiter\)\. A best\-effort fallback verdict is produced if extraction fails, so the pipeline degrades gracefully rather than crashing\.
## Appendix FFault tolerance
At jury scale, individual calls fail; the pipeline is designed to degrade rather than crash\. If some Phase 1 jurors fail, it proceeds with the remainder above a configured minimum\. A juror failing during deliberation has its turn skipped and is removed after repeated consecutive failures\. A moderator failure falls back to deterministic round\-robin turn order; a grounding\-checker failure accepts the contribution \(benefit of the doubt\); an extraction failure returns a best\-effort verdict aggregated from Phase 1\. API throttling is retried with exponential backoff, and a Phase 1 wall\-clock ceiling prevents one hung juror from stalling the fan\-out\.
## Appendix GFull results tables
This section collects the per\-configuration numbers behind the figures in[Section4](https://arxiv.org/html/2608.12585#S4)\. All numbers are step\-level, All\-records, minor\+, propagation\-aware Balanced F1 on a00–100100scale, scored against the Hard2Verify human labels\.[Table4](https://arxiv.org/html/2608.12585#A7.T4)corresponds to[Figure3](https://arxiv.org/html/2608.12585#S4.F3)\(each jury’s consensus against its jurors’ solo scores\),[Table5](https://arxiv.org/html/2608.12585#A7.T5)to[Figure4](https://arxiv.org/html/2608.12585#S4.F4)\(jury knockout\),[Table6](https://arxiv.org/html/2608.12585#A7.T6)to[Figure5](https://arxiv.org/html/2608.12585#S4.F5)\(moderator sweep\), and[Table7](https://arxiv.org/html/2608.12585#A7.T7)to[Figure6](https://arxiv.org/html/2608.12585#S4.F6)\(diversity control\)\. The four\-panel jury consensus is also reported inline as[Table1](https://arxiv.org/html/2608.12585#S4.T1)\.
Table 4:Per\-juror solo Phase 1 versus deliberated consensus for the four juries of[Figure3](https://arxiv.org/html/2608.12585#S4.F3)\(minor\+, All records, step\-level, propagation\-aware\)\. Within each panel the consensus row is the deliberated final verdict; the remaining rows are that panel’s jurors judging alone, sorted by Balanced F1\. The Bal\-F1 column carries a95%95\\%bootstrap confidence half\-width \(±\\pm,10,00010\{,\}000record\-level resamples\)\.Table 5:Knockout: consensus scores \(minor\+, All records, step\-level, propagation\-aware\) as each jury is reduced from five to two jurors, with Phase 1 held fixed \(each smaller jury reuses the full jury’s Phase 1 judgements; only Phase 2 is re\-run\)\. Quality holds down to three jurors on every panel; the sharpest drop is at two\.Table 6:Moderator sweep with the jury and Phase 1 held fixed: deliberation scores as only the moderator/verdict model varies\. Moderator choice changes Balanced F1 by2\.62\.6–3\.73\.7points\. Consolidation results are shown in[Figure5](https://arxiv.org/html/2608.12585#S4.F5)\.Table 7:Diversity control: homogeneous panels of three independent temperature\-1\.01\.0samples of one model with a same\-model moderator \(minor\+, All records, step\- level, propagation\-aware;n=200n=200\)\. Within each panel: the deliberated consensus followed by the three individual solo samples \(interchangeable draws of the same model, so their spread reflects sampling variability\)\. Even with zero model diversity the consensus rises well above every solo sample, and the gain is recall\-driven \(precision roughly flat or slightly lower\), but the ceiling tracks base\-model strength\.
## Appendix HDeliberation vs\. naive aggregation: full metrics
[Tables8](https://arxiv.org/html/2608.12585#A8.T8)and[9](https://arxiv.org/html/2608.12585#A8.T9)give the full metric breakdown \(Bal\. Acc, Bal\. F1, Acc\., Precision, Recall, F1\) underlying the comparison between the moderatedconsensusverdict and two naive Phase 1 aggregation baselines: theunionof every juror’s independently flagged steps, andmajority vote\(a step counts as flagged only if a strict majority of jurors — more than half — independently cited it\)\. All rows use theminor\+threshold \(positives = minor/major/fatal\) on the*All records*slice, scored step\-level against Hard2Verify human labels\.
Table 8:Moderated deliberation vs\. union and majority\-vote aggregation of independent Phase 1 juror judgements, for the four heterogeneous panels in the*Jury*strip figure \([Section4\.2](https://arxiv.org/html/2608.12585#S4.SS2)main text\)\. Majority vote collapses recall \(38–63%\) despite the highest precision of the three aggregation rules on every panel, dragging Bal\. F1 well below both union and consensus\. Union recovers most of consensus’s Bal\. F1 \(within 1–3 points on three of four panels\) but does so via a different precision/recall trade\-off in every panel: on Frontier, union’s recall exceeds consensus’s \(0\.884 vs\. 0\.871\) at a steep precision cost \(0\.722 vs\. 0\.775\); on Large OSS and Small/Medium OSS, union’s recall is comparable to or below consensus’s while precision trails by 2\.6–4\.1 points\. Deliberation’s bold entry per panel marks the better of consensus vs\. union on each column\.Table 9:Moderated consensus vs\. union and majority\-vote aggregation of independent Phase 1 juror judgements, for the two homogeneous \(3 identical jurors, same\-model moderator\) panels in the*Diversity of the panel matters*strip figure \([Section4\.5](https://arxiv.org/html/2608.12585#S4.SS5)main text\)\. For gpt\-oss\-120b, deliberation improves on union across every metric \(largest gain: recall 0\.741→\\to0\.786\)\. For qwen3\.6\-27b — the weakest juror model in this report — union’s Bal\. F1 \(0\.735\) and F1 \(0\.701\) both*exceed*consensus’s \(0\.725, 0\.698\): deliberation loses more true positives than it gains in precision for this panel, the only reversal of the consensus\-beats\-union pattern we observe across all six panels in this appendix\. Consensus’s Bal\. Acc \(0\.756\) is likewise essentially tied with union’s \(0\.754\)\.Deliberation’s effect on Bal\-F1 itself is systematically different depending on what the panel needs: for the frontier panel, where gpt\-5\.4 alone nearly saturates the task, unioning in four additional jurors’ raw claims increases recall above consensus \(88\.4 vs\. 87\.1\) but at a precision cost \(72\.2 vs\. 77\.5\), so deliberation trades a small amount of that recall back for a larger precision gain\.
For panels dominated by weaker jurors \(Large OSS, Small/Med OSS\), the same effect appears in milder form, where recall is slightly degraded and precision is higher in deliberation compared to union \(e\.g\. Small/Med OSS: \+4\.1 precision,−\-1\.1 recall, net \+1\.4 Bal\-F1\); and for a single strong model resampled three times \(3×\\timesgpt\-oss\-120b\), deliberation instead acts as a recall\-recovery mechanism, surfacing defects that no individual sample caught while precision is unchanged \(74\.1→\\rightarrow78\.6 recall at constant∼\\sim81% precision\)\. This adaptivity \(pruning spurious claims when the panel is noisy, trading recall for precision when a dominant juror over\-inflates the union, and recovering missed claims when the panel is merely under\-sampled, all on top of a consolidated and evidence\-backed final verdict\) is a capability neither union nor majority vote rule can offer\. Deliberation instead diagnoses and corrects the specific failure mode present in a given panel’s raw judgements, at the cost of the additional inference required to reach that single verdict\. These behaviors may be partly an artifact of Hard2Verify itself: on a benchmark where jurors are already reasonably well\-calibrated defect detectors, the marginal false positive is more common than the marginal false negative, so pruning has more room to help than recovery does\. On a harder task where models struggle to identify defects at all we would expect this balance to flip, with deliberation’s recall\-recovery role \(as seen here only in the homogeneous gpt\-oss\-120b panel\) becoming the dominant source of benefit rather than the exception\.
## Appendix IOn the difficulty of Hard2Verify
We chose Hard2Verify\[[25](https://arxiv.org/html/2608.12585#bib.bib6)\]as our benchmark because it is, to our knowledge, the hardest and most contamination\-resistant publicly available step\-level reasoning\-defect dataset\. Two properties support this: the*source*of its problems \(which governs both intrinsic difficulty and the chance of train–test contamination\) and the*length*of the reasoning being verified \(a segmentation\-independent proxy for how much there is to get wrong\)\. We contrast it throughout with ProcessBench\[[31](https://arxiv.org/html/2608.12585#bib.bib9)\], the closest prior step\-level error\-detection benchmark\.
##### Data source: difficulty and contamination\.
ProcessBench draws its problems from four established public test sets: GSM8K\[[7](https://arxiv.org/html/2608.12585#bib.bib7)\]\(grade\-school arithmetic\), MATH \(competition\), and OlympiadBench and Omni\-MATH \(olympiad\)\. Two of these tiers are effectively saturated by modern models, and—because the problems have been public for years and are widely used for training—even the harder tiers carry a real risk of train–test contamination: a verifier may have seen a given problem \(and its canonical solution\) during pre\-training, so strong verification scores can partly reflect memorised solutions rather than genuine step\-by\-step checking\. Hard2Verify is constructed to remove both issues\. Its problems are drawn from*very recent*open\-ended competition and olympiad mathematics—the regime in which frontier LLM systems only reached gold\-medal level at IMO 2025—chosen to postdate typical training cutoffs, so a high score is much harder to obtain by recall\. The solutions being verified are themselves generated by*frontier*models on open\-ended problems, rather than reformatted outputs of smaller open\-source models\. The result is a benchmark whose difficulty is uniformly at the frontier, with no easy floor and a reduced contamination surface\.
##### Record length\.
Length is a difficulty signal that does not depend on how a solution is segmented into steps: a longer solution has more places for a subtle error to hide and a longer dependency chain a verifier must follow\. We measured the length of every solution being verified—in tokens, using a byte\-level BPE tokenizer—for Hard2Verify and for each ProcessBench subset;[Table10](https://arxiv.org/html/2608.12585#A9.T10)reports the distributions\. Two patterns stand out\. First, within ProcessBench length tracks difficulty exactly as expected, rising from∼260\\sim 260tokens on grade\-school GSM8K to∼760\\sim 760on the olympiad\-level subsets\. Second, Hard2Verify solutions are far longer than any ProcessBench tier: a mean of∼2,000\\sim 2\{,\}000tokens \(median∼1,400\\sim 1\{,\}400\), roughly2\.7×2\.7\\timesthe olympiad\-level ProcessBench subsets and nearly8×8\\timesgrade\-school GSM8K, with a heavy tail \(90th percentile∼3,900\\sim 3\{,\}900tokens, maximum over12,00012\{,\}000\)\. ProcessBench’s*longest*olympiad solution is close to Hard2Verify’s*median*\. Longer, frontier\-generated, open\-ended reasoning is precisely the regime a single judge struggles to verify reliably, and where a deliberating jury has the most to add\.
Table 10:Length of the solution being verified, in byte\-level BPE tokens, for Hard2Verify and each ProcessBench subset\. Hard2Verify solutions are markedly longer than even the olympiad\-level ProcessBench subsets, a segmentation\-independent indication of its higher difficulty ceiling\. Absolute counts depend on the tokenizer; the cross\-dataset*ratios*do not\.
## Appendix JDefect taxonomy and per\-subcategory fingerprints
This appendix gives the methodology behind the shared defect taxonomy used in[Section5](https://arxiv.org/html/2608.12585#S5), the taxonomy itself, and the full per\-subcategory fingerprint table\.
##### How the taxonomy is built\.
The taxonomy is*emergent*\(induced from the data, not hand\-authored\) and*shared*\(induced once over the pooled defects of all three models, so every model’s distribution is over the same categories\)\. It is built by a single LLM in three passes over the pooled defect set\.\(1\) Grow\.The pooled defects are deterministically interleaved so each batch mixes models, then processed in batches\. Each grow step shows the LLM the current taxonomy plus a new batch of defect descriptions and asks for the updated taxonomy under a strict rule: preserve every existing category and subcategory verbatim, and add a node only when a defect fits nothing existing\. This is accretion rather than rewrite, so the taxonomy grows monotonically and converges as batches stop introducing novel error kinds\.\(2\) Consolidate\.One merge pass dedupes near\-identical nodes, keeps the set mutually exclusive and collectively exhaustive, and freezes the two\-layer \(category→\\rightarrowsubcategory\) shape into a flat list of leaves with stable ids\.\(3\) Assign\.A final pass labels every defect against the frozen leaves; assignment first resets any prior label so an omitted defect stays explicitly unassigned rather than retaining a stale id, and leftover unassigned defects are retried until covered\. The run reported here yields8 categories and 56 subcategory leaves\. Because grow and consolidate are LLM passes at low but non\-zero temperature, a re\-run can produce slightly different labels and counts; the taxonomy below is the result of*this*run, not a canonical fixed schema\.
##### The taxonomy\.
The eight top\-level categories and their subcategory leaves are:
Computation & algebraic manipulation \(8\)\.Numerical calculation error; misusing a previously computed or given value; modular arithmetic error; other arithmetic/computation error; sign error; incorrect expansion, simplification, or factoring; incorrect formula, identity, or substitution; invalid or mislabeled manipulation\.
Logical reasoning \(8\)\.Incorrect claim about a mathematical fact; faulty logical inference \(quantifiers, necessity/sufficiency\); failure to handle contradictions or dismiss/keep cases correctly; conflating distinct concepts; incorrect example construction or verification; reversing direction of inequality or bound; incorrect independence or factorization assumptions; other logical\-reasoning error\.
Combinatorial/Structural reasoning \(6\)\.Incorrect recurrence or counting formula; double\-counting or overcounting; missing cases or regions in decomposition; unjustified structural or completeness claims; invalid base cases or boundary conditions; incomplete enumeration of valid configurations\.
Geometric/Spatial reasoning \(7\)\.Incorrect coordinate or vector computations; flawed geometric arguments or property claims; incorrect transformation application; spatial orientation or ordering error; applying geometric formulas to invalid configurations; misidentifying geometric shapes or regions; other geometric/spatial error\.
Problem interpretation \(6\)\.Misreading or altering given conditions; substituting fabricated conditions for stated ones; incorrect mathematical modeling or constraint formulation; omitting stated constraints from the solution; misidentifying structural or geometric roles; misreading diagrams or code specifications\.
Proof methodology and rigor \(8\)\.Unproven assumptions used as established facts; incomplete case analysis or verification; overgeneralization or ad\-hoc argument as proof; answer asserted without derivation \(guessing or memorized recall\); insufficient or erroneous bounds/estimates; circular or self\-referential verification; applying proof techniques with violated prerequisites; approximate numerical evidence treated as exact proof\.
Self\-monitoring and error\-handling \(7\)\.Extended reasoning on an uncorrected false premise; self\-corrected error that does not propagate; failure to diagnose the source of a detected error; multiple contradictory values left unresolved; unsupported final answer disconnected from derivation; final answer contradicting correct intermediate work; other self\-monitoring/error\-handling failure\.
Output/generation failures \(6\)\.Truncated or incomplete solution; infinite loop or cycling without progress; verbose redundancy without progress; abandoning the problem for alternative interpretations; abandoning a productive approach without resolution; degenerate or garbled output\.
## Appendix Knemotron\-3\-super: subcategory distribution and worked examples
This appendix supports the close read of nemotron\-3\-super in[Section5\.2](https://arxiv.org/html/2608.12585#S5.SS2): the full subcategory \(leaf\) breakdown of its defects, and a worked example of a single trace with defects at each severity level\.
Figure 10:nemotron\-3\-super: flagged reasoning steps by taxonomy category, stacked by severity \(regex segmentation, AIME2026,1,9201\{,\}920traces\)\. Output/generation failures and self\-monitoring/error\-handling are the two largest categories by flagged\-step count; logical reasoning is the most fatal\-heavy, whereas proof methodology and rigor leans major/minor \(lapses in rigor rather than false claims\)\.### K\.1Subcategory \(leaf\) distribution
[Table11](https://arxiv.org/html/2608.12585#A11.T11)lists every non\-zero taxonomy leaf for nemotron\-3\-super, grouped by parent category and sorted by defect count within category\. Of the5656leaves in the shared taxonomy,5454are populated by at least one of nemotron\-3\-super’s3,3843\{,\}384defects \(a defect may cite several steps; this table counts defects, not flagged steps, so totals differ from[Figure10](https://arxiv.org/html/2608.12585#A11.F10), which counts flagged steps\)\. One leaf dominates:*extended reasoning on an uncorrected false premise*\(927927defects,27\.4%27\.4\\%of the model’s defects\), accounting for the bulk of the fatal mass seen in the self\-monitoring category of[Figure10](https://arxiv.org/html/2608.12585#A11.F10); the next\-largest leaves are*truncated or incomplete solution*\(254254,7\.5%7\.5\\%\) and*incorrect claim about a mathematical fact*\(218218,6\.4%6\.4\\%\)\.
Table 11:nemotron\-3\-super: every populated taxonomy leaf, grouped by category and sorted by defect count within category\. Percentages are of the model’s3,3843\{,\}384total defects\. Regex segmentation, AIME2026,1,9201\{,\}920traces\.SubcategoryCount% of defectsComputation & algebraic manipulation \(476 defects\)Numerical calculation error1581584\.7%4\.7\\%Incorrect formula, identity, or substitution1031033\.0%3\.0\\%Misusing a previously computed or given value99992\.9%2\.9\\%Invalid or mislabeled manipulation48481\.4%1\.4\\%Sign error27270\.8%0\.8\\%Incorrect expansion, simplification, or factoring26260\.8%0\.8\\%Modular arithmetic error15150\.4%0\.4\\%Logical reasoning \(554 defects\)Incorrect claim about a mathematical fact2182186\.4%6\.4\\%Conflating distinct concepts1041043\.1%3\.1\\%Failure to handle contradictions or dismiss/keep cases correctly82822\.4%2\.4\\%Faulty logical inference — quantifiers, necessity/sufficiency61611\.8%1\.8\\%Incorrect independence or factorization assumptions36361\.1%1\.1\\%Other logical\-reasoning error27270\.8%0\.8\\%Incorrect example construction or verification17170\.5%0\.5\\%Reversing direction of inequality or bound990\.3%0\.3\\%Combinatorial/Structural reasoning \(228 defects\)Unjustified structural or completeness claims60601\.8%1\.8\\%Incorrect recurrence or counting formula54541\.6%1\.6\\%Missing cases or regions in decomposition48481\.4%1\.4\\%Double\-counting or overcounting45451\.3%1\.3\\%Incomplete enumeration of valid configurations13130\.4%0\.4\\%Invalid base cases or boundary conditions880\.2%0\.2\\%Geometric/Spatial reasoning \(93 defects\)Flawed geometric arguments or property claims44441\.3%1\.3\\%Spatial orientation or ordering error32320\.9%0\.9\\%Incorrect coordinate or vector computations880\.2%0\.2\\%Applying geometric formulas to invalid configurations440\.1%0\.1\\%Incorrect transformation application330\.1%0\.1\\%Misidentifying geometric shapes or regions220\.1%0\.1\\%Problem interpretation \(113 defects\)Misreading or altering given conditions36361\.1%1\.1\\%Incorrect mathematical modeling or constraint formulation31310\.9%0\.9\\%Omitting stated constraints from the solution16160\.5%0\.5\\%Substituting fabricated conditions for stated ones14140\.4%0\.4\\%Misreading diagrams or code specifications14140\.4%0\.4\\%Misidentifying structural or geometric roles220\.1%0\.1\\%Proof methodology and rigor \(406 defects\)Incomplete case analysis or verification1301303\.8%3\.8\\%Unproven assumptions used as established facts1221223\.6%3\.6\\%Overgeneralization or ad\-hoc argument as proof51511\.5%1\.5\\%Applying proof techniques with violated prerequisites33331\.0%1\.0\\%Circular or self\-referential verification25250\.7%0\.7\\%Answer asserted without derivation — guessing or memorized recall23230\.7%0\.7\\%Insufficient or erroneous bounds/estimates20200\.6%0\.6\\%Approximate numerical evidence treated as exact proof220\.1%0\.1\\%Self\-monitoring and error\-handling \(1132 defects\)Extended reasoning on an uncorrected false premise92792727\.4%27\.4\\%Self\-corrected error that does not propagate1111113\.3%3\.3\\%Failure to diagnose the source of a detected error59591\.7%1\.7\\%Multiple contradictory values left unresolved16160\.5%0\.5\\%Other self\-monitoring/error\-handling failure880\.2%0\.2\\%Final answer contradicting correct intermediate work660\.2%0\.2\\%Unsupported final answer disconnected from derivation550\.1%0\.1\\%Output/generation failures \(382 defects\)Truncated or incomplete solution2542547\.5%7\.5\\%Abandoning a productive approach without resolution85852\.5%2\.5\\%Verbose redundancy without progress21210\.6%0\.6\\%Degenerate or garbled output10100\.3%0\.3\\%Abandoning the problem for alternative interpretations770\.2%0\.2\\%Infinite loop or cycling without progress550\.1%0\.1\\%The severity\-stacked, top\-three\-leaves\-per\-category view of this same data is given as[Figure8](https://arxiv.org/html/2608.12585#S5.F8)in the main text \([Section5\.2](https://arxiv.org/html/2608.12585#S5.SS2)\)\.
### K\.2Example defects by severity
To make the taxonomy and severity labels concrete, we walk through one nemotron\-3\-super record \(AIME2026 problem 28, sample 4\) on which the jury raised defects at all three non\-neutral severities\. The problem asks for the minimum size of a setSSof integers with exactly40404040“cousins”TT\(disjoint, same size, and pairable with elements differing by exactly11\)\. The trace runs to210210regex\-segmented steps and is truncated before reaching a final answer \(the correct answer is107107, from the factorization4040=23×5×1014040=2^\{3\}\\times 5\\times 101realised by components of sizes1,1,1,4,1001,1,1,4,100\)\.
The jury raises three defects on three different steps:
fatal: Truncated or incomplete solution \(Output/generation failures\)\.At\[STEP\-205\], the trace establishes a correct component\-based counting formula \(the number of cousins is the product of\(ki\+1\)\(k\_\{i\}\+1\)over path components\) by\[STEP\-201\], but never applies it to the target value40404040: it does not factor40404040, does not choose the factorization that minimises\|S\|\|S\|, and does not construct a concrete set\. It terminates mid\-sentence \(“ edges: 0\-\(\-1\), 1\- ”\) while enumerating the exampleS=\{0,1,3\}S=\\\{0,1,3\\\}, and the final steps break off while beginning thek≥2k\\geq 2case — so the problem is never answered despite the machinery being in place\.
major: Unproven assumptions used as established facts \(Proof methodology and rigor\)\.At\[STEP\-49\], the proof that the matching fromSStoTTis unique considers only22\-element swaps and concludes “you cannot swap assignments and keepTTthe same unlesss1=s2s\_\{1\}=s\_\{2\}”; it never rules out longer cycles \(e\.g\. a33\-cycles1→t2,s2→t3,s3→t1s\_\{1\}\{\\to\}t\_\{2\},\\,s\_\{2\}\{\\to\}t\_\{3\},\\,s\_\{3\}\{\\to\}t\_\{1\}\) that could produce the same setTT\. The uniqueness conclusion is in fact correct, but the argument as written has a gap\.
minor: Unproven assumptions used as established facts \(Proof methodology and rigor\)\.At\[STEP\-21\], after examiningS=\{0,1,2\}S=\\\{0,1,2\\\}the trace states “SScannot have cousins if it contains consecutive numbers”, which is too strong;SSmay contain consecutive pairs \(e\.g\.S=\{0,1\}S=\\\{0,1\\\}has cousinT=\{−1,2\}T=\\\{\-1,2\\\}\)\. The correct restriction, which the trace itself states one step later in\[STEP\-22\], is thatSScontains no three consecutive integers; the misstatement is immediately self\-corrected\.
## Appendix LReproducibility notes
The framework supports six model providers \(AWS Bedrock, OpenAI, Google Gemini, OpenRouter, OpenAI\-compatible local clusters, and a Bedrock Mantle gateway\), selected per model ID by prefix\. Per\-juror clients use separate connection pools to avoid TLS\-handshake races under parallel fan\-out\. The validation harness is parallel, resumable, and logs every LLM call \(caller, phase, model, prompts, response, reasoning text, token usage\) for audit\.
## Appendix MDeltaBench: full results and per\-domain analysis
This appendix expands the DeltaBench generalization result \([Section4\.6](https://arxiv.org/html/2608.12585#S4.SS6)\) with the complete metric set under both scoring protocols, the full aggregation ladder, and per\-domain breakdowns\.
##### Setup\.
DeltaBench\[[15](https://arxiv.org/html/2608.12585#bib.bib18)\]provides1,2361\{,\}236long chain\-of\-thought traces from o1\-style generators \(QwQ, DeepSeek\-R1\) with human annotations at the*section*level \(a section groups consecutive steps addressing one sub\-task\)\. We map each annotated section to one\[STEP\-k\]marker of our input format, discarding DeltaBench’s finer step indexing, and run our pipeline unchanged \(standard defect prompt, minor\+\+threshold, defect propagation on\)\. Ground truth follows DeltaBench’s own scoring convention: a section is positive if it appears in either the annotated*error*list or the annotated*unuseful*list \(their published evaluator unions the two\)\. We report two metric families side by side:
- •Ours: the corpus\-level step metrics used throughout this paper: Balanced Accuracy, Balanced F1, Accuracy, Precision, Recall, and F1, with error as the positive class, pooled over all steps of all records\.
- •Theirs: DeltaBench’s protocol, replicated exactly on our structured predictions: per\-sample precision/recall/F1 over the set of predicted vs\. annotated error sections, with predictions beyond the last annotated section discarded, then macro\-averaged across samples \(F1M\{\}\_\{\\text\{M\}\}\); micro variants pool TP/FP/FN over samples before computing the ratios\. Because our verdicts already carry structured step citations, no LLM\-based answer extraction is involved\.
The jury is the homogeneous3×3\\timesgpt\-oss\-120b panel with a same\-model moderator; the single\-model comparator is one opus\-4\.6 Phase 1 pass with the identical prompt\. All rows are computed on the same1,2261\{,\}226records \(“paired”\): opus could not complete1010of the1,2361\{,\}236traces \(the longest in the benchmark\) returning empty responses even at a24,00024\{,\}000\-token output ceiling after multiple retries; those records are excluded from both systems\. Since the excluded traces are concentrated in the hardest math records, the pairing, if anything, favors the single\-model baseline\.
##### Overall results\.
[Table12](https://arxiv.org/html/2608.12585#A13.T12)reports the full ladder\. There are three main observations\. First, the deliberated consensus beats every single system \(except gpt\-5\.4 where the performance is very close\) under both metric families:\+5\.8\+5\.8Balanced F1 over the best individual juror and\+3\.1\+3\.1over opus solo under our metrics;\+8\.2\+8\.2and\+4\.5\+4\.5macro\-F1 respectively under theirs\. Second, the aggregation ladder replicates the Hard2Verify pattern exactly: majority vote barely improves on a single juror \(38\.738\.7macro\-F1 vs\.39\.739\.7best solo\), the union of Phase 1 judgements captures most of the available gain \(48\.248\.2\), and deliberation lands beside the union \(47\.947\.9\) while achieving the highest recall of any method \(65\.765\.7macro\-recall\) — on this benchmark, at this severity threshold, deliberation and union are near\-equivalent on F1 and differ in their precision/recall trade\-off\. Third, for leaderboard context, the strongest critic reported by[15](https://arxiv.org/html/2608.12585#bib.bib18)is gpt\-4\-turbo at40\.840\.8macro\-F1: the jury exceeds it by77points using three small open\-weight models, and even our opus\-4\.6 single\-model baseline \(43\.443\.4\) exceeds it, indicating that part of the advantage comes from our defect\-prompt format and propagation convention, and the rest \(\+4\.5\+4\.5\) from deliberation\.
Table 12:DeltaBench, overall \(n=1,226n\{=\}1\{,\}226paired records\)\. Left block: our corpus\-level step metrics\. Right block: DeltaBench’s own protocol \(per\-sample macro\-averaged F1/precision/recall and micro\-F1\)\. The deliberated consensus beats every individual system except gpt\-5\.4 solo, which lands beside it under both metric families \(61\.861\.8vs\.61\.561\.5Balanced F1;47\.247\.2vs\.47\.947\.9F1M\{\}\_\{\\text\{M\}\}\) with an essentially identical recall profile — replicating the Hard2Verify finding that gpt\-5\.4 alone saturates what the panel achieves\. Union and consensus are near\-equivalent on F1, with consensus taking the highest recall\. DeltaBench’s best published critic scores40\.840\.8F1M\{\}\_\{\\text\{M\}\}\[[15](https://arxiv.org/html/2608.12585#bib.bib18)\]\.
## Appendix NDefect\-guided retry with jury feedback
This section details the downstream evaluation summarized in[Section6](https://arxiv.org/html/2608.12585#S6)\. Whereas the profiling analysis characterizes the types of defects produced by a reasoning model, this experiment asks whether the jury’s structured findings help that model correct its own solutions\. We evaluate retries against the AIME answer key, which provides an external success criterion independent of the jury\.
### N\.1Design
##### Data\.
For each of the3030AIME2026 problems, we generate6464fresh solutions with nemotron\-3\-super, for a total of1,9201\{,\}920traces\. Generation uses temperature0\.70\.7, a maximum of32,00032\{,\}000tokens, and high reasoning effort, matching[Section5\.1](https://arxiv.org/html/2608.12585#S5.SS1)\. All generation calls complete successfully, and exact\-answer scoring gives a corpus pass@1 of80\.0%80\.0\\%\(1,536/1,9201\{,\}536/1\{,\}920\)\.
We segment every trace with the regular\-expression procedure of[Section2\.1](https://arxiv.org/html/2608.12585#S2.SS1)and evaluate it with the profiling jury: glm\-5\.2, minimax\-m3, deepseek\-v4\-pro, and kimi\-k2\.6, with deepseek\-v4\-pro as moderator\. The jury produces a deliberated verdict for every trace and identifies3,3663\{,\}366non\-neutral defects\. In total,877877traces \(45\.7%45\.7\\%\) contain at least one defect with minor, major, or fatal severity\. We exclude neutral findings from the retry prompts because they do not identify an actionable error\. Among the384384wrong\-answer traces, the resulting gate flags372372\(96\.9%96\.9\\%recall\); the remaining1212receive no non\-neutral finding\.
##### Retry conditions\.
The same nemotron\-3\-super deployment retries every flagged trace at temperature0\.70\.7, with a maximum of32,00032\{,\}000generation tokens and high reasoning effort\. Each trace is evaluated under two conditions:
full\(rich feedback\)\.The model receives the problem, its previous\[STEP\-k\]\-marked trace, and every non\-neutral jury finding\. Each finding includes its cited step locations, severity, and verbatimwhat\_went\_wrongdiagnosis\.
steps\(anchors only\)\.The model receives the same problem and previous trace, but the findings are reduced to a deduplicated list of cited\[STEP\-k\]locations, with no severity labels or explanations\.
This design yields paired outcomes for all877877flagged traces and1,7541\{,\}754retry generations in total\. Every retry call completes, although4949stepsoutputs and2727fulloutputs contain no parseable integer and are scored as incorrect\. We first extract an answer from an explicitANSWERmarker and only then fall back to the reasoning text; this prevents a number in a truncated reasoning tail from being accepted as the answer\. Both prompts instruct the model to re\-solve the problem from scratch rather than patch individual steps\. The paired contrast therefore measures the incremental value of diagnosis and severity when both conditions already provide the same localization signal\.
### N\.2Prompts
The conditions use the same prompt scaffold and differ only in the review\-findings block\. Thestepsprompt is reproduced below\. Placeholders in braces are filled separately for each trace; when a finding cites multiple steps, every cited location is included in the deduplicated list\.
Youpreviouslyattemptedthefollowingcompetitionproblem\.
\#\#Problem
\{problem\}
\#\#Yourprevioussolution
\{marked\_trace\}
\#\#Reviewfindings
Anautomatedreviewofyoursolutionflaggedissuesatthe
followingsteps:
\-\[STEP\-14\]
\-\[STEP\-15\]
\#\#Task
Re\-solvetheproblemfromscratch\.Usetheflaggedlocationsto
decidewhichpartsofyourpreviousapproachtodistrust;donot
assumeunflaggedstepsarecorrect\.Showyourreasoning,thenend
with:
ANSWER:<integer\>
Listing 7:Thestepsretry prompt\.Thefullprompt replaces the location list with complete findings that include severity and the jury’s explanation:
\-\[STEP\-14\]\-\-FATAL
"In\[STEP\-14\],thetracewritesanincorrectintermediate
expression:’right:18\*\(2v\(v\+9\)/9\)=4\*2v\(v\+9\)’\.\.\."
Listing 8:Thefullarm’s findings block \(one entry shown\)\.Because every AIME answer is an integer in\[0,999\]\[0,999\], theANSWER: <integer\>contract matches the benchmark’s answer format\. Jury explanations sometimes state an explicit correction or corrected value\. Accordingly,fullrepresents the complete actionable feedback produced by the jury, including such corrections when present;stepsisolates the information supplied by localization alone\.
### N\.3Results
##### Outcome transitions\.
[Table13](https://arxiv.org/html/2608.12585#A14.T13)separates beneficial transitions \(an initially wrong answer becomes correct\) from regressions \(an initially correct answer becomes wrong\)\. Rich feedback improves both sides of this tradeoff: it corrects48\.7%48\.7\\%of the372372flagged wrong answers, compared with39\.2%39\.2\\%for step anchors, while reducing the regression rate from5\.3%5\.3\\%to3\.6%3\.6\\%among the505505flagged correct answers\.
Table 13:Retry outcome transitions by feedback condition \(6464samples per problem\)\. The non\-neutral\-defect gate selects372372of the384384wrong\-answer traces, so both conditions cover96\.9%96\.9\\%of wrong answers while retrying only877877of the1,9201\{,\}920traces\. Net is the number of fixes minus regressions\.
##### Matched comparison\.
Before retry, accuracy on the877877flagged traces is57\.6%57\.6\\%\(505/877505/877\)\. Step anchors raise observed retry accuracy to71\.2%71\.2\\%\(624/877624/877\), whereas rich feedback raises it to76\.2%76\.2\\%\(668/877668/877\)\. Relative to anchors alone, rich feedback therefore adds5\.05\.0percentage points, comprising3535additional fixes and nine fewer regressions\.
The paired outcomes are discordant on112112traces: rich feedback alone is correct on7878, whereas step anchors alone are correct on3434\. An exact McNemar test givesp=4×10−5p=4\\times 10^\{\-5\}\. Because the6464traces drawn for each problem are not independent, we also compute paired differences at the problem level\. The mean improvement is\+4\.1\+4\.1percentage points, with a bootstrap95%95\\%confidence interval of\[\+2\.3,\+6\.1\]\[\+2\.3,\+6\.1\]\. Among the2929problems with at least one flagged trace,fulloutperformsstepson1515, underperforms on one, and ties on1313\. Both analyses support a positive incremental effect from the richer feedback package\. They do not, however, separate the contribution ofwhat\_went\_wrongfrom that of severity\.
##### Defect\-gated retry policies\.
The experiment also supports a simple selective\-retry policy: retry a trace when the jury reports a qualifying defect, and otherwise retain the original answer\. The table below reports corpus accuracy over all1,9201\{,\}920traces, whose original pass@1 is0\.8000\.800\.
Table 14:Corpus accuracy under selective retry policies\. Traces that do not satisfy the gate retain their original answers\.
With the any\-defect gate, the system retries45\.7%45\.7\\%of the corpus while retaining the original answers for the1,0431\{,\}043unflagged traces\. Combining this gate with rich feedback yields88\.5%88\.5\\%corpus accuracy,8\.58\.5percentage points above the original pass@1\. This absolute improvement includes any generic benefit from a second attempt because the experiment has no blind retry condition\. The comparison under the shared gate is more direct: rich feedback exceeds step anchors by2\.32\.3percentage points on the full corpus \(0\.8850\.885vs\.0\.8620\.862\)\.
Restricting retries to fatal findings reduces the number of retries from877877to338338, or38\.5%38\.5\\%of the any\-defect volume\. With rich feedback, this policy still reaches85\.3%85\.3\\%accuracy,5\.35\.3percentage points above pass@1 but3\.23\.2points below the any\-defect policy\. The difference reflects wrong\-answer traces whose only findings have minor or major severity and are therefore not retried by the fatal\-only gate\.
### N\.4Limitations
1. 1\.The experiment covers one generator, one benchmark, and one retry sample per condition and trace\. The per\-problem analysis addresses dependence among traces from the same problem, but broader generalization remains untested\.
2. 2\.There is no blind\-retry condition\. Some improvement over the original answers may therefore come from making a second attempt rather than from the feedback itself\. The pairedfull\-versus\-stepscomparison controls for this shared retry opportunity\.
3. 3\.Thefullcondition bundles diagnosis and severity, and some diagnoses contain explicit corrections\. The experiment consequently measures the complete rich\-feedback package; it does not isolate the effect of either field or distinguish explanation from a supplied fix\.
4. 4\.The gate’s96\.9%96\.9\\%wrong\-answer recall is specific to this corpus and does not show that each flagged defect caused the corresponding wrong answer\. Failures caused by omissions may also escape detection\.
5. 5\.The jury findings are not human\-validated ground truth\. The external AIME answer key validates the correctness of retry outcomes, not the factual accuracy or severity of individual defect descriptions\.
### N\.5Artifacts
The traces, per\-trace deliberated verdicts, and extracted defects are stored incdd\_artifacts/defect\_profiling/aime2026\_nemotron64\_regex/\. Retry generations are stored as one JSON file per trace and condition, together with the extracted answer, correctness label, and prompt metadata, underretries/\\\{steps,full\\\}/\. The pipeline wrapperrun\_nemotron64\.py, retry harnessretry64\.py, and run logsrun\.logandretry64\.logare located incdd\_artifacts/defect\_profiling/\.Similar Articles
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
This paper proposes a framework that ensembles the reasoning structures of multiple LLMs by weighted merging of extracted Directed Acyclic Graphs (DAGs), enabling consensus reasoning with improved accuracy and interpretability across several benchmarks.
ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces
Introduces ReasoningFlow, a framework to capture discourse structures of large language model reasoning traces as directed acyclic graphs, enabling fine-grained analysis of reasoning behaviors like self-reflection and backtracking. Based on manual and automatic annotation of thousands of traces, it reveals structural similarities across models and that most erroneous steps do not contribute to final answers.
JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification
JuryProbe introduces an empirical diagnostic for assessing consensus risk in reference-free LLM judge panels for factuality checking, and uses a calibration-based routing policy to ground high-risk decisions with trusted references, reducing false accepts.
Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering
This paper introduces a multi-agent peer-reviewed reasoning method where multiple LLMs independently generate chain-of-thought reasoning and then evaluate each other's outputs to select the best answer. The method outperforms single-model reasoning and majority voting on medical QA benchmarks.
Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging
The study investigates whether LLM-as-a-Judge evaluators reliably assess psychological depth in LLM-generated stories, revealing that human preferences are heterogeneous while judges exhibit bias towards reasoning outputs based on surface features.