Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
Summary
This paper proposes a framework that ensembles the reasoning structures of multiple LLMs by weighted merging of extracted Directed Acyclic Graphs (DAGs), enabling consensus reasoning with improved accuracy and interpretability across several benchmarks.
View Cached Full Text
Cached at: 07/31/26, 10:02 AM
# Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
Source: [https://arxiv.org/html/2607.27783](https://arxiv.org/html/2607.27783)
Amruta Parulekar,Jinu Lee,Dilek Hakkani\-Tür,Hari Sundaram University of Illinois Urbana\-Champaign \{amp20, jinulee2, dilek, hs1\}@illinois\.edu
###### Abstract
Large Language Models \(LLMs\) explore problems through chain\-of\-thought, but this exploration is buried in unstructured prose\. On high\-stakes tasks, users cannot tell which steps are well\-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded\. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs \(DAGs\) extracted from reasoning chains\. We weight each step by how many traces independently attest to it, to return “Consensus Reasoning”\. Across six benchmarks spanning statutory interpretation, graduate\-level science, narrative multi\-hop reasoning, and first\-order logic, our ensemble outperforms a matched\-budget majority\-vote baseline, with a maximum accuracy gain of 3\.1% on MuSR\-MM \(narrative multi\-hop reasoning\)\. On a single model, the framework matches or exceeds self\-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph\. Ensemble weights correlate with LLM\-judge rankings of reasoning quality at Spearmanρ=0\.30\\rho=0\.30–0\.510\.51, and consensus subgraphs are preferred over alternatives leading to the majority\-vote answer in 54\.4–65\.4% of head\-to\-head comparisons across five of six datasets\. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem\.111Our code is available at[this link](https://anonymous.4open.science/r/REASONING-CONSENSUS-5058)
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
Amruta Parulekar, Jinu Lee, Dilek Hakkani\-Tür, Hari SundaramUniversity of Illinois Urbana\-Champaign\{amp20, jinulee2, dilek, hs1\}@illinois\.edu
## 1Introduction
We propose a framework that ensembles the*reasoning structure*of multiple LLMs to surface diverse, interpretable justifications for complex reasoning tasks\. Complex domain\-specific reasoning admits multiple legitimate lines of reasoning, and a thorough analysis must consider eachYanget al\.\([2026](https://arxiv.org/html/2607.27783#bib.bib69)\)\. These get lost in the unstructured chain\-of\-thought\(Weiet al\.,[2023](https://arxiv.org/html/2607.27783#bib.bib40)\)that LLMs return\. Users see a final answer and a free\-form rationale, but cannot tell which steps and conclusions are supported or rejected, or which arguments rely on a single uncertain inference\. Smaller models compound this by committing to a single solution path, missing alternative justifications that might be betterYanget al\.\([2025b](https://arxiv.org/html/2607.27783#bib.bib65)\)\. The result is brittle reasoningZenget al\.\([2026](https://arxiv.org/html/2607.27783#bib.bib70)\), that is especially costly in low\-resource, high\-stakes settingsDahlet al\.\([2024](https://arxiv.org/html/2607.27783#bib.bib17)\)where the model might provide the right answer but a wrong supporting reason, misleading users\.
Existing strategies for improving Chain\-of\-Thought \(CoT\) either constrain reasoning diversity or discard the structure that makes reasoning auditable\. Self\-consistency\(Wanget al\.,[2023](https://arxiv.org/html/2607.27783#bib.bib51)\)performs majority voting of only the final answer, and does not consider reasoning trace structure\. Tree/Graph of Thoughts\(Yaoet al\.,[2023](https://arxiv.org/html/2607.27783#bib.bib67); Bestaet al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib50)\)show more reasoning structure, but remain bounded by the single\-model ceiling\. Multi\-model methods broaden this at the cost of structural auditability: some extend exploration across models but still commit to a single chain\(Xuet al\.,[2025](https://arxiv.org/html/2607.27783#bib.bib42)\); others aggregate at the level of unstructured prose, making it hard to trace which model contributed what\(Wanget al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib61); Chenet al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib62)\); and search\-based ensembles require trained process reward models and ultimately return a single best chain\(Parket al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib63)\)\. Thus, the structure of disagreement and alternative support is lost\.
Our framework generalises Self\-consistencyWanget al\.\([2023](https://arxiv.org/html/2607.27783#bib.bib51)\)from answers to reasoning structure, and from a single model to many\. We sample CoT traces from multiple models and ensemble their underlying reasoning graphs\. We extract a Directed Acyclic Graph \(DAG\) from each trace and merge these DAGs while preserving bundles of jointly\-supporting premises, so that each merged node carries an “attestation count”\. Steps that appear across many traces accumulate weight and steps attested by only one trace are down\-weighted automatically, suppressing reasoning step errors\. We return the most\-supported, inspectable “consensus” subgraphs and conclusions with quantification of support received by each, and other possible answers as well as the amount of support for each\.
We evaluate our framework on six reasoning benchmarks spanning statutory interpretation \(SARA\), graduate\-level science \(GPQA Diamond\), narrative multi\-hop reasoning \(MuSR\-MM, MuSR\-OP, MuSR\-TA\), and first\-order logical entailment \(FOLIO\)\. Our accuracy\-weighted 4\-model ensemble outperforms a matched\-budget 20\-trace majority\-vote baseline on every dataset, with a maximum gain of 3\.1% on MuSR\-MM\. On a single model, we match or exceed self\-consistency while additionally exposing an inspectable justification graph\. Beyond accuracy, consensus subgraphs contain better explanations than single model traces, being preferred over alternatives leading to the majority\-vote answer in 54\.4–65\.4% of head\-to\-head LLM\-judge comparisons across five of six datasets\. Our subgraph ranking correlates with the judge at Spearmanρ=0\.30\\rho=0\.30–0\.510\.51\. Qualitatively, we observe that subgraphs can contain differing problem\-solving approaches, confirming that reasoning perspective diversity can be preserved through structural aggregation\. Our contributions:
\(1\) Structural ensembling of heterogeneous LLM reasoning\.We introduce a framework to aggregate reasoning at the DAG level rather than as unstructured prose\. This allows users in high\-stakes domains to inspect and override individual steps rather than treating the LLM as a black box\.
\(2\) Attestation\-weighted confidence without auxiliary training\.We derive step\-level confidence purely from cross\-trace agreement across different models\. This gives a training\-free alternative to Process Reward Models, scaling to low\-resource domains where step\-level labels are unavailable\.
\(3\) Accuracy gains with diverse, inspectable alternatives\.Our framework outperforms majority\-vote baselines on all six benchmarks and matches or exceeds self\-consistency on single\-model trace pools, while simultaneously surfacing most\-supported subgraphs and preserving competing conclusions that other traces explored\. By retaining and not collapsing the disagreement structure, our framework naturally extends to open\-ended settings like argument generation and creative writing, where the goal is not a single correct answer but a curated set of distinct, well\-supported perspectives\.
## 2Related Work
Recent “reasoning\-heavy” models like o1\(OpenAIet al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib49)\)and DeepSeek\-R1\(Guoet al\.,[2025](https://arxiv.org/html/2607.27783#bib.bib48)\)have popularised long\-form Chain\-of\-Thought \(CoT\) traces\(Weiet al\.,[2023](https://arxiv.org/html/2607.27783#bib.bib40)\)\. However, linear CoT collapses logical dependencies more naturally modeled as graphs\. To enrich the topology of asinglemodel’s reasoning, Tree\-of\-Thoughts\(Yaoet al\.,[2023](https://arxiv.org/html/2607.27783#bib.bib67)\)branched the chain to enable backtracking, and Graph\-of\-Thoughts\(Bestaet al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib50)\)generalised this to thought graphs with aggregation and feedback\. While these approaches show more reasoning structure, their diversity remains bounded by what one isolated model can produce\.
Subsequent methods expanded exploration at the expense of structural auditability\. Trace\-level aggregation, such as Self\-Consistency\(Wanget al\.,[2023](https://arxiv.org/html/2607.27783#bib.bib51)\), samples CoTs from one model and votes on the final answer\. Collaborative Beam Search\(Xuet al\.,[2025](https://arxiv.org/html/2607.27783#bib.bib42)\)extends this by selecting the best continuation at each step via multi\-model consensus\. At the response level, Mixture\-of\-Agents\(Wanget al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib61)\)and ReConcile\(Chenet al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib62)\)generate a single response from the critiques of different agents\. Finally, structured search methods like LE\-MCTS\(Parket al\.,[2024](https://arxiv.org/html/2607.27783#bib.bib63)\)and MoSA\(Yanget al\.,[2025b](https://arxiv.org/html/2607.27783#bib.bib65)\)use Monte Carlo tree search over multi\-LLM steps, while DER\(Huet al\.,[2025](https://arxiv.org/html/2607.27783#bib.bib41)\)trains a router to select the best expert output\. These approaches generate diversity, but it ultimately collapses at aggregation time\. Through voting, prose synthesis, and reward\-guided search, the systems force convergence into a single final chain or answer and alternative paths are discarded\.
A concurrent line of work avoids this compression by preserving disagreement structure rather than discarding it\. ARGORA\(Jinet al\.,[2026](https://arxiv.org/html/2607.27783#bib.bib68)\)organizes multi\-expert discussions into argumentation graphs to identify causal dependencies between competing claims, but relies on constrained interactive debate and causal interventions to evaluate argument fragility\. Our framework instead extracts and ensembles DAGs from independent, unconstrained CoT traces across models, preserving the full weighted graph of attested reasoning steps via statistical consensus rather than counterfactual intervention\. We never force convergence: when models disagree, both conclusions are retained with their supporting subgraphs\. Consensus emerges naturally from independent reasoners\. The resulting structure remains auditable at node level\.
## 3Methodology
Our pipeline \(Figure[1](https://arxiv.org/html/2607.27783#S3.F1)\) operates in five stages—trace extraction \(Section[3\.1](https://arxiv.org/html/2607.27783#S3.SS1)\), DAG extraction \(Section[3\.2](https://arxiv.org/html/2607.27783#S3.SS2)\), node merging \(Section[3\.3](https://arxiv.org/html/2607.27783#S3.SS3)\), ensemble weighting \(Section[3\.4](https://arxiv.org/html/2607.27783#S3.SS4)\), and finding consensus \(Section[3\.5](https://arxiv.org/html/2607.27783#S3.SS5)\)—to convert multiple free\-form reasoning chains into a weighted, auditable graph and obtain the consensus reasoning\. A query is given to multiple LLMs, which sample numerous reasoning traces; a Directed Acyclic Graph \(DAG\) is extracted from each trace, the DAGs are ensembled into a weighted graph and the conclusion with the highest support and its most\-supported reasoning trace are extracted\. All prompts are in Appendix[A](https://arxiv.org/html/2607.27783#A1)\.
### 3\.1Reasoning Trace Extraction
We sample multiple stochastic reasoning traces per query from each of several small, open\-source LLMs\. Each model is constrained to a fixed schema requiring a step\-by\-step reasoning trace and a final verdict tag\. For each generation, we separate the free\-form reasoning from the verdict, retaining both for downstream DAG extraction and ensembling\.
### 3\.2DAG Extraction
Figure 1:Detailed pipeline\.Step 1:Multiple LLMs sample reasoning traces for a query\.Step 2:Each trace is decomposed into a DAG with nodes labeled Planning \(P\), Fact \(F\), Reasoning \(Re\), and Conclusion \(C\)\. Typed node labels let independently generated DAGs be aligned and merged step\-by\-step rather than compared as opaque prose\.Step 3:Nodes are merged based on semantic similarity\. Our hybrid ensembling strategy filters candidate node pairs by type \(shown by color\) and embedding similarity \(thresholdτ′\\tau^\{\\prime\}\), then verifies surviving pairs with an LLM judge\. Verified pairs merge into a single node representing the same reasoning step attested by multiple traces\. The hybrid avoids both the false positives of similarity\-only merging and theO\(n2\)O\(n^\{2\}\)cost of LLM\-only merging\.Step 4:The ensemble nodes are weighted based on cross\-trace attestation, under simple weighting \(each trace casts a unit vote, so weights depend on raw attestation counts\) or accuracy\-weighted \(each vote scaled by its source model’s held\-out accuracy\)\. Red edges depict defeaters, gray supporters\. Edge numbers are attestation counts\. Accuracy weighting \(whole\-number for clarity, fractional in practice\) redistributes mass toward steps attested by stronger models\.Step 5:The highest supported answer and its highest supported justification is extracted\. The output preserves alternative reasoning paths and step\-level support that a single LLM’s chain\-of\-thought lacks\.Each reasoning trace is converted into a fine\-grained causal DAG using a small LLM as a structured extractor\. To keep the task tractable, the LLM emits nodes together with*pre\-grouped bundles*—each bundle is a list of node ids that must hold together to support a target—so the logical structure is encoded by the premise grouping in bundles\.
Each node carries an id, a node type, a short textual description, and bundle lists: a*support*list, whose bundles each independently justify the node, and a*defeats*list, whose bundles each independently suppress it\. Premises within a bundle are conjunctive\. Bundles within a list are disjunctive\.
Logical Reconstruction\.An AND/OR DAGG=\(V,𝒮\)G=\(V,\\mathcal\{S\}\)is reconstructed in a single deterministic pass: every support bundle becomes a positive support set, every defeater bundle a negated one, and each node’s gate is assigned by cardinality of its positive sets—Atomicif none,Andif one,Orif multiple\. Formally,𝒮\(v\)=\{\(S1,δ1\),\(S2,δ2\),…\}\\mathcal\{S\}\(v\)=\\\{\(S\_\{1\},\\delta\_\{1\}\),\(S\_\{2\},\\delta\_\{2\}\),\\dots\\\}, where each pair\(Sk,δk\)\(S\_\{k\},\\delta\_\{k\}\)consists of a predecessor setSk⊆VS\_\{k\}\\subseteq Vjointly supportingvvand a polarity flagδk∈\{0,1\}\\delta\_\{k\}\\in\\\{0,1\\\}indicating support \(δk=0\\delta\_\{k\}=0\) or defeasibility \(δk=1\\delta\_\{k\}=1\)\.
We verify that all node types are permitted, bundle members reference valid nodes, and the graph is acyclic\. Traces failing any check are discarded so malformed outputs do not pollute aggregation\.
DAG Structure\.Inspired byLeeet al\.\([2025](https://arxiv.org/html/2607.27783#bib.bib54)\), we use four labels:*Planning*\(wherein the LLM decides what to do next\),*Fact*\(factual statements from the input context\),*Reasoning*\(intermediate inference steps linking facts to an outcome\), and*Conclusion*\(the final determination\)\. This structure emphasizes procedural reasoning and is well suited to traces with meta\-reasoning about problem\-solving\. A final*Answer*node is added after the Conclusion, exactly matching the predicted label\. These DAGs provide the inputs to node merging\.
### 3\.3Node Merging
Our node\-merging strategy filters same node\-type candidate pairs by embedding similarity before LLM verification, balancing compute against merge precision\. Neither pure approach suffices\. An LLM judge, given each candidate pair’s text, labels, and local context yields accurate equivalence decisions but scales asO\(n2\)O\(n^\{2\}\)in the node count\. Embedding similarity, computed as cosine distance between dense vectors of node descriptions,
sim\(u,v\)=𝐞u⋅𝐞v‖𝐞u‖‖𝐞v‖\\text\{sim\}\(u,v\)=\\frac\{\\mathbf\{e\}\_\{u\}\\cdot\\mathbf\{e\}\_\{v\}\}\{\\\|\\mathbf\{e\}\_\{u\}\\\|\\\|\\mathbf\{e\}\_\{v\}\\\|\}\(1\)is cheaper but produces false positives, often merging surface\-similar yet substantively distinct nodes\.
Our hybrid method\.We combine the two: a permissive embedding thresholdτ′=0\.55\\tau^\{\\prime\}=0\.55prunes the search space, and surviving candidates are verified by the LLM judge\. To preserve logical operators, we merge nodes while keeping bundles intact\.
Bundle\-Preserving Merge Operator\.Once a clusteringπ:V→𝒞\\pi:V\\to\\mathcal\{C\}over all source\-trace nodes is fixed, we construct the merged ensemble DAGG∗=\(V∗,𝒮∗\)G^\{\*\}=\(V^\{\*\},\\mathcal\{S\}^\{\*\}\)by mapping each bundle throughπ\\piand aggregating identical bundles:
𝒮∗\(c\)=⋃v∈π−1\(c\)⋃\(S,δ\)∈𝒮\(v\)\{\(π\(S\),δ\)\}\\mathcal\{S\}^\{\*\}\(c\)=\\bigcup\_\{v\\in\\pi^\{\-1\}\(c\)\}\\bigcup\_\{\(S,\\delta\)\\in\\mathcal\{S\}\(v\)\}\\left\\\{\\left\(\\pi\(S\),\\delta\\right\)\\right\\\}\(2\)whereπ\(S\)=\{π\(u\):u∈S\}\\pi\(S\)=\\\{\\pi\(u\):u\\in S\\\}\. Two bundles match across traces iff they reduce to the same cluster\-id set underπ\\pi\. Identical bundles are combined, recording how many traces attested the same justification, while distinct bundles become separate support sets, recovering the OR\-of\-ANDs structure naturally\. Gate types are then re\-derived, and cycles introduced by clustering are resolved by removing the weakest support set along the cycle\. Next, the resulting ensemble DAG is assigned node weights, based on cross\-trace consensus\.
### 3\.4Weighted Ensemble of DAGs
We weight each merged node to amplify reasoning trajectories attested across multiple traces and suppress hallucinated ones, treating each trace as casting a*signed vote*on every merged node\.
LetSupp\(v\),Def\(v\)⊆M\\text\{Supp\}\(v\),\\text\{Def\}\(v\)\\subseteq Mdenote the source traces that contributed supporter and defeater bundles to merged nodevv, whereMMis the full set of traces, and letwtw\_\{t\}be the per\-trace weight\. We define the positive and negative attestation masses
α\+\(v\)=∑t∈Supp\(v\)wt∑t∈Mwt,α−\(v\)=∑t∈Def\(v\)wt∑t∈Mwt\.\\alpha^\{\+\}\(v\)\\;=\\;\\frac\{\\sum\_\{t\\in\\text\{Supp\}\(v\)\}w\_\{t\}\}\{\\sum\_\{t\\in M\}w\_\{t\}\},\\qquad\\alpha^\{\-\}\(v\)\\;=\\;\\frac\{\\sum\_\{t\\in\\text\{Def\}\(v\)\}w\_\{t\}\}\{\\sum\_\{t\\in M\}w\_\{t\}\}\.\(3\)so thatα\+\(v\),α−\(v\)∈\[0,1\]\\alpha^\{\+\}\(v\),\\alpha^\{\-\}\(v\)\\in\[0,1\]\. Traces that do not mentionvvcontribute to neither mass but inflate the denominator, thus reducing confidence\.
Node scoresW\(v\)∈\[0,1\]W\(v\)\\in\[0,1\]are computed topologically \(predecessors before successors\) and combine the predecessor support with the node’s own support\. The conjunctive strength of any bundleSSis the weakest\-link score over its members,
ϕ\(S\)=minu∈SW\(u\),\\phi\(S\)\\;=\\;\\min\_\{u\\in S\}W\(u\),\(4\)so a conjunction is as strong as its weakest member\.
The positive and negative aggregate signals are then assembled by gate\-aware combination over the disjunctive supporting and defeating bundles:
Φ\+\(v\)=\{1g\(v\)=Atomicϕ\(S\)g\(v\)=And,𝒮\+\(v\)=\{S\}maxS∈𝒮\+\(v\)ϕ\(S\)g\(v\)=Or\\Phi^\{\+\}\(v\)=\\begin\{cases\}1&g\(v\)=\\textsc\{Atomic\}\\\\ \\phi\(S\)&g\(v\)=\\textsc\{And\},\\,\\mathcal\{S\}^\{\+\}\(v\)=\\\{S\\\}\\\\ \\max\_\{S\\in\\mathcal\{S\}^\{\+\}\(v\)\}\\phi\(S\)&g\(v\)=\\textsc\{Or\}\\end\{cases\}\(5\)Φ−\(v\)=\{0𝒮−\(v\)=∅ϕ\(S\)𝒮−\(v\)=\{S\}maxS∈𝒮−\(v\)ϕ\(S\)\|𝒮−\(v\)\|\>1\\Phi^\{\-\}\(v\)=\\begin\{cases\}0&\\mathcal\{S\}^\{\-\}\(v\)=\\emptyset\\\\ \\phi\(S\)&\\mathcal\{S\}^\{\-\}\(v\)=\\\{S\\\}\\\\ \\max\_\{S\\in\\mathcal\{S\}^\{\-\}\(v\)\}\\phi\(S\)&\|\\mathcal\{S\}^\{\-\}\(v\)\|\>1\\end\{cases\}\(6\)
The final node score combines these with the attestation masses:
W\(v\)=max\(0,α\+\(v\)⋅Φ\+\(v\)−α−\(v\)⋅Φ−\(v\)\)\.W\(v\)\\;=\\;\\max\\\!\\Big\(0,\\;\\;\\alpha^\{\+\}\(v\)\\cdot\\Phi^\{\+\}\(v\)\\;\-\\;\\alpha^\{\-\}\(v\)\\cdot\\Phi^\{\-\}\(v\)\\Big\)\.\(7\)
Simple Weighting\.The baseline setswt=1w\_\{t\}=1for all traces, treating consensus as a raw vote count\.
Accuracy\-Weighting\.When per\-model accuracies are available from a held\-out test set, we replace uniform votes with model\-quality priors:wt=Waccm\(t\)w\_\{t\}=W\_\{\\text\{acc\}\}^\{m\(t\)\}, where∑mWaccm=1\\sum\_\{m\}W\_\{\\text\{acc\}\}^\{m\}=1\. Stronger models carry larger votes in bothα\+\(v\)\\alpha^\{\+\}\(v\)andα−\(v\)\\alpha^\{\-\}\(v\)\. With per\-node weights that quantify the support to each node computed, the ensemble is parsed to select the highest\-weighted conclusion and unfold its highest\-weighted supporting reasoning, as described next\.
### 3\.5Finding the Consensus
The consensus answer is the candidate conclusion whose answer node carries the highest weight:
c⋆=argmaxcW\(c\)\.c^\{\\star\}=\\arg\\max\_\{c\}W\(c\)\.\(8\)Ties are broken by the number of attesting traces\.
Consensus Subgraph\.Givenc⋆c^\{\\star\}, we recursively unfold a proof subgraph𝒫\(c⋆\)\\mathcal\{P\}\(c^\{\\star\}\)rooted atc⋆c^\{\\star\}by picking at each node the support set maximizingϕ\(S\)\\phi\(S\):
- •g\(v\)=Atomicg\(v\)=\\textsc\{Atomic\}: includevvand terminate\.
- •g\(v\)=Andg\(v\)=\\textsc\{And\}with unique support setSS: include all members and expand each\.
- •Ifg\(v\)=Org\(v\)=\\textsc\{Or\}with support sets\{S1,S2,…\}\\\{S\_\{1\},S\_\{2\},\\dots\\\}: selectS⋆=argmaxSϕ\(S\)S^\{\\star\}=\\arg\\max\_\{S\}\\phi\(S\), include all members, and recursively expand each\.
The resulting𝒫\(c⋆\)\\mathcal\{P\}\(c^\{\\star\}\)is the strongest supported complete justification tree forc⋆c^\{\\star\}, with every conjunct required to support every intermediate step preserved\. We report it as theconsensus reasoning\.
To surface alternative reasoning chains, we enumerate the top\-kksubgraphs for each candidate conclusion by varying the support set at eachOrnode, scoring each subgraph by its total node weight:
Ψ\(𝒫\)=∑v∈𝒫W\(v\)\.\\Psi\(\\mathcal\{P\}\)=\\sum\_\{v\\in\\mathcal\{P\}\}W\(v\)\.\(9\)Alongside the consensus, we return for each conclusion its top\-kkjustifications and its*net support*\(total node weight across the union of its proof subgraphs\) exposing alternative reasoning chains and relative standing of competing conclusions\.
### 3\.6Summary
Together, the five stages turn unstructured reasoning from heterogeneous models into a single auditable artifact\. Bundle\-preserving merging keeps each trace’s logical structure intact, attestation\-weighted scoring suppresses hallucinated steps without discarding them, and thus we expose the consensus conclusion, consensus reasoning and the alternatives that remain in contention\. The next section evaluates whether these design choices yield measurable accuracy, explainability and diversity gains over majority\-vote baselines\.
## 4Experimental Setup
Datasets\.We evaluate on six benchmarks spanning four regimes: statutory interpretation, graduate\-level science, narrative multi\-hop reasoning, and first\-order logic\. For each, we pick 50 examples for a held\-out split \(used to estimate per\-model accuracy priors\) and an evaluation split \(222, 148, 200, 206, 200, 153 samples\), with a fixed random seed \(42\) for reproducibility\. The entire pipeline runs on a single NVIDIA H200 GPU\.
\(1\) SARAHolzenbergeret al\.\([2020](https://arxiv.org/html/2607.27783#bib.bib39)\)is a 272 question statutory reasoning benchmark over US federal tax law\. It pairs a statute excerpt, a case scenario, and a hypothesis judgedEntailedorContradicted\.
\(2\) GPQA DiamondReinet al\.\([2023](https://arxiv.org/html/2607.27783#bib.bib56)\)is a 198 question graduate\-level multiple\-choice benchmark in biology, chemistry, and physics, testing generalization to high\-difficulty scientific reasoning\.
\(3–5\) MuSRSpragueet al\.\([2024](https://arxiv.org/html/2607.27783#bib.bib72)\)is a multi\-step soft reasoning benchmark over synthetic narratives, with three domains we evaluate separately:*Murder Mysteries*\(MuSR\-MM\-250\), inferring suspect, motive, and opportunity from narrative cues;*Object Placements*\(MuSR\-OP\-256\), tracking belief states across a sequence of actions;*Team Allocation*\(MuSR\-TA\-250\), assigning agents to tasks under capability, preference constraints\.
\(6\) FOLIOHanet al\.\([2024](https://arxiv.org/html/2607.27783#bib.bib71)\)is a 203 sample first\-order logic entailment benchmark with natural\-language premises, conclusions and three\-way labels with an explicit “insufficient evidence” option\.
Table 1:Accuracy of our framework against matched\-budget baselines across six benchmarks\.\(a\)*Heterogeneous ensemble*over four models, comparing our*Simple*\(wt=1w\_\{t\}=1\) and*Accuracy\-based*\(wtw\_\{t\}= source model’s held\-out accuracy\) weighting schemes to a 4\-model majority\-vote baseline and the four single\-model CoT accuracies\.\(b, c\)*Homogeneous ensemble*applied to a single weak model \(Llama 3\.2 3B\) and a single strong model \(Qwen 2\.5 32B\), each compared to self\-consistency and single\-chain CoT\. The consensus answer is the answer node that has the highest aggregated weight\. Values are mean±\\pmSD over 5 seeds\.Bold/underline= best / second\-best per column in each block\.Takeaway:on mixed models, our accuracy\-weighted framework exceeds the majority\-vote baseline on every dataset; on a single model, our framework exceeds single\-chain CoT and matches or exceeds self\-consistency\.Models\.We sample CoT from four small instruction\-tuned LLMs spanning distinct families and pretraining mixtures within a fixed parameter budget \(≤\\leq12B\): Qwen 2\.5 7B InstructQwenet al\.\([2025](https://arxiv.org/html/2607.27783#bib.bib57)\), Gemma 3 12B InstructTeamet al\.\([2025](https://arxiv.org/html/2607.27783#bib.bib58)\), Llama 3\.2 3B InstructGrattafioriet al\.\([2024](https://arxiv.org/html/2607.27783#bib.bib27)\), and Mistral 7B Instruct v0\.3Jianget al\.\([2023](https://arxiv.org/html/2607.27783#bib.bib59)\), each at temperature 0\.6\. For DAG extraction and LLM\-based node merging we use Qwen 2\.5 32B InstructQwenet al\.\([2025](https://arxiv.org/html/2607.27783#bib.bib57)\)as a stronger structured extractor at temperature 0\.3\. Node similarity uses all\-mpnet\-base\-v2Reimers and Gurevych \([2019](https://arxiv.org/html/2607.27783#bib.bib60)\), and the LLM judge in analysis is Qwen 3 32BYanget al\.\([2025a](https://arxiv.org/html/2607.27783#bib.bib29)\)\. All models are open\-weight, for reproducibility without proprietary access\.
Baselines\.We compare against majority vote over the same four\-model pool at a fixed 20\-trace budget, isolating the contribution of structural aggregation: any difference reflects how traces are combined, not the underlying reasoners or sampling budget\.
Evaluation\.We evaluate along three axes\.\(1\) Accuracy:The final answer parsed from the ensemble graph, compared against a majority\-vote baseline for both many\-model and single\-model ensembles\.\(2\) Explainability:The framework returns the most\-supported subgraphs, the support level for each conclusion, and the alternative conclusions reached\. We additionally quantify if ensemble weights track reasoning quality using Qwen 3 32B as an LLM judge\.\(3\) Diversity:We qualitatively check if alternative subgraphs present different solution perspectives\. To motivate the use of multiple model families for obtaining diverse reasoning chains, we conduct an experiment to see if multi\-model pools give more semantically diverse correct\-answer reasoning trajectories than single\-model pools \(Appendix[C](https://arxiv.org/html/2607.27783#A3)\)\.
## 5Experiments and Results
This section evaluates our framework\. Sections[5\.1](https://arxiv.org/html/2607.27783#S5.SS1)and[5\.2](https://arxiv.org/html/2607.27783#S5.SS2)report accuracy for multi\- and single\-model ensembles against majority\-vote baselines; Section[5\.3](https://arxiv.org/html/2607.27783#S5.SS3)quantifies if ensemble weights track reasoning quality via LLM\-judge rankings; Section[5\.4](https://arxiv.org/html/2607.27783#S5.SS4)illustrates graphs; Section[5\.5](https://arxiv.org/html/2607.27783#S5.SS5)has human evaluation\.
### 5\.1Heterogeneous Ensemble Accuracies
Table[1](https://arxiv.org/html/2607.27783#S4.T1)\(a\) reports our two weighting schemes against a matched\-budget 4\-model majority\-vote baseline \(20 chains\) and single\-model CoT\.
Different models lead on different datasets\.\.Strongest\-to\-weakest gaps within a dataset reach 25\.2%, confirming no single model is uniformly best and motivating heterogeneous ensembling\.
Accuracy\-weighted ensembling beats majority voting on every dataset\.With the largest gain of 3\.1% on MuSR\-MM \(59\.6% vs\. 56\.5%\)\. Simple weighting also beats majority vote on three datasets, indicating gains stem from structural aggregation rather than the accuracy prior alone, though the prior delivers the additional boost for consistent wins\.
The strongest single model often wins outright\.As expected for ensembling, averaging in weaker reasoners introduces noise structural aggregation may not filter\. The ensemble’s value here is inspectability that the single best model can not provide\.
### 5\.2Homogeneous Ensemble Accuracies
To isolate structural aggregation from cross\-model diversity, we apply our framework to single\-model traces using a weak reasoner and a stronger reasoner, 20 chains each\. Table[1](https://arxiv.org/html/2607.27783#S4.T1)\(b,c\) compares against self\-consistency and single\-chain CoT\.
Weak model: structural aggregation recovers gains\.On Llama 3\.2 3B, our method exceeds single\-chain CoT \(max\+8\.5%\+8\.5\\%on SARA and FOLIO\), and matches / exceeds self\-consistency \(max\+3\.0%\+3\.0\\%on SARA\)\.222Llama 3\.2 3B exhibits near\-deterministic majority voting on MuSR\-MM, making accuracy invariant across five seeds\.When individual traces are unreliable, structural aggregation extracts a more trustworthy verdict from cross\-trace consensus and exposes the supporting graph that voting discards\.
Table 2:Two complementary measures of consensus\-subgraph quality, judged by Qwen3\-32B\.\(a\)Per\-question Spearman correlation between judge\-ranked and weight\-ranked sampled subgraphs \(n=5n=5per question, 5 shuffled orderings averaged\)\.\(b\)Per\-question win rate of the consensus subgraph against a length\-matched random alternative trace leading to the majority\-vote answer \(ties excluded, ordering swapped\)\. Values are mean±\\pmSD across 5 seeds\.Boldmarks the higher value per column in each block\.Takeaway:ensemble weights correlate positively with judge\-rated reasoning quality, and the consensus subgraph is preferred over a random alternative\. FOLIO is the only exception under simple weighting, and accuracy weighting recovers preference to 52\.6%\.Strong model: ensemble matches self\-consistency and exceeds single\-chain CoT\.On Qwen 2\.5 32B, our framework is within±0\.6%\\pm 0\.6\\%of self\-consistency on all six benchmarks and matches or exceeds single\-chain CoT on every dataset, \(max\+3\.9%\+3\.9\\%on FOLIO\)\. A strong reasoner converges to the right answer in most chains, so any aggregation recovers it\. The framework’s contribution at this scale is inspectability, that self\-consistency cannot provide\.
Benefits hold without cross\-model diversity\.Our framework recovers self\-consistency’s accuracy while additionally producing a structured, weighted, auditable graph—demonstrating that its value is not contingent on heterogeneous pools\.
### 5\.3Evaluation of Consensus Graph Quality
Ensemble weights correlate positively with LLM\-judge rankings of reasoning quality across all datasets and weighting schemes, indicating that the weight signal reflects genuine reasoning quality\.
Expt A: Subgraph ranking\.For each question, we sampled 5 subgraphs from the ensemble, linearized each into a step\-by\-step chain, emitting nodes in topological order, and asked Qwen3\-32B to rank them by reasoning quality \(5 shuffled orderings per question, averaged to cancel presentation\-order bias\)\. Table[2](https://arxiv.org/html/2607.27783#S5.T2)\(a\) reports per\-question Spearman correlation between judge and weight rankings\. Correlation is moderate: weights and LLM\-judge measure overlapping but not identical notions of justification quality, and the judge is imperfect\.
Expt B: Pairwise win rate\.We sampled a subgraph leading to the majority\-vote answer, length\-matched it to the consensus subgraph to remove length bias, linearized both, and asked Qwen3\-32B which chain provided better reasoning \(each pair evaluated twice with swapped A/B order\)\. Table[2](https://arxiv.org/html/2607.27783#S5.T2)\(b\) reports the per\-question win rate of the consensus subgraph\. The consensus trace is preferred over 50% of the time, showing that the weighting procedure captures a meaningful signal for reasoning quality beyond mere plausible justification\.
### 5\.4Qualitative Analysis
To support qualitative explainability evaluation, our framework returns three components per query: the top\-kkmost\-supported reasoning subgraphs, the support level for each conclusion, and the alternative conclusions reached by the models\. Examples are provided in Appendix[B](https://arxiv.org/html/2607.27783#A2), where each shows a consensus graph alongside an alternative subgraph \(sampled from the same ensemble DAG\) and a random single\-trace CoT graph that reached the correct answer\. Across these examples, the consensus graph consistently provides the most detailed and well\-supported reasoning, the alternative graph offers a different justification for the same conclusion, and the single CoT is at times less thorough\.
### 5\.5Human Evaluation
We ran a small human evaluation \(Appendix[D](https://arxiv.org/html/2607.27783#A4)\) to complement our LLM\-as\-a\-Judge results\. Ten annotators were each shown 15 correct\-answer graph triplets \(consensus subgraph, alternative subgraph, single\-CoT graph\) randomly sampled from MuSR, and asked to pick the graph that best justified the answer\. The consensus, single\-CoT, alternative, and*none*options were chosen 41\.3%, 33\.3%, 20\.7%, and 4\.7% of the time, respectively, supporting our LLM\-judge preferring the consensus reasoning over alternatives\. The non\-trivial alternative\-graph share \(20\.7%\) also suggests that surfacing alternative reasoning chains is useful in its own right\.
## 6Discussion
Our results highlight three properties of structural ensembling\. Onaccuracy, edges attested by only one trace contribute little to the consensus score, so a single model’s spurious thought cannot dominate the aggregate\. This is what helps the framework beat majority voting on every dataset and recover gains over self\-consistency on the single\-weak\-model\. Oninterpretability, the output is a weighted graph rather than a free\-form rationale, so users can view the consensus reasoning, different possible conclusions, and different justifications from diverse perspectives that support any conclusion—a property no answer\-level baseline provides\. Ondiversity, preserving competing conclusions makes query produce multiple supported reasoning paths, and heterogeneous ensembles widen this due to distinct pretraining distributions\. This can also be observed qualitatively\.
These properties suggest a natural extension\. In open\-ended settings such as argument generation and creative writing, the goal is not to converge on a single answer but to surface a diverse set of well\-supported alternatives\. Our framework’s existing outputs—competing conclusions with their attested support graphs—map directly onto this use case: rather than report the highest\-weighted answer alone, the system can present the top\-kkalternative conclusions and let the user choose among reasoning paths grounded in different models’ perspectives\. This shifts the role of consensus from selection to curation, and we view it as the most promising direction for follow\-up work\.
## 7Conclusion
We introduced a framework for ensembling the reasoning structure—not just the answers—of multiple LLMs, by extracting per\-trace DAGs, merging them with a logic\-preserving operator, and weighting each node by cross\-trace attestation\. Across six benchmarks, our accuracy\-weighted ensemble outperforms matched\-budget majority voting on every dataset; applied to a single model, it matches self\-consistency\. Ensemble weights track LLM\-judge rankings of reasoning quality, and consensus subgraphs are consistently preferred over alternatives leading to the majority\-vote answer\. The resulting weightedReasoning Consensusgraph is auditable at the level of individual reasoning steps, exposes alternative conclusions that competing traces reached, alternative justifications for each conclusion, and naturally extends to open\-ended settings where curating diverse perspectives matters\.
## 8Limitations
Our pipeline relies on an extractor model faithfully converting free\-form chains of thought into typed, bundle\-structured DAGs\. Extraction errors would propagate through aggregation\. To mitigate this, we use Qwen 2\.5 32B, a substantially stronger model than the trace generators, at a low decoding temperature \(0\.3\) for deterministic structured output, and we discard any trace whose extraction fails validation \(acyclicity, valid node types, valid bundle references\)\. Cross\-trace attestation provides a second filter: a hallucinated step appearing in a single trace receives low weight, so even when individual extractions are imperfect, the aggregated graph remains dominated by steps that multiple independent traces converged on\.
We use Qwen 3 32B as a judge, and LLM judges exhibit known biases around position, length, and verbosity\. We control for the most common ones explicitly: every pairwise comparison is evaluated twice with swapped A/B order, every ranking experiment is repeated five times per question with shuffled presentations and averaged, and alternative subgraphs are length\-matched to the consensus subgraph to remove length confounds\. To further verify the judge, we ran a human evaluation on randomly sampled graph triplets on the MuSR dataset\.
Each query requires sampling 20 chains of thought, extracting a DAG per chain, computing pairwise embeddings, and running the LLM judge for node merging — more compute than single\-chain CoT or self\-consistency\. We deliberately target the high\-stakes setting \(legal, scientific\) where the cost of an unsupported answer outweighs the compute cost of our framework, and where users cannot rely on closed APIs\. To keep the framework optimized, we restrict all trace\-generator models to≤12\\leq 12B parameters, use a permissive embedding threshold \(τ′=0\.55\\tau^\{\\prime\}=0\.55\) to prune most node pairs before the LLM judge sees them, and run the entire pipeline on a single NVIDIA H200 GPU\.
## 9AI Involvement Disclosure
AI assistance was limited to language polishing for grammar and readability\. The conception, methodology, experiments, analyses, and interpretations were conducted fully by the authors\.
## References
- M\. Besta, N\. Blach, A\. Kubicek, R\. Gerstenberger, M\. Podstawski, L\. Gianinazzi, J\. Gajda, T\. Lehmann, H\. Niewiadomski, P\. Nyczyk, and T\. Hoefler \(2024\)Graph of thoughts: solving elaborate problems with large language models\.Proceedings of the AAAI Conference on Artificial Intelligence38\(16\),pp\. 17682–17690\.External Links:ISSN 2159\-5399,[Link](http://dx.doi.org/10.1609/aaai.v38i16.29720),[Document](https://dx.doi.org/10.1609/aaai.v38i16.29720)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p2.1),[§2](https://arxiv.org/html/2607.27783#S2.p1.1)\.
- ReConcile: round\-table conference improves reasoning via consensus among diverse llms\.External Links:2309\.13007,[Link](https://arxiv.org/abs/2309.13007)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p2.1),[§2](https://arxiv.org/html/2607.27783#S2.p2.1)\.
- M\. Dahl, V\. Magesh, M\. Suzgun, and D\. E\. Ho \(2024\)Large legal fictions: profiling legal hallucinations in large language models\.Journal of Legal Analysis16\(1\),pp\. 64–93\.External Links:[Document](https://dx.doi.org/10.1093/jla/laae003),ISSN 1946\-5319,[Link](http://dx.doi.org/10.1093/jla/laae003)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p6.1)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi, X\. Zhang, X\. Yu, Y\. Wu, Z\. F\. Wu, Z\. Gou, Z\. Shao, Z\. Li, Z\. Gao, A\. Liu, B\. Xue, B\. Wang, B\. Wu, B\. Feng, C\. Lu, C\. Zhao, C\. Deng, C\. Ruan, D\. Dai, D\. Chen, D\. Ji, E\. Li, F\. Lin, F\. Dai, F\. Luo, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Xu, H\. Ding, H\. Gao, H\. Qu, H\. Li, J\. Guo, J\. Li, J\. Chen, J\. Yuan, J\. Tu, J\. Qiu, J\. Li, J\. L\. Cai, J\. Ni, J\. Liang, J\. Chen, K\. Dong, K\. Hu, K\. You, K\. Gao, K\. Guan, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Zhao, L\. Wang, L\. Zhang, L\. Xu, L\. Xia, M\. Zhang, M\. Zhang, M\. Tang, M\. Zhou, M\. Li, M\. Wang, M\. Li, N\. Tian, P\. Huang, P\. Zhang, Q\. Wang, Q\. Chen, Q\. Du, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. J\. Chen, R\. L\. Jin, R\. Chen, S\. Lu, S\. Zhou, S\. Chen, S\. Ye, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. S\. Li, S\. Zhou, S\. Wu, T\. Yun, T\. Pei, T\. Sun, T\. Wang, W\. Zeng, W\. Liu, W\. Liang, W\. Gao, W\. Yu, W\. Zhang, W\. L\. Xiao, W\. An, X\. Liu, X\. Wang, X\. Chen, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yang, X\. Li, X\. Su, X\. Lin, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Sun, X\. Wang, X\. Song, X\. Zhou, X\. Wang, X\. Shan, Y\. K\. Li, Y\. Q\. Wang, Y\. X\. Wei, Y\. Zhang, Y\. Xu, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Yu, Y\. Zhang, Y\. Shi, Y\. Xiong, Y\. He, Y\. Piao, Y\. Wang, Y\. Tan, Y\. Ma, Y\. Liu, Y\. Guo, Y\. Ou, Y\. Wang, Y\. Gong, Y\. Zou, Y\. He, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. X\. Zhu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Y\. Tang, Y\. Zha, Y\. Yan, Z\. Z\. Ren, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Ma, Z\. Yan, Z\. Wu, Z\. Gu, Z\. Zhu, Z\. Liu, Z\. Li, Z\. Xie, Z\. Song, Z\. Pan, Z\. Huang, Z\. Xu, Z\. Zhang, and Z\. Zhang \(2025\)DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.External Links:ISSN 1476\-4687,[Link](http://dx.doi.org/10.1038/s41586-025-09422-z),[Document](https://dx.doi.org/10.1038/s41586-025-09422-z)Cited by:[§2](https://arxiv.org/html/2607.27783#S2.p1.1)\.
- S\. Han, H\. Schoelkopf, Y\. Zhao, Z\. Qi, M\. Riddell, W\. Zhou, J\. Coady, D\. Peng, Y\. Qiao, L\. Benson, L\. Sun, A\. Wardle\-Solano, H\. Szabo, E\. Zubova, M\. Burtell, J\. Fan, Y\. Liu, B\. Wong, M\. Sailor, A\. Ni, L\. Nan, J\. Kasai, T\. Yu, R\. Zhang, A\. R\. Fabbri, W\. Kryscinski, S\. Yavuz, Y\. Liu, X\. V\. Lin, S\. Joty, Y\. Zhou, C\. Xiong, R\. Ying, A\. Cohan, and D\. Radev \(2024\)FOLIO: natural language reasoning with first\-order logic\.External Links:2209\.00840,[Link](https://arxiv.org/abs/2209.00840)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p5.1)\.
- N\. Holzenberger, A\. Blair\-Stanek, and B\. V\. Durme \(2020\)A dataset for statutory reasoning in tax law entailment and question answering\.External Links:2005\.05257,[Link](https://arxiv.org/abs/2005.05257)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p2.1)\.
- J\. Hu, Y\. Wang, S\. Zhang, K\. Zhou, G\. Chen, Y\. Hu, B\. Xiao, and M\. Tan \(2025\)Efficient dynamic ensembling for multiple llm experts\.External Links:2412\.07448,[Link](https://arxiv.org/abs/2412.07448)Cited by:[§2](https://arxiv.org/html/2607.27783#S2.p2.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. Sayed \(2023\)Mistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p6.1)\.
- Y\. Jin, H\. Kim, K\. Kim, C\. Lee, and S\. Shin \(2026\)ARGORA: orchestrated argumentation for causally grounded llm reasoning and decision making\.External Links:2601\.21533,[Link](https://arxiv.org/abs/2601.21533)Cited by:[§2](https://arxiv.org/html/2607.27783#S2.p3.1)\.
- J\. Lee, S\. Mukherjee, D\. Hakkani\-Tur, and J\. Hockenmaier \(2025\)ReasoningFlow: semantic structure of complex reasoning traces\.External Links:2506\.02532,[Link](https://arxiv.org/abs/2506.02532)Cited by:[§3\.2](https://arxiv.org/html/2607.27783#S3.SS2.p5.1)\.
- OpenAI, :, A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney, A\. Iftimie, A\. Karpenko, A\. T\. Passos, A\. Neitz, A\. Prokofiev, A\. Wei, A\. Tam, A\. Bennett, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Duberstein, A\. Kondrich, A\. Mishchenko, A\. Applebaum, A\. Jiang, A\. Nair, B\. Zoph, B\. Ghorbani, B\. Rossen, B\. Sokolowsky, B\. Barak, B\. McGrew, B\. Minaiev, B\. Hao, B\. Baker, B\. Houghton, B\. McKinzie, B\. Eastman, C\. Lugaresi, C\. Bassin, C\. Hudson, C\. M\. Li, C\. de Bourcy, C\. Voss, C\. Shen, C\. Zhang, C\. Koch, C\. Orsinger, C\. Hesse, C\. Fischer, C\. Chan, D\. Roberts, D\. Kappler, D\. Levy, D\. Selsam, D\. Dohan, D\. Farhi, D\. Mely, D\. Robinson, D\. Tsipras, D\. Li, D\. Oprica, E\. Freeman, E\. Zhang, E\. Wong, E\. Proehl, E\. Cheung, E\. Mitchell, E\. Wallace, E\. Ritter, E\. Mays, F\. Wang, F\. P\. Such, F\. Raso, F\. Leoni, F\. Tsimpourlas, F\. Song, F\. von Lohmann, F\. Sulit, G\. Salmon, G\. Parascandolo, G\. Chabot, G\. Zhao, G\. Brockman, G\. Leclerc, H\. Salman, H\. Bao, H\. Sheng, H\. Andrin, H\. Bagherinezhad, H\. Ren, H\. Lightman, H\. W\. Chung, I\. Kivlichan, I\. O’Connell, I\. Osband, I\. C\. Gilaberte, I\. Akkaya, I\. Kostrikov, I\. Sutskever, I\. Kofman, J\. Pachocki, J\. Lennon, J\. Wei, J\. Harb, J\. Twore, J\. Feng, J\. Yu, J\. Weng, J\. Tang, J\. Yu, J\. Q\. Candela, J\. Palermo, J\. Parish, J\. Heidecke, J\. Hallman, J\. Rizzo, J\. Gordon, J\. Uesato, J\. Ward, J\. Huizinga, J\. Wang, K\. Chen, K\. Xiao, K\. Singhal, K\. Nguyen, K\. Cobbe, K\. Shi, K\. Wood, K\. Rimbach, K\. Gu\-Lemberg, K\. Liu, K\. Lu, K\. Stone, K\. Yu, L\. Ahmad, L\. Yang, L\. Liu, L\. Maksin, L\. Ho, L\. Fedus, L\. Weng, L\. Li, L\. McCallum, L\. Held, L\. Kuhn, L\. Kondraciuk, L\. Kaiser, L\. Metz, M\. Boyd, M\. Trebacz, M\. Joglekar, M\. Chen, M\. Tintor, M\. Meyer, M\. Jones, M\. Kaufer, M\. Schwarzer, M\. Shah, M\. Yatbaz, M\. Y\. Guan, M\. Xu, M\. Yan, M\. Glaese, M\. Chen, M\. Lampe, M\. Malek, M\. Wang, M\. Fradin, M\. McClay, M\. Pavlov, M\. Wang, M\. Wang, M\. Murati, M\. Bavarian, M\. Rohaninejad, N\. McAleese, N\. Chowdhury, N\. Chowdhury, N\. Ryder, N\. Tezak, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, P\. Chao, P\. Ashbourne, P\. Izmailov, P\. Zhokhov, R\. Dias, R\. Arora, R\. Lin, R\. G\. Lopes, R\. Gaon, R\. Miyara, R\. Leike, R\. Hwang, R\. Garg, R\. Brown, R\. James, R\. Shu, R\. Cheu, R\. Greene, S\. Jain, S\. Altman, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Hernandez, S\. Baker, S\. McKinney, S\. Yan, S\. Zhao, S\. Hu, S\. Santurkar, S\. R\. Chaudhuri, S\. Zhang, S\. Fu, S\. Papay, S\. Lin, S\. Balaji, S\. Sanjeev, S\. Sidor, T\. Broda, A\. Clark, T\. Wang, T\. Gordon, T\. Sanders, T\. Patwardhan, T\. Sottiaux, T\. Degry, T\. Dimson, T\. Zheng, T\. Garipov, T\. Stasi, T\. Bansal, T\. Creech, T\. Peterson, T\. Eloundou, V\. Qi, V\. Kosaraju, V\. Monaco, V\. Pong, V\. Fomenko, W\. Zheng, W\. Zhou, W\. McCabe, W\. Zaremba, Y\. Dubois, Y\. Lu, Y\. Chen, Y\. Cha, Y\. Bai, Y\. He, Y\. Zhang, Y\. Wang, Z\. Shao, and Z\. Li \(2024\)OpenAI o1 system card\.External Links:2412\.16720,[Link](https://arxiv.org/abs/2412.16720)Cited by:[§2](https://arxiv.org/html/2607.27783#S2.p1.1)\.
- S\. Park, X\. Liu, Y\. Gong, and E\. Choi \(2024\)Ensembling large language models with process reward\-guided tree search for better complex reasoning\.External Links:2412\.15797,[Link](https://arxiv.org/abs/2412.15797)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p2.1),[§2](https://arxiv.org/html/2607.27783#S2.p2.1)\.
- Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. Qiu \(2025\)Qwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p6.1)\.
- N\. Reimers and I\. Gurevych \(2019\)Sentence\-bert: sentence embeddings using siamese bert\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing,External Links:[Link](http://arxiv.org/abs/1908.10084)Cited by:[Appendix C](https://arxiv.org/html/2607.27783#A3.p1.1),[§4](https://arxiv.org/html/2607.27783#S4.p6.1)\.
- D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman \(2023\)GPQA: a graduate\-level google\-proof q&a benchmark\.External Links:2311\.12022,[Link](https://arxiv.org/abs/2311.12022)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p3.1)\.
- Z\. Sprague, X\. Ye, K\. Bostrom, S\. Chaudhuri, and G\. Durrett \(2024\)MuSR: testing the limits of chain\-of\-thought with multistep soft reasoning\.External Links:2310\.16049,[Link](https://arxiv.org/abs/2310.16049)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p4.1)\.
- G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. Hussenot \(2025\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p6.1)\.
- J\. Wang, J\. Wang, B\. Athiwaratkun, C\. Zhang, and J\. Zou \(2024\)Mixture\-of\-agents enhances large language model capabilities\.External Links:2406\.04692,[Link](https://arxiv.org/abs/2406.04692)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p2.1),[§2](https://arxiv.org/html/2607.27783#S2.p2.1)\.
- X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. Zhou \(2023\)Self\-consistency improves chain of thought reasoning in language models\.External Links:2203\.11171,[Link](https://arxiv.org/abs/2203.11171)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p2.1),[§1](https://arxiv.org/html/2607.27783#S1.p3.1),[§2](https://arxiv.org/html/2607.27783#S2.p2.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. Le, and D\. Zhou \(2023\)Chain\-of\-thought prompting elicits reasoning in large language models\.External Links:2201\.11903,[Link](https://arxiv.org/abs/2201.11903)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p1.1),[§2](https://arxiv.org/html/2607.27783#S2.p1.1)\.
- Y\. Xu, S\. Ren, and J\. Zhang \(2025\)Collaborative beam search: enhancing LLM reasoning via collective consensus\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 11398–11410\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.574),ISBN 979\-8\-89176\-332\-6,[Link](https://aclanthology.org/2025.emnlp-main.574/)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p2.1),[§2](https://arxiv.org/html/2607.27783#S2.p2.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, and K\. et al\. \(2025a\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4](https://arxiv.org/html/2607.27783#S4.p6.1)\.
- S\. Yang, Y\. Li, W\. Lam, and Y\. Cheng \(2025b\)Multi\-llm collaborative search for complex problem solving\.External Links:2502\.18873,[Link](https://arxiv.org/abs/2502.18873)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p1.1),[§2](https://arxiv.org/html/2607.27783#S2.p2.1)\.
- X\. Yang, X\. Zhang, H\. Hu, and F\. Ji \(2026\)Beyond accuracy: evaluating strategy diversity in llm mathematical reasoning\.External Links:2605\.09292,[Link](https://arxiv.org/abs/2605.09292)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p1.1)\.
- S\. Yao, D\. Yu, J\. Zhao, I\. Shafran, T\. L\. Griffiths, Y\. Cao, and K\. Narasimhan \(2023\)Tree of thoughts: deliberate problem solving with large language models\.External Links:2305\.10601,[Link](https://arxiv.org/abs/2305.10601)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p2.1),[§2](https://arxiv.org/html/2607.27783#S2.p1.1)\.
- X\. Zeng, J\. Lin, Y\. Yan, F\. Guo, L\. Shi, J\. Wu, and D\. Zhou \(2026\)HalluGuard: demystifying data\-driven and reasoning\-driven hallucinations in llms\.External Links:2601\.18753,[Link](https://arxiv.org/abs/2601.18753)Cited by:[§1](https://arxiv.org/html/2607.27783#S1.p1.1)\.
## Appendix APrompts
The prompts used in our pipeline are:
### A\.1Reasoning Trace Extraction
#### A\.1\.1Shared System Prompt
You are a careful reasoner\. For every question, you must:1\. Think step by step inside a single <reasoning\>\.\.\.</reasoning\> block\.2\. Put ONLY the final answer \(no explanation\) inside a <answer\>\.\.\.</answer\> tag\.Do not write anything outside these two tags\. Do not abbreviate your reasoning\.
#### A\.1\.2SARA \(Statutory Reasoning\)
You are reasoning about US federal tax law\.STATUTE:\{statute\}CASE SCENARIO:\{case\}HYPOTHESIS:\{question\}TASK: Decide whether the statute, applied to the case scenario,ENTAILS or CONTRADICTS the hypothesis\. Walk through:\(a\) which statutory provisions apply to the case,\(b\) what facts in the scenario trigger or fail to trigger those provisions,\(c\) what the provisions imply about the hypothesis\.Output exactly:<reasoning\>your step\-by\-step legal analysis</reasoning\><answer\>ENTAILED</answer\> or <answer\>CONTRADICTED</answer\>
#### A\.1\.3GPQA Diamond \(Graduate\-Level Science\)
You are answering a graduate\-level science question\. Exactly one option is correct\.QUESTION:\{question\}OPTIONS:A\. \{option\_a\}B\. \{option\_b\}C\. \{option\_c\}D\. \{option\_d\}TASK: Reason carefully through the underlying science\. Walk through:\(a\) what principles or equations the question is testing,\(b\) how to apply them to the specifics of this problem,\(c\) why one option follows and the others do not\.Output exactly:<reasoning\>your step\-by\-step scientific reasoning</reasoning\><answer\>A</answer\> \(or B, C, or D \-\- a single capital letter only\)
#### A\.1\.4MuSR \(Narrative Multi\-Hop Reasoning\)
\{internallinenumbers\*\}You are reasoning about a narrative scenario\. Read the narrative carefully, then choose the single correct answer\.NARRATIVE:\{narrative\}QUESTION:\{question\}OPTIONS:A\. \{choice\_a\}B\. \{choice\_b\}\.\.\.\(up to G\)TASK: Reason carefully through the evidence in the narrative\. Walk through:\(a\) what facts the narrative establishes,\(b\) how those facts bear on each candidate option,\(c\) why one option is best supported and the others are not\.Output exactly:<reasoning\>your step\-by\-step reasoning over the narrative</reasoning\><answer\>A</answer\> \(or B, C, \.\.\. up to \{last\_letter\} \-\- a single capital letter only\)
#### A\.1\.5FOLIO \(First\-Order Logic Entailment\)
You are reasoning about first\-order logical entailment in natural language\.PREMISES:\{premises\}CONCLUSION:\{conclusion\}\{internallinenumbers\*\}TASK: Decide whether the conclusion is logically entailed by the premises\. Use these labels:TRUE \-\- the conclusion follows necessarily from the premisesFALSE \-\- the negation of the conclusion follows from the premisesUNCERTAIN \-\- neither the conclusion nor its negation followsWalk through:\(a\) what each premise asserts,\(b\) what chain of inferences, if any, the conclusion would require,\(c\) whether the premises are sufficient to establish or refute it\.Output exactly:<reasoning\>your step\-by\-step logical analysis</reasoning\><answer\>TRUE</answer\> or <answer\>FALSE</answer\> or <answer\>UNCERTAIN</answer\>
### A\.2DAG Extraction
#### A\.2\.1Shared System Prompt
You are a structured\-reasoning extractor\. Given a reasoning trace andits final answer, you produce a Directed Acyclic Graph \(DAG\) in JSONthat captures the trace’s logical structure\. You output ONLY a singleJSON object inside <dag\>\.\.\.</dag\> tags\. No prose\. No markdown fences\.
#### A\.2\.2DAG Extraction User Prompt
TASK: Extract a Type\-B reasoning DAG from this trace\.NODE TYPES \(use exactly these strings\):\- "Planning": meta\-reasoning, problem decomposition, deciding what to check\- "Fact": factual statements drawn from the question/statute/scenario\- "Reasoning": intermediate inferential steps linking facts to outcome\- "Conclusion": the final determination
REQUIRED: the DAG must contain exactly ONE node of type "Conclusion", whose\`text\` expresses the final answer below \(paraphrase it however the trace did,but it must be the same verdict\)\. All other nodes must transitively supportthis Conclusion\.BUNDLE STRUCTURE: each node has\- "id": unique short string id \(e\.g\. "p1", "f1", "r1", "c1"\)\- "type": one of the four node types above\- "text": short natural\-language description \(one sentence\)\- "support": list of bundles\. Each bundle is a list of node ids thatJOINTLY \(conjunctive AND\) justify this node\. Differentbundles in the list are alternative \(disjunctive OR\)independent justifications\.\- "defeats": list of bundles whose joint truth would SUPPRESS this node\.Use \[\] if the trace mentions no counter\-considerations\.Do NOT invent defeaters\.RULES:\- Facts may have empty support \(they are atomic premises from the question\)\.\- Planning/Reasoning/Conclusion nodes MUST have \>=1 support bundle\.\- Every id referenced in a bundle must exist as a node in the graph\.\- The graph must be acyclic\.\- Do NOT add information the trace does not contain\.\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-QUESTION CONTEXT:\{question\_blob\}FINAL ANSWER \(must be the Conclusion node’s verdict\):\{final\_answer\}REASONING TRACE:\{reasoning\}\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-Output ONLY this:<dag\>\{"nodes": \[\{"id": "\.\.\.", "type": "Planning\|Fact\|Reasoning\|Conclusion", "text": "\.\.\.","support": \[\["id1","id2"\], \["id3"\]\], "defeats": \[\]\}\]\}</dag\>
### A\.3Node Merging
#### A\.3\.1Merge Judge System Prompt
You are a strict semantic equivalence judge for reasoning\-graph nodes\.Two nodes should merge ONLY IF they express the SAME underlying claim,fact, or inference, regardless of wording\. Lexical overlap is NOT enough\.If they cite different statutes, parties, events, rules, or drawdifferent inferences, they are DIFFERENT\.Answer with a single token: YES or NO\.
#### A\.3\.2Merge Judge User Prompt
Decide whether these two reasoning\-graph nodes should be merged into one\.Both are of type "\{ntype\}" \(already verified to match\)\.QUESTION CONTEXT \(helps interpret references like "the hypothesis"\):\{question\_blob\}NODE A: \{text\_a\}NODE B: \{text\_b\}Should A and B be merged because they assert the SAME underlyingclaim / fact / inference?Reply with exactly one token: YES or NO\.
Figure 2:Three reasoning graphs reaching the same conclusion \(Answer A\) for a MuSR Murder Mysteries question\.Consensus graph\(left\): the highest\-Ψ\\Psiproof subgraph of the consensus answer, aggregating attested reasoning across the heterogeneous ensemble\.Alternative graph\(middle\): a different proof subgraph for the same conclusion, sampled from the same ensemble DAG by varying the support set atOrnodes\.Single CoT graph\(right\): a random per\-trace DAG from one model whose chain reached the same answer\. The consensus graph integrates evidence from multiple traces; the alternative offers a substantively different justification \(eliminating Jose’s guilt vs\. establishing Bella’s guilt\); and the single CoT depends on fewer supporting facts\.
### A\.4LLM\-as\-a\-Judge
#### A\.4\.1Ranking Judge System Prompt
You are a strict evaluator of reasoning quality\. You are given severalcandidate reasoning chains that all argue toward the SAME final answer\.Rank them by the QUALITY of their reasoning and justification: are thesteps well\-grounded in facts, are the inferences logically sound, arethere gaps, leaps, or irrelevant steps\. Do NOT reward a chain forreaching the answer \-\- they all reach the same answer\. A chain can reacha correct answer through weak or wrong reasoning; rank that chain low\.The number of steps does NOT matter; do not prefer a chain merely because it is longer\.
#### A\.4\.2Ranking Judge User Prompt
You are given \{n\} candidate reasoning chains\. They ALL conclude thesame final answer\. Rank them from BEST to WORST reasoning quality\.QUESTION CONTEXT:\{question\_blob\}FINAL ANSWER \(every chain reaches this \-\- it is NOT what you are judging\):\{final\_answer\}\{chains\_block\}Rank all \{n\} chains from best reasoning to worst\. Output ONLY the ranking aschain numbers separated by ’\>’, best first\. For example: 2 \> 1 \> 3Do not output anything else\.
#### A\.4\.3Pairwise Judge System Prompt
You are a strict evaluator of reasoning quality\. You are given twocandidate reasoning chains, each arguing for a final answer to the samequestion\. Decide which chain is the better reasoning and justification:better\-grounded steps, sounder inferences, fewer gaps or irrelevantleaps\.Chain LENGTH IS IRRELEVANT: a longer chain with more steps is NOTautomatically better, and a shorter chain is NOT automatically worse\.Do not reward verbosity\. Judge only the soundness and grounding of thereasoning, not how many steps it has\.Answer with a single token: A or B\.
#### A\.4\.4Pairwise Judge User Prompt
Two candidate reasoning chains each argue for an answer to the samequestion\. Decide which chain is the better reasoning and justification\.QUESTION CONTEXT:\{question\_blob\}\-\-\- CHAIN A \(concludes: \{answer\_a\}\) \-\-\-\{chain\_a\}\-\-\- CHAIN B \(concludes: \{answer\_b\}\) \-\-\-\{chain\_b\}Which chain reasons better \-\- better\-grounded steps, sounder inferences,fewer gaps or irrelevant leaps? The number of steps does NOT matter; do notprefer a chain merely because it is longer\. Reply with exactly one token:A or B\.
## Appendix BVisualization
Warning: This appendix contains examples from MuSR\-Murder Mystery subset describing hypothetical murder\-related scenarios\.
Figure[2](https://arxiv.org/html/2607.27783#A1.F2)depicts an example of consensus, alternative and CoT graphs from MuSR\-MM that reach the same correct answer\. We can observe that the alternative graph solved the murder mystery by eliminating Jose guilty while the consensus graph solves it by asserting Bella’s guilt\. This shows us the diverse perspectives taken to solve the same problem by two different graphs extracted from our ensemble DAG\. The CoT graph is less detailed and does not take into account various facts that the consensus graph does\.
Figure[3](https://arxiv.org/html/2607.27783#A2.F3)also depicts an example of consensus, alternative and CoT graphs from MuSR\-MM that reach the same correct answer\. We can observe that the consensus graph reveals much more intricate reasoning as compared to the alternative or CoT graphs, which is beneficial for users audits\.
Figure[4](https://arxiv.org/html/2607.27783#A2.F4)compares a consensus graph with an alternative graph from the MUSR\-OP dataset that reach different answers\. The alternative graph argues that since tools are usually in a tool shed, they should be searched for in the shed first\. The consensus graph argues that since Emma had moved the tools to the backyard, she should look there first\. Surfacing multiple reasoning graphs and conclusions allows the users to audit them and decide which answer and justification they agree with\.
Figure 3:Three reasoning graphs reaching the same conclusion \(Answer A\) for a MuSR Murder Mysteries question\.Consensus graph\(left\): the highest\-Ψ\\Psiproof subgraph of the consensus answer, aggregating attested reasoning across the heterogeneous ensemble\.Alternative graph\(middle\): a different proof subgraph for the same conclusion, sampled from the same ensemble DAG by varying the support set atOrnodes\.Single CoT graph\(right\): a random per\-trace DAG from one model whose chain reached the same answer\. The consensus graph integrates evidence from multiple traces and we can observe here that both the alternative graph and CoT graph show less intricate reasoning when compared to the highest weighted consensus graph\.Figure 4:Three reasoning graphs reaching the same conclusion \(Answer A\) for a MuSR Object Placement question\.Consensus graph\(left\): the highest\-Ψ\\Psiproof subgraph of the consensus answer, aggregating attested reasoning across the heterogeneous ensemble\.Alternative conclusion graph\(middle\): a proof subgraph for a different conclusion, sampled from the same ensemble DAG\. The consensus graph integrates evidence from multiple traces; the alternative conclusion graph offers a substantively different reasoning and conclusion \(tools are usually in a tool shed so search there first vs\. they were moved to the backyard so search there first\)\.
## Appendix CSingle vs\. multi\-model diversity
Table 3:Semantic diversity of correct reasoning traces generated by individual model families and a heterogeneous ensemble\. Diversity is measured as mean pairwise1−cosine similarity1\-\\text\{cosine similarity\}between trace embeddings\. Mixed\-model setting achieves highest diversity, indicating a broader range of successful reasoning strategies\.We use multiple model families in our ensemble to obtain diverse reasoning perspectives that reach the correct answer\. To motivate this and verify that heterogeneous ensembles produce more diverse successful trajectories, we sampled eight reasoning traces \(for SARA and GPQA\-D\) either entirely from one model or in a mixed setting with two traces per model \(Qwen 2\.5 7B, Gemma 3 12B, Llama 3\.2 3B, Mistral 7B v0\.3\)\. We retained only traces reaching the correct final answer, embedded them withall\-MiniLM\-L6\-v2Reimers and Gurevych \([2019](https://arxiv.org/html/2607.27783#bib.bib60)\), and computed pairwise diversity as1−cosine similarity1\-\\text\{cosine similarity\}\. Table[3](https://arxiv.org/html/2607.27783#A3.T3)reports the mean and standard deviation of diversity, along with the average number of correct traces per question\. The mixed\-model setting achieves the highest diversity, supporting the claim that combining heterogeneous models yields a broader variety of successful reasoning paths than any single model\.
## Appendix DHuman annotation details
Annotators were students recruited on a volunteer basis\. Instruction:You are provided three reasoning graphs that lead to an answer, select the one that you think justifies the answer in the most logical, correct, coherent way\. If none justify the answer well, you can respond with "none"\. This will be used for research\. Warning: contains examples describing hypothetical murder\-related scenarios\.Similar Articles
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
This paper introduces Reasoning Jury, a system that uses a jury of open-weight LLMs with a moderated consensus mechanism to evaluate long reasoning traces, significantly outperforming frontier models at identifying reasoning defects while costing a fraction of the price.
Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation
This paper proposes SGR, a framework that enhances LLM stepwise reasoning by integrating external knowledge graphs through query-relevant subgraph generation, combining Cypher-based reasoning with collaborative reasoning integration. Experiments on CWQ, WebQSP, GrailQA, and KQA Pro show improved reasoning accuracy over standard prompting and knowledge-enhanced baselines.
Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty
This paper introduces structural uncertainty, a framework that evaluates LLM reasoning consistency by measuring the stability of self-preference rankings among sampled reasoning solutions, complementing traditional answer-dispersion methods for identifying unreliable reasoning.
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
LC-ERD is a framework that mines latent logic from LLM-generated reasoning chains to decompose global rewards into step-level signals, enabling self-evolving reasoning without human annotation. It addresses label noise, coarse supervision, and distributional collapse via variational logic potential and multi-agent value decomposition.
Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework
This paper introduces GraphEVAL, a graph-based framework for quantifying uncertainty in LLM reasoning, and proposes a new metric, Graph Reasoning Coherence Score (GRCS), that captures semantic-structural consensus and detects confident hallucinations. The authors also present Graph Self-Consistency (GSC), a decoding strategy that prioritizes reasoning fidelity over nominal accuracy.