小型LLM团队在不同编排架构下的精确生成-变换分解扩展性研究
摘要
本文提出了一种精确的生成-变换分解方法,用以分析小型7–9B LLM团队在八种编排架构下的扩展性表现。研究发现,扩展收益与任务高度相关:Proposer-Critic 架构在算术基准测试中表现最佳,但没有任何一种架构能在所有任务中全面占优。
arXiv:2609.36104v1 Announce Type: new
Abstract: Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks.
We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.
查看缓存全文
缓存时间: 2026/09/30 09:41
# An Exact Generate–Transform Decomposition of Small-LLM Team ScalingAcross Orchestration Architectures
Source: [https://arxiv.org/html/2609.36104](https://arxiv.org/html/2609.36104)
###### Abstract
Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear\. Sweeping eight agent orchestration architectures across five instruction\-tuned 7–9B models, five short\-answer benchmarks, and an executable\-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task\-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word\-problem benchmarks \(GSM8K, GSMHard\) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task\-averaged number conceals\. Proposer\-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget \(item\-clustered intervals exclude zero\), though it ranks among the weakest elsewhere, and no architecture wins across tasks\.
We explain these trajectories with an exact generate–transform decomposition\. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an*extensive*coverage dividend and an*intensive*transformation change\. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic\-guided transform converts, whereas the multiple\-choice benchmarks either saturate in coverage or fail to convert it, and on open\-ended code generative recovery nearly vanishes so accuracy tracks coverage\. At equal call budgets token cost still varies 2\.1×\\times\. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert\. Team scaling is a task\- and architecture\-specific bet, not a uniform lever\.
Jožef Stefan Institute, Ljubljana, Slovenia
## 1Introduction
Practitioners increasingly replace one LLM call with a team: agents debate\([Du et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib7);[Liang et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib8)\), pass proposals through aggregation layers\([Wang et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib9)\), or communicate over learned graphs\([Zhuge et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib10)\)\. A team is also a compute allocation: the extra calls can generate independent candidates, critique them, revise a single answer serially, or synthesize intermediate reports\. Choosing among these options matters: our largest workflows make 30 calls per problem, and architectures at that same call budget differ 2\.1×\\timesin observed tokens\. Yet most evidence fixes one team size and one base model while changing graph, roles, and synthesis together\. A favorable result at five calls on one model says little about whether the design will keep improving at 30, remain cost\-effective, or transfer to another model or task\.
The gap is increasingly consequential because adaptive systems already select collaboration modes, roles, models, or workflows per query\([Yue et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib14);[Zhang et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib15);[Yu 2026](https://arxiv.org/html/2609.36104#bib.bib16)\)\. They are optimizing over primitives whose scaling behavior is still poorly characterized\. A further call can act at two stages: it can make a correct candidate available, or it can help later computation preserve, repair, and synthesize the evidence already present\. Conflating the two hides which stage limits scaling\. Where scaling helps depends sharply on the task, and the architecture that gains most on one task can be among the weakest on another \(Section[5](https://arxiv.org/html/2609.36104#S5)\)\.
We divide each workflow into two functional stages\. The*generate*stage is the first tier that emits candidate answers\. The*transform*stage contains later critics, refiners, synthesizers, and the final manager\. We ask three questions in a common experimental protocol\.*First \(Q1\)*, when and where does scaling a team help, and does the answer depend on the task?*Second \(Q2\)*, which architecture, if any, captures the gains as the budget grows?*Third \(Q3\)*, do those gains arise because more calls cover more candidate answers, or because later calls transform the available information more effectively? We make three contributions:
1. 1\.Across eight architectures, five small LLMs, and six benchmarks, we show that the returns to team scaling \(Q1\) are sharply task\-dependent: scaling from three to thirty calls lifts accuracy by up to 17 points on the arithmetic word\-problem benchmarks \(GSM8K, GSMHard\) but by at most four on ARC, GPQA, and MMLU, a split that task\-averaged reporting hides, and Proposer\-Critic is the architecture that captures the arithmetic gains while ranking among the weakest elsewhere\.
2. 2\.We formulate an exact generate–transform decomposition that explains these trajectories\. Unlike fixed\-pool selection bounds, in which an uncovered item is lost, it retains generative recovery and so applies to hierarchical managers that synthesize answers outside the proposal pool\. It separates the value of added coverage from covered\-case success and recovery, and diagnoses, task by task, whether scaling is limited by missing coverage or by a transform that fails to convert it\.It can inform where and how to cut\.
3. 3\.Across the full sweep, the decomposition reveals that no architecture \(Q2\) wins over models and tasks, and equal\-budget token cost varies 2\.1×\\times\. The decomposition attributes the pervasive early saturation to a transformation term \(Q3\) that is compositional rather than a within\-item decline: added coverage falls on harder items that convert weakly, a pattern we confirm on both budget scaling and a controlled prompt\-only intervention\.
Section[2](https://arxiv.org/html/2609.36104#S2)summarizes related work, Section[3](https://arxiv.org/html/2609.36104#S3)elaborates on the architecture and decomposition, and Section[4](https://arxiv.org/html/2609.36104#S4)provides the experimental protocol, while Section[5](https://arxiv.org/html/2609.36104#S5)analyzes the results\. Finally, Section[7](https://arxiv.org/html/2609.36104#S7)concludes the paper\.
## 2Related Work
#### Debate, layers, and graphs\.
Multi\-agent debate repeatedly exposes agents to peer answers\([Du et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib7)\), and divergent roles can alter the resulting behavior\([Liang et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib8)\)\. Mixture\-of\-Agents instead passes generations through layers to an aggregator\([Wang et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib9)\), while GPTSwarm treats agent workflows as optimizable computational graphs\([Zhuge et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib10)\)\. Self\-consistency and Self\-Refine are useful limiting cases: independent proposal aggregation and serial revision, respectively\([Wang et al\. 2023](https://arxiv.org/html/2609.36104#bib.bib17);[Madaan et al\. 2023](https://arxiv.org/html/2609.36104#bib.bib18)\)\.
The closest broad comparisons are MultiAgentBench, which includes star, chain, tree, and graph coordination protocols on interactive scenarios\([Zhu et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib13)\), and the information\-propagation study of[Shen et al\. \(2025\)](https://arxiv.org/html/2609.36104#bib.bib11), which analyzes error and correct\-information diffusion as graph sparsity changes\.[Kim et al\. \(2025\)](https://arxiv.org/html/2609.36104#bib.bib12)standardize tools, prompts, and compute across five canonical tool\-using agent architectures and derive predictive principles \(capability, overhead, redundancy, error propagation\) at a fixed scale, rather than a full budget sweep with an exact coverage\-versus\-transformation accounting\. Adaptive systems select collaboration modes, roles, models, or workflows per query\([Yue et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib14);[Zhang et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib15);[Yu 2026](https://arxiv.org/html/2609.36104#bib.bib16)\)\. These works motivate architecture selection\. Our complementary goal is complete budget trajectories for fixed small\-model workflows, direct accounting between their first answer\-producing tier and final output, and an exact decomposition of how the two stages contribute to scaling\.
#### Scaling and diversity\.
MacNet organizes more than a thousand agents in DAGs and fits logistic performance curves, finding that topology affects the curve\([Qian et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib29)\)\.[Yang et al\. \(2026\)](https://arxiv.org/html/2609.36104#bib.bib30)instead derive architecture\-agnostic information bounds and an effective\-channel count, showing that heterogeneous channels can substitute for many homogeneous agents\. In LLM evaluation panels, correlated errors likewise reduce nine nominal judges to roughly two effective votes\([Kohli 2026](https://arxiv.org/html/2609.36104#bib.bib31)\)\. These results establish that nominal agent count is not informational count\. Repeated\-sampling coverage laws\([Brown et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib37)\)and voting\-based call scaling\([Chen et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib38)\)characterize this generate stage in isolation\. We add a common small\-model budget sweep across role\-asymmetric workflows and account for proposal generation and subsequent transformation separately\. A single topology\-wide effective team size is ambiguous here, because critics, refiners, and synthesizers produce dependent intermediate answers, so we use agreement only as a manipulation check in the matched Star/Persona\-Star comparison and rely on transfer events for cross\-architecture analysis\.
#### Oracle bounds, selection, and generative aggregation\.
Candidate\-pool oracles are established upper bounds for model selection\. SelectLLM, for example, analyzes the gap between a learned selector and an oracle over available models\([Maurya et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib2)\)\. LLM\-Blender makes the adjacent distinction between ranking candidates and generatively fusing them\([Jiang et al\. 2023](https://arxiv.org/html/2609.36104#bib.bib3)\)\. Concurrent fixed\-pool work further separates recoverable mass, selection\-signal quality, and harm to correct outputs\([Hu 2026](https://arxiv.org/html/2609.36104#bib.bib1)\)\. If the system must select from a fixed pool,O=0O=0impliesY=0Y=0\. Hierarchical managers can instead synthesize an answer absent from the initial pool\. Generative Self\-Aggregation has already demonstrated such successes when every sampled answer is wrong\([Li et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib4)\)\. Aggregation Fine\-Tuning and Recursive Self\-Aggregation likewise combine parallel proposal generation with sequential synthesis\([Li et al\. 2026](https://arxiv.org/html/2609.36104#bib.bib5);[Venkatraman et al\. 2025](https://arxiv.org/html/2609.36104#bib.bib6)\)\. We use the proposal boundary as a common diagnostic and jointly measure coverage, loss, and recovery as heterogeneous fixed workflows scale\.
The selection\-versus\-synthesis distinction is close to the selection bottleneck identified by[Maryanskyy et al\. \(2026\)](https://arxiv.org/html/2609.36104#bib.bib32), who show that generator diversity pays only when the downstream selector is sufficiently reliable\. Homogeneous debate also exhibits consensus collapse, in which aggregation discards correct answers already generated\([Bertalanič and Fortuna 2026](https://arxiv.org/html/2609.36104#bib.bib35)\)\. Our accounting places such discard and generative repair in one exact identity and evaluates both across parallel and serial workflows\.
#### Compute\-normalized evaluation\.
Extra agents also mean extra test\-time compute\. Under matched reasoning\-token budgets, single\-agent systems can match or exceed several multi\-agent designs on multi\-hop reasoning\([Tran and Kiela 2026](https://arxiv.org/html/2609.36104#bib.bib34)\)\. We therefore report both requested calls and observed tokens\.
## 3Orchestration Architectures and Decomposition Framework
We assume an orchestration architecture divides its node budget between two functions: creating candidate answers and transforming the resulting evidence\. LetNNbe the requested node budget, equivalently the team size and the number of LLM calls per problem, andLLthe number of nodes in the first tier that emits a parseable candidate answer\. Planners instructed not to answer, along with all later critics, refiners, and managers, are excluded fromLL\.
### 3\.1Orchestration architectures
Figure 1:The eight orchestration architectures, nodes colored by functional stage: generate \(theLLproposals\), transform \(critics, refiners, synthesizers, duelists\), the manager, and Diamond’s non\-answering planner\.Tournamentselects through pairwise duels,Treemerges via fan\-in\-three synthesis, andPersona\-Starshares Star’s graph but differs in prompts\. Counts are illustrative\.We evaluate eight directed acyclic graphs \(Figure[1](https://arxiv.org/html/2609.36104#S3.F1)\), each terminating in an active manager and grouped by how the first answer\-producing tier of sizeLLscales with the budget\.
- •Parallel, proposal\-expanding \(L=N−1L=N\-1\)\.StarroutesN−1N\-1unconstrained workers to a manager\.Persona\-Starkeeps Star’s graph and manager fixed while cycling six reasoning\-method prompts \(forward, backward, decomposition, step\-back, conservative, and contrarian\) across workers\([Wei et al\. 2022](https://arxiv.org/html/2609.36104#bib.bib19);[Zhou et al\. 2023](https://arxiv.org/html/2609.36104#bib.bib20);[Zheng et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib21)\), so only its roles differ, and all six are present fromN=7N=7\.
- •Hierarchical and filtered \(1<L<N−11<L<N\-1forN≥5N\\geq 5\)\.Proposer\-Criticpairs proposals with critics \(L=⌈\(N−1\)/2⌉L=\\lceil\(N\-1\)/2\\rceil\),Tournamentfilters proposals through a pairwise bracket,Treemerges them with fan\-in\-three synthesizers, andDiamondroutesmax\(1,⌊N/5⌋\)\\max\(1,\\lfloor N/5\\rfloor\)non\-answering planners to plan\-conditioned solvers \(L=N−max\(1,⌊N/5⌋\)−1L=N\-\\max\(1,\\lfloor N/5\\rfloor\)\-1\)\.
- •Serial depth \(L=1L=1\)\.Chainpasses one evolving solution throughN−2N\-2single\-parent refiners\.Cascading\-Chainlets each refiner read up to five preceding reports\.
Requested budgets areN∈\{2,3,5,7,10,15,20,30\}N\\in\\\{2,3,5,7,10,15,20,30\\\}\. Diamond starts atN=3N=3\. Tournament uses the largest complete pairwise bracket within the request, so its actual counts are 9, 19, and 29 at requested 10, 20, and 30, and all other conditions use the requested count\. Equal call budgets do not buy equal proposal budgets: atN=30N=30the first tier holdsL=29L=29proposals for both stars, 15 forProposer\-CriticandTournament, 22 forTree, 23 forDiamond, and 1 for both chains\. The remaining calls are critics, duelists, synthesizers, refiners, planners, and the manager\. Exact node compositions are tabulated in Appendix A\.
### 3\.2Proposal coverage and agreement
For valid proposal answersa1,…,aLva\_\{1\},\\ldots,a\_\{L\_\{v\}\}, exact\-answer pairwise agreement isApair=\(Lv2\)−1∑i<j𝟏\[ai=aj\]A\_\{\\rm pair\}=\\binom\{L\_\{v\}\}\{2\}^\{\-1\}\\sum\_\{i<j\}\\mathbf\{1\}\[a\_\{i\}=a\_\{j\}\]\. Agreement is undefined whenLv<2L\_\{v\}<2and is used only when the proposal boundary is held fixed, principally in the Star/Persona\-Star intervention\. Role\-asymmetric architectures are instead compared by loss, recovery, and final accuracy below\.
### 3\.3Generate–Transform decomposition
LetO=1O=1when at least one proposal is correct andY=1Y=1when the final answer is correct\. A final answer can then cross the proposal boundary in either direction: downstream computation can discard an available correct answer or recover when every proposal is wrong\. At a given budget, define the corresponding unconditional masses:
ℓ\\displaystyle\\ell=P\(O=1,Y=0\)\\displaystyle=P\(O=1,Y=0\)\(transfer loss\),\\displaystyle\\text\{\(transfer loss\)\},\(1\)r\\displaystyle r=P\(O=0,Y=1\)\\displaystyle=P\(O=0,Y=1\)\(transfer recovery\)\.\\displaystyle\\text\{\(transfer recovery\)\}\.\(2\)These two events give the exact endpoint identity
P\(Y=1\)−P\(O=1\)=r−ℓ\.P\(Y=1\)\-P\(O=1\)=r\-\\ell\.\(3\)Consequently, the oracle gapP\(O=1\)−P\(Y=1\)P\(O=1\)\-P\(Y=1\)\(proposal coverage minus final accuracy\) isℓ−r\\ell\-r, a*net*quantity, not the rate at which the final agent discards a correct proposal\. A selector restricted to the proposal pool hasr=0r=0\. Recovery can be nonzero here because critics, refiners, and managers are generative reasoners rather than passive voting rules\.
To analyze scaling, writeON=P\(O=1\)O\_\{N\}=P\(O=1\)\(proposal coverage, the accuracy of an oracle over the proposal pool\) andYN=P\(Y=1\)Y\_\{N\}=P\(Y=1\)at budgetNN, and define
sN=P\(Y=1∣O=1\),gN=P\(Y=1∣O=0\)\.s\_\{N\}=P\(Y=1\\mid O=1\),g\_\{N\}=P\(Y=1\\mid O=0\)\.\(4\)Heressis covered\-case success \(one minus conditional discard\),ggis generative recovery, andvN=sN−gNv\_\{N\}=s\_\{N\}\-g\_\{N\}is proposal leverage: the accuracy lift associated with an available correct proposal\. Total probability gives the response identity:
YN=ONsN\+\(1−ON\)gN=gN\+vNON\.Y\_\{N\}=O\_\{N\}s\_\{N\}\+\(1\-O\_\{N\}\)g\_\{N\}=g\_\{N\}\+v\_\{N\}O\_\{N\}\.\(5\)For any two conditionsAAandBB\(here, two team sizes of a single architecture\), writeΔx=xB−xA\\Delta x=x\_\{B\}\-x\_\{A\}andx¯=\(xA\+xB\)/2\\bar\{x\}=\(x\_\{A\}\+x\_\{B\}\)/2\. Applying the exact midpoint product identity to Eq\.[5](https://arxiv.org/html/2609.36104#S3.E5)yields:
ΔY\\displaystyle\\Delta Y=v¯ΔO⏟ℰ:extensive coveragedividend\\displaystyle=\\underbrace\{\\bar\{v\}\\Delta O\}\_\{\\mathcal\{E\}:\\ \\begin\{subarray\}\{c\}\\text\{extensive coverage\}\\\\ \\text\{dividend\}\\end\{subarray\}\}\+Δg\+O¯Δv⏟ℐ:intensivetransformation change\\displaystyle\\quad\+\\underbrace\{\\Delta g\+\\bar\{O\}\\Delta v\}\_\{\\mathcal\{I\}:\\ \\begin\{subarray\}\{c\}\\text\{intensive\}\\\\ \\text\{transformation change\}\\end\{subarray\}\}\(6\)
ThusΔY=ℰ\+ℐ\\Delta Y=\\mathcal\{E\}\+\\mathcal\{I\}exactly, without a fitted functional form\. We call Eq\.[6](https://arxiv.org/html/2609.36104#S3.E6)the two\-margin*Generate–Transform decomposition*: it is an exact accounting identity, not an empirical power curve\. Algebraically it is a symmetric rate–composition decomposition\([Kitagawa 1955](https://arxiv.org/html/2609.36104#bib.bib33)\)\. Our instantiation uses proposal coverage as the composition and the conditional success rates as the rate, keeping generative recovery inside the accounting rather than assuming a selector\.ℰ\\mathcal\{E\}weights the coverage change by midpoint leveragev¯\\bar\{v\}, so it credits only coverage the downstream stage can exploit\.ℐ=O¯Δs\+\(1−O¯\)Δg\\mathcal\{I\}=\\bar\{O\}\\Delta s\+\(1\-\\bar\{O\}\)\\Delta gmeasures transformation change at fixed midpoint opportunity weights\. Because the final answer is the manager’s own output rather than a tally over the workers,ssis the rate at which the manager returns an available correct answer and1−s1\-sthe rate at which it discards one\. A negativeℐ\\mathcal\{I\}records a fall in aggregate covered\-case success\. The paired coverage transitions below attribute this either to a same\-item decline or to harder newly\-covered items entering the covered population, and we find it is chiefly the latter\.
For a paired interventionA→BA\\to B, marginal conditional rates can still be misleading becauseAAandBBneed not cover the same items\. We therefore partition paired trials by their realized coverage transition\(OA,OB\)∈\{00,01,10,11\}\(O\_\{A\},O\_\{B\}\)\\in\\\{00,01,10,11\\\}\. Withπij=P\(OA=i,OB=j\)\\pi\_\{ij\}=P\(O\_\{A\}=i,O\_\{B\}=j\),
ΔY=∑i,j∈\{0,1\}πijE\[YB−YA∣OA=i,OB=j\]\.\\Delta Y=\\sum\_\{i,j\\in\\\{0,1\\\}\}\\pi\_\{ij\}E\[Y\_\{B\}\-Y\_\{A\}\\mid O\_\{A\}=i,O\_\{B\}=j\]\.\(7\)Each term is the stratum’s exact contribution to the paired effect\. Rate differences have long been separated into composition and conditional\-rate components\. Here the paired runs make the coverage transitions directly observable\.
For QA we additionally recompute deterministic plurality over the logged proposals using the experiment’s item\-seeded tie break\. This offline control invokes no LLM and distinguishes voting from active synthesis\.
## 4Experimental Protocol
#### Models and tasks\.
We evaluate Llama\-3\.1\-8B\-Instruct, Ministral\-3\-8B\-Instruct\-2512, NVIDIA\-Nemotron\-Nano\-9B\-v2, Qwen2\.5\-7B\-Instruct, and Qwen3\-8B\. Nemotron and Qwen3 run with thinking disabled, keeping the five models in a common standard\-inference setting\. We use QA as shorthand for five short\-answer benchmarks \(a small answer space, unlike open\-ended code\): ARC\-Challenge \(1,165 items\), GPQA \(198\), the arithmetic sets GSM8K \(1,319\) and GSMHard \(1,017\), and a 724\-item hard\-subject MMLU subset\([Clark et al\. 2018](https://arxiv.org/html/2609.36104#bib.bib22);[Rein et al\. 2024](https://arxiv.org/html/2609.36104#bib.bib23);[Cobbe et al\. 2021](https://arxiv.org/html/2609.36104#bib.bib24);[Gao et al\. 2023](https://arxiv.org/html/2609.36104#bib.bib25);[Hendrycks et al\. 2021](https://arxiv.org/html/2609.36104#bib.bib26)\)\. Functional code uses all 164 HumanEval problems scored with HumanEval\+ tests\([Chen and others 2021](https://arxiv.org/html/2609.36104#bib.bib27);[Liu et al\. 2023](https://arxiv.org/html/2609.36104#bib.bib28)\)\.
#### Inference protocol\.
Each cell has three runs with distinct prescribed node seeds\. vLLM serves each model\([Kwon et al\. 2023](https://arxiv.org/html/2609.36104#bib.bib36)\)\. Sampling uses temperature 0\.4 and nucleus probabilityp=0\.95p=0\.95\. Maximum generation is 1,024 tokens on QA and 2,048 on code\. Each parent report is capped at 120 tokens on QA and 512 on code\. The final answer is always the manager’s parsed output\. HumanEval\+ candidates execute in the EvalPlus sandbox\. Prompts and parsing rules are included in the code and summarized in Appendix B\.
#### Statistics\.
The analyzed sweep contains 4,179,735 QA and 154,980 code team trials, representing 48,498,195 and 1,798,260 individual agent responses\. We average the three runs within item and use 1,000 item\-clustered bootstrap replicates for 95% intervals\. Paired comparisons retain shared item and run identifiers\. Unless stated otherwise, aggregate numbers weight the 25 QA model–task cells equally\. Invalid \(unparseable\) proposals are rare, with an equal\-cell mean of 0\.026%\. They are excluded from agreement calculations but remain part of the requested\-call cost\.
Compatible logs provide a directN=1N=1condition for all 15 hard\-QA model–task cells \(GPQA, GSMHard, and hard\-subject MMLU across five models\) with one run, and for all five HumanEval\+ models with three runs\. It uses the ordinary task prompt without a persona, peer report, or architecture\-specific role\. A one\-run long\-reasoning control instead uses an extended\-verification prompt and raises the cap from 1,024 to 10,240 tokens\. It covers 15/15 hard\-QA cells\. Because prompt and cap change together and generation can stop early, this is a 10×\\times\-cap control, not a pure token effect\. Two dense\-communication controls run for three rounds and aggregate by plurality: full\-mesh Debate, in which every agent sees all peers each round, and a no\-peer Self control, in which each agent revises only its own previous answer\. Both cover the 15 hard\-QA cells atN=30N=30and the full HumanEval\+ budget sweep\. All controls are excluded from primary\-sweep counts\.
## 5Results
In this section we leverage the architecture and decomposition framework from Section[3](https://arxiv.org/html/2609.36104#S3), following the protocol from Section[4](https://arxiv.org/html/2609.36104#S4)to answerQ1\-Q3\. We first show where team scaling pays off and which architecture captures it \(Q1\), then read the accuracy–cost dependence across architectures \(Q2\), and finally use the generate–transform decomposition to explain the trajectories \(Q3\)\. Endpoint, prompt, and debate controls delimit the interpretation, and a controlled prompt\-only intervention is found in Appendix G\.1\.
Figure 2:Where team scaling pays off\. \(a\) The scaling gainΔY\\Delta Y\(theN=3N=3toN=30N=30accuracy change, Section[3\.3](https://arxiv.org/html/2609.36104#S3.SS3)\) for every architecture on every benchmark \(equal\-cell over five models\), grouped by task family, with Proposer\-Critic marked\. Returns concentrate on arithmetic and are near\-zero on ARC, GPQA, and MMLU for all architectures\. \(b\) Arithmetic accuracy versus team size: Proposer\-Critic \(red\) starts mid\-field and separates with scale\.Figure 3:Accuracy gain over a single agent \(the mean individual worker\), by task category: \(a\) arithmetic, \(b\) multiple\-choice, \(c\) code, each cell an equal mean over the category’s model–task cells with per\-panel color scales\. Gains on arithmetic \(25 to 40 points\) are two to three times those on multiple\-choice \(8 to 14\) or code, and Proposer\-Critic tops arithmetic atN=30N=30while trailing on multiple\-choice\. \(d\) QA tokens per problem versus team size: cost climbs at architecture\-dependent rates to a 2\.1×\\timesspread byN=30N=30\. Diamond is undefined atN=2N=2\.### 5\.1Scaling returns concentrate on arithmetic
We find that team scaling does not help uniformly\. Figure[2](https://arxiv.org/html/2609.36104#S5.F2)\(a\) plots every architecture’s accuracy gain fromN=3N=3toN=30N=30on each benchmark\. We notice the gains cluster by task: on the two arithmetic word\-problem benchmarks \(GSM8K, GSMHard\) the best architecture adds 12 to 17 points, on open\-ended code 4 to 6, and on ARC, GPQA, and MMLU at most about four points for*any*architecture\. A single task\-averaged number blends the large arithmetic effect with the near\-flat scaling of the others into a misleadingly uniform figure\.
Within the tasks that scale, one architecture stands out\. Proposer\-Critic is the steepest arithmetic scaler \(Figure[2](https://arxiv.org/html/2609.36104#S5.F2)\(b\)\): sixth of eight atN=3N=3, it overtakes the field byN≈10N\\approx 10and in aggregate atN=30N=30surpasses every other architecture on arithmetic by 5\.7 to 10\.2 points, all item\-clustered 95% intervals excluding zero \(the runner\-up margin is\+5\.7\+5\.7, CI\[5\.1,6\.4\]\[5\.1,6\.4\]\)\. The advantage is modal, not universal: Proposer\-Critic leads arithmetic for three to four of the five models, a chain for the rest, and it ranks last or near\-last on ARC, MMLU, and code\. The decomposition \(introduced in Section[3\.3](https://arxiv.org/html/2609.36104#S3.SS3)\) explains the split \(Section[5\.3](https://arxiv.org/html/2609.36104#S5.SS3)\): on arithmetic Proposer\-Critic’s transform converts the added coverage \(ℐ=\+8\.1\\mathcal\{I\}=\+8\.1on GSM8K,\+3\.0\+3\.0on GSMHard\), whereas on GPQA and MMLU it adds coverage but sheds most of it \(ℐ=−5\.0\\mathcal\{I\}=\-5\.0and−3\.6\-3\.6\), and ARC has little to add \(ℰ=\+1\.4\\mathcal\{E\}=\+1\.4\)\.Answer to Q1: The returns to team scaling are sharply task\-dependent\. FromN=3N=3toN=30N=30accuracy rises by up to 17 points on the arithmetic word\-problem benchmarks \(GSM8K, GSMHard\) and by 4 to 6 on open\-ended code, but by at most 4 on the multiple\-choice benchmarks \(ARC, GPQA, MMLU\) for every architecture\.
Figure 4:Two\-margin decomposition of theN=3→30N=3\\to 30accuracy change by task category: \(a\) arithmetic, \(b\) multiple\-choice, \(c\) code\. Each architecture’s change splits into an extensive coverage dividendℰ\\mathcal\{E\}\(blue\) and an intensive transformation changeℐ\\mathcal\{I\}\(red, extending left when the transform sheds coverage\), summing to the netΔY\\Delta Y\(diamond\), equal\-cell over each category’s model–task cells\. Coverage is added in every category, but the transform converts it only on arithmetic, where Proposer\-Critic alone posts a positiveℐ\\mathcal\{I\}and the largest net gain\. On multiple\-choice every proposal\-expanding design sheds most of its coverage and nets almost nothing, and on code the same shedding, with recovery near zero, leaves the net close to the coverage floor\.
### 5\.2Scaling is architecture\-, model\-, and cost\-dependent
Figure[3](https://arxiv.org/html/2609.36104#S5.F3)\(a–c\) gives each architecture’s gain over a single agent by budget and task category, and the trajectory shape itself is task\-dependent\. On multiple\-choice \(b\) and code \(c\) most of the lift is present byN=3N=3and the curves flatten byN=10N=10, whereas on arithmetic \(a\) they keep climbing toN=30N=30, where Proposer\-Critic reaches\+40\+40points over a single agent \(level, not slope\)\. Over theN=3→30N=3\\to 30range Proposer\-Critic improves in 21/25 QA model–task cells and Chain in 24/25, the most consistent signs\. Persona\-Star and Diamond reach similar aggregate gains through sharply different margin profiles \(Section[5\.3](https://arxiv.org/html/2609.36104#S5.SS3)\)\. We read theN=3N=3to 30 slope, not the intercept, as the scaling result, because the intercept is baseline\-sensitive and shrinks against a stronger single agent \(Section[5\.5](https://arxiv.org/html/2609.36104#S5.SS5)\)\.
Equal team sizes are not equal cost \(Figure[3](https://arxiv.org/html/2609.36104#S5.F3)\(d\)\)\. Token cost climbs withNNat architecture\-dependent rates: atN=30N=30mean QA cost ranges from 15\.1k tokens for Tournament to 31\.3k for Cascading\-Chain, whose accumulating context grows fastest, a 2\.1×\\timesspread\. Combining this cost with accuracy \(Appendix E\), Tournament, Star, Tree, and Proposer\-Critic form the accuracy–cost frontier, with Proposer\-Critic highest at 59\.1% and 17\.1k tokens\. A purpose\-built single\-agent control given a 10×\\timestoken cap gains 9\.18 points over the ordinary single agent and is matched on average by the post\-hoc best architecture at 13\.5×\\timesthe observed tokens, so extra test\-time compute in a single agent is itself competitive \(Appendix H\)\.
AtN=30N=30the level comparison splits sharply by category \(Appendix E\)\. On arithmetic the within\-model architecture spread is large \(mean 14\.0 points, up to 25\.6 for Llama\) and Proposer\-Critic leads four of the five models, a chain leading Ministral\. On multiple\-choice the spread collapses to a mean of 4\.3 points, every design lands within a few points, and a chain is marginally best\. No architecture wins across models in either category\. Model selection stays the larger lever on multiple\-choice, where architecture barely moves accuracy, whereas on arithmetic the architecture spread rivals it\. Teams beat a single agent by far more on arithmetic, where one can solve as little as 8% of items, than on multiple\-choice, a like\-for\-like margin a stronger single agent substantially narrows \(Appendix H\)\. The HumanEval\+ levels tell the same story \(Appendix E\): no architecture wins across models, and every design beats a single agent by \+6\.0 \(Diamond\) to \+14\.3 \(Tournament\) points\.
Answer to Q2: While no architecture wins across all tasks, we find thatProposer\-Criticcaptures the arithmetic gains, surpassing every other architecture in aggregate atN=30N=30on arithmetic \(GSM8K and GSMHard\), yet it ranks last or near\-last on ARC, MMLU, and code\. Equal\-budget token cost still varies 2\.1×\\times, andProposer\-Criticsits on the accuracy–cost frontier\.
### 5\.3Scaling gains split into coverage and transformation
Accuracy curves establish whether another call helped, but not what that call purchased\. We apply Eq\.[6](https://arxiv.org/html/2609.36104#S3.E6)within each model–task cell fromN=3→30N=3\\to 30, average the signed terms, and read them by task category \(Figure[4](https://arxiv.org/html/2609.36104#S5.F4)\)\. Added coverage is reliably valuable: across the six proposal\-expanding designs \(Section[3\.1](https://arxiv.org/html/2609.36104#S3.SS1)\) the extensive dividendℰ\\mathcal\{E\}is positive in every one of the 150 cells, from\+4\.9\+4\.9to\+14\.1\+14\.1points depending on category and design\. What separates the tasks is not coverage but whether the downstream transform converts it, and that is the intensive term\.
The intensive termℐ\\mathcal\{I\}has a task\-dependent sign \(Figure[4](https://arxiv.org/html/2609.36104#S5.F4)\)\. On the multiple\-choice benchmarks it is negative for nearly every width design \(87 of 90 cells,−2\.8\-2\.8to−10\.1\-10\.1points\) and cancels almost the whole dividend: Diamond turns a\+11\.5\+11\.5coverage gain into\+1\.4\+1\.4final points, and no width design nets above\+2\.3\+2\.3\. On arithmetic the transform sheds far less \(ℐ\\mathcal\{I\}negative in 43 of 60 cells, and only−0\.9\-0\.9to−5\.2\-5\.2where it is\), and Proposer\-Critic reverses it outright \(ℐ=\+5\.55\\mathcal\{I\}=\+5\.55\), converting its coverage and adding recovery for a\+14\.6\+14\.6\-point net gain, the largest of any design and category\. Pooled over both QA categories, Proposer\-Critic is the only width design whose intensive term does not significantly erode its dividend \(ℐ=\+0\.54\\mathcal\{I\}=\+0\.54, 95% CI\[−0\.02,1\.05\]\[\-0\.02,1\.05\], Appendix D\)\.
This intensive decline is consistent with composition rather than a same\-item transform decline\. Pairing every item acrossN=3N=3andN=30N=30by its coverage transition \(Eq\.[7](https://arxiv.org/html/2609.36104#S3.E7)\), on the stratum covered at both budgets final accuracy holds or rises for five of the six width designs \(\+2\.2\+2\.2to\+6\.7\+6\.7points, with Diamond flat and Tournament−0\.9\-0\.9\), so the drop in aggregate covered\-case success is carried by weakly\-converting newly\-covered items rather than by degradation on shared ones \(Appendix D\)\.
The serial chains invert this in every category\. Their first tier is a single proposal, soℰ\\mathcal\{E\}is effectively zero and their gains are entirely intensive, driven by recovery:ℐ=\+6\.35\\mathcal\{I\}=\+6\.35\(Chain\) and\+4\.15\+4\.15\(Cascading\-Chain\) on arithmetic,\+2\.42\+2\.42and\+1\.33\+1\.33on multiple\-choice\. An added call therefore has architecture\-dependent meaning\. Width buys candidate opportunity and must overcome an intensive headwind, whereas serial depth buys repeated repair of one evolving solution, and nominal team size alone obscures the distinction\.
The endpoint massesℓ\\ellandrrshow what the final transformation does with its proposals \(full table in the Appendix D\)\. Parallel architectures are loss\-dominated atN=30N=30: Diamond and Persona\-Star discard 27\.5% and 25\.5% of covered cases and recover under 5% of uncovered ones, whereas the chains reverse the pattern, recovering 29\.2% and 24\.9% of uncovered cases against 3% loss\. Proposer\-Critic nearly balances the two\.
This contrast shows up directly at the endpoints: the chains recover far more than they lose and finish above their proposal oracle, Proposer\-Critic nearly balances the two, and the five other width designs are loss\-dominated and finish below their proposal oracle\. The crossing boundary that formalizes it, with per\-architecture coverage values and its stability across budgets, is given in the Appendix D\. Loss\- and recovery\-dominance are not fixed architecture labels: the same model and architecture can switch by task\. With Diamond, Qwen3 recovers on 20\.8% of GSM8K trials but only 0\.3% on GPQA, so a one\-sided oracle gap would flag only the latter, yet Eq\.[3](https://arxiv.org/html/2609.36104#S3.E3)shows both are outcomes of the same transfer stage\.
#### Open\-ended code\.
The same decomposition \(Section[3\.3](https://arxiv.org/html/2609.36104#S3.SS3)\) applies, since the proposal boundary needs only executable correctness \(a correct program passes the tests\) and textual agreement is not used\. The width pattern reproduces on HumanEval\+ \(Figure[4](https://arxiv.org/html/2609.36104#S5.F4)\(c\)\): the six proposal\-expanding designs post an extensive dividend of \+5\.2 to \+18\.0 points fromN=3N=3to 30 offset by a negative intensive term, and Diamond adds the most coverage \(proposal oracle \+26\.1\) yet gains only \+6\.0 final points\. Two differences sharpen the account\. Generative recovery nearly disappears, withg=P\(Y=1∣O=0\)g=P\(Y\{=\}1\\mid O\{=\}0\)at 0–1% for Star, Persona\-Star, and Diamond against 15–31% across QA\. And every architecture finishes below its proposal oracle \(Y−OY\-Ofrom−1\.1\-1\.1to−19\.3\-19\.3\), whereas the QA chains finished 12–14 points above\. Open\-ended synthesis rarely yields a passing program the pool lacked, so the transform stage preserves or discards coverage but seldom recovers it, and accuracy tracks coverage minus loss\. On this one open\-ended benchmark the recovery that lets QA managers exceed their proposal oracle disappears, consistent with a small\-answer\-space affordance \(Appendix H\)\.Answer to Q3: The gains could come from extra calls more often making a correct answer available among the candidates \(coverage\), or from the final step converting those candidates into a correct final answer\. Both matter, but conversion is the bottleneck: coverage rises on every task, yet accuracy improves only where the downstream transform converts it\. Conversion succeeds on arithmetic, but on multiple\-choice and code the transform discards most of the newly available answers\.
### 5\.4Choosing an architecture: resolution beats prediction
\(1\) Scale only where a single call is weak\.Team scaling adds a lot where one call misses many items and little where it already scores well \(Section[5\.1](https://arxiv.org/html/2609.36104#S5.SS1)\): large gains on arithmetic and multi\-step reasoning, negligible ones on knowledge and multiple\-choice, where a single agent given a longer reasoning budget is competitive at far lower cost \(Section[5\.5](https://arxiv.org/html/2609.36104#S5.SS5)\)\. If one call already handles your task, do not scale\.
\(2\) If scaling helps, find the architecture by testing, not guessing\.If you can label a small development set, run several architectures at a small budget \(N=10N=10\), keep the best, and deploy it at your target budget \(N=30N=30\), landing within 1\.2 points of the best of the eight\. If you cannot label anything, deploy the task’s historically strongest architecture: Proposer\-Critic on arithmetic, a chain otherwise\. Rules that route on the decomposition’s own signals do no better than this default \(1\.55, 4\.02, and 2\.46 points below the best of the eight, versus the default’s 1\.59\), so the decomposition diagnoses scaling but cannot route it \(Appendix D\.1\)\.
### 5\.5Controls and robustness
Against a purpose\-built directN=1N=1call, a stronger single agent using a cleaner prompt \(and long reasoning for Nemotron\), the margin narrows: every fixedN=30N=30architecture still improves on the 15 hard\-QA cells \(\+3\.0 to \+6\.6\), the post\-hoc best by 8\.9 \(13/15, Appendix H\)\)\.
Three\-round full\-mesh debate beats directN=1N=1by 8\.3 points on the 15 hard\-QA cells, but the post\-hoc best sparse architecture is higher in 10/15 at 7\.3–24\.3×\\timesfewer tokens, and the no\-peer control edges it by 0\.69 point\. On code, where debate runs the full budget sweep, one round of peer exchange captures its entire benefit and beats no\-peer revision by 2\.4 points, yet debate still only ties the best sparse design at twice the calls\. Dense peer exchange is therefore a costly and task\-dependent baseline, not a free win \(Appendix I\)\.
Managers are not deterministic votes: final accuracy exceeds offline plurality by 7\.8–13\.3 points across the six proposal\-expanding architectures, and changing only the manager instruction can cut accuracy through lower recovery \(Appendix G\.2\)\.
## 6Limitations
Except for Star/Persona\-Star and Chain/Cascading\-Chain, graph and role prompts change together, so we rank orchestration bundles rather than graph structure alone\. Budgets match calls rather than tokens\. We study fixed, homogeneous 7–9B non\-thinking teams toN=30N=30on short\-answer and executable\-code tasks, not frontier or heterogeneous models, tool use, or long\-form generation\. The code results use a single benchmark, so the vanished\-recovery effect cannot be separated from domain, answer format, or execution\-based scoring\.
Agreement is answer\-space dependent and is used only for the matched Persona intervention\. The Generate–Transform decomposition is exact accounting, not causal mediation or a parametric forecast:ssandggcondition on populations that can change, soℐ\\mathcal\{I\}can mix composition and transformation\.
## 7Conclusion
Team scaling is not a uniform lever: its returns concentrate on arithmetic word problems, where Proposer\-Critic converts the added coverage at scale, while ARC, GPQA, and MMLU gain little for any architecture\. The generate–transform decomposition makes the difference legible, separating the coverage a workflow adds from whether its transform converts it\. No architecture wins across models and tasks\. The right question is where, and at what budget, scaling pays at all\.
## References
- Bertalanič and Fortuna \(2026\)B\. Bertalanič and C\. FortunaThe cost of consensus: isolated self\-correction prevails over unguided homogeneous multi\-agent debate\.InProceedings of the ACM Conference on AI and Agentic Systems,CAIS ’26,pp\. 311–329\.External Links:[Document](https://dx.doi.org/10.1145/3786335.3813137)Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p2.1)\.
- Brownet al\.\(2024\)B\. Brown, J\. Juravsky, R\. Ehrlich, R\. Clark, Q\. V\. Le, C\. Ré, and A\. MirhoseiniLarge language monkeys: scaling inference compute with repeated sampling\.arXiv preprint arXiv:2407\.21787\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2024\)L\. Chen, J\. Q\. Davis, B\. Hanin, P\. Bailis, I\. Stoica, M\. Zaharia, and J\. ZouAre more LLM calls all you need? towards the scaling properties of compound AI systems\.InProceedings of NeurIPS,Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2021\)M\. Chenet al\.Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px1.p1.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try ARC, the AI2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px1.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px1.p1.1)\.
- Duet al\.\(2024\)Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. MordatchImproving factuality and reasoning in language models through multiagent debate\.InProceedings of ICML,Cited by:[§1](https://arxiv.org/html/2609.36104#S1.p1.1),[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p1.1)\.
- Gaoet al\.\(2023\)L\. Gao, A\. Madaan, S\. Zhou, U\. Alon, P\. Liu, Y\. Yang, J\. Callan, and G\. NeubigPAL: program\-aided language models\.InProceedings of ICML,Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px1.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.InProceedings of ICLR,Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px1.p1.1)\.
- Hu \(2026\)J\. HuOracle gap and signal fidelity: a fixed\-pool diagnostic for test\-time collaboration\.arXiv preprint arXiv:2607\.17531\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p1.1)\.
- Jianget al\.\(2023\)D\. Jiang, X\. Ren, and B\. Y\. LinLLM\-Blender: ensembling large language models with pairwise ranking and generative fusion\.InProceedings of ACL,Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p1.1)\.
- Kimet al\.\(2025\)Y\. Kim, K\. Gu, C\. Park, C\. Park, S\. Schmidgall, A\. A\. Heydari, Y\. Yan, Z\. Zhang, Y\. Zhuang, Y\. Liu, M\. Malhotra, P\. P\. Liang, H\. W\. Park, Y\. Yang, X\. Xu, Y\. Du, S\. Patel, T\. Althoff, D\. McDuff, and X\. LiuTowards a science of scaling agent systems\.arXiv preprint arXiv:2512\.08296\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p2.1)\.
- Kitagawa \(1955\)E\. M\. KitagawaComponents of a difference between two rates\.Journal of the American Statistical Association50\(272\),pp\. 1168–1194\.External Links:[Document](https://dx.doi.org/10.1080/01621459.1955.10501299)Cited by:[§3\.3](https://arxiv.org/html/2609.36104#S3.SS3.p3.1)\.
- Kohli \(2026\)G\. KohliNine judges, two effective votes: correlated errors undermine LLM evaluation panels\.arXiv preprint arXiv:2605\.29800\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px2.p1.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with PagedAttention\.InProceedings of SOSP,Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2026\)Y\. Li, Z\. Wang, T\. Fu, G\. Cui, S\. Yang, and Y\. ChengThe best of both worlds: combining parallel and sequential inference scaling via aggregation fine\-tuning\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 31369–31389\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p1.1)\.
- Liet al\.\(2025\)Z\. Li, X\. Feng, Y\. Cai, Z\. Zhang, T\. Liu, C\. Liang, W\. Chen, H\. Wang, and T\. ZhaoLLMs can generate a better answer by aggregating their own responses\.arXiv preprint arXiv:2503\.04104\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p1.1)\.
- Lianget al\.\(2024\)T\. Liang, Z\. He, W\. Jiao, X\. Wang, Y\. Wang, R\. Wang, Y\. Yang, Z\. Tu, and S\. ShiEncouraging divergent thinking in large language models through multi\-agent debate\.InProceedings of EMNLP,Cited by:[§1](https://arxiv.org/html/2609.36104#S1.p1.1),[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2023\)J\. Liu, C\. S\. Xia, Y\. Wang, and L\. ZhangIs your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation\.InProceedings of NeurIPS,Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px1.p1.1)\.
- Madaanet al\.\(2023\)A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang, S\. Welleck, B\. P\. Majumder, S\. Gupta, A\. Yazdanbakhsh, and P\. ClarkSelf\-refine: iterative refinement with self\-feedback\.InProceedings of NeurIPS,Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p1.1)\.
- Maryanskyyet al\.\(2026\)A\. Maryanskyy, D\. Budnikov, and A\. T\. KaliyevWhen agents disagree: the selection bottleneck in multi\-agent LLM pipelines\.Applied Sciences16\(10\),pp\. 4914\.External Links:[Document](https://dx.doi.org/10.3390/app16104914)Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p2.1)\.
- Mauryaet al\.\(2025\)K\. K\. Maurya, K\. A\. Srivatsa, and E\. KochmarSelectLLM: query\-aware efficient selection algorithm for large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 20847–20863\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p1.1)\.
- Qianet al\.\(2025\)C\. Qian, Z\. Xie, Y\. Wang, W\. Liu, K\. Zhu, H\. Xia, Y\. Dang, Z\. Du, W\. Chen, C\. Yang, Z\. Liu, and M\. SunScaling large language model\-based multi\-agent collaboration\.InProceedings of ICLR,Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px2.p1.1)\.
- Reinet al\.\(2024\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGPQA: a graduate\-level google\-proof q&a benchmark\.InProceedings of COLM,Cited by:[§4](https://arxiv.org/html/2609.36104#S4.SS0.SSS0.Px1.p1.1)\.
- Shenet al\.\(2025\)X\. Shen, Y\. Liu, Y\. Dai, Y\. Wang, R\. Miao, Y\. Tan, S\. Pan, and X\. WangUnderstanding the information propagation effects of communication topologies in LLM\-based multi\-agent systems\.InProceedings of EMNLP,Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p2.1)\.
- Tran and Kiela \(2026\)D\. Tran and D\. KielaSingle\-agent LLMs outperform multi\-agent systems on multi\-hop reasoning under equal thinking token budgets\.arXiv preprint arXiv:2604\.02460\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px4.p1.1)\.
- Venkatramanet al\.\(2025\)S\. Venkatraman, V\. Jain, S\. Mittal, V\. Shah, J\. Obando\-Ceron, Y\. Bengio, B\. R\. Bartoldson, B\. Kailkhura, G\. Lajoie, G\. Berseth, N\. Malkin, and M\. JainRecursive self\-aggregation unlocks deep thinking in large language models\.arXiv preprint arXiv:2509\.26626\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px3.p1.1)\.
- Wanget al\.\(2024\)J\. Wang, J\. Wang, B\. Athiwaratkun, C\. Zhang, and J\. ZouMixture\-of\-agents enhances large language model capabilities\.InProceedings of COLM,Cited by:[§1](https://arxiv.org/html/2609.36104#S1.p1.1),[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain\-of\-thought reasoning in language models\.InProceedings of ICLR,Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. V\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InProceedings of NeurIPS,Cited by:[1st item](https://arxiv.org/html/2609.36104#S3.I1.i1.p1.1)\.
- Yanget al\.\(2026\)Y\. Yang, C\. Qu, M\. Wen, L\. Shi, Y\. Wen, W\. Zhang, A\. Wierman, and S\. GuUnderstanding agent scaling in LLM\-based multi\-agent systems via diversity\.arXiv preprint arXiv:2602\.03794\.Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px2.p1.1)\.
- Yu \(2026\)G\. YuAdaptOrch: task\-adaptive multi\-agent orchestration in the era of LLM performance convergence\.arXiv preprint arXiv:2602\.16873\.Cited by:[§1](https://arxiv.org/html/2609.36104#S1.p2.1),[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p2.1)\.
- Yueet al\.\(2025\)Y\. Yue, G\. Zhang, B\. Liu, G\. Wan, K\. Wang, D\. Cheng, and Y\. QiMasRouter: learning to route LLMs for multi\-agent systems\.InProceedings of ACL,Cited by:[§1](https://arxiv.org/html/2609.36104#S1.p2.1),[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p2.1)\.
- Zhanget al\.\(2025\)G\. Zhang, L\. Niu, J\. Fang, K\. Wang, L\. Bai, and X\. WangMulti\-agent architecture search via agentic supernet\.InProceedings of ICML,Cited by:[§1](https://arxiv.org/html/2609.36104#S1.p2.1),[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p2.1)\.
- Zhenget al\.\(2024\)H\. S\. Zheng, S\. Mishra, X\. Chen, H\. Cheng, E\. H\. Chi, Q\. V\. Le, and D\. ZhouTake a step back: evoking reasoning via abstraction in large language models\.InProceedings of ICLR,Cited by:[1st item](https://arxiv.org/html/2609.36104#S3.I1.i1.p1.1)\.
- Zhouet al\.\(2023\)D\. Zhou, N\. Schärli, L\. Hou, J\. Wei, N\. Scales, X\. Wang, D\. Schuurmans, C\. Cui, O\. Bousquet, Q\. Le, and E\. ChiLeast\-to\-most prompting enables complex reasoning in large language models\.InProceedings of ICLR,Cited by:[1st item](https://arxiv.org/html/2609.36104#S3.I1.i1.p1.1)\.
- Zhuet al\.\(2025\)K\. Zhu, H\. Du, Z\. Hong, X\. Yang, S\. Guo, Z\. Wang, Z\. Wang, C\. Qian, X\. Tang, H\. Ji, and J\. YouMultiAgentBench: evaluating the collaboration and competition of LLM agents\.InProceedings of ACL,Cited by:[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p2.1)\.
- Zhugeet al\.\(2024\)M\. Zhuge, W\. Wang, L\. Kirsch, F\. Faccio, D\. Khizbullin, and J\. SchmidhuberLanguage agents as optimizable graphs\.InProceedings of ICML,Cited by:[§1](https://arxiv.org/html/2609.36104#S1.p1.1),[§2](https://arxiv.org/html/2609.36104#S2.SS0.SSS0.Px1.p1.1)\.
## Supplementary Material
This supplement documents the experimental and statistical protocol and provides the evidence underlying the main paper’s three claims: the returns to team scaling are sharply task\-dependent \(large on the arithmetic word\-problem benchmarks, small on the multiple\-choice ones, with Proposer\-Critic capturing the arithmetic gains\), an exact generate–transform decomposition separates proposal coverage from downstream conversion and explains the split, and no architecture wins across models and tasks\. It also reports the direct and long\-reasoning single\-agent controls, HumanEval\+ results, full\-mesh debate and no\-peer revision comparison, plurality control, and manager\-prompt stress test used to delimit those claims\. The audited cell file and all result tables are regenerated from raw logs by versioned analysis scripts\.
## Appendix AArchitecture Specification
All eight conditions are directed acyclic workflows and all terminate in one LLM manager\. The manager is an active reasoner instructed to inspect its parent reports and emit a final answer, not cast a deterministic vote\. We call the conditions*orchestration architectures*because graph and role prompt jointly define most of them\.
#### Star and Persona\-Star\.
Star hasN−1N\-1parent\-free workers\. Persona\-Star changes only those worker prompts and cycles six methods by worker index: forward solving, backward checking, decomposition, step\-back abstraction, conservative calibration, and contrarian search\. The node count, edges, sampling settings, manager prompt, and model remain fixed\. At least one instance of every persona is present fromN=7N=7onward\.
#### Proposer\-Critic and Tournament\.
Proposer\-Critic allocates\(N−1\)/2\(N\-1\)/2pairs where possible\. Each critic sees one proposal and is instructed to find errors before answering\. The manager sees critic outputs plus an unpaired proposal whenN−1N\-1is odd\. Tournament begins with the largest worker count whose full pairwise bracket fits insideNN\. Each duelist sees two previous reports and selects or repairs one answer\. The final duelist is relabeled as manager\. Requested budgets 10, 20, and 30 therefore use 9, 19, and 29 actual calls\. Other requested budgets in the sweep match exactly\.
#### Tree and serial chains\.
Tree uses branching factor three\. Complete triples feed synthesizers and any orphan worker feeds the manager directly\. Chain begins with one worker\. Each of itsN−2N\-2refiners sees only the immediately preceding report\. The manager sees the complete sequence\. Cascading\-Chain changes one knob: each refiner sees up to the five most recent reports, while its manager still sees the complete sequence\. Both chains therefore haveL=1L=1regardless ofNN\. Their additional calls transform one evolving solution rather than add parallel proposals\.
#### Diamond\.
Diamond allocates⌊N/5⌋\\lfloor N/5\\rfloorplanners \(at least one\)\. Planners are explicitly forbidden to emit a final answer and are excluded fromLL\. Remaining non\-manager calls are balanced across plans\. These plan\-conditioned solvers form the first answer\-producing tier\. The manager sees every solver\.
Table S1:Exact architecture composition at requested budgetN=30N=30\.LLis the first answer\-producing tier used for proposal metrics\. Tournament rounds down to the largest pairwise bracket within the requested budget\.
## Appendix BPrompts, Inference, and Scoring
Every QA prompt requires a final parse in the formFINAL: <answer\>and a confidence\. Workers solve independently\. Critics receive one parent and must identify flaws before committing\. Refiners receive their parent reports and are instructed to verify and correct them\. Synthesizers and managers receive labeled reports, compare reasoning, and solve the question before emitting one final answer\. Duelists compare two reports, planners propose distinct approaches without answering, and plan\-solvers receive one plan and execute it\. Parent reports are token\-truncated, not character\-truncated\. The verbatim templates for every role are reproduced in Appendix[J](https://arxiv.org/html/2609.36104#A10)\.
Sampling temperature is 0\.4 and top\-ppis 0\.95\. QA generations are capped at 1,024 tokens with 120 tokens retained per parent report\. HumanEval\+ generations are capped at 2,048 tokens with 512 tokens per parent report and a 24,576\-token context cap\. A request seed mixes base seed 42, run id, requested budget, topology, item id, tier, and node id\. This prevents identical prompts within a tier from sharing an RNG stream\. Nemotron uses/no\_think, and Qwen3 passesenable\_thinking=False\.
The auxiliary long\-reasoning control remains a single agent but uses the foundation runner’s extended prompt, which requests multiple approaches and verification, and raises the QA generation cap from 1,024 to 10,240 tokens\. Because the instruction and cap both change and decoding can terminate early, we do not interpret it as a pure token\-budget treatment\.
The manager\-prompt stress test keeps Diamond’sN=30N=30DAG, upstream prompts, model, sampling settings, and seed recipe fixed and changes only the final manager instruction\. One variant explicitly tallies candidate answers before deciding whether to follow or override the mode\. The other critiques each distinct answer before synthesizing\. Baseline and variants were launched as separate jobs\. Seeded decoding therefore yields closely matched but not byte\-identical proposal pools, so we analyze the paired realized runs rather than describe this as frozen\-evidence replay\.
The primary sweeps were launched as one\-model Slurm jobs on NVIDIA H100 80GB GPUs with eight CPU cores and 80GB of host memory per job\. The code package records the Python dependency bounds and the model\-specific vLLM launch flags\. Because the original environment was not frozen to exact package patch versions, this part of the computational record is necessarily partial\.
MCQ answers are canonicalized to option letters and numeric answers to integer strings before scoring\. HumanEval candidates are parsed as Python functions and executed against both base and HumanEval\+ tests\. A candidate must pass both suites\. Textual plurality and answer entropy are not used as code metrics because distinct correct functions are not fungible strings\.
## Appendix CStatistical Protocol
The analyzed primary sweep contains 4,334,715 team trials and 50,296,455 agent responses\. Proposal coverage and agreement statistics use exactly the architecture’s declared first answer\-producing tierLL\. Later critics, refiners, and managers are excluded\.
For a model–architecture–task–budget cell, each item’s three runs are averaged first\. Confidence intervals resample unique items with replacement for 1,000 replicates\. Persona\-Star comparisons are paired on shared\(run,item\)\(\\text\{run\},\\text\{item\}\)keys before runs are averaged within item\. Aggregate tables weight model–task cells equally rather than letting GSM8K dominate GPQA by item count\. No multiplicity correction is applied to individual intervals\. The principal Persona agreement result is the uniform sign and the fact that all 25 intervals exclude zero\.
For the aggregate Persona\-Star effects, we use 5,000 fixed\-grid stratified replicates\. Each replicate resamples item identities within benchmark and carries all five models and all runs for a sampled question together, thereby preserving cross\-model dependence on the same item\. The resulting 25 cell effects are weighted equally\. Common draws are used for coverage, loss, recovery, and final accuracy, so the transfer identity closes in every replicate\. These intervals quantify item uncertainty conditional on the tested model–task grid\. They do not treat five models or five tasks as random samples from wider populations\.
For the paired coverage\-transition analysis, every shared\(run,item\)\(\\text\{run\},\\text\{item\}\)realization is assigned to one of\(OStar,OPersona\)∈\{00,01,10,11\}\(O\_\{\\rm Star\},O\_\{\\rm Persona\}\)\\in\\\{00,01,10,11\\\}\. Its unconditional contribution is the stratum share times the paired final\-accuracy difference within that stratum\. The four contributions sum exactly to the overall Persona\-Star effect\. Aggregate intervals use the same benchmark\-stratified item bootstrap, carrying all runs and five models for a sampled item together\. The strata are observed stochastic realizations, not latent causal types\.
The conflict probe is restricted toN=30N=30Star and Persona\-Star trials withO=1O=1\. Within each model–task–architecture cell it compares proposal agreement between final\-answer loss \(Y=0Y=0\) and preservation \(Y=1Y=1\)\. Bootstrap draws resample items and keep all runs of a sampled item together\. This controls proposal availability within an architecture but not item difficulty, and the two architectures cover different item populations\. The probe is therefore explicitly associational and is not used to infer a cross\-system manager effect\.
For the Diamond manager\-prompt stress test, baseline and each variant are paired by model, task, run, and item\. We reconstruct proposal coverageOOfrom the first\-tier correctness vector and apply the same loss–recovery identity\. Aggregate intervals use 5,000 benchmark\-stratified item\-bootstrap replicates, carry a sampled item jointly across all five models, and weight the 25 cells equally\. We also report exact proposal\-vector and coverage\-status agreement across the separately decoded pairs to delimit the intervention\.
For the dense communication controls, round accuracies first average all available runs within item\. Debate and no\-peer Self are paired on their common round\-2/round\-3 item support\. 1,000 bootstrap replicates resample items\. Cumulative cost sums each item’s mean prompt\-plus\-output tokens over rounds 1–3 before taking the cell mean\.
The direct\-baseline audit is separate from the primary sweep\. It retains 9,695 hard\-QA rows across all 15 model–task cells \(one run per item\) and 2,460 HumanEval\+ rows \(three runs for each of five models\)\. The condition uses a single agent, one model call, and the ordinary task prompt without persona, peer, or architecture\-specific instructions\. Consequently it answers whether collaboration improves over one clean direct generation\. The separate 10×\\times\-cap condition provides a stronger single\-agent reference on 14 hard\-QA cells, but does not exactly match observed team tokens or isolate the effect of tokens from its extended\-reasoning instruction\.
The endpoint identity is checked by construction on every audited row\. In the notation of the main paper,
net transfer=r−ℓ\\displaystyle=r\-\\ell=P\(Y=1\)−P\(O=1\)\.\\displaystyle=P\(Y=1\)\-P\(O=1\)\.
#### Code and data availability\.
All prompt templates, workflow implementations, analysis code, and per\-cell result files \(cell statistics, decomposition terms with closure errors, transfer masses, and bootstrap intervals\) are released with the paper, and each table and figure is produced by a named script in the release\.
## Appendix DExact Two\-Margin Scaling Decomposition
Table[S2](https://arxiv.org/html/2609.36104#A4.T2)expands Figure 4\(a,b\) of the main paper\. We compute the midpoint decomposition separately in every model–task cell and average the signed terms only afterward\. Consequently, its closure is not an identity that holds only for an aggregate “representative” system: every one of the 200 cell rows satisfiesΔY=ℰ\+ℐ\\Delta Y=\\mathcal\{E\}\+\\mathcal\{I\}to numerical precision\. Item\-clustered 95% intervals \(1,000 replicates\) sharpen the sign counts: every width design’s extensive term excludes zero above and the five eroding designs’ intensive terms exclude zero below, while Proposer\-Critic’s intensive interval\[−0\.02,1\.05\]\[\-0\.02,1\.05\]straddles zero\. A per\-cell count \(111 of 150 significantly negative\) agrees but is uncorrected for multiplicity, so the equal\-cell intervals are the primary evidence\.
Table S2:Exact two\-margin decomposition of QA scaling from requestedN=3N=3toN=30N=30, in equal\-cell percentage points\. The intensive term is split into changes in covered\-case success and recovery\. Every row satisfiesΔY=ℰ\+ℐ\\Delta Y=\\mathcal\{E\}\+\\mathcal\{I\}before rounding\. The last three columns count signs across 25 model–task cells\.Table[S3](https://arxiv.org/html/2609.36104#A4.T3)tests whether the negative intensive term is a same\-item transform decline or a composition effect of the growing covered population\. We pair every trial acrossN=3N=3andN=30N=30on \(task, run, item\), stratify by the coverage transition\(O3,O30\)\(O\_\{3\},O\_\{30\}\), and read the stratum covered at both budgets, which holds item difficulty fixed and varies only the tier width the manager faces\. There final accuracy rises for Star, Persona\-Star, Proposer\-Critic, and Tree, is flat for Diamond, and falls only for Tournament, whose pairwise bracket can discard a correct finalist\. The aggregate fall in covered\-case success is therefore dominated by the newly\-covered stratum \(contributionc01c\_\{01\}\): harder items that convert weakly, mirroring the matched Star–Persona\-Star intervention \(Appendix[G\.1](https://arxiv.org/html/2609.36104#A7.SS1)\) rather than a manager that degrades on answers it already had\. The same pattern holds for theN=10→30N=10\\to 30contrast\.
Table S3:Paired coverage\-transition decomposition of QA scaling fromN=3N=3toN=30N=30, equal\-cell means over 25 model–task cells\. Each trial is paired across the two budgets by \(task, run, item\) and stratified by its coverage transition\(O3,O30\)\(O\_\{3\},O\_\{30\}\); match rate is 100%\.π11\\pi\_\{11\}is the share of trials covered at both budgets andπ01\\pi\_\{01\}the newly\-covered share \(percent\)\.Y311Y^\{11\}\_\{3\}andY3011Y^\{11\}\_\{30\}are final accuracies on the both\-covered stratum andΔY11\\Delta Y\_\{11\}their difference;c01c\_\{01\}andc11c\_\{11\}are the strata contributions to the totalΔY\\Delta Y\(points\)\. The both\-covered stratum holds item difficulty fixed, so a non\-negativeΔY11\\Delta Y\_\{11\}argues against a same\-item transform decline: accuracy there rises for the four proposal\-expanding designs Star, Persona\-Star, Proposer\-Critic, and Tree, is flat for Diamond, and falls only for Tournament\. The negative intensive term of Table[S2](https://arxiv.org/html/2609.36104#A4.T2)is therefore consistent with the weak conversion of harder newly\-covered items rather than the transform discarding answers it previously returned\.The intensive split also clarifies how the architectures differ\. For five width designs, both covered\-case success and recovery contribute negatively on average\. Proposer\-Critic combines a small negative covered\-case component with a positive recovery component\. The resulting mean intensive term is positive despite being negative in 18/25 individual cells\. Chain’s gain is almost entirely increased recovery, while Cascading\-Chain obtains smaller positive contributions from both intensive components\.
Table[S4](https://arxiv.org/html/2609.36104#A4.T4)gives the complete endpoint transfer quantities underlying Section 5\.3 of the main paper, including opportunity\-normalized discard and recovery\. Its columns are equal means of the 25 cellwise rates, so ratios of the displayed aggregate columns need not reproduce the displayed conditional rates\.
Table S4:Proposal\-to\-team transfer atN=30N=30on QA \(equal mean over 25 model–task cells\)\.OOandYYare proposal coverage and final accuracy\. Loss and recovery are unconditional masses needed by the accounting identity; the conditional columns measure downstream behavior given the corresponding opportunity\. Net is recovery minus loss\.Writingv=s−gv=s\-gexposes the response formY=g\+vOY=g\+vO\. SubtractingOOgivesY−O=g−\(1−v\)OY\-O=g\-\(1\-v\)O, hence the exact crossing pointO⋆=g/\(1−v\)=g/\(1−s\+g\)O^\{\\star\}=g/\(1\-v\)=g/\(1\-s\+g\)\. Table[S5](https://arxiv.org/html/2609.36104#A4.T5)instead pools the equal\-cell joint masses before calculating conditional rates\. This choice makes the displayed response law close exactly and therefore differs slightly from Table[S4](https://arxiv.org/html/2609.36104#A4.T4)\.
Table S5:Response\-law boundary atN=30N=30\. Rates pool the equal\-cell joint masses, soY=g\+\(s−g\)OY=g\+\(s\-g\)Oholds exactly for every row\. A generative transformer finishes above its proposal oracle exactly whenO<O⋆=g/\(1−s\+g\)O<O^\{\\star\}=g/\(1\-s\+g\)\. All values are percentages except proposal leveragev=s−gv=s\-g\.The boundary is diagnostic rather than a learned forecast, but which side of it a cell falls on can be estimated before the largest budget\. We label each cell by whetherO<O⋆O<O^\{\\star\}at an earlier budget and test that label atN=30N=30without using the intervening budgets\. Table[S6](https://arxiv.org/html/2609.36104#A4.T6)shows agreement rising from 76\.0% atB=3B=3to 93\.0% atB=7B=7and 96\.0% atB=10B=10\. This is held\-out\-budget persistence on the same model–task cells and benchmark items, not a deployable router\. The next subsection tests selection under model\- and item\-hold\-out directly\.
### D\.1Selecting an architecture under strict hold\-out
We test whether the two behaviors support a deployable selector under a model\- and item\-held\-out protocol: leave one model out, split each task’s questions into a probe half and a scoring half shared across models, fit every rule only on the training models’ probe half, and measure each rule’s accuracy gap to the hindsight best of the eight architectures on the held\-out model’s scoring items \(20 splits\)\. The strongest no\-probe baseline, deploying each task’s training\-best architecture, has a 1\.59\-point gap\. Single\-probe rules derived from the decomposition do not beat it: routing a Star probe by the net oracle gap, by coverage against a fittedO⋆O^\{\\star\}, or by both gives 1\.55, 4\.02, and 2\.46 points, the first a statistical tie and the other two significantly worse\. A leak\-free pilot over all eight architectures atN=10N=10, deploying the winner atN=30N=30, reaches 1\.22 points \(sd 0\.26\), below the baseline across the 20 paired splits \(pairedt=5\.2t=5\.2\)\. Restricting the pilot to one archetype per behavior \(Proposer\-Critic and Chain\) reaches 0\.94 \(sd 0\.16\) but selects that menu with hindsight\. The choice is therefore resolvable by a labeled pilot but not predictable from a cheap probe, consistent with a transform already near\-optimal given its proposals\.
Table S6:Held\-out\-budget stability of the oracle\-crossing label\. The labelO<O⋆O<O^\{\\star\}\(final accuracy above proposal coverage\) at calibration budgetBBpredicts the corresponding label atN=30N=30\. Cells tied at either endpoint are excluded\. Because the boundary exactly encodes the sign ofY−OY\-Oat each budget, this evaluates label persistence rather than an independently fitted classifier\.For architectures with a positive extensive term, the aggregate dividend realizationη=ΔY/ℰ\\eta=\\Delta Y/\\mathcal\{E\}is 33\.0% \(Star\), 38\.7% \(Persona\-Star\), 108\.1% \(Proposer\-Critic\), 25\.6% \(Tournament\), 53\.9% \(Tree\), and 35\.1% \(Diamond\)\. We leaveη\\etaundefined for the chains: their extensive terms are numerically near zero, so a ratio is unstable and their intensive behavior is already the informative description\.
Per\-cell decompositions with endpointO,s,gO,s,gvalues, closure errors, the Persona\-Star contrast, and pooled response parameters are released\.
## Appendix EPer\-model level accuracy
Tables[S7](https://arxiv.org/html/2609.36104#A5.T7)and[S8](https://arxiv.org/html/2609.36104#A5.T8)report team accuracy atN=30N=30for every architecture and model, split by task category for QA and given separately for code\. They carry the per\-model detail that the main paper’s gain figure averages over: no architecture wins across all five models in any category, and the within\-model architecture spread is large on arithmetic \(mean 14\.0 points\) but small on multiple\-choice \(mean 4\.3\) and on code\.
Table S7:Team accuracy \(%\) atN=30N=30, equal\-cell mean over each category’s tasks, by model\. On arithmetic the architecture spread is large \(mean 14 points\) and Proposer\-Critic leads four of five models, whereas on multiple\-choice every design is within a few points \(mean spread 4\)\.*Single*is the mean individual worker, one non\-thinking call\. Bold marks the best architecture within a model and category\.Table S8:HumanEval\+ pass rate \(%\) atN=30N=30by architecture and model, with the mean individual\-worker one\-call baseline \(Single\)\. Bold is the best architecture within a model\.#### Accuracy–cost frontier\.
Requested budget matches calls, not tokens\. AtN=30N=30the equal\-cell mean QA cost per problem ranges from 15\.1k tokens for Tournament to 31\.3k for Cascading\-Chain, whose accumulating context grows fastest, a 2\.1×\\timesspread at fixedNN\. Read against final accuracy \(theYYcolumn of Table[S4](https://arxiv.org/html/2609.36104#A4.T4)\), the non\-dominated accuracy–cost set is Tournament \(55\.4%, 15\.1k\), Star \(56\.4%, 15\.3k\), Tree \(57\.3%, 16\.0k\), and Proposer\-Critic \(59\.1%, 17\.1k\); Proposer\-Critic is both the most accurate design overall and the most accurate point on that frontier\. Persona\-Star, Chain, Diamond, and Cascading\-Chain are each dominated by a cheaper design at equal or higher accuracy\.
## Appendix FWithin\-Task Structure and Task\-Type Specialization
The main\-paper accuracies marginalize over items within a task\. Three cuts test whether that hides conclusion\-flipping structure\. It does not, and one surfaces a result the aggregates omit\.
*Task\-type specialization emerges with scale\.*Table[S9](https://arxiv.org/html/2609.36104#A6.T9)tracks the Proposer\-Critic arithmetic advantage against team size \(the accuracy trajectories are shown in the main paper\): below zero atN=3N=3, indistinguishable from zero throughN=7N=7, and significantly positive fromN=10N=10to\+5\.7\+5\.7points over the runner\-up atN=30N=30, where it beats all seven other architectures with intervals excluding zero\. Because task identity is free, this refines the task\-conditioned policy into an interpretable prior: on arithmetic, scale a Proposer\-Critic team, while on the multiple\-choice tasks the scaling gains are small for every design\.
Table S9:Task\-type specialization with scale on the two arithmetic word\-problem benchmarks \(GSM8K, GSMHard\), equal\-cell over 5 models×\\times2 tasks\. PC is Proposer\-Critic accuracy \(%\); runner\-up is the best non\-PC architecture at that budget \(chosen on the full sample\);Δ\\Deltais their difference \(points\) with an item\-clustered 95% bootstrap interval \(2,000 replicates resampling questions within benchmark\)\.∗/†mark intervals excluding zero above/below\. PC is significantly behind the field\-best atN=3N=3and significantly ahead fromN=10N=10; atN=30N=30it beats every other architecture \(all seven intervals exclude zero,\+5\.7\+5\.7to\+10\.2\+10\.2\)\.*The pattern is not a pooled\-subject artifact\.*Table[S10](https://arxiv.org/html/2609.36104#A6.T10)splits MMLU\-hard into its five constituent subjects: accuracies span only 54–62% and Chain leads in four of five, so the no\-universal\-winner picture holds within the task rather than arising from averaging heterogeneous subjects\.
Table S10:MMLU\-hard accuracy \(%\) atN=30N=30by subject and architecture \(pooled over 5 models\)\. The five hard subjects behave alike \(accuracy 54–62%\) and Chain leads in four of five, so the no\-universal\-winner picture is not an artifact of pooling subjects\.*Architecture spread does not grow with problem difficulty\.*Binning GSM8K by the number of calculator steps in its annotated solution, and GSMHard by the same step count recovered through a digit\-stripped content join to GSM8K \(94\.4% of items matched\), the spread across architectures widens with difficulty on GSM8K \(from about 6 to 20 points\) but is flat on GSMHard \(7–9 points at every level\)\. The GSM8K pattern is a ceiling effect, since its easy items saturate near 80%, rather than evidence that architecture matters more on harder problems\.
*Why the returns concentrate on arithmetic\.*The generate–transform decomposition locates the cause in its two margins \(Star proposal tier, equal\-cell over models,N=3→30N=3\\to 30\)\. ARC coverage is already saturated \(OOrises only from 89\.6 to 91\.2%\), so there is no extensive room\. GPQA and MMLU do add coverage \(OOup 13\.7 and 10\.0 points\), but the transform barely converts it \(ΔY=0\.0\\Delta Y=0\.0and\+1\.5\+1\.5\)\. Only on arithmetic does coverage both grow and convert \(OOup 23\.9 and 14\.8 points,ΔY=\+5\.4\\Delta Y=\+5\.4and\+4\.3\+4\.3for Star, larger for Proposer\-Critic\)\. The task\-averaged QA number superimposes these patterns, which is why it understates arithmetic and overstates the multiple\-choice tasks\.
## Appendix GControlled Intervention and Downstream Controls
### G\.1Persona\-Star intervention
The matched Star–Persona\-Star comparison is a controlled prompt\-only probe of the extensive margin, holding the graph and manager fixed and changing only the worker prompts\. Its two\-margin decomposition \(Figure[S1](https://arxiv.org/html/2609.36104#A7.F1)\(a\)\) readsΔY=ℰ\+ℐ\\Delta Y=\\mathcal\{E\}\+\\mathcal\{I\}withℰ=\+5\.08\\mathcal\{E\}=\+5\.08\(the 8\.13\-point coverage gain weighted by midpoint leverage\) andℐ=−3\.72\\mathcal\{I\}=\-3\.72\(−3\.50\-3\.50from covered\-case success and−0\.22\-0\.22from recovery\), giving\+1\.36\+1\.36points\. The paired coverage transitions \(Figure[S1](https://arxiv.org/html/2609.36104#A7.F1)\(b\)\) localize the shortfall exactly as theN=3→30N=3\\to 30budget scaling does: the 10\.53% of trials Persona newly covers gain only\+0\.94\+0\.94point, whereas on the 59\.14% both systems cover Persona is\+1\.19\+1\.19points*more*accurate \(95% CI\[0\.41,1\.98\]\[0\.41,1\.98\]\)\. Newly created availability, not degradation on shared items, is what converts weakly\.
Figure S1:Matched Star–Persona\-Star prompt intervention atN=30N=30\. \(a\) The intervention’s two\-margin decomposition: an extensive coverage dividend of\+5\.08\+5\.08points is mostly offset by a−3\.72\-3\.72intensive transformation change, leaving\+1\.36\+1\.36final points\. \(b\) Final accuracy split by which system has a correct proposal available \(neither, Persona\-Star only, Star only, or both\), with equal\-cell shares annotated\. The margin Persona newly covers converts weakly, and where both cover, Persona\-Star is slightly more accurate, so the aggregate intensive term is not a same\-item causal effect\.Table S11:Matched Star–Persona\-Star intervention atN=30N=30, averaged over five QA tasks\. The two agreement columns report exact\-answer pairwise agreement\. The remaining columns are Persona\-minus\-Star percentage\-point changes\.Table S12:Paired Persona\-Star minus Star effects atN=30N=30\. Brackets are item\-clustered 95% bootstrap intervals\.Table S13:Paired Persona\-Star minus Star effects atN=30N=30\. Sign counts summarize 25 item\-clustered model–task intervals; the aggregate interval uses 5,000 benchmark\-stratified item\-bootstrap replicates, carries each sampled item across five models, and weights cells equally\. The means closeΔY=ΔO\+Δr−Δℓ\\Delta Y=\\Delta O\+\\Delta r\-\\Delta\\ell\.Table S14:Paired coverage transitions for Persona\-Star versus Star atN=30N=30\. Each row is a realized\(OStar,OPersona\)\(O\_\{\\rm Star\},O\_\{\\rm Persona\}\)stratum\. ConditionalΔY\\Delta Ycompares final accuracy on the same paired trials\. Contribution is stratum share times conditionalΔY\\Delta Y, computed within each of the 25 model–task cells and then averaged, so it need not equal the product of the displayed aggregate columns\. It sums to the overall \+1\.36\-point effect\.Table S15:Proposal conflict and downstream loss atN=30N=30, restricted to trials with an available correct proposal \(O=1O=1\)\. Means weight the 25 model–task cells equally\. The agreement gap is preserved minus lost and is positive in every cell\.Table[S11](https://arxiv.org/html/2609.36104#A7.T11)summarizes the matched intervention by model\. Table[S12](https://arxiv.org/html/2609.36104#A7.T12)reports all 25 paired cells underlying the uniform agreement result\. Every agreement interval excludes zero in the negative direction, whereas only seven final\-accuracy intervals exclude zero positively and none exclude zero negatively\. Table[S13](https://arxiv.org/html/2609.36104#A7.T13)reports both cellwise heterogeneity and the fixed\-grid uncertainty of each equal\-cell headline effect\. The 8\.13\-point proposal\-coverage gain is positive in 23/25 cells, loss rises by 5\.85 points, and recovery changes by−0\.91\-0\.91\. Using unrounded means,ΔY=ΔO\+Δr−Δℓ=1\.36\\Delta Y=\\Delta O\+\\Delta r\-\\Delta\\ell=1\.36points\. Individual cell intervals are descriptive and are not treated as a multiplicity\-corrected family of tests\.
Table[S14](https://arxiv.org/html/2609.36104#A7.T14)separates composition from paired within\-stratum performance\. Persona newly creates coverage on 10\.53% of trials and loses it on only 2\.40%, yielding the 8\.13\-point net coverage increase\. Yet the newly covered stratum contributes only \+0\.94 point to final accuracy\. On the 59\.14% of trials covered by both systems, Persona is \+1\.19 points more accurate, not worse\. The larger marginal Persona discard rate therefore partly reflects a harder covered population\.
Table[S15](https://arxiv.org/html/2609.36104#A7.T15)conditions on correct\-proposal availability\. Lost cases have lower agreement in all 25 Star and all 25 Persona cells\. The item\-clustered interval for the agreement gap excludes zero in 24/25 Star and 25/25 Persona cells\. This within\-system association is compatible with conflict making aggregation harder, but item difficulty remains an uncontrolled common cause\. Unlike the paired transition table, it does not identify why the two systems’ marginal discard rates differ\.
### G\.2Downstream transformation controls
Table S16:Diamond manager\-prompt stress test atN=30N=30\. Each entry is the prompt variant minus the baseline manager in percentage points, with a 95% benchmark\-stratified item\-bootstrap interval\. The DAG, upstream instructions, model, sampling settings, and seed recipe are unchanged; separately launched decoding is not exact frozen\-evidence replay\.Table S17:Active final\-agent synthesis versus deterministic plurality over the same proposal tier atN=30N=30\(equal mean over QA cells\)\. Plurality is an offline control and invokes no additional LLM\.Table[S16](https://arxiv.org/html/2609.36104#A7.T16)tests two simple downstream interventions on Diamond\. Proposal coverage is stable: tally–decide changesOOby−0\.02\-0\.02points \(95% CI\[−0\.15,0\.11\]\[\-0\.15,0\.11\]\) and critique–synthesize by \+0\.08 \(\[−0\.10,0\.26\]\[\-0\.10,0\.26\]\)\. Nevertheless, final accuracy falls by 3\.18 and 1\.46 points, respectively\. Most of each decline is reduced recovery \(−2\.67\-2\.67and−1\.27\-1\.27points\), with only small loss increases \(\+0\.49 and \+0\.27\)\. Thus neither generic instruction repairs the measured bottleneck\. Explicitly emphasizing selection or critique can suppress constructive synthesis\. Across paired runs, coverage status agrees on 97\.7% and 96\.4% of trials, while the complete proposal\-answer vector agrees on 70\.8% and 60\.0%\. The stable aggregate coverage supports a downstream interpretation, but the lack of exact proposal replay prevents a stronger causal claim\.
Table[S17](https://arxiv.org/html/2609.36104#A7.T17)recomputes exact plurality from the logged proposal answers\. Ties use the same MD5\-derived item/run seed as the experiment\. For the serial designs plurality overL=1L=1is simply the initial worker\. Their final\-minus\-vote difference is therefore identical to net recovery from the initial proposal\. For all architectures, the control isolates what a no\-LLM vote would obtain from the same first answer\-producing tier\.
## Appendix HDirect Baseline and HumanEval\+ Robustness
Table S18:FixedN=30N=30sparse architectures versus the direct one\-call baseline on 15 hard\-QA model–task cells\. Accuracy and change are equal\-cell means in percentage points\. The final row is a descriptive post\-hoc upper envelope, not a deployable policy\.Table[S18](https://arxiv.org/html/2609.36104#A8.T18)compares each fixed architecture to a purpose\-built direct single\-agent baseline, stronger than the ordinary single agent of Table[S7](https://arxiv.org/html/2609.36104#A5.T7)\(it uses a cleaner prompt, and long reasoning for Nemotron\)\. Every architecture has a positive equal\-cell mean, but only 9–12 of 15 individual cells improve, so the like\-for\-like lift over the ordinary single agent shrinks against this stronger one\. The best\-observed row is selected after seeing the grid and is therefore an upper envelope, not a fair fixed\-policy estimate\.
Table S19:One\-agent 10×\\times\-cap long\-reasoning control on shared hard\-QA items\. The maximum generation length rises from 1,024 to 10,240 tokens and the prompt requests extended verification; hence this is a combined prompt\-and\-cap control, not a pure token intervention\.Δ\\Deltais Long minus Direct with a paired item\-bootstrap 95% CI\. No Ministral–GPQA long\-cap run is available\.Table[S19](https://arxiv.org/html/2609.36104#A8.T19)pairs the ordinary and long\-reasoning single\-agent runs on item id\. Equal weighting over its 14 model–task cells gives\+9\.18\+9\.18points \(fixed\-grid, benchmark\-stratified 95% CI\[8\.05,10\.35\]\[8\.05,10\.35\]\), while observed prompt\-plus\-output tokens rise from 528 to 1,995 per item\. On exactly these items, every fixedN=30N=30sparse architecture has lower equal\-cell mean accuracy than the long control\. The post\-hoc per\-cell sparse upper envelope is effectively tied \(−0\.16\-0\.16points, 9/14 cell wins\) while using 13\.5×\\timesas many observed tokens on average\. This is evidence that the ordinary direct baseline understates a stronger single\-agent alternative, not evidence that the generation cap alone causes the gain\.
Table S20:Exact two\-margin decomposition on the open\-ended HumanEval\+ task, equal\-cell means over the five models\.ℰ\\mathcal\{E\},ℐ\\mathcal\{I\}, andΔY\\Delta Yare the requestedN=3→30N=3\\to 30change in percentage points \(ΔY=ℰ\+ℐ\\Delta Y=\\mathcal\{E\}\+\\mathcal\{I\}exactly\)\.OO, generative recoveryg=P\(Y=1∣O=0\)g=P\(Y\{=\}1\\mid O\{=\}0\), andY−OY\-Oare atN=30N=30\. Coverage rises as on QA, butggnearly vanishes for the wide star\-family and every architecture finishes below its proposal oracle \(Y<OY<O\)\.Table[S20](https://arxiv.org/html/2609.36104#A8.T20)applies the exact two\-margin decomposition to HumanEval\+, using per\-candidate test execution for the proposal boundary\. The width structure matches QA: proposal\-expanding designs post a positive extensive term offset by a negative intensive term, and the chains add essentially no coverage\. What differs is downstream\. Generative recoveryggis 0–1% for the wide star\-family and at most 9% elsewhere, so no architecture finishes above its proposal oracle, unlike the QA chains\. On open\-ended output the transform stage seldom synthesizes a correct program absent from its pool, leaving final accuracy close to coverage minus loss\.
## Appendix IDense Debate and No\-Peer Revision
Table S21:Direct one\-call accuracy andN=30N=30three\-round controls\. In Self, each agent revises only its own prior answer; Debate exposes every agent to the other 29 reports\. Self R3 and both Debate columns use their common item support; the 95% item bootstrap pairs the two round\-3 outcomes\. Both cost ratios accumulate rounds 1–3\. All 15 hard\-QA cells are available\.Rows use requestedN=30N=30and the three hard\-QA tasks, with all 15 model–task cells complete\. Full\-mesh Debate exposes each agent to the other 29 reports\. In Self, every agent instead revises only its own preceding answer\. Both run for three rounds and aggregate the 30 answers by plurality\.
Debate round 3 beats the direct call by 8\.3 points on average and in 14/15 cells, but the post\-hoc sparse upper envelope is higher in 10/15\. Round 3 minus round 2 is positive/zero/negative in 11/1/3 cells\. Against the more diagnostic Self control, item\-paired Debate round 3 is 0\.69 point lower on average, is higher in only 5/15 cells, and consumes 2\.5–6\.7×\\timesas many cumulative tokens\. Peer exchange therefore does not explain the gross gain over one direct call uniformly\. The final cost column divides cumulative Debate tokens through rounds 1–3 by those of the highest\-accuracy observed sparse architecture\. Because that sparse choice is made after seeing the grid, it remains a descriptive upper envelope rather than a deployable router\.
#### Code, full budget sweep\.
On HumanEval\+, Debate and Self additionally run the full sweepN∈\{2,…,30\}N\\in\\\{2,\\dots,30\\\}for all five models, three rounds each, and the picture reverses in two ways\. First, one round of peer exchange captures Debate’s entire benefit: equal\-cell accuracy atN=30N=30is 73\.6% initial, 76\.1% after one round, and 76\.1% after two, so the second round adds nothing, unlike the 11/15 hard\-QA cells where it helped\. Second, peer content now helps: atN=30N=30Debate beats the no\-peer Self control by 2\.4 points, positive for all five models, the opposite of the hard\-QA result\. Even so, one\-round Debate \(60 calls\) only ties the best sparse architecture \(30 calls\), winning on Llama and Ministral and losing on Nemotron, Qwen2\.5, and Qwen3 \(mean\+1\.0\+1\.0point\)\. Dense peer exchange therefore helps more on executable code than on hard QA, but remains a roughly2×2\\times\-cost baseline that does not beat the sparse frontier\.
## Appendix JPrompt Templates
Every node receives a task\-specific template with a strict output contract\. Braced placeholders are filled at run time:\{question\}\(and\{choices\}for multiple choice\), the parent reports\{proposal\}/\{workers\}/\{reports\}, and the counts\{b\}/\{m\}/\{M\}\. Three output contracts appear, shown by the three worker templates below: multiple\-choice \(FINAL: <A, B, C, or D\>\), numeric \(FINAL: <integer\>\), and code \(onepythonfence preceded byCONF\)\. For a given role the multiple\-choice and numeric templates differ only in the task noun and the format block, so we reproduce the numeric template per role; the code template uses the code contract\. The exact set for all three modalities is in the released code\.
#### Worker, multiple\-choice contract\.
Answerthemultiple\-choicequestion\.Beconcise\.
Question:\{question\}
\{choices\}
FormatEXACTLY\(answerandconfidenceFIRST,thenreasoning\):
FINAL:<A,B,C,orD\>
CONF:<0\-100\>
RATIONALE:<briefreasoning,max4lines\>
#### Worker, numeric contract\.
Solvethismathproblemstepbystep\.Beconcise\.
Problem:\{question\}
FormatEXACTLY\(answerandconfidenceFIRST,thenreasoning\):
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<briefreasoning,max4lines\>
#### Worker, code contract\.
\{question\}
FormatEXACTLY:
CONF:<0\-100\>
SOLUTION:
‘‘‘python
<fullfunctiondefinitionhere\>
‘‘‘
#### Personas \(Persona\-Star\)\.
Each persona replaces only the worker’s leading directive; the format block is the worker’s\. The six numeric directives:
forward:SolvethismathproblembyworkingFORWARDfromthegivenvaluesstepbystep\.Beconcise\.
backward:SolvethismathproblembyworkingBACKWARDS:proposeaplausiblenumericalanswer,thencheckwhetheritsatisfieseveryconstraintintheproblem\.Adjustuntilconsistent\.Beconcise\.
decomposer:SolvethismathproblembyfirstDECOMPOSINGitintoasequenceofsimplersubproblems\.Solveeachsubprobleminorder;theanswertothelastyieldsthefinalanswer\.Beconcise\.
stepback:SolvethismathproblembyfirstSTEPPINGBACK:identifythegeneralmethod,formula,ortheoremthisproblemrequires,thenapplyittothespecificnumbers\.Beconcise\.
conservative:SolvethismathproblemCONSERVATIVELY:commitonlyafteratleasttwoindependentverificationpasses\(e\.g\.dimensionalcheck,sanitybound,alternativederivation\)agreeontheanswer\.ReportlowCONFwhenchecksdisagree\.Beconcise\.
contrarian:SolvethismathproblemasaCONTRARIAN\.Arriveatanobviousanswer,thendeliberatelyattempttofindanerrorinyourreasoning\.Ifyoufindarealflaw,revise;otherwisereporttheoriginalanswerandnotethatyourattemptedfalsificationfailed\.Beconcise\.
On code the six personas are recast as coding strategies \(leading line each\):
forward:Approach:implementthefunctionFORWARD—translatethespecdirectlyintocode,handlingeachrequirementintheorderitappearsinthesignatureanddocstring\.
backward:Approach:workBACKWARDfromtheexamples—figureouttheexpectedoutputforthedocstring’sexampleinputsfirst,thenwritecodethatreproducesthemandgeneralizestotherest\.
decomposer:Approach:DECOMPOSEthetaskinto2\-3smallersteps\(helpercomputationsorsub\-cases\),solveeach,thencomposethemintothefinalfunction\.
stepback:Approach:STEPBACKfirst—namethegeneralalgorithmordatastructurethiscallsfor\(sorting,hashing,acounter,twopointers,dynamicprogramming,\.\.\.\),thenimplementthatapproach\.
conservative:Approach:codeCONSERVATIVELY—handleedgecasesexplicitly\(emptyinput,zero,negatives,boundaries\),mentallyrunthedocstringexamplesbeforecommitting,andreportlowerCONFifanycaseisuncertain\.
contrarian:Approach:asaCONTRARIAN,writetheobviousimplementation,thendeliberatelyhuntfortheinputthatbreaksit\(off\-by\-one,emptycase,aliasing,overflow\)\.Ifyoufindone,fixit;otherwisekeepitandnotetheattackthatfailed\.
#### Critic \(Proposer\-Critic\)\.
YouareaCRITIC\.Aworkerproducedthesolutionbelow\.Findtheflaws—checkarithmetic,challengeassumptions,lookformissedcases\.Ifaftercritiquetheoriginalanswerstillholds,saysoexplicitlyandreportit;otherwisereporttheansweryourcritiquesupports\.
Problem:\{question\}
Workersolutiontocritique:
\{proposal\}
FormatEXACTLY\(yourownanswerafterthecritique\):
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<listflawsorconfirmsoundness,max4lines\>
#### Synthesizer \(Tree\)\.
YouareaSYNTHESIZER\.Youhavereceived\{b\}independentworkersolutionstothesameproblem\.Extractthemostconsistentcalculationchain\.Donotsimplyvote—verifythearithmeticoftheansweryoureport\.
Problem:\{question\}
Workersolutions\(independent,nocommunicationbetweenthem\):
\{workers\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<integratedreasoning,verifythearithmetic,max6lines\>
#### Refiner \(Chain, Cascading\-Chain\)\.
YouareaREFINER\.The\{b\}previousworker\(s\)belowsolvedthisprobleminsequence\(oldestfirst\)\.Readthem,thenproduceyourownsolution\.Youmayagree,disagree,orextend—butdonotjustcopy\.
Problem:\{question\}
Previousworkersolutions\(chronological\):
\{workers\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<yourrefinedreasoning,verifyarithmetic,max4lines\>
#### Duelist \(Tournament\)\.
YouareaJUDGEinatournamentbracket\.Youwillseesolutionsfrom\{b\}competitors\.Comparethemcritically—checkarithmetic,identifyreasoninggaps—thencommittoasingleanswer\.Ifbothagree,verify\.Iftheydisagree,pickthestrongersolutionandexplainwhy\.
Problem:\{question\}
Competitorsolutions:
\{workers\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<whichsolutionprevailsandwhy,max4lines\>
#### Duelist, bye \(Tournament, one parent\)\.
Onecontestantadvancedviabyeinthistournamentround;theirsolutionisbelow\.Yourjobistoverifytheirarithmeticandlogic,NOTpickbetweenalternatives\.Ifsound,reportthesameanswerwithyourownconfidence\.Ifyoufindaflaw,reporttheansweryourchecksupports\.
Problem:\{question\}
Contestantsolution:
\{workers\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<verificationresult;soundorflawed,withtheload\-bearingcheck,max4lines\>
#### Planner \(Diamond\)\.
YouarePLANNER\#\{i\_one\_indexed\}of\{M\}\.YourjobistodescribeONEmethodforsolvingthismathproblem—doNOTsolveit\.Theother\{M\_minus\_one\}plannersareindependentlyproducingtheirownmethods;yourmethodshouldbeDIFFERENTfromthemostobviousapproachasolverwoulddefaultto\.Keepitshortandspecific\(3\-5sentences\)\.DoNOToutputafinalnumber\.
Problem:\{question\}
FormatEXACTLY:
APPROACH:<nameofmethod,e\.g\.,’unitconversionthenratio’,’setupequationinx’,’workbackwardsfromtotal’\>
STEPS:<3\-5brief,actionablestepsasolvershouldtake,ononeortwolineseach\>
#### Plan\-solver \(Diamond\)\.
Aplannerhasproposedthefollowingmethodforthisproblem\.Executethemethodtoarriveattheanswer\.Youmaydeviateifyouidentifyaflaw—butstateexplicitlywhenyoudo\.
Problem:\{question\}
Plan:
\{workers\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<executedplan,noteanydeviations,max4lines\>
#### Manager, baseline \(all topologies\)\.
YouaretheFINALADJUDICATOR\.Youwillsee\{m\}downstreamreports\(workers,refiners,duelists,plan\-solvers,synthesizers,orcriticsdependingontheteamstructure\)\.YouareexplicitlyempoweredtoOVERRIDEthemajorityiftheirarithmeticiswrong\.Beskeptical\.Re\-deriveifneeded\.
Problem:\{question\}
Downstreamreports:
\{reports\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<audit\-stylereasoning,identifyanycomputationalerrors,max6lines\>
#### Manager stress\-test variants \(Appendix[G\.1](https://arxiv.org/html/2609.36104#A7.SS1)is the Persona probe; these back the Diamond manager\-prompt test of Appendix G\.2\)\.
Same input format as the baseline manager, changing only the synthesis instruction\.
#### Manager, tally\-decide\.
YouaretheFINALADJUDICATOR\.Youwillsee\{m\}downstreamreports\.YourdecisionMUSTfollowthistwo\-stepprocedure:
STEP1—TALLY:GroupthereportsbytheirFINALanswer\.Writetheexacttallyinyourrationale\(e\.g\.,"42:3,48:5,50:1"\)\.Identifythemodalanswer\.
STEP2—DECIDE:Ifthemodalanswerissupportedbysoundarithmetic,committoit\.ONLYOVERRIDEthemodalanswerifyoucanidentifyaspecificarithmeticorreasoningerrorinthesupportingreports\.
Problem:\{question\}
Downstreamreports:
\{reports\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<Step1tally;Step2decision\(modaloroverridewithnamederror\);max6lines\>
#### Manager, critique\-synthesize\.
YouaretheFINALADJUDICATOR\.Youwillsee\{m\}downstreamreports\.YourdecisionMUSTfollowthistwo\-stepprocedure:
STEP1—CRITIQUE:ForeachdistinctFINALanswerpresent,identifyONEarithmeticorreasoningweaknessinthesupportingreports\(or"checksout"\)\.
STEP2—SYNTHESIZE:Amongthesurviving\(uncritiqued\)candidateanswers,committoone\.Re\-deriveifnecessary;donotsplitthedifference\.
Problem:\{question\}
Downstreamreports:
\{reports\}
FormatEXACTLY:
FINAL:<integer\>
CONF:<0\-100\>
RATIONALE:<onecritiquelineperdistinctanswer;survivingdecision;max8lines\>相似文章
用 LLM 优化 LLM:面向测试时扩展的智能体发现方法
本文提出了 AutoTTS,这是一种环境驱动的框架,通过将测试时扩展(TTS)策略的发现过程形式化为控制器合成,自动发现用于大型语言模型(LLM)的测试时扩展策略。该框架在数学推理基准测试上展示了更优的准确率-成本权衡,且计算开销极小。
LLM的下一个重大‘扩展时代’是什么
作者回顾了过去LLM的扩展时代(模型规模、思维链、智能体),并推测下一个重大扩展维度,提出递归自我改进可能是一个突破。
Lean软件规模定律(阅读时间17分钟)
该研究提案探讨了不同编程语言中代码库大小如何影响编码LLM的困惑度,并以Lean作为形式语言的测试案例。它表明Lean可能具有更优的缩放指数,从而使大规模软件更安全、更可靠。
OrchSLM: 探析小型语言模型编排的动态特性
OrchSLM引入了一个路由框架,用于在非交互式智能体管道中编排小型语言模型,探究任务结构和模型组成等设计选择如何影响性能。
LLM智能体系统中技能的规模化定律
本文识别了LLM智能体系统中技能库的两个耦合规模化定律:路由准确率随库大小呈对数衰减,执行动态表现出救援效应。这些定律在15个模型和超过百万次决策中得到验证,且定律指导的优化显著提升了性能。