Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection
Summary
The paper proposes a heterogeneity-driven framework for selecting complementary LLM teams by profiling error decorrelation and predictive divergence, then greedily optimizing a quality–complementarity objective, showing gains over quality-only top-k baselines across multiple benchmarks.
View Cached Full Text
Cached at: 10/01/26, 09:43 AM
# Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection
Source: [https://arxiv.org/html/2609.38274](https://arxiv.org/html/2609.38274)
###### Abstract
The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter\-member error resonance and predictive differences\. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interpretable, and optimizable, leaving team composition to rely on heuristics\. We propose a heterogeneity\-driven team selection framework that performs offline profiling to characterize individual capability along with two complementary signals: one captures decorrelation in error patterns to reduce co\-failures, while the other measures divergence in predictive behavior to capture strategy diversity\. We formulate team selection as a standardized quality–complementarity combinatorial objective and apply an efficient greedy search to select a small team from a candidate pool\. Experiments across multiple benchmarks demonstrate that our framework consistently outperforms quality\-only baselines under controlled candidate pools and team sizes, establishing reusable selection principles for multi\-LLM systems\.
Fudan Institute on Networking Systems of AI
## 1Introduction
Large language models \(LLMs\) are increasingly deployed as*teams*: multiple models or agents collaborate through voting, debate, or orchestration to improve reasoning accuracy and robustness\([Du et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib4);[Liang et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib5);[Wang et al\., 2025a](https://arxiv.org/html/2609.38274#bib.bib10)\)\. In practical deployments, however, collaboration gains are tightly constrained by inference budget: each additional member introduces extra model calls and tokens, increasing end\-to\-end latency and cost\. Real systems must therefore keep teams small, drawing from a limited pool of deployable candidates\. These constraints make*member selection*a primary design decision: under a fixed budget, which models should we call so that the team is both individually strong and behaviorally complementary?
A natural starting point is to pick the top\-kkmodels by individual accuracy, but as illustrated in Figure[1](https://arxiv.org/html/2609.38274#S1.F1)\(a\), this strategy ignores a second axis that bounds team performance:*similarity across members*\. Models can exhibit error resonance, failing on the same instances and concentrating probability on the same incorrect options; even when they agree on a label, they may share highly similar predictive behavior \(confidence and preferences over alternatives\), offering little strategy diversity\. Top\-kk\-by\-accuracy selection therefore tends to yield a strong yet redundant team whose blind spots largely overlap\. These two sources of redundancy \(co\-failureandpredictive similarity\) motivate selection criteria that go beyond quality and capture inter\-model complementarity directly\.
Figure 1:Illustration of heterogeneity\-aware team selection\.While the value of heterogeneous teaming is widely recognized, practical team composition still relies largely on heuristics such as mixing model families\. Most recent work has instead focused on*downstream*combination mechanisms such as output fusion, reranking, routing, and cascading, all of which take a callable model set and budget as given\([Jiang et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib29);[Chen et al\., 2024b](https://arxiv.org/html/2609.38274#bib.bib35);[Shnitzer et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib36);[Ong et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib37)\)\. The*upstream*question of which models to select in the first place remains comparatively underexplored\. We address this gap by formulating team selection as an offline, budgeted optimization problem: composing a compact and complementary team from development data, independently of the inference\-time aggregation rule\. This requires complementarity signals that are computable, interpretable, and tractable to optimize, together with a selection procedure that scales to combinatorial search over candidate teams\.
Driven by these requirements, we proposeC2\-MAS\(Covariance\-minimized andCapability\-maximizedMulti\-AgentSystem\), an offline profiling\-and\-selection pipeline that directly targets the two sources of redundancy identified above\. Given development data, we profile each candidate to estimate its individual capability together with two complementary heterogeneity signals:\(1\) error decorrelation, which addresses co\-failure by capturing whether models fail on different instances; and\(2\) distributional divergence, which addresses predictive similarity by measuring disagreement in choice distributions over alternatives\. When multiple tasks are available, we average pairwise estimates across tasks to stabilize the complementarity signal\. We then select a size\-kkteam by optimizing a standardized quality–complementarity objective with an efficient multi\-start greedy search\.
Experiments on 13 multiple\-choice benchmarks with 11 open\-weight models show that heterogeneity signals yield consistent gains across aggregation rules \(Figure[1](https://arxiv.org/html/2609.38274#S1.F1)\(b\)\)\. On the 7 primary benchmarks, C2\-MAS reaches 75\.52% accuracy underStacking, improving over the strongest selection baseline \(Quality\-Only, 74\.72%\) by \+0\.80 pp*without regressing on any task*, and over the random\-k baseline by \+3\.58 pp\. On 6 additional held\-out benchmarks, the gains persist under distribution shift, suggesting that the proposed signals capture transferable complementarity rather than task\-specific artifacts\. Our key contributions are summarized as follows:
❶Problem Formulation:We frame team selection as an upstream, budget\-constrained optimization problem, decoupled from downstream aggregation mechanisms\.
❷Practical Solution:We propose two interpretable heterogeneity metrics and a standardized quality–complementarity objective with efficient greedy search\. Our analysis connects the two signals to collective errors and identifies conditions for accuracy gains over quality\-only selection\.
❸Empirical Validation:We demonstrate consistent, non\-regressive improvements across multiple benchmarks and aggregation rules, with gains that persist on held\-out tasks under distribution shift\.
## 2Related Work
### 2\.1Multi\-Agent and Multi\-Model Collaboration
LLM\-based multi\-agent systems organize collaboration through predefined roles and workflows\([Li et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib11);[Hong et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib12);[Wu et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib13)\), and refine interaction through debate and voting\([Du et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib4);[Choi et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib7)\), communication structure optimization, and dynamic team formation\([Zhuge et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib21);[Zhang et al\., 2025b](https://arxiv.org/html/2609.38274#bib.bib22);[Liu et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib26)\)\. Beyond agent interaction, multi\-model methods perform fusion at the output, reasoning, or decoding stage\([Jiang et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib29);[Wang et al\., 2025a](https://arxiv.org/html/2609.38274#bib.bib10);[Huang et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib9)\), while routing and cascading select models per input to balance quality and cost\([Chen et al\., 2024b](https://arxiv.org/html/2609.38274#bib.bib35);[Shnitzer et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib36);[Ong et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib37)\)\. Distinct from these approaches, we focus on task\-level offline team selection, constructing fixed\-size teams whose membership is determined independently of the subsequent aggregation rule\.
### 2\.2Ensemble Diversity and Team Selection
Classical ensemble theory emphasizes individual accuracy and low error correlation, motivating diversity measures and validation\-driven ensemble selection\([Hansen and Salamon, 1990](https://arxiv.org/html/2609.38274#bib.bib42);[Krogh and Vedelsby, 1994](https://arxiv.org/html/2609.38274#bib.bib43);[Kuncheva and Whitaker, 2003](https://arxiv.org/html/2609.38274#bib.bib44);[Windeatt, 2005](https://arxiv.org/html/2609.38274#bib.bib45);[Caruana et al\., 2004](https://arxiv.org/html/2609.38274#bib.bib46)\)\. Recent LLM ensemble selection methods exploit error or semantic diversity\([Tekin et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib74);[Cohen et al\., 2026](https://arxiv.org/html/2609.38274#bib.bib75)\), or adopt information\-theoretic and label\-based surrogate objectives\([Turkmen et al\., 2026](https://arxiv.org/html/2609.38274#bib.bib76);[Zhang et al\., 2026](https://arxiv.org/html/2609.38274#bib.bib77)\)\. Building on this line of work, we jointly quantify error correlation and distributional divergence through a single offline profiling pass, combine these signals with model quality for team selection, and evaluate their utility and generality across multiple aggregation rules\.
## 3Methodology
C2\-MAS selects compact LLM teams through offline profiling and team search, followed by inference\-time aggregation \(Figure[2](https://arxiv.org/html/2609.38274#S3.F2)\)\. For each task, development data provide quality and heterogeneity estimates \([Section3\.2](https://arxiv.org/html/2609.38274#S3.SS2)\) for the standardized selection objective \([Section3\.3](https://arxiv.org/html/2609.38274#S3.SS3)\); the selected team’s predictions are then combined by an aggregator \([Section3\.4](https://arxiv.org/html/2609.38274#S3.SS4)\)\.
### 3\.1Preliminaries
We consider a collection of tasks𝒯\\mathcal\{T\}\. Each taskt∈𝒯t\\in\\mathcal\{T\}has a label set𝒴t\\mathcal\{Y\}\_\{t\}\(e\.g\.,\{A,B,C,D\}\\\{A,B,C,D\\\}\) and two disjoint splits: a development setDtdevD\_\{t\}^\{\\mathrm\{dev\}\}and a held\-out test setDttestD\_\{t\}^\{\\mathrm\{test\}\}\. We assume a fixed pool ofmmcandidate modelsℳ=\{1,…,m\}\\mathcal\{M\}=\\\{1,\\ldots,m\\\}and select a teamSt⊆ℳS\_\{t\}\\subseteq\\mathcal\{M\}of sizekkfor each tasktt\. Throughout,DtdevD\_\{t\}^\{\\mathrm\{dev\}\}is used only for candidate profiling, complementarity estimation, and training an aggregator \(if needed\), whileDttestD\_\{t\}^\{\\mathrm\{test\}\}is used only for final evaluation\.
To obtain comparable choice distributions, we use a short prompt to elicit a single option from the model: “Complete the sentence ‘The correct answer is ’ by outputting exactly one letter from:<options\>’’\.111See[SectionD\.1](https://arxiv.org/html/2609.38274#A4.SS1)for full details\.Using the fixed probe prefix “The correct answer is ”, for each instancexxwe query the model’s log\-probabilityℓi\(y∣x\)\\ell\_\{i\}\(y\\mid x\)for each labely∈𝒴ty\\in\\mathcal\{Y\}\_\{t\}\(i\.e\., the next\-token log\-probability under this prefix\) and normalize within𝒴t\\mathcal\{Y\}\_\{t\}:
pit\(y∣x\)=exp\(ℓi\(y∣x\)\)∑y′∈𝒴texp\(ℓi\(y′∣x\)\)\.p\_\{i\}^\{t\}\(y\\mid x\)=\\frac\{\\exp\(\\ell\_\{i\}\(y\\mid x\)\)\}\{\\sum\_\{y^\{\\prime\}\\in\\mathcal\{Y\}\_\{t\}\}\\exp\(\\ell\_\{i\}\(y^\{\\prime\}\\mid x\)\)\}\.\(1\)By construction,pit\(⋅∣x\)p\_\{i\}^\{t\}\(\\cdot\\mid x\)is a proper distribution over𝒴t\\mathcal\{Y\}\_\{t\}\. Based on this distribution, modeliipredictsy^i\(x\)=argmaxy∈𝒴tpit\(y∣x\)\\hat\{y\}\_\{i\}\(x\)=\\argmax\_\{y\\in\\mathcal\{Y\}\_\{t\}\}p\_\{i\}^\{t\}\(y\\mid x\), and we define the correctness indicatorCi\(x\)=𝕀\[y^i\(x\)=y\(x\)\]C\_\{i\}\(x\)=\\mathbb\{I\}\[\\hat\{y\}\_\{i\}\(x\)=y\(x\)\], which equals11if the prediction is correct and00otherwise\. These definitions provide a unified interface for the subsequent quality estimation, complementarity signals, and team selection\.
Figure 2:Overview of C2\-MAS framework\. Given a candidate pool ofmmmodels, we \(1\) profile each model on development data to obtain quality estimates and choice distributions; \(2\) compute two heterogeneity index metrics \(HIerr\\mathrm\{HI\}\_\{err\}andHIdist\\mathrm\{HI\}\_\{dist\}\) and stabilize them via cross\-task pooling; \(3\) select a team of sizekkby optimizing a composite objective that balances quality and complementarity using multi\-start greedy search; and \(4\) aggregate the selected team’s predictions to produce final outputs, evaluated on held\-out test data\.
### 3\.2Profiling Signals
FromDtdevD\_\{t\}^\{\\mathrm\{dev\}\}, we estimate per\-model quality\{qit\}\\\{q\_\{i\}^\{t\}\\\}and two heterogeneity index \(HI\) metrics: error decorrelationHIerr\\mathrm\{HI\}\_\{err\}from an association matrixRtR^\{t\}, and distributional divergenceHIdist\\mathrm\{HI\}\_\{dist\}from a choice\-divergence matrixJtJ^\{t\}\. These signals serve as inputs to the subsequent team selection procedure\.
#### Quality\.
We measure the individual capability of modeliion taskttby its development\-set accuracy, denoted asqitq\_\{i\}^\{t\}, computed as:
qit=1\|Dtdev\|∑x∈DtdevCi\(x\)\.q\_\{i\}^\{t\}=\\frac\{1\}\{\|D\_\{t\}^\{\\mathrm\{dev\}\}\|\}\\sum\_\{x\\in D\_\{t\}^\{\\mathrm\{dev\}\}\}C\_\{i\}\(x\)\.\(2\)
#### Error Decorrelation \(HIerr\\mathrm\{HI\}\_\{err\}\)\.
To capture complementarity at the error level, we summarize whether two models tend to succeed and fail on the same instances\. For a pair\(i,j\)\(i,j\), we form a2×22\\times 2contingency table onDtdevD\_\{t\}^\{\\mathrm\{dev\}\}:
Nabt\(i,j\)=∑x∈Dtdev𝕀\[Ci\(x\)=a,Cj\(x\)=b\],N\_\{ab\}^\{t\}\(i,j\)=\\sum\_\{x\\in D\_\{t\}^\{\\mathrm\{dev\}\}\}\\mathbb\{I\}\\\!\\left\[C\_\{i\}\(x\)=a,\\;C\_\{j\}\(x\)=b\\right\],\(3\)wherea,b∈\{0,1\}a,b\\in\\\{0,1\\\}\. LetDijt=N11tN00t\+N10tN01tD\_\{ij\}^\{t\}=N\_\{11\}^\{t\}N\_\{00\}^\{t\}\+N\_\{10\}^\{t\}N\_\{01\}^\{t\}, we then compute Yule’sQQassociation statistic and define the pairwise error\-correlation matrixRt∈ℝm×mR^\{t\}\\in\\mathbb\{R\}^\{m\\times m\}as:
Rijt=\{N11tN00t−N10tN01tDijt,Dijt\>0,0,Dijt=0\.R\_\{ij\}^\{t\}=\\begin\{cases\}\\frac\{N\_\{11\}^\{t\}N\_\{00\}^\{t\}\-N\_\{10\}^\{t\}N\_\{01\}^\{t\}\}\{D\_\{ij\}^\{t\}\},&D\_\{ij\}^\{t\}\>0,\\\\ 0,&D\_\{ij\}^\{t\}=0\.\\end\{cases\}\(4\)where we writeNabt≡Nabt\(i,j\)N\_\{ab\}^\{t\}\\equiv N\_\{ab\}^\{t\}\(i,j\)for readability\. Given a teamSS, we aggregate pairwise correlations into an error\-decorrelation score:
HIerr\(S,Rt\)=1−1\(\|S\|2\)∑\{i,j\}⊆SRijt\.\\mathrm\{HI\}\_\{err\}\(S;R^\{t\}\)=1\-\\frac\{1\}\{\\binom\{\|S\|\}\{2\}\}\\sum\_\{\\\{i,j\\\}\\subseteq S\}R\_\{ij\}^\{t\}\.\(5\)HigherHIerr\\mathrm\{HI\}\_\{err\}indicates that team members make mistakes on different instances\.
#### Distributional Divergence \(HIdist\\mathrm\{HI\}\_\{dist\}\)\.
Correctness correlations omit how models distribute probability over answer choices\. We therefore measure pairwise disagreement directly in the choice distribution space using the Jensen–Shannon divergence \(JSD\)\([Lin, 1991](https://arxiv.org/html/2609.38274#bib.bib47);[Endres and Schindelin, 2003](https://arxiv.org/html/2609.38274#bib.bib48)\), computed with base\-2 logarithms\. For a pair\(i,j\)\(i,j\), we average per\-instance JSD onDtdevD\_\{t\}^\{\\mathrm\{dev\}\}to obtain a matrixJt∈ℝm×mJ^\{t\}\\in\\mathbb\{R\}^\{m\\times m\}:
Jijt=1\|Dtdev\|∑x∈DtdevJSD\(pit\(⋅∣x\),pjt\(⋅∣x\)\)\.J\_\{ij\}^\{t\}=\\frac\{1\}\{\|D\_\{t\}^\{\\mathrm\{dev\}\}\|\}\\sum\_\{x\\in D\_\{t\}^\{\\mathrm\{dev\}\}\}\\mathrm\{JSD\}\\\!\\left\(p\_\{i\}^\{t\}\(\\cdot\\mid x\),\\,p\_\{j\}^\{t\}\(\\cdot\\mid x\)\\right\)\.\(6\)Given a teamSS, we aggregate pairwise divergences into a distributional\-divergence score:
HIdist\(S,Jt\)=1\(\|S\|2\)∑\{i,j\}⊆SJijt\.\\mathrm\{HI\}\_\{dist\}\(S;J^\{t\}\)=\\frac\{1\}\{\\binom\{\|S\|\}\{2\}\}\\sum\_\{\\\{i,j\\\}\\subseteq S\}J\_\{ij\}^\{t\}\.\(7\)A largerHIdist\\mathrm\{HI\}\_\{dist\}indicates greater disagreement in the models’ choice distributions\.
#### Cross\-Task Pooling for Stable HI\.
Pairwise heterogeneity estimates can be noisy when\|Dtdev\|\|D\_\{t\}^\{\\mathrm\{dev\}\}\|is limited, which may make team selection overly sensitive to idiosyncrasies of a single task’s development split\. When a collection of tasks𝒯\\mathcal\{T\}is available, we pool pairwise heterogeneity across tasks to obtain a shared selection signal\. For each pairwise matrix, we compute the pooled estimates by averaging over tasks:
R¯=1\|𝒯\|∑u∈𝒯Ru,J¯=1\|𝒯\|∑u∈𝒯Ju\.\\bar\{R\}=\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{u\\in\\mathcal\{T\}\}R^\{u\},\\qquad\\bar\{J\}=\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{u\\in\\mathcal\{T\}\}J^\{u\}\.\(8\)In our experiments, we useR¯\\bar\{R\}andJ¯\\bar\{J\}as the heterogeneity signals for team selection\.
### 3\.3Heterogeneity\-Aware Team Selection
Given the development\-set profiling results, our goal is to select a teamSt⊆ℳS\_\{t\}\\subseteq\\mathcal\{M\}of sizekkthat balances individual capability and inter\-model complementarity\. We combine task\-specific quality estimates\{qit\}\\\{q\_\{i\}^\{t\}\\\}with the cross\-task pooled heterogeneity matricesR¯\\bar\{R\}andJ¯\\bar\{J\}, which summarize correctness dependence and differences in choice distributions, respectively\.
We formulate team selection as maximizing a composite score comprising average member quality, error decorrelationHIerr\(S,R¯\)\\mathrm\{HI\}\_\{err\}\(S;\\bar\{R\}\), and distributional divergenceHIdist\(S,J¯\)\\mathrm\{HI\}\_\{dist\}\(S;\\bar\{J\}\)\. Since these terms may have different scales across tasks and candidate pools, we standardize each term with a standard\-error\-scaledzz\-score before combining them\.
Letu\(S\)u\(S\)denote a team\-level quantity computed as an average \(e\.g\., the average quality,HIerr\(S,R¯\)\\mathrm\{HI\}\_\{err\}\(S;\\bar\{R\}\), orHIdist\(S,J¯\)\\mathrm\{HI\}\_\{dist\}\(S;\\bar\{J\}\)\), and letnu\(S\)n\_\{u\}\(S\)be the number of items being averaged \(e\.g\.,\|S\|\|S\|for member\-level quantities and\(\|S\|2\)\\binom\{\|S\|\}\{2\}for pairwise quantities\)\. We define:
z\(u\(S\)\)=u\(S\)−μuσu/nu\(S\),z\(u\(S\)\)=\\frac\{u\(S\)\-\\mu\_\{u\}\}\{\\sigma\_\{u\}/\\sqrt\{n\_\{u\}\(S\)\}\},\(9\)where\(μu,σu\)\(\\mu\_\{u\},\\sigma\_\{u\}\)are computed over the corresponding base items in the candidate pool: the set\{qit\}\\\{q\_\{i\}^\{t\}\\\}for quality, the set\{1−R¯ij\}\\\{1\-\\bar\{R\}\_\{ij\}\\\}for error decorrelation, and the set\{J¯ij\}\\\{\\bar\{J\}\_\{ij\}\\\}for distributional divergence\.
The C2\-MAS selection score for taskttcan then be formulated as:
scoret\(S\)=z\(1\|S\|∑i∈Sqit\)\+λ1z\(HIerr\(S,R¯\)\)\+λ2z\(HIdist\(S,J¯\)\),\\mathrm\{score\}\_\{t\}\(S\)=z\\\!\\left\(\\frac\{1\}\{\|S\|\}\\sum\_\{i\\in S\}q\_\{i\}^\{t\}\\right\)\+\\lambda\_\{1\}\\,z\\\!\\left\(\\mathrm\{HI\}\_\{err\}\(S;\\bar\{R\}\)\\right\)\+\\lambda\_\{2\}\\,z\\\!\\left\(\\mathrm\{HI\}\_\{dist\}\(S;\\bar\{J\}\)\\right\),\(10\)\{wrapfloat\}
algorithmr0\.55Multi\-start forward greedy selection for tasktt
1:Input:Candidates
ℳ\\mathcal\{M\}, team size
kk, qualities
\{qit\}\\\{q\_\{i\}^\{t\}\\\}, pooled matrices
R¯,J¯\\bar\{R\},\\bar\{J\}, weights
\(λ1,λ2\)\(\\lambda\_\{1\},\\lambda\_\{2\}\)
2:Output:Selected team
St∗S\_\{t\}^\{\*\}
3:
St∗←∅S\_\{t\}^\{\*\}\\leftarrow\\emptyset,
best←−∞\\mathrm\{best\}\\leftarrow\-\\infty
4:foreach seed
s∈ℳs\\in\\mathcal\{M\}do
7:
j∗←argmaxj∈ℳ∖Sscoret\(S∪\{j\}\)j^\{\*\}\\leftarrow\\argmax\_\{j\\in\\mathcal\{M\}\\setminus S\}\\;\\mathrm\{score\}\_\{t\}\(S\\cup\\\{j\\\}\)
8:
S←S∪\{j∗\}S\\leftarrow S\\cup\\\{j^\{\*\}\\\}
9:endwhile
10:if
scoret\(S\)\>best\\mathrm\{score\}\_\{t\}\(S\)\>\\mathrm\{best\}then
11:
best←scoret\(S\)\\mathrm\{best\}\\leftarrow\\mathrm\{score\}\_\{t\}\(S\),
St∗←SS\_\{t\}^\{\*\}\\leftarrow S
12:endif
13:endfor
whereλ1,λ2≥0\\lambda\_\{1\},\\lambda\_\{2\}\\geq 0are trade\-off coefficients forHIerr\\mathrm\{HI\}\_\{err\}andHIdist\\mathrm\{HI\}\_\{dist\}\. The selection objective ismax\|S\|=kscoret\(S\)\\max\_\{\|S\|=k\}\\mathrm\{score\}\_\{t\}\(S\)\.
To efficiently solve this combinatorial optimization problem, we adopt a multi\-start forward greedy procedure: starting from a seed model, we iteratively add the candidate that yields the largest score improvement until reaching sizekk, and we repeat this process for each seed inℳ\\mathcal\{M\}, returning the best team found\. The full procedure is summarized inAlgorithm[3\.3](https://arxiv.org/html/2609.38274#S3.SS3)\.
### 3\.4Aggregation
Our method is agnostic to the downstream aggregator, as team selection relies solely on capability and complementarity statistics obtained through offline profiling\. To examine whether the gains from team selection persist across aggregation mechanisms, we evaluate the same selected teams usingChoice\-Soft,PoE,DS, andStacking\.Choice\-SoftandPoEuse fixed rules based on the mean and normalized product of member probabilities, respectively, whereasDSandStackinglearn aggregation parameters from development data to account for members’ predictive characteristics and reliability differences\. Full definitions and training details are provided in[SectionD\.5](https://arxiv.org/html/2609.38274#A4.SS5)\.
### 3\.5Theoretical Insights
We analyze how correctness dependence and choice\-distribution divergence jointly affect team accuracy, and when their benefits compensate for a reduction in individual model quality\.
#### A choice\-probability model\.
Fix a task distribution and three\-member teams on four choices\. With the true answer listed first, a correct member outputspc=\(a,u,u,u\)p^\{\\rm c\}=\(a,u,u,u\); an incorrect member outputspw=\(a,b,c,c\)p^\{w\}=\(a,b,c,c\)with the last three coordinates permuted to placebbon a wrong answerww\. Assume
0<c<a<b,a\+b\+2c=1,u=\(b\+2c\)/3,a\>\(2b\+u\)/3\.0<c<a<b,\\quad a\+b\+2c=1,\\quad u=\(b\+2c\)/3,\\quad a\>\(2b\+u\)/3\.\(11\)Members’ correctness and preferred wrong answers may be arbitrarily dependent\. For a teamSS, writeq¯S=13∑i∈Sℙ\(Ci=1\)\\bar\{q\}\_\{S\}=\\frac\{1\}\{3\}\\sum\_\{i\\in S\}\\mathbb\{P\}\(C\_\{i\}=1\),eS=1−q¯Se\_\{S\}=1\-\\bar\{q\}\_\{S\}, and letDSD\_\{S\}be the mean pairwise probability of both members being incorrect\. LetKSK\_\{S\}be the mean probability of both being incorrect*and preferring the same wrong answer*, and letJSJ\_\{S\}be their mean JSD on this task\. Defined=JSD\(pw,pw′\)d=\\mathrm\{JSD\}\(p^\{w\},p^\{w^\{\\prime\}\}\)forw≠w′w\\neq w^\{\\prime\},h=JSD\(pc,pw\)h=\\mathrm\{JSD\}\(p^\{\\rm c\},p^\{w\}\), andθ=2h/d\\theta=2h/d\.
###### Proposition 3\.1\(Shared errors and wrong\-answer concentration\)\.
Under \([11](https://arxiv.org/html/2609.38274#S3.E11)\),0<θ<10<\\theta<1and
KS=θeS\+\(1−θ\)DS−JSd\.K\_\{S\}=\\theta e\_\{S\}\+\(1\-\\theta\)D\_\{S\}\-\\frac\{J\_\{S\}\}\{d\}\.\(12\)For eitherChoice\-SoftorPoE,ℙ\(y^S≠Y\)≤KS\\mathbb\{P\}\(\\hat\{y\}\_\{S\}\\neq Y\)\\leq K\_\{S\}\. Both aggregators fail exactly when all three members prefer the same wrong answer\.
#### Accuracy gains over quality\-only selection\.
LetSQS\_\{Q\}be the size\-three Quality\-Only team, and abbreviate its quantities by a subscriptQQ\. Define the baseline correctionεQ=min\{KQ,eQ−\(DQ\+KQ\)/2\}≥0\\varepsilon\_\{Q\}=\\min\\\{K\_\{Q\},\\,e\_\{Q\}\-\(D\_\{Q\}\+K\_\{Q\}\)/2\\\}\\geq 0\.
###### Proposition 3\.2\(A sufficient condition for improvement\)\.
Under \([11](https://arxiv.org/html/2609.38274#S3.E11)\), the teamSSreturned by Algorithm[3\.3](https://arxiv.org/html/2609.38274#S3.SS3), evaluated with the sameChoice\-SoftorPoEaggregator asSQS\_\{Q\}, satisfies
Acc\(S\)−Acc\(SQ\)≥\(1−θ\)\(DQ−DS\)\+JS−JQd−θ\(q¯Q−q¯S\)−εQ\.\\operatorname\{Acc\}\(S\)\-\\operatorname\{Acc\}\(S\_\{Q\}\)\\geq\(1\-\\theta\)\(D\_\{Q\}\-D\_\{S\}\)\+\\frac\{J\_\{S\}\-J\_\{Q\}\}\{d\}\-\\theta\(\\bar\{q\}\_\{Q\}\-\\bar\{q\}\_\{S\}\)\-\\varepsilon\_\{Q\}\.\(13\)A positive right\-hand side guarantees strictly higher team accuracy\.
## 4Experiments
We evaluate C2\-MAS on various benchmarks to address the following research questions:
RQ1 \(Effectiveness and Generalization\)\.Does heterogeneity\-aware team selection yield consistent gains over quality\-only baselines, both on the primary task set and on out\-of\-distribution \(OOD\) benchmarks?
RQ2 \(Mechanism\)\.How do error decorrelation \(HIerr\\mathrm\{HI\}\_\{err\}\) and distributional divergence \(HIdist\\mathrm\{HI\}\_\{dist\}\) individually contribute to team performance, and what underlies these gains?
RQ3 \(Robustness\)\.Is C2\-MAS robust to the choice of heterogeneity weights\(λ1,λ2\)\(\\lambda\_\{1\},\\lambda\_\{2\}\)and to limited profiling data?
### 4\.1Setup
#### Tasks and Benchmarks\.
We evaluate C2\-MAS on 7 primary datasets and test its OOD generalization on 6 held\-out datasets\. The 13 public multiple\-choice benchmarks span four domains: \(1\)*math reasoning*, AQUA\-RAT\([Ling et al\., 2017](https://arxiv.org/html/2609.38274#bib.bib49)\)and MathQA\([Amini et al\., 2019](https://arxiv.org/html/2609.38274#bib.bib50)\); \(2\)*natural science*, ARC\-Challenge\([Clark et al\., 2018](https://arxiv.org/html/2609.38274#bib.bib51)\)and OpenBookQA\([Mihaylov et al\., 2018](https://arxiv.org/html/2609.38274#bib.bib52)\); \(3\)*reading comprehension and logical reasoning*, RACE\([Lai et al\., 2017](https://arxiv.org/html/2609.38274#bib.bib53)\), ReClor\([Yu et al\., 2020](https://arxiv.org/html/2609.38274#bib.bib54)\), and LogiQA2\([Liu et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib55)\); and \(4\)*general or domain knowledge*, MMLU\([Hendrycks et al\., 2021](https://arxiv.org/html/2609.38274#bib.bib56)\), MMLU\-Pro\([Wang et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib57)\), C\-EVAL\([Huang et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib58)\), CommonsenseQA\([Talmor et al\., 2019](https://arxiv.org/html/2609.38274#bib.bib59)\), MedQA\([Jin et al\., 2021](https://arxiv.org/html/2609.38274#bib.bib60)\), and StrategyQA\([Geva et al\., 2021](https://arxiv.org/html/2609.38274#bib.bib61)\)\. For efficiency, we randomly subsample both the dev and test splits to 1,068 examples each\.
#### Models\.
We usem=11m=11open\-weight chat models in the 7–9B parameter class for the candidate pool\. The pool includes Llama\-3\.1\-8B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib62)\), Qwen3\-8B\([Yang et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib63)\), Gemma\-2\-9B\-IT\([Team et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib64)\), Ministral\-3\-8B\-Instruct\([Liu et al\., 2026](https://arxiv.org/html/2609.38274#bib.bib65)\), Granite\-3\.3\-8B\-Instruct\([IBM Granite Team, 2024](https://arxiv.org/html/2609.38274#bib.bib66)\), GLM\-4\-9B\-Chat\([GLM et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib67)\), Nemotron\-Nano\-9B\-v2\([Basant et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib68)\), EuroLLM\-9B\-Instruct\([Martins et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib69)\), OLMo\-3\-7B\-Instruct\([Olmo et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib70)\), Apertus\-8B\-Instruct\([Apertus et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib71)\), and InternLM3\-8B\-Instruct\([Cai et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib72)\)\.
#### Baselines\.
We compare: \(1\)*Self\-Consistency*\([Wang et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib73)\): the best single model on the dev set sampledk=3k=3times with stochastic decoding and aggregated by majority vote; \(2\)*Random\-k*: uniformly samplek=3k=3models without replacement and report mean accuracy over 100 samples; \(3\)*Caruana*\([Caruana et al\., 2004](https://arxiv.org/html/2609.38274#bib.bib46)\): forward ensemble selection onDtdevD\_\{t\}^\{\\mathrm\{dev\}\}with replacement \(allow repeats to express non\-uniform member frequencies\); \(4\)*Quality\-Only*: selectk=3k=3by maximizing mean development accuracy \(i\.e\.,λ1=λ2=0\\lambda\_\{1\}=\\lambda\_\{2\}=0\); and \(5\)*C2\-MAS*: our heterogeneity\-aware selection withk=3k=3\.
#### Implementation details\.
We select teams by maximizing the standardized objective in[Equation10](https://arxiv.org/html/2609.38274#S3.E10)with weights\(λ1,λ2\)=\(0\.13,0\.05\)\(\\lambda\_\{1\},\\lambda\_\{2\}\)=\(0\.13,0\.05\)\. For aggregation, DS uses additive smoothing withε=10−3\\varepsilon=10^\{\-3\}when estimating the class prior and confusion matrices; Stacking trains a linear combiner onDtdevD\_\{t\}^\{\\mathrm\{dev\}\}and evaluates onDttestD\_\{t\}^\{\\mathrm\{test\}\}\(L2L\_\{2\}coefficientβ=0\.003\\beta=0\.003\); the other aggregators are parameter\-free\. For statistical testing, we use paired bootstrap with 2,000 resamples to compute confidence intervals \(CI\) for accuracy differences on the test set \(paired by instance\)\.
Table 1:Test accuracy \(%\) on the 7 primary benchmarks; teams of sizek=3k=3are selected fromm=11m=11candidates\.↑\\uparrow: gain over Random\-kk\(pp\);bold: best within each aggregation block\. Held\-out and full results:[Tables5](https://arxiv.org/html/2609.38274#A5.T5)and[E\.1](https://arxiv.org/html/2609.38274#A5.SS1)\. ARC\-C: ARC\-Challenge; CSQA: CommonsenseQA; OBQA: OpenBookQA\.MethodDatasetAvg\.ARC\-CCSQALogiQA2MedQAMMLUMMLU\-ProOBQAClosed\-Source Models \(for reference\)Gemini\-2\.5\-Flash92\.7092\.7081\.6581\.6560\.9660\.9677\.4377\.4381\.3781\.3758\.6158\.6192\.7092\.7077\.9277\.92GPT\-4o94\.1094\.1083\.3383\.3364\.9864\.9884\.7484\.7484\.6484\.6450\.1950\.1991\.8591\.8579\.1279\.12Single\-Model AggregationSelf\-Consistency89\.1489\.1478\.6578\.6567\.3267\.3261\.7061\.7073\.4173\.4144\.7644\.7688\.5888\.5871\.9471\.94Team Selection BaselinesChoice\-SoftRandom\-kk87\.6187\.6178\.8778\.8762\.4962\.4960\.5060\.5072\.0872\.0843\.1543\.1585\.0885\.0869\.9769\.97Caruana89\.9889\.9881\.6581\.6570\.5170\.5163\.1163\.1173\.6973\.6948\.2248\.2290\.54\\mathbf\{90\.54\}73\.9673\.96Quality\-Only91\.48\\mathbf\{91\.48\}81\.4681\.4670\.88\\mathbf\{70\.88\}66\.39\\mathbf\{66\.39\}76\.59\\mathbf\{76\.59\}47\.0047\.0089\.7089\.7074\.7974\.79C2\-MAS \(Ours\)91\.48\\mathbf\{91\.48\}↑\\uparrow3\.8783\.90\\mathbf\{83\.90\}↑\\uparrow5\.0270\.88\\mathbf\{70\.88\}↑\\uparrow8\.3966\.39\\mathbf\{66\.39\}↑\\uparrow5\.8976\.59\\mathbf\{76\.59\}↑\\uparrow4\.5148\.50\\mathbf\{48\.50\}↑\\uparrow5\.3590\.4590\.45↑\\uparrow5\.3775\.45\\mathbf\{75\.45\}↑\\uparrow5\.48PoERandom\-kk87\.5887\.5879\.2079\.2063\.0463\.0460\.8160\.8172\.3072\.3043\.4143\.4184\.7984\.7970\.1670\.16Caruana90\.5490\.5483\.0583\.0571\.44\\mathbf\{71\.44\}65\.0765\.0774\.0674\.0649\.34\\mathbf\{49\.34\}90\.82\\mathbf\{90\.82\}74\.9174\.91Quality\-Only91\.57\\mathbf\{91\.57\}82\.6882\.6870\.6970\.6965\.36\\mathbf\{65\.36\}76\.69\\mathbf\{76\.69\}48\.6948\.6989\.3389\.3375\.0075\.00C2\-MAS \(Ours\)91\.57\\mathbf\{91\.57\}↑\\uparrow3\.9983\.90\\mathbf\{83\.90\}↑\\uparrow4\.7070\.6970\.69↑\\uparrow7\.6565\.36\\mathbf\{65\.36\}↑\\uparrow4\.5576\.69\\mathbf\{76\.69\}↑\\uparrow4\.3948\.9748\.97↑\\uparrow5\.5689\.5189\.51↑\\uparrow4\.7275\.24\\mathbf\{75\.24\}↑\\uparrow5\.08DSRandom\-kk87\.5687\.5678\.0078\.0062\.9062\.9060\.8760\.8771\.5271\.5242\.4642\.4684\.9684\.9669\.7569\.75Caruana86\.2486\.2479\.4979\.4968\.8268\.8263\.9563\.9572\.9472\.9447\.47\\mathbf\{47\.47\}89\.6189\.6172\.6572\.65Quality\-Only90\.92\\mathbf\{90\.92\}81\.1881\.1870\.32\\mathbf\{70\.32\}65\.17\\mathbf\{65\.17\}75\.47\\mathbf\{75\.47\}47\.2847\.2889\.5189\.5174\.2674\.26C2\-MAS \(Ours\)90\.92\\mathbf\{90\.92\}↑\\uparrow3\.3682\.87\\mathbf\{82\.87\}↑\\uparrow4\.8770\.32\\mathbf\{70\.32\}↑\\uparrow7\.4265\.17\\mathbf\{65\.17\}↑\\uparrow4\.3075\.47\\mathbf\{75\.47\}↑\\uparrow3\.9547\.1047\.10↑\\uparrow4\.6490\.45\\mathbf\{90\.45\}↑\\uparrow5\.4974\.61\\mathbf\{74\.61\}↑\\uparrow4\.86StackingRandom\-kk89\.3089\.3080\.2080\.2065\.5465\.5462\.8262\.8273\.5573\.5544\.8244\.8287\.3587\.3571\.9471\.94Caruana90\.5490\.5482\.7782\.7770\.60\\mathbf\{70\.60\}64\.4264\.4273\.8873\.8848\.7848\.7891\.57\\mathbf\{91\.57\}74\.6574\.65Quality\-Only91\.67\\mathbf\{91\.67\}81\.7481\.7470\.60\\mathbf\{70\.60\}65\.54\\mathbf\{65\.54\}76\.12\\mathbf\{76\.12\}47\.1947\.1990\.1790\.1774\.7274\.72C2\-MAS \(Ours\)91\.67\\mathbf\{91\.67\}↑\\uparrow2\.3784\.18\\mathbf\{84\.18\}↑\\uparrow3\.9870\.60\\mathbf\{70\.60\}↑\\uparrow5\.0665\.54\\mathbf\{65\.54\}↑\\uparrow2\.7276\.12\\mathbf\{76\.12\}↑\\uparrow2\.5748\.97\\mathbf\{48\.97\}↑\\uparrow4\.1591\.57\\mathbf\{91\.57\}↑\\uparrow4\.2275\.52\\mathbf\{75\.52\}↑\\uparrow3\.58
### 4\.2Effectiveness of C2\-MAS \(RQ1\)
We report primary results on 7 datasets in Table[1](https://arxiv.org/html/2609.38274#S4.T1)and OOD generalization on the remaining 6 in[Table5](https://arxiv.org/html/2609.38274#A5.T5)\.
#### Primary Results\.
Under strict development\-test separation, C2\-MAS delivers small but consistent improvements over the Quality\-Only baseline\. As shown in Table[1](https://arxiv.org/html/2609.38274#S4.T1), C2\-MAS reaches 75\.52% with theStackingaggregator, outperforming Quality\-Only \(74\.72%\) by\+0\.80\+0\.80pp \(95% CI: \[0\.44, 1\.16\]\), while also surpassing both Self\-Consistency and Random\-kkby\+3\.58\+3\.58pp\.
#### OOD Generalization\.
To assess robustness under distribution shifts, we evaluate C2\-MAS on 6 held\-out benchmarks\. C2\-MAS achieves 66\.74% average accuracy, surpassing Quality\-Only \(66\.43%, \+0\.31 pp\) and significantly outperforming Self\-Consistency \(64\.45%,\+2\.29\+2\.29pp\) and Random\-kk\(61\.12%,\+5\.62\+5\.62pp\), confirming that heterogeneity signals capture intrinsic complementarity that transfers beyond the primary task set\.
#### Non\-regression and Consistency\.
Beyond average gains, safety is critical for practical deployment\. UnderStacking, the wins come from the three benchmarks where C2\-MAS selects a different team from Quality\-Only \(CommonsenseQA, MMLU\-Pro, and OpenBookQA\), with a mean gain of\+1\.87\+1\.87pp\. On the remaining four benchmarks the teams are identical and predictions match Quality\-Only by construction \([SectionE\.2](https://arxiv.org/html/2609.38274#A5.SS2)\)\. In each differing case, C2\-MAS replaces a redundant high\-accuracy member with a more complementary one rather than trading quality for diversity\. C2\-MAS thus acts as a*safe plugin*: it preserves the strongest baseline’s performance and adds gains where complementarity helps\. The pattern holds across all four aggregators, with the largest margin under Stacking where the learned combiner benefits most from heterogeneous inputs\.
### 4\.3Mechanism Analysis: Why Heterogeneity Matters? \(RQ2\)
While the main results confirm the effectiveness of our approach, we further investigate the source of these gains through quantitative ablation studies \([Table2](https://arxiv.org/html/2609.38274#S4.T2)\) and qualitative visualization \([Figure4](https://arxiv.org/html/2609.38274#S4.F4)\)\.
\(a\)Sensitivity of
heterogeneity weights\.
\(b\)Accuracy vs\. profiling size\.\(c\)Selection stability \(Jaccard vs\. full\-dev team\)\.
Figure 3:Robustness and data efficiency of C2\-MAS\. \(a\) Test accuracy underStackingacross heterogeneity weights;\(0,0\)\(0,0\)denotes Quality\-Only and⋆\\starmarks the configuration used in our main experiments\. \(b–c\) Test accuracy and selection stability versus profiling size, with stability measured by mean Jaccard similarity to the full\-dev team\.HIerr\\mathrm\{HI\}\_\{\\mathrm\{err\}\}HIdist\\mathrm\{HI\}\_\{\\mathrm\{dist\}\}Choice\-SoftPoEDSStackingOverall✗✗74\.7975\.0074\.2674\.7274\.69✓✗75\.2475\.2074\.6475\.2775\.09✓✓75\.4575\.2474\.6175\.5275\.21Table 2:Ablation study of C2\-MAS\. The first row corresponds to the Quality\-Only baseline\.#### Contribution of Profiling Signals\.
Table[2](https://arxiv.org/html/2609.38274#S4.T2)presents the performance impact of removing individual heterogeneity signals\. \(1\)*Error Decorrelation*: WithHIerr\\mathrm\{HI\}\_\{err\}, all four aggregators improve over Quality\-Only, supporting the value of reducing shared errors\. \(2\)*Distributional Divergence*: AddingHIdist\\mathrm\{HI\}\_\{dist\}further improves overall accuracy, indicating that choice probabilities contain complementary information\. \(3\)*Synergy*: The combined signals achieve the highest overall score, consistent with their distinct roles in profiling complementarity\.
Figure 4:MDS visualization of model heterogeneity\.#### Visualization of the Heterogeneity Landscape\.
To examine the limitations of Quality\-Only selection, we visualize model heterogeneity using multidimensional scaling \(MDS\)\. The embedding uses pairwise Yule’sQQand choice\-space JSD averaged across the seven primary benchmarks\. Models with similar prediction patterns cluster together\. As shown in Figure[4](https://arxiv.org/html/2609.38274#S4.F4), high\-accuracy models \(e\.g\., Qwen3, Gemma\-2\) tend to cluster tightly in the center\. This reveals a key phenomenon: the strongest models are often the most similar, likely sharing knowledge boundaries and blind spots\. Consequently, a Quality\-Only strategy that simply selects the top\-kkmodels often results in a highly redundant team\. In contrast, distinct models with slightly lower accuracy \(e\.g\., GLM\-4\) are scattered along the periphery\. C2\-MAS excels by identifying and recruiting these “outlier” members while maintaining a quality baseline\. These diverse perspectives fill the blind spots of strong models, thereby breaking the performance ceiling of homogeneous teams\.
### 4\.4Robustness and Efficiency \(RQ3\)
For real\-world deployment, sensitivity to hyperparameters and dependence on cold\-start data are key concerns\. We verify the robustness of C2\-MAS as follows\.
#### Sensitivity to Weights\.
Figure[3](https://arxiv.org/html/2609.38274#S4.F3)\(a\) shows broad regions of high accuracy away from the Quality\-Only origin\(0,0\)\(0,0\)\. The selection score is affine in\(λ1,λ2\)\(\\lambda\_\{1\},\\lambda\_\{2\}\), so the greedy output is constant within regions where its comparison signs remain fixed \([SectionC\.5](https://arxiv.org/html/2609.38274#A3.SS5)\)\. This explains why nearby weights can select the same team\. The observed high\-accuracy regions indicate low sensitivity to local weight changes in the evaluated range\.
#### Data Efficiency and Stability\.
Figures[3](https://arxiv.org/html/2609.38274#S4.F3)\(b\) and \(c\) examine profiling data size\. C2\-MAS retains its advantage with a few hundred development instances, while Jaccard similarity to the team selected using the full dev set rises rapidly and saturates early\. Together, these results show that reliable team composition and competitive accuracy can be achieved with limited development data, supporting the practical efficiency of our framework\.
## 5Conclusion
We proposed C2\-MAS, a heterogeneity\-driven framework for selecting complementary LLM teams under budget constraints\. By explicitly modelingerror decorrelationanddistributional divergence, our approach quantifies inter\-member complementarity beyond individual capability\. Extensive experiments validate the effectiveness, robustness, and efficiency of C2\-MAS, demonstrating that it consistently outperforms quality\-only baselines by recruiting complementary models to fill the blind spots of top\-performing candidates\. Future work could extend heterogeneity profiling to open\-ended generation and explore how complementary teams collaborate through multi\-turn debate and iterative verification\.
## Ethics Statement
Our experiments use only publicly available multiple\-choice benchmarks under their original licenses, and we evaluate open\-weight models alongside two closed\-source APIs \(GPT\-4o, Gemini\-2\.5\-Flash\) accessed through their official endpoints in accordance with the providers’ terms of use\. No human subjects, crowdsourced annotation, or personally identifiable information are involved\.
A risk specific to aggregation\-based methods is that biases or errors present in individual models may be amplified, attenuated, or obscured when their outputs are combined\. Although our heterogeneity metrics are designed to diversify error patterns, they do not directly target social bias, and practitioners deploying C2\-MAS should pair team selection with bias auditing of the resulting predictions\.
## Reproducibility Statement
The candidate models, benchmarks, selection weights, and aggregation settings are described in[Sections4\.1](https://arxiv.org/html/2609.38274#S4.SS1)and[3](https://arxiv.org/html/2609.38274#S3)\. The appendix provides the choice\-scoring implementation, paired\-bootstrap procedure, hyperparameter selection protocol, and complete results \([AppendicesD](https://arxiv.org/html/2609.38274#A4)and[E\.1](https://arxiv.org/html/2609.38274#A5.SS1)\)\.
## Authors and Affiliations
Liangyu Teng1, Hengsong Liu1, Juncen Guo1, Jingyu Zhang1, Yang Liu2, Jing Liu3, Liang Song1
1College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China2Tongji University, Shanghai, China3College of Future Information Technology, Fudan University, Shanghai, China
## References
- Aminiet al\.\(2019\)A\. Amini, S\. Gabriel, S\. Lin, R\. Koncel\-Kedziorski, Y\. Choi, and H\. HajishirziMathqa: towards interpretable math word problem solving with operation\-based formalisms\.InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: Human language technologies, volume 1 \(long and short papers\),pp\. 2357–2367\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Apertuset al\.\(2025\)P\. Apertus, A\. Hernández\-Cano, A\. Hägele, A\. H\. Huang, A\. Romanou, A\. Solergibert, B\. Pasztor, B\. Messmer, D\. Garbaya, E\. F\. Ďurech,et al\.Apertus: democratizing open and compliant llms for global language environments\.arXiv preprint arXiv:2509\.14233\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Basantet al\.\(2025\)A\. Basant, A\. Khairnar, A\. Paithankar, A\. Khattar, A\. Renduchintala, A\. Malte, A\. Bercovich, A\. Hazare, A\. Rico, A\. Ficek,et al\.Nvidia nemotron nano 2: an accurate and efficient hybrid mamba\-transformer reasoning model\.arXiv preprint arXiv:2508\.14444\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Bian and Chen \(2022\)Y\. Bian and H\. ChenWhen does diversity help generalization in classification ensembles?\.IEEE Transactions on Cybernetics52\(9\),pp\. 9059–9075\.External Links:[Document](https://dx.doi.org/10.1109/TCYB.2021.3053165)Cited by:[Appendix C](https://arxiv.org/html/2609.38274#A3.p1.1)\.
- Caiet al\.\(2024\)Z\. Cai, M\. Cao, H\. Chen, K\. Chen, K\. Chen, X\. Chen, X\. Chen, Z\. Chen, Z\. Chen, P\. Chu,et al\.Internlm2 technical report\.arXiv preprint arXiv:2403\.17297\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Caruanaet al\.\(2004\)R\. Caruana, A\. Niculescu\-Mizil, G\. Crew, and A\. KsikesEnsemble selection from libraries of models\.InProceedings of the twenty\-first international conference on Machine learning,pp\. 18\.Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px3.p1.1)\.
- Chenet al\.\(2024a\)J\. Chen, S\. Saha, and M\. BansalReConcile: round\-table conference improves reasoning via consensus among diverse LLMs\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 7066–7085\.External Links:[Link](https://aclanthology.org/2024.acl-long.381/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.381)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Chenet al\.\(2024b\)L\. Chen, M\. Zaharia, and J\. ZouFrugalGPT: how to use large language models while reducing cost and improving performance\.Transactions on Machine Learning Research\.Note:Featured CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=cSimKw5p6R)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1),[§1](https://arxiv.org/html/2609.38274#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Chenet al\.\(2024c\)W\. Chen, Y\. Su, J\. Zuo, C\. Yang, C\. Yuan, C\. Chan, H\. Yu, Y\. Lu, Y\. Hung, C\. Qian, Y\. Qin, X\. Cong, R\. Xie, Z\. Liu, M\. Sun, and J\. ZhouAgentVerse: facilitating multi\-agent collaboration and exploring emergent behaviors\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 20094–20136\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/578e65cdee35d00c708d4c64bce32971-Paper-Conference.pdf)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Chenet al\.\(2024d\)X\. Chen, R\. Aksitov, U\. Alon, J\. Ren, K\. Xiao, P\. Yin, S\. Prakash, C\. Sutton, X\. Wang, and D\. ZhouUniversal self\-consistency for large language models\.InICML 2024 Workshop on In\-Context Learning,External Links:[Link](https://openreview.net/forum?id=LjsjHF7nAN)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Chenet al\.\(2025\)Z\. Chen, X\. Lu, J\. Li, P\. Chen, Z\. Li, K\. Sun, Y\. Luo, Q\. Mao, M\. Li, L\. Xiao, D\. Yang, X\. Huang, Y\. Ban, H\. Sun, and P\. S\. YuHarnessing multiple large language models: a survey on llm ensemble\.arXiv preprint arXiv:2502\.18036\.External Links:[Link](https://arxiv.org/abs/2502.18036)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Choiet al\.\(2025\)H\. K\. Choi, X\. Zhu, and S\. LiDebate or vote: which yields better decisions in multi\-agent large language models?\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=iUjGNJzrF1)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try arc, the ai2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Cohenet al\.\(2026\)S\. Cohen, N\. C\. Inger, N\. Goldshlager, B\. Shapira, and L\. RokachDFPE: a diverse fingerprint ensemble for enhancing LLM performance\.InFindings of the Association for Computational Linguistics: EACL 2026,V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 5326–5336\.External Links:[Link](https://aclanthology.org/2026.findings-eacl.282/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-eacl.282),ISBN 979\-8\-89176\-386\-9Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Dekonincket al\.\(2025\)J\. Dekoninck, M\. Baader, and M\. VechevA unified approach to routing and cascading for LLMs\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=AAl89VNNy1)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Dinget al\.\(2025\)D\. Ding, A\. Mallick, S\. Zhang, C\. Wang, D\. Madrigal, M\. D\. C\. H\. Garcia, M\. Xia, L\. V\. S\. Lakshmanan, Q\. Wu, and V\. RühleBEST\-route: adaptive LLM routing with test\-time optimal compute\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=tFBIbCVXkG)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Duet al\.\(2024\)Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. MordatchImproving factuality and reasoning in language models through multiagent debate\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§1](https://arxiv.org/html/2609.38274#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Endres and Schindelin \(2003\)D\. M\. Endres and J\. E\. SchindelinA new metric for probability distributions\.IEEE Transactions on Information theory49\(7\),pp\. 1858–1860\.Cited by:[§3\.2](https://arxiv.org/html/2609.38274#S3.SS2.SSS0.Px3.p1.1)\.
- Estornell and Liu \(2024\)A\. Estornell and Y\. LiuMulti\-llm debate: framework, principals, and interventions\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 28938–28964\.External Links:[Document](https://dx.doi.org/10.52202/079017-0911),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/32e07a110c6c6acf1afbf2bf82b614ad-Paper-Conference.pdf)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Gevaet al\.\(2021\)M\. Geva, D\. Khashabi, E\. Segal, T\. Khot, D\. Roth, and J\. BerantDid aristotle use a laptop? a question answering benchmark with implicit reasoning strategies\.Transactions of the Association for Computational Linguistics9,pp\. 346–361\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- GLMet al\.\(2024\)T\. GLM, A\. Zeng, B\. Xu, B\. Wang, C\. Zhang, D\. Yin, D\. Zhang, D\. Rojas, G\. Feng, H\. Zhao,et al\.Chatglm: a family of large language models from glm\-130b to glm\-4 all tools\.arXiv preprint arXiv:2406\.12793\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Hansen and Salamon \(1990\)L\. K\. Hansen and P\. SalamonNeural network ensembles\.IEEE transactions on pattern analysis and machine intelligence12\(10\),pp\. 993–1001\.Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ)Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Honget al\.\(2024\)S\. Hong, M\. Zhuge, J\. Chen, X\. Zheng, Y\. Cheng, J\. Wang, C\. Zhang, Z\. Wang, S\. K\. S\. Yau, Z\. Lin, L\. Zhou, C\. Ran, L\. Xiao, C\. Wu, and J\. SchmidhuberMetaGPT: meta programming for a multi\-agent collaborative framework\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VtmBAGCN7o)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Huet al\.\(2024\)Q\. J\. Hu, J\. Bieker, X\. Li, N\. Jiang, B\. Keigwin, G\. Ranganath, K\. Keutzer, and S\. K\. UpadhyayRouterbench: a benchmark for multi\-llm routing system\.arXiv preprint arXiv:2403\.12031\.Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Huanget al\.\(2024\)Y\. Huang, X\. Feng, B\. Li, Y\. Xiang, H\. Wang, T\. Liu, and B\. QinEnsemble learning for heterogeneous large language models with deep parallel collaboration\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 119838–119860\.External Links:[Document](https://dx.doi.org/10.52202/079017-3808),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/d8a6eb79f8ccaacbe7198a5caf3a0323-Paper-Conference.pdf)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Huanget al\.\(2023\)Y\. Huang, Y\. Bai, Z\. Zhu, J\. Zhang, J\. Zhang, T\. Su, J\. Liu, C\. Lv, Y\. Zhang, Y\. Fu,et al\.C\-eval: a multi\-level multi\-discipline chinese evaluation suite for foundation models\.Advances in Neural Information Processing Systems36,pp\. 62991–63010\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- IBM Granite Team \(2024\)IBM Granite TeamGranite 3\.0 language models\.Note:Technical reportExternal Links:[Link](https://github.com/ibm-granite/granite-3.0-language-models)Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Jianget al\.\(2023\)D\. Jiang, X\. Ren, and B\. Y\. LinLLM\-blender: ensembling large language models with pairwise ranking and generative fusion\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 14165–14178\.External Links:[Link](https://aclanthology.org/2023.acl-long.792/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.792)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1),[§1](https://arxiv.org/html/2609.38274#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Jinet al\.\(2021\)D\. Jin, E\. Pan, N\. Oufattole, W\. Weng, H\. Fang, and P\. SzolovitsWhat disease does this patient have? a large\-scale open domain question answering dataset from medical exams\.Applied Sciences11\(14\)\.External Links:[Link](https://www.mdpi.com/2076-3417/11/14/6421),ISSN 2076\-3417,[Document](https://dx.doi.org/10.3390/app11146421)Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Krogh and Vedelsby \(1994\)A\. Krogh and J\. VedelsbyNeural network ensembles, cross validation, and active learning\.Advances in neural information processing systems7\.Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Kuncheva and Whitaker \(2003\)L\. I\. Kuncheva and C\. J\. WhitakerMeasures of diversity in classifier ensembles and their relationship with the ensemble accuracy\.Machine learning51\(2\),pp\. 181–207\.Cited by:[§C\.1](https://arxiv.org/html/2609.38274#A3.SS1.p1.1),[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Laiet al\.\(2017\)G\. Lai, Q\. Xie, H\. Liu, Y\. Yang, and E\. HovyRACE: large\-scale reading comprehension dataset from examinations\.InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing,pp\. 785–794\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Liet al\.\(2023\)G\. Li, H\. A\. Al Kader Hammoud, H\. Itani, D\. Khizbullin, and B\. GhanemCAMEL: communicative agents for "mind" exploration of large language model society\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Liet al\.\(2024\)Y\. Li, Y\. Du, J\. Zhang, L\. Hou, P\. Grabowski, Y\. Li, and E\. IeImproving multi\-agent debate with sparse communication topology\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 7281–7294\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.427/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.427)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Lianget al\.\(2024\)T\. Liang, Z\. He, W\. Jiao, X\. Wang, Y\. Wang, R\. Wang, Y\. Yang, S\. Shi, and Z\. TuEncouraging divergent thinking in large language models through multi\-agent debate\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 17889–17904\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.992/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.992)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§1](https://arxiv.org/html/2609.38274#S1.p1.1)\.
- Lin \(1991\)J\. LinDivergence measures based on the shannon entropy\.IEEE Transactions on Information Theory37\(1\),pp\. 145–151\.External Links:[Document](https://dx.doi.org/10.1109/18.61115)Cited by:[§C\.4](https://arxiv.org/html/2609.38274#A3.SS4.p1.1),[§D\.1](https://arxiv.org/html/2609.38274#A4.SS1.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.38274#S3.SS2.SSS0.Px3.p1.1)\.
- Linget al\.\(2017\)W\. Ling, D\. Yogatama, C\. Dyer, and P\. BlunsomProgram induction by rationale generation: learning to solve and explain algebraic word problems\.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),R\. Barzilay and M\. Kan \(Eds\.\),Vancouver, Canada,pp\. 158–167\.External Links:[Link](https://aclanthology.org/P17-1015/),[Document](https://dx.doi.org/10.18653/v1/P17-1015)Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Liuet al\.\(2026\)A\. H\. Liu, K\. Khandelwal, S\. Subramanian, V\. Jouault, A\. Rastogi, A\. Sadé, A\. Jeffares, A\. Jiang, A\. Cahill, A\. Gavaudan,et al\.Ministral 3\.arXiv preprint arXiv:2601\.08584\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Liuet al\.\(2023\)H\. Liu, J\. Liu, L\. Cui, Z\. Teng, N\. Duan, M\. Zhou, and Y\. ZhangLogiqa 2\.0—an improved dataset for logical reasoning in natural language understanding\.IEEE/ACM Transactions on Audio, Speech, and Language Processing31,pp\. 2947–2962\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Liuet al\.\(2024\)Z\. Liu, Y\. Zhang, P\. Li, Y\. Liu, and D\. YangA dynamic LLM\-powered agent network for task\-oriented agent collaboration\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=XII0Wp1XA9)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Luet al\.\(2024\)J\. Lu, Z\. Pang, M\. Xiao, Y\. Zhu, R\. Xia, and J\. ZhangMerge, ensemble, and cooperate\! a survey on collaborative strategies in the era of large language models\.arXiv preprint arXiv:2407\.06089\.Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Martinset al\.\(2025\)P\. H\. Martins, J\. Alves, P\. Fernandes, N\. M\. Guerreiro, R\. Rei, A\. Farajian, M\. Klimaszewski, D\. M\. Alves, J\. Pombal, N\. Boizard,et al\.Eurollm\-9b: technical report\.arXiv preprint arXiv:2506\.04079\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Mihaylovet al\.\(2018\)T\. Mihaylov, P\. Clark, T\. Khot, and A\. SabharwalCan a suit of armor conduct electricity? a new dataset for open book question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,pp\. 2381–2391\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Olmoet al\.\(2025\)T\. Olmo, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison,et al\.Olmo 3\.arXiv preprint arXiv:2512\.13961\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Onget al\.\(2025\)I\. Ong, A\. Almahairi, V\. Wu, W\. Chiang, T\. Wu, J\. E\. Gonzalez, M\. W\. Kadous, and I\. StoicaRouteLLM: learning to route LLMs from preference data\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=8sSqNntaMr)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1),[§1](https://arxiv.org/html/2609.38274#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- O’Brien and Lewis \(2023\)S\. O’Brien and M\. LewisContrastive decoding improves reasoning in large language models\.arXiv preprint arXiv:2309\.09117\.Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Parket al\.\(2023\)J\. S\. Park, J\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology,UIST ’23,New York, NY, USA\.External Links:ISBN 9798400701320,[Link](https://doi.org/10.1145/3586183.3606763),[Document](https://dx.doi.org/10.1145/3586183.3606763)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Parket al\.\(2025\)S\. Park, X\. Liu, Y\. Gong, and E\. ChoiEnsembling large language models with process reward\-guided tree search for better complex reasoning\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 10256–10277\.External Links:[Link](https://aclanthology.org/2025.naacl-long.515/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.515),ISBN 979\-8\-89176\-189\-6Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Piaoet al\.\(2025\)J\. Piao, Y\. Yan, J\. Zhang, N\. Li, J\. Yan, X\. Lan, Z\. Lu, Z\. Zheng, J\. Y\. Wang, D\. Zhou,et al\.Agentsociety: large\-scale simulation of llm\-driven generative agents advances understanding of human behaviors and society\.arXiv preprint arXiv:2502\.08691\.Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Qianet al\.\(2024\)C\. Qian, W\. Liu, H\. Liu, N\. Chen, Y\. Dang, J\. Li, C\. Yang, W\. Chen, Y\. Su, X\. Cong, J\. Xu, D\. Li, Z\. Liu, and M\. SunChatDev: communicative agents for software development\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15174–15186\.External Links:[Link](https://aclanthology.org/2024.acl-long.810/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.810)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Qianet al\.\(2025\)C\. Qian, Z\. Xie, Y\. Wang, W\. Liu, K\. Zhu, H\. Xia, Y\. Dang, Z\. Du, W\. Chen, C\. Yang, Z\. Liu, and M\. SunScaling large language model\-based multi\-agent collaboration\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 41488–41505\.External Links:[Link](https://openreview.net/forum?id=K3n5jPkrU6)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Shenet al\.\(2024\)Z\. Shen, H\. Lang, B\. Wang, Y\. Kim, and D\. SontagLearning to decode collaboratively with multiple language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 12974–12990\.External Links:[Link](https://aclanthology.org/2024.acl-long.701/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.701)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Shnitzeret al\.\(2024\)T\. Shnitzer, A\. Ou, M\. Silva, K\. Soule, Y\. Sun, J\. Solomon, N\. Thompson, and M\. YurochkinLarge language model routing with benchmark datasets\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Zb0ajZ7vAt)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1),[§1](https://arxiv.org/html/2609.38274#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Talmoret al\.\(2019\)A\. Talmor, J\. Herzig, N\. Lourie, and J\. BerantCommonsenseQA: a question answering challenge targeting commonsense knowledge\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 4149–4158\.External Links:[Link](https://aclanthology.org/N19-1421/),[Document](https://dx.doi.org/10.18653/v1/N19-1421)Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Tanget al\.\(2006\)E\. K\. Tang, P\. N\. Suganthan, and X\. YaoAn analysis of diversity measures\.Machine Learning65,pp\. 247–271\.External Links:[Document](https://dx.doi.org/10.1007/s10994-006-9449-2)Cited by:[§C\.1](https://arxiv.org/html/2609.38274#A3.SS1.p1.1),[Appendix C](https://arxiv.org/html/2609.38274#A3.p1.1)\.
- Tanget al\.\(2025\)J\. Tang, H\. Gao, X\. Pan, L\. Wang, H\. Tan, D\. Gao, Y\. Chen, X\. Chen, Y\. Lin, Y\. Li, B\. Ding, J\. Zhou, J\. Wang, and J\. WenGenSim: a general social simulation platform with large language model based agents\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(System Demonstrations\),N\. Dziri, S\. \(\. Ren, and S\. Diao \(Eds\.\),Albuquerque, New Mexico,pp\. 143–150\.External Links:[Link](https://aclanthology.org/2025.naacl-demo.15/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-demo.15),ISBN 979\-8\-89176\-191\-9Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Teamet al\.\(2024\)G\. Team, M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé,et al\.Gemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Tekinet al\.\(2024\)S\. F\. Tekin, F\. Ilhan, T\. Huang, S\. Hu, and L\. LiuLLM\-TOPLA: efficient LLM ensemble by maximising diversity\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 11951–11966\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.698),[Link](https://aclanthology.org/2024.findings-emnlp.698/)Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Turkmenet al\.\(2026\)Y\. Turkmen, B\. Buyukates, and M\. BastopcuDon’t always pick the highest\-performing model: an information theoretic view of LLM ensemble selection\.arXiv preprint arXiv:2602\.08003\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.08003),[Link](https://arxiv.org/abs/2602.08003)Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Wanget al\.\(2025a\)J\. Wang, J\. Wang, B\. Athiwaratkun, C\. Zhang, and J\. Y\. ZouMixture\-of\-agents enhances large language model capabilities\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 33944–33963\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/5434be94e82c54327bb9dcaf7fca52b6-Paper-Conference.pdf)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1),[§1](https://arxiv.org/html/2609.38274#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. V\. Le, E\. H\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.InThe Eleventh International Conference on Learning Representations,Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px3.p1.1)\.
- Wanget al\.\(2024\)Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.Mmlu\-pro: a more robust and challenging multi\-task language understanding benchmark\.Advances in Neural Information Processing Systems37,pp\. 95266–95290\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Wanget al\.\(2025b\)Z\. Wang, S\. Moriyama, W\. Wang, B\. Gangopadhyay, and S\. TakamatsuTalk structurally, act hierarchically: a collaborative framework for llm multi\-agent systems\.arXiv preprint arXiv:2502\.11098\.Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Wanget al\.\(2025c\)Z\. Wang, M\. Azmat, A\. Li, R\. Horesh, and M\. YurochkinSpeculate, then collaborate: fusing knowledge of language models during decoding\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=XCBYIfu9Fs)Cited by:[§B\.2](https://arxiv.org/html/2609.38274#A2.SS2.p1.1)\.
- Windeatt \(2005\)T\. WindeattDiversity measures for multiple classifier system analysis and design\.Information fusion6\(1\),pp\. 21–36\.Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Woodet al\.\(2023\)D\. Wood, T\. Mu, A\. M\. Webb, H\. W\. J\. Reeve, M\. Luján, and G\. BrownA unified theory of diversity in ensemble learning\.Journal of Machine Learning Research24\(359\),pp\. 1–49\.External Links:[Link](https://www.jmlr.org/papers/v24/23-0041.html)Cited by:[Appendix C](https://arxiv.org/html/2609.38274#A3.p1.1)\.
- Wuet al\.\(2024\)Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu, A\. H\. Awadallah, R\. W\. White, D\. Burger, and C\. WangAutoGen: enabling next\-gen LLM applications via multi\-agent conversations\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=BAakY1hNKS)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px2.p1.1)\.
- Yuet al\.\(2020\)W\. Yu, Z\. Jiang, Y\. Dong, and J\. FengReClor: a reading comprehension dataset requiring logical reasoning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HJgJtT4tvB)Cited by:[§4\.1](https://arxiv.org/html/2609.38274#S4.SS1.SSS0.Px1.p1.1)\.
- Yule \(1900\)G\. U\. YuleOn the association of attributes in statistics\.Philosophical Transactions of the Royal Society of London\. Series A194,pp\. 257–319\.Cited by:[§C\.1](https://arxiv.org/html/2609.38274#A3.SS1.p1.1)\.
- Zhanget al\.\(2025a\)G\. Zhang, Y\. Yue, Z\. Li, S\. Yun, G\. Wan, K\. Wang, D\. Cheng, J\. Yu, and T\. ChenCut the crap: an economical communication pipeline for llm\-based multi\-agent systems\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 75389–75428\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/bbc461518c59a2a8d64e70e2c38c4a0e-Paper-Conference.pdf)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Zhanget al\.\(2025b\)G\. Zhang, Y\. Yue, X\. Sun, G\. Wan, M\. Yu, J\. Fang, K\. Wang, T\. Chen, and D\. ChengG\-designer: architecting multi\-agent communication topologies via graph neural networks\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=LpE54NUnmO)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
- Zhanget al\.\(2026\)Y\. Zhang, K\. Lu, Y\. Zhang, J\. Gao, L\. Xia, and F\. YuMixture of complementary agents for robust LLM ensemble\.arXiv preprint arXiv:2605\.24048\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.24048),[Link](https://arxiv.org/abs/2605.24048)Cited by:[§2\.2](https://arxiv.org/html/2609.38274#S2.SS2.p1.1)\.
- Zhouet al\.\(2024\)X\. Zhou, H\. Zhu, L\. Mathur, R\. Zhang, H\. Yu, Z\. Qi, L\. Morency, Y\. Bisk, D\. Fried, G\. Neubig, and M\. SapSOTOPIA: interactive evaluation for social intelligence in language agents\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 40975–41019\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/b3075b88e583a0e98d8b24338a613060-Paper-Conference.pdf)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1)\.
- Zhugeet al\.\(2024\)M\. Zhuge, W\. Wang, L\. Kirsch, F\. Faccio, D\. Khizbullin, and J\. SchmidhuberGPTSwarm: language agents as optimizable graphs\.InForty\-first International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=uTC9AFXIhg)Cited by:[§B\.1](https://arxiv.org/html/2609.38274#A2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.38274#S2.SS1.p1.1)\.
## Appendix Contents
[\*app:theory](https://arxiv.org/html/2609.38274#A3)\.[C](https://arxiv.org/html/2609.38274#A3)
## Appendix ANotations
Table 3:Notations and DefinitionsNotationDefinition𝒯\\mathcal\{T\}Set of tasks used for cross\-task pooling\.ttTask / benchmark index\.Dtdev,DttestD\_\{t\}^\{\\mathrm\{dev\}\},\\,D\_\{t\}^\{\\mathrm\{test\}\}Development and test splits for tasktt\(strict dev→\\rightarrowtest\)\.𝒴t\\mathcal\{Y\}\_\{t\}Label / option set for tasktt\(multiple\-choice\)\.x,yx,\\,yAn instance and its gold label \(y∈𝒴ty\\in\\mathcal\{Y\}\_\{t\}\)\.ℳ=\{1,…,m\}\\mathcal\{M\}=\\\{1,\\ldots,m\\\}Index set of themmcandidate agents \(callable LLMs\)\.mmNumber of candidate agents\.kkTeam size budget\.S⊆ℳS\\subseteq\\mathcal\{M\}A team of agents, with\|S\|=k\|S\|=k\.i∈ℳi\\in\\mathcal\{M\}Index of a candidate agent\.pi\(⋅∣x\)p\_\{i\}\(\\cdot\\mid x\)Agentii’s aligned choice distribution over𝒴t\\mathcal\{Y\}\_\{t\}for instancexx\(task index omitted when clear\)\.y^i\(x\)\\hat\{y\}\_\{i\}\(x\)Predicted label of agentii,y^i\(x\)=argmaxy∈𝒴tpi\(y∣x\)\\hat\{y\}\_\{i\}\(x\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\_\{t\}\}p\_\{i\}\(y\\mid x\)\.Ci\(x\)∈\{0,1\}C\_\{i\}\(x\)\\in\\\{0,1\\\}Correctness indicator of agentiion examplexx\.Nab\(i,j\)N\_\{ab\}\(i,j\)Count of examples where\(Ci\(x\),Cj\(x\)\)=\(a,b\)\(C\_\{i\}\(x\),C\_\{j\}\(x\)\)=\(a,b\)fora,b∈\{0,1\}a,b\\in\\\{0,1\\\}\.R\(i,j\)R\(i,j\)Error\-correlation / complementarity statistic between agentsiiandjj\(e\.g\., Yule’sQQ\)\.qitq\_\{i\}^\{t\}Estimated individual quality of agentiion tasktt\(development accuracy\)\.R¯\\bar\{R\}Pooled Yule’sQQassociation matrix estimated from development data\.J¯\\bar\{J\}Pooled distribution\-divergence matrix estimated from development data \(across tasks if applicable\)\.HIerr\(S,R¯\)\\mathrm\{HI\}\_\{\\mathrm\{err\}\}\(S;\\bar\{R\}\)Error\-level heterogeneity / complementarity of teamSS\(computed fromR¯\\bar\{R\}\)\.HIdist\(S,J¯\)\\mathrm\{HI\}\_\{\\mathrm\{dist\}\}\(S;\\bar\{J\}\)Distribution\-level heterogeneity / disagreement of teamSS\(computed fromJ¯\\bar\{J\}\)\.z\(⋅\)z\(\\cdot\)Standard\-error\-scaled normalization of a team average using candidate\-pool base\-item statistics\.λ1,λ2\\lambda\_\{1\},\\,\\lambda\_\{2\}Objective weights trading off capability and complementarity\.scoret\(S\)\\mathrm\{score\}\_\{t\}\(S\)Standardized selection objective for tasktt\(see[Equation10](https://arxiv.org/html/2609.38274#S3.E10)\)\.St∗S\_\{t\}^\{\*\}Size\-kkteam returned by multi\-start greedy search onscoret\(S\)\\mathrm\{score\}\_\{t\}\(S\)\.pcst\(y∣x\)p\_\{\\textsc\{cs\}\}^\{t\}\(y\\mid x\)Choice\-Soft aggregation distribution \(mean of members’pitp\_\{i\}^\{t\}\)\.ppoet\(y∣x\)p\_\{\\textsc\{poe\}\}^\{t\}\(y\\mid x\)PoE aggregation distribution \(product of members’pitp\_\{i\}^\{t\}, renormalized\)\.πt\(y\)\\pi^\{t\}\(y\)Class prior estimated fromDtdevD\_\{t\}^\{\\mathrm\{dev\}\}\(used in DS\)\.Ait\(a,b\)A\_\{i\}^\{t\}\(a,b\)DS confusion matrix entry estimatingℙ\(y^i=b∣y=a\)\\mathbb\{P\}\(\\hat\{y\}\_\{i\}=b\\mid y=a\)onDtdevD\_\{t\}^\{\\mathrm\{dev\}\}\.ε\\varepsilonAdditive smoothing constant for DS priors/confusions \(e\.g\.,ε=10−3\\varepsilon=10^\{\-3\}\)\.pdst\(y∣x\)p\_\{\\textsc\{ds\}\}^\{t\}\(y\\mid x\)DS posterior \(up to proportionality\) combiningπt\\pi^\{t\}and\{Ait\}\\\{A\_\{i\}^\{t\}\\\}\.ϕ\(x\)\\phi\(x\)Stacking feature vector formed by concatenating members’pit\(⋅∣x\)p\_\{i\}^\{t\}\(\\cdot\\mid x\)\.W,bW,\\,bLinear stacking combiner parameters\.β\\betaL2L\_\{2\}regularization coefficient for stacking\.πW,b\(y∣x\)\\pi\_\{W,b\}\(y\\mid x\)Stacking aggregation distribution,πW,b\(y∣x\)=softmax\(W⊤ϕ\(x\)\+b\)y\\pi\_\{W,b\}\(y\\mid x\)=\\mathrm\{softmax\}\(W^\{\\top\}\\phi\(x\)\+b\)\_\{y\}\.y^St\(x\)\\hat\{y\}\_\{S\_\{t\}\}\(x\)Final team prediction onxxunder a chosen aggregator\.q¯S,eS\\bar\{q\}\_\{S\},\\,e\_\{S\}Mean member accuracy and error probability in the fixed\-task theoretical model\.DS,KS,JSD\_\{S\},\\,K\_\{S\},\\,J\_\{S\}Mean pairwise double\-fault probability, same\-wrong\-answer probability, and JSD in that model\.d,h,θd,\\,h,\\,\\thetaDivergence between distinct wrong\-answer templates, divergence between correct and incorrect templates, andθ=2h/d\\theta=2h/d\.εQ\\varepsilon\_\{Q\}Baseline correction in the accuracy comparison with Quality\-Only\.
## Appendix BAdditional Related Works
### B\.1Multi\-Agent LLM Systems
Much prior work builds multi\-agent collaboration on predefined roles and fixed workflows, ranging from role\-playing paradigms\([Li et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib11);[Hong et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib12)\)and workflow\-based dialogue orchestration\([Wu et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib13);[Qian et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib14);[Chen et al\., 2024c](https://arxiv.org/html/2609.38274#bib.bib15)\)to multi\-agent simulation environments for modeling collective and social behaviors\([Park et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib16);[Zhou et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib17);[Tang et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib18);[Piao et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib19)\)\. To improve reliability, studies model interaction as structured group decision processes such as debate\([Du et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib4);[Liang et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib5)\), vote\-based aggregation\([Choi et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib7)\), and consensus mechanisms\([Chen et al\., 2024a](https://arxiv.org/html/2609.38274#bib.bib20)\), and analyze protocol interventions to understand mechanisms and failure modes\([Estornell and Liu, 2024](https://arxiv.org/html/2609.38274#bib.bib6)\)\. More recent work treats the interaction structure as an optimization target, using learnable communication graphs\([Zhuge et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib21);[Zhang et al\., 2025b](https://arxiv.org/html/2609.38274#bib.bib22)\), sparse or hierarchical messaging\([Li et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib23);[Qian et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib8);[Wang et al\., 2025b](https://arxiv.org/html/2609.38274#bib.bib24)\), pruning and economical communication pipelines\([Zhang et al\., 2025a](https://arxiv.org/html/2609.38274#bib.bib25)\), or dynamic team formation\([Liu et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib26)\)to improve scalability and efficiency\.
### B\.2Multi\-Model Collaboration, Output Fusion, and Routing
Multi\-model collaboration is often organized as an ensemble pipeline that generates candidates, selects among them, and fuses outputs\([Lu et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib27);[Chen et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib28)\)\. Methods operate at the output level via pairwise ranking\([Jiang et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib29)\), consistency\-based selection\([Chen et al\., 2024d](https://arxiv.org/html/2609.38274#bib.bib30)\), or consensus aggregation\([Chen et al\., 2024a](https://arxiv.org/html/2609.38274#bib.bib20)\); during reasoning through iterative refinement\([Wang et al\., 2025a](https://arxiv.org/html/2609.38274#bib.bib10)\)or tree\-search\-based decision making\([Park et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib31)\); or at decoding time by combining token\-level distributions\([Huang et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib9);[O’Brien and Lewis, 2023](https://arxiv.org/html/2609.38274#bib.bib32);[Shen et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib33);[Wang et al\., 2025c](https://arxiv.org/html/2609.38274#bib.bib34)\)\. For deployment, routing and cascading dynamically choose models per input to trade quality for cost\([Chen et al\., 2024b](https://arxiv.org/html/2609.38274#bib.bib35);[Shnitzer et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib36);[Ong et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib37)\), with systematic benchmarks characterizing the quality–cost trade\-off\([Hu et al\., 2024](https://arxiv.org/html/2609.38274#bib.bib38);[Ding et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib41)\)and unified analyses of when such strategies are effective\([Dekoninck et al\., 2025](https://arxiv.org/html/2609.38274#bib.bib39)\)\.
## Appendix CTheoretical Analysis and Proofs
The analysis follows the two questions in[Section3\.5](https://arxiv.org/html/2609.38274#S3.SS5): how the two heterogeneity signals characterize collective errors, and when their benefits compensate for a reduction in member quality\. Classical ensemble analyses relate diversity to margins and loss decompositions\([Tang et al\., 2006](https://arxiv.org/html/2609.38274#bib.bib1);[Bian and Chen, 2022](https://arxiv.org/html/2609.38274#bib.bib2);[Wood et al\., 2023](https://arxiv.org/html/2609.38274#bib.bib3)\)\. We develop the connection through a choice\-probability model, prove an accuracy comparison with Quality\-Only, and relate its terms to the standardized selection objective\. Further results cover soft recovery, profiling perturbations, cross\-task pooling, and the member\-specific information available to Stacking\. JSD uses base\-22logarithms; cross\-entropy and the auxiliary logarithmic calculations below use natural logarithms\.
### C\.1Correctness\-Level Dependence
Correctness correlations summarize which examples models solve together\. We first record their connection to shared errors, using the classical diversity statistics studied by[Yule \(1900\)](https://arxiv.org/html/2609.38274#bib.bib40);[Kuncheva and Whitaker \(2003\)](https://arxiv.org/html/2609.38274#bib.bib44);[Tang et al\. \(2006\)](https://arxiv.org/html/2609.38274#bib.bib1)\. Fix a task distribution𝒟\\mathcal\{D\}; an empirical development distribution is also allowed\. Writeqi=𝔼\[Ci\]q\_\{i\}=\\mathbb\{E\}\[C\_\{i\}\],Ei=1−CiE\_\{i\}=1\-C\_\{i\}, andDFij=ℙ\(Ei=Ej=1\)\\mathrm\{DF\}\_\{ij\}=\\mathbb\{P\}\(E\_\{i\}=E\_\{j\}=1\)\.
###### Lemma C\.1\(Fixed\-marginal dependence\)\.
For fixedqi,qj∈\(0,1\)q\_\{i\},q\_\{j\}\\in\(0,1\), Yule’sQQ, correctness covariance, and double\-fault probability are strictly increasing functions of the joint\-correct probabilityt=ℙ\(Ci=Cj=1\)t=\\mathbb\{P\}\(C\_\{i\}=C\_\{j\}=1\)over the interior of its feasible interval\. In particular,
DFij=\(1−qi\)\(1−qj\)\+Cov\(Ci,Cj\)\.\\mathrm\{DF\}\_\{ij\}=\(1\-q\_\{i\}\)\(1\-q\_\{j\}\)\+\\operatorname\{Cov\}\(C\_\{i\},C\_\{j\}\)\.\(14\)
###### Proof\.
The four joint probabilities arep11=tp\_\{11\}=t,p10=qi−tp\_\{10\}=q\_\{i\}\-t,p01=qj−tp\_\{01\}=q\_\{j\}\-t, andp00=1−qi−qj\+tp\_\{00\}=1\-q\_\{i\}\-q\_\{j\}\+t, wheremax\{0,qi\+qj−1\}≤t≤min\{qi,qj\}\\max\\\{0,q\_\{i\}\+q\_\{j\}\-1\\\}\\leq t\\leq\\min\\\{q\_\{i\},q\_\{j\}\\\}\. ThusCov\(Ci,Cj\)=t−qiqj\\operatorname\{Cov\}\(C\_\{i\},C\_\{j\}\)=t\-q\_\{i\}q\_\{j\}andDFij=1−qi−qj\+t\\mathrm\{DF\}\_\{ij\}=1\-q\_\{i\}\-q\_\{j\}\+t, giving[Equation14](https://arxiv.org/html/2609.38274#A3.E14); both have derivative one\. In the interior, all four cells are positive\. For the odds ratioO\(t\)=p11p00/\(p10p01\)O\(t\)=p\_\{11\}p\_\{00\}/\(p\_\{10\}p\_\{01\}\),
ddtlogO\(t\)=1t\+11−qi−qj\+t\+1qi−t\+1qj−t\>0\.\\frac\{d\}\{dt\}\\log O\(t\)=\\frac\{1\}\{t\}\+\\frac\{1\}\{1\-q\_\{i\}\-q\_\{j\}\+t\}\+\\frac\{1\}\{q\_\{i\}\-t\}\+\\frac\{1\}\{q\_\{j\}\-t\}\>0\.\(15\)SinceQ=\(O−1\)/\(O\+1\)Q=\(O\-1\)/\(O\+1\)anddQ/dO=2/\(O\+1\)2\>0dQ/dO=2/\(O\+1\)^\{2\}\>0,QQis strictly increasing as well\. Endpoint values follow by continuity\. Finally,Cov\(Ei,Ej\)=Cov\(Ci,Cj\)\\operatorname\{Cov\}\(E\_\{i\},E\_\{j\}\)=\\operatorname\{Cov\}\(C\_\{i\},C\_\{j\}\)becauseEi=1−CiE\_\{i\}=1\-C\_\{i\}\. ∎
With individual qualities held fixed,HIerr\\mathrm\{HI\}\_\{\\mathrm\{err\}\}therefore favors pairs with fewer shared errors\. For two members, an oracle that succeeds whenever either member is correct has error probabilityDFij\\mathrm\{DF\}\_\{ij\}\. Distributional analysis below examines how soft aggregation can recover information even on examples where every member’s top prediction is wrong\.
#### Classical majority\-vote bound\.
For completeness, the standard one\-sided variance argument gives another interpretation of correctness dependence\. For binary\-label prediction, letk≥3k\\geq 3be odd and let correctness indicators have common meanp∈\(1/2,1\)p\\in\(1/2,1\)and common pairwise correlationρ\\rho, and putC¯=k−1∑iCi\\bar\{C\}=k^\{\-1\}\\sum\_\{i\}C\_\{i\}\. Expanding the variance and applying Cantelli’s inequality gives
V\(ρ\):=Var\(C¯\)\\displaystyle V\(\\rho\):=\\operatorname\{Var\}\(\\bar\{C\}\)=p\(1−p\)k\(1\+\(k−1\)ρ\),\\displaystyle=\\frac\{p\(1\-p\)\}\{k\}\\bigl\(1\+\(k\-1\)\\rho\\bigr\),\(16\)ℙ\(majority vote errs\)=ℙ\(C¯≤1/2\)\\displaystyle\\mathbb\{P\}\(\\text\{majority vote errs\}\)=\\mathbb\{P\}\(\\bar\{C\}\\leq 1/2\)≤V\(ρ\)V\(ρ\)\+\(p−1/2\)2\.\\displaystyle\\leq\\frac\{V\(\\rho\)\}\{V\(\\rho\)\+\(p\-1/2\)^\{2\}\}\.Indeed, the variance containskkdiagonal termsp\(1−p\)p\(1\-p\)andk\(k−1\)k\(k\-1\)off\-diagonal termsρp\(1−p\)\\rho p\(1\-p\), all divided byk2k^\{2\}\. Cantelli’s inequality applied toC¯−p≤−\(p−1/2\)\\bar\{C\}\-p\\leq\-\(p\-1/2\)yields the second line\. WritingA=\(p−1/2\)2\>0A=\(p\-1/2\)^\{2\}\>0, the derivative of this upper bound isAp\(1−p\)\(k−1\)/\[k\(V\(ρ\)\+A\)2\]\>0Ap\(1\-p\)\(k\-1\)/\[k\(V\(\\rho\)\+A\)^\{2\}\]\>0throughout the feasible correlation range\. This is a classical bound on binary majority voting; the subsequent results analyze the choice distributions and selection rule used by C2\-MAS directly\.
### C\.2Joint Characterization of Collective Failure
We work on one fixed distribution𝒟\\mathcal\{D\}of labeled examples\. The same identities hold for the uniform distribution on a finite set\. Task indices are suppressed within this analysis; the target\-task and pooled profiles are distinguished explicitly in[SectionC\.7](https://arxiv.org/html/2609.38274#A3.SS7)\.
#### Model and pairwise quantities\.
All candidate members follow \([11](https://arxiv.org/html/2609.38274#S3.E11)\), with common constantsa,b,c,ua,b,c,u\. For each example, letwi≠yw\_\{i\}\\neq ydenote the preferred answer of an incorrect member\. Writing the true answer first, the distributions are
pc=\(a,u,u,u\),pw1=\(a,b,c,c\),pw2=\(a,c,b,c\),pw3=\(a,c,c,b\)\.p^\{\\rm c\}=\(a,u,u,u\),\\qquad p^\{w\_\{1\}\}=\(a,b,c,c\),\\quad p^\{w\_\{2\}\}=\(a,c,b,c\),\\quad p^\{w\_\{3\}\}=\(a,c,c,b\)\.The correctness indicators and thewiw\_\{i\}have an arbitrary joint distribution across members and examples\. Sincec<u<bc<u<banda\>\(2b\+u\)/3a\>\(2b\+u\)/3, we havea\>ua\>u, so these templates have the stated unique preferred answers\. The parameter region has nonempty interior; for example,\(a,b,c\)=\(0\.42,0\.48,0\.05\)\(a,b,c\)=\(0\.42,0\.48,0\.05\)satisfies all inequalities\.
For a three\-member team, define
q¯S\\displaystyle\\bar\{q\}\_\{S\}=13∑i∈Sqi,eS=1−q¯S,\\displaystyle=\\frac\{1\}\{3\}\\sum\_\{i\\in S\}q\_\{i\},\\qquad e\_\{S\}=1\-\\bar\{q\}\_\{S\},DS\\displaystyle D\_\{S\}=13∑\{i,j\}⊂Sℙ\(Ci=Cj=0\),\\displaystyle=\\frac\{1\}\{3\}\\sum\_\{\\\{i,j\\\}\\subset S\}\\mathbb\{P\}\(C\_\{i\}=C\_\{j\}=0\),KS\\displaystyle K\_\{S\}=13∑\{i,j\}⊂Sℙ\(Ci=Cj=0,wi=wj\),\\displaystyle=\\frac\{1\}\{3\}\\sum\_\{\\\{i,j\\\}\\subset S\}\\mathbb\{P\}\(C\_\{i\}=C\_\{j\}=0,\\ w\_\{i\}=w\_\{j\}\),JS\\displaystyle J\_\{S\}=13∑\{i,j\}⊂S𝔼\[JSD2\(pi,pj\)\]\.\\displaystyle=\\frac\{1\}\{3\}\\sum\_\{\\\{i,j\\\}\\subset S\}\\mathbb\{E\}\[\\mathrm\{JSD\}\_\{2\}\(p\_\{i\},p\_\{j\}\)\]\.\(17\)The event definingKSK\_\{S\}evaluateswi,wjw\_\{i\},w\_\{j\}only when both members are incorrect\. HereJSJ\_\{S\}is the population counterpart of the task\-specificHIdist\\mathrm\{HI\}\_\{\\mathrm\{dist\}\}\. The pairwise terms inDSD\_\{S\}are determined by the marginal qualities and Yule’sQQon nondegenerate correctness tables, by[LemmaC\.1](https://arxiv.org/html/2609.38274#A3.Thmtheorem1)\.
#### Divergence constants\.
Let
g\(x,z\)=12\[xlog22xx\+z\+zlog22zx\+z\]\.g\(x,z\)=\\frac\{1\}\{2\}\\left\[x\\log\_\{2\}\\frac\{2x\}\{x\+z\}\+z\\log\_\{2\}\\frac\{2z\}\{x\+z\}\\right\]\.The two relevant divergences are
d=2g\(b,c\)\>0,h=g\(b,u\)\+2g\(c,u\)\>0,θ=2hd\.d=2g\(b,c\)\>0,\\qquad h=g\(b,u\)\+2g\(c,u\)\>0,\\qquad\\theta=\\frac\{2h\}\{d\}\.\(18\)The true\-answer coordinate contributes zero to both divergences\. Two correct members, or two incorrect members preferring the same answer, have identical distributions\. A correct–incorrect pair contributeshh, and two incorrect members preferring different answers contributedd\.
###### Lemma C\.2\(Positive coefficients in the joint decomposition\)\.
Forb\>c\>0b\>c\>0andu=\(b\+2c\)/3u=\(b\+2c\)/3, the constants in \([18](https://arxiv.org/html/2609.38274#A3.E18)\) satisfy0<2h<d0<2h<d\.
###### Proof\.
Positivity follows from the distinct templates\. To proved−2h\>0d\-2h\>0, sett=b/c\>1t=b/c\>1andφ\(z\)=zlnz\\varphi\(z\)=z\\ln z\. Expanding the scalar divergences and cancelling the terms containinglnc\\ln cgives
d−2h=cln2F\(t\),d\-2h=\\frac\{c\}\{\\ln 2\}F\(t\),where
F\(t\)=2φ\(2t\+13\)\+4φ\(t\+56\)−3φ\(t\+23\)−2φ\(t\+12\)\.F\(t\)=2\\varphi\\\!\\left\(\\frac\{2t\+1\}\{3\}\\right\)\+4\\varphi\\\!\\left\(\\frac\{t\+5\}\{6\}\\right\)\-3\\varphi\\\!\\left\(\\frac\{t\+2\}\{3\}\\right\)\-2\\varphi\\\!\\left\(\\frac\{t\+1\}\{2\}\\right\)\.Direct differentiation yieldsF\(1\)=F′\(1\)=0F\(1\)=F^\{\\prime\}\(1\)=0and
F′′\(t\)=−2t2\+7t\+13\(t\+2\)\(t\+1\)\(2t\+1\)\(t\+5\)\.F^\{\\prime\\prime\}\(t\)=\\frac\{\-2t^\{2\}\+7t\+13\}\{\(t\+2\)\(t\+1\)\(2t\+1\)\(t\+5\)\}\.\(19\)This derivative is positive for1<t<t∗=\(7\+153\)/41<t<t\_\{\*\}=\(7\+\\sqrt\{153\}\)/4and negative fort\>t∗t\>t\_\{\*\}\. ThusF′F^\{\\prime\}first increases from zero and then decreases to
limt→∞F′\(t\)=53ln2−ln3\>0,\\lim\_\{t\\to\\infty\}F^\{\\prime\}\(t\)=\\frac\{5\}\{3\}\\ln 2\-\\ln 3\>0,where the final inequality is equivalent to32\>2732\>27\. ConsequentlyF′\(t\)\>0F^\{\\prime\}\(t\)\>0for everyt\>1t\>1, and integration from11givesF\(t\)\>0F\(t\)\>0\. Dividing0<2h<d0<2h<dbyddproves0<θ<10<\\theta<1\. ∎
###### Proof of Proposition[3\.1](https://arxiv.org/html/2609.38274#S3.Thmtheorem1)\.
For a pair\(i,j\)\(i,j\), writeDij=ℙ\(Ci=Cj=0\)D\_\{ij\}=\\mathbb\{P\}\(C\_\{i\}=C\_\{j\}=0\)andKij=ℙ\(Ci=Cj=0,wi=wj\)K\_\{ij\}=\\mathbb\{P\}\(C\_\{i\}=C\_\{j\}=0,w\_\{i\}=w\_\{j\}\)\. The probability that exactly one member is incorrect is\(1−qi\)\+\(1−qj\)−2Dij\(1\-q\_\{i\}\)\+\(1\-q\_\{j\}\)\-2D\_\{ij\}\. The pairwise divergence calculation above therefore gives
𝔼\[JSD2\(pi,pj\)\]=h\[\(1−qi\)\+\(1−qj\)−2Dij\]\+d\(Dij−Kij\)\.\\mathbb\{E\}\[\\mathrm\{JSD\}\_\{2\}\(p\_\{i\},p\_\{j\}\)\]=h\\bigl\[\(1\-q\_\{i\}\)\+\(1\-q\_\{j\}\)\-2D\_\{ij\}\\bigr\]\+d\(D\_\{ij\}\-K\_\{ij\}\)\.\(20\)Each member occurs in two of the three unordered pairs\. Averaging yieldsJS=2h\(eS−DS\)\+d\(DS−KS\)J\_\{S\}=2h\(e\_\{S\}\-D\_\{S\}\)\+d\(D\_\{S\}\-K\_\{S\}\)\. Rearranging and using[LemmaC\.2](https://arxiv.org/html/2609.38274#A3.Thmtheorem2)establishes \([12](https://arxiv.org/html/2609.38274#S3.E12)\)\.
It remains to characterize aggregation\. Every member assignsaato the true answer\. If all three are incorrect and prefer the same wrong answer, that answer has scoreb\>ab\>aunder Choice\-Soft andb3\>a3b^\{3\}\>a^\{3\}under PoE\. Both aggregators therefore err\.
In every other case, fix an incorrect answerrr\. If at least one member is correct, the three probabilities assigned torrhave sum at most2b\+u2b\+uand product at mostb2ub^\{2\}u\. If all members are incorrect but their preferred answers are not all equal, at most two assignbbtorr, so the corresponding bounds are2b\+c<2b\+u2b\+c<2b\+uandb2c<b2ub^\{2\}c<b^\{2\}u\. Thus the true answer has a strictly larger Choice\-Soft score by \([11](https://arxiv.org/html/2609.38274#S3.E11)\)\. For PoE, the same condition and the arithmetic–geometric mean inequality imply
a3\>\(2b\+u3\)3≥b2u\.a^\{3\}\>\\left\(\\frac\{2b\+u\}\{3\}\\right\)^\{3\}\\geq b^\{2\}u\.Hence both aggregators err exactly on the event that all three members prefer the same wrong answer\. On that event, every pair contributes one to the same\-wrong\-answer indicator\. Pointwise,
𝟏\{y^S≠y\}≤13∑\{i,j\}⊂S𝟏\{Ci=Cj=0,wi=wj\}\.\\mathbf\{1\}\\\{\\hat\{y\}\_\{S\}\\neq y\\\}\\leq\\frac\{1\}\{3\}\\sum\_\{\\\{i,j\\\}\\subset S\}\\mathbf\{1\}\\\{C\_\{i\}=C\_\{j\}=0,\\ w\_\{i\}=w\_\{j\}\\\}\.Taking expectations givesℙ\(y^S≠Y\)≤KS\\mathbb\{P\}\(\\hat\{y\}\_\{S\}\\neq Y\)\\leq K\_\{S\}\. ∎
### C\.3Accuracy Gains over Quality\-Only Selection
We derive a lower bound on the baseline’s error probability to accompany the upper bound in[Proposition3\.1](https://arxiv.org/html/2609.38274#S3.Thmtheorem1)\. Their combination gives an accuracy comparison between teams\.
###### Lemma C\.3\(Error bounds from pairwise profiles\)\.
Under \([11](https://arxiv.org/html/2609.38274#S3.E11)\), letℛS\\mathcal\{R\}\_\{S\}be the error probability of either Choice\-Soft or PoE\. Then
max\{0,DS\+3KS−2eS2\}≤ℛS≤min\{KS,1−3eS\+3DS\}\.\\max\\left\\\{0,\\frac\{D\_\{S\}\+3K\_\{S\}\-2e\_\{S\}\}\{2\}\\right\\\}\\leq\\mathcal\{R\}\_\{S\}\\leq\\min\\\{K\_\{S\},\\ 1\-3e\_\{S\}\+3D\_\{S\}\\\}\.\(21\)Both endpoints are attainable for every feasible triple\(eS,DS,KS\)\(e\_\{S\},D\_\{S\},K\_\{S\}\)\. In particular,
0≤KS−ℛS≤εS:=min\{KS,eS−DS\+KS2\}\.0\\leq K\_\{S\}\-\\mathcal\{R\}\_\{S\}\\leq\\varepsilon\_\{S\}:=\\min\\left\\\{K\_\{S\},\\ e\_\{S\}\-\\frac\{D\_\{S\}\+K\_\{S\}\}\{2\}\\right\\\}\.\(22\)
###### Proof\.
Suppress the team subscript\. There are seven possible error patterns, summarized below\. The last four columns give the contribution of each pattern to the mean member error, double\-fault probability, same\-wrong\-answer probability, and aggregation error\.
ProbabilityError patterneeDDKKℛ\\mathcal\{R\}p0p\_\{0\}No member incorrect00000000p1p\_\{1\}One member incorrect1/31/3000000p2sp\_\{2s\}Two incorrect, same answer2/32/31/31/31/31/300p2dp\_\{2d\}Two incorrect, different answers2/32/31/31/30000p3sp\_\{3s\}Three incorrect, same answer11111111p21p\_\{21\}Three incorrect, two answers11111/31/300p111p\_\{111\}Three incorrect, three answers11110000These probabilities are nonnegative and sum to one\. The table gives
e\\displaystyle e=p1\+2p2s\+2p2d3\+p3s\+p21\+p111,\\displaystyle=\\frac\{p\_\{1\}\+2p\_\{2s\}\+2p\_\{2d\}\}\{3\}\+p\_\{3s\}\+p\_\{21\}\+p\_\{111\},D\\displaystyle D=p2s\+p2d3\+p3s\+p21\+p111,\\displaystyle=\\frac\{p\_\{2s\}\+p\_\{2d\}\}\{3\}\+p\_\{3s\}\+p\_\{21\}\+p\_\{111\},K\\displaystyle K=p2s3\+p3s\+p213,ℛ=p3s\.\\displaystyle=\\frac\{p\_\{2s\}\}\{3\}\+p\_\{3s\}\+\\frac\{p\_\{21\}\}\{3\},\\qquad\\mathcal\{R\}=p\_\{3s\}\.\(23\)In particular,K−ℛ=\(p2s\+p21\)/3≥0K\-\\mathcal\{R\}=\(p\_\{2s\}\+p\_\{21\}\)/3\\geq 0\. A second identity is
e−D\+K2−\(K−ℛ\)=p13\+p2d\+p1112≥0\.e\-\\frac\{D\+K\}\{2\}\-\(K\-\\mathcal\{R\}\)=\\frac\{p\_\{1\}\}\{3\}\+\\frac\{p\_\{2d\}\+p\_\{111\}\}\{2\}\\geq 0\.Sinceℛ≥0\\mathcal\{R\}\\geq 0, these two identities imply \([22](https://arxiv.org/html/2609.38274#A3.E22)\) and the lower bound in \([21](https://arxiv.org/html/2609.38274#A3.E21)\)\. Finally,1−3e\+3D=p0\+p3s\+p21\+p111≥ℛ1\-3e\+3D=p\_\{0\}\+p\_\{3s\}\+p\_\{21\}\+p\_\{111\}\\geq\\mathcal\{R\}gives the second upper bound\.
To show attainability, fix any feasible\(e,D,K\)\(e,D,K\)and choose anyffbetween the two endpoints in \([21](https://arxiv.org/html/2609.38274#A3.E21)\)\. Set
τ=max\{f,2D−e,0\},z=max\{0,3\(K−D\+τ−f\)\}\.\\tau=\\max\\\{f,\\,2D\-e,\\,0\\\},\\qquad z=\\max\\\{0,\\,3\(K\-D\+\\tau\-f\)\\\}\.Feasibility gives0≤K≤D≤e0\\leq K\\leq D\\leq eandD≥2e−1D\\geq 2e\-1\. The bounds onffthen imply
f≤τ≤min\{D,1−3e\+3D,f\+32\(D−K\)\}\.f\\leq\\tau\\leq\\min\\\{D,\\ 1\-3e\+3D,\\ f\+\\tfrac\{3\}\{2\}\(D\-K\)\\\}\.For the final inequality, the only additional comparison is2D−e≤f\+32\(D−K\)2D\-e\\leq f\+\\tfrac\{3\}\{2\}\(D\-K\), which is exactly the stated lower bound onff\. It follows that0≤z≤min\{τ−f,3\(K−f\)\}0\\leq z\\leq\\min\\\{\\tau\-f,\\,3\(K\-f\)\\\}\. Define
p3s\\displaystyle p\_\{3s\}=f,\\displaystyle=f,p21\\displaystyle p\_\{21\}=z,\\displaystyle=z,p111\\displaystyle p\_\{111\}=τ−f−z,\\displaystyle=\\tau\-f\-z,p2s\\displaystyle p\_\{2s\}=3\(K−f\)−z,\\displaystyle=3\(K\-f\)\-z,p2d\\displaystyle p\_\{2d\}=3\(D−τ\)−p2s,\\displaystyle=3\(D\-\\tau\)\-p\_\{2s\},p1\\displaystyle p\_\{1\}=3e−6D\+3τ,\\displaystyle=3e\-6D\+3\\tau,p0\\displaystyle p\_\{0\}=1−3e\+3D−τ\.\\displaystyle=1\-3e\+3D\-\\tau\.Every term is nonnegative; forp2dp\_\{2d\}this follows fromz≥3\(K−D\+τ−f\)z\\geq 3\(K\-D\+\\tau\-f\)\. They sum to one, and substitution into \([23](https://arxiv.org/html/2609.38274#A3.E23)\) recovers\(e,D,K,ℛ\)=\(e,D,K,f\)\(e,D,K,\\mathcal\{R\}\)=\(e,D,K,f\)\. Each pattern is realizable using the three incorrect choices\. Thus the endpoints are sharp given these aggregate pairwise quantities\. ∎
###### Proof of Proposition[3\.2](https://arxiv.org/html/2609.38274#S3.Thmtheorem2)\.
WriteℛS=1−Acc\(S\)\\mathcal\{R\}\_\{S\}=1\-\\operatorname\{Acc\}\(S\), with the same aggregator for both teams\. By[LemmaC\.3](https://arxiv.org/html/2609.38274#A3.Thmtheorem3),ℛQ≥KQ−εQ\\mathcal\{R\}\_\{Q\}\\geq K\_\{Q\}\-\\varepsilon\_\{Q\}andℛS≤KS\\mathcal\{R\}\_\{S\}\\leq K\_\{S\}\. Therefore
Acc\(S\)−Acc\(SQ\)=ℛQ−ℛS≥KQ−KS−εQ\.\\operatorname\{Acc\}\(S\)\-\\operatorname\{Acc\}\(S\_\{Q\}\)=\\mathcal\{R\}\_\{Q\}\-\\mathcal\{R\}\_\{S\}\\geq K\_\{Q\}\-K\_\{S\}\-\\varepsilon\_\{Q\}\.Applying \([12](https://arxiv.org/html/2609.38274#S3.E12)\) to each team and usingeQ−eS=−\(q¯Q−q¯S\)e\_\{Q\}\-e\_\{S\}=\-\(\\bar\{q\}\_\{Q\}\-\\bar\{q\}\_\{S\}\)gives \([13](https://arxiv.org/html/2609.38274#S3.E13)\)\. Its positive right\-hand side implies a strict accuracy improvement\. If the baseline members make identical predictions, thenDQ=KQ=eQD\_\{Q\}=K\_\{Q\}=e\_\{Q\}, henceεQ=0\\varepsilon\_\{Q\}=0\. ∎
The correction quantifies how much pairwise error concentration can overstate the baseline’s aggregate error\. Its exact excess is\(p2s\+p21\)/3\(p\_\{2s\}\+p\_\{21\}\)/3: these are cases with a same\-wrong\-answer pair that the third member or soft aggregation resolves\. The bound on this excess uses only\(eQ,DQ,KQ\)\(e\_\{Q\},D\_\{Q\},K\_\{Q\}\), withKQK\_\{Q\}supplied by the joint decomposition\. The full interval in \([21](https://arxiv.org/html/2609.38274#A3.E21)\) also gives the sharper comparison
Acc\(S\)−Acc\(SQ\)≥max\{0,DQ\+3KQ−2eQ2\}−min\{KS,1−3eS\+3DS\}\.\\operatorname\{Acc\}\(S\)\-\\operatorname\{Acc\}\(S\_\{Q\}\)\\geq\\max\\left\\\{0,\\frac\{D\_\{Q\}\+3K\_\{Q\}\-2e\_\{Q\}\}\{2\}\\right\\\}\-\\min\\\{K\_\{S\},\\ 1\-3e\_\{S\}\+3D\_\{S\}\\\}\.Equation \([13](https://arxiv.org/html/2609.38274#S3.E13)\) separates this comparison into the two complementarity gains and the member\-quality difference\.
#### Perturbations of the probability templates\.
The joint mechanism extends to a neighborhood of the common templates\. Suppose each actual distribution is within total variationδ\\deltaof its template on every example, where
0≤δ<c,2δ<min\{b−a,a−2b\+u3\}\.0\\leq\\delta<c,\\qquad 2\\delta<\\min\\left\\\{b\-a,\\ a\-\\frac\{2b\+u\}\{3\}\\right\\\}\.\(24\)A coordinate changes by at mostδ\\delta\. An incorrect member retains its preferred answer becauseb−a\>2δb\-a\>2\\delta; a correct member retains its preferred answer becausea−u\>a−\(2b\+u\)/3\>2δa\-u\>a\-\(2b\+u\)/3\>2\\delta\. For every pattern with successful aggregation, the correct Choice\-Soft score is at leasta−δa\-\\deltaand every wrong score is at most\(2b\+u\)/3\+δ\(2b\+u\)/3\+\\delta\. The same inequality gives\(a−δ\)3\>\(b\+δ\)2\(u\+δ\)\(a\-\\delta\)^\{3\}\>\(b\+\\delta\)^\{2\}\(u\+\\delta\)by the arithmetic–geometric mean inequality, preserving PoE recovery\. The common wrong answer still wins in the remaining pattern\. Thuse,D,Ke,D,Kand both aggregation accuracies are unchanged\.
By the JSD gradient argument in \([49](https://arxiv.org/html/2609.38274#A3.E49)\), the observed divergenceJS′J^\{\\prime\}\_\{S\}satisfies
\|JS′−JS\|≤ηδ,ηδ=2δlog21c−δ\.\|J^\{\\prime\}\_\{S\}\-J\_\{S\}\|\\leq\\eta\_\{\\delta\},\\qquad\\eta\_\{\\delta\}=2\\delta\\log\_\{2\}\\frac\{1\}\{c\-\\delta\}\.Consequently, definingKS′=θeS\+\(1−θ\)DS−JS′/dK^\{\\prime\}\_\{S\}=\\theta e\_\{S\}\+\(1\-\\theta\)D\_\{S\}\-J^\{\\prime\}\_\{S\}/dgives\|KS′−KS\|≤ηδ/d\|K^\{\\prime\}\_\{S\}\-K\_\{S\}\|\\leq\\eta\_\{\\delta\}/dandℛS≤KS′\+ηδ/d\\mathcal\{R\}\_\{S\}\\leq K^\{\\prime\}\_\{S\}\+\\eta\_\{\\delta\}/d\. The mapK↦min\{K,e−\(D\+K\)/2\}K\\mapsto\\min\\\{K,e\-\(D\+K\)/2\\\}is11\-Lipschitz\. After clipping its value below at zero, the baseline correction computed withKQ′K^\{\\prime\}\_\{Q\}differs fromεQ\\varepsilon\_\{Q\}by at mostηδ/d\\eta\_\{\\delta\}/d\. ReplacingJS,JQ,εQJ\_\{S\},J\_\{Q\},\\varepsilon\_\{Q\}in \([13](https://arxiv.org/html/2609.38274#S3.E13)\) by these perturbed quantities therefore preserves the lower bound after subtracting3ηδ/d3\\eta\_\{\\delta\}/d\.
### C\.4Soft Recovery from Dispersed Incorrect Choices
This section isolates recovery on examples where every member is incorrect, extending the mechanism to general team sizes, larger label sets, and perturbed choice distributions\. We work on a fixed distribution of labeled examples; the same calculations apply to the uniform distribution on a finite profiling set\. Throughout this section, Jensen–Shannon divergence is measured in bits\([Lin, 1991](https://arxiv.org/html/2609.38274#bib.bib47)\):
JSD2\(p,q\)=12∑rprlog22prpr\+qr\+12∑rqrlog22qrpr\+qr\.\\mathrm\{JSD\}\_\{2\}\(p,q\)=\\frac\{1\}\{2\}\\sum\_\{r\}p\_\{r\}\\log\_\{2\}\\frac\{2p\_\{r\}\}\{p\_\{r\}\+q\_\{r\}\}\+\\frac\{1\}\{2\}\\sum\_\{r\}q\_\{r\}\\log\_\{2\}\\frac\{2q\_\{r\}\}\{p\_\{r\}\+q\_\{r\}\}\.\(25\)
#### Incorrect\-choice occupancy and aggregation margins\.
Considerk≥2k\\geq 2members andL≥4L\\geq 4choices\. On an example with true labelyy, memberiihas an incorrect preferred choicewi≠yw\_\{i\}\\neq yand distribution
pi\(r\)=\{a,r=y,b,r=wi,c,r∉\{y,wi\},0<c<a<b,a\+b\+\(L−2\)c=1\.p\_\{i\}\(r\)=\\begin\{cases\}a,&r=y,\\\\ b,&r=w\_\{i\},\\\\ c,&r\\notin\\\{y,w\_\{i\}\\\},\\end\{cases\}\\qquad 0<c<a<b,\\qquad a\+b\+\(L\-2\)c=1\.\(26\)Thus every member predicts incorrectly, while the true label is each member’s second choice\. For each incorrect labelrr, let
nr=∑i=1k𝟏\{wi=r\},vr=nrk,∑r≠yvr=1\.n\_\{r\}=\\sum\_\{i=1\}^\{k\}\\mathbf\{1\}\\\{w\_\{i\}=r\\\},\\qquad v\_\{r\}=\\frac\{n\_\{r\}\}\{k\},\\qquad\\sum\_\{r\\neq y\}v\_\{r\}=1\.\(27\)Define the divergence between two members with different incorrect preferred choices and the mean pairwise divergence on this example by
d\\displaystyle d=blog22bb\+c\+clog22cb\+c\>0,\\displaystyle=b\\log\_\{2\}\\frac\{2b\}\{b\+c\}\+c\\log\_\{2\}\\frac\{2c\}\{b\+c\}\>0,\(28\)jS\(x\)\\displaystyle j\_\{S\}\(x\)=1\(k2\)∑i<jJSD2\(pi,pj\)\.\\displaystyle=\\frac\{1\}\{\\binom\{k\}\{2\}\}\\sum\_\{i<j\}\\mathrm\{JSD\}\_\{2\}\(p\_\{i\},p\_\{j\}\)\.\(29\)
###### Lemma C\.4\(Occupancy and strict recovery\)\.
Under \([26](https://arxiv.org/html/2609.38274#A3.E26)\),
jS\(x\)=dk2−∑r≠ynr2k\(k−1\)=dkk−1\(1−∑r≠yvr2\)\.j\_\{S\}\(x\)=d\\,\\frac\{k^\{2\}\-\\sum\_\{r\\neq y\}n\_\{r\}^\{2\}\}\{k\(k\-1\)\}=d\\,\\frac\{k\}\{k\-1\}\\left\(1\-\\sum\_\{r\\neq y\}v\_\{r\}^\{2\}\\right\)\.\(30\)Choice\-Soft assigns the true label a strictly larger score than every incorrect label if and only if
maxr≠yvr<a−cb−c\.\\max\_\{r\\neq y\}v\_\{r\}<\\frac\{a\-c\}\{b\-c\}\.\(31\)For PoE, the corresponding necessary and sufficient condition is
maxr≠yvr<log2\(a/c\)log2\(b/c\)\.\\max\_\{r\\neq y\}v\_\{r\}<\\frac\{\\log\_\{2\}\(a/c\)\}\{\\log\_\{2\}\(b/c\)\}\.\(32\)Equality in either threshold gives a tie between the true label and a highest scoring incorrect label for that aggregator\.
###### Proof\.
Whenwi=wjw\_\{i\}=w\_\{j\}, the two distributions coincide\. Whenwi≠wjw\_\{i\}\\neq w\_\{j\}, they differ only at these two incorrect labels, with probabilities\(b,c\)\(b,c\)and\(c,b\)\(c,b\)\. Substitution into \([25](https://arxiv.org/html/2609.38274#A3.E25)\) gives
JSD2\(pi,pj\)=d1\{wi≠wj\}\.\\mathrm\{JSD\}\_\{2\}\(p\_\{i\},p\_\{j\}\)=d\\,\\mathbf\{1\}\\\{w\_\{i\}\\neq w\_\{j\}\\\}\.\(33\)The number of unordered pairs with different incorrect preferred choices is
\(k2\)−∑r≠y\(nr2\)=k2−∑r≠ynr22,\\binom\{k\}\{2\}\-\\sum\_\{r\\neq y\}\\binom\{n\_\{r\}\}\{2\}=\\frac\{k^\{2\}\-\\sum\_\{r\\neq y\}n\_\{r\}^\{2\}\}\{2\},which proves \([30](https://arxiv.org/html/2609.38274#A3.E30)\)\.
Choice\-Soft assigns probabilityaato the true label andc\+\(b−c\)vrc\+\(b\-c\)v\_\{r\}to incorrect labelrr\. Its margin against its strongest incorrect competitor is therefore
a−c−\(b−c\)maxr≠yvr\.a\-c\-\(b\-c\)\\max\_\{r\\neq y\}v\_\{r\}\.\(34\)This margin is positive exactly under \([31](https://arxiv.org/html/2609.38274#A3.E31)\)\. For PoE, the unnormalized scores areaka^\{k\}for the true label andbnrck−nrb^\{n\_\{r\}\}c^\{k\-n\_\{r\}\}for incorrect labelrr\. The base\-two log score ratio, divided bykk, against the strongest incorrect competitor is
log2\(a/c\)−maxr≠yvrlog2\(b/c\)\.\\log\_\{2\}\(a/c\)\-\\max\_\{r\\neq y\}v\_\{r\}\\log\_\{2\}\(b/c\)\.\(35\)Sinceb\>c\>0b\>c\>0, this expression is positive exactly under \([32](https://arxiv.org/html/2609.38274#A3.E32)\)\. The common positive PoE normalizer does not affect the comparison\. These expressions also establish the statements about ties\. ∎
The divergence records the squared concentration∑rvr2\\sum\_\{r\}v\_\{r\}^\{2\}of incorrect preferred choices, whereas recovery depends on their largest concentrationmaxrvr\\max\_\{r\}v\_\{r\}\. This distinction explains both the information supplied by choice\-space divergence and the information lost by averaging it\. Within this family, the PoE threshold is larger than the Choice\-Soft threshold: strict concavity oflog2\\log\_\{2\}gives
log2\(a/c\)log2\(b/c\)\>a−cb−c\(c<a<b\)\.\\frac\{\\log\_\{2\}\(a/c\)\}\{\\log\_\{2\}\(b/c\)\}\>\\frac\{a\-c\}\{b\-c\}\\qquad\(c<a<b\)\.
#### Three\-member recovery and sharp bounds\.
###### Proposition C\.5\(Distributional heterogeneity and recovery\)\.
Letk=3k=3and suppose \([26](https://arxiv.org/html/2609.38274#A3.E26)\) holds on the eventHHthat all three members predict incorrectly, withℙ\(H\)\>0\\mathbb\{P\}\(H\)\>0anda\>\(2b\+c\)/3a\>\(2b\+c\)/3\. DefineU=𝔼\[jS\(X\)∣H\]/dU=\\mathbb\{E\}\[j\_\{S\}\(X\)\\mid H\]/d, withjS,dj\_\{S\},dfrom \([29](https://arxiv.org/html/2609.38274#A3.E29)\) and \([28](https://arxiv.org/html/2609.38274#A3.E28)\)\. For either Choice\-Soft or PoE, the conditional recovery probabilityG=ℙ\(y^S=Y∣H\)G=\\mathbb\{P\}\(\\hat\{y\}\_\{S\}=Y\\mid H\)satisfies the sharp bounds
U≤G≤min\{1,3U/2\}\.U\\leq G\\leq\\min\\\{1,3U/2\\\}\.\(36\)Both aggregators recover exactly when the three preferred wrong answers are not all identical\.
###### Proof of Proposition[C\.5](https://arxiv.org/html/2609.38274#A3.Thmtheorem5)\.
Condition onHH, the event that all three members predict incorrectly\. If all three incorrect preferred choices coincide, that choice receives scoreb\>ab\>aunder Choice\-Soft and unnormalized scoreb3\>a3b^\{3\}\>a^\{3\}under PoE, so both aggregators fail\. Otherwise, no incorrect label is preferred by more than two members\. The conditiona\>\(2b\+c\)/3a\>\(2b\+c\)/3makes the Choice\-Soft margin strictly positive even for an incorrect label preferred by two members\. Moreover,
a3\>\(2b\+c3\)3≥b2c,a^\{3\}\>\\left\(\\frac\{2b\+c\}\{3\}\\right\)^\{3\}\\geq b^\{2\}c,so PoE also recovers in this case\. Thus both aggregators recover onHHexactly when the three incorrect preferred choices are not all equal\.
Letr3r\_\{3\},r21r\_\{21\}, andr111r\_\{111\}denote, conditional onHH, the probabilities of all three choices being equal, exactly two being equal, and all three being different, respectively\. They are nonnegative and sum to one\. Equation \([33](https://arxiv.org/html/2609.38274#A3.E33)\) gives mean pairwise divergences00,2d/32d/3, andddin these three cases\. Consequently, the quantitiesUUandGGin the proposition satisfy
U=23r21\+r111,G=r21\+r111\.U=\\frac\{2\}\{3\}r\_\{21\}\+r\_\{111\},\\qquad G=r\_\{21\}\+r\_\{111\}\.\(37\)In particular,
G−U=13r21≥0,32U−G=12r111≥0,G\-U=\\frac\{1\}\{3\}r\_\{21\}\\geq 0,\\qquad\\frac\{3\}\{2\}U\-G=\\frac\{1\}\{2\}r\_\{111\}\\geq 0,andG≤1G\\leq 1, provingU≤G≤min\{1,3U/2\}U\\leq G\\leq\\min\\\{1,3U/2\\\}\.
Both endpoints are attainable for everyU∈\[0,1\]U\\in\[0,1\]\. The lower endpoint follows from\(r3,r21,r111\)=\(1−U,0,U\)\(r\_\{3\},r\_\{21\},r\_\{111\}\)=\(1\-U,0,U\)\. For the upper endpoint, take
\(r3,r21,r111\)=\{\(1−3U/2,3U/2,0\),0≤U≤2/3,\(0,3\(1−U\),3U−2\),2/3≤U≤1\.\(r\_\{3\},r\_\{21\},r\_\{111\}\)=\\begin\{cases\}\(1\-3U/2,\\,3U/2,\\,0\),&0\\leq U\\leq 2/3,\\\\ \(0,\\,3\(1\-U\),\\,3U\-2\),&2/3\\leq U\\leq 1\.\\end\{cases\}\(38\)There are at least three incorrect labels becauseL≥4L\\geq 4, so all three patterns can be realized\. These constructions use the samea,b,ca,b,cand can be combined with any fixed distribution of predictions outsideHH\.
Finally, changing the joint pattern of\(w1,w2,w3\)\(w\_\{1\},w\_\{2\},w\_\{3\}\)onHHleaves each correctness indicator equal to zero there\. WithHHand all predictions outsideHHfixed, the entire correctness vector is unchanged\. Hence every member’s accuracy and every pair’s correctness contingency table are unchanged, as is Yule’sQQ, including the zero\-denominator convention in \([4](https://arxiv.org/html/2609.38274#S3.E4)\)\. ∎
#### Relation to overall accuracy and the profiled divergence\.
Letτ=ℙ\(H\)\>0\\tau=\\mathbb\{P\}\(H\)\>0, and write
Jall=𝔼\[jS\(X\)\],Jout=𝔼\[𝟏\{Hc\}jS\(X\)\]\.J\_\{\\mathrm\{all\}\}=\\mathbb\{E\}\[j\_\{S\}\(X\)\],\\qquad J\_\{\\mathrm\{out\}\}=\\mathbb\{E\}\[\\mathbf\{1\}\\\{H^\{c\}\\\}j\_\{S\}\(X\)\]\.\(39\)HereJallJ\_\{\\mathrm\{all\}\}is precisely the population version ofHIdist\(S,Jt\)\\mathrm\{HI\}\_\{\\mathrm\{dist\}\}\(S;J^\{t\}\)for the fixed task distribution\. For aggregator𝒜∈\{cs,poe\}\\mathcal\{A\}\\in\\\{\\textsc\{cs\},\\textsc\{poe\}\\\}, defineAout𝒜=ℙ\(Hc,y^𝒜\(X\)=Y\)A\_\{\\mathrm\{out\}\}^\{\\mathcal\{A\}\}=\\mathbb\{P\}\(H^\{c\},\\hat\{y\}\_\{\\mathcal\{A\}\}\(X\)=Y\)and letAcc𝒜=ℙ\(y^𝒜\(X\)=Y\)\\operatorname\{Acc\}\_\{\\mathcal\{A\}\}=\\mathbb\{P\}\(\\hat\{y\}\_\{\\mathcal\{A\}\}\(X\)=Y\)\. Partitioning byHHgives
Jall=Jout\+τdU,Acc𝒜=Aout𝒜\+τG\.J\_\{\\mathrm\{all\}\}=J\_\{\\mathrm\{out\}\}\+\\tau dU,\\qquad\\operatorname\{Acc\}\_\{\\mathcal\{A\}\}=A\_\{\\mathrm\{out\}\}^\{\\mathcal\{A\}\}\+\\tau G\.\(40\)The sharp conditional bounds therefore imply
Aout𝒜\+Jall−Joutd\\displaystyle A\_\{\\mathrm\{out\}\}^\{\\mathcal\{A\}\}\+\\frac\{J\_\{\\mathrm\{all\}\}\-J\_\{\\mathrm\{out\}\}\}\{d\}≤Acc𝒜\\displaystyle\\leq\\operatorname\{Acc\}\_\{\\mathcal\{A\}\}\(41\)≤Aout𝒜\+min\{τ,3\(Jall−Jout\)2d\}\.\\displaystyle\\leq A\_\{\\mathrm\{out\}\}^\{\\mathcal\{A\}\}\+\\min\\\!\\left\\\{\\tau,\\frac\{3\(J\_\{\\mathrm\{all\}\}\-J\_\{\\mathrm\{out\}\}\)\}\{2d\}\\right\\\}\.These bounds remain sharp when the predictions outsideHHare fixed\.
For example, suppose that with probabilityq∈\(0,1\)q\\in\(0,1\)all three members output the same positive distribution whose unique preferred choice is correct, and the remaining probability1−q1\-qfollows the model onHH\. Thenqi=qq\_\{i\}=qandQij=1Q\_\{ij\}=1for every pair, whileJout=0J\_\{\\mathrm\{out\}\}=0andAout𝒜=qA\_\{\\mathrm\{out\}\}^\{\\mathcal\{A\}\}=q\. Equation \([41](https://arxiv.org/html/2609.38274#A3.E41)\) reduces to
q\+Jalld≤Acc𝒜≤q\+min\{1−q,3Jall2d\}\.q\+\\frac\{J\_\{\\mathrm\{all\}\}\}\{d\}\\leq\\operatorname\{Acc\}\_\{\\mathcal\{A\}\}\\leq q\+\\min\\\!\\left\\\{1\-q,\\frac\{3J\_\{\\mathrm\{all\}\}\}\{2d\}\\right\\\}\.\(42\)Thus the existing choice\-space divergence can quantify recovery opportunities while the full correctness profile is held fixed\. The calculation applies task by task; the pooledHIdist\\mathrm\{HI\}\_\{\\mathrm\{dist\}\}is the arithmetic mean of the corresponding taskwise divergences\.
#### Stability under perturbations of member distributions\.
The recovery mechanism persists when the members’ probabilities differ from the common template\. We use total variation in the conventionTV\(p,q\)=12‖p−q‖1\\operatorname\{TV\}\(p,q\)=\\tfrac\{1\}\{2\}\\\|p\-q\\\|\_\{1\}\.
###### Corollary C\.6\(Perturbed soft recovery\)\.
Letk=3k=3, and letpi0p\_\{i\}^\{0\}have the form \([26](https://arxiv.org/html/2609.38274#A3.E26)\) on an eventHHof positive probability\. Suppose the actual distributionspip\_\{i\}satisfy, almost surely onHH,
TV\(pi,pi0\)≤ε\(i=1,2,3\),0≤ε<c,b−a\>2ε,\\operatorname\{TV\}\(p\_\{i\},p\_\{i\}^\{0\}\)\\leq\\varepsilon\\quad\(i=1,2,3\),\\qquad 0\\leq\\varepsilon<c,\\qquad b\-a\>2\\varepsilon,\(43\)and
a−2b\+c3\>2ε\.a\-\\frac\{2b\+c\}\{3\}\>2\\varepsilon\.\(44\)Each member still has unique preferred choicewi≠yw\_\{i\}\\neq y, and both Choice\-Soft and PoE recover onHHexactly when\(w1,w2,w3\)\(w\_\{1\},w\_\{2\},w\_\{3\}\)are not all equal\. LetGGbe this conditional recovery probability and set
JH=𝔼\[13∑i<jJSD2\(pi,pj\)\|H\],η=2εlog21c−ε\.J\_\{H\}=\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{3\}\\sum\_\{i<j\}\\mathrm\{JSD\}\_\{2\}\(p\_\{i\},p\_\{j\}\)\\,\\middle\|\\,H\\right\],\\qquad\\eta=2\\varepsilon\\log\_\{2\}\\frac\{1\}\{c\-\\varepsilon\}\.\(45\)Withddas in \([28](https://arxiv.org/html/2609.38274#A3.E28)\),
max\{0,JH−ηd\}≤G≤min\{1,3\(JH\+η\)2d\}\.\\max\\\!\\left\\\{0,\\frac\{J\_\{H\}\-\\eta\}\{d\}\\right\\\}\\leq G\\leq\\min\\\!\\left\\\{1,\\frac\{3\(J\_\{H\}\+\\eta\)\}\{2d\}\\right\\\}\.\(46\)For PoE alone, the conclusions hold with \([44](https://arxiv.org/html/2609.38274#A3.E44)\) replaced by
\(a−ε\)3\>\(b\+ε\)2\(c\+ε\)\.\(a\-\\varepsilon\)^\{3\}\>\(b\+\\varepsilon\)^\{2\}\(c\+\\varepsilon\)\.\(47\)
###### Proof\.
For probability distributions, a total\-variation bound ofε\\varepsilonimplies an absolute difference of at mostε\\varepsilonin every coordinate\. Therefore
pi\(wi\)≥b−ε\>a\+ε≥pi\(y\),p\_\{i\}\(w\_\{i\}\)\\geq b\-\\varepsilon\>a\+\\varepsilon\\geq p\_\{i\}\(y\),andpi\(wi\)\>pi\(r\)p\_\{i\}\(w\_\{i\}\)\>p\_\{i\}\(r\)for every other incorrect labelrr, sincea\>ca\>c\. Thus eachwiw\_\{i\}remains the unique preferred choice\. If all threewiw\_\{i\}agree, their common label strictly exceeds the true label for every member, so both the arithmetic mean and the product favor an incorrect label\.
If thewiw\_\{i\}are not all equal, any incorrect label is preferred by at most two members\. Its Choice\-Soft score is at most\(2b\+c\)/3\+ε\(2b\+c\)/3\+\\varepsilon, whereas the true label’s score is at leasta−εa\-\\varepsilon\. Condition \([44](https://arxiv.org/html/2609.38274#A3.E44)\) makes the latter strictly larger\. For PoE, any incorrect label has unnormalized score at most\(b\+ε\)2\(c\+ε\)\(b\+\\varepsilon\)^\{2\}\(c\+\\varepsilon\), while the true label has score at least\(a−ε\)3\(a\-\\varepsilon\)^\{3\}\. All factors are positive becauseε<c<a\\varepsilon<c<a\. Condition \([47](https://arxiv.org/html/2609.38274#A3.E47)\) therefore suffices\. It is also implied by \([44](https://arxiv.org/html/2609.38274#A3.E44)\), since
a−ε\>2\(b\+ε\)\+\(c\+ε\)3≥\(\(b\+ε\)2\(c\+ε\)\)1/3\.a\-\\varepsilon\>\\frac\{2\(b\+\\varepsilon\)\+\(c\+\\varepsilon\)\}\{3\}\\geq\\bigl\(\(b\+\\varepsilon\)^\{2\}\(c\+\\varepsilon\)\\bigr\)^\{1/3\}\.This proves the stated recovery rule for both aggregators\.
To bound the change in divergence, writeF\(p,q\)=JSD2\(p,q\)F\(p,q\)=\\mathrm\{JSD\}\_\{2\}\(p,q\)andmε=c−ε\>0m\_\{\\varepsilon\}=c\-\\varepsilon\>0\. Along the straight line from any template pair\(p0,q0\)\(p^\{0\},q^\{0\}\)to its perturbed pair\(p,q\)\(p,q\), every coordinate of both distributions belongs to\[mε,1\]\[m\_\{\\varepsilon\},1\]\. Direct differentiation gives
∂F∂pr=12log22prpr\+qr,∂F∂qr=12log22qrpr\+qr\.\\frac\{\\partial F\}\{\\partial p\_\{r\}\}=\\frac\{1\}\{2\}\\log\_\{2\}\\frac\{2p\_\{r\}\}\{p\_\{r\}\+q\_\{r\}\},\\qquad\\frac\{\\partial F\}\{\\partial q\_\{r\}\}=\\frac\{1\}\{2\}\\log\_\{2\}\\frac\{2q\_\{r\}\}\{p\_\{r\}\+q\_\{r\}\}\.\(48\)On this line,
mε≤2prpr\+qr≤1mε,m\_\{\\varepsilon\}\\leq\\frac\{2p\_\{r\}\}\{p\_\{r\}\+q\_\{r\}\}\\leq\\frac\{1\}\{m\_\{\\varepsilon\}\},and the same inequalities hold after interchangingprp\_\{r\}andqrq\_\{r\}\. Consequently, both gradient vectors have infinity norm at most12log2\(1/mε\)\\tfrac\{1\}\{2\}\\log\_\{2\}\(1/m\_\{\\varepsilon\}\)\. Integrating the directional derivative along the line yields
\|F\(p,q\)−F\(p0,q0\)\|\\displaystyle\|F\(p,q\)\-F\(p^\{0\},q^\{0\}\)\|≤12log21mε\(‖p−p0‖1\+‖q−q0‖1\)\\displaystyle\\leq\\frac\{1\}\{2\}\\log\_\{2\}\\frac\{1\}\{m\_\{\\varepsilon\}\}\\bigl\(\\\|p\-p^\{0\}\\\|\_\{1\}\+\\\|q\-q^\{0\}\\\|\_\{1\}\\bigr\)≤2εlog21c−ε=η\.\\displaystyle\\leq 2\\varepsilon\\log\_\{2\}\\frac\{1\}\{c\-\\varepsilon\}=\\eta\.\(49\)Averaging over the three pairs and then conditioning onHHpreserves this bound\. IfUUdenotes the normalized conditional divergence of the template distributions, we obtain\|JH−dU\|≤η\|J\_\{H\}\-dU\|\\leq\\eta\. The recovered cases are exactly the same occupancy patterns as for the templates, so \([37](https://arxiv.org/html/2609.38274#A3.E37)\) givesU≤G≤min\{1,3U/2\}U\\leq G\\leq\\min\\\{1,3U/2\\\}\. Combining these inequalities with\(JH−η\)/d≤U≤\(JH\+η\)/d\(J\_\{H\}\-\\eta\)/d\\leq U\\leq\(J\_\{H\}\+\\eta\)/dproves \([46](https://arxiv.org/html/2609.38274#A3.E46)\)\. ∎
Every template satisfying the strict condition in Proposition[C\.5](https://arxiv.org/html/2609.38274#A3.Thmtheorem5)admits a positive perturbation radius: any
0<ε<min\{c,b−a2,12\(a−2b\+c3\)\}0<\\varepsilon<\\min\\\!\\left\\\{c,\\frac\{b\-a\}\{2\},\\frac\{1\}\{2\}\\left\(a\-\\frac\{2b\+c\}\{3\}\\right\)\\right\\\}satisfies the corollary\. Hence recovery is stable on a neighborhood of member distributions, with an explicit error allowance for the observed JSD\.
### C\.5Stability of Multi\-Start Selection
We analyze two finite profiling outputs for the same target task, candidate set, team size, and heterogeneity weights\. Each output is standardized separately\. LetP=\(m2\)P=\\binom\{m\}\{2\},Lr=\(r2\)L\_\{r\}=\\binom\{r\}\{2\}, andE\(S\)=\{\{i,j\}:i,j∈S,i<j\}E\(S\)=\\\{\\\{i,j\\\}:i,j\\in S,\\ i<j\\\}\. Edge vectors contain thePPupper\-triangular entries, each counted once; all edge norms below use this convention\. For a vectorx∈ℝdx\\in\\mathbb\{R\}^\{d\}, define population standardization by
Cd=Id−𝟏𝟏⊤d,σ\(x\)=‖Cdx‖2d,Z\(x\)=\{Cdx/σ\(x\),σ\(x\)\>0,0,σ\(x\)=0\.C\_\{d\}=I\_\{d\}\-\\frac\{\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\}\{d\},\\qquad\\sigma\(x\)=\\frac\{\\\|C\_\{d\}x\\\|\_\{2\}\}\{\\sqrt\{d\}\},\\qquad Z\(x\)=\\begin\{cases\}C\_\{d\}x/\\sigma\(x\),&\\sigma\(x\)\>0,\\\\ 0,&\\sigma\(x\)=0\.\\end\{cases\}\(50\)Thus‖Z\(x\)‖2=d\\\|Z\(x\)\\\|\_\{2\}=\\sqrt\{d\}for every nonconstant vector\. The corresponding quality and combined edge signals are
ui=Z\(qt\)i,vij=λ1Z\(1−R¯\)ij\+λ2Z\(J¯\)ij,u\_\{i\}=Z\(q^\{t\}\)\_\{i\},\\qquad v\_\{ij\}=\\lambda\_\{1\}Z\(1\-\\bar\{R\}\)\_\{ij\}\+\\lambda\_\{2\}Z\(\\bar\{J\}\)\_\{ij\},\(51\)where matrix standardization acts on the upper\-triangular vector\. The quality statistics are computed over themmcandidates, and each HI channel is standardized over itsPPbase pairs\. For\|S\|=r≥2\|S\|=r\\geq 2, the implemented objective in its active normalization regime is exactly
F\(S\)=1r∑i∈Sui\+1Lr∑e∈E\(S\)ve,F\(\{i\}\)=ui\.F\(S\)=\\frac\{1\}\{\\sqrt\{r\}\}\\sum\_\{i\\in S\}u\_\{i\}\+\\frac\{1\}\{\\sqrt\{L\_\{r\}\}\}\\sum\_\{e\\in E\(S\)\}v\_\{e\},\\qquad F\(\\\{i\\\}\)=u\_\{i\}\.\(52\)In particular,uualways uses the target task’s quality estimates\. Write the second profiling output asu′,v′u^\{\\prime\},v^\{\\prime\}, its score asF′F^\{\\prime\}, and the changes asδu=u′−u\\delta u=u^\{\\prime\}\-u,δv=v′−v\\delta v=v^\{\\prime\}\-v, andδF\(S\)=F′\(S\)−F\(S\)\\delta F\(S\)=F^\{\\prime\}\(S\)\-F\(S\)\.
#### Normalization in the implementation\.
The default normalizer uses population standard deviation \(std\_ddof=0\)\. It sets a term to zero whenσ/nitems≤10−12\\sigma/\\sqrt\{n\_\{\\rm items\}\}\\leq 10^\{\-12\}\. Hence \([52](https://arxiv.org/html/2609.38274#A3.E52)\) agrees with the numerical implementation when constant channels are zero and every nonconstant channel remains active at each relevant team size\. Sufficient conditions areσ\(qt\)\>10−12k\\sigma\(q^\{t\}\)\>10^\{\-12\}\\sqrt\{k\}for a nonconstant quality channel andσ\(h\)\>10−12Lk\\sigma\(h\)\>10^\{\-12\}\\sqrt\{L\_\{k\}\}for either nonconstant pooled HI channel, in each profiling output\. A channel suppressed at every relevant size can instead be represented by zero in the effective profile\. If a channel changes activation with team size, the same comparison argument below applies to the effective size\-dependent vectorsu\(r\),v\(r\)u^\{\(r\)\},v^\{\(r\)\}, using their perturbations at sizer=s\+1r=s\+1for an extension and at sizer=kr=kfor the final comparison\.
###### Proposition C\.7\(Stability of multi\-start selection\)\.
Fix the candidates,2≤k≤m2\\leq k\\leq m, weights, and tie\-breaking, with each nonconstant standardization denominator above the numerical cutoff at every team size in both profiles\. At a reference state of sizes<ks<k, the change in the score difference between any two candidate extensions is at most
Bs=2‖δu‖2\+2‖δv‖2s\+1\.B\_\{s\}=\\frac\{\\sqrt\{2\}\\\|\\delta u\\\|\_\{2\}\+2\\\|\\delta v\\\|\_\{2\}\}\{\\sqrt\{s\+1\}\}\.\(53\)If every winning extension in every seed’s reference path leads its alternatives by more thanBsB\_\{s\}, and the final winner leads every other distinct terminal team by more than2\(‖δu‖2\+‖δv‖2\)\\sqrt\{2\}\(\\\|\\delta u\\\|\_\{2\}\+\\\|\\delta v\\\|\_\{2\}\), all paths and the returned team remain unchanged\.
###### Proof of Proposition[C\.7](https://arxiv.org/html/2609.38274#A3.Thmtheorem7)\.
*Candidate comparisons\.*At a reference stateSSwith\|S\|=s<k\|S\|=s<k, consider two remaining candidatesjjandℓ\\ell\. Define
GS\(j,ℓ\)=F\(S∪\{j\}\)−F\(S∪\{ℓ\}\)\.G\_\{S\}\(j,\\ell\)=F\(S\\cup\\\{j\\\}\)\-F\(S\\cup\\\{\\ell\\\}\)\.Both extended teams have sizes\+1s\+1\. Expanding \([52](https://arxiv.org/html/2609.38274#A3.E52)\) cancels the quality of the existing members and all edges withinSS, giving
GS\(j,ℓ\)=uj−uℓs\+1\+∑i∈S\(vij−viℓ\)Ls\+1\.G\_\{S\}\(j,\\ell\)=\\frac\{u\_\{j\}\-u\_\{\\ell\}\}\{\\sqrt\{s\+1\}\}\+\\frac\{\\sum\_\{i\\in S\}\(v\_\{ij\}\-v\_\{i\\ell\}\)\}\{\\sqrt\{L\_\{s\+1\}\}\}\.Comparing marginal gains gives the same expression, since the current scoreF\(S\)F\(S\)also cancels\. Subtracting the corresponding expression for the two profiles yields the exact perturbation
δGS\(j,ℓ\)=δuj−δuℓs\+1\+∑i∈S\(δvij−δviℓ\)Ls\+1\.\\delta G\_\{S\}\(j,\\ell\)=\\frac\{\\delta u\_\{j\}\-\\delta u\_\{\\ell\}\}\{\\sqrt\{s\+1\}\}\+\\frac\{\\sum\_\{i\\in S\}\(\\delta v\_\{ij\}\-\\delta v\_\{i\\ell\}\)\}\{\\sqrt\{L\_\{s\+1\}\}\}\.\(54\)The quality coefficient vector has two nonzero entries, each of magnitude1/s\+11/\\sqrt\{s\+1\}, and therefore norm2/\(s\+1\)\\sqrt\{2/\(s\+1\)\}\. The edge coefficient vector has2s2snonzero entries: thessedges connectingjjtoSSand thessedges connectingℓ\\elltoSS\. These edges are distinct becausej,ℓ∉Sj,\\ell\\notin S\. Its norm is
2sLs\+1=2s\+1\.\\sqrt\{\\frac\{2s\}\{L\_\{s\+1\}\}\}=\\frac\{2\}\{\\sqrt\{s\+1\}\}\.Cauchy–Schwarz and the triangle inequality therefore give
\|δGS\(j,ℓ\)\|≤2‖δu‖2\+2‖δv‖2s\+1=Bs\.\|\\delta G\_\{S\}\(j,\\ell\)\|\\leq\\frac\{\\sqrt\{2\}\\\|\\delta u\\\|\_\{2\}\+2\\\|\\delta v\\\|\_\{2\}\}\{\\sqrt\{s\+1\}\}=B\_\{s\}\.\(55\)
*Complete\-team comparisons\.*For two size\-kkteamsA,BA,B, subtraction similarly gives
δ\(F\(A\)−F\(B\)\)=\\displaystyle\\delta\\bigl\(F\(A\)\-F\(B\)\\bigr\)=\{\}∑i∈A∖Bδui−∑i∈B∖Aδuik\\displaystyle\\frac\{\\sum\_\{i\\in A\\setminus B\}\\delta u\_\{i\}\-\\sum\_\{i\\in B\\setminus A\}\\delta u\_\{i\}\}\{\\sqrt\{k\}\}\(56\)\+∑e∈E\(A\)∖E\(B\)δve−∑e∈E\(B\)∖E\(A\)δveLk\.\\displaystyle\+\\frac\{\\sum\_\{e\\in E\(A\)\\setminus E\(B\)\}\\delta v\_\{e\}\-\\sum\_\{e\\in E\(B\)\\setminus E\(A\)\}\\delta v\_\{e\}\}\{\\sqrt\{L\_\{k\}\}\}\.Letr=\|A∩B\|r=\|A\\cap B\|\. The quality difference has2\(k−r\)2\(k\-r\)nonzero coefficients\. Moreover,E\(A\)∩E\(B\)=E\(A∩B\)E\(A\)\\cap E\(B\)=E\(A\\cap B\), so each team hasLk−LrL\_\{k\}\-L\_\{r\}edges absent from the other\. The same norm calculation gives the overlap\-sensitive bound
\|δ\(F\(A\)−F\(B\)\)\|\\displaystyle\\left\|\\delta\\bigl\(F\(A\)\-F\(B\)\\bigr\)\\right\|≤2−2rk‖δu‖2\+2−2LrLk‖δv‖2\\displaystyle\\leq\\sqrt\{2\-\\frac\{2r\}\{k\}\}\\,\\\|\\delta u\\\|\_\{2\}\+\\sqrt\{2\-\\frac\{2L\_\{r\}\}\{L\_\{k\}\}\}\\,\\\|\\delta v\\\|\_\{2\}\(57\)≤2\(‖δu‖2\+‖δv‖2\)=Bfin\.\\displaystyle\\leq\\sqrt\{2\}\\bigl\(\\\|\\delta u\\\|\_\{2\}\+\\\|\\delta v\\\|\_\{2\}\\bigr\)=B\_\{\\rm fin\}\.
*Preservation of every seed run\.*Fix one seed\. Both runs start with the same singleton\. Suppose inductively that they have reached the same reference stateSS, and letjSj\_\{S\}be the reference winner at this state\. For every competing candidateℓ\\ell, the assumed reference margin and \([55](https://arxiv.org/html/2609.38274#A3.E55)\) imply
GS′\(jS,ℓ\)=GS\(jS,ℓ\)\+δGS\(jS,ℓ\)≥GS\(jS,ℓ\)−Bs\>0\.G^\{\\prime\}\_\{S\}\(j\_\{S\},\\ell\)=G\_\{S\}\(j\_\{S\},\\ell\)\+\\delta G\_\{S\}\(j\_\{S\},\\ell\)\\geq G\_\{S\}\(j\_\{S\},\\ell\)\-B\_\{s\}\>0\.The perturbed run therefore choosesjSj\_\{S\}as well\. Induction to sizekkproves that this seed reaches the same terminal team\. Applying the argument to every seed preserves all terminal teams\.
Let𝒞\\mathcal\{C\}be the set of*distinct*terminal teams in the reference run, and letSgS\_\{\\rm g\}be its returned team\. Repeated occurrences of the same team among different seeds do not create competing teams\. For anyT∈𝒞∖\{Sg\}T\\in\\mathcal\{C\}\\setminus\\\{S\_\{\\rm g\}\\\}, the assumed final margin and \([57](https://arxiv.org/html/2609.38274#A3.E57)\) give
F′\(Sg\)−F′\(T\)≥F\(Sg\)−F\(T\)−Bfin\>0\.F^\{\\prime\}\(S\_\{\\rm g\}\)\-F^\{\\prime\}\(T\)\\geq F\(S\_\{\\rm g\}\)\-F\(T\)\-B\_\{\\rm fin\}\>0\.Thus the perturbed multi\-start run returnsSgS\_\{\\rm g\}\. If there is only one distinct terminal team, this final condition is vacuous\. The induction follows exactly the seed paths and terminal\-team comparisons performed by Algorithm[3\.3](https://arxiv.org/html/2609.38274#S3.SS3)\. ∎
#### A sharper certificate using signed comparisons\.
The norm bounds can be replaced by the exact projected perturbations\. At a reference stateSS, write
wS\(j\)=δujs\+1\+∑i∈SδvijLs\+1,γS\(ℓ\)=GS\(jS,ℓ\)\.w\_\{S\}\(j\)=\\frac\{\\delta u\_\{j\}\}\{\\sqrt\{s\+1\}\}\+\\frac\{\\sum\_\{i\\in S\}\\delta v\_\{ij\}\}\{\\sqrt\{L\_\{s\+1\}\}\},\\qquad\\gamma\_\{S\}\(\\ell\)=G\_\{S\}\(j\_\{S\},\\ell\)\.The reference winner remains a strict winner at this state if and only if
maxℓ∈ℳ∖\(S∪\{jS\}\)\{wS\(ℓ\)−wS\(jS\)−γS\(ℓ\)\}<0\.\\max\_\{\\ell\\in\\mathcal\{M\}\\setminus\(S\\cup\\\{j\_\{S\}\\\}\)\}\\bigl\\\{w\_\{S\}\(\\ell\)\-w\_\{S\}\(j\_\{S\}\)\-\\gamma\_\{S\}\(\\ell\)\\bigr\\\}<0\.\(58\)Indeed, the expression inside the braces is exactlyF′\(S∪\{ℓ\}\)−F′\(S∪\{jS\}\)F^\{\\prime\}\(S\\cup\\\{\\ell\\\}\)\-F^\{\\prime\}\(S\\cup\\\{j\_\{S\}\\\}\)\. The corresponding final condition is
maxT∈𝒞∖\{Sg\}\{δF\(T\)−δF\(Sg\)−\[F\(Sg\)−F\(T\)\]\}<0\.\\max\_\{T\\in\\mathcal\{C\}\\setminus\\\{S\_\{\\rm g\}\\\}\}\\bigl\\\{\\delta F\(T\)\-\\delta F\(S\_\{\\rm g\}\)\-\[F\(S\_\{\\rm g\}\)\-F\(T\)\]\\bigr\\\}<0\.\(59\)An empty maximum is−∞\-\\infty\. These conditions retain cancellation between quality and HI perturbations and require only reference\-winner comparisons\. Applying them at every reference state, followed by \([59](https://arxiv.org/html/2609.38274#A3.E59)\), gives the same induction\. Repeated reference states can share a single condition\. Strict inequalities remove tie ambiguity; a non\-strict comparison is sufficient when a retained tie is resolved in favor of the reference choice by the fixed priority rule\.
Both versions are deterministic statements about finite profiles, including estimates that share development examples or models\. Their conditions preserve the selected team relative to the reference profile\.
#### Piecewise\-constant selection as weights vary\.
Fix the profiling outputs and vary\(λ1,λ2\)∈ℝ≥02\(\\lambda\_\{1\},\\lambda\_\{2\}\)\\in\\mathbb\{R\}\_\{\\geq 0\}^\{2\}\. Every candidate comparison and complete\-team comparison is affine in these two weights\. There are finitely many possible states and teams, so the nontrivial comparison equalities form a finite collection of lines\. Within each open cell of their arrangement, all comparison signs are fixed\. Induction over the seed runs and the final comparison then shows that the returned team is constant throughout the cell\. Identically tied comparisons are handled by the fixed priority rule\. A team’s selection region may be a union of cells and need not be convex\. This accounts for constant\-selection regions in a weight sweep; one color in an accuracy heatmap can also represent several teams with the same accuracy\.
### C\.6Pooling Geometry and Selection Perturbations
For each source taskτ∈𝒯\\tau\\in\\mathcal\{T\}, collect the upper\-triangular HI entries as
h1τ=\(1−Rijτ\)i<j,h2τ=\(Jijτ\)i<j,h¯ℓ=1\|𝒯\|∑τ∈𝒯hℓτ,ℓ∈\{1,2\}\.h\_\{1\}^\{\\tau\}=\(1\-R\_\{ij\}^\{\\tau\}\)\_\{i<j\},\\qquad h\_\{2\}^\{\\tau\}=\(J\_\{ij\}^\{\\tau\}\)\_\{i<j\},\\qquad\\bar\{h\}\_\{\\ell\}=\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}\}h\_\{\\ell\}^\{\\tau\},\\quad\\ell\\in\\\{1,2\\\}\.Thush¯1=1−R¯\\bar\{h\}\_\{1\}=1\-\\bar\{R\}follows from averaging the task\-level Yule\-QQstatistics in \([8](https://arxiv.org/html/2609.38274#S3.E8)\)\. In this notation the selector usesv=λ1Z\(h¯1\)\+λ2Z\(h¯2\)v=\\lambda\_\{1\}Z\(\\bar\{h\}\_\{1\}\)\+\\lambda\_\{2\}Z\(\\bar\{h\}\_\{2\}\)while retainingu=Z\(qt\)u=Z\(q^\{t\}\)from the target task\. All identities below use population standard deviation as in \([50](https://arxiv.org/html/2609.38274#A3.E50)\)\. Their use in the numerical selector inherits the active\-channel conditions above; a numerically suppressed channel contributes zero to its effective profile\.
#### The direction of the standardized pool\.
For one HI channel, abbreviatehτ=hℓτh^\{\\tau\}=h\_\{\\ell\}^\{\\tau\}, and set
στ=σ\(hτ\),zτ=Z\(hτ\),W=∑τ∈𝒯στzτ\.\\sigma\_\{\\tau\}=\\sigma\(h^\{\\tau\}\),\\qquad z\_\{\\tau\}=Z\(h^\{\\tau\}\),\\qquad W=\\sum\_\{\\tau\\in\\mathcal\{T\}\}\\sigma\_\{\\tau\}z\_\{\\tau\}\.Constant task vectors haveστ=0\\sigma\_\{\\tau\}=0and contribute zero toWW\. IfW≠0W\\neq 0, then
Z\(h¯\)=PW‖W‖2\.Z\(\\bar\{h\}\)=\\sqrt\{P\}\\,\\frac\{W\}\{\\\|W\\\|\_\{2\}\}\.\(60\)To prove the identity, center the average usingCPC\_\{P\}from \([50](https://arxiv.org/html/2609.38274#A3.E50)\):
CPh¯=1\|𝒯\|∑τ∈𝒯CPhτ=W\|𝒯\|,σ\(h¯\)=‖W‖2\|𝒯\|P\.C\_\{P\}\\bar\{h\}=\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}\}C\_\{P\}h^\{\\tau\}=\\frac\{W\}\{\|\\mathcal\{T\}\|\},\\qquad\\sigma\(\\bar\{h\}\)=\\frac\{\\\|W\\\|\_\{2\}\}\{\|\\mathcal\{T\}\|\\sqrt\{P\}\}\.Their ratio gives \([60](https://arxiv.org/html/2609.38274#A3.E60)\)\. IfW=0W=0, the pooled vector is constant and its standardized value is zero\. Equal weighting of the raw task vectors therefore combines their standardized directions with weights proportional to their within\-task dispersion across model pairs\. Task\-wide offsets vanish under centering, whereas opposing centered directions can cancel inWW\.
#### Positive\-affine invariance\.
Suppose, for each channelℓ\\ell, the pooled and target\-task vectors obey
h¯ℓ=αℓhℓt\+βℓ𝟏,αℓ\>0\.\\bar\{h\}\_\{\\ell\}=\\alpha\_\{\\ell\}h\_\{\\ell\}^\{t\}\+\\beta\_\{\\ell\}\\mathbf\{1\},\\qquad\\alpha\_\{\\ell\}\>0\.\(61\)ThenZ\(h¯ℓ\)=Z\(hℓt\)Z\(\\bar\{h\}\_\{\\ell\}\)=Z\(h\_\{\\ell\}^\{t\}\)\. For a nonconstant vector this follows becauseCPh¯ℓ=αℓCPhℓtC\_\{P\}\\bar\{h\}\_\{\\ell\}=\\alpha\_\{\\ell\}C\_\{P\}h\_\{\\ell\}^\{t\}andσ\(h¯ℓ\)=αℓσ\(hℓt\)\\sigma\(\\bar\{h\}\_\{\\ell\}\)=\\alpha\_\{\\ell\}\\sigma\(h\_\{\\ell\}^\{t\}\); if the target vector is constant, both standardized vectors are zero\. Consequently, selection with pooled HI and selection with target\-task HI have identical scores at every team size when both use the sameqtq^\{t\}\. Their greedy paths and returned teams are identical, including ties resolved by the same rule\. This identity applies to the numerical implementation when the scaling preserves the active or zero status of each channel at every relevant team size\.
A sufficient source\-level condition for \([61](https://arxiv.org/html/2609.38274#A3.E61)\) ishℓτ=ατℓhℓ⋆\+βτℓ𝟏h\_\{\\ell\}^\{\\tau\}=\\alpha\_\{\\tau\\ell\}h\_\{\\ell\}^\{\\star\}\+\\beta\_\{\\tau\\ell\}\\mathbf\{1\}withατℓ\>0\\alpha\_\{\\tau\\ell\}\>0for the source tasks and the target\. Averaging preserves this common centered direction\. The condition concerns each channel separately and imposes no sharing requirement on target qualitiesqtq^\{t\}\. Different target tasks can therefore select different teams\. Positive\-affine agreement preserves standardized spacings as well as pair order; agreement of pair order alone does not preserve sums of edge values\.
#### Direction changes bound the effective edge perturbation\.
Consider any two nonconstant vectorsh,h′∈ℝPh,h^\{\\prime\}\\in\\mathbb\{R\}^\{P\}, and let
ρ\(h,h′\)=Z\(h\)⊤Z\(h′\)P\.\\rho\(h,h^\{\\prime\}\)=\\frac\{Z\(h\)^\{\\top\}Z\(h^\{\\prime\}\)\}\{P\}\.Their standardized norms are bothP\\sqrt\{P\}, so expansion of the squared distance gives
‖Z\(h′\)−Z\(h\)‖22=2P\(1−ρ\(h,h′\)\)\.\\\|Z\(h^\{\\prime\}\)\-Z\(h\)\\\|\_\{2\}^\{2\}=2P\\bigl\(1\-\\rho\(h,h^\{\\prime\}\)\\bigr\)\.\(62\)When the pooled vectors are nonconstant in both profiles for each channel, letdℓ=Z\(h¯ℓ′\)−Z\(h¯ℓ\)d\_\{\\ell\}=Z\(\\bar\{h\}^\{\\prime\}\_\{\\ell\}\)\-Z\(\\bar\{h\}\_\{\\ell\}\)andρℓ=ρ\(h¯ℓ,h¯ℓ′\)\\rho\_\{\\ell\}=\\rho\(\\bar\{h\}\_\{\\ell\},\\bar\{h\}^\{\\prime\}\_\{\\ell\}\)\. Then
δv=λ1d1\+λ2d2,‖δv‖2≤2P∑ℓ=12λℓ1−ρℓ\.\\delta v=\\lambda\_\{1\}d\_\{1\}\+\\lambda\_\{2\}d\_\{2\},\\qquad\\\|\\delta v\\\|\_\{2\}\\leq\\sqrt\{2P\}\\sum\_\{\\ell=1\}^\{2\}\\lambda\_\{\\ell\}\\sqrt\{1\-\\rho\_\{\\ell\}\}\.\(63\)The direct norm‖λ1d1\+λ2d2‖2\\\|\\lambda\_\{1\}d\_\{1\}\+\\lambda\_\{2\}d\_\{2\}\\\|\_\{2\}in the selection proposition retains any cancellation between channels\. If a channel is constant, the exact differencedℓd\_\{\\ell\}remains defined by the zero convention, without requiring a correlation value\. Changing only the HI pool leavesδu=0\\delta u=0; re\-estimating the development profiles can also change the target quality vector, in which case itsδu\\delta uterm is retained\.
The pool\-to\-target alignment can also be expressed using task\-to\-task alignments\. For one channel, letρτν=zτ⊤zν/P\\rho\_\{\\tau\\nu\}=z\_\{\\tau\}^\{\\top\}z\_\{\\nu\}/Pfor nonconstant task vectors\. For a nonconstant target vector andW≠0W\\neq 0, substitution into \([60](https://arxiv.org/html/2609.38274#A3.E60)\) gives
ρ\(ht,h¯\)=∑τστρtτ∑τ,νστσνρτν,\\rho\(h^\{t\},\\bar\{h\}\)=\\frac\{\\sum\_\{\\tau\}\\sigma\_\{\\tau\}\\rho\_\{t\\tau\}\}\{\\sqrt\{\\sum\_\{\\tau,\\nu\}\\sigma\_\{\\tau\}\\sigma\_\{\\nu\}\\rho\_\{\\tau\\nu\}\}\},\(64\)where zero\-dispersion source vectors are omitted from the sums\. Indeed,zt⊤W=P∑τστρtτz\_\{t\}^\{\\top\}W=P\\sum\_\{\\tau\}\\sigma\_\{\\tau\}\\rho\_\{t\\tau\}and‖W‖22=P∑τ,νστσνρτν\\\|W\\\|\_\{2\}^\{2\}=P\\sum\_\{\\tau,\\nu\}\\sigma\_\{\\tau\}\\sigma\_\{\\nu\}\\rho\_\{\\tau\\nu\}\. Equations \([62](https://arxiv.org/html/2609.38274#A3.E62)\)–\([64](https://arxiv.org/html/2609.38274#A3.E64)\) connect task agreement to the perturbation that enters the actual selection comparisons\. Entrywise variance reduction under resampling is a different quantity: after standardization, stability also depends on the pooled direction and the decision margins\.
### C\.7From Selection Scores to Accuracy Gains
We connect the accuracy comparison in[Proposition3\.2](https://arxiv.org/html/2609.38274#S3.Thmtheorem2)to the quality and HI terms in the implemented objective\. The connection separates the curvature of the Q\-to\-error relation, the choice of objective weights, and the change from a target\-task profile to a pooled development profile\. All teams in this subsection have size three and follow \([11](https://arxiv.org/html/2609.38274#S3.E11)\)\.
#### From Yule’s Q to a positive three\-term score\.
Writeei=1−qie\_\{i\}=1\-q\_\{i\}\. For a pair with error probabilitiesx,zx,zand Yule’s coefficientrr, letf\(x,z,r\)f\(x,z,r\)denote its double\-fault probability\. On positive contingency tables it is the unique solutionssof
Φ\(s,x,z\):=lns\(1−x−z\+s\)\(x−s\)\(z−s\)=ln1\+r1−r\.\\Phi\(s,x,z\):=\\ln\\frac\{s\(1\-x\-z\+s\)\}\{\(x\-s\)\(z\-s\)\}=\\ln\\frac\{1\+r\}\{1\-r\}\.\(65\)The feasible interior ismax\{0,x\+z−1\}<s<min\{x,z\}\\max\\\{0,x\+z\-1\\\}<s<\\min\\\{x,z\\\}, with0<x,z<10<x,z<1and−1<r<1\-1<r<1\. The left\-hand side increases from−∞\-\\inftyto\+∞\+\\inftyon this interval\. Define
T=∂sΦ=1s\+11−x−z\+s\+1x−s\+1z−s\>0\.T=\\partial\_\{s\}\\Phi=\\frac\{1\}\{s\}\+\\frac\{1\}\{1\-x\-z\+s\}\+\\frac\{1\}\{x\-s\}\+\\frac\{1\}\{z\-s\}\>0\.Implicit differentiation gives
∂xf=\(1−x−z\+s\)−1\+\(x−s\)−1T∈\(0,1\),∂rf=2\(1−r2\)T\>0,\\partial\_\{x\}f=\\frac\{\(1\-x\-z\+s\)^\{\-1\}\+\(x\-s\)^\{\-1\}\}\{T\}\\in\(0,1\),\\qquad\\partial\_\{r\}f=\\frac\{2\}\{\(1\-r^\{2\}\)T\}\>0,\(66\)and the expression for∂zf\\partial\_\{z\}ffollows by symmetry\.
Choose an interior reference\(e∗,e∗,r∗\)\(e\_\{\*\},e\_\{\*\},r\_\{\*\}\)and putf∗=f\(e∗,e∗,r∗\)f\_\{\*\}=f\(e\_\{\*\},e\_\{\*\},r\_\{\*\}\),ae=∂xf\(e∗,e∗,r∗\)=∂zf\(e∗,e∗,r∗\)a\_\{e\}=\\partial\_\{x\}f\(e\_\{\*\},e\_\{\*\},r\_\{\*\}\)=\\partial\_\{z\}f\(e\_\{\*\},e\_\{\*\},r\_\{\*\}\), andar=∂rf\(e∗,e∗,r∗\)a\_\{r\}=\\partial\_\{r\}f\(e\_\{\*\},e\_\{\*\},r\_\{\*\}\)\. For each actual pair define its residual by the exact identity
Dij=f∗\+ae\[\(ei−e∗\)\+\(ej−e∗\)\]\+ar\(Rijt−r∗\)\+ρij\.D\_\{ij\}=f\_\{\*\}\+a\_\{e\}\[\(e\_\{i\}\-e\_\{\*\}\)\+\(e\_\{j\}\-e\_\{\*\}\)\]\+a\_\{r\}\(R^\{t\}\_\{ij\}\-r\_\{\*\}\)\+\\rho\_\{ij\}\.\(67\)This definition also applies to boundary tables using their actualDijD\_\{ij\}\. If the segments from the reference to\(ei,ej,Rijt\)\(e\_\{i\},e\_\{j\},R^\{t\}\_\{ij\}\)lie in a compact interior region on which‖∇2f‖op≤M\\\|\\nabla^\{2\}f\\\|\_\{\\rm op\}\\leq M, Taylor’s theorem gives the quantitative bound
\|ρij\|≤M2\[\(ei−e∗\)2\+\(ej−e∗\)2\+\(Rijt−r∗\)2\]\.\|\\rho\_\{ij\}\|\\leq\\frac\{M\}\{2\}\\left\[\(e\_\{i\}\-e\_\{\*\}\)^\{2\}\+\(e\_\{j\}\-e\_\{\*\}\)^\{2\}\+\(R^\{t\}\_\{ij\}\-r\_\{\*\}\)^\{2\}\\right\]\.\(68\)Such a finiteMMexists on every compact region of positive contingency tables, by the implicit function theorem and \([66](https://arxiv.org/html/2609.38274#A3.E66)\)\.
LetHS=HIerr\(S,Rt\)=1−13∑\{i,j\}⊂SRijtH\_\{S\}=\\mathrm\{HI\}\_\{\\mathrm\{err\}\}\(S;R^\{t\}\)=1\-\\frac\{1\}\{3\}\\sum\_\{\\\{i,j\\\}\\subset S\}R^\{t\}\_\{ij\}\. Averaging \([67](https://arxiv.org/html/2609.38274#A3.E67)\) over a team, then substituting into \([12](https://arxiv.org/html/2609.38274#S3.E12)\), yields
−KS=c0\+Aqq¯S\+AHHS\+AJJS\+ηS,\-K\_\{S\}=c\_\{0\}\+A\_\{q\}\\bar\{q\}\_\{S\}\+A\_\{H\}H\_\{S\}\+A\_\{J\}J\_\{S\}\+\\eta\_\{S\},\(69\)wherec0c\_\{0\}is independent of the team and
Aq=θ\+2\(1−θ\)ae\>0,AH=\(1−θ\)ar\>0,AJ=1/d\>0,A\_\{q\}=\\theta\+2\(1\-\\theta\)a\_\{e\}\>0,\\qquad A\_\{H\}=\(1\-\\theta\)a\_\{r\}\>0,\\qquad A\_\{J\}=1/d\>0,\(70\)with
ηS=−1−θ3∑\{i,j\}⊂Sρij,\|ηS\|≤τS:=1−θ3∑\{i,j\}⊂S\|ρij\|\.\\eta\_\{S\}=\-\\frac\{1\-\\theta\}\{3\}\\sum\_\{\\\{i,j\\\}\\subset S\}\\rho\_\{ij\},\\qquad\|\\eta\_\{S\}\|\\leq\\tau\_\{S\}:=\\frac\{1\-\\theta\}\{3\}\\sum\_\{\\\{i,j\\\}\\subset S\}\|\\rho\_\{ij\}\|\.To see the coefficient of quality explicitly, the average of\(ei−e∗\)\+\(ej−e∗\)\(e\_\{i\}\-e\_\{\*\}\)\+\(e\_\{j\}\-e\_\{\*\}\)over the three pairs is2\(eS−e∗\)2\(e\_\{S\}\-e\_\{\*\}\)\. The averageRijtR^\{t\}\_\{ij\}is1−HS1\-H\_\{S\}, andeS=1−q¯Se\_\{S\}=1\-\\bar\{q\}\_\{S\}\. These substitutions give \([70](https://arxiv.org/html/2609.38274#A3.E70)\)\. Thus the local directions of all three terms agree with the signs used in the selection objective, with a second\-order remainder in an interior neighborhood\.
#### Normalization and fixed objective weights\.
LetFtF\_\{t\}be the score computed using target\-task qualities and target\-task HI, with the base\-item normalization in \([9](https://arxiv.org/html/2609.38274#S3.E9)\)\. For size three, it has the exact affine form
Ft\(S\)=cF\+wqq¯S\+wHHS\+wJJS,wq=3σq,wH=λ13σH,wJ=λ23σJ\.F\_\{t\}\(S\)=c\_\{F\}\+w\_\{q\}\\bar\{q\}\_\{S\}\+w\_\{H\}H\_\{S\}\+w\_\{J\}J\_\{S\},\\qquad w\_\{q\}=\\frac\{\\sqrt\{3\}\}\{\\sigma\_\{q\}\},\\quad w\_\{H\}=\\frac\{\\lambda\_\{1\}\\sqrt\{3\}\}\{\\sigma\_\{H\}\},\\quad w\_\{J\}=\\frac\{\\lambda\_\{2\}\\sqrt\{3\}\}\{\\sigma\_\{J\}\}\.\(71\)Here each standard deviation is over its corresponding candidate\-pool base items, andcFc\_\{F\}collects their means\. A constant channel has coefficient zero; nonconstant channels satisfy the active\-normalization conditions in[SectionC\.5](https://arxiv.org/html/2609.38274#A3.SS5)\. Choose any scaleα\>0\\alpha\>0and define
𝐀=\(Aq,AH,AJ\),𝐰=\(wq,wH,wJ\),XS=\(q¯S,HS,JS\),𝐛=𝐀−α𝐰\.\\mathbf\{A\}=\(A\_\{q\},A\_\{H\},A\_\{J\}\),\\quad\\mathbf\{w\}=\(w\_\{q\},w\_\{H\},w\_\{J\}\),\\quad X\_\{S\}=\(\\bar\{q\}\_\{S\},H\_\{S\},J\_\{S\}\),\\quad\\mathbf\{b\}=\\mathbf\{A\}\-\\alpha\\mathbf\{w\}\.Forwq\>0w\_\{q\}\>0, takingα=Aq/wq\\alpha=A\_\{q\}/w\_\{q\}matches the quality coefficient exactly\. The scale is used only to compare the two mathematical quantities; the selector retains its original weights\. For two teams define the one\-sided coefficient correction
Ewt\(S,T\)=\[−𝐛⊤\(XS−XT\)\]\+,\[z\]\+=max\{z,0\}\.E\_\{\\rm wt\}\(S,T\)=\\bigl\[\-\\mathbf\{b\}^\{\\top\}\(X\_\{S\}\-X\_\{T\}\)\\bigr\]\_\{\+\},\\qquad\[z\]\_\{\+\}=\\max\\\{z,0\\\}\.\(72\)Subtracting \([69](https://arxiv.org/html/2609.38274#A3.E69)\) and using \([71](https://arxiv.org/html/2609.38274#A3.E71)\), together with𝐛⊤\(XS−XT\)≥−Ewt\(S,T\)\\mathbf\{b\}^\{\\top\}\(X\_\{S\}\-X\_\{T\}\)\\geq\-E\_\{\\rm wt\}\(S,T\), gives
KT−KS≥α\[Ft\(S\)−Ft\(T\)\]−Ewt\(S,T\)−τS−τT\.K\_\{T\}\-K\_\{S\}\\geq\\alpha\[F\_\{t\}\(S\)\-F\_\{t\}\(T\)\]\-E\_\{\\rm wt\}\(S,T\)\-\\tau\_\{S\}\-\\tau\_\{T\}\.\(73\)This form keeps the specified values ofλ1,λ2\\lambda\_\{1\},\\lambda\_\{2\}and measures their effect through the explicit coefficient differences𝐛\\mathbf\{b\}\.
#### Pooled development profiles and the returned team\.
WriteF^\\widehat\{F\}for the actual score based on estimated target qualities and pooled HI\. Let\(ut,vt\)\(u\_\{t\},v\_\{t\}\)and\(u^,v^\)\(\\widehat\{u\},\\widehat\{v\}\)be the corresponding standardized signals from \([51](https://arxiv.org/html/2609.38274#A3.E51)\), withRt,JtR^\{t\},J^\{t\}used forvtv\_\{t\}\. Putδu=u^−ut\\delta u=\\widehat\{u\}\-u\_\{t\}andδv=v^−vt\\delta v=\\widehat\{v\}\-v\_\{t\}\. For two teams with overlapo=\|S∩T\|o=\|S\\cap T\|, \([57](https://arxiv.org/html/2609.38274#A3.E57)\) specializes to
Bprof\(S,T\)\\displaystyle B\_\{\\rm prof\}\(S,T\):=2−2o3‖δu‖2\+2−2\(o2\)3‖δv‖2,\\displaystyle:=\\sqrt\{2\-\\frac\{2o\}\{3\}\}\\\|\\delta u\\\|\_\{2\}\+\\sqrt\{2\-\\frac\{2\\binom\{o\}\{2\}\}\{3\}\}\\\|\\delta v\\\|\_\{2\},\(74\)\|\[F^\(S\)−F^\(T\)\]−\[Ft\(S\)−Ft\(T\)\]\|\\displaystyle\\left\|\[\\widehat\{F\}\(S\)\-\\widehat\{F\}\(T\)\]\-\[F\_\{t\}\(S\)\-F\_\{t\}\(T\)\]\\right\|≤Bprof\(S,T\)\.\\displaystyle\\leq B\_\{\\rm prof\}\(S,T\)\.Every normalization is recomputed on its own profile, so this expression includes changes in both base\-item values and normalization statistics\.
###### Proposition C\.8\(Selection\-score condition for an accuracy gain\)\.
Under \([11](https://arxiv.org/html/2609.38274#S3.E11)\) and the active\-normalization conditions above, letSgS\_\{\\rm g\}be the output of multi\-start greedy selection onF^\\widehat\{F\}, and letSQS\_\{Q\}be the Quality\-Only team from the same candidate pool\. For either Choice\-Soft or PoE,
Acc\(Sg\)−Acc\(SQ\)≥\\displaystyle\\operatorname\{Acc\}\(S\_\{\\rm g\}\)\-\\operatorname\{Acc\}\(S\_\{Q\}\)\\geq\{\}α\[F^\(Sg\)−F^\(SQ\)\]−εQ\\displaystyle\\alpha\[\\widehat\{F\}\(S\_\{\\rm g\}\)\-\\widehat\{F\}\(S\_\{Q\}\)\]\-\\varepsilon\_\{Q\}\(75\)−αBprof\(Sg,SQ\)−Ewt\(Sg,SQ\)−τSg−τSQ\.\\displaystyle\-\\alpha B\_\{\\rm prof\}\(S\_\{\\rm g\},S\_\{Q\}\)\-E\_\{\\rm wt\}\(S\_\{\\rm g\},S\_\{Q\}\)\-\\tau\_\{S\_\{\\rm g\}\}\-\\tau\_\{S\_\{Q\}\}\.In particular, a score advantage exceeding the displayed correction terms guarantees a strict accuracy improvement\.
###### Proof\.
The proof of[Proposition3\.2](https://arxiv.org/html/2609.38274#S3.Thmtheorem2)givesAcc\(Sg\)−Acc\(SQ\)≥KQ−KSg−εQ\\operatorname\{Acc\}\(S\_\{\\rm g\}\)\-\\operatorname\{Acc\}\(S\_\{Q\}\)\\geq K\_\{Q\}\-K\_\{S\_\{\\rm g\}\}\-\\varepsilon\_\{Q\}\. Apply \([73](https://arxiv.org/html/2609.38274#A3.E73)\) withS=SgS=S\_\{\\rm g\}andT=SQT=S\_\{Q\}\. Then \([74](https://arxiv.org/html/2609.38274#A3.E74)\) impliesFt\(Sg\)−Ft\(SQ\)≥F^\(Sg\)−F^\(SQ\)−Bprof\(Sg,SQ\)F\_\{t\}\(S\_\{\\rm g\}\)\-F\_\{t\}\(S\_\{Q\}\)\\geq\\widehat\{F\}\(S\_\{\\rm g\}\)\-\\widehat\{F\}\(S\_\{Q\}\)\-B\_\{\\rm prof\}\(S\_\{\\rm g\},S\_\{Q\}\)\. Substitution proves \([75](https://arxiv.org/html/2609.38274#A3.E75)\)\. ∎
To make the role of multi\-start search explicit, let𝒞\\mathcal\{C\}be its set of distinct terminal teams\. Algorithm[3\.3](https://arxiv.org/html/2609.38274#S3.SS3)satisfiesF^\(Sg\)=maxT∈𝒞F^\(T\)\\widehat\{F\}\(S\_\{\\rm g\}\)=\\max\_\{T\\in\\mathcal\{C\}\}\\widehat\{F\}\(T\)\. Set
Emax=maxT∈𝒞\{αBprof\(T,SQ\)\+Ewt\(T,SQ\)\+τT\+τSQ\}\.E\_\{\\max\}=\\max\_\{T\\in\\mathcal\{C\}\}\\\{\\alpha B\_\{\\rm prof\}\(T,S\_\{Q\}\)\+E\_\{\\rm wt\}\(T,S\_\{Q\}\)\+\\tau\_\{T\}\+\\tau\_\{S\_\{Q\}\}\\\}\.Therefore the sufficient condition
α\[maxT∈𝒞F^\(T\)−F^\(SQ\)\]\>εQ\+Emax\\alpha\\left\[\\max\_\{T\\in\\mathcal\{C\}\}\\widehat\{F\}\(T\)\-\\widehat\{F\}\(S\_\{Q\}\)\\right\]\>\\varepsilon\_\{Q\}\+E\_\{\\max\}\(76\)ensures that the returned team outperforms Quality\-Only\. The maximization here is over the teams actually generated by the seed runs\.
Finally, the HI component of the profile change is exactly
δv=λ1\[Z\(1−R¯^\)−Z\(1−Rt\)\]\+λ2\[Z\(J¯^\)−Z\(Jt\)\]\.\\delta v=\\lambda\_\{1\}\[Z\(1\-\\widehat\{\\bar\{R\}\}\)\-Z\(1\-R^\{t\}\)\]\+\\lambda\_\{2\}\[Z\(\\widehat\{\\bar\{J\}\}\)\-Z\(J^\{t\}\)\]\.With exact target qualities,δu=0\\delta u=0\. Positive\-affine agreement between each pooled channel and its target counterpart makesδv=0\\delta v=0by \([61](https://arxiv.org/html/2609.38274#A3.E61)\)\. More generally, \([62](https://arxiv.org/html/2609.38274#A3.E62)\) and \([63](https://arxiv.org/html/2609.38274#A3.E63)\) bound its norm through the standardized profile directions\. These terms connect target\-task accuracy to the pooled signals used by the selector, while \([72](https://arxiv.org/html/2609.38274#A3.E72)\) and \([68](https://arxiv.org/html/2609.38274#A3.E68)\) account separately for the objective weights and the curvature of the correctness association\.
### C\.8Member\-Specific Information in Regularized Stacking
The concatenated probability features used byStackingcontain both the team’s mean prediction and differences between its members\. We analyze the additional reduction in regularized cross\-entropy available from these member\-specific differences after the mean prediction is already available to a linear softmax classifier\.
#### Setting and comparison class\.
Fix a teamSSof sizek≥2k\\geq 2and relabel its members as1,…,k1,\\ldots,k\. Letc≥2c\\geq 2be the number of classes, and letpi\(X\)∈Δc−1p\_\{i\}\(X\)\\in\\Delta^\{c\-1\}be each member’s aligned choice distribution\. All expectations in this subsection are taken under one fixed distributionDDover\(X,Y\)\(X,Y\)\. This may be a population distribution or the empirical distribution of a development set\. Define the cross\-entropy loss on logits using natural logarithms:
ℓ\(z,y\)=log\(∑a=1cexp\(za\)\)−zy\.\\ell\(z,y\)=\\log\\\!\\left\(\\sum\_\{a=1\}^\{c\}\\exp\(z\_\{a\}\)\\right\)\-z\_\{y\}\.Forβ\>0\\beta\>0, the optimal value of the Stacking objective is
Fβ\(S\)=infW1,…,Wk,b\{𝔼Dℓ\(∑i=1kWi⊤pi\(X\)\+b,Y\)\+β2∑i=1k‖Wi‖F2\},F\_\{\\beta\}\(S\)=\\inf\_\{W\_\{1\},\\ldots,W\_\{k\},b\}\\left\\\{\\mathbb\{E\}\_\{D\}\\ell\\\!\\left\(\\sum\_\{i=1\}^\{k\}W\_\{i\}^\{\\top\}p\_\{i\}\(X\)\+b,Y\\right\)\+\\frac\{\\beta\}\{2\}\\sum\_\{i=1\}^\{k\}\\\|W\_\{i\}\\\|\_\{F\}^\{2\}\\right\\\},\(77\)whereWi∈ℝc×cW\_\{i\}\\in\\mathbb\{R\}^\{c\\times c\}are the blocks of the weight matrix in[Equations101](https://arxiv.org/html/2609.38274#A4.E101)and[102](https://arxiv.org/html/2609.38274#A4.E102), andb∈ℝcb\\in\\mathbb\{R\}^\{c\}is unpenalized\. For an empiricalDD, this is exactly the average cross\-entropy plus theL2L\_\{2\}penalty in[Equation102](https://arxiv.org/html/2609.38274#A4.E102)\. In the implementation, the data\-fit gradient is averaged over the development examples and the penalty contributes𝚕𝟸W\\mathtt\{l2\}\\,Wto the weight gradient, so the coefficient here isβ=𝚕𝟸\\beta=\\mathtt\{l2\}\.
Letp¯\(X\)=k−1∑i=1kpi\(X\)\\bar\{p\}\(X\)=k^\{\-1\}\\sum\_\{i=1\}^\{k\}p\_\{i\}\(X\)\. The comparison class uses only this mean probability vector as the input to a linear softmax classifier:
Gβ/k\(p¯\)=minM,b\{𝔼Dℓ\(M⊤p¯\(X\)\+b,Y\)\+β2k‖M‖F2\}\.G\_\{\\beta/k\}\(\\bar\{p\}\)=\\min\_\{M,b\}\\left\\\{\\mathbb\{E\}\_\{D\}\\ell\(M^\{\\top\}\\bar\{p\}\(X\)\+b,Y\)\+\\frac\{\\beta\}\{2k\}\\\|M\\\|\_\{F\}^\{2\}\\right\\\}\.\(78\)This is a restricted class within Stacking: settingWi=M/kW\_\{i\}=M/kreproduces its logits and penalty\. Thus the factorβ/k\\beta/kfollows from the original regularizer\. The comparison class is used only for analysis; its input equals theChoice\-Softmean, followed by a learned linear softmax map\. Assume that[Equation78](https://arxiv.org/html/2609.38274#A3.E78)has a finite minimizer\(M0,b0\)\(M\_\{0\},b\_\{0\}\)\. For example, positive probability for every class is sufficient whenβ\>0\\beta\>0, after fixing𝟏⊤b=0\\mathbf\{1\}^\{\\top\}b=0to remove the irrelevant common shift in the logits\.
Define the mean classifier, the concatenated member differences, and their residual alignment by
π0\(X\)\\displaystyle\\pi\_\{0\}\(X\)=softmax\(M0⊤p¯\(X\)\+b0\),\\displaystyle=\\operatorname\{softmax\}\(M\_\{0\}^\{\\top\}\\bar\{p\}\(X\)\+b\_\{0\}\),\(79\)δϕ\(X\)\\displaystyle\\delta\\phi\(X\)=\[p1\(X\)−p¯\(X\);…;pk\(X\)−p¯\(X\)\],\\displaystyle=\[p\_\{1\}\(X\)\-\\bar\{p\}\(X\);\\ldots;p\_\{k\}\(X\)\-\\bar\{p\}\(X\)\],Γ\\displaystyle\\Gamma=𝔼D\[δϕ\(X\)\(eY−π0\(X\)\)⊤\],\\displaystyle=\\mathbb\{E\}\_\{D\}\\\!\\left\[\\delta\\phi\(X\)\(e\_\{Y\}\-\\pi\_\{0\}\(X\)\)^\{\\top\}\\right\],whereeYe\_\{Y\}is the one\-hot vector of the true class\. Also set
Σδ=𝔼D\[δϕ\(X\)δϕ\(X\)⊤\],Lδ=β\+12λmax\(Σδ\)\.\\Sigma\_\{\\delta\}=\\mathbb\{E\}\_\{D\}\[\\delta\\phi\(X\)\\delta\\phi\(X\)^\{\\top\}\],\\qquad L\_\{\\delta\}=\\beta\+\\frac\{1\}\{2\}\\lambda\_\{\\max\}\(\\Sigma\_\{\\delta\}\)\.\(80\)All these moments are finite because the probability features are bounded\. To use the same distribution throughout the analysis, write
HD\(S\)=HIdist\(S,JD\),JijD=𝔼D\[JSD2\(pi\(X\),pj\(X\)\)\],H\_\{D\}\(S\)=\\mathrm\{HI\}\_\{dist\}\(S;J^\{D\}\),\\qquad J^\{D\}\_\{ij\}=\\mathbb\{E\}\_\{D\}\\\!\\left\[\\operatorname\{JSD\}\_\{2\}\(p\_\{i\}\(X\),p\_\{j\}\(X\)\)\\right\],whereJSD2\\operatorname\{JSD\}\_\{2\}uses logarithms to base22andHDH\_\{D\}follows the pairwise averaging rule in[Equation7](https://arxiv.org/html/2609.38274#S3.E7)\.
###### Theorem C\.9\(Additional information available to regularized Stacking\)\.
Under the setting above, letΔS=Gβ/k\(p¯\)−Fβ\(S\)\\Delta\_\{S\}=G\_\{\\beta/k\}\(\\bar\{p\}\)\-F\_\{\\beta\}\(S\)\. Then
‖Γ‖F22Lδ≤ΔS≤‖Γ‖F22β≤2ln2\(k−1\)βHD\(S\)\.\\frac\{\\\|\\Gamma\\\|\_\{F\}^\{2\}\}\{2L\_\{\\delta\}\}\\;\\leq\\;\\Delta\_\{S\}\\;\\leq\\;\\frac\{\\\|\\Gamma\\\|\_\{F\}^\{2\}\}\{2\\beta\}\\;\\leq\\;\\frac\{2\\ln 2\\,\(k\-1\)\}\{\\beta\}\\,H\_\{D\}\(S\)\.\(81\)Consequently,ΔS\>0\\Delta\_\{S\}\>0if and only ifΓ≠0\\Gamma\\neq 0\.
###### Proof\.
*Step 1: separate the mean and member\-specific weights\.*For arbitrary Stacking weights, define
M=∑i=1kWi,Ui=Wi−Mk,U=\[U1;…;Uk\]\.M=\\sum\_\{i=1\}^\{k\}W\_\{i\},\\qquad U\_\{i\}=W\_\{i\}\-\\frac\{M\}\{k\},\\qquad U=\[U\_\{1\};\\ldots;U\_\{k\}\]\.Then∑iUi=0\\sum\_\{i\}U\_\{i\}=0, and every collection of weights admits this decomposition\. Conversely, anyMMandUUsatisfying∑iUi=0\\sum\_\{i\}U\_\{i\}=0define weightsWi=M/k\+UiW\_\{i\}=M/k\+U\_\{i\}\. The logits and penalty obey
∑iWi⊤pi\+b\\displaystyle\\sum\_\{i\}W\_\{i\}^\{\\top\}p\_\{i\}\+b=M⊤p¯\+U⊤δϕ\+b,\\displaystyle=M^\{\\top\}\\bar\{p\}\+U^\{\\top\}\\delta\\phi\+b,\(82\)∑i‖Wi‖F2\\displaystyle\\sum\_\{i\}\\\|W\_\{i\}\\\|\_\{F\}^\{2\}=‖M‖F2k\+‖U‖F2\.\\displaystyle=\\frac\{\\\|M\\\|\_\{F\}^\{2\}\}\{k\}\+\\\|U\\\|\_\{F\}^\{2\}\.\(83\)The first identity uses∑iUi=0\\sum\_\{i\}U\_\{i\}=0; the cross terms in the second identity vanish for the same reason\. Therefore, with
Ψ\(M,b,U\)=𝔼Dℓ\(M⊤p¯\+U⊤δϕ\+b,Y\)\+β2k‖M‖F2\+β2‖U‖F2,\\Psi\(M,b,U\)=\\mathbb\{E\}\_\{D\}\\ell\(M^\{\\top\}\\bar\{p\}\+U^\{\\top\}\\delta\\phi\+b,Y\)\+\\frac\{\\beta\}\{2k\}\\\|M\\\|\_\{F\}^\{2\}\+\\frac\{\\beta\}\{2\}\\\|U\\\|\_\{F\}^\{2\},we haveFβ\(S\)=infM,b,U:∑iUi=0Ψ\(M,b,U\)F\_\{\\beta\}\(S\)=\\inf\_\{M,b,U:\\,\\sum\_\{i\}U\_\{i\}=0\}\\Psi\(M,b,U\)andGβ/k\(p¯\)=Ψ\(M0,b0,0\)G\_\{\\beta/k\}\(\\bar\{p\}\)=\\Psi\(M\_\{0\},b\_\{0\},0\)\. In particular,ΔS≥0\\Delta\_\{S\}\\geq 0\.
*Step 2: bound the improvement using the residual gradient\.*The finite optimum of the mean classifier satisfies
𝔼D\[p¯\(π0−eY\)⊤\]\+βkM0=0,𝔼D\[π0−eY\]=0\.\\mathbb\{E\}\_\{D\}\[\\bar\{p\}\(\\pi\_\{0\}\-e\_\{Y\}\)^\{\\top\}\]\+\\frac\{\\beta\}\{k\}M\_\{0\}=0,\\qquad\\mathbb\{E\}\_\{D\}\[\\pi\_\{0\}\-e\_\{Y\}\]=0\.\(84\)If𝟏⊤b=0\\mathbf\{1\}^\{\\top\}b=0is imposed, stationarity initially gives the second equation on that subspace\. The bias gradient always has coordinate sum zero, so this also makes its full gradient zero\. Differentiation under the expectation is justified by bounded probability features and bounded first derivatives of cross\-entropy with respect to logits\.
Apply convexity to the cross\-entropy term at\(M0,b0,0\)\(M\_\{0\},b\_\{0\},0\), and expand the quadratic penalties\. The first\-order terms inM−M0M\-M\_\{0\}andb−b0b\-b\_\{0\}cancel by[Equation84](https://arxiv.org/html/2609.38274#A3.E84), leaving
Ψ\(M,b,U\)≥Gβ/k\(p¯\)−⟨Γ,U⟩F\+β2k‖M−M0‖F2\+β2‖U‖F2\.\\Psi\(M,b,U\)\\geq G\_\{\\beta/k\}\(\\bar\{p\}\)\-\\langle\\Gamma,U\\rangle\_\{F\}\+\\frac\{\\beta\}\{2k\}\\\|M\-M\_\{0\}\\\|\_\{F\}^\{2\}\+\\frac\{\\beta\}\{2\}\\\|U\\\|\_\{F\}^\{2\}\.\(85\)Here⟨A,B⟩F=tr\(A⊤B\)\\langle A,B\\rangle\_\{F\}=\\operatorname\{tr\}\(A^\{\\top\}B\)\. Completing the square inUUgives
−⟨Γ,U⟩F\+β2‖U‖F2=β2‖U−Γβ‖F2−‖Γ‖F22β\.\-\\langle\\Gamma,U\\rangle\_\{F\}\+\\frac\{\\beta\}\{2\}\\\|U\\\|\_\{F\}^\{2\}=\\frac\{\\beta\}\{2\}\\left\\\|U\-\\frac\{\\Gamma\}\{\\beta\}\\right\\\|\_\{F\}^\{2\}\-\\frac\{\\\|\\Gamma\\\|\_\{F\}^\{2\}\}\{2\\beta\}\.Taking the infimum in[Equation85](https://arxiv.org/html/2609.38274#A3.E85)yieldsFβ\(S\)≥Gβ/k\(p¯\)−‖Γ‖F2/\(2β\)F\_\{\\beta\}\(S\)\\geq G\_\{\\beta/k\}\(\\bar\{p\}\)\-\\\|\\Gamma\\\|\_\{F\}^\{2\}/\(2\\beta\), proving the upper bound onΔS\\Delta\_\{S\}\. This argument uses the quadratic penalty on the weights and the bias stationarity condition; it requires no strong convexity in the unpenalized bias\.
*Step 3: exhibit a feasible improvement\.*For a probability vectorπ\\pi, the Hessian of cross\-entropy with respect to logits isB\(π\)=diag\(π\)−ππ⊤B\(\\pi\)=\\operatorname\{diag\}\(\\pi\)\-\\pi\\pi^\{\\top\}\. For everyv∈ℝcv\\in\\mathbb\{R\}^\{c\},
v⊤B\(π\)v\\displaystyle v^\{\\top\}B\(\\pi\)v=∑a<bπaπb\(va−vb\)2\\displaystyle=\\sum\_\{a<b\}\\pi\_\{a\}\\pi\_\{b\}\(v\_\{a\}\-v\_\{b\}\)^\{2\}≤2∑aπa\(1−πa\)va2≤12‖v‖22\.\\displaystyle\\leq 2\\sum\_\{a\}\\pi\_\{a\}\(1\-\\pi\_\{a\}\)v\_\{a\}^\{2\}\\leq\\frac\{1\}\{2\}\\\|v\\\|\_\{2\}^\{2\}\.The last inequality usesπa\(1−πa\)≤1/4\\pi\_\{a\}\(1\-\\pi\_\{a\}\)\\leq 1/4, so0⪯B\(π\)⪯I/20\\preceq B\(\\pi\)\\preceq I/2\. Fix\(M,b\)=\(M0,b0\)\(M,b\)=\(M\_\{0\},b\_\{0\}\)and writef\(U\)=Ψ\(M0,b0,U\)f\(U\)=\\Psi\(M\_\{0\},b\_\{0\},U\)\. For any matrix directionVV, its second directional derivative obeys
D2f\(U\)\[V,V\]\\displaystyle D^\{2\}f\(U\)\[V,V\]≤β‖V‖F2\+12𝔼D‖V⊤δϕ‖22\\displaystyle\\leq\\beta\\\|V\\\|\_\{F\}^\{2\}\+\\frac\{1\}\{2\}\\mathbb\{E\}\_\{D\}\\\|V^\{\\top\}\\delta\\phi\\\|\_\{2\}^\{2\}=β‖V‖F2\+12tr\(V⊤ΣδV\)≤Lδ‖V‖F2\.\\displaystyle=\\beta\\\|V\\\|\_\{F\}^\{2\}\+\\frac\{1\}\{2\}\\operatorname\{tr\}\(V^\{\\top\}\\Sigma\_\{\\delta\}V\)\\leq L\_\{\\delta\}\\\|V\\\|\_\{F\}^\{2\}\.Since∇f\(0\)=−Γ\\nabla f\(0\)=\-\\Gamma, this global curvature bound implies
f\(U\)≤Gβ/k\(p¯\)−⟨Γ,U⟩F\+Lδ2‖U‖F2\.f\(U\)\\leq G\_\{\\beta/k\}\(\\bar\{p\}\)\-\\langle\\Gamma,U\\rangle\_\{F\}\+\\frac\{L\_\{\\delta\}\}\{2\}\\\|U\\\|\_\{F\}^\{2\}\.PartitionΓ\\Gammainto blocksΓi∈ℝc×c\\Gamma\_\{i\}\\in\\mathbb\{R\}^\{c\\times c\}\. Their sum is zero:
∑iΓi=𝔼D\[\(∑i\(pi−p¯\)\)\(eY−π0\)⊤\]=0\.\\sum\_\{i\}\\Gamma\_\{i\}=\\mathbb\{E\}\_\{D\}\\\!\\left\[\\left\(\\sum\_\{i\}\(p\_\{i\}\-\\bar\{p\}\)\\right\)\(e\_\{Y\}\-\\pi\_\{0\}\)^\{\\top\}\\right\]=0\.ThusU=Γ/LδU=\\Gamma/L\_\{\\delta\}satisfies the constraint∑iUi=0\\sum\_\{i\}U\_\{i\}=0and gives
Fβ\(S\)≤f\(Γ/Lδ\)≤Gβ/k\(p¯\)−‖Γ‖F22Lδ\.F\_\{\\beta\}\(S\)\\leq f\(\\Gamma/L\_\{\\delta\}\)\\leq G\_\{\\beta/k\}\(\\bar\{p\}\)\-\\frac\{\\\|\\Gamma\\\|\_\{F\}^\{2\}\}\{2L\_\{\\delta\}\}\.This proves the lower bound onΔS\\Delta\_\{S\}\.
*Step 4: relate the residual gradient to pairwise JSD\.*First,‖eY−π0‖22≤2\\\|e\_\{Y\}\-\\pi\_\{0\}\\\|\_\{2\}^\{2\}\\leq 2\. The triangle inequality and Cauchy–Schwarz therefore give
‖Γ‖F2≤𝔼D‖δϕ‖22𝔼D‖eY−π0‖22≤2𝔼D‖δϕ‖22\.\\\|\\Gamma\\\|\_\{F\}^\{2\}\\leq\\mathbb\{E\}\_\{D\}\\\|\\delta\\phi\\\|\_\{2\}^\{2\}\\,\\mathbb\{E\}\_\{D\}\\\|e\_\{Y\}\-\\pi\_\{0\}\\\|\_\{2\}^\{2\}\\leq 2\\mathbb\{E\}\_\{D\}\\\|\\delta\\phi\\\|\_\{2\}^\{2\}\.\(86\)For completeness, the constants connecting this energy to JSD follow directly from Pinsker’s inequality\. For probability vectorsp,qp,q, putm=\(p\+q\)/2m=\(p\+q\)/2and use natural logarithms forKL\\operatorname\{KL\}andJSDnat\\operatorname\{JSD\}\_\{\\mathrm\{nat\}\}\. Then
JSDnat\(p,q\)\\displaystyle\\operatorname\{JSD\}\_\{\\mathrm\{nat\}\}\(p,q\)=12KL\(p∥m\)\+12KL\(q∥m\)\\displaystyle=\\frac\{1\}\{2\}\\operatorname\{KL\}\(p\\\|m\)\+\\frac\{1\}\{2\}\\operatorname\{KL\}\(q\\\|m\)≥14‖p−m‖12\+14‖q−m‖12=18‖p−q‖12\.\\displaystyle\\geq\\frac\{1\}\{4\}\\\|p\-m\\\|\_\{1\}^\{2\}\+\\frac\{1\}\{4\}\\\|q\-m\\\|\_\{1\}^\{2\}=\\frac\{1\}\{8\}\\\|p\-q\\\|\_\{1\}^\{2\}\.Becausep−qp\-qhas coordinate sum zero, its positive and negative parts have equal total massaa\. Hence‖p−q‖22≤2a2=‖p−q‖12/2\\\|p\-q\\\|\_\{2\}^\{2\}\\leq 2a^\{2\}=\\\|p\-q\\\|\_\{1\}^\{2\}/2\. Converting to bits yields
JSD2\(p,q\)=JSDnat\(p,q\)ln2≥‖p−q‖224ln2\.\\operatorname\{JSD\}\_\{2\}\(p,q\)=\\frac\{\\operatorname\{JSD\}\_\{\\mathrm\{nat\}\}\(p,q\)\}\{\\ln 2\}\\geq\\frac\{\\\|p\-q\\\|\_\{2\}^\{2\}\}\{4\\ln 2\}\.\(87\)Finally, expanding the squared norms around the mean gives
𝔼D‖δϕ‖22\\displaystyle\\mathbb\{E\}\_\{D\}\\\|\\delta\\phi\\\|\_\{2\}^\{2\}=𝔼D∑i‖pi−p¯‖22=1k∑i<j𝔼D‖pi−pj‖22\\displaystyle=\\mathbb\{E\}\_\{D\}\\sum\_\{i\}\\\|p\_\{i\}\-\\bar\{p\}\\\|\_\{2\}^\{2\}=\\frac\{1\}\{k\}\\sum\_\{i<j\}\\mathbb\{E\}\_\{D\}\\\|p\_\{i\}\-p\_\{j\}\\\|\_\{2\}^\{2\}≤4ln2k∑i<jJijD=2ln2\(k−1\)HD\(S\)\.\\displaystyle\\leq\\frac\{4\\ln 2\}\{k\}\\sum\_\{i<j\}J^\{D\}\_\{ij\}=2\\ln 2\\,\(k\-1\)H\_\{D\}\(S\)\.Together with[Equation86](https://arxiv.org/html/2609.38274#A3.E86), this gives‖Γ‖F2≤4ln2\(k−1\)HD\(S\)\\\|\\Gamma\\\|\_\{F\}^\{2\}\\leq 4\\ln 2\\,\(k\-1\)H\_\{D\}\(S\)and proves the final inequality in[Equation81](https://arxiv.org/html/2609.38274#A3.E81)\. IfΓ=0\\Gamma=0, the upper bound andΔS≥0\\Delta\_\{S\}\\geq 0giveΔS=0\\Delta\_\{S\}=0\. IfΓ≠0\\Gamma\\neq 0, the lower bound is strictly positive because0<β≤Lδ<∞0<\\beta\\leq L\_\{\\delta\}<\\infty\. ∎
#### Interpretation\.
Pairwise JSD quantifies the size of the probability differences available to Stacking\. The matrixΓ\\Gammaidentifies whether those differences align with the predictive residual left by the mean classifier\. The theorem concerns the additional value of member\-specific features beyond that classifier, measured by the optimal regularized cross\-entropy underDD\. Improvements already present in the mean probability vector are part of the comparison class\. The theorem complements the soft\-aggregation analysis by characterizing the gain from retaining member identity in a learned linear combiner\.
## Appendix DImplementation Details
### D\.1Choice\-Space Aligned Scoring Implementation
For each examplexxwith label space𝒴t=\{y1,…,y\|𝒴t\|\}\\mathcal\{Y\}\_\{t\}=\\\{y\_\{1\},\\ldots,y\_\{\|\\mathcal\{Y\}\_\{t\}\|\}\\\}, we:
1. 1\.Force the model to generate each choice tokenyky\_\{k\}\(e\.g\., “A”, “B”, “C”, “D”\)
2. 2\.Extract the log\-probabilityℓi\(yk\|x\)\\ell\_\{i\}\(y\_\{k\}\|x\)from the model’s output distribution
3. 3\.Normalize across𝒴t\\mathcal\{Y\}\_\{t\}: pi\(yk\|x\)=exp\(ℓi\(yk\|x\)\)∑y′∈𝒴texp\(ℓi\(y′\|x\)\)p\_\{i\}\(y\_\{k\}\|x\)=\\frac\{\\exp\(\\ell\_\{i\}\(y\_\{k\}\|x\)\)\}\{\\sum\_\{y^\{\\prime\}\\in\\mathcal\{Y\}\_\{t\}\}\\exp\(\\ell\_\{i\}\(y^\{\\prime\}\|x\)\)\}\(88\)
This ensures∑y∈𝒴tpi\(y\|x\)=1\\sum\_\{y\\in\\mathcal\{Y\}\_\{t\}\}p\_\{i\}\(y\|x\)=1with negligible OTHER mass\.
#### Prompt Format\.
We use a chat\-based prompt format with two messages:
Profiling Prompt TemplateUser Message:[⬇](data:text/plain;base64,W1F1ZXN0aW9uXQoKT3B0aW9uczoKW09wdGlvbiAxXQpbT3B0aW9uIDJdCi4uLgoKQ29tcGxldGUgdGhlIHNlbnRlbmNlICdUaGUgY29ycmVjdCBhbnN3ZXIgaXMgJyBieSBvdXRwdXR0aW5nCmV4YWN0bHkgb25lIGxldHRlciBmcm9tOiBBLCBCLCBDLCBELiBPdXRwdXQgbm90aGluZyBlbHNlLg==)\[Question\]Options:\[Option1\]\[Option2\]\.\.\.Completethesentence’Thecorrectansweris’byoutputtingexactlyoneletterfrom:A,B,C,D\.Outputnothingelse\.Assistant Message \(prefix\):\(Note: there is a trailingspaceafter “is”\.\)[⬇](data:text/plain;base64,VGhlIGNvcnJlY3QgYW5zd2VyIGlz)Thecorrectansweris
#### Why normalize within the choice space\.
Outside\-choice token mass can introduce divergence unrelated to preferences among the actual answer options\. The following entropy identity\([Lin, 1991](https://arxiv.org/html/2609.38274#bib.bib47)\)isolates this effect\.
###### Lemma D\.1\(Divergence due to outside\-choice mass\)\.
Let𝒴\\mathcal\{Y\}be the evaluation label space \(e\.g\.,\{A,B,C,D\}\\\{A,B,C,D\\\}\) and𝒴′=𝒴∪\{OTHER\}\\mathcal\{Y\}^\{\\prime\}=\\mathcal\{Y\}\\cup\\\{\\text\{OTHER\}\\\}be an extended space\. Suppose two agents have distributions:
pi\\displaystyle p\_\{i\}=\(1−αi\)q\+αiδOTHER\\displaystyle=\(1\-\\alpha\_\{i\}\)q\+\\alpha\_\{i\}\\delta\_\{\\text\{OTHER\}\}\(89\)pj\\displaystyle p\_\{j\}=\(1−αj\)q\+αjδOTHER\\displaystyle=\(1\-\\alpha\_\{j\}\)q\+\\alpha\_\{j\}\\delta\_\{\\text\{OTHER\}\}\(90\)whereqqis a distribution over𝒴\\mathcal\{Y\}\(identical for both agents\),δOTHER\\delta\_\{\\text\{OTHER\}\}is a point mass on OTHER, andαi,αj∈\[0,1\]\\alpha\_\{i\},\\alpha\_\{j\}\\in\[0,1\]\.
Then:
JSD\(pi,pj\)=JSD\(\[αi,1−αi\],\[αj,1−αj\]\)\\mathrm\{JSD\}\(p\_\{i\},p\_\{j\}\)=\\mathrm\{JSD\}\\big\(\[\\alpha\_\{i\},1\-\\alpha\_\{i\}\],\[\\alpha\_\{j\},1\-\\alpha\_\{j\}\]\\big\)\(91\)
Thus all divergence arises from the difference in OTHER mass, although their common choice distribution isqq\(the conditional distribution wheneverαi,αj<1\\alpha\_\{i\},\\alpha\_\{j\}<1\)\.
###### Proof\.
Letα¯=\(αi\+αj\)/2\\bar\{\\alpha\}=\(\\alpha\_\{i\}\+\\alpha\_\{j\}\)/2, and letHHdenote Shannon entropy in bits\. SinceqqandδOTHER\\delta\_\{\\text\{OTHER\}\}have disjoint supports,
H\(pi\)\\displaystyle H\(p\_\{i\}\)=H\(αi\)\+\(1−αi\)H\(q\),\\displaystyle=H\(\\alpha\_\{i\}\)\+\(1\-\\alpha\_\{i\}\)H\(q\),\(92\)H\(pj\)\\displaystyle H\(p\_\{j\}\)=H\(αj\)\+\(1−αj\)H\(q\),\\displaystyle=H\(\\alpha\_\{j\}\)\+\(1\-\\alpha\_\{j\}\)H\(q\),\(93\)H\(pi\+pj2\)\\displaystyle H\\\!\\left\(\\frac\{p\_\{i\}\+p\_\{j\}\}\{2\}\\right\)=H\(α¯\)\+\(1−α¯\)H\(q\)\.\\displaystyle=H\(\\bar\{\\alpha\}\)\+\(1\-\\bar\{\\alpha\}\)H\(q\)\.\(94\)Substituting these identities into the entropy expression for JSD cancels theH\(q\)H\(q\)terms:
JSD\(pi,pj\)=H\(αi\+αj2\)−H\(αi\)\+H\(αj\)2=JSD\(\[αi,1−αi\],\[αj,1−αj\]\)\.\\mathrm\{JSD\}\(p\_\{i\},p\_\{j\}\)=H\\left\(\\frac\{\\alpha\_\{i\}\+\\alpha\_\{j\}\}\{2\}\\right\)\-\\frac\{H\(\\alpha\_\{i\}\)\+H\(\\alpha\_\{j\}\)\}\{2\}=\\mathrm\{JSD\}\\big\(\[\\alpha\_\{i\},1\-\\alpha\_\{i\}\],\[\\alpha\_\{j\},1\-\\alpha\_\{j\}\]\\big\)\.\(95\)whereH\(p\)=−plog2p−\(1−p\)log2\(1−p\)H\(p\)=\-p\\log\_\{2\}p\-\(1\-p\)\\log\_\{2\}\(1\-p\)is the binary entropy\.
Forαi=0\.1\\alpha\_\{i\}=0\.1andαj=0\.9\\alpha\_\{j\}=0\.9, the divergence is1−H\(0\.1\)≈0\.5311\-H\(0\.1\)\\approx 0\.531bits, entirely due to the difference in OTHER mass\. ∎
Choice\-space normalization removes this outside\-mass contribution before computingJtJ^\{t\}\.
### D\.2Paired Bootstrap Procedure
For comparing two methodsAAandBBon a test set ofNNexamples:
1. 1\.Compute per\-example correctness:cA\(xn\),cB\(xn\)∈\{0,1\}c\_\{A\}\(x\_\{n\}\),c\_\{B\}\(x\_\{n\}\)\\in\\\{0,1\\\}forn=1,…,Nn=1,\\ldots,N
2. 2\.Forb=1,…,Bb=1,\\ldots,B\(we useB=2000B=2000\): 1. \(a\)Sample indices\{i1,…,iN\}\\\{i\_\{1\},\\ldots,i\_\{N\}\\\}uniformly with replacement from\{1,…,N\}\\\{1,\\ldots,N\\\} 2. \(b\)Compute bootstrap accuracies: AccA\(b\)\\displaystyle\\text\{Acc\}\_\{A\}^\{\(b\)\}=1N∑n=1NcA\(xin\)\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}c\_\{A\}\(x\_\{i\_\{n\}\}\)\(96\)AccB\(b\)\\displaystyle\\text\{Acc\}\_\{B\}^\{\(b\)\}=1N∑n=1NcB\(xin\)\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}c\_\{B\}\(x\_\{i\_\{n\}\}\)\(97\) 3. \(c\)Compute bootstrap difference:Δ\(b\)=AccA\(b\)−AccB\(b\)\\Delta^\{\(b\)\}=\\text\{Acc\}\_\{A\}^\{\(b\)\}\-\\text\{Acc\}\_\{B\}^\{\(b\)\}
3. 3\.Compute 95% confidence interval:\[percentile\(\{Δ\(b\)\},2\.5\),percentile\(\{Δ\(b\)\},97\.5\)\]\[\\text\{percentile\}\(\\\{\\Delta^\{\(b\)\}\\\},2\.5\),\\text\{percentile\}\(\\\{\\Delta^\{\(b\)\}\\\},97\.5\)\]
4. 4\.Mark as significant if the CI excludes 0
The paired structure \(using the same bootstrap sample for both methods\) preserves the correlation betweencAc\_\{A\}andcBc\_\{B\}, providing more statistical power than unpaired bootstrap\.
### D\.3Hyperparameter Selection
We maintain strict development–test separation when selecting the objective weights\. We first perform a grid search using only development splits to identify a stable high\-performing region for the two heterogeneity weights, and then fix\(λ1,λ2\)=\(0\.13,0\.05\)\(\\lambda\_\{1\},\\lambda\_\{2\}\)=\(0\.13,0\.05\)for all datasets and aggregation rules\. No test\-set performance is used for per\-task hyperparameter tuning\.
To examine whether this choice reflects a task\-specific artifact, we further conduct leave\-one\-task\-out hyperparameter selection over the seven primary benchmarks\. For each held\-out benchmark, the weight pair is selected using the development splits of the remaining six benchmarks and then applied to the held\-out benchmark\. Across the held\-out tasks, the grid contains broad high\-performing plateaus rather than isolated optima: multiple neighboring weight pairs achieve the same development\-set performance\. The fixed weight pair used in our main experiments lies within this high\-performing region for 6 out of 7 held\-out tasks\. This indicates that the selected weights capture a stable quality–heterogeneity trade\-off rather than a task\-specific tuning artifact\.
### D\.4Computational Cost of Team Selection
Withmmcandidates and target sizekk, Algorithm[3\.3](https://arxiv.org/html/2609.38274#S3.SS3)triesmmseeds and at mostk−1k\-1additions per seed\. If within\-team quality and pairwise sums are cached, evaluating an extension from a size\-ssteam needssspairwise lookups plus constant\-time quality and normalization operations\. Evaluating all remaining candidates therefore costsO\(ms\)O\(ms\)per round andO\(m∑s=1k−1s\)=O\(mk2\)O\(m\\sum\_\{s=1\}^\{k\-1\}s\)=O\(mk^\{2\}\)per seed\. The total cost isO\(m2k2\)O\(m^\{2\}k^\{2\}\), withO\(m2\)O\(m^\{2\}\)storage for the pairwise matrices\. These are operation counts for cached score evaluation\. At the experimental sizem=11m=11,k=3k=3, exhaustive enumeration is also feasible; multi\-start search extends the same scoring rule to larger pools\.
### D\.5Aggregation Rules
#### Choice\-Soft\.
Choice\-Softaggregates by averaging the members’ aligned choice distributions\. For a teamStS\_\{t\}and instancexx, we compute the mean distribution as follows:
pcst\(y∣x\)=1\|St\|∑i∈Stpit\(y∣x\)\.p\_\{\\textsc\{cs\}\}^\{t\}\(y\\mid x\)=\\frac\{1\}\{\|S\_\{t\}\|\}\\sum\_\{i\\in S\_\{t\}\}p\_\{i\}^\{t\}\(y\\mid x\)\.\(98\)We then predicty^St\(x\)=argmaxy∈𝒴tpcst\(y∣x\)\\hat\{y\}\_\{S\_\{t\}\}\(x\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\_\{t\}\}p\_\{\\textsc\{cs\}\}^\{t\}\(y\\mid x\)\.
#### PoE\.
PoE\(Product of Experts\) combines distributions multiplicatively, rewarding labels that receive consistently high probability across members\. We form the aggregated distribution as:
ppoet\(y∣x\)=∏i∈Stpit\(y∣x\)∑y′∈𝒴t∏i∈Stpit\(y′∣x\),p\_\{\\textsc\{poe\}\}^\{t\}\(y\\mid x\)=\\frac\{\\prod\_\{i\\in S\_\{t\}\}p\_\{i\}^\{t\}\(y\\mid x\)\}\{\\sum\_\{y^\{\\prime\}\\in\\mathcal\{Y\}\_\{t\}\}\\prod\_\{i\\in S\_\{t\}\}p\_\{i\}^\{t\}\(y^\{\\prime\}\\mid x\)\},\(99\)and predicty^St\(x\)=argmaxy∈𝒴tppoet\(y∣x\)\\hat\{y\}\_\{S\_\{t\}\}\(x\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\_\{t\}\}p\_\{\\textsc\{poe\}\}^\{t\}\(y\\mid x\)\.
#### DS\.
DS\(Dawid–Skene\) models each member as a noisy labeler with an agent\-specific confusion matrix estimated onDtdevD\_\{t\}^\{\\mathrm\{dev\}\}\. Lety^i\(x\)=argmaxy∈𝒴tpit\(y∣x\)\\hat\{y\}\_\{i\}\(x\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\_\{t\}\}p\_\{i\}^\{t\}\(y\\mid x\)denote the label predicted by modelii\. We estimate a priorπt\(y\)\\pi^\{t\}\(y\)and, for each modelii, a confusion matrixAit∈ℝ\|𝒴t\|×\|𝒴t\|A\_\{i\}^\{t\}\\in\\mathbb\{R\}^\{\|\\mathcal\{Y\}\_\{t\}\|\\times\|\\mathcal\{Y\}\_\{t\}\|\}via additive smoothing \(with a small constantϵ\\epsilon\), whereAit\(a,b\)A\_\{i\}^\{t\}\(a,b\)estimatesℙ\(y^i=b∣y=a\)\\mathbb\{P\}\(\\hat\{y\}\_\{i\}=b\\mid y=a\)\. At inference time, we compute the posterior up to proportionality:
pdst\(y∣x\)∝πt\(y\)∏i∈StAit\(y,y^i\(x\)\)\.p\_\{\\textsc\{ds\}\}^\{t\}\(y\\mid x\)\\;\\propto\\;\\pi^\{t\}\(y\)\\prod\_\{i\\in S\_\{t\}\}A\_\{i\}^\{t\}\\\!\\left\(y,\\,\\hat\{y\}\_\{i\}\(x\)\\right\)\.\(100\)We then renormalize over𝒴t\\mathcal\{Y\}\_\{t\}and predicty^St\(x\)=argmaxy∈𝒴tpdst\(y∣x\)\\hat\{y\}\_\{S\_\{t\}\}\(x\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\_\{t\}\}p\_\{\\textsc\{ds\}\}^\{t\}\(y\\mid x\)\.
#### Stacking\.
Stackingtrains a lightweight combiner onDtdevD\_\{t\}^\{\\mathrm\{dev\}\}\. For a teamStS\_\{t\}and instancexx, we concatenate members’ choice distributions:
ϕ\(x\)=\[pi1t\(⋅∣x\);…;pikt\(⋅∣x\)\]∈ℝk\|𝒴t\|\.\\phi\(x\)=\\big\[p\_\{i\_\{1\}\}^\{t\}\(\\cdot\\mid x\);\\ldots;p\_\{i\_\{k\}\}^\{t\}\(\\cdot\\mid x\)\\big\]\\in\\mathbb\{R\}^\{k\|\\mathcal\{Y\}\_\{t\}\|\}\.\(101\)We learn a linear classifier parametrized byW∈ℝk\|𝒴t\|×\|𝒴t\|W\\in\\mathbb\{R\}^\{k\|\\mathcal\{Y\}\_\{t\}\|\\times\|\\mathcal\{Y\}\_\{t\}\|\}andb∈ℝ\|𝒴t\|b\\in\\mathbb\{R\}^\{\|\\mathcal\{Y\}\_\{t\}\|\}viaL2L\_\{2\}\-regularized cross\-entropy\. LetπW,b\(y∣x\)=softmax\(W⊤ϕ\(x\)\+b\)y\\pi\_\{W,b\}\(y\\mid x\)=\\mathrm\{softmax\}\(W^\{\\top\}\\phi\(x\)\+b\)\_\{y\}\. We train the stacking combiner by minimizing:
minW,bℒ\(W,b\)=\\displaystyle\\min\_\{W,b\}\\;\\mathcal\{L\}\(W,b\)=\{\}1\|Dtdev\|∑\(x,y\)∈Dtdev−logπW,b\(y∣x\)\\displaystyle\\frac\{1\}\{\|D\_\{t\}^\{\\mathrm\{dev\}\}\|\}\\sum\_\{\(x,y\)\\in D\_\{t\}^\{\\mathrm\{dev\}\}\}\-\\log\\pi\_\{W,b\}\(y\\mid x\)\(102\)\+β2‖W‖F2\.\\displaystyle\+\\frac\{\\beta\}\{2\}\\\|W\\\|\_\{F\}^\{2\}\.At test time, we predicty^St\(x\)=argmaxy∈𝒴tπW,b\(y∣x\)\\hat\{y\}\_\{S\_\{t\}\}\(x\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\_\{t\}\}\\pi\_\{W,b\}\(y\\mid x\)\.
## Appendix ESupplementary Results
This section provides additional results that complement the main experiments\. We first report the complete benchmark tables omitted from the main text for space, then analyze the team compositions selected by C2\-MAS, and finally provide robustness and ablation analyses for team size and objective validity\.
### E\.1Complete Benchmark Results
[SectionE\.1](https://arxiv.org/html/2609.38274#A5.SS1)reports the full per\-benchmark accuracy across all 11 open\-weight candidate models, the Self\-Consistency baseline, and C2\-MAS under four aggregation rules, separated into \(a\) 7 primary benchmarks and \(b\) 6 OOD benchmarks\.
Table[E\.1](https://arxiv.org/html/2609.38274#A5.SS1): Complete per\-benchmark results across all 13 benchmarks\. We report the full open\-weight candidate pool \(m=11m=11\), the single\-model Self\-Consistency baseline, and C2\-MAS under four aggregation rules\. \(a\) covers the seven primary benchmarks used in[Table1](https://arxiv.org/html/2609.38274#S4.T1); \(b\) covers the six OOD benchmarks, where heterogeneity matrices are pooled only on the primary set and applied to OOD teams \(dev→\\rightarrowtest\)\.Avg\.reports the average accuracy within each panel\. Abbreviations: ARC\-C: ARC\-Challenge; CSQA: CommonsenseQA; OBQA: OpenBookQA\.\(a\) Primary benchmarks \(7\)\.MethodDatasetAvg\.ARC\-CCSQALogiQA2MedQAMMLUMMLU\-ProOBQAOpen\-Weight Individual ModelsLlama\-3\.1\-8B80\.1580\.1575\.2875\.2854\.0354\.0361\.9861\.9867\.7067\.7037\.9237\.9281\.7481\.7465\.5465\.54Qwen3\-8B89\.6189\.6179\.7879\.7867\.2367\.2361\.0561\.0571\.7271\.7245\.7945\.7983\.2483\.2471\.2071\.20Gemma\-2\-9B88\.4888\.4878\.5678\.5663\.2063\.2059\.7459\.7474\.4474\.4444\.0144\.0185\.4985\.4970\.5670\.56Ministral\-3\-8B79\.4079\.4062\.0862\.0859\.9359\.9355\.5255\.5267\.8867\.8841\.2041\.2068\.4568\.4562\.0762\.07Granite\-3\.3\-8B80\.9080\.9072\.6672\.6653\.0053\.0051\.2251\.2265\.4565\.4535\.0235\.0280\.5280\.5262\.6862\.68GLM\-4\-9B85\.8685\.8674\.4474\.4456\.5556\.5555\.8155\.8164\.6164\.6131\.8431\.8484\.9384\.9364\.8664\.86Nemotron\-Nano\-9B88\.9588\.9574\.7274\.7265\.0765\.0760\.2160\.2172\.9472\.9445\.8845\.8888\.9588\.9570\.9670\.96EuroLLM\-9B75\.9475\.9472\.8572\.8545\.8845\.8850\.0050\.0059\.9359\.9326\.9726\.9773\.8873\.8857\.9257\.92OLMo\-3\-7B72\.4772\.4771\.5471\.5442\.8842\.8841\.4841\.4856\.1856\.1832\.8732\.8766\.5766\.5754\.8654\.86Apertus\-8B72\.0072\.0065\.5465\.5446\.8246\.8248\.2248\.2255\.7155\.7129\.6829\.6868\.6368\.6355\.2355\.23InternLM3\-8B88\.2088\.2075\.9475\.9467\.1367\.1359\.6459\.6471\.1671\.1642\.7042\.7084\.3684\.3669\.8869\.88Single\-Model AggregationSelf\-Consistency89\.1489\.1478\.6578\.6567\.3267\.3261\.7061\.7073\.4173\.4144\.7644\.7688\.5888\.5871\.9471\.94C2\-MAS Team SelectionC2\-MAS \(Choice\-Soft\)91\.4891\.4883\.9083\.9070\.8870\.8866\.3966\.3976\.5976\.5948\.5048\.5090\.4590\.4575\.4575\.45C2\-MAS \(PoE\)91\.5791\.5783\.9083\.9070\.6970\.6965\.3665\.3676\.6976\.6948\.9748\.9789\.5189\.5175\.2475\.24C2\-MAS \(DS\)90\.9290\.9282\.8782\.8770\.3270\.3265\.1765\.1775\.4775\.4747\.1047\.1090\.4590\.4574\.6174\.61C2\-MAS \(Stacking\)91\.6791\.6784\.1884\.1870\.6070\.6065\.5465\.5476\.1276\.1248\.9748\.9791\.5791\.5775\.5275\.52Table[E\.1](https://arxiv.org/html/2609.38274#A5.SS1)\(continued\)\.\(b\) OOD benchmarks \(6\)\.MethodDatasetAvg\.AQUA\-RATC\-EvalMathQARACEReClorStrategyQAOpen\-Weight Individual ModelsLlama\-3\.1\-8B33\.9933\.9949\.5349\.5333\.2433\.2460\.7760\.7762\.4562\.4567\.3267\.3251\.2251\.22Qwen3\-8B42\.6042\.6075\.0975\.0940\.0740\.0763\.1163\.1181\.4681\.4667\.8867\.8861\.7061\.70Gemma\-2\-9B30\.9030\.9055\.3455\.3432\.2132\.2164\.5164\.5177\.0677\.0670\.1370\.1355\.0355\.03Ministral\-3\-8B31\.3731\.3758\.8058\.8027\.1527\.1558\.8058\.8067\.7967\.7958\.5258\.5250\.4150\.41Granite\-3\.3\-8B32\.4032\.4045\.5145\.5131\.0931\.0960\.9660\.9666\.4866\.4864\.8964\.8950\.2250\.22GLM\-4\-9B30\.8130\.8165\.7365\.7330\.4330\.4360\.3960\.3961\.8961\.8967\.4267\.4252\.7852\.78Nemotron\-Nano\-9B45\.3245\.3254\.2154\.2140\.7340\.7364\.1464\.1482\.1282\.1263\.2063\.2058\.2958\.29EuroLLM\-9B27\.8127\.8146\.5446\.5424\.9124\.9156\.9356\.9356\.4656\.4664\.3364\.3346\.1646\.16OLMo\-3\-7B33\.6133\.6137\.9237\.9230\.8130\.8158\.7158\.7155\.2455\.2463\.6763\.6746\.6646\.66Apertus\-8B27\.0627\.0647\.1947\.1928\.4628\.4653\.7553\.7550\.0950\.0961\.7061\.7044\.7144\.71InternLM3\-8B32\.1232\.1284\.0884\.0829\.3129\.3163\.7663\.7672\.4772\.4762\.5562\.5557\.3857\.38Single\-Model AggregationSelf\-Consistency45\.5145\.5181\.4681\.4639\.2339\.2367\.6067\.6081\.6581\.6571\.2571\.2564\.4564\.45C2\-MAS Team SelectionC2\-MAS \(Choice\-Soft\)45\.6945\.6980\.2480\.2441\.9541\.9567\.9867\.9884\.1884\.1869\.7669\.7664\.9764\.97C2\-MAS \(PoE\)45\.2245\.2278\.9378\.9341\.5741\.5767\.7967\.7984\.5584\.5569\.6669\.6664\.6264\.62C2\-MAS \(DS\)44\.1944\.1980\.5280\.5242\.0442\.0468\.3568\.3583\.9083\.9070\.7970\.7964\.9764\.97C2\-MAS \(Stacking\)46\.7246\.7285\.9685\.9643\.9143\.9168\.3568\.3584\.8384\.8370\.6970\.6966\.7466\.74Table 5:OOD test accuracy \(%\) on the remaining 6 benchmarks\. For C2\-MAS, we pool heterogeneity matrices using only the seven primary benchmarks in[Table1](https://arxiv.org/html/2609.38274#S4.T1)and apply them to select teams on these 6 benchmarks \(dev→\\rightarrowtest\)\.↑\\uparrowindicates gain over Random\-kk\(pp\)\.Boldmarks the best within each aggregation block\.MethodDatasetAvg\.AQUA\-RATC\-EvalMathQARACEReClorStrategyQAClosed\-Source Models \(for reference\)Gemini\-2\.5\-Flash26\.1226\.1273\.4173\.4130\.1530\.1563\.8663\.8679\.2179\.2162\.7362\.7355\.9155\.91GPT\-4o29\.9629\.9672\.4772\.4729\.4929\.4965\.0765\.0784\.3684\.3680\.0680\.0660\.2460\.24Single\-Model AggregationSelf\-Consistency45\.5145\.5181\.4681\.4639\.2339\.2367\.6067\.6081\.6581\.6571\.2571\.2564\.4564\.45Team Selection BaselinesChoice\-SoftRandom\-kk39\.1039\.1065\.7165\.7136\.8336\.8366\.3166\.3175\.5575\.5566\.6366\.6358\.3658\.36Caruana47\.00\\mathbf\{47\.00\}84\.55\\mathbf\{84\.55\}40\.9240\.9267\.2367\.2383\.7183\.7170\.13\\mathbf\{70\.13\}65\.59\\mathbf\{65\.59\}Quality\-Only45\.3245\.3280\.2480\.2442\.88\\mathbf\{42\.88\}67\.98\\mathbf\{67\.98\}82\.9682\.9669\.7669\.7664\.8664\.86C2\-MAS \(Ours\)45\.6945\.69↑\\uparrow6\.5980\.2480\.24↑\\uparrow14\.5341\.9541\.95↑\\uparrow5\.1267\.98\\mathbf\{67\.98\}↑\\uparrow1\.6784\.18\\mathbf\{84\.18\}↑\\uparrow8\.6369\.7669\.76↑\\uparrow3\.1364\.9764\.97↑\\uparrow6\.61PoERandom\-kk39\.2439\.2466\.6366\.6336\.7036\.7066\.4766\.4776\.3876\.3866\.7666\.7658\.7058\.70Caruana47\.00\\mathbf\{47\.00\}83\.24\\mathbf\{83\.24\}41\.4841\.4868\.07\\mathbf\{68\.07\}84\.64\\mathbf\{84\.64\}69\.94\\mathbf\{69\.94\}65\.73\\mathbf\{65\.73\}Quality\-Only46\.7246\.7278\.9378\.9342\.04\\mathbf\{42\.04\}67\.7967\.7984\.4684\.4669\.6669\.6664\.9364\.93C2\-MAS \(Ours\)45\.2245\.22↑\\uparrow5\.9878\.9378\.93↑\\uparrow12\.3041\.5741\.57↑\\uparrow4\.8767\.7967\.79↑\\uparrow1\.3284\.5584\.55↑\\uparrow8\.1769\.6669\.66↑\\uparrow2\.9064\.6264\.62↑\\uparrow5\.92DSRandom\-kk38\.7138\.7168\.3868\.3836\.3336\.3366\.6466\.6475\.4775\.4768\.2968\.2958\.9758\.97Caruana44\.7644\.7684\.08\\mathbf\{84\.08\}39\.7939\.7967\.2367\.2383\.1583\.1570\.79\\mathbf\{70\.79\}64\.97\\mathbf\{64\.97\}Quality\-Only44\.94\\mathbf\{44\.94\}80\.5280\.5240\.3640\.3668\.35\\mathbf\{68\.35\}82\.8782\.8770\.79\\mathbf\{70\.79\}64\.6464\.64C2\-MAS \(Ours\)44\.1944\.19↑\\uparrow5\.4880\.5280\.52↑\\uparrow12\.1442\.04\\mathbf\{42\.04\}↑\\uparrow5\.7168\.35\\mathbf\{68\.35\}↑\\uparrow1\.7183\.90\\mathbf\{83\.90\}↑\\uparrow8\.4370\.79\\mathbf\{70\.79\}↑\\uparrow2\.5064\.97\\mathbf\{64\.97\}↑\\uparrow6\.00StackingRandom\-kk41\.0441\.0471\.9071\.9038\.4338\.4367\.2167\.2178\.8178\.8169\.3569\.3561\.1261\.12Caruana46\.3546\.3584\.9384\.9343\.2643\.2668\.45\\mathbf\{68\.45\}83\.8083\.8070\.69\\mathbf\{70\.69\}66\.2566\.25Quality\-Only46\.5446\.5485\.96\\mathbf\{85\.96\}43\.6343\.6368\.3568\.3583\.4383\.4370\.69\\mathbf\{70\.69\}66\.4366\.43C2\-MAS \(Ours\)46\.72\\mathbf\{46\.72\}↑\\uparrow5\.6885\.96\\mathbf\{85\.96\}↑\\uparrow14\.0643\.91\\mathbf\{43\.91\}↑\\uparrow5\.4868\.3568\.35↑\\uparrow1\.1484\.83\\mathbf\{84\.83\}↑\\uparrow6\.0270\.69\\mathbf\{70\.69\}↑\\uparrow1\.3466\.74\\mathbf\{66\.74\}↑\\uparrow5\.62
### E\.2Team Composition Analysis
To further examine how heterogeneity\-aware selection changes the selected team, we report in[Table6](https://arxiv.org/html/2609.38274#A5.T6)the benchmarks where C2\-MAS selects a different team from Quality\-Only underStackingaggregation\. On the seven primary benchmarks, the two methods select identical teams on four benchmarks, with a mean Jaccard similarity of0\.790\.79\. The three differing teams share the same pattern: C2\-MAS keeps the strongest shared anchor \(Nemotron, Gemma\-2, or Qwen3\) and substitutes one model from the central high\-accuracy cluster with a peripheral one \(GLM\-4, Ministral\-3\)\. This pattern aligns with the heterogeneity landscape in[Figure4](https://arxiv.org/html/2609.38274#S4.F4), where peripheral models provide complementary error coverage despite slightly lower individual accuracy\.
Table 6:Team composition changes made by C2\-MAS relative to Quality\-Only underStackingaggregation\. We report only benchmarks where the selected teams differ\.Δ\\DeltaAcc\. denotes the accuracy difference \(pp\) between C2\-MAS and Quality\-Only; arrows indicate gains or losses\.BenchmarkQuality\-Only TeamC2\-MAS TeamΔ\\DeltaAcc\.CommonsenseQA![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/gemma.png)Gemma\-2![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/qwen.png)Qwen3![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/internlm.png)InternLM3![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/gemma.png)Gemma\-2![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/qwen.png)Qwen3![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/glm.png)GLM\-4↑2\.43\\uparrow 2\.43MMLU\-Pro![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/nemotron.png)Nemotron![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/qwen.png)Qwen3![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/gemma.png)Gemma\-2![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/nemotron.png)Nemotron![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/qwen.png)Qwen3![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/ministral.png)Ministral\-3↑1\.78\\uparrow 1\.78OpenBookQA![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/nemotron.png)Nemotron![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/gemma.png)Gemma\-2![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/internlm.png)InternLM3![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/nemotron.png)Nemotron![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/gemma.png)Gemma\-2![[Uncaptioned image]](https://arxiv.org/html/2609.38274v1/figures/logos/glm.png)GLM\-4↑1\.40\\uparrow 1\.40
### E\.3Robustness and Ablation Analysis
We examine the robustness of C2\-MAS along two axes\.[Figure5](https://arxiv.org/html/2609.38274#A5.F5)sweeps team sizek∈\{2,3,4,5\}k\\in\\\{2,3,4,5\\\}under all four aggregation rules and shows that C2\-MAS maintains a consistent margin over Quality\-Only and Caruana across budgets, confirming that the gains reported withk=3k=3are not an artifact of a specific team size\.[Figure6](https://arxiv.org/html/2609.38274#A5.F6)then assesses whether the optimization objective is a faithful proxy for downstream accuracy: across the 13 benchmarks, the Spearman rank correlation between objective scores and ground\-truth accuracy over all\(113\)\\binom\{11\}\{3\}teams averagesρ=0\.751\\rho=0\.751, supporting the validity of our selection criterion\.
\(a\)Team size sweep \(kk\): Choice\-Soft\.\(b\)Team size sweep \(kk\): PoE\.\(c\)Team size sweep \(kk\): DS\.\(d\)Team size sweep \(kk\): Stacking\.
Figure 5:Performance scaling with team size \(kk\) across different aggregators\. We compare the average test accuracy of C2\-MAS \(ours\) against Quality\-Only and Caruana baselines as the team sizekkvaries from 2 to 5\. The evaluation is performed under four distinct inference\-time aggregation strategies: \(a\)Choice\-Soft, \(b\)PoE, \(c\)DS, and \(d\)Stacking\. C2\-MAS demonstrates consistent improvements over baselines across all strategies, particularly inStackingandDS, indicating that our heterogeneity\-aware selection is robust to the choice of downstream combination mechanism\.Figure 6:Validity of the optimization objective\. To assess the reliability of our selection criterion, we compute the Spearman rank correlation coefficient \(ρ\\rho\) between the team scores derived from our objective function and their actual downstream test accuracy, calculated across all\(113\)\\binom\{11\}\{3\}possible team combinations for each task\. The histogram illustrates the distribution of these correlations across the 13 benchmarks\. With a high mean correlation of0\.751, the results confirm that our proposed objective serves as an effective proxy for ground\-truth performance, allowing C2\-MAS to accurately rank candidate teams based solely on profiling data\.
### E\.4Stability of Pooled Heterogeneity Estimates
[Figure7](https://arxiv.org/html/2609.38274#A5.F7)reports the variability of pairwise HI entries under repeated subsampling of development examples\. The reduced variability after pooling complements the selection analysis:[SectionC\.6](https://arxiv.org/html/2609.38274#A3.SS6)maps changes in pooled, centered signals to changes in standardized pairwise scores, and[PropositionC\.7](https://arxiv.org/html/2609.38274#A3.Thmtheorem7)identifies when those changes preserve the selected team\.
Figure 7:Empirical stability of pooled HI under dev subsampling\. We subsample 70% of each task’s dev set for 50 repetitions, recompute per\-task HI matrices \(Yule’s\-QQRRand JSD\), and form pooled matrices from these noisy estimates\. Pooling substantially reduces the across\-repetition standard deviation of pairwise entries, complementing the selection\-stability analysis in[SectionsC\.5](https://arxiv.org/html/2609.38274#A3.SS5)and[C\.6](https://arxiv.org/html/2609.38274#A3.SS6)\.
## Appendix FLimitations
Our current formulation ofHIerr\\mathrm\{HI\}\_\{err\}andHIdist\\mathrm\{HI\}\_\{dist\}targets discriminative tasks with a shared finite label space, and does not directly cover open\-ended generation, code synthesis, or multi\-turn interactive settings\. However, the decoupled pipeline extends naturally to generative settings by redesigning the metrics:HIerr\\mathrm\{HI\}\_\{err\}can be adapted to pass@kkdecorrelation \(co\-failures on unit tests\), andHIdist\\mathrm\{HI\}\_\{dist\}to embedding\-based distances \(e\.g\., BERTScore\) between open\-ended outputs\. The selection procedure itself is unchanged, and we leave empirical study to future work\.
Our profiling also assumes access to representative development data for each candidate model\. Although cross\-task pooling reduces estimation noise and improves transfer to held\-out benchmarks, the selected team may still be sensitive to distribution shift when development tasks poorly match deployment conditions\.
Finally, C2\-MAS addresses offline team selection rather than adaptive per\-instance routing or interactive collaboration\. The selected team is fixed for a task and can be paired with different aggregation rules, but our experiments do not study dynamic protocols such as debate, tool use, or iterative verification\. Studying how the heterogeneity signals interact with such protocols is a natural next step\.Similar Articles
You're Hired: Strategic Model Selection for LLM Collaboration
This paper introduces a taxonomy of 9 model selection algorithms for multi-LLM collaboration, showing that capability-aware selection strategies outperform random or heuristic team assembly by up to 36.1% across math, coding, QA, and reasoning tasks.
Mixture of Complementary Agents for Robust LLM Ensemble
Proposes a framework for selecting complementary LLMs as proposers in ensemble systems, reformulating proposer selection as a combinatorial problem and exploring greedy algorithms for efficient performance-cost trade-offs.
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
This paper audits five diversity measures for LLM ensembles, finding that their associations with majority-vote gain are heavily entangled with model capability and are unstable after controlling for capability. The only robust signal is a modest residual pairwise co-failure association.
Which Pairs to Compare for LLM Post-Training?
This paper studies the problem of selecting which completion pairs to label for human preference feedback in LLM post-training. It formulates comparison curation as a sampling-design problem, provides theoretical bounds on DPO's policy optimality gap, and proposes practical sampling designs that improve sample efficiency over common heuristics on synthetic and real benchmarks.
Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
This paper investigates whether aggregating probability estimates from multiple LLMs exhibits a wisdom-of-crowds effect, finding that learned aggregators outperform individual models and that training cutoff contamination is a pervasive confound in such evaluations.