You're Hired: Strategic Model Selection for LLM Collaboration
Summary
This paper introduces a taxonomy of 9 model selection algorithms for multi-LLM collaboration, showing that capability-aware selection strategies outperform random or heuristic team assembly by up to 36.1% across math, coding, QA, and reasoning tasks.
View Cached Full Text
Cached at: 10/01/26, 09:45 AM
# You’re Hired: Strategic Model Selection for LLM Collaboration
Source: [https://arxiv.org/html/2609.38816](https://arxiv.org/html/2609.38816)
Zongwan Cao Ziyuan Yang11footnotemark:1Shangbin Feng11footnotemark:1††thanks:equal contributionAffiliation:University of Washington University of Southern California Allen Institute for AI\{zongwanc,ziyuan86\}@uw\.edushangbin@cs\.washington\.edu
###### Abstract
While multi\-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models \(LLMs\), existing systems remain bottlenecked on pre\-defined and hand\-crafted model pools\. In this work, we investigate the problem of*model selection in multi\-LLM systems*\. We propose and systematically evaluate a taxonomy of 9 selection algorithms ranging from diversity of model descriptions, capability\-aware behavioral diversity, and LLM\-based recruiters\. We conduct extensive experiments across two candidate pools of 10 and 32 models, deployed in four model collaboration algorithms, and evaluated across tasks spanning math, coding, QA, and reasoning\. Results demonstrate that successful selection algorithms greatly outperform random or heuristics\-based teams such as merely selecting the models with top individual performance, by up to 36\.1% across settings\. Specifically, capability\- and training\-based selection strategies alleviate selection variance and achieve the best performance, which we recommend to employ before deploying real\-world multi\-LLM systems\. Further analysis reveals that larger candidate pools pose greater challenges to shallow selection heuristics, while algorithms grounded in interacting with candidate models and understanding model capability robustly filter out misaligned, unsafe models, as well as generalizing to novel, out\-of\-distribution tasks\. Together, we establish that principled and informed team selection is critical and present strong model selection algorithms for assembling effective multi\-LLM systems\.
## 1Introduction
Recent advances in Large Language Models \(LLMs\) have fueled a transition from single\-model inference to multi\-model collaborative systems\. Frameworks such as multi\-agent debate\([Du et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib5)\), dynamic routing\([Feng et al\., 2026a](https://arxiv.org/html/2609.38816#bib.bib20);[Ong et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib4)\), and parameter fusion\([Yu et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib26);[Yadav et al\., 2023](https://arxiv.org/html/2609.38816#bib.bib25)\)demonstrate that combining multiple LLMs can overcome individual biases\([Liang et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib16);[Zhao et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib3)\), reduce hallucinations\([Du et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib5)\), and push the frontier of task performance\([Feng et al\., 2026a](https://arxiv.org/html/2609.38816#bib.bib20);[Feng et al\., 2026c](https://arxiv.org/html/2609.38816#bib.bib21)\)\. Yet, the effectiveness of these collaborative architectures rests on a critical, often overlooked assumption:a well\-composed team\. With over 2 million open language models available\([Wolf et al\., 2020](https://arxiv.org/html/2609.38816#bib.bib2)\), assembling a synergistic pool of collaborator models beyond manual cherry\-picking remains an open challenge\.
Currently, multi\-LLM teams are typically assembled via ad\-hoc heuristics, uniform random sampling, or by simply grouping the largest available models\. Yet these practices overlook a central challenge: models that are individually strong or superficially diverse may not necessarily form an effective team\([Kuncheva and Whitaker, 2003](https://arxiv.org/html/2609.38816#bib.bib13);[Sanchez and Hitaj, 2025](https://arxiv.org/html/2609.38816#bib.bib1)\)\. This motivates a fundamental question:*How should we select models to form an effective collaborative team?*In this work, we systematically investigate the science of model selection in multi\-LLM systems\. We formalize the candidate selection problem and investigate a comprehensive taxonomy of selection strategies that differ in the signals used to characterize candidate models and the procedures used to form teams\. Specifically, we investigate and propose five categories of algorithms:*standard baselines*such as employing models with best solo scores,*stated diversity*from model descriptions,*capability\-aware diversity*capturing observed behavioral patterns,*hybrid*approaches ensembling description and capability diversity, as well as*LLM\-based recruiters*\.
We evaluate these selectors across two candidate pools of varying heterogeneity with 10 and 32 models, diverse benchmarks, and four distinct collaboration mechanisms\. Specifically, our evaluation encompasses tasks spanning mathematical problem solving, code generation, question answering, and complex reasoning\. Our results reveal four key insights: First, randomly assembled teams underperform and introduce substantial performance variance, highlighting the necessity of fine\-grained team selection; Second, capability\-aware strategies frequently build stronger teams that reliably outperform the best single model and heuristics\-based baselines; Third, capability\-aware selection is the safest default as they consistently emerged as the top performers\. If adversarial robustness is the priority, quality\-filtered approaches \(such as Agentic Top\-kk, Nested Diversity\) are recommended to select out the malicious models\. Furthermore, we demonstrate that effective selection strategies scales robustly in redundant ecosystems and generalizes to out\-of\-distribution tasks\. Ultimately, who you select for a multi\-LLM team matters just as much as how they collaborate, and our algorithms provide a practical blueprint for hiring the right models for the job\.
Figure 1:Overview of our work\. Given a candidate model pool, we study how to select a team for downstream collaboration\. Our selectors consist of standard baselines, stated diversity, capability\-aware\-diversity, hybrid approaches, and LLM\-based recruiters\. The selected models are then composed into a team and evaluated across diverse collaboration paradigms\.
## 2Related Work
#### Multi\-LLM Collaboration\.
The paradigm of combining multiple language models takes several forms\. Routing mechanisms aim to dynamically assign inputs to the most suitable expert\([Ong et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib4)\)\. Generative fusion\([Jiang et al\., 2023](https://arxiv.org/html/2609.38816#bib.bib10)\)and multi\-agent debate\([Du et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib5)\)allow models to iteratively critique and refine each other’s outputs\. At the parameter level, techniques like DARE\([Yu et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib26)\)and TIES\([Yadav et al\., 2023](https://arxiv.org/html/2609.38816#bib.bib25)\)merge the weights of distinct models into a single unified model\. While these methods provide the infrastructure for collaboration, they generally assume the input ensemble is pre\-determined, leaving the question of optimal team composition largely unaddressed\.
#### Model Diversity and Selection\.
Diversity is a key consideration in ensemble construction, as complementary models can reduce errors\([Kuncheva and Whitaker, 2003](https://arxiv.org/html/2609.38816#bib.bib13)\)\. Recent works shows that diverse expertise across LLM agents can improve collective performance\([Zhang et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib37)\), while structural metadata can be used to characterize model relationships in the broader LLM ecosystem\([Li et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib14);[Laufer et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib15)\)\. However, these structural proxies do not directly characterize behavioral complementarity\. Our work bridges this gap by adapting combinatorial search algorithms to fine\-grained, behavioral signals, ensuring that selected models are diverse and effective\.
#### Capability Modeling\.
Quantifying the intrinsic abilities of LLMs has evolved beyond simple benchmark leaderboards\([Liang et al\., 2023](https://arxiv.org/html/2609.38816#bib.bib18)\)\. Methods such as Item Response Theory \(IRT\), originally developed for scaling NLP evaluation\([Lalor et al\., 2016](https://arxiv.org/html/2609.38816#bib.bib17)\), are recently adapted to map LLMs into continuous latent ability spaces based on item\-level success patterns\([Chen et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib24)\)\. Similarly, attribution confusability\([Sun et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib23)\)measures behavioral overlap by analyzing whether a classifier can distinguish between the text generated by different models\. We leverage these advanced capability representations not for leaderboard ranking, but as the underlying signals for our diversity\-seeking selection algorithms\.
## 3Methodology
#### Problem Formulation
Letℳ=\{m1,…,mP\}\\mathcal\{M\}=\\\{m\_\{1\},\\ldots,m\_\{P\}\\\}denote a pool ofPPcandidate language models, Given a target team sizenn, a selectorσ\\sigmachooses a subset of models from the model pool to form teamTT:
σ\(ℳ,n\)=T,T⊆ℳ,\|T\|=n,\\sigma\(\\mathcal\{M\},n\)=T,\\qquad T\\subseteq\\mathcal\{M\},\\;\|T\|=n,\(1\)
Given a collaboration methodc∈𝒞c\\in\\mathcal\{C\}and a taskd∈𝒟d\\in\\mathcal\{D\}, the performance of the selected teamTTunder methodccon taskddis denoted byscore\(T,c,d\)\\textit\{score\}\(T,c,d\)\. In this work, we study how the choice of selectorσ\\sigmaaffectsscoreunder different collaboration methods and tasks\.
Table 1:Selection strategy taxonomy\.Selectorσ\\sigmaSignal𝒵\\mathcal\{Z\}ProcedureI\. Standard BaselinesRandomNoneRandom samplingTop sizeParameter countDirect rankingTop solo scoreSolo performanceDirect rankingII\. Stated DiversityDescription diversityModel descriptionExact \(Alg\.[1](https://arxiv.org/html/2609.38816#alg1)\)III\. Capability\-Aware DiversityIdiosyncrasiesAttribution behaviorGreedy \(Alg\.[2](https://arxiv.org/html/2609.38816#alg2)\)IRT abilityIRT abilityGreedy \(Alg\.[2](https://arxiv.org/html/2609.38816#alg2)\)Performance profileBenchmark profileExact \(Alg\.[1](https://arxiv.org/html/2609.38816#alg1)\)IV\. HybridCombinedDescription \+ IRTExact \(Alg\.[1](https://arxiv.org/html/2609.38816#alg1)\)NestedDescription \+ IRTFilter\+\+Alg\.[1](https://arxiv.org/html/2609.38816#alg1)V\. LLM\-based recruitersLLM promptModel descriptionsDirect selectionSFT ClassifierPredicted team scoreTeam rankingAgentic top\-kkInterview scoreDirect ranking
#### Selection Strategies
Selecting a teamTTrequires information about candidate models that can inform the selection decision\. We consider three common forms of such information: \(i\)*stated*metadata, such as model cards\([Mitchell et al\., 2019](https://arxiv.org/html/2609.38816#bib.bib11)\); \(ii\)*demonstrated behavior*via item\-level evaluation\([Lalor et al\., 2016](https://arxiv.org/html/2609.38816#bib.bib17);[Chen et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib24)\)or behavioral fingerprinting\([Sun et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib23)\); and \(iii\) external*judgment*from strong LLMs or human raters\([Zheng et al\., 2023](https://arxiv.org/html/2609.38816#bib.bib12)\)\. Based on these observations, we instantiate our selector by two components:Signal𝒵\\mathcal\{Z\}andProcedure\. The signal captures candidate\-level attributes, pairwise similarities, or whole\-team judgments, while the procedure maps𝒵\\mathcal\{Z\}to teamTTvia ranking, search \(Algorithms[1](https://arxiv.org/html/2609.38816#alg1)and[2](https://arxiv.org/html/2609.38816#alg2)\), filtering, or direct selection\. Based on how these signals are constructed and used, we organize the resulting strategies into five families\. Table[1](https://arxiv.org/html/2609.38816#S3.T1)summarizes their signals and selection procedures\.
### 3\.1Signal Construction
Different selectors rely on different information about the candidate models\. The information𝒵\\mathcal\{Z\}may include model descriptions, observed capabilities, or behavioral patterns\. Each strategy constructs𝒵\\mathcal\{Z\}differently and then applies a corresponding selection procedure to form the teamTT\.
For strategies that explicitly seek model diversity, we convert𝒵\\mathcal\{Z\}into a pairwise similarity matrix𝐒=\[si,j\]∈ℝP×P\\mathbf\{S\}=\[s\_\{i,j\}\]\\in\\mathbb\{R\}^\{P\\times P\}, wheresi,js\_\{i,j\}measures the similarity between candidatesmim\_\{i\}andmjm\_\{j\}under the corresponding signal\.
#### Baselines\.
We consider three reference methods that do not explicitly model complementarity\.*Random*uses no signal\.*Top size*uses model parameter counts as the selection signal\.*Top solo score*uses performance on the development set of target task of each model as signal\.
#### Stated diversity\.
We next use models’ stated descriptions to characterize their different intended specializations\. For*Description diversity*, we define the selection signal as𝒵=\{desc\(mi\)\}i=1P\\mathcal\{Z\}=\\\{\\operatorname\{desc\}\(m\_\{i\}\)\\\}\_\{i=1\}^\{P\}\. We then encode description using a sentence transformerϕ\\phi, and construct the pairwise similarity matrix assi,j=cos\(ϕ\(desc\(mi\),ϕ\(desc\(mj\)\)CLOSECLOSEs\_\{i,j\}=\\cos\(\\phi\(\\operatorname\{desc\}\(m\_\{i\}\),\\phi\(\\operatorname\{desc\}\(m\_\{j\}\)\)\.
#### Capability\-aware diversity\.
We characterize model diversity through observed capabilities and behavioral patterns\.*Idiosyncrasies*and*IRT*also derive behavioral signals from model responses on the development set of evaluation tasks, and*Performance Profile*characterizes models on a held\-out suite of benchmark tasks spanning math, code, QA, and reasoning, rather from the evaluation tasks\.
- •*Idiosyncrasies*: Following[Sun et al\. \(2025\)](https://arxiv.org/html/2609.38816#bib.bib23), we train a text classifier to predict which candidate model generated a given response\. Letci,jc\_\{i,j\}denote the probability that the classifier predicts modeljjfor a response generated by modelii\. We use the resulting confusion probabilities as the selection signal𝒵=\{ci,j\}i,j=1P\\mathcal\{Z\}=\\\{c\_\{i,j\}\\\}\_\{i,j=1\}^\{P\}, and construct the pairwise similarity assi,j=12\(ci,j\+cj,i\)s\{i,j\}=\\frac\{1\}\{2\}\(c\_\{i,j\}\+c\_\{j,i\}\)\.
- •*IRT Ability*: Following[Chen et al\. \(2025\)](https://arxiv.org/html/2609.38816#bib.bib24), we train an IRT model to estimate an ability vectorθi\\theta\_\{i\}for each candidate from its item\-level successes and failures\. We use these ability vectors as the selection signal𝒵=\{θi\}i=1P\\mathcal\{Z\}=\\\{\\theta\_\{i\}\\\}\_\{i=1\}^\{P\}, and construct the pairwise similarity assi,j=cos\(θi,θj\)s\_\{i,j\}=\\cos\(\\theta\_\{i\},\\theta\_\{j\}\)\.
- •*Performance Profile*: We represent each candidate model by its performance vectorpi∈ℝ10p\_\{i\}\\in\\mathbb\{R\}^\{10\}across held\-out benchmarks\. We use these performance vectors as the selection signal𝒵=\{pi\}i=1P\\mathcal\{Z\}=\\\{p\_\{i\}\\\}\_\{i=1\}^\{P\}, and construct the pairwise similarity assi,j=cos\(pi,pj\)s\_\{i,j\}=\\cos\(p\_\{i\},p\_\{j\}\)\.
#### Hybrid\.
We combine signals from*Description Diversity*and*IRT Ability*in two ways\. For*Combined*, we concatenate the description embeddingdesc\(mi\)\\operatorname\{desc\}\(m\_\{i\}\)and IRT ability embeddingθi\\theta\_\{i\}into a joint representationhi=\[vi;θi\]h\_\{i\}=\[v\_\{i\};\\theta\_\{i\}\]\. We use these joint representations as the selection signal,𝒵=\{hi\}i=1P\\mathcal\{Z\}=\\\{h\_\{i\}\\\}\_\{i=1\}^\{P\}and computesi,j=cos\(hi,hj\)s\_\{i,j\}=\\cos\(h\_\{i\},h\_\{j\}\)\.*Nested*retains the description and IRT signals separately and use them sequentially\.
#### LLM\-based recruiters\.
Finally, rather than constructing a fixed pairwise similarity matrix𝐒\\mathbf\{S\}, these methods use learned or LLM\-elicited judgments as selection signals\. For*LLM prompt*, the signal consists of model descriptions𝒵=\{desc\(mi\)\}i=1P\\mathcal\{Z\}=\\\{\\operatorname\{desc\}\(m\_\{i\}\)\\\}\_\{i=1\}^\{P\}\. For*SFT Classifier*,𝒵\\mathcal\{Z\}consists of predicted team scores produced by a learned regression model from the descriptions of the team members\. For*Agentic top\-kk*, the signal consists of scalar interview scores,𝒵=\{qi\}i=1P\\mathcal\{Z\}=\\\{q\_\{i\}\\\}\_\{i=1\}^\{P\}, whereqiq\_\{i\}is assigned to candidatemim\_\{i\}by an LLM recruiter\. We provide details in Appendix[C](https://arxiv.org/html/2609.38816#A3)\.
### 3\.2Selection Procedures
Given a selection signal𝒵\\mathcal\{Z\}, each selector applies a corresponding procedure to form a teamTTof sizenn\. Depending on the selector, this procedure takes the form of direct sampling or ranking, diversity\-based search, sequential filtering, or LLM\-based selection\.
#### Baselines\.
The baseline selectors require no complex search\.*Random*uniformly samples a team ofnncandidates from model pool\.*Top size*and*Top solo score*directly rank candidates by their respective scalar signals and select the topnncandidates\.
#### Diversity\-based selection\.
For selectors that construct a pairwise similarity matrix𝐒\\mathbf\{S\}, we define team diversity using a max–min dispersion criterion, with average pairwise similarity as a tie\-breaker\. Specifically, we define the*most*diverse team as the one that minimizes the maximum pairwise similarity, whereas the*least*diverse team maximizes it\. We additionally compare against average\-dispersion objectives in Appendix[B](https://arxiv.org/html/2609.38816#A2)\. We consider two selection procedures:
- •*Exact diversity search*\(Algorithm[1](https://arxiv.org/html/2609.38816#alg1)\): We enumerate all candidate teams of sizennand select a team according to the diversity criterion\. We use this procedure for*Description Diversity*and*Performance Profile*, where pairwise similarity provides a direct diversity signal\.
- •*Capability\-seeded greedy search*\(Algorithm[2](https://arxiv.org/html/2609.38816#alg2)\): We incorporate candidate capability into the search and use this procedure for*Idiosyncrasies*and*IRT Ability*, where weak models may produce unusual errors and therefore appear distinctive\. The search starts from the two highest\-performing candidates and greedily adds the remaining members according to the diversity criterion, reducing the risk of selecting weak models solely for their distinctiveness\.
#### Hybrid\.
These methods combine the signals above in two ways\.*Combined*applies Algorithm[1](https://arxiv.org/html/2609.38816#alg1)to the joint representationhih\_\{i\}\.*Nested*first retains the stronger half of the candidate pool based on*IRT Ability*scoresθi\\theta\_\{i\}, then applies Algorithm[1](https://arxiv.org/html/2609.38816#alg1)to the remaining candidates using the description\-based similarity signal\.
#### LLM\-based recruiters\.
These procedures bypass pairwise similarity matrices and select teams using learned or zero\-shot LLM recruiters:*LLM prompt*: directly selects a team from candidate descriptions using a zero\-shot LLM recruiter\.*SFT Classifier*: enumerates all teams of sizenn, scores each with a fine\-tuned LLM\-based regressor, and selects the highest\-scoring team\.*Agentic top\-kk*: conducts multi\-turn LLM interviews with each candidate, ranks candidates by their interview scores, and selects the topnnmodels\. We provide further details of these procedures in Appendix[C](https://arxiv.org/html/2609.38816#A3)\.
Algorithm 1Exact Diversity Search0:Similarity matrix
𝐒∈ℝP×P\\mathbf\{S\}\\in\\mathbb\{R\}^\{P\\times P\}, team size
nn, mode
∈\{most,least\}\\in\\\{\\textsc\{most\},\\textsc\{least\}\\\}
1:
T⋆←∅T^\{\\star\}\\leftarrow\\varnothing
2:foreach
T⊆\{1,…,P\}T\\subseteq\\\{1,\\ldots,P\\\}with
\|T\|=n\|\{\}T\|\{\}=ndo
3:
smax\(T\)←maxi,j∈T,i≠jsi,js\_\{\\text\{max\}\}\(T\)\\leftarrow\\max\_\{i,j\\in T,\\,i\\neq j\}s\_\{i,j\}
4:
savg\(T\)←1/\(n2\)∑i,j∈T,i<jsi,js\_\{\\text\{avg\}\}\(T\)\\leftarrow 1/\\binom\{n\}\{2\}\\sum\_\{i,j\\in T,\\,i<j\}s\_\{i,j\}
5:
s\(T\)←\(smax\(T\),savg\(T\)\)s\(T\)\\leftarrow\(s\_\{\\text\{max\}\}\(T\),s\_\{\\text\{avg\}\}\(T\)\)
6:ifmode
=most=\\textsc\{most\}and
s\(T\)<s\(T⋆\)s\(T\)<s\(T^\{\\star\}\)then
7:
T⋆←TT^\{\\star\}\\leftarrow T
8:elseifmode
=least=\\textsc\{least\}and
s\(T\)\>s\(T⋆\)s\(T\)\>s\(T^\{\\star\}\)then
9:
T⋆←TT^\{\\star\}\\leftarrow T
10:endif
11:endfor
12:return
T⋆T^\{\\star\}
Algorithm 2Capability\-Seeded Greedy Search0:Similarity matrix
𝐒∈ℝP×P\\mathbf\{S\}\\in\\mathbb\{R\}^\{P\\times P\}, team size
nn, scores
\{cap\(i\)\}\\\{\\textit\{cap\}\(i\)\\\}, mode
∈\{most,least\}\\in\\\{\\textsc\{most\},\\textsc\{least\}\\\}
1:
seed←\\textit\{seed\}\\leftarrowtop\-2 indices by
cap\(⋅\)\\textit\{cap\}\(\\cdot\)
2:
T←seedT\\leftarrow\\textit\{seed\}
3:
R←\{1,…,P\}∖TR\\leftarrow\\\{1,\\ldots,P\\\}\\setminus T
4:while
\|T\|<n\|\{\}T\|\{\}<ndo
5:for
c∈Rc\\in Rdo
6:
s\(c\)←maxj∈Tsc,js\(c\)\\leftarrow\\max\_\{j\\in T\}s\_\{c,j\}
7:endfor
8:
c⋆←argminc∈Rs\(c\)c^\{\\star\}\\leftarrow\\arg\\min\_\{c\\in R\}s\(c\)ifmode
=most=\\textsc\{most\}else
argmaxc∈Rs\(c\)\\arg\\max\_\{c\\in R\}s\(c\)
9:
T←T∪\{c⋆\}T\\leftarrow T\\cup\\\{c^\{\\star\}\\\}
10:
R←R∖\{c⋆\}R\\leftarrow R\\setminus\\\{c^\{\\star\}\\\}
11:endwhile
12:return
TT
Table 2:Main collaborative performance results across all datasets and mechanisms forPool 1\(top\) andPool 2\(bottom\)\. The highest mean score in each column per pool isbolded\. Cells are colored to highlight the efficacy of adaptive selection based on mean performance:Light Greenindicates the team beats the*Random*baseline,Medium Greenindicates the team beats all standard baselines \(*Random*,*Top Solo*, and*Top Size*when applicable\), andDark Greenindicates the team beats all standard baselines plus the*Single Best Model*\.Pool 1 \(Homogeneous Candidates, 10→\\rightarrow4 models\)
Pool 2 \(Heterogeneous Candidates, 32→\\rightarrow4 models\)
## 4Experiment Details
#### Collaboration Methods
We evaluate selected teamTTunder four collaboration methods:*prompt routing*, which asks an LLM to route each query to the team member whose description best matches it;*multi\-agent refinement*\([Du et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib5)\), in which team members iteratively critique and revise a shared draft;*weight merging*, which fuses team members’ parameters into a single model via DARE\([Yu et al\., 2024](https://arxiv.org/html/2609.38816#bib.bib26)\)composed with TIES\([Yadav et al\., 2023](https://arxiv.org/html/2609.38816#bib.bib25)\); and*LLM blender*\([Jiang et al\., 2023](https://arxiv.org/html/2609.38816#bib.bib10)\), which pairwise\-ranks team members’ outputs and generatively fuses the top\-ranked subsets\.
#### Candidate models\.
We evaluate selectors across two distinct candidate poolsℳ\\mathcal\{M\}\. Pool 1 consists of 10 Qwen2\.5\-7B\-Instruct\([Qwen Team, 2024](https://arxiv.org/html/2609.38816#bib.bib36)\)models trained on different instruction\-tuning corpora\([Jiang et al\., 2025](https://arxiv.org/html/2609.38816#bib.bib34)\)\. Pool 2 consists of 32 independently contributed systems from the participatory ecosystem of[Feng et al\. \(2026c\)](https://arxiv.org/html/2609.38816#bib.bib21), spanning diverse architectures, scales, and training objectives\. For both pools, our goal is to select a team ofn=4n=4candidate models for collaboration\. Full model details are provided in Appendix[D](https://arxiv.org/html/2609.38816#A4)\.
#### Datasets\.
We evaluate collaborative performance on three main datasets: GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.38816#bib.bib7)\), TruthfulQA\([Lin et al\., 2022](https://arxiv.org/html/2609.38816#bib.bib35)\), and MBPP\([Austin et al\., 2021](https://arxiv.org/html/2609.38816#bib.bib6)\)\. Auxiliary datasets are used to construct the*Performance Profile*signal and train the*SFT Classifier*\.*Idiosyncrasies*,*IRT*, and*Top Solo Score*use the development set of the evaluation datasets to construct their selection signals, while all reported collaborative performance is evaluated on the test splits\. Also, the*SFT Classifier*is trained exclusively on teams sampled from Pool 1\. When applied to Pool 2, it scores all candidate teams without further training or adaptation, constituting a cross\-pool out\-of\-distribution transfer setting\. Dataset roles, sizes, and splits are detailed in Appendix[E](https://arxiv.org/html/2609.38816#A5)\.
#### Implementation\.
We useall\-MiniLM\-L6\-v2to encode model descriptions andall\-mpnet\-base\-v2for the learned behavioral representations\. Qwen2\.5\-7B\-Instruct serves as the backbone for LLM\-based selectors and LLM\-Blender\. Unless otherwise specified, generation uses temperature0\.70\.7, top\-pp0\.90\.9, and a maximum response length of 512 tokens \(1024 for MBPP\)\. Full selector and collaboration configurations are provided in Appendix[F](https://arxiv.org/html/2609.38816#A6)\.
#### Variance estimation\.
For deterministic selectors, we estimate test\-set uncertainty usingB=10,000B=10\{,\}000nonparametric bootstrap resamples over item\-level outcomes and report the mean with 95% percentile intervals\. For the*Random*selector, we measure selection variance over five independently sampled teams and report the corresponding mean and 95% percentile interval\.
## 5Results
Table[2](https://arxiv.org/html/2609.38816#S3.T2)reports the collaborative performance for all selection strategies\. The results reveal four key findings in multi\-LLM team composition:
Blind hiring masks severe variance\.Randomly selecting a team leads to considerable performance swings depending on the luck of the draw, yielding standard deviations that make the system unreliable\. For instance, Random selection on GSM8K yields intervals of±25\.3\\pm 25\.3under LLM\-Blender in Pool 2 and±13\.4\\pm 13\.4under DARE\-TIES in Pool 1\. In contrast, intentional selection strategies effectively eliminate this selection\-process variance, replacing unpredictable draws with deterministic, stable teams that exhibit tight confidence intervals \(typically±1\.0\\pm 1\.0to±2\.5\\pm 2\.5\)\.
Intentional selection can build stronger teams\.By systematically vetting candidates, capability\-aware selection strategies outperform naive heuristics\. In Pool 2, simply choosing the largest models \(*Top Size*\) yields catastrophic failures in generative and parameter\-fusion settings, scoring a mere14\.414\.4on GSM8K under DARE\-TIES compared to*Random*’s28\.428\.4\. Conversely, capability\-aware strategies can identify substantially stronger team compositions\. Methods like*Idiosyncrasies \(least\)*and*Nested \(most\)*frequently hit the highest performance tier, outscoring the*Random*baseline, the*Top Solo Score*baseline, and even the*Single Best Model*available in the entire pool\.
Team selection is critical: suboptimal teams actually undermine the benefit of collaboration methods\.Outperforming the single strongest individual candidate remains a high, context\-dependent bar\. When the best solo model is exceptionally strong, assembling a team that actually surpasses it requires highly precise composition\. In Pool 2, the single best model achieves82\.982\.9on GSM8K\. The majority of selection strategies fail to clear this threshold under Prompt Routing or DARE\-TIES\. However, specific optimized pairings, such as*Nested \(least\)*under Multiagent Refine or*Idiosyncrasies \(least\)*under LLM\-Blender demonstrate that a well\-composed team can still push past a highly capable solo expert\. The added value of collaboration depends heavily on both the inherent difficulty of the task and the baseline competence of the available pool\.
There is no one\-size\-fits\-all hiring strategy\.The success of a selection method fluctuates significantly depending on the downstream collaboration mechanism it feeds into\. A strategy that recruits a state\-of\-the\-art team for one collaborative framework can fail completely under another\. For example, in Pool 1,*Nested \(most\)*is the dominant strategy for Multiagent Refine on GSM8K \(72\.972\.9\), yet it struggles significantly when applied to MBPP under LLM\-Blender \(31\.631\.6\)\. Conversely,*Combined \(most\)*thrives under DARE\-TIES and LLM\-Blender for MBPP \(scoring59\.359\.3and58\.358\.3, respectively\) but falls to45\.645\.6under Multiagent Refine\. This confirms that team composition cannot be decoupled from team interaction; the underlying signal used to select collaborators must be directly tailored to how the team will ultimately merge or fuse their outputs\.
## 6Analysis
### 6\.1Robustness to Malicious Models
0\.50\.50\.60\.60\.70\.70\.80\.80\.90\.9*SFT Rec\.**Random**Desc\. Div\.**Idiosync\.**Perf\. Prof\.**Comb\. Div\.**IrtNet Div\.**LLM Prompt**Nested Div\.**Ag\. Top\-K*0\.6360\.6360\.6500\.6500\.6500\.6500\.6540\.6540\.6730\.6730\.7040\.7040\.7250\.7250\.7750\.7750\.8130\.8130\.8250\.825Avg\. Fraction CleanFigure 2:Average fraction of clean models in selected teams across contamination levels and randomized injection trials\. The red dashed line marks the Random baseline\.To evaluate robustness to misaligned candidates, we introduce six misaligned models from[Yang et al\. \(2026\)](https://arxiv.org/html/2609.38816#bib.bib22)\. Their descriptions follow the same innocuous one\-line format as aligned models, providing no explicit signal of misalignment to text\-based selectors\. At each contamination levelrr, we injectrrmisaligned models into a pool of 10 models and repeat the experiment 10 times with randomized removal and injection orders\. We measure the fraction of aligned models in the selected four\-model team with an expected fraction of\(10−r\)/10\(10\-r\)/10under random selection\.
Figure[2](https://arxiv.org/html/2609.38816#S6.F2)reveals a contrast between methods that passively rely on indirect signals and those that explicitly evaluate model capability\. Unanchored diversity\-based selectors are particularly vulnerable\. Semantic selectors like*Description Diversity*perform near a random baseline when model descriptions are misleading\. A full round\-by\-round breakdown can be found in Appendix[G](https://arxiv.org/html/2609.38816#A7)\.
In contrast, some selection methods equipped with explicit capability filters or agentic evaluation outperform the random baseline and are better equipped to avoid compromised candidates\.*Agentic Top\-K*achieves the highest clean\-model fraction, consistent with its interactive interview process providing a stronger capability signal\. Similarly,*Nested*first filters out lower\-capability candidates before applying diversity\-based selection, reducing the exposure of the downstream diversity search to malicious candidates\. Overall, these results suggest that robustness benefits from selection in observed model behavior or capability, rather than relying solely on stated descriptions\.
### 6\.2Effects of Pool\-Size Scaling
81624320\.40\.40\.50\.50\.60\.60\.70\.7Available Models \(PP\)TruthfulQA ScorePrompt Routing81624320\.40\.40\.50\.50\.60\.60\.70\.7Available Models \(PP\)LLM\-Blender*IRT\-Net \(most\)**Desc\. \(most\)**Nested \(most\)**Perf\. Prof\. \(least\)**Random*Figure 3:Pool\-size scaling on Pool 2\.Expanding the uncurated pool without strict capability constraints causes team performance to saturate or even degrade\.To evaluate how selection strategies scale with candidate pool size, we vary the number of available models in Pool 2 asP∈\{8,16,24,32\}P\\in\\\{8,16,24,32\\\}while fixing the selected team size ton=4n=4\. This setting allows us to examine how different selectors behave as the number of candidate collaborators increases\. Figure[3](https://arxiv.org/html/2609.38816#S6.F3)shows how collaborative performance changes as the candidate pool expands\. Initially, increasing the pool from small \(P=8P=8\) to moderately large \(P=16P=16\) helps some selection methods build stronger teams, suggesting that additional candidates provide useful opportunities for identifying complementary models\. However, further expanding the pool \(P=32P=32\) yields diminishing or negative returns for several strategies\. Semantic diversity methods tend to plateau, while some unconstrained approaches degrade at larger pool sizes, particularly LLM\-Blender\.
The pronounced drop of the*Random*baseline atP=32P=32is consistent with the expanded pool containing substantially weaker candidates\. Consequently, methods that prioritize diversity without accounting for baseline capability become susceptible to selecting these weak outliers into the team\. Overall, these results suggest that simply scaling up the candidate pool is not sufficient to improve collaboration\. As the pool expands, selection must balance diversity with candidate capability to avoid incorporating weak but superficially distinctive models\.
### 6\.3Effects of Pool Diversity Scaling
4×84\\times 88×48\\times 416×216\\times 232×132\\times 10\.30\.30\.40\.40\.50\.50\.60\.60\.70\.7Pool Composition \(kkdistinct models×\\timesmmcopies\)TruthfulQA ScorePrompt Routing4×84\\times 88×48\\times 416×216\\times 232×132\\times 10\.30\.30\.40\.40\.50\.50\.60\.60\.70\.7Pool Composition \(kkdistinct models×\\timesmmcopies\)LLM\-Blender*IRT\-Net \(most\)**Desc\. \(most\)**Nested \(most\)**Perf\. Profile \(least\)**Random*Figure 4:Pool diversity scaling on Pool 2\.Evaluating a 32\-slot pool across varying compositions ofkkdistinct models×\\timesmmcopies\. Naive diversity \(*Desc\.*\) saturates askkincreases, while behavioral diversity \(*IRT\-net*\) degrades at32×132\\times 1, risking the inclusion of low\-quality outliers\.To isolate the effect of candidate diversity from pool size, we fix the pool at 32 candidate slots and the selected team size atn=4n=4, while varying the number of distinct modelskkand the number of copies per modelmm:\(k,m\)∈\{\(4,8\),\(8,4\),\(16,2\),\(32,1\)\}\(k,m\)\\in\\\{\(4,8\),\(8,4\),\(16,2\),\(32,1\)\\\}\. Here, we operationalize pool diversity by the number of distinct model identities while holding the total number of candidate slots fixed\. We evaluate downstream performance on TruthfulQA\. Figure[4](https://arxiv.org/html/2609.38816#S6.F4)shows how collaborative performance changes as the number of distinct candidates increases\.
For several selectors, replacing duplicate candidates with distinct models initially improves performance, but the gains tend to saturate askkincreases\. At the highest\-diversity setting \(32×132\\times 1\), some selectors even degrade, indicating that greater distinctness does not necessarily translate into more useful team composition\. Once the pool already covers a broad range of capabilities, additional distinct models may instead introduce weaker candidates\. Overall, these results suggest that reducing redundancy can benefit collaboration, but maximizing distinctness alone is insufficient; effective selection must balance candidate diversity with capability\.
Table 3:OOD evaluation on Pool 2\.The highest mean score in each column per pool is bolded\.
### 6\.4Does Capability\-Aware Selection Generalize Out\-of\-Distribution?
The behavioral and learned selectors in Table[2](https://arxiv.org/html/2609.38816#S3.T2)construct their selection signals from task\-specific or auxiliary data\. To examine whether these signals transfer to unseen task domains, we evaluate selected methods on CoCoNot \(safety\) and GPQA\-Diamond \(graduate\-level science\)\. we evaluate three learned behavioral selectors—*IRT\-Net*,*Idiosyncrasies*, and*SFT Classifier*—on two unseen task domains\.*Performance Profile*is excluded from this analysis because their construction suite includes GPQA\-Diamond and CoCoNot \(Table[3](https://arxiv.org/html/2609.38816#S6.T3)\)\.
The*SFT Classifier*transfers consistently across the two unseen domains, outperforming*Random*in all four settings and achieving the highest mean performance in three of them\. In contrast, representation\-based behavioral signals show more scenario\-dependent transfer:*IRT\-Net*and*Idiosyncrasies*underperform*Random*in some Prompt Routing settings, while both improve over*Random*on GPQA\-Diamond under LLM\-Blender\.
## 7Conclusion
In this work, we formalized the problem of collaborator selection in multi\-LLM systems and introduced a comprehensive taxonomy of selection strategies\. Through extensive evaluation across diverse candidate pools, tasks, and collaboration mechanisms, we demonstrated that the composition of an LLM team is just as critical as the method by which they collaborate\. Relying on blind random sampling or simplistic heuristics introduces severe variance and limits the potential of multi\-agent frameworks\. Our findings highlight the importance of grounding selection in observed model capabilities rather than relying solely on model descriptions or simple heuristics\. These principled strategies successfully scale with pool size, resist the inclusion of misaligned models, and generalize to novel task distributions\. Ultimately, as the ecosystem of open\-source models continues to expand, building effective collaborative AI systems will require shifting our focus from merely*how*models interact to*who*is hired for the job in the first place\.
### AI use statement
In this work, we used a generative AI tool to assist with drafting and editing manuscript text, and to help identify and verify citations for prior work discussed throughout the paper\. We did not use generative AI tools to design the selection algorithms, collaboration methods, or experimental protocols presented in this work, nor to produce the underlying experimental results: all reported scores were obtained by running the described selection methods and MoCo collaboration methods on the stated benchmarks, independent of any AI tool\.
All AI\-assisted text was reviewed and edited by the authors, and all factual and numerical claims traceable to the underlying experimental data were checked against that data by the authors\. We take responsibility for the final content of this work, including all text, claims, and artifacts produced with the aid of generative AI\.
### Ethics statement
This work does not involve human subjects; the “interview” procedure in Section[I](https://arxiv.org/html/2609.38816#A9)is an automated exchange between two language models and involves no human participants, data, or annotation\.
The adversarial\-model\-detection study deliberately incorporates six models fine\-tuned to be misaligned, drawn from the existing, published Among Us benchmark rather than newly created for this paper\. These models are used strictly to test whether selection strategies can identify and exclude them from a collaborating team; they are not released, deployed, or used for any purpose beyond this detection benchmark, and no new misaligned model is introduced by this work\.
All base models and benchmarks used in this work are obtained from publicly released, appropriately licensed sources, and we do not introduce new data collected from or about individuals\. We are not aware of a direct harmful application of this work’s findings\.
### Reproducibility Statement
All code, datasets, and experiment logs required to reproduce the reported results will be released at a repository upon publication\.
## References
- Austinet al\.\(2021\)J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le, and C\. SuttonProgram synthesis with large language models\.External Links:2108\.07732,[Link](https://arxiv.org/abs/2108.07732)Cited by:[Table 5](https://arxiv.org/html/2609.38816#A6.T5.2.4.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px3.p1.1)\.
- Brahmanet al\.\(2024\)F\. Brahman, S\. Kumar, V\. Balachandran, P\. Dasigi, V\. Pyatkin, A\. Ravichander, S\. Wiegreffe, N\. Dziri, K\. Chandu, J\. Hessel, Y\. Tsvetkov, N\. A\. Smith, Y\. Choi, and H\. HajishirziThe art of saying no: contextual noncompliance in language models\.InThe Thirty\-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=f1UL4wNlw6)Cited by:[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.10.1),[Table 8](https://arxiv.org/html/2609.38816#A6.T8.2.3.1)\.
- Chenet al\.\(2025\)J\. Chen, C\. Wang, G\. Zhang, P\. Ye, L\. Bai, W\. Hu, Y\. Qu, and S\. HuLearning compact representations of llm abilities via item response theory\.External Links:2510\.00844,[Link](https://arxiv.org/abs/2510.00844)Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px3.p1.1),[2nd item](https://arxiv.org/html/2609.38816#S3.I1.i2.p1.1),[§3](https://arxiv.org/html/2609.38816#S3.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. de Oliveira Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder, M\. Pavlov, A\. Power, L\. Kaiser, M\. Bavarian, C\. Winter, P\. Tillet, F\. P\. Such, D\. Cummings, M\. Plappert, F\. Chantzis, E\. Barnes, A\. Herbert\-Voss, W\. H\. Guss, A\. Nichol, A\. Paino, N\. Tezak, J\. Tang, I\. Babuschkin, S\. Balaji, S\. Jain, W\. Saunders, C\. Hesse, A\. N\. Carr, J\. Leike, J\. Achiam, V\. Misra, E\. Morikawa, A\. Radford, M\. Knight, M\. Brundage, M\. Murati, K\. Mayer, P\. Welinder, B\. McGrew, D\. Amodei, S\. McCandlish, I\. Sutskever, and W\. ZarembaEvaluating large language models trained on code\.External Links:2107\.03374,[Link](https://arxiv.org/abs/2107.03374)Cited by:[Table 6](https://arxiv.org/html/2609.38816#A6.T6.2.4.1),[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.9.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try arc, the ai2 reasoning challenge\.External Links:1803\.05457,[Link](https://arxiv.org/abs/1803.05457)Cited by:[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.2.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.External Links:2110\.14168,[Link](https://arxiv.org/abs/2110.14168)Cited by:[Table 5](https://arxiv.org/html/2609.38816#A6.T5.2.2.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px3.p1.1)\.
- Duet al\.\(2024\)Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. MordatchImproving factuality and reasoning in language models through multiagent debate\.InForty\-first International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=zj7YuTE4t8)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1),[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px1.p1.1)\.
- Fenget al\.\(2026a\)S\. Feng, Y\. Bai, Z\. Yang, Y\. Wang, Z\. Tan, J\. Yan, Z\. Lei, W\. Ding, W\. Shi, H\. Wang,et al\.MoCo: a one\-stop shop for model collaboration research\.External Links:2601\.21257,[Link](https://arxiv.org/abs/2601.21257)Cited by:[§F\.3](https://arxiv.org/html/2609.38816#A6.SS3.p1.1),[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.11.1),[§1](https://arxiv.org/html/2609.38816#S1.p1.1)\.
- Fenget al\.\(2026b\)S\. Feng, W\. Ding, A\. Liu, Z\. Wang, W\. Shi, Y\. Wang, S\. Z\. Shen, X\. Han, H\. Lang, C\. Lee,et al\.When one llm drools, multi\-llm collaboration rules\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 17048–17063\.External Links:[Link](https://aclanthology.org/2026.acl-long.775/)Cited by:[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.11.1)\.
- Fenget al\.\(2026c\)S\. Feng, Y\. Wang, W\. Shi, L\. Zettlemoyer, Y\. Choi, and Y\. TsvetkovScaling participation in modular ai systems\.External Links:2606\.07812,[Link](https://arxiv.org/abs/2606.07812)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px2.p1.1)\.
- Gemaet al\.\(2025\)A\. P\. Gema, J\. O\. J\. Leang, G\. Hong, A\. Devoto, A\. C\. M\. Mancino, R\. Saxena, X\. He, Y\. Zhao, X\. Du, M\. R\. Ghasemi Madani, C\. Barale, R\. McHardy, J\. Harris, J\. Kaddour, E\. Van Krieken, and P\. MinerviniAre we done with MMLU?\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 5069–5096\.External Links:[Link](https://aclanthology.org/2025.naacl-long.262/)Cited by:[Table 6](https://arxiv.org/html/2609.38816#A6.T6.2.3.1),[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.4.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the MATH dataset\.InThirty\-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track \(Round 2\),External Links:[Link](https://openreview.net/forum?id=7Bywt2mQsCe)Cited by:[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.6.1)\.
- Jianget al\.\(2023\)D\. Jiang, X\. Ren, and B\. Y\. LinLLM\-blender: ensembling large language models with pairwise comparison and generative fusion\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14165–14178\.External Links:[Link](https://aclanthology.org/2023.acl-long.792/)Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px1.p1.1)\.
- Jianget al\.\(2025\)Y\. Jiang, W\. Ding, S\. Feng, G\. Durrett, and Y\. TsvetkovSparta alignment: collectively aligning multiple language models through combat\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=nTfhNThKX2)Cited by:[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px2.p1.1)\.
- Jinet al\.\(2021\)D\. Jin, E\. Pan, N\. Oufattole, W\. Weng, H\. Fang, and P\. SzolovitsWhat disease does this patient have? A large\-scale open domain question answering dataset from medical exams\.Applied Sciences\.Cited by:[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.7.1)\.
- Kuncheva and Whitaker \(2003\)L\. I\. Kuncheva and C\. J\. WhitakerMeasures of diversity in classifier ensembles and their relationship with the ensemble accuracy\.Machine Learning51\(2\),pp\. 181–207\.Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p2.1),[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px2.p1.1)\.
- Laloret al\.\(2016\)J\. P\. Lalor, H\. Wu, and H\. YuBuilding an evaluation scale using item response theory\.InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing,pp\. 648–657\.External Links:[Link](https://aclanthology.org/D16-1062/)Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.38816#S3.SS0.SSS0.Px2.p1.1)\.
- Lauferet al\.\(2025\)B\. Laufer, H\. Oderinwale, and J\. KleinbergAnatomy of a machine learning ecosystem: 2 million models on Hugging Face\.External Links:2508\.06811,[Link](https://arxiv.org/abs/2508.06811)Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2024\)Z\. Li, H\. van der Wilk, D\. Zhan, M\. Khosla, A\. Bozzon, and R\. HaiModel selection with model zoo via graph learning\.In2024 IEEE 40th International Conference on Data Engineering \(ICDE\),pp\. 1296–1309\.Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px2.p1.1)\.
- Lianget al\.\(2023\)P\. Liang, R\. Bommasani, T\. Lee,et al\.Holistic evaluation of language models\.External Links:2211\.09110,[Link](https://arxiv.org/abs/2211.09110)Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px3.p1.1)\.
- Lianget al\.\(2024\)T\. Liang, Z\. He, W\. Jiao, X\. Wang, Y\. Wang, R\. Wang, Y\. Yang, Z\. Tu, and S\. ShiEncouraging divergent thinking in large language models through multi\-agent debate\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 17889–17904\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.992/)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1)\.
- Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3214–3252\.External Links:[Link](https://aclanthology.org/2022.acl-long.229/)Cited by:[Table 5](https://arxiv.org/html/2609.38816#A6.T5.2.3.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px3.p1.1)\.
- Mallenet al\.\(2023\)A\. Mallen, A\. Asai, V\. Zhong, R\. Das, D\. Khashabi, and H\. HajishirziWhen not to trust language models: investigating effectiveness of parametric and non\-parametric memories\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9802–9822\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.546),[Link](https://aclanthology.org/2023.acl-long.546/)Cited by:[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.8.1)\.
- Mitchellet al\.\(2019\)M\. Mitchell, S\. Wu, A\. Zaldivar, P\. Barnes, L\. Vasserman, B\. Hutchinson, E\. Spitzer, I\. D\. Raji, and T\. GebruModel cards for model reporting\.InProceedings of the Conference on Fairness, Accountability, and Transparency \(FAT\*\),pp\. 220–229\.External Links:[Link](https://doi.org/10.1145/3287560.3287596)Cited by:[§3](https://arxiv.org/html/2609.38816#S3.SS0.SSS0.Px2.p1.1)\.
- Onget al\.\(2025\)I\. Ong, A\. Almahairi, V\. Wu, W\. Chiang, T\. Wu, J\. E\. Gonzalez, M\. W\. Kadous, and I\. StoicaRouteLLM: learning to route LLMs from preference data\.InThe Thirteenth International Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=8sSqNntaMr)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1),[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px1.p1.1)\.
- Qwen Team \(2024\)Qwen TeamQwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px2.p1.1)\.
- Reinet al\.\(2024\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGPQA: a graduate\-level google\-proof q&a benchmark\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Ti67584b98)Cited by:[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.5.1),[Table 8](https://arxiv.org/html/2609.38816#A6.T8.2.2.1)\.
- Sanchez and Hitaj \(2025\)H\. Sanchez and B\. HitajLLM chemistry estimation for multi\-llm recommendation\.External Links:2510\.03930,[Link](https://arxiv.org/abs/2510.03930)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p2.1)\.
- Sunet al\.\(2025\)M\. Sun, Y\. Yin, Z\. Xu, J\. Z\. Kolter, and Z\. LiuIdiosyncrasies in large language models\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=FCZ3jVzmTZ)Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px3.p1.1),[1st item](https://arxiv.org/html/2609.38816#S3.I1.i1.p1.1),[§3](https://arxiv.org/html/2609.38816#S3.SS0.SSS0.Px2.p1.1)\.
- Suzgunet al\.\(2023\)M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. Le, E\. Chi, D\. Zhou, and J\. WeiChallenging BIG\-bench tasks and whether chain\-of\-thought can solve them\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 13003–13051\.External Links:[Link](https://aclanthology.org/2023.findings-acl.824/)Cited by:[Table 6](https://arxiv.org/html/2609.38816#A6.T6.2.2.1),[Table 7](https://arxiv.org/html/2609.38816#A6.T7.2.3.1)\.
- Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz,et al\.Transformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations,pp\. 38–45\.External Links:[Link](https://aclanthology.org/2020.emnlp-demos.6/)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1)\.
- Yadavet al\.\(2023\)P\. Yadav, D\. Tam, L\. Choshen, C\. Raffel, and M\. BansalTIES\-merging: resolving interference when merging models\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=xtaX3WyCj1)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1),[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2026\)Z\. Yang, W\. Ding, S\. Feng, and Y\. TsvetkovAmong us: measuring and mitigating malicious contributions in model collaboration systems\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15969–15988\.External Links:[Link](https://aclanthology.org/2026.acl-long.725/)Cited by:[§6\.1](https://arxiv.org/html/2609.38816#S6.SS1.p1.1)\.
- Yuet al\.\(2024\)L\. Yu, B\. Yu, H\. Yu, F\. Huang, and Y\. LiLanguage models are super mario: absorbing abilities from homologous models as a free lunch\.InForty\-first International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=fq0NaiU8Ex)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1),[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.38816#S4.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2025\)K\. Zhang, W\. Yao, Z\. Liu, Y\. Feng, Z\. Liu, R\. R\. N, T\. Lan, L\. Li, R\. Lou, J\. Xu, B\. Pang, Y\. Zhou, S\. Heinecke, S\. Savarese, H\. Wang, and C\. XiongDiversity empowers intelligence: integrating expertise of software engineering agents\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=cKlzKs3Nnb)Cited by:[§2](https://arxiv.org/html/2609.38816#S2.SS0.SSS0.Px2.p1.1)\.
- Zhaoet al\.\(2025\)J\. Zhao, F\. M\. Plaza\-del\-Arco, B\. Genchel, and A\. C\. CurryLanguage model council: democratically benchmarking foundation models on highly subjective tasks\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 12395–12450\.External Links:[Link](https://aclanthology.org/2025.naacl-long.617/)Cited by:[§1](https://arxiv.org/html/2609.38816#S1.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.InThirty\-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=uccHPGDlao)Cited by:[§3](https://arxiv.org/html/2609.38816#S3.SS0.SSS0.Px2.p1.1)\.
## Appendix ALimitations
While principled team selection improves multi\-LLM systems, several limitations remain\.
Computational overhead\.Exact dispersion search scales combinatorially \(\(Pn\)\\binom\{P\}\{n\}\), becoming intractable for large pools or teams\. While greedy approximations and LLM\-based recruiters accelerate inference, their upfront signal\-generation costs remain high: IRT and idiosyncrasies classifiers require evaluating every candidate on full proxy datasets, andAgentic top\-kkrequires multi\-turn interviews per member\.
Decoupled selection and collaboration\.We treat downstream collaboration mechanisms as fixed black boxes using default configurations\. We do not jointly optimize the selection algorithm and the collaboration mechanism \(e\.g\., co\-training the SFT Classifier alongside DARE\-TIES scaling weights\)\.
Model scale\.We evaluate open\-weight models up to 14B parameters\. Collaboration dynamics may shift significantly with frontier\-class models \(70B\+ or closed APIs\), which might exhibit less behavioral dispersion or unlock novel synergistic benefits during refinement that smaller pools cannot capture\.
## Appendix BDesign Choice: Max\-Min vs\. Average Dispersion
#### Algorithmic Difference\.
Algorithm[1](https://arxiv.org/html/2609.38816#alg1)prioritizes max\-min dispersion: it evaluates a team by its maximum pairwise similarity and breaks ties using average pairwise similarity\. Conversely, average dispersion optimizesavgSim\\mathrm\{avgSim\}as the primary objective and usesmaxSim\\mathrm\{maxSim\}as the tie\-breaker\.
#### Performance Comparison\.
In a head\-to\-head evaluation across 16 settings \(Table[4](https://arxiv.org/html/2609.38816#A2.T4)\), 7 generated identical teams where minor score variations resulted solely from generation sampling noise\. Across the 9 settings where the objectives selected different teams, average dispersion won slightly more frequently \(5 vs\. 4\) by small margins \(≤0\.101\\leq 0\.101\)\. However, max\-min achieved the two largest performance gains\. Thus, max\-min trades negligible typical\-case margins for vital worst\-case protection against severe performance collapses in redundant candidate pools\.
#### Rationale for Max\-Min\.
We select max\-min primary dispersion because a team’s functional complementarity is bottlenecked by its most redundant pair; optimizing average similarity allows near\-duplicate models as long as other pairs are highly distinct\. Max\-min acts as a robust worst\-case safeguard against severe performance degradation, while using average similarity as a tie\-breaker preserves overall team structure\.
Algorithm 3Exact Max\-Min Diversity Search \(Algorithm[1](https://arxiv.org/html/2609.38816#alg1), repeated for comparison\)0:
sim∈ℝP×P\\mathrm\{sim\}\\in\\mathbb\{R\}^\{P\\times P\}, team size
nn, mode
∈\{most,least\}\\in\\\{\\textsc\{most\},\\textsc\{least\}\\\}
1:
T⋆←∅T^\{\\star\}\\leftarrow\\varnothing
2:foreach
T⊆\{1,…,P\}T\\subseteq\\\{1,\\dots,P\\\}with
\|T\|=n\|T\|=ndo
3:
maxSim\(T\)←maxi,j∈Tsim\(i,j\)\\mathrm\{\\textbf\{maxSim\}\}\(T\)\\leftarrow\\max\_\{i,j\\in T\}\\mathrm\{sim\}\(i,j\)
4:
avgSim\(T\)←meani,j∈Tsim\(i,j\)\\mathrm\{avgSim\}\(T\)\\leftarrow\\mathrm\{mean\}\_\{i,j\\in T\}\\,\\mathrm\{sim\}\(i,j\)
5:ifmode
=most=\\textsc\{most\}and
\(maxSim\(𝐓\),avgSim\(T\)\)\(\\mathbf\{\\mathrm\{maxSim\}\(T\)\},\\mathrm\{avgSim\}\(T\)\)lex\. smaller than incumbentthen
6:
T⋆←TT^\{\\star\}\\leftarrow T
7:elseifmode
=least=\\textsc\{least\}and
\(maxSim\(𝐓\),avgSim\(T\)\)\(\\mathbf\{\\mathrm\{maxSim\}\(T\)\},\\mathrm\{avgSim\}\(T\)\)lex\. larger than incumbentthen
8:
T⋆←TT^\{\\star\}\\leftarrow T
9:endif
10:endfor
11:return
T⋆T^\{\\star\}
Algorithm 4Exact Average\-Similarity Diversity Search0:
sim∈ℝP×P\\mathrm\{sim\}\\in\\mathbb\{R\}^\{P\\times P\}, team size
nn, mode
∈\{most,least\}\\in\\\{\\textsc\{most\},\\textsc\{least\}\\\}
1:
T⋆←∅T^\{\\star\}\\leftarrow\\varnothing
2:foreach
T⊆\{1,…,P\}T\\subseteq\\\{1,\\dots,P\\\}with
\|T\|=n\|T\|=ndo
3:
avgSim\(T\)←meani,j∈Tsim\(i,j\)\\mathrm\{\\textbf\{avgSim\}\}\(T\)\\leftarrow\\mathrm\{mean\}\_\{i,j\\in T\}\\,\\mathrm\{sim\}\(i,j\)
4:
maxSim\(T\)←maxi,j∈Tsim\(i,j\)\\mathrm\{maxSim\}\(T\)\\leftarrow\\max\_\{i,j\\in T\}\\mathrm\{sim\}\(i,j\)
5:ifmode
=most=\\textsc\{most\}and
\(avgSim\(𝐓\),maxSim\(T\)\)\(\\mathbf\{\\mathrm\{avgSim\}\(T\)\},\\mathrm\{maxSim\}\(T\)\)lex\. smaller than incumbentthen
6:
T⋆←TT^\{\\star\}\\leftarrow T
7:elseifmode
=least=\\textsc\{least\}and
\(avgSim\(𝐓\),maxSim\(T\)\)\(\\mathbf\{\\mathrm\{avgSim\}\(T\)\},\\mathrm\{maxSim\}\(T\)\)lex\. larger than incumbentthen
8:
T⋆←TT^\{\\star\}\\leftarrow T
9:endif
10:endfor
11:return
T⋆T^\{\\star\}
Table 4:Max\-min vs\. average dispersion score comparison\. Tasks are evaluated using LLM\-Blender on TruthfulQA\. Higher scores per setting are inbold\.
## Appendix CAdditional Selection Method Details
#### LLM Prompt\.
The LLM\-prompt recruiter directly selects a team from candidate model descriptions\. Given the descriptions𝒵=\{desc\(mi\)\}i=1P\\mathcal\{Z\}=\\\{\\operatorname\{desc\}\(m\_\{i\}\)\\\}\_\{i=1\}^\{P\}and the target team sizenn, we prompt an LLM recruiter with the full candidate pool and ask it to selectnnmodels that would form an effective collaborative team\. The recruiter is instructed to consider both individual capabilities and potential complementarity among team members, and returns the selected model identifiers directly\. We provide the full prompt in Appendix[H](https://arxiv.org/html/2609.38816#A8)\.
#### SFT Classifier\.
The SFT Classifier predicts the collaborative performance of a candidate team directly from its member descriptions\. Given a candidate teamT=\{mi1,…,min\}T=\\\{m\_\{i\_\{1\}\},\\ldots,m\_\{i\_\{n\}\}\\\}, we concatenate the descriptions of its members and use a learned scorerfϕf\_\{\\phi\}to produce a predicted team score,
y^T=fϕ\(desc\(mi1\),…,desc\(min\)\)\.\\hat\{y\}\_\{T\}=f\_\{\\phi\}\\big\(\\operatorname\{desc\}\(m\_\{i\_\{1\}\}\),\\ldots,\\operatorname\{desc\}\(m\_\{i\_\{n\}\}\)\\big\)\.Training targets are obtained by executing sampled teams under the collaboration methods and recording their downstream performance\. At inference time, the scorer evaluates each candidate team of sizenn, and teams are ranked by their predicted scores for selection\.
#### Agentic top\-kk\.
The agentic recruiter evaluates candidates individually through adaptive multi\-turn interviews\. For each candidatemim\_\{i\}, an LLM interviewer iteratively selects a capability axis and generates a question conditioned on the interview history\. After turns, the complete transcript is evaluated to obtain an initial interview score\. Because independently assigned scores tend to be compressed, we subsequently perform comparative re\-scoring by presenting multiple anonymized candidate transcripts jointly to the interviewer\. This produces the selection signal
𝒵=\{qi\}i=1P,\\mathcal\{Z\}=\\\{q\_\{i\}\\\}\_\{i=1\}^\{P\},whereqiq\_\{i\}denotes the calibrated interview score of candidatemim\_\{i\}\. Candidates are then ranked byqiq\_\{i\}, and the topnncandidates form the selected team\. The full interview and scoring protocol is provided in Appendix[I](https://arxiv.org/html/2609.38816#A9)\.
## Appendix DAdditional Candidate Pools Details
In our experiment, Pool 1 contains 10 Qwen2\.5\-7B models that share the same backbone but differ in their fine\-tuning corpora, providing a controlled setting in which model variation primarily reflects post\-training data\. Pool 2 contains 32 models spanning diverse source architectures, scales, training objectives, and corpora, providing a substantially more heterogeneous candidate pool\.
Table[9](https://arxiv.org/html/2609.38816#A9.T9)lists all 10 Pool 1 models; Table[10](https://arxiv.org/html/2609.38816#A9.T10)lists all 32 Pool 2 models\. For Pool 2, size is the parameter count of the*original*contributed model each checkpoint was distilled from and it is used by the*Top\-size*baseline\. Every deployed Pool\-2 checkpoint is itself a uniform Qwen2\.5\-7B distillation regardless of this value\. Two sizes \(parti\_6,parti\_21\) and two base architectures \(parti\_6,parti\_15\) were not stated in the original repository name and were confirmed against the corresponding Hugging Face model card; where a base architecture could not be confirmed from either source, we mark it “unspecified” rather than assume one\. “ID” links the deployed distilled checkpoint; “Original source” links the pre\-distillation contributed model\.
## Appendix EAdditional Data Details
We use four groups of datasets for distinct experimental purposes: main collaboration evaluation, training the SFT Classifier, constructing model performance profiles, and out\-of\-distribution \(OOD\) evaluation\. Tables[5](https://arxiv.org/html/2609.38816#A6.T5),[6](https://arxiv.org/html/2609.38816#A6.T6),[7](https://arxiv.org/html/2609.38816#A6.T7), and[8](https://arxiv.org/html/2609.38816#A6.T8)summarize the datasets used for each role\.
The main collaboration datasets are summarized in Table[5](https://arxiv.org/html/2609.38816#A6.T5)\. Each dataset is split into disjoint development and test sets, with all final collaborative performance reported on the test set\. The development set is used only to construct selection signals for methods that require task\-specific behavioral information\. Specifically,*Top solo score*ranks candidate models by their individual performance on the development set, while*Idiosyncrasies*and*IRT*construct their behavioral signals from model responses on the development set\. The selected teams are then evaluated exclusively on the held\-out test set\. None of the selectors evaluated in the OOD experiment uses GPQA\-Diamond or CoCoNot to construct its selection signal\.
The three main evaluation datasets are disjoint from the auxiliary datasets used for SFT training and performance profiling\. For training the SFT\-Recruiter, we use BBH, MMLU\-Redux, and HumanEval\. For constructing*Performance Profile*, we use a separate suite of held\-out benchmarks: ARC\-Challenge, BBH, MMLU\-Redux, GPQA\-Diamond, MATH, MedQA, PopQA, HumanEval, CoCoNot, and Human\-Interest\. For OOD evaluation, we use GPQA\-Diamond and CoCoNot\.
## Appendix FAdditional Experimental Details
### F\.1Collaboration Method Configurations
For*Prompt routing*, each query is routed according to the descriptions of the team members\.*Multi\-agent refine*performs three refinement rounds\. For*Weight merging*, we use DARE\-TIES with Qwen2\.5\-7B\-Instruct as the backbone and*average*as the merging mode\. For*LLM\-Blender*, we use Qwen2\.5\-7B\-Instruct as the backbone, set top\-kkto33, and use the first team member as the fuser\.
### F\.2Selection Method Configurations
#### Idiosyncrasies\.
We useall\-mpnet\-base\-v2with full fine\-tuning, mean pooling, and a linear classification head\. We train the classifier for 5 epochs with batch size 32, learning rate2×10−42\\times 10^\{\-4\}, AdamW, and a cosine learning\-rate schedule\. We fit the Idiosyncrasies classifier using item\-level outcomes from the development split of each main evaluation dataset\.
#### IRT Ability\.
We learn a 232\-dimensional ability embedding for each model using a 12\-expert dense MoE \(4×4\\timesthe number of datasets\) over a frozen 768\-dimensional query encoder\. The model is trained for 50 epochs with batch size 256, learning rate10−310^\{\-3\}, Adam, a cosine learning\-rate schedule, binary cross\-entropy loss, and embedding weight decay of0\.10\.1\. We fit the IRT model using item\-level success/failure outcomes from the development split of each main evaluation dataset\.
#### SFT Classifier\.
We fine\-tuneQwen2\.5\-7B\-Instructwith LoRA \(r=16r=16,α=32\\alpha=32, dropout0\.050\.05, targeting the\{q,k,v,o\}\\\{q,k,v,o\\\}projections\) and a sequence\-classification head with sigmoid output\. We train for 40 epochs with batch size 4, learning rate10−410^\{\-4\}, and MSE loss, using 6\-fold group cross\-validation by team identity\. At inference time, the scorer evaluates all\(104\)=210\\binom\{10\}\{4\}=210candidate teams in Pool 1\. For Pool 2, it evaluates all\(324\)=35,960\\binom\{32\}\{4\}=35\{,\}960teams, constituting an out\-of\-distribution transfer setting relative to Pool 1\.
#### Agentic top\-kk\.
We useQwen2\.5\-7B\-Instructas the interviewer\. Each candidate undergoes a four\-turn adaptive interview covering five axes: reasoning, code, factual knowledge, calibrated honesty, and open\-ended conversational quality\. Interviewer questions and verdicts use greedy decoding, while candidate responses are sampled with temperature0\.70\.7and top\-pp0\.90\.9\. Initial independent ratings on a 1–10 scale are followed by comparative re\-scoring using anonymized batches of up to 10 candidates\. For Pool 2, whose 32 candidates exceed the single\-batch limit, scores are averaged over three independently re\-batched comparative rounds, with residual ties resolved by a final strict\-ranking pass\. Full interview prompts and scoring details are provided in Appendix[I](https://arxiv.org/html/2609.38816#A9)\.
### F\.3Compute and Implementation
All experiments run on a mixture of NVIDIA A40, A100, and L40S GPUs with bfloat16 precision\. Collaboration methods execute through MoCo\([Feng et al\., 2026a](https://arxiv.org/html/2609.38816#bib.bib20)\), installed unmodified from its published release and used strictly as a black box\. Selection methods are implemented in PyTorch and HuggingFace Transformers, utilizing PEFT for LoRA adapters and thesentence\-transformerslibrary for embedding backbones\.
Table 5:Main evaluation datasetsTable 6:Datasets used to train the SFT Classifier\.Table 7:Datasets used to build the performance\-profile signal\.Table 8:Datasets used in the out\-of\-distribution generalization experiment\.
## Appendix GRound\-by\-Round Malicious Model Detection
Figure[5](https://arxiv.org/html/2609.38816#A7.F5)provides the detailed, round\-by\-round trajectory for every selection method as the number of malicious models in the 10\-candidate pool increases from 1 to 6\.
665544332211000\.50\.511Bad Models \(rr\)1\. Agentic Top\-K665544332211000\.50\.511Bad Models \(rr\)2\. Nested Diverse665544332211000\.50\.511Bad Models \(rr\)3\. LLM Prompt665544332211000\.50\.511Bad Models \(rr\)4\. IrtNet Diverse665544332211000\.50\.511Bad Models \(rr\)5\. Combined Diverse665544332211000\.50\.511Bad Models \(rr\)6\. Performance Profile665544332211000\.50\.511Bad Models \(rr\)7\. Idiosyncrasies665544332211000\.50\.511Bad Models \(rr\)8\. Description Diverse665544332211000\.50\.511Bad Models \(rr\)9\. SFT ClassifierFigure 5:Round\-by\-round robustness of every selection method against the theoretical*Random*baseline \(dashed gray line\)\. Theyy\-axis is the fraction of clean models in the final selected team, and thexx\-axis shows the number of malicious models in the 10\-candidate pool\. Methods that fall below the dashed line are actively selecting malicious models at a higher rate than blind chance\.
## Appendix HLLM\-Prompt Template
For the*LLM\-prompt*selection method in Family V, we query the instruction\-tuned judge to elicit a team\-level judgment in a single shot\. The exact prompt template provided to the model is shown below\. Variables such as the total pool size \(\{P\}\), the enumerated candidate list \(\{LIST\_OF\_MODELS\}\), and the target team size \(\{N\}\) are populated dynamically per experiment\.
> You are selecting models to collaborate on a task\. Below is a pool of \{P\} candidate models, each described by its fine\-tuning specialty: \{LIST\_OF\_MODELS\} Select exactly \{N\} models from the list above that would work well together as a collaborating team \-\-\- consider both individual strengths and complementary diversity between them\. Respond with ONLY a comma\-separated list of the chosen indices, and nothing else\. Do not explain your reasoning\. Do not repeat an index\. Example format for selecting 3 models: 2, 5, 7
The enumerated\{LIST\_OF\_MODELS\}is formatted as a newline\-separated list, where each line consists of the candidate’s integer index and its one\-line natural language description \(e\.g\.,0: \[Description of model 0\]\)\. To ensure the selector’s outputs are perfectly reproducible, the prompt is evaluated using greedy decoding\.
## Appendix IAgentic Interview Protocol
Here is the full protocol behind theAgentic top\-kkselector \(Section[3](https://arxiv.org/html/2609.38816#S3), Family V\), including how questions are generated and how the initial verdict is produced and corrected\.
### I\.1Turn\-by\-turn question generation
The interviewer \(Qwen2\.5\-7B\-Instruct\) conducts a fixed four\-turn interview with each candidate independently\. Each turn, it is shown the transcript so far and asked to pick one of five axes not yet covered \(step\-by\-step reasoning, code/programming ability, factual knowledge, calibrated honesty, open\-ended conversational quality\) and pose one new question probing it, under the following instruction \(reproduced verbatim, with dynamic slots in angle brackets\):
> You are interviewing a candidate AI model to decide whether to hire it onto a collaborative team\. Interview so far:⟨\\langletranscript so far⟩\\rangle\. Candidate axes to probe across the whole interview:⟨\\langleaxis list⟩\\rangle\. Axes already covered so far:⟨⋅⟩\\langle\\cdot\\rangle\. Axes NOT yet covered:⟨⋅⟩\\langle\\cdot\\rangle\. Pick ONE axis from the NOT yet covered list above and ask exactly ONE new question that probes it\. This candidate is likely a competent instruction\-tuned model, so a generic ‘‘explain concept X’’ or ‘‘describe a time you did Y’’ question will get a generically good answer from almost any such model and won’t tell us anything useful\. Instead, design the question to actively expose a gap or error if one exists: give it a subtly flawed premise to catch, a genuine edge case, a problem with a specific checkable right answer, or a request that tempts an overconfident wrong answer instead of an honest ‘‘I don’t know\.’’ Respond in EXACTLY this format, nothing else: AXIS:⟨\\langleaxis name⟩\\rangleQUESTION:⟨\\langlethe question⟩\\rangle
The interviewer generates both the axis choice and the question greedily \(temperature00, for reproducibility\); the candidate then answers as part of a genuine multi\-turn conversation \(full history carried forward each turn\), sampled at temperature0\.70\.7, top\-pp0\.90\.9, up to512512new tokens\. Axis and question are parsed from anAXIS: \.\.\. QUESTION: \.\.\.format\.
### I\.2Initial verdict
After four turns, the interviewer reads the complete transcript and issues a verdict under an explicit rubric:
> Rate this candidate’s overall quality as a hire on a scale of 1\-10\. Most instruction\-tuned models can produce fluent, well\-structured answers\. Do not reward fluency or length by itself\. Grade strictly against these bands, anchored on correctness and depth relative to what a knowledgeable human expert would say: 9\-10: Answers are technically flawless, catch subtleties or edge cases a generic competent model would miss, and any claimed uncertainty is honest and warranted\. 7\-8: Answers are correct and well\-organized but stay at a generic, textbook level\. 5\-6: At least one answer has a real inaccuracy, an unsupported claim stated with false confidence, or misses the point of the question\. 3\-4: Multiple answers are wrong, evasive, or fail to engage with what was actually asked\. 1\-2: Answers are incoherent, largely incorrect, or the candidate asserts obviously false claims confidently\.
The verdict is parsed strictly as \(SCORE: <1\-\-10\>/REASON: <\.\.\.\>\)\.
### I\.3Correcting for score compression
Independent, per\-candidate verdicts compress toward a lenient, narrow band in practice: on Pool 1, all ten candidates initially scored99or9\.59\.5out of1010despite visibly different answer quality, even after the rubric\-anchoring above\. We correct this with a*comparative*re\-scoring pass: batches of up to1010candidates’ full transcripts are shown to the interviewer side\-by\-side, anonymized as Candidate A/B/C/…, under an instruction that explicitly forbids defaulting to identical scores:
> Compare them against each other directly \-\-\- do not score each one in isolation\. \[…\] Critically: you MUST differentiate between candidates whose answers differ in correctness, depth, or honesty about uncertainty\. Giving multiple candidates the identical score is only acceptable if their answers are genuinely indistinguishable in quality \-\-\- re\-read the transcripts and look for a concrete difference \(a missed edge case, an unsupported claim, more precise reasoning\) before doing so\.
This is the score actually used byAgentic top\-kk; the initial per\-candidate verdict above is only a Phase\-1 draft\.
### I\.4Tie\-break refinement
If, after averaging, multiple candidates still land on the exact same score, a final pass forces a strict ranking within each tied group \(no two candidates may share a rank\) and spreads that group’s score upward into a fractional span according to rank, capped strictly below both the next distinct score already present in the data and the scale’s ceiling of1010\. This ensures no two final scores can collide, and no spread can push a score out of range\.
Table 9:Pool 1: 10 Qwen2\.5\-7B checkpoints, differing only in fine\-tuning corpusTable 10:Pool 2: 32 models distilled from independently contributed systems\. Base architecture is read from the original repository name where stated; “unspecified” means neither the name nor the source paper states it, and we did not assume one\.Similar Articles
Which Models Work Well Together? Measuring Heterogeneity for LLM Team Selection
The paper proposes a heterogeneity-driven framework for selecting complementary LLM teams by profiling error decorrelation and predictive divergence, then greedily optimizing a quality–complementarity objective, showing gains over quality-only top-k baselines across multiple benchmarks.
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
This paper formalizes deliberative collaboration for LLM agents under partial observability, introduces a scalable benchmark across multiple domains, and systematically evaluates representative LLMs, finding that complex tasks remain challenging while deliberation can enable error correction.
The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate
This paper formulates the collaboration tax in LLM multi-agent systems, measuring it across tasks and models to reveal predictable mechanisms and practical interventions.
Can LLM Teams Play What? Where? When?
This paper investigates whether team-based interaction improves LLM performance in the quiz game 'What? Where? When?' (ChGK). Using six recent open LLMs on a 2025 dataset of 572 questions, they show that team strategies (voting, silent captain, talkative captain) outperform single models by up to 20 percentage points, with the best team achieving 44.23% accuracy, approaching human performance.
Mixture of Complementary Agents for Robust LLM Ensemble
Proposes a framework for selecting complementary LLMs as proposers in ensemble systems, reformulating proposer selection as a combinatorial problem and exploring greedy algorithms for efficient performance-cost trade-offs.