FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

arXiv cs.LG Papers

Summary

FlexRouter 提出一种显式建模模型互补性的 LLM 路由框架,将路由建模为基于覆盖的子集选择问题,并用 Determinantal Point Processes 对模型能力与冗余进行联合建模,同时通过边缘化失败集合的目标函数直接优化答案覆盖率。在 RouterEval 基准上,该方法在域内与域外任务上以更低冗余实现了更高的覆盖率,且推理成本灵活可控。

arXiv:2609.38585v1 Announce Type: new Abstract: Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
Original Article
View Cached Full Text

Cached at: 10/02/26, 09:50 AM

# FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing
Source: [https://arxiv.org/html/2609.38585](https://arxiv.org/html/2609.38585)
Harry YangAffiliation:Independent ResearcherEmail:[hdardiry@vt\.edu](mailto:)Tiankai YangAffiliation:University of Southern CaliforniaSamyadeep BasuAffiliation:Adobe ResearchHongjie ChenAffiliation:Dolby LabsAndy ZhaoAffiliation:University of Southern CaliforniaFranck DernoncourtAffiliation:Adobe ResearchRyan A\. RossiAffiliation:Adobe ResearchHoda Eldardiry††thanks:Corresponding author\.Affiliation:Virginia Tech

###### Abstract

Existing Large Language Model \(LLM\) routing methods score LLMs independently to select top\-kkmodels\. However, this ignores model correlations and enforces a rigid computational budget\. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success\. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity\. FlexRouter optimizes foranswer coverage, maximizing the probability that at least one selected model yields a correct response\. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one\. We formulate routing as a coverage\-oriented subset selection problem and model the routing policy using Determinantal Point Processes \(DPPs\), which naturally capture both model competence and redundancy\. To directly optimize coverage without requiring a ground\-truth target subset, we introduce a training objective based on marginalizing over failure sets\. During inference, we employ a greedy strategy based on marginal log\-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget\. Extensive experiments on the large\-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in\-domain and out\-of\-domain tasks than strong baselines while maintaining flexible inference cost\.

## 1Introduction

Figure 1:Testing Success@10 score of baselines and FlexRouter on in\-domain and out\-of\-domain tasks\.Large Language Models \(LLMs\) are increasingly deployed as pools of heterogeneous models that differ in capability, specialization, and inference cost\([Chen et al\., 2023](https://arxiv.org/html/2609.38585#bib.bib24)\)\. A core problem in using LLMs systematically is*routing*: given an input query, which LLM\(s\) should be invoked to maximize answer quality under a compute budget\. Routing is closely related to earlier methods such as classic conditional computation and mixture\-of\-experts, where a gating mechanism activates a small fraction of experts per input\([Jacobs et al\., 1991](https://arxiv.org/html/2609.38585#bib.bib9);[Shazeer et al\., 2017](https://arxiv.org/html/2609.38585#bib.bib10)\)\. In modern LLM serving, routing has become especially important since no single model dominates across all tasks\([Srivatsa et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib7)\), and the candidate set can be large and rapidly evolving\([Hu et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib5);[Huang et al\., 2025a](https://arxiv.org/html/2609.38585#bib.bib4)\)\.

Question AnsweringQuery:*The car wash is only 100m away from my house, should I walk or drive?*Top\-kk•Walk…•Walk…•Walk…Ours•Walk…•Drive…•Drive…

Math reasoningQuery:*If you flip a fair coin twice, what is the expected number of heads*Top\-kk•11•11•11Ours•1 by enumerating outcomes•1 by linearity of expectation•1 by symmetry

CodingQuery:*Write a Python palindrome checker \(ignore space and case\)\.*Top\-kk•s == s\[::\-1\]•s == s\[::\-1\]•s == s\[::\-1\]Ours•normalize \+ reverse•remove spaces \+ lowercase•s == s\[::\-1\]

Figure 2:Independent top\-kkrouting often selects highly similar models, leading to redundant outputs and shared failure modes\. FlexRouter instead selects complementary models that cover different interpretations, reasoning paths, and implementations, improving answer coverage across tasks\.A common strategy in LLM routing is to assign each candidate model a predicted score and select the top\-kkmodels independently\([Zhuang et al\., 2025](https://arxiv.org/html/2609.38585#bib.bib2);[Chen et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib6)\)\. Although simple and effective, this design overlooks an important characteristic of modern LLM systems:*model correlation*\. Models trained with similar data or architectures often exhibit similar strengths and failure modes\. As a result, independent selection of the highest\-scoring models can lead to*redundant selections*that produce nearly identical outputs\. When these outputs are incorrect, the entire subset fails, limiting the probability of obtaining at least one correct answer\. As illustrated in Figure[2](https://arxiv.org/html/2609.38585#S1.F2), SOTA routing models that do not take diversity into account tend to produce multiple similar responses that fail in the same way, while our approach that selects a diverse set of models yields varied reasoning paths and increases the likelihood that at least one response is correct\. Such diversity is particularly important in real\-world systems, where multiple candidate responses are generated and a downstream verifier, reranker, or human user selects the final answer\. In these settings, the key objective is not that*all*selected models are similarly correct but that they could be diverse and*at least one*of them is correct\. This does not by itself solve final answer selection, but it provides a necessary candidate pool for downstream selection\. If all selected models fail, no downstream selector can recover the correct answer\.

Motivated by this observation, we formulate LLM routing as a*coverage\-oriented subset selection*problem\. Given an input query, our objective is to select a subset of models that maximizes the likelihood of obtaining at least one correct response while minimizing redundant selections\. To achieve this, we proposeFlexRouter, a routing framework that explicitly models both model competence and inter\-model correlations\. We parameterize the routing policy using Determinantal Point Processes \(DPPs\)\([Han et al\., 2017](https://arxiv.org/html/2609.38585#bib.bib23)\)\. DPPs naturally define a distribution over subsets, inherently favoring selections that are high\-quality and diverse\. Unlike independent scoring methods, DPPs penalize the joint selection of highly similar models, encouraging complementary model sets\.

A central challenge in this framework is learning under set\-valued objectives\([Dietterich et al\., 1997](https://arxiv.org/html/2609.38585#bib.bib36);[Cour et al\., 2011](https://arxiv.org/html/2609.38585#bib.bib37)\)\. For any given query, there is generally no unique ground\-truth subset, as any combination containing at least one successful model is acceptable\. Because standard supervised learning requires fixed targets, we must instead optimize directly for answer coverage\. To bridge this gap, we introduce a coverage\-aligned training objective designed to maximize the probability of sampling a subset with at least one correct model\. This objective is derived by marginalizing over*failure sets*, yielding a tractable objective in terms of DPP determinants\.

![Refer to caption](https://arxiv.org/html/2609.38585v1/flexrouter_overview.png)Figure 3:FlexRouter overview\. For a query, the router encoder produces embeddings that define query\-dependent competence for each LLM, while learned model embeddings capture pairwise similarity between models\. These signals define a DPP kernel whose diagonal entries favor strong models and whose off\-diagonals penalize redundant selections, thus encouraging individually strong yet mutually diverse model sets\. Greedy selection with early stopping returns an adaptive\-size subset\.At inference time, FlexRouter performs greedy subset selection using marginal log\-determinant gains\([Nemhauser et al\., 1978](https://arxiv.org/html/2609.38585#bib.bib35)\)and employs an adaptive stopping rule\. This allows the router to dynamically determine the number of models to invoke for each query, allocating more resources to difficult queries while avoiding unnecessary computation on easier ones\([Han et al\., 2021](https://arxiv.org/html/2609.38585#bib.bib34)\)\.

We evaluate FlexRouter on the large\-scaleRouterEvalbenchmark\([Huang et al\., 2025a](https://arxiv.org/html/2609.38585#bib.bib4)\), which provides extensive model–query performance records across a broad range of tasks\. FlexRouter consistently achieves higher answer coverage than strong routing baselines with lower inference cost\. Moreover, FlexRouter reduces redundancy in the selected subsets\. These results highlight the value of explicitly modeling model correlations and performing adaptive subset selection for LLM routing at scale\.

#### Contributions\.

Our main contributions are: \(i\) we formulate LLM routing as a*coverage\-oriented subset selection*problem that emphasizes the probability of obtaining at least one correct answer; \(ii\) we proposeFlexRouter, a DPP\-based subset routing framework that jointly models model quality and redundancy to select complementary and diverse model sets; \(iii\) we derive a tractable*coverage training objective*via failure\-set marginalization that directly optimizes the probability of selecting at least one correct model; \(iv\) we develop an adaptive greedy inference strategy that selects variable\-size subsets without requiring a fixed budget; \(v\) we demonstrate consistent improvements onRouterEval\([Huang et al\., 2025a](https://arxiv.org/html/2609.38585#bib.bib4)\), including higher coverage with lower redundancy under flexible inference cost\.

## 2Methodology

### 2\.1Problem Formulation

Letℳ=\{1,…,M\}\\mathcal\{M\}=\\\{1,\\dots,M\\\}be a set of available LLMs\. For a given input queryxx, the objective of the router is to select a subset of modelsS⊆ℳS\\subseteq\\mathcal\{M\}that maximizes the probability of generating a correct response while minimizing redundancy\. We assume access to a dataset of query\-response pairs for each model, where each queryxxis associated with a binary label vectory∈\{0,1\}M\{y\}\\in\\\{0,1\\\}^\{M\}\. Here,yi=1y\_\{i\}=1if modeliianswersxxcorrectly, andyi=0y\_\{i\}=0otherwise\.

For a specific queryxx, we partition the model space into aCorrect SetCxC\_\{x\}and aFailure SetFxF\_\{x\}:

Cx=\{i∈ℳ∣yi=1\},Fx=\{i∈ℳ∣yi=0\}\.C\_\{x\}=\\\{i\\in\\mathcal\{M\}\\mid y\_\{i\}=1\\\},\\quad F\_\{x\}=\\\{i\\in\\mathcal\{M\}\\mid y\_\{i\}=0\\\}\.\(1\)The routing task is successful if the selected subsetSScontains at least one model fromCxC\_\{x\}\. Any subset that intersectsCxC\_\{x\}is considered successful under this objective, as it contains at least one correct model\. We define the reward function as the indicator of coverage:

R\(x,S\)=𝕀\[S∩Cx≠∅\]=1−∏i∈S\(1−yi\)\.R\(x,S\)=\\mathbb\{I\}\[S\\cap C\_\{x\}\\neq\\emptyset\]=1\-\\prod\_\{i\\in S\}\(1\-y\_\{i\}\)\.\(2\)

### 2\.2Complementary Model Sets Parameterization

To capture both the individual capabilities of models and their pairwise correlations, we model the subset selection probability using a Determinantal Point Process \(DPP\) which defines a distribution over subsets that naturally favors high\-quality yet non\-redundant selections, aligning with the coverage\-based routing objective\. A DPP is fully parameterized by a positive semi\-definite \(PSD\)LL\-ensemble matrixLx∈ℝM×ML\_\{x\}\\in\\mathbb\{R\}^\{M\\times M\}\. The probability of selecting a subsetSSis given by:

𝒫Lx​\(S\)=det\(\(Lx\)S\)det\(I\+Lx\),\\mathcal\{P\}\_\{L\_\{x\}\}\(S\)=\\frac\{\\det\(\(L\_\{x\}\)\_\{S\}\)\}\{\\det\(I\+L\_\{x\}\)\},\(3\)where\(Lx\)S\(L\_\{x\}\)\_\{S\}denotes the submatrix ofLxL\_\{x\}indexed by the elements ofSS, andIIis the identity matrix\. To enable end\-to\-end learning, we propose a complementary parameterization for the kernelLxL\_\{x\}\. This construction ensures the matrix is inherently PSD and explicitly models the trade\-off between model performance and redundancy\.

#### Query and Model Embeddings\.

LetE⁡\(x\)=vx∈ℝdE\(x\)=\{v\}\_\{x\}\\in\\mathbb\{R\}^\{d\}be the query embedding generated by a router encoder whereE⁡\(⋅\)E\(\\cdot\)is a shared text encoder that maps the input query into a dense representation\. Each modeliiis assigned a learnable embeddingui∈ℝd\{u\}\_\{i\}\\in\\mathbb\{R\}^\{d\}representing its functional capabilities\. We compute a scalar quality scoreqi​\(x\)∈ℝ\+q\_\{i\}\(x\)\\in\\mathbb\{R\}^\{\+\}for each model, representing the confidence that modeliican answer queryxx\. This is parameterized as:

qi​\(x\)=σ⁡\(vx⊤​ui\),q\_\{i\}\(x\)=\\sigma\\left\(\{v\}\_\{x\}^\{\\top\}\{u\}\_\{i\}\\right\),\(4\)whereqi​\(x\)q\_\{i\}\(x\)captures query\-dependent model competence\.

We define a similarity matrixK∈ℝM×M\{K\}\\in\\mathbb\{R\}^\{M\\times M\}based on the cosine similarity between model embeddings, which captures the static correlation between models \(i\.e\., models with similar architectures or training data will have high similarity\):

Ki​j=ui⊤​uj‖ui‖​‖uj‖\.K\_\{ij\}=\\frac\{\{u\}\_\{i\}^\{\\top\}\{u\}\_\{j\}\}\{\\\|\{u\}\_\{i\}\\\|\\\|\{u\}\_\{j\}\\\|\}\.\(5\)

#### Kernel Construction\.

The final query\-dependent kernelLxL\_\{x\}is constructed as:

Lx=diag⁡\(q⁡\(x\)\)⋅K⋅diag⁡\(q⁡\(x\)\)L\_\{x\}=\\operatorname\{diag\}\(\{q\}\(x\)\)\\cdot\{K\}\\cdot\\operatorname\{diag\}\(\{q\}\(x\)\)\(6\)Element\-wise, this corresponds to\(Lx\)i​j=qi​\(x\)​qj​\(x\)​Ki​j\(L\_\{x\}\)\_\{ij\}=q\_\{i\}\(x\)q\_\{j\}\(x\)K\_\{ij\}\. The diagonal entries\(Lx\)i​i=qi​\(x\)2\(L\_\{x\}\)\_\{ii\}=q\_\{i\}\(x\)^\{2\}capture the individual quality of each model, while the off\-diagonal entries penalize the joint selection of correlated models\. Consequently, the determinant of a subset reflects both the quality and diversity of the selected models, as highly correlated models reduce the determinant\.

### 2\.3Learn to Maximize Coverage

The standard maximum likelihood estimate for DPPs assumes access to an observed target subset\. In our setting, the supervision does not specify a unique optimal routing decision\. Any subset that intersects the correct setCxC\_\{x\}is considered successful\. As a result, there is no single ground\-truth subset to maximize likelihood against\. Instead, we propose a Coverage Loss that directly maximizes the probability of obtaining at least one correct answer\. Furthermore, in traditional DPPs, the observed supervision is a subset valued random variableY⊆ℳY\\subseteq\\mathcal\{M\}, which enables likelihood based learning vialog⁡Pr⁡\(Y=S∗\)\\log\\Pr\(Y=S^\{\*\}\)\. In contrast, our setting observes only a binary correctness vector𝐲∈\{0,1\}M\\mathbf\{y\}\\in\\\{0,1\\\}^\{M\}, which is not a realization ofYYand therefore admits no valid DPP maximum likelihood objective\.

To efficiently compute the probability of success, we consider its complement, i\.e\., the probability that a sampled subsetYYfails to cover the query \(i\.e\.,YYis entirely contained within the failure setFxF\_\{x\}\)\. This is given by the marginal probability of the failure set \(see Appendix[B](https://arxiv.org/html/2609.38585#A2), Proposition[3](https://arxiv.org/html/2609.38585#Thmtheorem3)for detailed analysis\):

P⁡\(Y⊆Fx\)=∑S⊆Fx𝒫Lx​\(S\)=det\(I\+\(Lx\)Fx\)det\(I\+Lx\)\.P\(Y\\subseteq F\_\{x\}\)=\\sum\_\{S\\subseteq F\_\{x\}\}\\mathcal\{P\}\_\{L\_\{x\}\}\(S\)=\\frac\{\\det\(I\+\(L\_\{x\}\)\_\{F\_\{x\}\}\)\}\{\\det\(I\+L\_\{x\}\)\}\.\(7\)The probability of success \(Hit\) is thereforePsucc​\(x\)=1−P⁡\(Y⊆Fx\)P\_\{\\text\{succ\}\}\(x\)=1\-P\(Y\\subseteq F\_\{x\}\)\. We minimize the negative log\-probability of success as:

ℒhit​\(x\)=−log⁡\(1−det\(I\+\(Lx\)Fx\)det\(I\+Lx\)\)\.\\mathcal\{L\}\_\{\\text\{hit\}\}\(x\)=\-\\log\\left\(1\-\\frac\{\\det\(I\+\(L\_\{x\}\)\_\{F\_\{x\}\}\)\}\{\\det\(I\+L\_\{x\}\)\}\\right\)\.\(8\)
#### Auxiliary Supervision\.

In addition to the coverage objective, we include a binary cross\-entropy \(BCE\) loss on the predicted model\-wise correctness scoresqi​\(x\)q\_\{i\}\(x\)to provide direct supervision for individual model predictions\. This auxiliary loss stabilizes training by guiding the encoderE⁡\(⋅\)E\(\\cdot\)and the model embeddings\{ui\}\\\{u\_\{i\}\\\}to produce accurate estimates of per\-model correctness\. The overall training objective is:

ℒ=ℒhit\+λ​∑i=1MBCE⁡\(qi​\(x\),yi\),\\mathcal\{L\}=\\mathcal\{L\}\_\{\\text\{hit\}\}\+\\lambda\\sum\_\{i=1\}^\{M\}\\mathrm\{BCE\}\(q\_\{i\}\(x\),y\_\{i\}\),\(9\)whereλ\\lambdacontrols the strength of the auxiliary supervision\.

### 2\.4Flexible Router Inference

At inference time, we seek the subsetSSthat maximizes the determinant that corresponds to selecting the most probable subset under the learned DPP\. This is also known as MAP inference\([Gillenwater et al\., 2012](https://arxiv.org/html/2609.38585#bib.bib28)\)\. However, exact MAP inference is NP\-hard\([Civril and Magdon\-Ismail, 2009](https://arxiv.org/html/2609.38585#bib.bib29)\)\. We employ a greedy algorithm with an adaptive stopping condition\.

We initializeS=∅S=\\emptyset\. At each step, we select the modeli∉Si\\notin Sthat provides the maximal multiplicative gain to the determinant volume\. The marginal gaingi​\(S\)g\_\{i\}\(S\)is defined as:

gi​\(S\)=det\(\(Lx\)S∪\{i\}\)det\(\(Lx\)S\)\.g\_\{i\}\(S\)=\\frac\{\\det\(\(L\_\{x\}\)\_\{S\\cup\\\{i\\\}\}\)\}\{\\det\(\(L\_\{x\}\)\_\{S\}\)\}\.\(10\)Using the Schur complement, this can be computed efficiently as:

gi​\(S\)=\(Lx\)i​i−\(Lx\)i,S​\[\(Lx\)S\]−1​\(Lx\)S,i\.g\_\{i\}\(S\)=\(L\_\{x\}\)\_\{ii\}\-\(L\_\{x\}\)\_\{i,S\}\[\(L\_\{x\}\)\_\{S\}\]^\{\-1\}\(L\_\{x\}\)\_\{S,i\}\.\(11\)
Stopping Condition:Unlike top\-kkrouting which enforces a fixed budget, FlexRouter dynamically determines the subset size\. We stop adding models when the marginal gaingi​\(S\)g\_\{i\}\(S\)falls below a thresholdτ≥0\\tau\\geq 0relative to the initial gain\. Crucially, because the DPP log\-determinant objective is non\-monotone, adding highly correlated models can actually decrease the overall subset score\. This adaptive stopping rule naturally stops selection when candidates offer no unique contributions \(see Appendix[B](https://arxiv.org/html/2609.38585#A2), Proposition[1](https://arxiv.org/html/2609.38585#Thmtheorem1)\-[2](https://arxiv.org/html/2609.38585#Thmtheorem2)for formal proofs and intuition\), allowing the router to adaptively balance coverage and computational cost\.

## 3Experiments

We evaluate FlexRouter on large\-scale LLM routing benchmarks to answer the following questions: \(i\) Does FlexRouter improve coverage over independent\-scoring baselines? \(ii\) Does modeling complementarity reduce redundancy in the selected model sets? \(iii\) Can flexible routing reduce inference cost while maintaining high coverage? \(iv\) Does the learned routing policy generalize to unseen tasks?

### 3\.1Benchmarks and Baselines

We evaluate FlexRouter onRouterEval\([Huang et al\., 2025b](https://arxiv.org/html/2609.38585#bib.bib3)\), a large\-scale LLM routing benchmark that provides binary correctness labels for each query–model pair\. We consider two settings with shared candidate pools: amedium\-poolsetting with 3811 candidate LLMs and alarge\-poolsetting with 5000 candidate LLMs\.

Datasets\.For the medium\-pool setting, we train and evaluate in\-domain on BBH\([Suzgun et al\., 2022](https://arxiv.org/html/2609.38585#bib.bib21)\), MATH\([Hendrycks et al\., 2021b](https://arxiv.org/html/2609.38585#bib.bib20)\), and GPQA\([Rein et al\., 2023](https://arxiv.org/html/2609.38585#bib.bib19)\), and test out\-of\-domain generalization on IFEval\([Zhou et al\., 2023](https://arxiv.org/html/2609.38585#bib.bib18)\)and MuSR\([Sprague et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib17)\)\. For the large\-pool setting, the in\-domain tasks are MMLU\([Hendrycks et al\., 2021a](https://arxiv.org/html/2609.38585#bib.bib12)\), HellaSwag\([Zellers et al\., 2019](https://arxiv.org/html/2609.38585#bib.bib15)\), GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.38585#bib.bib11)\), and ARC\([Clark et al\., 2018](https://arxiv.org/html/2609.38585#bib.bib13)\), while TruthfulQA\([Lin et al\., 2022](https://arxiv.org/html/2609.38585#bib.bib16)\)and WinoGrande\([Sakaguchi et al\., 2021](https://arxiv.org/html/2609.38585#bib.bib14)\)are used for out\-of\-domain evaluation\. For in\-domain evaluation, we report results on held\-out 20% test splits of the training tasks\. For out\-of\-domain evaluation, the router is trained only on the in\-domain tasks and evaluated on unseen tasks\. Details are in Table[7](https://arxiv.org/html/2609.38585#A3.T7)and Appendix[C\.1](https://arxiv.org/html/2609.38585#A3.SS1)\.

Baselines\.We compare FlexRouter with strong routing baselines\.Independent scoringranks models by predicted correctness and selects the top\-kkcandidates independently, without modeling inter\-model correlations\.EmbedLLM\([Zhuang et al\., 2025](https://arxiv.org/html/2609.38585#bib.bib2)\)learns query\-dependent model scores in a shared embedding space, but still performs independent top\-kkselection\.BaRP\([Wei et al\., 2025](https://arxiv.org/html/2609.38585#bib.bib1)\)formulates routing as a multi\-objective contextual bandit and learns an adaptive policy under bandit feedback\. We also report theRef\. score, which is the performance of a representative strong single model on each benchmark and serves as a single\-model reference\. Additional baseline details are deferred to Appendix[C\.2](https://arxiv.org/html/2609.38585#A3.SS2)\. For fixed\-budget comparisons, all routing methods are evaluated with the same maximum selection budget\.

We also evaluate additional coverage\- and diversity\-oriented baselines, including Random\-k, MaxDiversity, and MMR\. We defer their definitions and results to Appendix[D\.2](https://arxiv.org/html/2609.38585#A4.SS2)\.

### 3\.2Evaluation Metrics

Let𝒬\\mathcal\{Q\}denote the set of queries\. We evaluate routing quality usingSuccess@kk, the fraction of queries for which at least one selected model is correct:

Success@k=1\|𝒬\|∑q∈𝒬\[∃m∈𝒮q:yq,m=1\]\.\\text\{Success@\}k\\;=\\;\\frac\{1\}\{\|\\mathcal\{Q\}\|\}\\sum\_\{q\\in\\mathcal\{Q\}\}\\mathbf\{1\}\\\!\\left\[\\,\\exists\\,m\\in\\mathcal\{S\}\_\{q\}:y\_\{q,m\}=1\\right\]\.\(12\)
To measure redundancy, we reportILD@kk\(Intra\-List Diversity\), defined as the average pairwise cosine distance among the selected models in a fixed pretrained embedding space:

ILD@​k=1\|𝒬\|​∑q∈𝒬1\(\|𝒮q\|2\)​∑i,j∈𝒮qi<j\(1−𝐮i⊤​𝐮j‖𝐮i‖​‖𝐮j‖\)\.\\text\{ILD@\}k\\;=\\;\\frac\{1\}\{\|\\mathcal\{Q\}\|\}\\sum\_\{q\\in\\mathcal\{Q\}\}\\frac\{1\}\{\\binom\{\|\\mathcal\{S\}\_\{q\}\|\}\{2\}\}\\sum\_\{\\begin\{subarray\}\{c\}i,j\\,\\in\\,\\mathcal\{S\}\_\{q\}\\\\ i<j\\end\{subarray\}\}\\left\(1\-\\frac\{\\mathbf\{u\}\_\{i\}^\{\\top\}\\mathbf\{u\}\_\{j\}\}\{\\\|\\mathbf\{u\}\_\{i\}\\\|\\,\\\|\\mathbf\{u\}\_\{j\}\\\|\}\\right\)\.\(13\)Higher ILD indicates less redundant and more complementary selections\. Further metric details are provided in Appendix[C\.3](https://arxiv.org/html/2609.38585#A3.SS3)\.

### 3\.3Implementation Details

FlexRouter uses a shared query encoder to map each query into a representation, and learnable model embeddings to represent candidate models\. Queries are encoded using RoBERTa\-base and projected into a shared 128\-dimensional space, which is used to construct the query\-dependent DPP kernel\. The model is trained with Adam for up to 100 epochs\. The training objective combines the coverage loss with an auxiliary binary cross\-entropy loss with weightλ=1\.0\\lambda=1\.0\. At inference time, model selection is performed using a greedy MAP procedure based on marginal log\-determinant gains, with adaptive stopping and a maximum subset size ofk=10k=10\. All experiments are conducted on NVIDIA A100 80GB GPUs\. Details and computational complexity are provided in Appendix[C\.4](https://arxiv.org/html/2609.38585#A3.SS4)\.

Downstream Answer Selection\.FlexRouter focuses on candidate\-pool construction rather than final answer selection\. Its objective is to select a complementary subset of models such that at least one selected model is likely to produce a correct response for the given task\.

Therefore, Success@k should be interpreted as a routing\-stage coverage metric, not as final deployed accuracy\. In practice, the routed candidates can be passed to a task\-specific verifier, reward model, LLM\-as\-judge reranker, self\-consistency or aggregation method, or human user\. However, these downstream mechanisms are not guaranteed to recover a lone correct response in every setting\. A full end\-to\-end evaluation with a concrete selector is therefore complementary to our work and remains an important future direction\.

### 3\.4Routing Performance

Table 1:Success@10 on themedium\-pool setting\(3811 models\)\. Ref\. score is the reference average accuracy for each task of a representative LLM, such as GPT\-4\. Best results arebolded; second\-best areunderlined\.Table 2:Success@10 on thelarge\-pool setting\(5000 models\)\. Best results arebolded; second\-best areunderlined\.Tables[1](https://arxiv.org/html/2609.38585#S3.T1)and[2](https://arxiv.org/html/2609.38585#S3.T2)report routing performance measured by Success@10 on the two RouterEval settings\. Across both settings, FlexRouter achieves the highest average success rate, demonstrating the benefit of selecting complementary model subsets\.

Medium\-pool setting results\.The medium\-pool setting \(Table[1](https://arxiv.org/html/2609.38585#S3.T1)\) contains fewer candidate models and several challenging reasoning benchmarks\. FlexRouter achieves the best overall performance with an average Success@10 of0\.8632\. The improvement is most pronounced on GPQA, where FlexRouter substantially outperforms both baselines\. The method also achieves the highest performance on both out\-of\-domain tasks, suggesting that the learned routing policy generalizes well to unseen tasks\.

Large\-pool setting results\.On the large\-pool setting \(Table[2](https://arxiv.org/html/2609.38585#S3.T2)\), FlexRouter achieves the best overall performance with an average Success@10 of0\.9914\. The method obtains the highest score on five of the six tasks and performs particularly well on the out\-of\-domain benchmarks TruthfulQA and WinoGrande\. These results indicate that explicitly modeling correlations between models helps identify complementary candidates and increases the likelihood that at least one selected model produces a correct answer\.

EmbedLLM achieves the highest score on MATH\. This behavior likely reflects that MATH contains a small number of specialist models with strong performance, where selecting the highest\-scoring models independently can already perform well\. Nevertheless, FlexRouter consistently achieves the best average performance across tasks, indicating that modeling model complementarity provides a robust advantage across diverse benchmarks\.

Additional coverage analysis\.To further analyze the behavior across smaller budgets and the composition of selected pools, Appendix[D\.1](https://arxiv.org/html/2609.38585#A4.SS1)reports Success@k curves, Avg\-Correct@10, and Zero\-Correct Rate\. These analyses show that FlexRouter’s advantage appears in the multi\-model setting and that it reduces all\-wrong selected pools\.

### 3\.5Subset Diversity

Table 3:ILD@10 in the fixed pre\-trained embedding space on themedium\-pool setting\. Higher is better\. Best results arebolded\.Table 4:ILD@10 in the fixed pre\-trained embedding space on thelarge\-pool setting\. Higher is better\. Best results arebolded\.We next examine whether the selected model subsets exhibit higher diversity\. Tables[3](https://arxiv.org/html/2609.38585#S3.T3)and[4](https://arxiv.org/html/2609.38585#S3.T4)report ILD@10, which measures the average pairwise cosine distance between embeddings of the selected models\. Higher values indicate that the router selects models that are less redundant and more complementary\.

Across both settings, FlexRouter consistently produces the most diverse model subsets\. On thelarge\-pool setting\(Table[4](https://arxiv.org/html/2609.38585#S3.T4)\), FlexRouter achieves the highest average ILD@10 and obtains the best result on five of the six tasks\. In particular, the diversity gains are most pronounced on ARC, TruthfulQA, and WinoGrande, where the selected models are substantially more dissimilar than those chosen by the baselines\. Although BaRP produces relatively high diversity score on MMLU, its selections are nearly constant across tasks, suggesting that it relies on a fixed set of models rather than adapting to query\-specific complementarity\.

The difference is even more pronounced on themedium\-pool setting\(Table[3](https://arxiv.org/html/2609.38585#S3.T3)\), where FlexRouter achieves substantially higher ILD@10 on every task\. In contrast, EmbedLLM and BaRP produce relatively low diversity scores, indicating that their routing strategies tend to select similar models\.

These results provide empirical evidence that FlexRouter effectively captures model complementarity during routing\. By explicitly modeling correlations between models, the router selects subsets that are both high\-quality and diverse, which directly supports the coverage improvements observed in Section[3\.4](https://arxiv.org/html/2609.38585#S3.SS4)\.

### 3\.6Generalization to Unseen Tasks

We next examine whether the learned routing policy generalizes to tasks that are not observed during training\. Across both RouterEval settings, FlexRouter consistently achieves the best Success@10 on the out\-of\-domain benchmarks according to Tables[1](https://arxiv.org/html/2609.38585#S3.T1)and[2](https://arxiv.org/html/2609.38585#S3.T2), indicating that the routing strategy learned on the in\-domain tasks transfers effectively to unseen evaluation settings\. The diversity analysis further supports this observation\. As shown in Tables[3](https://arxiv.org/html/2609.38585#S3.T3)and[4](https://arxiv.org/html/2609.38585#S3.T4), FlexRouter selects substantially more diverse model subsets on the out\-of\-domain tasks compared with both baselines\. This behavior suggests that the router captures general patterns of model complementarity rather than memorizing task\-specific preferences\. Together, these results indicate that modeling correlations between models enables FlexRouter to construct complementary model subsets that remain effective even when routing queries from previously unseen tasks\.

### 3\.7Ablation on Kernel Design

The DPP kernel requires a similarity measure between model embeddings\. Our default design usescosine similarity, while an alternative is aGaussian RBF kernel\(as detailed in Appendix[A](https://arxiv.org/html/2609.38585#A1)\)Si​j=exp⁡\(−γ​‖𝐮i−𝐮j‖2\)S\_\{ij\}=\\exp\(\-\\gamma\\\|\\mathbf\{u\}\_\{i\}\-\\mathbf\{u\}\_\{j\}\\\|^\{2\}\)with a trainable bandwidthγ\\gamma\. Table[5](https://arxiv.org/html/2609.38585#S3.T5)reports results on both RouterEval settings atk=10k=10\. On the large\-pool setting, the RBF kernel achieves slightly higher in\-domain Success@10, while cosine performs better on the out\-of\-domain tasks and consistently produces more diverse model subsets\. On the medium\-pool setting, cosine achieves higher in\-domain success and again selects more diverse subsets, while RBF attains slightly higher out\-of\-domain success\. Overall, cosine similarity consistently yields higher diversity while maintaining competitive success rates across both settings\. We therefore adopt cosine similarity as the default kernel for FlexRouter\.

Table 5:Kernel design ablation atk=10k=10\. ILD is measured in the fixed embedding space\.
### 3\.8Adaptive Subset Size via Stopping Thresholdτ\\tau

FlexRouter allows adaptive control of inference cost through a stopping thresholdτ\\tauin the greedy MAP procedure\. A model is added only if its marginal log\-determinant gain exceedsτ\\tautimes the first gain, allowing the router to stop early when additional models provide limited benefit\. Figure[4](https://arxiv.org/html/2609.38585#S3.F4)illustrates the resulting coverage–cost trade\-off on the combined in\-domain test set of the large\-pool setting \(kmax=10k\_\{\\max\}=10\)\. Whenτ\\tauis small, the router selects nearly all candidate models and achieves the highest coverage\. Asτ\\tauincreases, the average subset size decreases while Success@10 gradually declines\. A favorable operating point appears aroundτ≈0\.2\\tau\\approx 0\.2, where the router reduces the average subset size substantially while maintaining high coverage\. Beyond this point, further reductions in subset size lead to a sharper drop in coverage\. Exact numerical results are reported in Table[6](https://arxiv.org/html/2609.38585#A3.T6)\.

Figure 4:Success@10 and Average Subset Size over Stopping Thresholdτ\\tau\. Whenτ≈0\.2\\tau\\approx 0\.2, the router reduces the average subset size substantially while maintaining high coverage\.These results demonstrate that the stopping threshold provides a simple mechanism for controlling the coverage–cost trade\-off\. Importantly, the subset size varies across queries, indicating that FlexRouter adapts its inference budget to query difficulty rather than using a fixed number of models\.

## 4Limitations and Scope

FlexRouter focuses on routing\-stage candidate\-pool construction rather than final answer selection\. Success@k measures whether the selected subset contains at least one correct model, so it should not be interpreted as final deployed accuracy\. The selected candidates may be passed to a verifier, reranker, aggregation method, or human user, but these downstream mechanisms are not guaranteed to recover a correct answer in every case\. Jointly evaluating FlexRouter with concrete downstream selectors is an important future direction\.

Our diversity analysis measures model\-level diversity using fixed model embeddings, which does not necessarily imply output\-level semantic diversity for every query\. Since RouterEval provides correctness labels but not full generated responses for all query–model pairs, response\-level diversity and answer\-agreement analysis are left for future work\.

Our cost analysis uses the number of invoked models, or average subset size, as a proxy for inference cost\. Real deployment cost also depends on model\-specific latency, token price, output length, batching, and hardware constraints\. Extending FlexRouter to optimize such model\-specific costs is a useful future direction\.

## 5Conclusion

We presentedFlexRouter, a coverage\-oriented routing framework that selects complementary subsets of LLMs under flexible inference budgets\. By modeling routing as a subset selection problem with determinantal point processes, FlexRouter explicitly accounts for model correlations and reduces redundant model selections\. A coverage\-aligned training objective enables learning without requiring a ground\-truth target subset while an adaptive greedy inference procedure determines subset sizes on a per\-query basis without enforcing a fixed budget\. Experiments and ablation studies on theRouterEvalbenchmark demonstrate that FlexRouter achieves higher coverage than strong baselines while maintaining flexible inference cost and lower redundancy, highlighting the benefits of correlation\-aware and adaptive subset selection for LLM routing at scale\.

## Ethics Statement

This work studies routing among existing LLMs and does not introduce new models or training data\. As such, it inherits potential biases and risks present in the underlying models\. By promoting diverse model selection, our method may surface a broader range of outputs, including both useful alternatives and undesirable content\. In practice, it should be combined with downstream safeguards such as verification, filtering, or human oversight\. Our approach may improve efficiency by reducing redundant model usage, but adaptive routing can also increase cost for difficult queries if not properly controlled\.

## References

- Affandiet al\.\(2014\)R\. H\. Affandi, E\. Fox, R\. Adams, and B\. TaskarLearning the parameters of determinantal point process kernels\.InInternational Conference on Machine Learning,pp\. 1224–1232\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px3.p1.1)\.
- Chenet al\.\(2023\)L\. Chen, M\. Zaharia, and J\. ZouFrugalgpt: how to use large language models while reducing cost and improving performance\.arXiv preprint arXiv:2305\.05176\.Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p1.1)\.
- Chenet al\.\(2024\)S\. Chen, W\. Jiang, B\. Lin, J\. T\. Kwok, and Y\. ZhangRouterDC: query\-based router by dual contrastive learning for assembling large language models\.InAdvances in Neural Information Processing Systems,Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38585#S1.p2.1)\.
- Choet al\.\(2019\)S\. Cho, C\. Li, D\. Yu, H\. Foroosh, and F\. LiuMulti\-document summarization with determinantal point processes and contextualized representations\.InProceedings of the 2nd Workshop on New Frontiers in Summarization,pp\. 98–103\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px3.p1.1)\.
- Civril and Magdon\-Ismail \(2009\)A\. Civril and M\. Magdon\-IsmailOn selecting a maximum volume sub\-matrix of a matrix and related problems\.Theoretical Computer Science410\(47\-49\),pp\. 4801–4811\.Cited by:[§2\.4](https://arxiv.org/html/2609.38585#S2.SS4.p1.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try arc, the ai2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p4.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p4.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Couret al\.\(2011\)T\. Cour, B\. Sapp, and B\. TaskarLearning from partial labels\.The Journal of Machine Learning Research12,pp\. 1501–1536\.Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p4.1)\.
- Denget al\.\(2020\)Q\. Deng, K\. Wang, M\. Zhao, Z\. Zou, R\. Wu, J\. Tao, C\. Fan, and L\. ChenPersonalized bundle recommendation in online games\.InProceedings of the 29th ACM international conference on information & knowledge management,pp\. 2381–2388\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px3.p1.1)\.
- Dietterichet al\.\(1997\)T\. G\. Dietterich, R\. H\. Lathrop, and T\. Lozano\-PérezSolving the multiple instance problem with axis\-parallel rectangles\.Artificial intelligence89\(1\-2\),pp\. 31–71\.Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p4.1)\.
- Gillenwateret al\.\(2012\)J\. Gillenwater, A\. Kulesza, and B\. TaskarNear\-optimal map inference for determinantal point processes\.Advances in Neural Information Processing Systems25\.Cited by:[§2\.4](https://arxiv.org/html/2609.38585#S2.SS4.p1.1)\.
- Hanet al\.\(2017\)I\. Han, P\. Kambadur, K\. Park, and J\. ShinFaster greedy map inference for determinantal point processes\.InInternational Conference on Machine Learning,pp\. 1384–1393\.Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p3.1)\.
- Hanet al\.\(2021\)Y\. Han, G\. Huang, S\. Song, L\. Yang, H\. Wang, and Y\. WangDynamic neural networks: a survey\.IEEE transactions on pattern analysis and machine intelligence44\(11\),pp\. 7436–7456\.Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p5.1)\.
- Hendryckset al\.\(2021a\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p4.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Hendryckset al\.\(2021b\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the math dataset\.External Links:2103\.03874,[Link](https://arxiv.org/abs/2103.03874)Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Huet al\.\(2024\)Q\. J\. Hu, J\. Bieker, X\. Li, N\. Jiang, B\. Keigwin, G\. Ranganath, K\. Keutzer, and S\. K\. UpadhyayRouterBench: a benchmark for multi\-LLM routing system\.arXiv preprint arXiv:2403\.12031\.Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p1.1)\.
- Huanget al\.\(2025a\)Z\. Huang, G\. Ling, V\. S\. Liang, Y\. Lin, Y\. Chen, S\. Zhong, H\. Wu, and L\. LinRouterEval: a comprehensive benchmark for routing LLMs to explore model\-level scaling up in LLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Cited by:[§1](https://arxiv.org/html/2609.38585#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38585#S1.p1.1),[§1](https://arxiv.org/html/2609.38585#S1.p6.1)\.
- Huanget al\.\(2025b\)Z\. Huang, G\. Ling, Y\. Lin, Y\. Chen, S\. Zhong, H\. Wu, and L\. LinRouterEval: a comprehensive benchmark for routing LLMs to explore model\-level scaling up in LLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 3860–3887\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.208/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.208),ISBN 979\-8\-89176\-335\-7Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p1.1)\.
- Jacobset al\.\(1991\)R\. A\. Jacobs, M\. I\. Jordan, S\. J\. Nowlan, and G\. E\. HintonAdaptive mixtures of local experts\.Neural Computation3\(1\),pp\. 79–87\.External Links:[Document](https://dx.doi.org/10.1162/neco.1991.3.1.79)Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p1.1)\.
- Jianget al\.\(2023\)D\. Jiang, X\. Ren, and B\. Y\. LinLlm\-blender: ensembling large language models with pairwise ranking and generative fusion\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14165–14178\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px2.p1.1)\.
- Kulesza and Taskar \(2012\)A\. Kulesza and B\. TaskarDeterminantal point processes for machine learning\.Foundations and Trends® in Machine Learning\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px3.p1.1)\.
- Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulQA: measuring how models mimic human falsehoods\.External Links:2109\.07958,[Link](https://arxiv.org/abs/2109.07958)Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p4.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Nemhauseret al\.\(1978\)G\. L\. Nemhauser, L\. A\. Wolsey, and M\. L\. FisherAn analysis of approximations for maximizing submodular set functions—i\.Mathematical programming14\(1\),pp\. 265–294\.Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p5.1)\.
- OpenAIet al\.\(2024\)OpenAI, J\. Achiam, S\. Adler, S\. Agarwal,et al\.GPT\-4 technical report\.External Links:2303\.08774,[Link](https://arxiv.org/abs/2303.08774)Cited by:[§C\.2](https://arxiv.org/html/2609.38585#A3.SS2.p4.1)\.
- Reinet al\.\(2023\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGPQA: a graduate\-level google\-proof q&a benchmark\.External Links:2311\.12022,[Link](https://arxiv.org/abs/2311.12022)Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Sakaguchiet al\.\(2021\)K\. Sakaguchi, R\. Le Bras, C\. Bhagavatula, and Y\. ChoiWinoGrande: an adversarial winograd schema challenge at scale\.Communications of the ACM\.Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p4.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Shazeeret al\.\(2017\)N\. Shazeer, A\. Mirhoseini, K\. Maziarz, A\. Davis, Q\. Le, G\. Hinton, and J\. DeanOutrageously large neural networks: the sparsely\-gated mixture\-of\-experts layer\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/1701.06538)Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p1.1)\.
- Spragueet al\.\(2024\)Z\. Sprague, X\. Ye, K\. Bostrom, S\. Chaudhuri, and G\. DurrettMuSR: testing the limits of chain\-of\-thought with multistep soft reasoning\.External Links:2310\.16049,[Link](https://arxiv.org/abs/2310.16049)Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Srivatsaet al\.\(2024\)K\. A\. Srivatsa, K\. K\. Maurya, and E\. KochmarHarnessing the power of multiple minds: lessons learned from LLM routing\.InProceedings of the Fifth Workshop on Insights from Negative Results in NLP,External Links:[Document](https://dx.doi.org/10.18653/v1/2024.insights-1.15)Cited by:[§1](https://arxiv.org/html/2609.38585#S1.p1.1)\.
- Suzgunet al\.\(2022\)M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. V\. Le, E\. H\. Chi, D\. Zhou, and J\. WeiChallenging big\-bench tasks and whether chain\-of\-thought can solve them\.arXiv preprint arXiv:2210\.09261\.Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Wanget al\.\(2024\)J\. Wang, J\. Wang, B\. Athiwaratkun, C\. Zhang, and J\. ZouMixture\-of\-agents enhances large language model capabilities\.arXiv preprint arXiv:2406\.04692\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2022\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.arXiv preprint arXiv:2203\.11171\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px2.p1.1)\.
- Weiet al\.\(2025\)W\. Wei, T\. Yang, H\. Chen, Y\. Zhao, F\. Dernoncourt, R\. A\. Rossi, and H\. EldardiryLearning to route llms from bandit feedback: one policy, many trade\-offs\.External Links:2510\.07429,[Link](https://arxiv.org/abs/2510.07429)Cited by:[§C\.2](https://arxiv.org/html/2609.38585#A3.SS2.p3.1),[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p3.1)\.
- Wilhelmet al\.\(2018\)M\. Wilhelm, A\. Ramanathan, A\. Bonomo, S\. Jain, E\. H\. Chi, and J\. GillenwaterPractical diversified recommendations on youtube with determinantal point processes\.InProceedings of the 27th ACM International Conference on Information and Knowledge Management,pp\. 2165–2173\.Cited by:[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px3.p1.1)\.
- Zellerset al\.\(2019\)R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. ChoiHellaswag: can a machine really finish your sentence?\.arXiv preprint arXiv:1905\.07830\.Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p4.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Zhouet al\.\(2023\)J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. HouInstruction\-following evaluation for large language models\.External Links:2311\.07911,[Link](https://arxiv.org/abs/2311.07911)Cited by:[§C\.1](https://arxiv.org/html/2609.38585#A3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p2.1)\.
- Zhuanget al\.\(2025\)R\. Zhuang, T\. Wu, Z\. Wen, A\. Li, J\. Jiao, and K\. RamchandranEmbedLLM: learning compact representations of large language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Fs9EabmQrJ)Cited by:[§C\.2](https://arxiv.org/html/2609.38585#A3.SS2.p2.1),[Appendix E](https://arxiv.org/html/2609.38585#A5.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38585#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.38585#S3.SS1.p3.1)\.

## Appendix

## Appendix AAlternative Similarity Modeling

#### Gaussian\.

We model the correlation between models using a Gaussian Radial Basis Function \(RBF\) kernel\. This captures the intuition that models with similar embeddings \(close in the functional space\) are redundant\. We define the similarity matrixK∈ℝM×M\{K\}\\in\\mathbb\{R\}^\{M\\times M\}as:

Ki​j=exp⁡\(−γ​‖ui−uj‖22\),K\_\{ij\}=\\exp\\left\(\-\\gamma\\\|\{u\}\_\{i\}\-\{u\}\_\{j\}\\\|\_\{2\}^\{2\}\\right\),\(1\)whereui,uj\{u\}\_\{i\},\{u\}\_\{j\}are the learnable model embeddings andγ\>0\\gamma\>0is a trainable scaling parameter \(inverse bandwidth\)\. This formulation guarantees thatK\{K\}is positive semi\-definite and bounded between\(0,1\]\(0,1\]\.

## Appendix BTheoretical Results

###### Proposition 1\(Submodularity of Log\-Determinant\)\.

LetLx≻0L\_\{x\}\\succ 0be a positive definite kernel matrix overℳ\\mathcal\{M\}\. Definef\(S\)=logdet\(\(Lx\)S\)f\(S\)=\\log\\det\(\(L\_\{x\}\)\_\{S\}\)forS⊆ℳS\\subseteq\\mathcal\{M\}, where\(Lx\)S\(L\_\{x\}\)\_\{S\}denotes the principal submatrix ofLxL\_\{x\}indexed bySS, andf⁡\(∅\)=0f\(\\emptyset\)=0by convention\. Thenffis submodular: for allS⊆T⊆ℳS\\subseteq T\\subseteq\\mathcal\{M\}andi∉Ti\\notin T,

f⁡\(S∪\{i\}\)−f⁡\(S\)≥f⁡\(T∪\{i\}\)−f⁡\(T\)\.f\(S\\cup\\\{i\\\}\)\-f\(S\)\\geq f\(T\\cup\\\{i\\\}\)\-f\(T\)\.\(2\)However,ffis*not*monotone in general\. The marginal gain of adding modeliito setSSis

f⁡\(S∪\{i\}\)−f⁡\(S\)=log⁡\(Lx\)i​i⋅S,f\(S\\cup\\\{i\\\}\)\-f\(S\)=\\log\(L\_\{x\}\)\_\{ii\\cdot S\},\(3\)where\(Lx\)i​i⋅S=\(Lx\)i​i−\(Lx\)i,S​\[\(Lx\)S\]−1​\(Lx\)S,i\(L\_\{x\}\)\_\{ii\\cdot S\}=\(L\_\{x\}\)\_\{ii\}\-\(L\_\{x\}\)\_\{i,S\}\[\(L\_\{x\}\)\_\{S\}\]^\{\-1\}\(L\_\{x\}\)\_\{S,i\}is the Schur complement\. This quantity is positive \(sinceLx≻0L\_\{x\}\\succ 0\) but may be less than one, in which caselog⁡\(Lx\)i​i⋅S<0\\log\(L\_\{x\}\)\_\{ii\\cdot S\}<0andffdecreases\.

###### Proof\.

For anyS⊆ℳS\\subseteq\\mathcal\{M\}andi∉Si\\notin S, the determinant formula for bordered matrices gives

det\(\(Lx\)S∪\{i\}\)det\(\(Lx\)S\)=\(Lx\)i​i−\(Lx\)i,S​\[\(Lx\)S\]−1​\(Lx\)S,i=\(Lx\)i​i⋅S\.\\frac\{\\det\(\(L\_\{x\}\)\_\{S\\cup\\\{i\\\}\}\)\}\{\\det\(\(L\_\{x\}\)\_\{S\}\)\}=\(L\_\{x\}\)\_\{ii\}\-\(L\_\{x\}\)\_\{i,S\}\[\(L\_\{x\}\)\_\{S\}\]^\{\-1\}\(L\_\{x\}\)\_\{S,i\}=\(L\_\{x\}\)\_\{ii\\cdot S\}\.\(4\)ForS⊆TS\\subseteq T, the Schur complement satisfies\(Lx\)i​i⋅T≤\(Lx\)i​i⋅S\(L\_\{x\}\)\_\{ii\\cdot T\}\\leq\(L\_\{x\}\)\_\{ii\\cdot S\}, since conditioning on a larger set can only reduce the residual variance in a positive semi\-definite matrix\. Taking logarithms preserves the inequality, establishing submodularity\.

SinceLx≻0L\_\{x\}\\succ 0, all Schur complements are strictly positive, sof⁡\(S∪\{i\}\)−f⁡\(S\)=log⁡\(Lx\)i​i⋅S\>−∞f\(S\\cup\\\{i\\\}\)\-f\(S\)=\\log\(L\_\{x\}\)\_\{ii\\cdot S\}\>\-\\infty\. However, nothing prevents\(Lx\)i​i⋅S<1\(L\_\{x\}\)\_\{ii\\cdot S\}<1\. ∎

The non\-monotonicity off\(S\)=logdet\(\(Lx\)S\)f\(S\)=\\log\\det\(\(L\_\{x\}\)\_\{S\}\)has a natural interpretation in the routing context: adding a model that is highly correlated with the existing selection can reduce the overall determinant, reflecting the fact that redundant coverage does not improve the DPP objective\.

###### Proposition 2\(Greedy Determinant Maximality with Adaptive Stopping\)\.

LetLx≻0L\_\{x\}\\succ 0and letf\(S\)=logdet\(\(Lx\)S\)f\(S\)=\\log\\det\(\(L\_\{x\}\)\_\{S\}\)as above\. Consider the greedy algorithm that initializesS=∅S=\\emptysetand iteratively selects

i∗=arg⁡maxi∉S​gi​\(S\)=arg⁡maxi∉S​\(Lx\)i​i⋅S,i^\{\*\}=\\arg\\max\_\{i\\notin S\}\\;g\_\{i\}\(S\)=\\arg\\max\_\{i\\notin S\}\\;\(L\_\{x\}\)\_\{ii\\cdot S\},\(5\)terminating whengi∗​\(S\)≤1g\_\{i^\{\*\}\}\(S\)\\leq 1for the best remaining candidate\. LetSgreedyS\_\{\\mathrm\{greedy\}\}denote the output\. Then for allT⊇SgreedyT\\supseteq S\_\{\\mathrm\{greedy\}\}withT≠SgreedyT\\neq S\_\{\\mathrm\{greedy\}\},

det\(\(Lx\)T\)≤det\(\(Lx\)Sgreedy\)\.\\det\(\(L\_\{x\}\)\_\{T\}\)\\leq\\det\(\(L\_\{x\}\)\_\{S\_\{\\mathrm\{greedy\}\}\}\)\.\(6\)That is,SgreedyS\_\{\\mathrm\{greedy\}\}globally maximizesdet\(\(Lx\)S\)\\det\(\(L\_\{x\}\)\_\{S\}\)over all supersets of itself\.

###### Proof\.

When the algorithm terminates,gi​\(Sgreedy\)≤1g\_\{i\}\(S\_\{\\mathrm\{greedy\}\}\)\\leq 1for alli∉Sgreedyi\\notin S\_\{\\mathrm\{greedy\}\}, meaning\(Lx\)i​i⋅Sgreedy≤1\(L\_\{x\}\)\_\{ii\\cdot S\_\{\\mathrm\{greedy\}\}\}\\leq 1\. For any supersetT=Sgreedy∪\{i1,…,ir\}T=S\_\{\\mathrm\{greedy\}\}\\cup\\\{i\_\{1\},\\ldots,i\_\{r\}\\\}, the determinant telescopes as

det\(\(Lx\)T\)=det\(\(Lx\)Sgreedy\)​∏j=1r\(Lx\)ij​ij⋅Sgreedy∪\{i1,…,ij−1\}\.\\det\(\(L\_\{x\}\)\_\{T\}\)=\\det\(\(L\_\{x\}\)\_\{S\_\{\\mathrm\{greedy\}\}\}\)\\prod\_\{j=1\}^\{r\}\(L\_\{x\}\)\_\{i\_\{j\}i\_\{j\}\\cdot S\_\{\\mathrm\{greedy\}\}\\cup\\\{i\_\{1\},\\ldots,i\_\{j\-1\}\\\}\}\.\(7\)By submodularity offf\(Theorem[1](https://arxiv.org/html/2609.38585#Thmtheorem1)\),\(Lx\)ij​ij⋅Sgreedy∪\{i1,…,ij−1\}≤\(Lx\)ij​ij⋅Sgreedy≤1\(L\_\{x\}\)\_\{i\_\{j\}i\_\{j\}\\cdot S\_\{\\mathrm\{greedy\}\}\\cup\\\{i\_\{1\},\\ldots,i\_\{j\-1\}\\\}\}\\leq\(L\_\{x\}\)\_\{i\_\{j\}i\_\{j\}\\cdot S\_\{\\mathrm\{greedy\}\}\}\\leq 1, so each factor in the product is at most11, givingdet\(\(Lx\)T\)≤det\(\(Lx\)Sgreedy\)\\det\(\(L\_\{x\}\)\_\{T\}\)\\leq\\det\(\(L\_\{x\}\)\_\{S\_\{\\mathrm\{greedy\}\}\}\)\. ∎

This shows that the greedy algorithm with adaptive stopping finds a set that no superset can improve upon\. In particular,SgreedyS\_\{\\mathrm\{greedy\}\}is a global maximizer ofdet\(\(Lx\)S\)\\det\(\(L\_\{x\}\)\_\{S\}\)over all setsS⊇SgreedyS\\supseteq S\_\{\\mathrm\{greedy\}\}, includingℳ\\mathcal\{M\}itself\.

###### Proposition 3\(Coverage\-Diversity Equivalence\)\.

LetCx⊆ℳC\_\{x\}\\subseteq\\mathcal\{M\}denote the correct set andFx=ℳ∖CxF\_\{x\}=\\mathcal\{M\}\\setminus C\_\{x\}the failure set\. The hit probability decomposes as

Psucc​\(x\)=1−P⁡\(Y⊆Fx\)=∑S⊆ℳS∩Cx≠∅𝒫Lx​\(S\)\.P\_\{\\mathrm\{succ\}\}\(x\)=1\-P\(Y\\subseteq F\_\{x\}\)=\\sum\_\{\\begin\{subarray\}\{c\}S\\subseteq\\mathcal\{M\}\\\\ S\\cap C\_\{x\}\\neq\\emptyset\\end\{subarray\}\}\\mathcal\{P\}\_\{L\_\{x\}\}\(S\)\.\(8\)Consequently, minimizing the coverage lossℒhit​\(x\)\\mathcal\{L\}\_\{\\mathrm\{hit\}\}\(x\)is equivalent to maximizing the total DPP probability mass assigned to subsets that contain at least one correct model\.

###### Proof\.

By the law of total probability,

1=∑S⊆ℳ𝒫Lx​\(S\)=∑S⊆ℳS∩Cx≠∅𝒫Lx​\(S\)\+∑S⊆ℳS∩Cx=∅𝒫Lx​\(S\)\.1=\\sum\_\{S\\subseteq\\mathcal\{M\}\}\\mathcal\{P\}\_\{L\_\{x\}\}\(S\)=\\sum\_\{\\begin\{subarray\}\{c\}S\\subseteq\\mathcal\{M\}\\\\ S\\cap C\_\{x\}\\neq\\emptyset\\end\{subarray\}\}\\mathcal\{P\}\_\{L\_\{x\}\}\(S\)\+\\sum\_\{\\begin\{subarray\}\{c\}S\\subseteq\\mathcal\{M\}\\\\ S\\cap C\_\{x\}=\\emptyset\\end\{subarray\}\}\\mathcal\{P\}\_\{L\_\{x\}\}\(S\)\.\(9\)The second sum equalsP⁡\(Y⊆Fx\)P\(Y\\subseteq F\_\{x\}\)sinceS∩Cx=∅S\\cap C\_\{x\}=\\emptysetif and only ifS⊆FxS\\subseteq F\_\{x\}\. Therefore,

Psucc​\(x\)=1−P⁡\(Y⊆Fx\)=∑S⊆ℳS∩Cx≠∅𝒫Lx​\(S\)\.P\_\{\\mathrm\{succ\}\}\(x\)=1\-P\(Y\\subseteq F\_\{x\}\)=\\sum\_\{\\begin\{subarray\}\{c\}S\\subseteq\\mathcal\{M\}\\\\ S\\cap C\_\{x\}\\neq\\emptyset\\end\{subarray\}\}\\mathcal\{P\}\_\{L\_\{x\}\}\(S\)\.\(10\)Sinceℒhit​\(x\)=−log⁡Psucc​\(x\)\\mathcal\{L\}\_\{\\mathrm\{hit\}\}\(x\)=\-\\log P\_\{\\mathrm\{succ\}\}\(x\)and−log\-\\logis monotone decreasing, minimizingℒhit​\(x\)\\mathcal\{L\}\_\{\\mathrm\{hit\}\}\(x\)is equivalent to maximizingPsucc​\(x\)P\_\{\\mathrm\{succ\}\}\(x\)\. ∎

## Appendix CExperiment Details

Table 6:Effect of the relative stopping thresholdτ\\tauon coverage and inference cost \(combined ID test, large\-pool setting,kmax=10k\_\{\\max\}=10\)\.τ\\tauis expressed as a fraction of the first marginal log\-determinant gain\. Size Reduction and Coverage Drop are relative toτ=0\\tau=0\.### C\.1Benchmarks and Datasets

We conduct experiments onRouterEval\([Huang et al\., 2025b](https://arxiv.org/html/2609.38585#bib.bib3)\), a large\-scale benchmark designed for evaluating LLM routing methods\. RouterEval aggregates multiple widely used reasoning and knowledge benchmarks and records the performance of thousands of candidate LLMs on each prompt\. For every\(query,model\)\(\\text\{query\},\\text\{model\}\)pair, the benchmark provides a binary correctness label indicating whether the candidate model answers the query correctly\. This design enables systematic evaluation of routing algorithms under large candidate model pools\.

Our experiments cover two RouterEval settings: themedium\-poolsetting and thelarge\-poolsetting\. The medium\-pool setting contains3811candidate LLMs, while the large\-pool setting contains5000candidate LLMs\. Within each setting, all tasks share the same candidate model pool, enabling consistent comparison across routing methods\.

For themedium\-pool setting, we use three tasks for in\-domain training and evaluation:BBH\([Suzgun et al\., 2022](https://arxiv.org/html/2609.38585#bib.bib21)\),MATH\([Hendrycks et al\., 2021b](https://arxiv.org/html/2609.38585#bib.bib20)\), andGPQA\([Rein et al\., 2023](https://arxiv.org/html/2609.38585#bib.bib19)\)\. We evaluate out\-of\-domain generalization on two additional tasks that are not observed during training:IFEval\([Zhou et al\., 2023](https://arxiv.org/html/2609.38585#bib.bib18)\)andMuSR\([Sprague et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib17)\)\. These tasks cover diverse reasoning settings including multi\-step reasoning, mathematical reasoning, graduate\-level question answering, instruction\-following evaluation, and multi\-step reasoning over structured contexts\.

For thelarge\-pool setting, the in\-domain tasks includeMMLU\([Hendrycks et al\., 2021a](https://arxiv.org/html/2609.38585#bib.bib12)\),HellaSwag\([Zellers et al\., 2019](https://arxiv.org/html/2609.38585#bib.bib15)\),GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.38585#bib.bib11)\), andARC\([Clark et al\., 2018](https://arxiv.org/html/2609.38585#bib.bib13)\)\. Out\-of\-domain evaluation is conducted onTruthfulQA\([Lin et al\., 2022](https://arxiv.org/html/2609.38585#bib.bib16)\)andWinoGrande\([Sakaguchi et al\., 2021](https://arxiv.org/html/2609.38585#bib.bib14)\)\. These datasets span knowledge\-intensive reasoning, commonsense inference, mathematical problem solving, and truthfulness evaluation\. Table[7](https://arxiv.org/html/2609.38585#A3.T7)summarizes the datasets used in our experiments\.

Following prior work, we evaluate FlexRouter both in\-domain and out\-of\-domain\. For the in\-domain tasks, the router is evaluated on the held\-out 20% test splits of the tasks used for training, measuring its ability to handle unseen instances from familiar tasks\. For the out\-of\-domain tasks, the router is trained on the in\-domain tasks but evaluated on additional tasks that are entirely excluded from training\. This setup assesses the router’s ability to generalize to previously unseen tasks\.

DatasetCategory\#Prompts\#LLMsBBHComplex Reasoning57613811MATHMathematical Reasoning13243811GPQAGraduate\-level QA11923811IFEvalInstruction Following5413811MuSRMulti\-step Reasoning7563811MMLUKnowledge140425000HellaSwagCommonsense Reasoning100425000GSM8KMathematical Reasoning13195000ARCScience QA11725000TruthfulQATruthfulness Evaluation8175000WinoGrandeCommonsense Reasoning12675000Table 7:Statistics of datasets used in our work\. Each dataset consists of prompts evaluated across a shared pool of candidate LLMs within each setting\.
### C\.2Baselines

We compare FlexRouter with several strong routing baselines that represent commonly used model selection strategies usingIndependent scoring, which ranks candidate models according to predicted correctness scores and selects the top\-kkmodels with the highest scores\. Each model is evaluated independently, and correlations between models are not considered during selection\. This setting represents a common strategy used in many routing and ensemble systems where models are chosen solely based on their estimated accuracy\.

EmbedLLM\([Zhuang et al\., 2025](https://arxiv.org/html/2609.38585#bib.bib2)\)learns query\-dependent scores by jointly embedding queries and candidate models into a shared representation space\. The router predicts the likelihood that each model will answer a query correctly based on these embeddings\. However, model selection is still performed independently for each candidate, without explicitly modeling correlations between models\.

BaRP\([Wei et al\., 2025](https://arxiv.org/html/2609.38585#bib.bib1)\)is a learned routing method that models routing as a multi\-objective contextual bandit problem\. It uses REINFORCE on bandit feedback to learn a routing policy that balances performance and cost when selecting models\. BaRP serves as a strong learned baseline for adaptive model selection under resource constraints\.

Reference model \(Ref\. score\)\.The Ref\. score corresponds to the performance of a representative high\-capacity LLM evaluated directly on the benchmark without routing\. This reference model is not a routing method and does not perform model selection\. Instead, it reflects the performance of a single strong model such as GPT\-4\([OpenAI et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib22)\)\. The reference score provides a useful point of comparison because it represents the outcome of always relying on a single capable model\.

For all baselines, we evaluate performance under identical inference budgets\. Each routing method selects the same number of candidate models per query, ensuring that performance differences reflect the quality of the routing strategy rather than differences in computational cost\.

### C\.3Evaluation Metrics

We evaluate routing performance using metrics that capture both success probability and model diversity\.

Success@kkmeasures the fraction of queries for which the selected subset contains at least one model that produces a correct answer\. This metric directly reflects the routing objective of maximizing the probability that at least one selected model succeeds\.

Success@k=1\|𝒬\|∑q∈𝒬\[∃m∈𝒮q:yq,m=1\]\.\\text\{Success@\}k\\;=\\;\\frac\{1\}\{\|\\mathcal\{Q\}\|\}\\sum\_\{q\\in\\mathcal\{Q\}\}\\mathbf\{1\}\\\!\\left\[\\,\\exists\\,m\\in\\mathcal\{S\}\_\{q\}:y\_\{q,m\}=1\\right\]\.\(11\)
Here𝒬\\mathcal\{Q\}denotes the set of queries,𝒮q\\mathcal\{S\}\_\{q\}is the subset of models selected for queryqq, andyq,m∈\{0,1\}y\_\{q,m\}\\in\\\{0,1\\\}indicates whether modelmmanswers queryqqcorrectly\.

ILD@kk\(Intra\-List Diversity\) measures the average pairwise cosine distance among the embeddings of the selected models\. Larger values indicate lower redundancy and greater complementarity among the selected models\.

ILD@​k=1\|𝒬\|​∑q∈𝒬1\(\|𝒮q\|2\)​∑i,j∈𝒮qi<j\(1−𝐮i⊤​𝐮j‖𝐮i‖​‖𝐮j‖\),\\text\{ILD@\}k\\;=\\;\\frac\{1\}\{\|\\mathcal\{Q\}\|\}\\sum\_\{q\\in\\mathcal\{Q\}\}\\frac\{1\}\{\\binom\{\|\\mathcal\{S\}\_\{q\}\|\}\{2\}\}\\sum\_\{\\begin\{subarray\}\{c\}i,j\\,\\in\\,\\mathcal\{S\}\_\{q\}\\\\ i<j\\end\{subarray\}\}\\left\(1\-\\frac\{\\mathbf\{u\}\_\{i\}^\{\\top\}\\mathbf\{u\}\_\{j\}\}\{\\\|\\mathbf\{u\}\_\{i\}\\\|\\,\\\|\\mathbf\{u\}\_\{j\}\\\|\}\\right\),\(12\)
where𝐮m\\mathbf\{u\}\_\{m\}denotes the embedding of modelmm\. The metric ranges from00to22, where00indicates that all selected models are identical in the embedding space and22indicates that model embeddings are maximally dissimilar\. We compute ILD only for queries with at least two selected models\. To ensure fair comparison across routing methods, the embeddings𝐮m\\mathbf\{u\}\_\{m\}are taken from a fixed pre\-trained embedding space derived from model metadata\. As a result, ILD reflects model diversity independently of each method’s internal scoring mechanism\.

### C\.4Implementation Details

FlexRouter uses a shared query encoder together with learnable model embeddings to parameterize the routing policy\. Queries are encoded usingRoBERTa\-base, and both query and model representations are projected into a shared latent space with dimension 128\. The resulting representations are used to construct the query\-dependent DPP kernel that defines the subset selection distribution\.

The model is trained using the Adam optimizer with a learning rate of10−310^\{\-3\}\. Training proceeds for up to 100 epochs with early stopping based on validation Success@kk, using a patience of 15 epochs\. The training objective combines the coverage\-oriented objective and an auxiliary binary cross\-entropy loss\. The weight for auxiliary supervision is set toλ=1\.0\\lambda=1\.0for all experiments\.

During inference, FlexRouter selects models using a greedy maximum a posteriori procedure based on marginal log\-determinant gains of the DPP objective\. Model selection proceeds iteratively and stops when the marginal gain becomes non\-positive or when the maximum subset sizek=10k=10is reached\. This adaptive stopping rule allows the router to adjust the number of selected models according to query difficulty while maintaining a bounded inference budget\. All experiments were conducted on NVIDIA A100 80GB GPUs\.

While the kernel is of sizeM×MM\\times M, greedy inference only requires computing marginal gains with respect to the current subset and maintaining a Cholesky factorization, resulting inO⁡\(k2​M\)O\(k^\{2\}M\)complexity\. In practice, we further restrict evaluation to a top\-K0K\_\{0\}candidate pool for efficiency\.

## Appendix DAdditional Results

The results in this section are from additional runs conducted for further analysis\. They follow the same evaluation protocol as the main experiments, but may differ slightly from Tables 1\-2 due to stochastic training and rerunning the models\. The qualitative trends remain consistent\.

### D\.1Additional Coverage Analysis

The main experiments report Success@10 because all routing methods are evaluated under the same maximum budgetkmax=10k\_\{\\max\}=10\. Since RouterEval contains thousands of candidate LLMs in both the medium\-pool and large\-pool settings, selecting 10 models is still a compact subset of the full candidate pool\. To further examine whether the advantage of FlexRouter only appears atk=10k=10, we rerun the experiments and report average Success@k across different budgets in Table[8](https://arxiv.org/html/2609.38585#A4.T8)\.

FlexRouter trails EmbedLLM atk=1k=1, which is expected because the method is not optimized for selecting a single strongest model\. However, FlexRouter overtakes independent selection once routing becomes genuinely multi\-model, e\.g\., fromk=3k=3in the medium\-pool setting and fromk=2k=2in the large\-pool setting\. This supports the intended use case of FlexRouter: constructing a compact candidate pool whose selected models cover complementary strengths rather than redundantly selecting similar high\-scoring models\.

Pool composition\.Success@k measures whether the selected pool contains at least one correct model, but it does not show how many correct models are selected or whether improvements come from reducing all\-wrong pools\. We therefore report two additional pool\-composition metrics\. Avg\-Correct@10 measures the average number of correct models in the selected pool, while Zero\-Correct Rate measures the fraction of queries for which all selected models are incorrect\. The results are shown in Table[9](https://arxiv.org/html/2609.38585#A4.T9)\.

The results show a trade\-off between correct\-candidate density and candidate\-pool coverage\. EmbedLLM obtains slightly higher Avg\-Correct@10, but it also leaves more queries with zero correct candidates\. In contrast, FlexRouter reduces the Zero\-Correct Rate while maintaining a high number of correct models in the selected pool\. Compared with EmbedLLM, FlexRouter reduces the Zero\-Correct Rate from 0\.160 to 0\.128 in the medium\-pool setting and from 0\.022 to 0\.008 in the large\-pool setting\. This suggests that independent top\-k selection may concentrate correct models on some queries while leaving more queries completely uncovered, whereas FlexRouter reduces the number of all\-wrong selected pools\.

Table 8:Average Success@k across different model\-call budgets\. FlexRouter is designed for complementary subset construction rather than single\-model routing\. It becomes stronger once the setting allows multiple models to be selected\.Table 9:Pool\-composition analysis atk=10k=10\. Avg\-Correct@10 measures the average number of correct models in the selected subset, while Zero\-Correct Rate measures the fraction of queries for which the selected subset contains no correct model\.
### D\.2Coverage\-Oriented Baselines

The main experiments compare FlexRouter with strong routing baselines, including EmbedLLM and BaRP\. However, these methods mainly score candidate models independently and are not explicitly designed to optimize Success@k\. To further test whether FlexRouter’s gains come from the DPP coverage objective rather than from adding any diversity heuristic, we compare with three additional coverage\- and diversity\-oriented baselines\.

#### Random\-k\.

Random\-k uniformly selectskkmodels from the candidate pool\.

#### MaxDiversity\.

MaxDiversity greedily selects models to maximize pairwise embedding diversity\. We initialize the selected set with the highest predicted\-quality model and then greedily add the model that maximizes the minimum pairwise distance to the already selected models\.

#### MMR\.

MMR greedily balances predicted model quality and redundancy\. At each step, it selects

i∗=arg⁡maxi∉S​α​qi​\(x\)−\(1−α\)​maxj∈S⁡sim⁡\(i,j\),i^\{\*\}=\\arg\\max\_\{i\\notin S\}\\alpha q\_\{i\}\(x\)\-\(1\-\\alpha\)\\max\_\{j\\in S\}\\mathrm\{sim\}\(i,j\),\(13\)whereqi​\(x\)q\_\{i\}\(x\)is the predicted quality score andsim⁡\(i,j\)\\mathrm\{sim\}\(i,j\)is cosine similarity in the fixed model\-embedding space\. This baseline is closest to a diversity\-aware top\-k heuristic, since it applies a post\-hoc pairwise penalty to independently predicted quality scores\.

Table 10:Coverage\-oriented baselines across medium\- and large\-pool settings\. FlexRouter outperforms Random\-k, MaxDiversity, and MMR across different budgets\.FlexRouter consistently outperforms the coverage\-oriented baselines in both RouterEval settings\. The comparison with MMR is especially informative: MMR applies diversity as a post\-hoc penalty to independently predicted quality scores, while FlexRouter jointly learns query\-dependent model quality and redundancy through the DPP coverage objective\. These results suggest that the improvement is not simply due to adding a generic diversity heuristic, but comes from learning complementarity under the coverage\-oriented subset objective\.

Table 11:Supervision ablation in the medium\-pool setting\. FlexRouter remains effective when trained with fewer labeled queries\.

### D\.3Supervision Ablation

RouterEval provides dense binary correctness labels for each query–model pair\. To study how much query\-level supervision FlexRouter needs, we train on 25%, 50%, and 100% of the available training queries in the medium\-pool setting, while evaluating all checkpoints on the same held\-out test set\. For each retained query, we keep all model\-level correctness labels\. Therefore, this ablation measures query\-level supervision efficiency rather than model\-level label sparsity\.

FlexRouter degrades gracefully as the number of labeled training queries decreases\. With only 25% of the training queries, it achieves 0\.860 average Success@10, within 2\.4 percentage points of the full\-supervision setting\. With 50% of the training queries, it achieves 0\.865 average Success@10, within 1\.9 percentage points of the full\-supervision setting\. This suggests that FlexRouter can learn useful quality\-diversity structure without using the full query pool\.

This experiment studies query\-level label sparsity\. It does not address model\-level label sparsity, where only a subset of candidate models is labeled for each query\. That setting is closer to matrix completion or bandit\-style exploration and remains an important direction for scalable routing\.

## Appendix ERelated Work

#### LLM Routing\.

Routing across LLMs has emerged as a practical approach for improving performance under heterogeneous model capabilities and costs\. A common strategy is to predict model\-wise correctness scores and select the top\-kkmodels independently\. Methods such as EmbedLLM\([Zhuang et al\., 2025](https://arxiv.org/html/2609.38585#bib.bib2)\)learn joint representations of queries and models to estimate query\-dependent performance\. Other recent approaches, such as RouterDC\([Chen et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib6)\)and BaRP\([Wei et al\., 2025](https://arxiv.org/html/2609.38585#bib.bib1)\), formulate routing as representation learning or contextual bandit problems\. However, these methods typically score models independently and do not explicitly account for correlations or shared failure modes, which can lead to redundant selections\.

#### Multi\-Model Ensembles and Verification\.

Our coverage\-oriented objective aligns with multi\-generation LLM pipelines where a downstream verifier or ranker selects the final answer\. This paradigm is widely used in ensemble methods like LLM\-Blender\([Jiang et al\., 2023](https://arxiv.org/html/2609.38585#bib.bib25)\), which ranks outputs from multiple diverse models, and self\-consistency frameworks\([Wang et al\., 2022](https://arxiv.org/html/2609.38585#bib.bib26)\)that aggregate multiple reasoning paths\. More recently, frameworks like Mixture\-of\-Agents\([Wang et al\., 2024](https://arxiv.org/html/2609.38585#bib.bib27)\)have demonstrated that layering multiple LLMs can significantly enhance generation quality through collaborative refinement\. While these methods benefit from diverse candidate pools, they typically rely on fixed, manually curated model sets\. FlexRouter automates and optimizes the construction of these complementary candidate sets on a per\-query basis\.

#### Determinantal Point Processes\.

Determinantal Point Processes\([Kulesza and Taskar, 2012](https://arxiv.org/html/2609.38585#bib.bib8)\)have been widely used for subset selection tasks that require balancing quality and diversity, including document summarization\([Cho et al\., 2019](https://arxiv.org/html/2609.38585#bib.bib33)\), recommendation\([Wilhelm et al\., 2018](https://arxiv.org/html/2609.38585#bib.bib32)\), and information retrieval\([Affandi et al\., 2014](https://arxiv.org/html/2609.38585#bib.bib30);[Deng et al\., 2020](https://arxiv.org/html/2609.38585#bib.bib31)\)\. A DPP defines a distribution over subsets where the determinant of a kernel matrix captures both item quality and pairwise diversity\. Prior work has explored learning DPP kernels from data and applying DPPs to subset selection problems with fixed ground\-truth subsets\. In contrast, our setting involves weak supervision where no single optimal subset is observed\. We therefore design a coverage\-oriented objective that directly optimizes the probability of selecting at least one correct model, while leveraging the diversity\-inducing properties of DPPs to model complementarity among LLMs\.

Similar Articles

FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

Hugging Face Daily Papers

FlexRouter proposes a coverage-oriented LLM routing framework that uses Determinantal Point Processes to model complementarity among models, maximizing the probability that at least one selected model answers correctly while avoiding redundant selections and fixed budgets.

Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

Hugging Face Daily Papers

The paper introduces RouteFM, a foundation model for LLM routing that learns reusable routing capabilities via episodic pretraining across heterogeneous environments, allowing a frozen router to adapt to new domains, modalities, and candidate pools through behavioral context alone. RouteFM outperforms the strongest baseline by 2.23 quality points on MMR-Bench with only eight observations per candidate, supporting a 'pretrain once, route anywhere' paradigm.

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Hugging Face Daily Papers

The paper proposes SaveRouter, a sparse-supervision LLM routing framework that selectively acquires query-model feedback and shares capability information across related queries, cutting supervision costs while maintaining competitive routing quality and reducing break-even deployment volume by 1.9-9.5x.