LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

arXiv cs.CL Papers

Summary

LLMRouter presents a unified formulation of LLM routing as a sequential decision process, along with an open-source infrastructure and benchmark (xRouteBench) for developing, evaluating, and deploying LLM routers. Empirical results show learned routers achieve 14.6% relative improvement over the strongest fixed-model baseline.

arXiv:2608.06867v1 Announce Type: new Abstract: No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for constructing routing supervision and evaluating routers jointly on response quality and inference cost. The resulting benchmark, xRouteBench, spans generic LLM, memory-augmented, vision, time-series, and personalized routing tasks. We further introduce LLMRouter, an open-source modular infrastructure with more than 16 representative routers. Our empirical study shows that learned routers outperform the strongest fixed-model baseline by 14.6% relatively, lightweight routers become more competitive under tight cost constraints, and user-conditioned routing consistently improves personalization.
Original Article
View Cached Full Text

Cached at: 08/10/26, 08:04 AM

# Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Source: [https://arxiv.org/html/2608.06867](https://arxiv.org/html/2608.06867)
Tao Feng1\*, Fangxu Yu2\*, Haozhen Zhang3\*, Zhongjie Dai1, Liangqi Yuan4, Zijie Lei1, Weizhi Zhang5,Kunlun Zhu1,Haodong Yue1,Keyang Xuan1,Ge Liu1,Jiaxuan You1 1University of Illinois Urbana\-Champaign,2University of Maryland, College Park, 3Nanyang Technological University,4Purdue University,5University of Illinois Chicago

[Project](https://ulab-uiuc.github.io/LLMRouter/)![[Uncaptioned image]](https://arxiv.org/html/2608.06867v1/x1.png)[xRouteBench](https://huggingface.co/datasets/ulab-ai/xRouteBench)![[Uncaptioned image]](https://arxiv.org/html/2608.06867v1/x2.png)[Code](https://github.com/ulab-uiuc/LLMRouter)

###### Abstract

No single large language model \(LLM\) is optimal across all queries and budget constraints, making model routing essential for cost\-effective LLM deployment\. Existing routers span binary quality predictors, cost\-aware cascades, graph\-based routers, and agentic routers, yet their diverse formalisms and incompatible implementations, coupled with the absence of a standardized evaluation pipeline, hinder fair comparison and further extension\. In this paper, we present a unified formulation of LLM routing as a sequential decision process\. Under this formulation, a router can be characterized in terms of five types of components: context encoders, model encoders, scoring functions, decision rules, and learning signals\. Existing methods can then be organized into three families of single\-turn, multi\-turn, and personalized routing\. Building on this formulation, we develop an automated pipeline that constructs routing supervision by systematically running a pool of candidate models across benchmarks and evaluates routers in terms of both response quality and inference cost under a unified protocol\. The resulting benchmark,xRouteBench, spans generic LLM tasks, memory\-augmented, vision \(image and video\), time\-series, and personalized routing scenarios\. Grounded in the formulation and pipeline, we presentLLMRouter, an open\-source infrastructure for standardized and modular implementation of LLM routers, where users can add a new router by implementing only a routing method and a loss function and access built\-in implementations of more than 16 representative routers spanning all three families\. Using the library and benchmark, we conduct a systematic empirical study of LLM routing and find that learned routers achieve a 14\.6% relative improvement over the strongest fixed\-model baseline, router rankings reverse in favor of lightweight designs under tighter cost constraints, and user\-conditioned routing delivers consistent personalization gains\.

††footnotetext:\*Equal contribution\.## 1Introduction

The rapid proliferation of large language models \(LLMs\) has created a heterogeneous ecosystem of models with widely varying costs and task\-specific capabilities, ranging from frontier systems to substantially cheaper open\-weight alternatives\. Since no single model is optimal across all queries and budget constraints, model routing, which determines which model should handle each query, has become essential for cost\-effective LLM deployment\. Beyond cost efficiency, routing also matches each query to the candidate model best suited to it and adapts model choice to user\-specific preferences \(Figure[1](https://arxiv.org/html/2608.06867#S2.F1)\)\. Rich research has been devoted to this problem, from binary routers that arbitrate between a weak and a strong model\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339); Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\)and cost\-aware cascades\(Chenet al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib71); Aggarwalet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib335)\), to reward\-guided ensembles, contrastive and graph\-based routers\(Chenet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib295); Fenget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib304)\), personalized routers that adapt to individual users\(Xieet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib317); Daiet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib319)\), and agentic routers trained with reinforcement learning\(Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296); Fenget al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib19)\)\.

Despite this rapid progress, the field still lacks a unified foundation on which these diverse approaches can be developed and compared, mainly due to two obstacles\. First, existing routers are developed under distinct formalisms, released as separate codebases with incompatible interfaces, trained with different supervision, and tuned for different candidate pools, making it difficult to isolate the design elements that actually drive performance or determine whether observed differences stem from the routers themselves or from their broader experimental stacks\. Second, evaluating a router is fundamentally more demanding than evaluating a single model, as constructing routing supervision and enabling standardized evaluation require running every candidate model on every benchmark query and scoring each response with task\-specific metrics\. Existing benchmarks precompute candidate responses for fixed model pools\(Huet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib338); Huanget al\.,[2025b](https://arxiv.org/html/2608.06867#bib.bib334)\), but are limited to single\-turn routing and provide no pipeline for generating supervision for new benchmarks or candidate pools\. Consequently, multi\-turn and personalized routing still lack a standardized, cost\-aware evaluation framework, making routing methods difficult to compare, reuse, and transfer from offline studies to real applications\.

In this paper, we introduceLLMRouter, a unified infrastructure for developing, evaluating, and deploying routing policies over heterogeneous LLM backends\. Within this infrastructure, a router is characterized by five types of components: context encoders, model encoders, scoring functions, decision rules, and learning signals\. This abstraction accommodates existing routers, which we group into three families of single\-turn, multi\-turn \(including agentic\), and personalized routing\. It also reduces the effort to add a new router to implementing a routing method and a loss function, while data construction, training, inference, and evaluation apply unchanged, so switching the router, candidate pool, or training objective requires only a configuration change rather than reimplementation\.LLMRouterincludes more than 16 representative routers spanning all three families, and can expose any of them as an OpenAI\-compatible server for deployment on messaging platforms via OpenClaw\(OpenClaw,[2026](https://arxiv.org/html/2608.06867#bib.bib344)\)or through a ComfyUI\-based visual interface for code\-free prototyping\.

LLMRouterfurther automates the construction of routing supervision and evaluation, the main obstacle to comparing routers on equal footing\. Its pipeline assembles queries from established benchmarks, dispatches each query to a pool of 18 candidate models spanning a broad price range, and scores every response with task\-specific metrics while recording token\-level cost\. Every router is then evaluated on the same queries, candidate pool, and metrics, enabling direct comparison of their quality\-cost trade\-offs\. With this pipeline, we constructxRouteBench, a benchmark that spans generic LLM tasks, memory\-augmented, vision \(image and video\), time\-series, and personalized scenarios under one protocol\.

LeveragingLLMRouterand xRouteBench, we conduct a systematic empirical study of LLM routing across the three families under one protocol\. Our study surfaces four findings: \(i\)no single router dominates, as the best router varies across tasks and cost budgets\. Strong average performance therefore reflects consistency across scenarios rather than dominance in any single setting\. \(ii\)learned routing still outperforms the strongest fixed\-model baseline, because always selecting the largest model incurs the highest cost yet delivers only mediocre performance, whereas learned routers select smaller, cheaper models for many queries that the largest model answers incorrectly\. \(iii\)multi\-turn routing does not consistently outperform single\-turn routing, as additional rounds of decomposition and aggregation often add cost and redundant information, and their benefit hinges on the capability of the base model that performs them\. \(iv\)personalization pays off, but only when user context is modeled well, as a user\-conditioned router ranks first under both the persona judge and real human preferences, yet the two settings favor different personalized designs\. We also release the library and benchmark in the hope of fostering more systematic progress in LLM routing\.

## 2A Unified Formulation of LLM Routing

### 2\.1Routing as a Sequential Decision Process

Existing routers take seemingly incompatible forms, from binary routers that arbitrate between a weak and a strong model\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339); Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\)and cost\-aware cascades\(Chenet al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib71); Aggarwalet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib335)\)to graph\-based routers\(Fenget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib304); Yuet al\.,[2026a](https://arxiv.org/html/2608.06867#bib.bib299)\)and agentic routers trained with reinforcement learning\(Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296)\), yet they can all be formulated as a sequential decision process\. At steptt, the router observes a statest=\(q,u,ht\)s\_\{t\}=\(q,u,h\_\{t\}\), consisting of the input queryqq, an optional user contextuu\(e\.g\., a user identifier with past interactions and feedback\), and the interaction historyhth\_\{t\}accumulated so far, and takes an actionat∈ℳ∪\{⊥\}a\_\{t\}\\in\\mathcal\{M\}\\cup\\\{\\bot\\\}\. The dispatch actionat=ma\_\{t\}=msends the state to candidatemmfrom the poolℳ=\{m1,…,mK\}\\mathcal\{M\}=\\\{m\_\{1\},\\dots,m\_\{K\}\\\}and appends its responseyty\_\{t\}to the history,ht\+1=ht⊕yth\_\{t\+1\}=h\_\{t\}\\oplus y\_\{t\}, while the terminating actionat=⊥a\_\{t\}=\\botends the episode and aggregates the collected responses into the final answeryy; single\-turn routing is the special case that terminates after one dispatch\. The goal of routing is a policyπ\\piwhose trajectoryτ=\(a1,…,aT\)\\tau=\(a\_\{1\},\\dots,a\_\{T\}\)produces high\-quality answers at low inference cost:

π⋆=arg⁡maxπ⁡𝔼q,τ∼π​\[perf​\(y∣q\)−λ⋅c​\(τ\)\],\\pi^\{\\star\}\\;=\\;\\arg\\max\_\{\\pi\}\\;\\mathbb\{E\}\_\{q,\\;\\tau\\sim\\pi\}\\big\[\\,\\mathrm\{perf\}\(y\\mid q\)\\;\-\\;\\lambda\\cdot c\(\\tau\)\\,\\big\],\(1\)whereperf​\(y∣q\)\\mathrm\{perf\}\(y\\mid q\)aggregates task\-specific quality metrics \(e\.g\., accuracy, F1, or an LLM\-judged score\),c​\(τ\)c\(\\tau\)sums the monetary or token cost of every call in the routing trajectoryτ\\tau, andλ≥0\\lambda\\geq 0controls the performance–cost trade\-off\. Under this formulation, a router is characterized by the choice of acontext encoderEqE\_\{q\}that encodes the routing state, amodel encoderEmE\_\{m\}that encodes each LLM candidate, ascoring functionggand adecision ruleddthat turn the context and model representations into a routing action, and alearning signalℒ\\mathcal\{L\}that fits these components toward the optimal policy\. This section elaborates on each component and demonstrates how existing routers fall into three families, namely single\-turn, multi\-turn \(including agentic\), and personalized, with specific designs of these five components \(Table[1](https://arxiv.org/html/2608.06867#S2.T1)\)\. We provide an overview in Figure[1](https://arxiv.org/html/2608.06867#S2.F1)\.

Context encoder\.The context encoderEqE\_\{q\}maps the routing statests\_\{t\}to the representation on which the routing decision is based, and its output takes one of two forms\. i\)Embedding\-based: the state is represented as a vector\. Rating\-based routers degenerate to a constant that ignores the query and routes by global model quality,kkNN\-style routers use an off\-the\-shelf sentence embedding\(Huet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib338); Shnitzeret al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib76)\), discriminative routers train a lightweight encoder over frozen embeddings\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339); Stripeliset al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib337)\), and personalized routers condition the representation on user and session nodes of a heterogeneous interaction graph\(Xieet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib317); Daiet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib319)\)\. ii\)Text\-based: the state is kept in natural language\. Cascades append the draft response and a verification confidence to the query\(Chenet al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib71); Aggarwalet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib335)\), and fine\-tuned LM routers, exemplified by Router\-R1, verbalize the whole state directly in the prompt, leaving its representation to the model’s forward pass\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294); Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296)\)\. Which portion of the stateEqE\_\{q\}reads is precisely what separates the three router families, and any router is personalized by swapping in a user\-conditionedEqE\_\{q\}while inheriting the remaining components\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/x3.png)Figure 1:Overview of LLM routing\.Routing is driven by three needs \(left\), namely cost efficiency, capability matching, and user preference\. Our unified formulation \(right\) casts all of them as one decision process: a context encoderEqE\_\{q\}represents the routing state of query, persona, and interaction history, a model encoderEmE\_\{m\}represents each candidate, and the router dispatches the query or its sub\-queries to selected models and aggregates their responses into the answer\. The single\-turn, multi\-turn, and personalized families differ only in which part of the state they observe\.Model encoder\.The model encoderEmE\_\{m\}encodes each candidate in the pool\. i\)Static metadata: the simplest choice describes a candidate by its model size, capability description, and pricing\. ii\)Historical profiles: most routers instead profile candidates by their past behavior, representing a model by the set of embedded queries it has previously solved \(kkNN\), a scalar rating \(Elo\), or a latent factor fit by matrix factorization\. iii\)Learned embeddings: stronger routers learn model embeddings jointly with the context encoder\(Chenet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib295); Fenget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib304); Zhuanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib320)\)\. iv\)Verbalized description: fine\-tuned LM routers instead name the candidates directly in the prompt\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294); Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296)\)\.

Scoring function and decision rule\.The scoring functionggmeasures the compatibility between the encoded state and each candidate, and the decision ruleddconverts the resulting scores into a routing action\. Instantiations ofggtrack the encoders, from embedding similarity inkkNN\-style routers and a bilinear product in factorization\-based ones, to a classification head\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339); Stripeliset al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib337)\), message passing over a query–model graph\(Fenget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib304)\), and next\-token logits in fine\-tuned LM routers that foldEqE\_\{q\},EmE\_\{m\}, andgginto one forward pass\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\)\. Fordd, greedyarg⁡max\\arg\\maxis the default choice, yet it is optimal for Eq\.[1](https://arxiv.org/html/2608.06867#S2.E1)only whenλ=0\\lambda=0\. Cost\-aware rules instead threshold the predicted quality gap between a cheap and an expensive model\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339)\), accept or escalate in cascades\(Chenet al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib71); Aggarwalet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib335)\), or sample for exploration in online settings\(Daiet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib81)\), and multi\-turn routers further equipddwith the terminating action⊥\\bot\(Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296)\)\.

Learning signal\.The learning signalℒ\\mathcal\{L\}specifies how the components above are fit toward Eq\.[1](https://arxiv.org/html/2608.06867#S2.E1)\. Non\-parametric routers require no training and rely purely on stored interactions\. Supervised routers fit pointwise correctness labels harvested by running the candidate pool over benchmark queries, preference\-based routers learn from pairwise comparisons such as human votes from Chatbot Arena\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\)or contrastive objectives that pull queries toward the models that solve them\(Chenet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib295)\), and agentic routers directly optimize trajectory\-level rewards with reinforcement learning\(Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296)\)\. In every case,ℒ\\mathcal\{L\}is a surrogate for the same objective\. What differs is not the goal but the form in whichperf\\mathrm\{perf\}is observable, measured for every candidate by supervised routers, returned only at the end of a trajectory for agentic ones, and revealed only through comparisons when quality is user\-specific\.

Table 1:Instantiation of the unified routing formulation for the three router families\.For each family, the table specifies the routing statess, the context and model encodersEqE\_\{q\}andEmE\_\{m\}, the routing action defined by the scoring functionggand decision ruledd, and the learning signalℒ\\mathcal\{L\}used to optimize response quality and inference cost\.FamilyStatessEncodersEq,EmE\_\{q\},E\_\{m\}Routing action\(scoringgg, decisiondd\)Learning signalℒ\\mathcal\{L\}\(surrogate of Eq\.[1](https://arxiv.org/html/2608.06867#S2.E1)\)Single\-turn\(q\)\(q\)Eq​\(q\),Em​\(m\)E\_\{q\}\(q\),\\;E\_\{m\}\(m\)a=arg⁡maxm∈ℳ⁡g​\(Eq​\(q\),Em​\(m\)\)a=\\arg\\max\_\{m\\in\\mathcal\{M\}\}\\,g\\big\(E\_\{q\}\(q\),\\,E\_\{m\}\(m\)\\big\)fitggto per\-candidate rewardperf​\(ym∣q\)−λ​cm\\mathrm\{perf\}\(y\_\{m\}\\mid q\)\-\\lambda\\,c\_\{m\}Multi\-turn\(q,ht\)\(q,h\_\{t\}\)Eq​\(q,ht\),Em​\(m\)E\_\{q\}\(q,h\_\{t\}\),\\;E\_\{m\}\(m\)at∼d​\(\{g​\(Eq​\(q,ht\),Em​\(m\)\)\}m\)a\_\{t\}\\sim d\\big\(\\\{g\(E\_\{q\}\(q,h\_\{t\}\),\\,E\_\{m\}\(m\)\)\\\}\_\{m\}\\big\)maximize episode return𝔼τ​\[perf​\(y∣q\)−λ​c​\(τ\)\]\\mathbb\{E\}\_\{\\tau\}\\big\[\\mathrm\{perf\}\(y\\mid q\)\-\\lambda\\,c\(\\tau\)\\big\]Personalized\(q,u,ht\)\(q,u,h\_\{t\}\)Eq​\(q,u,ht\),Em​\(m\)E\_\{q\}\(q,u,h\_\{t\}\),\\;E\_\{m\}\(m\)a=arg⁡maxm∈ℳ⁡g​\(Eq​\(q,u,ht\),Em​\(m\)\)a=\\arg\\max\_\{m\\in\\mathcal\{M\}\}\\,g\\big\(E\_\{q\}\(q,u,h\_\{t\}\),\\,E\_\{m\}\(m\)\\big\)fitggto comparisonsm\+≻um−m^\{\+\}\\\!\\succ\_\{u\}\\\!m^\{\-\}observingperfu\\mathrm\{perf\}\_\{u\}

### 2\.2Automatic Evaluation of LLM Routing

Evaluating a router is substantially more demanding than evaluating a single model\. Under Eq\.[1](https://arxiv.org/html/2608.06867#S2.E1), a router must be judged on both the quality of its answers and the cost spent to obtain them, and constructing its supervision requires knowing how every candidate performs on every query under task\-specific metrics\. In current practice, these elements are assembled manually for a single benchmark and fixed candidate pool\(Huet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib338); Huanget al\.,[2025b](https://arxiv.org/html/2608.06867#bib.bib334)\), requiring fresh engineering for every new task or candidate pool\.

LLMRouterautomates this process end\-to\-end with a three\-stage pipeline: i\)Query Curation: queries are sampled from source benchmarks, normalized into a unified schema, and split into training and test sets; ii\)Response Collection: each query is dispatched to every candidate in the pool, which is declared in a single configuration file, and responses are collected together with their token counts; iii\)Metric Scoring and Pricing: every response is scored with its task metric and priced from its token counts\. The product is a dense query–model matrix of performance and cost that serves at once as routing supervision and as the test bed, so a new task or candidate pool enters through a configuration change rather than a re\-engineered stack\.

Evaluation reuses this path, except that a test query goes only to the candidate the router selects rather than to the whole pool\. Every router therefore faces the same queries, candidate pool, and metrics, so measured differences reflect the routing policy rather than the surrounding stack\. For the multi\-turn and agentic families, every decomposition and aggregation call is priced into the trajectory cost\.

Metrics\.Task quality is assessed using built\-in metrics aligned with standard benchmark conventions, including exact and close matching, multiple\-choice accuracy, token\-level F1, mathematical answer verification, and execution\-based code evaluation\.LLMRouteralso supports optional LLM\-based judging and exposes a weighted objective that balances performance and cost, allowing routers and trainers to target performance\-first, cost\-sensitive, or hybrid operating points\.

## 3xRouteBench: A Multi\-Scenario Benchmark for LLM Routing

Existing routing benchmarks cover only a subset of the settings captured by the formulation in §[2\.1](https://arxiv.org/html/2608.06867#S2.SS1)\. RouterBench\(Huet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib338)\)precomputes candidate responses for single\-turn text queries over a fixed candidate pool, but does not cover settings in which input\-token costs dominate or only a subset of candidate models can process the input\. RouterEval\(Huanget al\.,[2025b](https://arxiv.org/html/2608.06867#bib.bib334)\)aggregates large\-scale performance records but likewise focuses on single\-turn text tasks and evaluates response quality independently of inference cost\. Recent vision–language routing benchmarks\(Huanget al\.,[2025a](https://arxiv.org/html/2608.06867#bib.bib326); Maet al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib327)\)extend routing evaluation to image inputs, but remain limited to image question answering and do not cover video, long\-context, or modality\-selection settings\. Preference data from Chatbot Arena\(Zhenget al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib329)\)provides population\-level signals but lacks the persistent user context needed to supervise user\-conditioned routing\. To close these gaps, we construct xRouteBench to evaluate, under a unified cost\-aware protocol, regimes in which routing decisions fundamentally differ: long\-context inputs for which input\-token costs dominate, image and video inputs that only a subset of candidate models can process, time\-series inputs with multiple modality encodings, and tasks with user\-specific quality preferences\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/x4.png)Figure 2:Task composition of xRouteBench\.The benchmark covers generic LLM tasks, memory, vision, time\-series, and personalized routing, with percentages indicating the proportion of test queries contributed by each dataset\.Design principle\.All tasks are constructed by theLLMRouterdata engine and share a common query schema, supervision format, and evaluation protocol\. Each non\-text asset is converted by a transformation script into a self\-contained textual query, with an optional pointer to the source image, video, or time series\. This separates routing from perception by ensuring that text\-only and multimodal candidates receive the same textual input\. For every query, the protocol jointly evaluates response quality and inference cost, exposing regimes in which input\-token costs dominate answer\-generation costs\. Adding a new application requires only a transformation script and a registered metric\.

Tracks\.xRouteBench spans five tracks comprising 4,767 instances\. Figure[2](https://arxiv.org/html/2608.06867#S3.F2)provides an overview of task distribution\. Specifically, we include: \(i\)Generic LLM Tasksmix established knowledge and commonsense QA \(MMLU\(Hendryckset al\.,[2020](https://arxiv.org/html/2608.06867#bib.bib315)\), MMLU\-Pro\(Wanget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib345)\), ARC\-Challenge\(Clarket al\.,[2018](https://arxiv.org/html/2608.06867#bib.bib311)\), OpenBookQA\(Mihaylovet al\.,[2018](https://arxiv.org/html/2608.06867#bib.bib312)\), CommonsenseQA\(Talmoret al\.,[2019](https://arxiv.org/html/2608.06867#bib.bib314)\), BoolQ\(Clarket al\.,[2019](https://arxiv.org/html/2608.06867#bib.bib333)\), HellaSwag\(Zellerset al\.,[2019](https://arxiv.org/html/2608.06867#bib.bib330)\), SQuAD\(Rajpurkaret al\.,[2016](https://arxiv.org/html/2608.06867#bib.bib331)\)\), mathematical reasoning \(GSM8K\(Cobbeet al\.,[2021](https://arxiv.org/html/2608.06867#bib.bib209)\), MATH\(Hendryckset al\.,[2021](https://arxiv.org/html/2608.06867#bib.bib309)\), AIME\), and code generation \(MBPP\(Austinet al\.,[2021](https://arxiv.org/html/2608.06867#bib.bib306)\), HumanEval\(Chen,[2021](https://arxiv.org/html/2608.06867#bib.bib203)\)\), the conventional single\-turn text setting\. \(ii\)Memoryroutes long\-horizon conversational QA over hundreds of accumulated turns \(LoCoMo\(Maharanaet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib336)\), LongMemEval\(Wuet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib305)\)\), where token cost is governed by the history rather than the answer\. \(iii\)Visioncovers image\-grounded mathematical reasoning \(Geometry3K\(Luet al\.,[2021](https://arxiv.org/html/2608.06867#bib.bib298)\), MathVista\(Luet al\.,[2024b](https://arxiv.org/html/2608.06867#bib.bib308)\)\) and egocentric video understanding \(Charades\-Ego\(Sigurdssonet al\.,[2018](https://arxiv.org/html/2608.06867#bib.bib307)\)\)\. \(iv\)TimeSeriescovers time\-series reasoning \(TSRBench\(Yuet al\.,[2026b](https://arxiv.org/html/2608.06867#bib.bib264)\)\), with each series rendered as both text and image so the router also selects a modality encoding\. \(v\)Personalizeddraws open\-ended prompts from Chatbot Arena and MT\-Bench\(Zhenget al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib329)\), each tied to a user persona and scored by a persona\-conditioned LLM judge, so its supervision is preference feedback rather than pointwise correctness\.

## 4TheLLMRouterLibrary

![Refer to caption](https://arxiv.org/html/2608.06867v1/x5.png)Figure 3:Architecture ofLLMRouter\.The system consists of six modules that support routing data construction, router implementation and training, inference, evaluation, and deployment\.LLMRouterties the formulation of §[2\.1](https://arxiv.org/html/2608.06867#S2.SS1), the evaluation protocol of §[2\.2](https://arxiv.org/html/2608.06867#S2.SS2), and xRouteBench into one executable system organized as six modules around a single query–model matrix \(Figure[3](https://arxiv.org/html/2608.06867#S4.F3)\)\. Its organizing principle is that the five components of a router are the only thing a user writes, while data construction, training, inference, evaluation, and deployment are shared infrastructure that operates on any router unchanged\. Swapping a router, a candidate pool, or a training objective is therefore a configuration change rather than a reimplementation\.

Data Engine\.The data engine implements the three\-stage construction pipeline of §[2\.2](https://arxiv.org/html/2608.06867#S2.SS2), turning a declared task list and candidate pool into the query–model matrix that supervises and tests every router\. Adding a task requires only a prompt template and a registered metric, and adding a candidate requires only its endpoint and per\-token price\.

Router Library\.The library implements more than 16 routers spanning all three families under a unifiedMetaRouterinterface, which hides mechanisms as different as nearest\-neighbor retrieval in akkNN router and autoregressive decoding in a fine\-tuned LM router behind one call\. Adding a new router requires subclassingMetaRouterand implementing eitherroute\_singleor its batched counterpart,route\_batch\. Within this method, the context encoderEqE\_\{q\}, model encoderEmE\_\{m\}, scoring functiongg, and decision ruleddmap a routing state to a routing action\. A short YAML file specifies the router’s candidate pool and objective weights, after which the router can be invoked by name using the same commands as any built\-in method\. Personalization and component ablations can then be performed with a one\-line change, without forking the codebase\. Figure[4](https://arxiv.org/html/2608.06867#S4.F4)shows the complete code needed to design a new router\.

fromllmrouter\.modelsimportMetaRouter,BaseTrainer

classMyRouter\(MetaRouter\):

defroute\_single\(self,query\):

s=self\.encode\_state\(query\)

scores=self\.score\(s,self\.models\)

query\["model\_name"\]=self\.decide\(scores\)

returnquery

classMyRouterTrainer\(BaseTrainer\):

defloss\_func\(self,outputs,batch\):

returnmy\_objective\(outputs,batch\)

router=MyRouter\(yaml\_path="my\_router\.yaml"\)

trainer=MyRouterTrainer\(router\)

trainer\.train\(\)

answer=router\.route\_single\(\{"query":"\.\.\."\}\)

Figure 4:The five components of the routing formulation map onto two classes inLLMRouter\. A router subclassesMetaRouterand implementsroute\_single\(orroute\_batch\), where the context encoderEqE\_\{q\}, model encoderEmE\_\{m\}, scoring functiongg, and decision ruleddturn a state into a selected model; the learning signalℒ\\mathcal\{L\}lives in aBaseTrainersubclass\.Trainer\.Training is decoupled from routing through aBaseTrainer\. Itsloss\_funcdefines the learning signalℒ\\mathcal\{L\}as a pointwise loss, pairwise loss, or trajectory\-level reward, while itstrainloop uses this signal to optimize the router for the weighted objective in Eq\.[1](https://arxiv.org/html/2608.06867#S2.E1)\. A router and its trainer are paired but can be swapped independently, allowing the same scorer to be trained with different forms of supervision without modifying its routing code\. Non\-parametric routers bypass this module\.

Route Engine\.At inference time, the route engine drives any router through the same call, dispatching the query to the selected candidate\. For multi\-turn policies, it repeats the decision step until a termination action is produced and aggregates the collected responses into the final answer\.

Evaluation\.The evaluation module implements the evaluation protocol of §[2\.2](https://arxiv.org/html/2608.06867#S2.SS2), scoring each router on the same test queries, candidate pool, and metrics, and sweeping the trade\-off weightλ\\lambdato trace its performance–cost frontier\.

Deployment\.In addition to training and inference commands,LLMRoutercan expose any router as an OpenAI\-compatible server that integrates with OpenClaw\(OpenClaw,[2026](https://arxiv.org/html/2608.06867#bib.bib344)\)for deployment on messaging platforms such as Slack and Discord\. A routing memory persists the interaction historyhhacross turns, while a ComfyUI\-based canvas supports code\-free prototyping\. The same router evaluated offline can therefore serve live single\-agent and multi\-agent traffic without modification\.

## 5Experiments

### 5\.1Experimental Setups

Benchmarks\.We evaluate routers across the five xRouteBench tracks defined in §[3](https://arxiv.org/html/2608.06867#S3): Generic LLM Tasks, memory, vision, time\-series, and personalized\. Together, they comprise eight test sets, with full statistics reported in Table[6](https://arxiv.org/html/2608.06867#A1.T6)\. For each query in the memory track, we retrieve up to five memory items as context and score responses using token\-level F1\.

LLM Candidates\.The candidate pool contains 18 open\-weight models served through two providers \(i\.e\., Together API111[https://www\.together\.ai/](https://www.together.ai/)and NVIDIA NIM API222[https://build\.nvidia\.com/](https://build.nvidia.com/)\), spanning 7B to 671B parameters\. It covers Gemma\-2\-9B\(Teamet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib303)\); Mistral\-7B, Mistral\-Small\-24B, Mixtral\-8x7B, and Mixtral\-8x22B\(Jianget al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib328);[2024](https://arxiv.org/html/2608.06867#bib.bib171)\); Qwen2\.5\-7B, Qwen3\-Next\-80B, and Qwen3\-Coder\(Yanget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib146);[2025](https://arxiv.org/html/2608.06867#bib.bib68)\); Llama\-3\-8B, Llama\-3\-70B, Llama\-3\.3\-70B, and Llama\-4\-Maverick\(Grattafioriet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib158); Adcocket al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib325)\); GPT\-OSS\-20B and GPT\-OSS\-120B\(Agarwalet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib67)\); RNJ\-1\-15B\(Callahanet al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib323)\); and the two 671B models DeepSeek\-V3\.1\(Liuet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib160)\)and Cogito\-v2\(Deep Cogito,[2025](https://arxiv.org/html/2608.06867#bib.bib324)\)\. Per\-token pricing is given in Table[9](https://arxiv.org/html/2608.06867#A3.T9)in the appendix\.

Implemented Routers\.LLMRouterimplements more than 16 routers covering three families\. \(i\)Single\-turn routersincludekkNNRouter\(Li,[2025](https://arxiv.org/html/2608.06867#bib.bib332)\), SVMRouter, MLPRouter, EloRouter, and MFRouter\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294); Shnitzeret al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib76)\), as well as RouterDC\(Chenet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib295)\), Hybrid LLM\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339)\), AutoMix\(Aggarwalet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib335)\), GraphRouter\(Fenget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib304)\), CausalLM\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\), and two rule\-based baselines that always select the smallest or largest model; \(ii\)Multi\-turn routersinclude Router\-R1\(Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296)\)together withkkNN\-based and LLM\-based multi\-round routers; and \(iii\)Personalized routersinclude GMTRouter\(Xieet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib317)\)and PersonalizedRouter\(Daiet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib319)\)\.

Evaluation Protocol\.We score each router by a weighted rewardα⋅perf−β⋅cost\\alpha\\cdot\\mathrm\{perf\}\-\\beta\\cdot\\mathrm\{cost\}\. We sweep five weight settings from the quality\-only\(α,β\)=\(1\.0,0\.0\)\(\\alpha,\\beta\)=\(1\.0,0\.0\)to the heavily cost\-weighted\(0\.2,0\.8\)\(0\.2,0\.8\)\. Multi\-round and RL\-based routers cannot optimize this weighted objective and are therefore run once under a single configuration\. On the personalized track, answers are scored by a persona\-conditioned LLM judge \(DeepSeek\-V3\.1\) as win, tie, or loss \(11,0\.50\.5,0\), and the judge’s own cost is excluded from the reported cost\.

### 5\.2Main Results

Table 2:Results on xRouteBench under the performance\-first setting\(α,β\)=\(1\.0,0\.0\)\(\\alpha,\\beta\)=\(1\.0,0\.0\)\.Scores are reported across the Generic LLM Tasks, memory, vision, and time\-series tracks, together with their average\. Following the original implementations where applicable, all multi\-turn routers use Qwen2\.5\-3B\-Instruct as the base model\. Top two results are highlighted inboldandunderline\.RouterGeneric LLM TasksMemoryVisionTimeSeriesAvgLoCoMoLongMemEvalGeometry3KMathVistaVideo\\rowcolorbarGrayRule\-based baselinesSmallest\-LLM57\.5525\.4436\.7727\.8735\.0033\.3349\.6137\.94Largest\-LLM70\.2926\.5935\.5737\.7033\.0022\.2245\.6738\.72\\rowcolorbarGraySingle\-turn routerskkNNRouter71\.3725\.2438\.7431\.1541\.0029\.6351\.9741\.30SVMRouter74\.2127\.6438\.6842\.6247\.0029\.6355\.9145\.10MLPRouter68\.1226\.7832\.2727\.8734\.0029\.6356\.6939\.34MFRouter67\.2324\.4934\.9140\.9829\.0022\.2251\.9738\.69EloRouter64\.1525\.7037\.2745\.9050\.0025\.9363\.7844\.68Hybrid LLM64\.6825\.8936\.5632\.7937\.0033\.3351\.1840\.20RouterDC80\.5624\.9336\.7716\.3924\.0025\.9345\.6736\.32GraphRouter80\.5425\.9433\.9342\.6250\.0022\.2262\.9945\.46CausalLM66\.9025\.4037\.6024\.6034\.0033\.3345\.7038\.22\\rowcolorbarGrayMulti\-turn routersRouter\-R135\.6424\.6017\.2814\.7518\.0022\.2223\.6222\.30kkNN\-MultiRound13\.9924\.7018\.3216\.3930\.0025\.9333\.0723\.20LLM\-MultiRound12\.9824\.6017\.4414\.2931\.0325\.9330\.3322\.37

Table[5\.2](https://arxiv.org/html/2608.06867#S5.SS2)and Table[3](https://arxiv.org/html/2608.06867#S5.T3)report the results under the performance\-only setting\. Based on these results, we have the following key observations:

No single router dominates across all tasks: The winner router varies across tasks\. For example, RouterDC performs best on Generic LLM Tasks and SVMRouter on LoCoMo\. Though GraphRouter attains the best average on xRouteBench, it did not consistently outperform other routers in all tasks\.

Multi\-turn routing does not consistently outperform single\-turn routing: Across all benchmarks, multiple rounds of routing and aggregation provide no consistent gain over a single routing decision\. For many queries, one well\-chosen route is sufficient, whereas additional rounds introduce redundant information and computational overhead\. Multi\-turn routers also rely on a base model \(Qwen2\.5\-3B\-Instruct\) to decompose queries and aggregate responses, making their performance sensitive to the capabilities of this model\. These results highlight the need for better sufficiency estimation, early stopping, and more effective decomposition and aggregation\.

Table 3:Performance comparison on the personalized track\.Top two results are highlighted inboldandunderline\.RouterAcc\.RouterAcc\.GMTRouter68\.78RouterDC56\.44PersonalizedRouter67\.86MFRouter54\.39EloRouter66\.40MLPRouter52\.93GraphRouter65\.23kkNNRouter51\.76SVMRouter65\.08CausalLM46\.78Largest\-LLM58\.05Router\-R145\.46Hybrid LLM57\.91Smallest\-LLM42\.53
Conditioning on user context helps, but how it is modeled matters: Table[3](https://arxiv.org/html/2608.06867#S5.T3)reports persona\-judge accuracy on the personalized track\. GMTRouter achieves the highest accuracy of 68\.78, outperforming PersonalizedRouter \(67\.86\) and the best user\-agnostic router, EloRouter \(66\.40\)\. The strong performance of both personalized methods confirms the benefit of conditioning routing decisions on user context, while GMTRouter’s additional 0\.92\-point gain over PersonalizedRouter shows that the way user context is encoded and integrated remains important\.

### 5\.3Performance–Cost Trade\-offs

![Refer to caption](https://arxiv.org/html/2608.06867v1/x6.png)Figure 5:Router rankings across the Generic LLM Tasks, memory, vision, and time\-series tracks as the cost weightβ\\betaincreases\.Each cell gives a router’s rank under the weighted performance–cost objective, with smaller rank values indicating better performance\.No single router is best under every performance–cost trade\-off\.Figure[5](https://arxiv.org/html/2608.06867#S5.F5)ranks the routers by reward within each category as the cost weightβ\\betagrows, and the rankings shift dramatically along the sweep\. RouterDC tops Generic LLM Tasks when only quality matters, yet falls to tenth of eleven under the most cost\-sensitive setting; EloRouter leads Vision and TimeSeries atβ=0\\beta=0but drops out of the lead once cost enters the objective\. However, weakness at one operating point does not imply weakness at another, as MLPRouter sits near the bottom of Vision under the quality\-first setting yet becomes the best choice there for everyβ≥0\.4\\beta\\geq 0\.4\. Therefore, it is practical to choose the router that matches the performance–cost requirements of the deployment at hand\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/x7.png)Figure 6:Performance–cost trade\-offs of routers averaged across the xRouteBench tracks\.Each point represents an operating setting with a different cost weightβ\\beta, where higher performance and lower per\-query inference cost are preferred\.Increasing inference cost leads to improved performance\.Figure[6](https://arxiv.org/html/2608.06867#S5.F6)presents the trade\-off between performance and inference cost\. For most routers, performance and cost exhibit a clear positive correlation, with the operating points rising from the low\-cost to the high\-cost end\. This is because increasing the inference budget unlocks more powerful and expensive models, which perform better\. Meanwhile, always calling the largest model incurs the highest cost yet delivers only mediocre performance and is dominated by the learned routes, since many queries that the largest model fails are solved by smaller and cheaper ones\. This confirms that no single model covers all queries, which is exactly the headroom that routing exploits\.

### 5\.4Routing in Deployment: Real Users and Multi\-Agent Systems

Table 4:Router performance on held\-out real\-user sessions collected through the Slack deployment\.Accuracy measures how often each router’s model selection agrees with the users’ pairwise preferences\.RouterAcc\.RouterAcc\.PersonalizedRouter83\.05RouterDC65\.25EloRouter82\.20kkNNRouter60\.17MLPRouter78\.81kkNN\-MultiRound60\.17SVMRouter77\.12Smallest\-LLM55\.08Hybrid LLM73\.73MFRouter51\.69GMTRouter70\.70Largest\-LLM41\.53GraphRouter67\.17CausalLM27\.97The deployment layer ofLLMRoutercarries routers beyond static benchmarks\. We study two settings it enables: routing for real users served through OpenClaw and routing inside multi\-agent systems\.

Routing for real users\.Using the OpenClaw server ofLLMRouter, we deploy the routing stack behind Slack and collect live preference feedback\. 15 users contribute 40 sessions of 1 to 12 turns, totaling 234 pairwise records\. For each query, two models are sampled from a pool of ten candidates, their answers are shown in randomized positions, and the user marks one as better or declares a tie\. We split by session into 32 training sessions and 8 test sessions, train every router on the human training split, and score how often its selection matches the human preference on held\-out sessions\. Table[4](https://arxiv.org/html/2608.06867#S5.T4)shows that PersonalizedRouter leads at 83\.05, while the fine\-tuned CausalLM router ranks last\. The simulated ranking does not fully transfer, as GMTRouter, the winner under the persona judge, drops to sixth on real users, showing that it matters to validate personalized routers against real feedback\.

Routing inside multi\-agent systems\.A multi\-agent system \(MAS\) is conventionally instantiated with a single base model shared by every agent\. We instead treat model choice as a per\-agent decision, where a router receives the prompt of each functional node and selects the most suitable LLM for that call\. Since node prompts differ substantially, covering planning, execution, and verification instructions, this setting stress\-tests how well a router trained on ordinary queries generalizes\. Following MultiAgentBench\(Zhuet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib322)\)and GraphPlanner\(Fenget al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib19)\), we instantiate five coordination topologies, namely Star, Tree, Graph, Chain, and Plan\-Exec\-Sum, illustrated in Figure[7](https://arxiv.org/html/2608.06867#S5.F7)with node roles detailed in Appendix[E](https://arxiv.org/html/2608.06867#A5)\. The learning\-based routers are trained on the Generic LLM Tasks training split, each test query is then fed through the full MAS, and the final MAS answer is scored by the task metric\. Table[5](https://arxiv.org/html/2608.06867#S5.T5)shows that routing every node pays off, as six of the seven learned routers beat always selecting the largest model on average, with MFRouter attaining the best average of 76\.48 against 71\.48\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/x8.png)Figure 7:Representative multi\-agent system architectures and coordination topologies:\(a\) star\-based centralized coordination, \(b\) hierarchical tree\-based delegation, \(c\) graph\-based peer interaction, \(d\) sequential chain collaboration, and \(e\) planner–executor–summarizer workflow\.Table 5:Router performance on the Generic LLM Tasks test split when each node in a multi\-agent system is routed independently\.Results are reported across five coordination topologies, with the final column showing the average performance\.RouterStarTreeGraphChainPlan\-Exec\-SumAvgLargest\-LLM69\.0067\.0077\.2069\.0075\.2071\.48kkNNRouter74\.8078\.6078\.6076\.6071\.8076\.08SVMRouter76\.2075\.6080\.0074\.4075\.2076\.28MLPRouter75\.4076\.6076\.8078\.0071\.4075\.64MFRouter75\.4074\.2081\.0078\.6073\.2076\.48EloRouter73\.8072\.4078\.6076\.6075\.2075\.32GraphRouter68\.2070\.8066\.2072\.0069\.0069\.24RouterDC77\.6079\.6074\.2072\.0076\.2075\.92## 6Related Work

LLM Routing\.Prior routers can be read along the axes of our formulation\. Single\-turn routers differ mainly in how they encode a query and score candidates, from quality predictors that arbitrate between a weak and a strong model\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339); Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\)and classifiers over frozen embeddings\(Shnitzeret al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib76); Stripeliset al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib337); Li,[2025](https://arxiv.org/html/2608.06867#bib.bib332)\), to reward\-guided rankings\(Luet al\.,[2024a](https://arxiv.org/html/2608.06867#bib.bib301)\), prompt\-conditioned preference models\(Fricket al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib321)\), contrastive query–model matching\(Chenet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib295)\), learned model embeddings\(Zhuanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib320)\), graph\-based scorers\(Fenget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib304)\), and online methods under bandit feedback\(Daiet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib81); Wanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib4)\)\. Multi\-turn routers instead enrich the state across rounds, whether as cascades that escalate upon failed verification\(Chenet al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib71); Aggarwalet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib335)\)or as agentic routers that decompose a query and route sub\-queries with reinforcement learning\(Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296); Fenget al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib19)\)\. Personalized routers add the user to the state and learn from preference feedback\(Xieet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib317); Daiet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib319); Yuet al\.,[2026a](https://arxiv.org/html/2608.06867#bib.bib299)\)\. Each is developed in its own formalism and evaluated on its own stack;LLMRouterinstead expresses them as instantiations of a single sequential decision process behind one interface, so single\-turn, multi\-turn, and personalized routers meet on the same performance–cost frontier\.

Routing Benchmarks and Evaluation\.RouterBench\(Huet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib338)\)precomputes candidate responses over a fixed pool, RouterEval\(Huanget al\.,[2025b](https://arxiv.org/html/2608.06867#bib.bib334)\)aggregates large\-scale performance records for routing study, preference data from Chatbot Arena has served as routing supervision\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\), and recent benchmarks extend routing to vision–language pools\(Huanget al\.,[2025a](https://arxiv.org/html/2608.06867#bib.bib326); Maet al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib327)\)\. However, existing benchmarks each target a single scenario, one\-shot text or image QA, with fixed pools, and provide no pipeline for constructing supervision on new tasks or candidate sets\.LLMRoutercloses this gap with automatic supervision construction and cost\-aware evaluation over configurable pools, and xRouteBench spans generic LLM tasks, memory\-augmented, vision \(image and video\), time\-series, and personalized scenarios under one protocol\.

## 7Conclusion

We introducedLLMRouter, a unified framework for LLM routing that casts single\-turn, multi\-turn, and personalized routing as instances of a common sequential decision process\.LLMRouteralso provides an automatic pipeline for constructing routing supervision and evaluation for new tasks and candidate pools, the multi\-scenario xRouteBench benchmark, and an open\-source library that implements more than 16 routers behind a unified interface and supports deployment to real users and multi\-agent systems\. We hopeLLMRouterwill serve as a common foundation for developing, evaluating, and deploying LLM routers\.

## References

- A\. Adcock, A\. Srivastava, A\. Dubey, A\. Jauhri, A\. Pande, A\. Pandey, A\. Sharma, A\. Kadian, A\. Kumawat, A\. Kelsey,et al\.\(2026\)The llama 4 herd: architecture, training, evaluation, and deployment notes\.arXiv preprint arXiv:2601\.11659\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- S\. Agarwal, L\. Ahmad, J\. Ai, S\. Altman, A\. Applebaum, E\. Arbus, R\. K\. Arora, Y\. Bai, B\. Baker, H\. Bao,et al\.\(2025\)Gpt\-oss\-120b & gpt\-oss\-20b model card\.arXiv preprint arXiv:2508\.10925\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- P\. Aggarwal, A\. Madaan, A\. Anand, S\. P\. Potharaju, S\. Mishra, P\. Zhou, A\. Gupta, D\. Rajagopal, K\. Kappaganthu, Y\. Yang,et al\.\(2024\)Automix: automatically mixing language models\.Advances in Neural Information Processing Systems37,pp\. 131000–131034\.Cited by:[§B\.2](https://arxiv.org/html/2608.06867#A2.SS2.p4.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p1.15),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le,et al\.\(2021\)Program synthesis with large language models\.arXiv preprint arXiv:2108\.07732\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- E\. A\. M\. Callahan, A\. Chaluvaraju, A\. Gordić, D\. Gupta, Y\. Jain, P\. Monk, M\. Pust, T\. Romanski, P\. Rushton, A\. Shehper, D\. Shivaprasad, S\. Srivastava, A\. Thomas, A\. Tripathy, A\. Velingker, and A\. Vaswani \(2026\)Rnj\-1\-5\-Instruct\.Note:Long\-context Instruction\-tuned model releaseExternal Links:[Link](https://huggingface.co/EssentialAI/rnj-1-5-instruct)Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- Frugalgpt: how to use large language models while reducing cost and improving performance\.arXiv preprint arXiv:2305\.05176\.Cited by:[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p1.15),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- M\. Chen \(2021\)Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- S\. Chen, W\. Jiang, B\. Lin, J\. Kwok, and Y\. Zhang \(2024\)Routerdc: query\-based router by dual contrastive learning for assembling large language models\.Advances in Neural Information Processing Systems37,pp\. 66305–66328\.Cited by:[§B\.2](https://arxiv.org/html/2608.06867#A2.SS2.p3.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p3.2),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p5.3),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- C\. Clark, K\. Lee, M\. Chang, T\. Kwiatkowski, M\. Collins, and K\. Toutanova \(2019\)Boolq: exploring the surprising difficulty of natural yes/no questions\.InProceedings of the 2019 conference of the north American chapter of the association for computational linguistics: Human language technologies, volume 1 \(long and short papers\),pp\. 2924–2936\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. Tafjord \(2018\)Think you have solved question answering? try arc, the ai2 reasoning challenge\.arXiv:1803\.05457v1\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- X\. Dai, J\. Li, X\. Liu, A\. Yu, and J\. Lui \(2024\)Cost\-effective online multi\-llm selection with versatile reward models\.arXiv preprint arXiv:2405\.16587\.Cited by:[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- Z\. Dai, T\. Feng, and J\. You \(2025\)PersonalizedRouter: personalized llm routing via graph\-based user preference modeling\.arXiv preprint arXiv:2511\.16883\.Cited by:[§B\.4](https://arxiv.org/html/2608.06867#A2.SS4.p1.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- Deep Cogito \(2025\)Cogito v2 preview: deepseek 671b moe\.Note:[https://huggingface\.co/deepcogito/cogito\-v2\-preview\-deepseek\-671B\-MoE](https://huggingface.co/deepcogito/cogito-v2-preview-deepseek-671B-MoE)Hugging Face model card\. Accessed: 2026\-07\-28Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- D\. Ding, A\. Mallick, C\. Wang, R\. Sim, S\. Mukherjee, V\. Ruhle, L\. V\. Lakshmanan, and A\. H\. Awadallah \(2024\)Hybrid llm: cost\-efficient and quality\-aware query routing\.arXiv preprint arXiv:2404\.14618\.Cited by:[§B\.2](https://arxiv.org/html/2608.06867#A2.SS2.p4.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p1.15),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- T\. Feng, Y\. Shen, and J\. You \(2024\)Graphrouter: a graph\-based router for llm selections\.arXiv preprint arXiv:2410\.03834\.Cited by:[§B\.2](https://arxiv.org/html/2608.06867#A2.SS2.p3.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p1.15),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p3.2),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- T\. Feng, H\. Zhang, Z\. Lei, P\. Han, and J\. You \(2026\)GraphPlanner: graph memory\-augmented agentic routing for multi\-agent llms\.arXiv preprint arXiv:2604\.23626\.Cited by:[Appendix E](https://arxiv.org/html/2608.06867#A5.p1.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§5\.4](https://arxiv.org/html/2608.06867#S5.SS4.p3.1),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- E\. Frick, C\. Chen, J\. Tennyson, T\. Li, W\. Chiang, A\. N\. Angelopoulos, and I\. Stoica \(2025\)Prompt\-to\-leaderboard\.arXiv preprint arXiv:2502\.14855\.Cited by:[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- T\. Ge, X\. Chan, X\. Wang, D\. Yu, H\. Mi, and D\. Yu \(2024\)Scaling synthetic data creation with 1,000,000,000 personas\.arXiv preprint arXiv:2406\.20094\.Cited by:[§A\.6](https://arxiv.org/html/2608.06867#A1.SS6.p2.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2020\)Measuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt \(2021\)Measuring mathematical problem solving with the math dataset\.arXiv preprint arXiv:2103\.03874\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- Q\. J\. Hu, J\. Bieker, X\. Li, N\. Jiang, B\. Keigwin, G\. Ranganath, K\. Keutzer, and S\. K\. Upadhyay \(2024\)Routerbench: a benchmark for multi\-llm routing system\.arXiv preprint arXiv:2403\.12031\.Cited by:[§1](https://arxiv.org/html/2608.06867#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§2\.2](https://arxiv.org/html/2608.06867#S2.SS2.p1.1),[§3](https://arxiv.org/html/2608.06867#S3.p1.1),[§6](https://arxiv.org/html/2608.06867#S6.p2.1)\.
- Z\. Huang, B\. Lin, J\. Zhang, J\. Wang, Y\. Liu, N\. Lu, T\. Li, and X\. Huang \(2025a\)VL\-routerbench: a benchmark for vision\-language model routing\.arXiv preprint arXiv:2512\.23562\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p1.1),[§6](https://arxiv.org/html/2608.06867#S6.p2.1)\.
- Z\. Huang, G\. Ling, Y\. Lin, Y\. Chen, S\. Zhong, H\. Wu, and L\. Lin \(2025b\)Routereval: a comprehensive benchmark for routing llms to explore model\-level scaling up in llms\.arXiv preprint arXiv:2503\.10657\.Cited by:[§1](https://arxiv.org/html/2608.06867#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.06867#S2.SS2.p1.1),[§3](https://arxiv.org/html/2608.06867#S3.p1.1),[§6](https://arxiv.org/html/2608.06867#S6.p2.1)\.
- G\. Izacard, M\. Caron, L\. Hosseini, S\. Riedel, P\. Bojanowski, A\. Joulin, and E\. Grave \(2022\)Unsupervised dense information retrieval with contrastive learning\.Transactions on Machine Learning Research\.Cited by:[§A\.2](https://arxiv.org/html/2608.06867#A1.SS2.p2.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. d\. l\. Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier,et al\.\(2023\)Mistral 7b\.arXiv preprint arXiv:2310\.06825\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Roux, A\. Mensch, B\. Savary, C\. Bamford, D\. S\. Chaplot, D\. d\. l\. Casas, E\. B\. Hanna, F\. Bressand,et al\.\(2024\)Mixtral of experts\.arXiv preprint arXiv:2401\.04088\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- Y\. Li \(2025\)Rethinking predictive modeling for llm routing: when simple knn beats complex learned routers\.arXiv preprint arXiv:2505\.12601\.Cited by:[§B\.2](https://arxiv.org/html/2608.06867#A2.SS2.p2.1),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.\(2024\)Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- K\. Lu, H\. Yuan, R\. Lin, J\. Lin, Z\. Yuan, C\. Zhou, and J\. Zhou \(2024a\)Routing to the expert: efficient reward\-guided ensemble of large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 1964–1974\.Cited by:[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- P\. Lu, H\. Bansal, T\. Xia, J\. Liu, C\. Li, H\. Hajishirzi, H\. Cheng, K\. Chang, M\. Galley, and J\. Gao \(2024b\)Mathvista: evaluating mathematical reasoning of foundation models in visual contexts\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 23439–23554\.Cited by:[§A\.4](https://arxiv.org/html/2608.06867#A1.SS4.p1.1),[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- P\. Lu, R\. Gong, S\. Jiang, L\. Qiu, S\. Huang, X\. Liang, and S\. Zhu \(2021\)Inter\-gps: interpretable geometry problem solving with formal language and symbolic reasoning\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 6774–6786\.Cited by:[§A\.4](https://arxiv.org/html/2608.06867#A1.SS4.p1.1),[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- H\. Ma, G\. Lai, and H\. Ye \(2026\)MMR\-bench: a comprehensive benchmark for multimodal llm routing\.arXiv preprint arXiv:2601\.17814\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p1.1),[§6](https://arxiv.org/html/2608.06867#S6.p2.1)\.
- A\. Maharana, D\. Lee, S\. Tulyakov, M\. Bansal, F\. Barbieri, and Y\. Fang \(2024\)Evaluating very long\-term conversational memory of llm agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13851–13870\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- T\. Mihaylov, P\. Clark, T\. Khot, and A\. Sabharwal \(2018\)Can a suit of armor conduct electricity? a new dataset for open book question answering\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 2381–2391\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- I\. Ong, A\. Almahairi, V\. Wu, W\. Chiang, T\. Wu, J\. E\. Gonzalez, M\. W\. Kadous, and I\. Stoica \(2024\)Routellm: learning to route llms with preference data\.arXiv preprint arXiv:2406\.18665\.Cited by:[§B\.2](https://arxiv.org/html/2608.06867#A2.SS2.p3.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p1.15),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p3.2),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p5.3),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1),[§6](https://arxiv.org/html/2608.06867#S6.p2.1)\.
- OpenClaw \(2026\)OpenClaw: personal AI assistant\.Note:[https://github\.com/openclaw/openclaw](https://github.com/openclaw/openclaw)GitHub repository, accessed August 3, 2026Cited by:[§1](https://arxiv.org/html/2608.06867#S1.p3.1),[§4](https://arxiv.org/html/2608.06867#S4.p7.1)\.
- P\. Rajpurkar, J\. Zhang, K\. Lopyrev, and P\. Liang \(2016\)Squad: 100,000\+ questions for machine comprehension of text\.InProceedings of the 2016 conference on empirical methods in natural language processing,pp\. 2383–2392\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- T\. Shnitzer, A\. Ou, M\. Silva, K\. Soule, Y\. Sun, J\. Solomon, N\. Thompson, and M\. Yurochkin \(2023\)Large language model routing with benchmark datasets\.arXiv preprint arXiv:2309\.15789\.Cited by:[§B\.2](https://arxiv.org/html/2608.06867#A2.SS2.p2.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- G\. A\. Sigurdsson, A\. Gupta, C\. Schmid, A\. Farhadi, and K\. Alahari \(2018\)Charades\-ego: a large\-scale dataset of paired third and first person videos\.arXiv preprint arXiv:1804\.09626\.Cited by:[Figure 9](https://arxiv.org/html/2608.06867#A1.F9.12.1),[Figure 9](https://arxiv.org/html/2608.06867#A1.F9.13.1),[§A\.5](https://arxiv.org/html/2608.06867#A1.SS5.p1.1),[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- D\. Stripelis, Z\. Xu, Z\. Hu, A\. D\. Shah, H\. Jin, Y\. Yao, J\. Zhang, T\. Zhang, S\. Avestimehr, and C\. He \(2024\)Tensoropera router: a multi\-model router for efficient llm inference\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 452–462\.Cited by:[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- A\. Talmor, J\. Herzig, N\. Lourie, and J\. Berant \(2019\)Commonsenseqa: a question answering challenge targeting commonsense knowledge\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),pp\. 4149–4158\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière,et al\.\(2025\)Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.Cited by:[§A\.3](https://arxiv.org/html/2608.06867#A1.SS3.p2.1)\.
- G\. Team, M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé,et al\.\(2024\)Gemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- X\. Wang, Y\. Liu, W\. Cheng, X\. Zhao, Z\. Chen, W\. Yu, Y\. Fu, and H\. Chen \(2025\)Mixllm: dynamic routing in mixed large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 10912–10922\.Cited by:[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang, T\. Li, M\. Ku, K\. Wang, A\. Zhuang, R\. Fan, X\. Yue, and W\. Chen \(2024\)MMLU\-pro: a more robust and challenging multi\-task language understanding benchmark\.arXiv preprint arXiv:2406\.01574\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- D\. Wu, H\. Wang, W\. Yu, Y\. Zhang, K\. Chang, and D\. Yu \(2024\)Longmemeval: benchmarking chat assistants on long\-term interactive memory\.arXiv preprint arXiv:2410\.10813\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- E\. Xie, Y\. Sun, T\. Feng, and J\. You \(2025\)GMTRouter: personalized llm router over multi\-turn user interactions\.arXiv preprint arXiv:2511\.08590\.Cited by:[§B\.4](https://arxiv.org/html/2608.06867#A2.SS4.p1.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei,et al\.\(2024\)Qwen2\. 5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p2.1)\.
- F\. Yu, T\. Feng, D\. Min, L\. Cheng, G\. Liu, and T\. Zhou \(2026a\)TSRouter: dynamic modality\-model selection for time series reasoning\.arXiv preprint arXiv:2607\.08940\.Cited by:[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p1.15),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- F\. Yu, X\. Guo, L\. Yuan, H\. Kang, H\. Zhao, L\. Qin, F\. Huang, B\. Hu, and T\. Zhou \(2026b\)TSRBench: a comprehensive multi\-task multi\-modal time series reasoning benchmark for generalist models\.arXiv preprint arXiv:2601\.18744\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. Choi \(2019\)Hellaswag: can a machine really finish your sentence?\.InProceedings of the 57th annual meeting of the association for computational linguistics,pp\. 4791–4800\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- H\. Zhang, T\. Feng, and J\. You \(2025\)Router\-r1: teaching llms multi\-round routing and aggregation via reinforcement learning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,Cited by:[§B\.3](https://arxiv.org/html/2608.06867#A2.SS3.p1.1),[§1](https://arxiv.org/html/2608.06867#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p1.15),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p2.5),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p3.2),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p4.12),[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p5.3),[§5\.1](https://arxiv.org/html/2608.06867#S5.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§3](https://arxiv.org/html/2608.06867#S3.p1.1),[§3](https://arxiv.org/html/2608.06867#S3.p3.1)\.
- K\. Zhu, H\. Du, Z\. Hong, X\. Yang, S\. Guo, D\. Z\. Wang, Z\. Wang, C\. Qian, R\. Tang, H\. Ji,et al\.\(2025\)Multiagentbench: evaluating the collaboration and competition of llm agents\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8580–8622\.Cited by:[Appendix E](https://arxiv.org/html/2608.06867#A5.p1.1),[§5\.4](https://arxiv.org/html/2608.06867#S5.SS4.p3.1)\.
- R\. Zhuang, T\. Wu, Z\. Wen, A\. Li, J\. Jiao, and K\. Ramchandran \(2025\)Embedllm: learning compact representations of large language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 76913–76926\.Cited by:[§2\.1](https://arxiv.org/html/2608.06867#S2.SS1.p3.2),[§6](https://arxiv.org/html/2608.06867#S6.p1.1)\.

Appendix

## Appendix ABenchmark Details

Table[6](https://arxiv.org/html/2608.06867#A1.T6)lists the eight test sets of xRouteBench, their categories, sizes, and metrics, totaling 4,767 test queries\. The Generic LLM Tasks track further decomposes into 13 subtasks, and the memory datasets are retrieved with top\-5 RAG over turn pairs\.

Table 6:The eight test sets of xRouteBench\.Sizes are the number of test queries; metrics are exact match \(EM\), multiple\-choice accuracy \(MC\), token\-level F1, execution\-based code pass rate, math answer matching, and a persona\-conditioned LLM judge\.CategoryTest setContent\#TestMetricGeneric LLM TasksGeneric mix13 subtasks3,729EM/MC/F1/GSM8K/MATH/codeMemoryLoCoMolong\-conversation QA314F1LongMemEvallong\-term memory QA101F1TimeSeriesTimeSeries7 reasoning skills127MCVisionGeometry3Kgeometry math \(image\)61EMMathVistavisual math reasoning100EM/MCCharades\-Egoegocentric video27EMPersonalizedChatbot Arena / MT\-Benchpreference prompts308LLM judgeTotal4,767

Table[7](https://arxiv.org/html/2608.06867#A1.T7)details the 13 subtasks that make up the Generic LLM Tasks track\.

Table 7:Composition of the Generic LLM Tasks track\.The table lists the 13 subtasks, their target skills, and the number of test queries, totaling 3,729 examples\.SubtaskSkill\#TestMBPPcode generation500MATHmathematical reasoning500GSM8Kmathematical reasoning500MMLU\-Proknowledge QA500OpenBookQAknowledge QA500ARC\-Challengeknowledge QA500MMLUknowledge QA500CommonsenseQAcommonsense QA50BoolQcommonsense QA50SQuADreading comprehension50HellaSwagcommonsense QA50HumanEvalcode generation16AIME \(2020–2024\)competition math13### A\.1Generic LLM Tasks

The Generic LLM Tasks track is deliberately a mixture rather than a single task family\. It places knowledge questions, commonsense inference, reading comprehension, mathematical reasoning, and code generation behind the same routing interface\. All examples use text input, but an ARC\-Challenge question, an AIME problem, and a Python synthesis task reward very different capabilities and impose different output constraints\. The track therefore asks whether a router can recognize these distinctions within an apparently uniform text setting, rather than defaulting to one model for every natural language request\. Table[7](https://arxiv.org/html/2608.06867#A1.T7)lists the 13 source benchmarks and their test sizes\.

We serialize each example as one fixed text query before collecting candidate responses\. A system instruction specific to the task states the required response format, and the user problem follows it\. Choice items include their options\. Mathematics items request a final answer in the expected form\. Code generation items provide the programming task and tests, while SQuAD supplies its passage and question\. We score each response with the source benchmark’s native objective\. We use choice accuracy for knowledge and commonsense tasks, token\-level F1 for SQuAD, answer matching for GSM8K, MATH, and AIME, and execution\-based pass rates for MBPP and HumanEval\.

### A\.2Long Context Conversational Memory

The Memory track fixes retrieval and asks a narrower routing question\. Given the same retrieved conversational evidence, which model is most likely to use it correctly? LoCoMo and LongMemEval contain questions whose supporting facts may be separated from the query by many turns or sessions\. Some can be answered by recovering one explicit fact, whereas others require combining facts across sessions or resolving a later update against an earlier statement\. This distinction is useful for routing because it separates the cost of carrying a long history from the ability to interpret the evidence selected from it\.

We construct every memory query through the same retrieval pipeline\. Each conversation history is divided into adjacent turn pairs that retain speaker and date information\. A fixed Contriever\(Izacardet al\.,[2022](https://arxiv.org/html/2608.06867#bib.bib346)\)encoder embeds the question and the turn pairs, and the five most similar pairs are inserted into a brief answer prompt\. LongMemEval additionally includes the date of the question, which provides the temporal reference needed for its time\-sensitive items\. Every candidate therefore receives the same retrieved evidence for a query\. Differences in score reflect its use of that evidence rather than a different retrieval result\. We evaluate both datasets with token\-level F1\.

### A\.3Time Series Pattern Reasoning

The TimeSeries track is built from TSRBench and covers anomaly detection, similarity analysis, noise understanding, pattern recognition, inductive reasoning, causality analysis, and event prediction\. These problems are numerical, but they often require recognizing structure that is easier to see as a shape than as a list of values\. A local spike may signal an anomaly, while periodicity, changes in trend, or agreement between two series emerge over a longer range\. The track asks whether a router can distinguish the models that handle these different forms of temporal reasoning well\.

We convert every series into a common text query before collecting candidate responses\. First, we render each series as a line chart and pass the chart to a fixed Gemma\-3\-27B\-IT\(Teamet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib260)\)captioner, which produces a paragraph description of its visible temporal pattern\. We append these descriptions and the raw numerical values to the original question and its answer options\. The raw values preserve exact numerical evidence and are truncated after 200 values when necessary, while the descriptions make higher\-level structure explicit\. The charts are used only in this offline captioning step\. Every candidate model receives the same text query rather than an image\. Models return an option letter, and we score the responses with multiple\-choice accuracy\.

### A\.4Visual Mathematical Reasoning

The Visual Mathematical Reasoning track routes image\-grounded mathematics drawn from Geometry3K\(Luet al\.,[2021](https://arxiv.org/html/2608.06867#bib.bib298)\)and MathVista\(Luet al\.,[2024b](https://arxiv.org/html/2608.06867#bib.bib308)\), where a geometry diagram or a scientific figure carries information the question depends on\. The strongest candidate on these problems can differ from the strongest on text mathematics, so the track probes whether a router follows per\-model strength as the query type changes\. We render each problem to text before it reaches the pool, describing its image with a fixed vision\-language model and appending the description to the question, so every candidate reasons over one identical query\. Rendering the image once holds perception constant across the pool and turns visual mathematics into a query the whole pool can answer, so the track scores the reasoning quality of each candidate on the same input\. These rendered queries read as machine\-written descriptions of a figure, a distribution that departs from the natural questions the routers are trained on, so the track also tests how well a router generalizes when the query type shifts\.

A single frozen vision\-language model, Gemma\-3\-27B\-IT by default, produces every description, so all candidates see the same rendering of a given image\. Its prompt asks it to report every visible number, symbol, angle, length, and relationship and to withhold the solution, which keeps the answer out of the query and leaves the reasoning to the routed model\. The returned text is appended to the original problem, from which Geometry3K’s image placeholder is removed, to form the routed query\. Geometry3K is graded by math answer matching against its numeric answer, and MathVista is graded by exact match on its open\-ended items and by multiple\-choice accuracy on its choice items, following the question\-type field of each query\. Every query keeps its ground\-truth answer and, where present, its answer choices, so the scored responses populate the same query–model matrix of quality and cost that supervises every other track \(§[2\.2](https://arxiv.org/html/2608.06867#S2.SS2)\)\. Figure[8](https://arxiv.org/html/2608.06867#A1.F8)shows one example from each dataset with the description Gemma\-3\-27B\-IT produces for it\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/geometry3k.png)

\(a\) Geometry3K, which asks forxx\.

Gemma\-3\-27B\-IT descriptionThe diagram shows a circle with two intersecting chords\. One chord is divided into segments of length44and66, and the other into segments of lengthxxand88\. The labelxxmarks one of these segments, and the two chords cross at a point inside the circle\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/mathvista.png)

\(b\) MathVista, which asks for the spring compressiondd\.

Gemma\-3\-27B\-IT descriptionA block of massmmslides to the left along a horizontal frictionless surface with velocityvvand approaches a coiled spring of constantkkfixed to a wall\. Labels mark where the block first touches the spring and where it stops, anddddenotes the distance between these two points\. A caption states that the spring force does negative work, decreasing speed and kinetic energy\.

Figure 8:One example from each dataset in the Visual Reasoning track, shown with the description that Gemma\-3\-27B\-IT produces for its image\.Geometry3K \(a\) provides a geometry diagram, and MathVista \(b\) provides a scientific figure\. Each description is appended to the problem text to form the query the router sees\.
### A\.5Multi\-View Video Recognition

The video track builds on Charades\-Ego\(Sigurdssonet al\.,[2018](https://arxiv.org/html/2608.06867#bib.bib307)\), which records each activity simultaneously from a first\-person egocentric camera worn by the actor and a third\-person exocentric camera facing the scene \(Figure[9](https://arxiv.org/html/2608.06867#A1.F9)\)\. The two viewpoints carry different evidence, and a deployment often captures only one of them, so the routing decision here depends on which views a query provides\. We describe every available view in text and route the merged description, so the whole candidate pool answers the same query whether one view or both are present\. This makes the video track the regime in xRouteBench where the available modality varies from query to query, and it tests whether a router still selects the right model when the visual evidence is partial\.

We build three classification tasks from the annotations, predicting the activity, the verb, or the object of the depicted action, each answered by a compact identifier drawn from the task label inventory, with the verb and object labels recovered from the Charades action mapping\. For each paired clip we match the two views by their shared identifier and select the action segment whose egocentric and exocentric occurrences overlap most closely in time, which keeps both views on the same moment, and we then keep the first\-person view, the third\-person view, or both at random to populate the single\-view and dual\-view regimes\. From each retained view we sample frames inside the aligned window and pass them to the vision\-language model, which returns a structured description of the actor’s motion, the handled objects, and the scene\. These descriptions are merged into one query that states the task, lists the candidate identifiers with their names, and requests an identifier as the answer, holding every candidate to the same text input\.

time→\\rightarrowFirst\-person

![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/ego_t1.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/ego_t2.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/ego_t3.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/ego_t4.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/ego_t5.jpg)Third\-person

![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/exo_t1.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/exo_t2.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/exo_t3.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/exo_t4.jpg)![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/charades/exo_t5.jpg)
Gemma\-3\-27B\-IT descriptionPlaceholder for the structured description the vision\-language model produces from the available views\.

Figure 9:A multi\-view sample from Charades\-Ego\(Sigurdssonet al\.,[2018](https://arxiv.org/html/2608.06867#bib.bib307)\), recorded from a first\-person egocentric camera \(top\) and a third\-person exocentric camera \(bottom\)\.Each row is a viewpoint and each column is a sampled time step\. Our transformation samples these frames, describes the available views with a vision\-language model, and routes the resulting text query, randomly withholding one view to form single\-view and dual\-view queries\.
### A\.6Personalized Dialogue Preference

The Personalized track combines open\-ended dialogue from MT\-Bench and Chatbot Arena\. Each query retains its original message sequence, including preceding system, user, and assistant messages\. MT\-Bench contributes instruction following conversations in which a later turn depends on an earlier exchange, while Chatbot Arena contributes natural user requests across diverse conversational settings\. Rather than assuming that one answer is universally better, the track constructs supervision from the preference that a particular user profile would express after seeing two answers\.

The user profiles used for preference elicitation are drawn from PersonaHub\(Geet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib343)\)\. We sample 200 personas from this resource for data collection\. Figure[10](https://arxiv.org/html/2608.06867#A1.F10)presents ten examples from the full persona set used during data collection\.

Ten example personas used to condition the judge1\.A 71\-year\-old retired nurse from Italy, volunteering in hospice care and advocating for compassionate end\-of\-life support\.6\.A 34\-year\-old scientist from London who is a social media influencer\.2\.A 54\-year\-old divorced mother from Spain, running a successful winery and promoting sustainable viticulture practices\.7\.A 41\-year\-old scientist from London who loves hiking\.3\.A 63\-year\-old retired teacher from China, teaching calligraphy and preserving the art form for future generations\.8\.An 87\-year\-old World War II veteran from Poland, sharing stories of his experiences and advocating for peace\.4\.A 68\-year\-old retired engineer from Japan, practicing ikebana and teaching the art to younger generations\.9\.A 31\-year\-old social worker from Colombia, supporting victims of domestic violence and fighting for gender equality\.5\.A 21\-year\-old photographer from Paris who spends weekends volunteering\.10\.A 23\-year\-old aspiring musician from Brazil, fusing traditional and modern sounds and promoting cultural exchange through music\.

Figure 10:Ten examples from the 200 PersonaHub personas sampled for preference collection\.For each comparison, DeepSeek\-V3\.1 receives one selected persona as its role specification when judging the two candidate answers\. The candidate models do not receive the persona\.For each dialogue, we sample two models at random from the candidate pool of 18 models and generate one response from each under the same conversation history\. DeepSeek\-V3\.1 is conditioned on the selected persona to act as the judge, comparing the two answers and returning a preference for either response or a tie\. The persona is supplied to the judge only, not to either candidate model\. Each comparison is converted into a pairwise supervision record whose label is the final judged user preference\. The router learns from these preference outcomes rather than from the persona description itself\. This construction keeps the candidate responses comparable while making the routing target sensitive to the user utility represented by the judge\.

## Appendix BRouter Details

The routers inLLMRoutershare the candidate pool and task interface described above, but they expose different information to the routing decision\. Rule\-based baselines select from a fixed property of the pool, while single\-turn methods make one model selection from the query and logged routing outcomes\. Multi\-turn methods additionally condition on intermediate calls and may decide whether further routing is useful\. Personalized methods add a user and interaction history, so their target is the model a particular user is most likely to prefer rather than one model that is best on average\. Table[B](https://arxiv.org/html/2608.06867#A2)summarizes the 17 built\-in implementations along these differences\.

Table 8:Built\-in routers inLLMRouter, grouped by routing family\.For each method, the table summarizes the information available to the routing decision \(State\) and the rule used to select or aggregate candidate models \(Selection\)\.RouterStateSelection\\rowcolorbarGrayRule\-based baselinesSmallest\-LLMcandidate parameter countsalways selects the smallest candidateLargest\-LLMcandidate parameter countsalways selects the largest candidate\\rowcolorbarGraySingle\-turn routerskNNRouterquery embedding and nearby logged queriesvotes over the models preferred by nearest neighborsSVMRouterquery embeddingkernel classifier predicts a candidateMLPRouterquery embeddingMLP classifier predicts a candidateMFRouterquery and model latent factorsranks candidates by their interaction scoreEloRouterlogged pairwise model outcomesalways selects the highest\-rated candidateRouterDCquery and candidate representationscontrastive query–model matching scoreHybrid LLMquery embedding and a small/large model pairpredicts whether the small model is sufficientAutoMixsmall\-model draft and verification signalaccepts the draft or escalates to the large modelGraphRouterquery–model interaction graphpredicts performance on query–model edgesCausalLM Routertextual query and candidate listgenerates the selected model name\\rowcolorbarGrayMulti\-turn routersRouter\-R1query and accumulated search resultsiteratively searches specialists or terminates and aggregateskNN\-MultiRoundsub\-queries and their embeddingsroutes each sub\-query with kNN and aggregates the answersLLM\-MultiRoundtextual query, decomposition, and candidate listan LLM chooses routes for sub\-queries and aggregates\\rowcolorbarGrayPersonalized routersGMTRouteruser, session, query, model, and response interactionspredicts user\-conditioned model preferencePersonalizedRouteruser features, task description, query, and modelpredicts preference for a user–query pair
### B\.1Rule\-Based Baselines

Smallest\-LLM and Largest\-LLM provide fixed reference points for the learned routers\. They ignore the query and always select the candidate with the smallest or largest declared parameter count, respectively\. These rules make no attempt to identify per\-query model strengths, but they expose the two simple deployment policies against which query\-aware routing should be compared: consistently favoring the smallest available model or consistently favoring the largest one\.

### B\.2Single\-Turn Routers

Single\-turn routers terminate after one model selection, so their differences lie in how they represent a query and score candidates\. EloRouter is query\-independent, but it replaces a fixed parameter\-count rule with a global ranking estimated from pairwise outcomes in the routing data\. It therefore captures which model is strongest on average while deliberately discarding the variation between individual queries\.

kNNRouter, SVMRouter, and MLPRouter instead make the selection query\-specific from its embedding\. kNNRouter retrieves similar training queries and transfers their observed best\-model choices through a vote or distance\-weighted vote, requiring no fitted decision function beyond the stored examples\(Li,[2025](https://arxiv.org/html/2608.06867#bib.bib332); Shnitzeret al\.,[2023](https://arxiv.org/html/2608.06867#bib.bib76)\)\. SVMRouter and MLPRouter fit discriminative boundaries over the same embedding space, using a kernel classifier and a multilayer perceptron, respectively\. MFRouter takes a different view of the logged query–model matrix: it learns latent representations for both sides and selects the model with the strongest query–model interaction\. These methods all learn from per\-query outcomes, but differ in whether the candidate is represented by neighboring solved examples, a classifier label, or a learned latent factor\.

The remaining single\-turn methods enrich this compatibility score in different ways\. RouterDC learns query–model matching with dual contrastive objectives over query–model, query–query, and cluster\-level relations\(Chenet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib295)\)\. GraphRouter instead passes information over a query–model interaction graph and predicts the quality of an unobserved query–model edge\(Fenget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib304)\)\. CausalLM Router verbalizes the query and the available candidates, then fine\-tunes a causal language model to generate the selected model name\(Onget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib294)\)\. In this case, the query representation, candidate representation, and selection score are combined within the language model rather than implemented as separate embedding modules\.

Hybrid LLM and AutoMix are cost\-aware two\-model cascades\. Hybrid LLM learns from the quality gap between the smallest and largest candidates, then uses a thresholded prediction to decide whether the smaller model is sufficient\(Dinget al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib339)\)\. AutoMix first obtains a draft from the smaller model and a verification signal for that draft; it retains the draft when the signal is reliable and otherwise escalates the query to the larger model\(Aggarwalet al\.,[2024](https://arxiv.org/html/2608.06867#bib.bib335)\)\. We list both methods with the single\-turn family because they return one final model answer through a fixed cascade, rather than repeatedly choosing among arbitrary candidates and aggregating an open\-ended history\. Their extra calls are nevertheless part of the inference cost\.

### B\.3Multi\-Turn Routers

Multi\-turn routers treat intermediate responses as part of the routing state\. Router\-R1 is a pretrained routing agent that reasons over the query and the results gathered so far\. At each step it can issue a<search\>call to a suitable specialist, incorporate the returned evidence, or terminate and produce an aggregate answer; its policy is trained with trajectory\-level reinforcement learning rather than pointwise best\-model labels\(Zhanget al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib296)\)\. The cost of its reasoning, search, and final aggregation calls is included in the routed trajectory\.

The two multi\-round baselines use decomposition more explicitly\. kNN\-MultiRound first breaks a complex query into a small set of sub\-queries, applies the kNN routing rule to each one, and combines the resulting answers\. LLM\-MultiRound uses an LLM to express both the decomposition and the model choice in text, then executes the selected sub\-queries and aggregates their outputs\. Both methods can assign different candidates to different parts of a problem, but they also introduce decomposition and aggregation calls that a one\-shot router does not pay for\.

### B\.4Personalized Routers

Personalized routers replace the single global notion of quality with a user\-conditioned preference\. GMTRouter represents users, sessions, queries, candidate models, and responses as nodes in a heterogeneous interaction graph, allowing a decision for the current query to draw on earlier interactions from the same user\(Xieet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib317)\)\. PersonalizedRouter similarly uses a graph\-based scorer, while explicitly incorporating user features and task descriptions alongside the query and candidate representations\(Daiet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib319)\)\. Both methods are trained from comparisons between candidate answers, so a high score means that a model is predicted to be preferred for this user and query, not simply that it has the highest population\-average task score\.

## Appendix CCandidate Pool and Pricing

Table[9](https://arxiv.org/html/2608.06867#A3.T9)lists the 18 candidate models and their per\-token prices\. Prices are in USD per 1M tokens, with input and output priced separately; the blended average is their mean\. The input price spans a roughly25×25\\timesrange, from $0\.05 to $1\.25 per 1M tokens\. The 17\-model no\-cogito pool drops the most expensive model \(cogito\-v2\-1\-671b\)\.

Table 9:The 18 candidate LLMs, sorted by blended average price\.Prices in USD per 1M tokens\.\#ModelParamsInputOutputService1gemma\-2\-9b\-it9B0\.100\.10NVIDIA2llama\-3\-8b\-instruct\-lite8B0\.100\.10Together3gpt\-oss\-20b20B0\.050\.20Together4rnj\-1\-instruct15B0\.150\.15Together5mistral\-7b\-instruct\-v0\.37B0\.200\.20NVIDIA6mistral\-small\-3\-24b\-instruct24B0\.100\.30Together7qwen2\.5\-7b\-instruct7B0\.200\.20NVIDIA8qwen2\.5\-7b\-instruct\-turbo7B0\.300\.30Together9gpt\-oss\-120b120B0\.150\.60Together10llama\-4\-maverick402B0\.270\.85Together11mixtral\-8x7b\-instruct\-v0\.146\.7B0\.600\.60NVIDIA12qwen3\-next\-80b\-a3b\-instruct80B0\.151\.50Together13qwen3\-coder\-next200B0\.501\.20Together14llama\-3\.3\-70b\-instruct\-turbo70B0\.880\.88Together15llama3\-70b\-instruct70B0\.900\.90NVIDIA16deepseek\-v3\.1671B0\.601\.70Together17mixtral\-8x22b\-instruct\-v0\.1140\.6B1\.201\.20NVIDIA18cogito\-v2\-1\-671b671B1\.251\.25Together
## Appendix DHuman Preference Collection on Slack

The human preference dataset of §[5\.4](https://arxiv.org/html/2608.06867#S5.SS4)is collected through a Slack application built on the OpenClaw server ofLLMRouter, as illustrated in Figure[11](https://arxiv.org/html/2608.06867#A4.F11)\. A user asks a question directly in Slack, and the system samples two LLMs at random from the ten\-candidate pool and generates one answer with each\. The conversation view then presents the two responses as anonymized Answer A and Answer B, with their positions randomly shuffled, and the user clicks a button to mark A better, B better, or a tie\. In multi\-turn sessions the user follows up freely, and every turn is labeled in the same way\. Fifteen users contributed 40 sessions and 234 pairwise preference records in total\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/slack_platform.png)Figure 11:Slack interface for collecting pairwise user preferences\.Two anonymized model responses are shown as Answer A and Answer B in randomized order, and users select A, B, or a tie for each interaction turn\.
## Appendix EMulti\-Agent Topologies

Table[10](https://arxiv.org/html/2608.06867#A5.T10)summarizes the five coordination topologies used in §[5\.4](https://arxiv.org/html/2608.06867#S5.SS4)\. Star, Tree, Graph, and Chain follow MultiAgentBench\(Zhuet al\.,[2025](https://arxiv.org/html/2608.06867#bib.bib322)\), and Plan\-Exec\-Sum follows the static router design of GraphPlanner\(Fenget al\.,[2026](https://arxiv.org/html/2608.06867#bib.bib19)\)with width 3 and depth 1\. Star and Plan\-Exec\-Sum differ in two respects: in Star, the planner produces free\-form subtasks and consolidates the actors’ outputs itself, whereas in Plan\-Exec\-Sum the planner decomposes the query into three atomic sub\-queries and a dedicated summarizer merges the executors’ answers\. Every node is replaced by a router that receives the node’s prompt and selects a model from the candidate pool for that call\. We adopt the topology of GraphPlanner without training its planner, since the training requires live answers from candidate models that have since been retired\.

Table 10:The five multi\-agent topologies, their structures, and the number of LLM calls per query\.TopologyStructureLLM callsStarplanner decomposes→\\rightarrow3 actors in parallel→\\rightarrowplanner consolidates6Treeroot planner→\\rightarrow2 sub\-planners refine→\\rightarrow2 actors→\\rightarrowroot consolidates7Graph3 actors answer independently→\\rightarrowone full\-communication revision round7Chain3 agents relay sequentially, each verifying and improving the previous answer4Plan\-Exec\-Sumplanner emits 3 atomic sub\-queries→\\rightarrow3 executors→\\rightarrowsummarizer merges6

## Appendix FPrompt Usage

This section reports the fixed prompt templates used to construct and evaluate xRouteBench, as well as the templates used by routers that make language\-model calls internally\. Curly brackets denote instance\-dependent fields\. Each prompt is presented in a titled box, combining the fixed instruction and the instance fields in the order in which the model reads them, rather than splitting them into separate system and user blocks\. In the implementation, the fixed instruction is sent as a system message when the API supports that role; otherwise, the two parts are concatenated with\[System Instruction\]and\[User Query\]delimiters\. Thus, all candidate models receive the same logical prompt for a given query\.

### F\.1Benchmark\-Query Templates

#### Generic LLM Tasks\.

The 13 Generic LLM Tasks use seven task\-specific templates\. The multiple\-choice instruction is shared by MMLU, MMLU\-Pro, BoolQ, and HellaSwag; CommonsenseQA, OpenBookQA, and ARC\-Challenge use the same instruction but retain their source choice formatting\.

Standard multiple choiceAnswer the following multiple\-choice question by selecting the correct option \(A, B, C, or D\)\. You MUST put your final answer letter in a parenthesis\.Question: \{question\}Options: A\. \{choice\_1\} B\. \{choice\_2\} C\. \{choice\_3\} D\. \{choice\_4\}

Source\-formatted multiple choiceAnswer the following multiple\-choice question by selecting the correct option \(A, B, C, or D\)\. You MUST put your final answer letter in a parenthesis\.Question: \{question\} \(A\) \{choice\_1\} \(B\) \{choice\_2\} \(C\) \{choice\_3\} \.\.\.

GSM8KAnswer the following math question step by step\.Question: \{question\}

MATH and AIMEAnswer the following math question\. Make sure to put the answer \(and only answer\) inside \\boxed \{\}\.Question: \{question\}

MBPPYou are an expert Python programmer\. Implement the function with no irrelevant words or comments\. Put your code in this format: \[BEGIN\] <Your Code\> \[Done\]Task: \{task\_description\}Your code should pass these tests: \{tests\}

HumanEvalYou are an expert Python programmer\. Implement the function body only\. Do not repeat the function signature or docstring\. Put the code in this format: \[BEGIN\] <Your Code \- function body only\> \[Done\]Complete the following function: \{function\_signature\_and\_docstring\}

SQuADAnswer the question using the provided context\. If the answer is not in the context, output noanswer\.\{source query containing the context and question\}

#### Conversational memory\.

For both memory datasets, Contriever retrieves the top five turn\-pair chunks before the following prompt is formed\. The retrieval result is fixed across candidate models for each query\. LoCoMo uses:

LoCoMoYou answer questions about conversation history\. Reply with a short phrase only\.Below is relevant information from the conversation history:\{retrieved turn\-pair chunks\}Based on the above context, write an answer in the form of a short phrase for the following question\. Answer with exact words from the context whenever possible\.Question: \{question\} Short answer:

If no chunk is retrieved, the context field is replaced byNo relevant information available\.LongMemEval uses the analogous template, with its question date made explicit:

LongMemEvalYou answer questions based on chat history\. Reply with a short phrase only\.I will give you several history chats between you and a user\. Please answer the question based on the relevant chat history\.History Chats:\#\#\#\#\#\# Session \{index\}: Session Content: \{retrieved chunk\}\.\.\.Current Date: \{current\_date\}Question: \{question\}Short Answer:

Here, an empty history is represented byNo relevant chat history available\.

#### Visual mathematical reasoning\.

Geometry3K and MathVista first use one frozen vision\-language model to turn the image into text\. The captioning call is:

Image captioningDescribe the image for solving the problem\. Include all visible text, numbers, symbols, angles, lengths, and geometric relationships\. Be concise and factual\. Do not include introductory phrases\.\{image\}

The resulting caption is then given to every candidate model with the following answer prompt:

Visual mathematical reasoningYou are an expert in solving math and geometry problems\. Think step\-by\-step, and when you are ready to provide the final answer, you MUST enclose it in \\boxed \{\}\.For example, if the answer is 42, write \\boxed \{42\}\. If it is a multiple choice question and the option is C, write \\boxed \{C\}\.Image Description: \{caption\}Answer the following question based on the provided image description: \{question\}

#### Time\-series reasoning\.

We first render every series as an image and request a one\-paragraph caption\. The captioner receives the following system instruction and one instruction sampled uniformly from 20 equivalent phrasings \(the first phrasing is shown\):

Time\-series captioningYou are a time series captioner\.Write a paragraph that analyzes the time series\.\{time\-series plot\}

The other phrasings differ only in wording \(e\.g\., “Create a detailed description of the time series in one paragraph”\) and impose the same one\-paragraph description requirement\. The answer prompt combines the caption with the raw values:

Time\-series reasoningYou are a time series analysis expert\. Answer the multiple\-choice question by providing the correct option letter \(e\.g\., A, B, C, or D\)\. Be concise\.1\. Time Series Description: \(1\) \{series\_name\}: \{caption\}, \.\.\.2\. Raw time series: \(1\) \{series\_name\}: \[\{values\}\], \.\.\.based on the one of the following information to reason the final answer: \{original\_question\}

For perception and causality\-analysis items, this two\-source block replaces the<ts\><ts/\>marker in the source question; it is appended for the other two task types\.

#### Multi\-view video recognition\.

Five frames from each available egocentric or exocentric view are summarized independently before classification\. The video captioner receives:

Video captioningYou are given multiple frames sampled from a short time window of a video\.Return ONLY a JSON object with the following keys \(no extra text\):"motion": string,"objects": \[string\],"summary": stringDefinitions:motion: describe what the person is doing \(hands/arms/body\) in one short phrase\.objects: list the main target objects the person is acting on / interacting with \(not background clutter\)\.summary: a short description of the scene in one sentence\.\{video frames\}

Each candidate model then receives the two structured captions, the relevant label inventory, and this classification prompt:

Video classificationYou are a helpful AI assistant tasked with answering questions about videos based on descriptions\.Read the provided text carefully and return ONLY the ID \(e\.g\. c001, v003, o012\) representing the correct answer\.Task: Identify the \{activity/verb/object\} from video descriptions\.Determine the correct id using the structured descriptions below\.First\-person view description: \{ego\_caption\}Third\-person view description: \{exo\_caption\}Available \{activity/verb/object\}s: \{id\}: \{label\} \.\.\.Provide your answer as the ID \(e\.g\., c001, v003, o012\)\.

#### Personalized dialogue preference\.

For MT\-Bench, the original multi\-turn user/assistant message sequence is preserved; for Chatbot Arena, the original user message is preserved\. No additional fixed instruction is inserted before a candidate answers\. Two candidate responses are then compared by a judge conditioned on a sampled persona:

Persona preference judgeYou are simulating the following user persona: \{persona\}\.As this persona, evaluate which assistant response you would personally prefer\.\[User Question\] \{question\}\[Assistant 1\] \{answer\_1\}\[Assistant 2\] \{answer\_2\}Output ONLY one of the following tokens exactly: 1 2 Tie Do not output anything else\.

### F\.2Router\-Internal Templates

The preceding templates construct the query–model response matrix used by all routers\. Most single\-turn routers consume this matrix without another LLM prompt\. The following templates are used only by the routers that perform additional language\-model calls during routing or aggregation\.

#### kNN\-MultiRound and LLM\-MultiRound\.

Both multi\-round routers decompose a query, obtain sub\-answers, and aggregate them\. kNN\-MultiRound selects a model for each sub\-query with nearest\-neighbour lookup, whereas LLM\-MultiRound asks an LLM to choose the model\. Their shared decomposition template is:

Multi\-round decompositionGiven the query ’\{query\}’, decompose it into as many as 4 meaningful sub\-queries \(minimum 1, maximum 4\)\.Try to cover the full scope of the original query by breaking it down into multiple specific and distinct sub\-tasks whenever possible\.Aim for the maximum number of high\-quality sub\-queries without introducing redundancy\.Each sub\-query should be clear, self\-contained, and semantically coherent\.Output only the decomposed sub\-queries, one per line\. Do not include any other text, explanations, or headers in your output\. Each line should contain only one sub\-query\.

LLM\-MultiRound appends the candidate names and descriptions, then replaces the last three lines by the following routing constraint:

Multi\-round model selectionYou will then be provided with descriptions of the following Large Language Models \(LLMs\): \{model\_list\}\.\{model\_descriptions\}For each sub\-query, select the single LLM that is most likely to generate the highest\-quality response, regardless of cost or efficiency\.Focus entirely on maximizing effectiveness and providing the most accurate and relevant output\. Base your decision strictly on the descriptions of the models\.Output only the decomposed sub\-queries and the full name of the selected LLM for each\. Output exactly one sub\-query and one LLM per line\.Each line must be formatted as follows:<sub\-query\>: <LLM name\>Use a colon ’:’ as the separator\. Do not include any other text, explanations, or headers in your output\.

Each selected model is queried with:

Multi\-round specialistYou are a helpful assistant\. You are participating in a multi\-agent reasoning process, where a base model delegates sub\-questions to specialized models like you\.Your task is to do your absolute best to either:Answer the question directly, if possible, and provide a brief explanation;Offer helpful and relevant context, background knowledge, or insights related to the question, even if you cannot fully answer it\.If you are completely unable to answer the question or provide any relevant or helpful information, you must:Clearly state that you are unable to assist with this question\.Explicitly instruct the base model to consult other LLMs for further assistance\.Keep your response clear, concise, and informative \(preferably under 512 tokens\)\.Stay strictly on\-topic\.Do not include irrelevant or generic content\.Here is the sub\-question for you to assist with: \{sub\_query\}

For non\-multiple\-choice tasks, the aggregator receives:

Multi\-round aggregationYou are given a question along with auxiliary information, which consists of several sub\-questions derived from the original question and their respective answers\.Use this information to answer the original question if relevant, but make your own reasoning step by step before arriving at the final answer\.Important: Your final answer MUST be clearly marked and enclosed within <answer\> and </answer\> tags at the end of your response\. No other part of the output should be inside these tags\.Auxiliary Information: \{sub\_query\_and\_response\_pairs\}Question: \{query\}Let’s think step by step\.

For multiple\-choice tasks, it instead receives:

Multi\-round multiple\-choice aggregationYou are given a multiple\-choice question and supporting sub\-answers\. Use the information only if helpful\.Question: \{original\_query\}Supporting information: \{sub\_query\_and\_response\_pairs\}Rules:Select exactly one option: A, B, C, D, or E\.Output only the letter in <answer\> tags\.No explanation\.Format:<answer\>A</answer\>

#### Router\-R1\.

Router\-R1 appends descriptions of the available models to the following template\. Its agent can request a specialist model within<search\>tags; the returned text is inserted as<information\>in the next round\.

Router\-R1Answer the given question\.Every time you receive new information, you must first conduct reasoning inside <think\> \.\.\. </think\>\.After reasoning, if you find you lack some knowledge, you can call a specialized LLM by writing a query inside <search\> LLM\-Name:Your\-Query </search\>\.STRICT FORMAT RULES for <search\>:Use an exact model name from: \{candidate\_model\_names\}\.Replace Your\-Query with a concrete question\.Never copy model descriptions into <search\>\.Never output the placeholder format literally\.Before each LLM call, reason about why external information is needed and which model is appropriate\.The returned response appears between <information\> and </information\>\.Different models may be called multiple times\.\{model\_descriptions\}If no further external knowledge is needed, output the final answer inside <answer\> \.\.\. </answer\>\.Do not output empty answer tags\.Question: \{question\}

#### CausalLM\.

The CausalLM router is trained to complete the selected model name after this prefix:

CausalLMYou are an intelligent router that selects the best Large Language Model \(LLM\) for a given query\.Available LLMs: \{model\_1\}, \{model\_2\}, \.\.\.Based on the query content, complexity, and requirements, predict which LLM would provide the best response\.Query: \{query\}Best LLM:

#### AutoMix\.

AutoMix first asks a small model to answer with the following template:

AutoMix answer generationYou are given a question\. Answer the question as concisely as you can, using a single phrase if possible\.Question: \{question\}Answer: The answer is’

It uses the following few\-shot verifier:

AutoMix verifierQuestion: Whose lost work was discovered in a dusty attic in 1980?AI Generated Answer: ShakespeareInstruction: Your task is to evaluate if the AI Generated Answer is correct, based on the provided question\. Provide the judgement and reasoning for each case\. Choose between Correct or Incorrect\.Evaluation: The lost work of Shakespeare was discovered in 1980 in a dusty attic\.Verification Decision: The AI generated answer is Correct\.\-\-\-Question: In which month does the celestial event, the Pink Moon, occur?AI Generated Answer: JulyInstruction: Your task is to evaluate if the AI Generated Answer is correct, based on the provided question\. Provide the judgement and reasoning for each case\. Choose between Correct or Incorrect\.Evaluation: The Pink Moon is unique to the month of April\.Verification Decision: The AI generated answer is Incorrect\.\-\-\-Question: Who is believed to have painted the Mona Lisa in the early 16th century?AI Generated Answer: Vincent van GoghInstruction: Your task is to evaluate if the AI Generated Answer is correct, based on the provided question\. Provide the judgement and reasoning for each case\. Choose between Correct or Incorrect\.Evaluation: The Mona Lisa was painted by Leonardo da Vinci in the early 16th century\.Verification Decision: The AI generated answer is Incorrect\.\-\-\-Question: How far away is the planet Kepler\-442b?AI Generated Answer: 1,100 light\-yearsInstruction: Your task is to evaluate if the AI Generated Answer is correct, based on the provided question\. Provide the judgement and reasoning for each case\. Choose between Correct or Incorrect\.Evaluation: The Kepler\-442b is located 1,100 light\-years away\.Verification Decision: The AI generated answer is Correct\.\-\-\-Question: \{question\}AI Generated Answer: \{generated\_answer\}Instruction: Your task is to evaluate if the AI Generated Answer is correct, based on the provided question\. Provide the judgement and reasoning for each case\. Choose between Correct or Incorrect\.Evaluation:

The verifier’s estimated correctness determines whether the query is escalated to the larger model\.

#### OpenClaw deployment router\.

When the deployed server uses prompt\-based routing, it selects one model with:

OpenClaw deployment routerYou are an intelligent LLM router\. Choose the most suitable model for the user’s query\.Available models: \{model\_descriptions\}Rules:1\. Simple greetings/daily chat \-\> cheaper models \(8b, 9b size\)2\. Q&A/knowledge retrieval \-\> chatqa models3\. Instruction following/structured output \-\> mistral models4\. Code generation/technical questions \-\> nemotron or larger models5\. Complex reasoning/deep analysis \-\> 70b or larger modelsIMPORTANT: Only return the model name, nothing else\!Model names: \{model\_list\}User query: \{query\}

For image and video inputs, OpenClaw first uses the following preprocessing prompts; the returned text is included in the router query\.

OpenClaw image and video preprocessingImage input:Describe this image concisely in 2\-\-3 sentences\.Video input:Describe what you see in these video frames\.

### F\.3Multi\-Agent Topology Templates

The multi\-agent topologies in §[E](https://arxiv.org/html/2608.06867#A5)route every LLM\-call node independently from the prompt assigned to that node\. Repeated nodes use the same template and differ only in their indexed fields\. The final node is shared by all five topologies and is given after the topology\-specific templates\.

#### Star\.

The central planner decomposes the problem for\{n\_actors\}actors, then consolidates their reports\.

Star plannerYou are the central planner of a team of \{n\_actors\} agents working on the problem below\. Decompose the problem into exactly \{n\_actors\} subtasks, one per agent, so that their combined results solve the problem\. Output EXACTLY \{n\_actors\} lines, one subtask per line, no numbering or extra words\.Problem: \{user\_query\}

Star actorYou are agent \{i\+1\} in a team\. Solve the following subtask assigned by your planner\. Be concise and factual\.Original problem: \{user\_query\}Your subtask: \{subtask\}

Star planner consolidationYou are the central planner\. Your agents reported the results below\. Consolidate them into a single coherent analysis of the problem\.Problem: \{user\_query\}Agent reports: Subtask: \{subtask\_1\} Result: \{result\_1\}Subtask: \{subtask\_2\} Result: \{result\_2\}\.\.\.

#### Tree\.

The root planner assigns two complementary sub\-problems to team leads\. Each lead refines its branch into one task for a worker, after which the root planner consolidates the two reports\.

Tree root plannerYou are the top\-level planner\. Split the problem below into exactly 2 complementary sub\-problems for your two subordinate team leads\. Output EXACTLY 2 lines, one sub\-problem per line, no numbering\.Problem: \{user\_query\}

Tree sub\-plannerYou are team lead \{b\+1\}\. Refine the sub\-problem below into one concrete task for your worker agent\. Output ONE line only\.Original problem: \{user\_query\}Your sub\-problem: \{branch\_task\}

Tree workerYou are a worker agent\. Solve the task below concisely and factually\.Original problem: \{user\_query\}Your task: \{task\}

Tree root consolidationYou are the top\-level planner\. Consolidate your two teams’ reports into a single coherent analysis of the problem\.Problem: \{user\_query\}Team reports: Sub\-problem: \{branch\_task\_1\} Task: \{task\_1\} Result: \{result\_1\}Sub\-problem: \{branch\_task\_2\} Task: \{task\_2\} Result: \{result\_2\}

#### Graph\.

Each of\{n\_actors\}agents first answers independently\. In each subsequent communication round, an agent receives the other agents’ current answers and revises its own answer\.

Graph actor: initial roundYou are agent \{i\+1\} of \{n\_actors\} independent agents\. Solve the problem below on your own\. Show brief reasoning, then your answer\.Problem: \{user\_query\}

Graph actor: revision roundYou are agent \{i\+1\} of \{n\_actors\}\. You communicated with all other agents; their current answers are below\. Reconsider and give your revised reasoning and answer\.Problem: \{user\_query\}Your previous answer: \{answers\[i\]\}Agent \{j\+1\}’s current answer: \{answers\[j\]\}\.\.\.

#### Chain\.

The first agent answers independently\. Each later agent verifies and improves the immediately preceding decision before passing its result onward\.

Chain: first agentYou are agent 1 in a chain of \{n\_agents\} agents\. Solve the problem below\. Show brief reasoning, then your answer\. Your output will be passed to the next agent\.Problem: \{user\_query\}

Chain: later agentYou are agent \{i\+1\} in a chain of \{n\_agents\} agents\. The previous agent passed you their decision below\. Verify it, fix any mistakes, and pass on your improved reasoning and answer\.Problem: \{user\_query\}Previous agent’s decision: \{decision\}

#### Plan\-Exec\-Sum\.

The GraphPlanner\-style topology first decomposes the query into\{width\}atomic sub\-queries\. Each executor receives its sub\-query without an additional instruction, and the summarizer combines the results\.

Plan\-Exec\-Sum plannerYou are a query decomposition assistant\. Your task is to decompose the user’s query into exactly \{width\} atomic and independent sub\-queries\.Keep each sub\-query self\-contained and non\-overlapping\.Prefer factual, directly answerable units\.Output EXACTLY \{width\} lines, one sub\-query per line, no numbering or extra words\.User query: \{user\_query\}

Plan\-Exec\-Sum executor\{sub\_query\}

Plan\-Exec\-Sum summarizerYou are a professional summarizer\. Summarize the following content into a concise, coherent paragraph without bullet points\. Make it fluent and logically connected\.Content: Sub\-query: \{sub\_query\_1\} Answer: \{answer\_1\}Sub\-query: \{sub\_query\_2\} Answer: \{answer\_2\}\.\.\.Summary:

#### Shared final node\.

Every topology uses the same final node\. Its system instruction is restored from the original task’s\[System Instruction\]block, so the required answer format remains task\-specific\. When there is team context, the final node receives the following user prompt:

Shared final node: user prompt with team context\{user\_query\}Below is supporting analysis from your team\. Use it if helpful, but follow the answer format required above\.\[Team Analysis\] \{context\}

Without team context, the final node receives\{user\_query\}alone\.

## Appendix GThe ComfyUI Visual Interface

LLMRouterexposes its full routing pipeline as a graph on the ComfyUI canvas, where every node is a library component and every edge is an artifact that flows between components\. Two input nodes supply the source benchmarks and the candidate poolℳ\\mathcal\{M\}, a data\-engine node consumes them and emits the query–model matrix as a single edge, and each router node consumes that matrix and emits its evaluation\. The matrix records the per\-query performance and cost of every candidate, the supervision that every router in §[2\.1](https://arxiv.org/html/2608.06867#S2.SS1)is trained on\. The graph traces the evaluation protocol of §[2\.2](https://arxiv.org/html/2608.06867#S2.SS2)from left to right, and the router nodes appear in menu groups named after the three router families, so the taxonomy of §[2\.1](https://arxiv.org/html/2608.06867#S2.SS1)is visible in the node menu\. The canvas therefore renders the architecture of Figure[3](https://arxiv.org/html/2608.06867#S4.F3)as a diagram a user reads and edits, and replacing a router replaces one node while the data\-engine and evaluation nodes stay in place, so the modularity of the library becomes visible on the canvas\.

The data nodes realize the three stages of the data engine\. The Select Datasets and Select LLMs nodes declare the query source and the candidate pool, and the Generate Data node runs query curation, response collection, and scoring in one call and writes the query–model matrix to a self\-contained data directory\. The router nodes cover every built\-in router and are grouped in the menu by router family, and each router node reads its defaults from the same YAML configuration that the command line uses and renders every hyperparameter as a typed widget whose value, range, and options come from that file, so a canvas node and its scripted counterpart run one configuration\. On execution a router node writes a runtime configuration, dispatches training and evaluation through the shared router registry, and returns a summary of the query count, the success count, the average performance, and the routing distribution over candidates\. The Generate Data node hashes the selected datasets, candidate pool, and sample size into a metadata record and reuses an existing data directory when the record and its files are unchanged, which skips the costly response\-collection stage on repeated runs\.

![Refer to caption](https://arxiv.org/html/2608.06867v1/figs/comfyui.png)Figure 12:The ComfyUI interface ofLLMRouter\.The source benchmarks and the candidate pool enter at the left, the data\-engine node produces the query–model matrix, and each router node consumes the matrix and reports its evaluation, so the graph traces the routing pipeline from data construction to evaluation\. Each router node exposes the hyperparameters of its library configuration as typed widgets\.

Similar Articles

LLMRouter open-sources 16+ router library with xRouteBench

Reddit r/ArtificialInteligence

A new paper on arXiv introduces an open-source library called LLMRouter with over 16 router implementations and a benchmark xRouteBench, demonstrating that learned routers can outperform fixed-model baselines by 14.6%.

LLM Routers Have Become a Service Category of Their Own

Reddit r/ArtificialInteligence

LLM routers are evolving from a niche infrastructure trick into a mainstream service category, enabling users to automatically select the most cost-effective model for each request as frontier model costs rise.

Everyone is building LLM routers, we deprecated ours

Hacker News Top

Manifest explains why it deprecated its LLM router, arguing that model routing introduces unpredictability, breaks behavior consistency, and that prompt complexity cannot be inferred from the prompt alone, making caching and deliberate model selection more effective for most use cases.