Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing
Summary
The paper proposes SaveRouter, a sparse-supervision LLM routing framework that selectively acquires query-model feedback and shares capability information across related queries, cutting supervision costs while maintaining competitive routing quality and reducing break-even deployment volume by 1.9-9.5x.
View Cached Full Text
Cached at: 09/30/26, 04:17 PM
Paper page - Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing
Source: https://huggingface.co/papers/2609.37402 Published on Sep 29
·
Submitted byhttps://huggingface.co/Gene-Liu
Luciuson Sep 30
Abstract
Largelanguagemodel(LLM)routingreducesservingcostbyassigningeachquerytoanappropriatemodelwhilepreservingresponsequality.Learningsucharouter,however,oftenrequiresexecutingmultiplecandidatemodelsonhistoricalqueriestocollectquery--modelqualityfeedback,creatinganontrivialsupervisioncostbeforedeployment.Existingworklargelyfocusesonserving-timeefficiency,overlookingwhethertheresultingsavingsaresufficienttorecoverthisupfrontexpenditure.Wefurtherobservethatroutingqualityoftensaturateswellbeforeallquery--modelfeedbackiscollected,suggestingthatdensesupervisioncanbeeconomicallyover-provisioned.WeproposeSaveRouter,asparse-supervisionroutingframeworkthatselectivelyacquiresinformativemodelfeedbackandsharescapabilityinformationacrossrelatedqueries,whileretainingquery-levelrefinementforfine-grainedrouting.Weevaluateroutingbyjointlyaccountingforsupervisionexpenditureandsubsequentserving-timesavings.Acrossfourroutingbenchmarks,themainsettingusesonlyabout33--41%ofavailabletrainingfeedbackwhilemaintainingcompetitiveorbetterroutingquality,andreducesthebreak-evendeploymentvolumebyapproximately1.9--9.5timescomparedwiththefastestconventionalrouter.Furtheranalysisshowsthatacquiringmoresupervisionisnotalwayseconomicallypreferable:thesupervisionlevelthatminimizesservingcostcandifferfromtheonethatachievestheearliestpayback.Ourcodeispubliclyavailableathttps://github.com/LAMDA-Model-Reuse/SaveRouter.
View arXiv pageView PDFGitHub13Add to collection
Get this paper in your agent:
hf papers read 2609\.37402
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.37402 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.37402 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.37402 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
When Should LLMs Search? Counterfactual Supervision for Search Routing
This paper formulates the decision of when to use search in LLMs as an instance-level search-routing problem, using counterfactual supervision to train and improve routing policies, achieving macro-F1 improvements on Gemma and Qwen models.
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
LLMRouter presents a unified formulation of LLM routing as a sequential decision process, along with an open-source infrastructure and benchmark (xRouteBench) for developing, evaluating, and deploying LLM routers. Empirical results show learned routers achieve 14.6% relative improvement over the strongest fixed-model baseline.
From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing
This paper proposes DARS, a framework that constructs routing supervision from a distributional view of model behavior to address the unreliability of single-shot labels in LLM routing.
Online Learning for Cost-Efficient LLM Routing (6 minute read)
Ramp Router uses EWMA for failure rates and Thompson sampling for latency to select the cheapest LLM model and service tier meeting deadlines, achieving 30% cost savings without performance loss.
Learning Agent Routing From Early Experience
This paper introduces BoundaryRouter, a training-free framework that optimizes LLM agent usage by routing queries to either lightweight inference or full agent execution based on early experience. It also presents RouteBench, a benchmark for evaluating routing performance, showing significant improvements in speed and accuracy.