Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Hugging Face Daily Papers Papers

Summary

The paper proposes SaveRouter, a sparse-supervision LLM routing framework that selectively acquires query-model feedback and shares capability information across related queries, cutting supervision costs while maintaining competitive routing quality and reducing break-even deployment volume by 1.9-9.5x.

Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/SaveRouter.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:17 PM

Paper page - Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Source: https://huggingface.co/papers/2609.37402 Published on Sep 29

·

Submitted byhttps://huggingface.co/Gene-Liu

Luciuson Sep 30

Abstract

Largelanguagemodel(LLM)routingreducesservingcostbyassigningeachquerytoanappropriatemodelwhilepreservingresponsequality.Learningsucharouter,however,oftenrequiresexecutingmultiplecandidatemodelsonhistoricalqueriestocollectquery--modelqualityfeedback,creatinganontrivialsupervisioncostbeforedeployment.Existingworklargelyfocusesonserving-timeefficiency,overlookingwhethertheresultingsavingsaresufficienttorecoverthisupfrontexpenditure.Wefurtherobservethatroutingqualityoftensaturateswellbeforeallquery--modelfeedbackiscollected,suggestingthatdensesupervisioncanbeeconomicallyover-provisioned.WeproposeSaveRouter,asparse-supervisionroutingframeworkthatselectivelyacquiresinformativemodelfeedbackandsharescapabilityinformationacrossrelatedqueries,whileretainingquery-levelrefinementforfine-grainedrouting.Weevaluateroutingbyjointlyaccountingforsupervisionexpenditureandsubsequentserving-timesavings.Acrossfourroutingbenchmarks,themainsettingusesonlyabout33--41%ofavailabletrainingfeedbackwhilemaintainingcompetitiveorbetterroutingquality,andreducesthebreak-evendeploymentvolumebyapproximately1.9--9.5timescomparedwiththefastestconventionalrouter.Furtheranalysisshowsthatacquiringmoresupervisionisnotalwayseconomicallypreferable:thesupervisionlevelthatminimizesservingcostcandifferfromtheonethatachievestheearliestpayback.Ourcodeispubliclyavailableathttps://github.com/LAMDA-Model-Reuse/SaveRouter.

View arXiv pageView PDFGitHub13Add to collection

Get this paper in your agent:

hf papers read 2609\.37402

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.37402 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37402 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37402 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Learning Agent Routing From Early Experience

arXiv cs.CL

This paper introduces BoundaryRouter, a training-free framework that optimizes LLM agent usage by routing queries to either lightweight inference or full agent execution based on early experience. It also presents RouteBench, a benchmark for evaluating routing performance, showing significant improvements in speed and accuracy.