Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

Hugging Face Daily Papers Papers

Summary

The paper introduces RouteFM, a foundation model for LLM routing that learns reusable routing capabilities via episodic pretraining across heterogeneous environments, allowing a frozen router to adapt to new domains, modalities, and candidate pools through behavioral context alone. RouteFM outperforms the strongest baseline by 2.23 quality points on MMR-Bench with only eight observations per candidate, supporting a 'pretrain once, route anywhere' paradigm.

Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspective, learning a reusable routing capability that generalizes across tasks, candidate models, and deployment conditions. To this end, we introduce RouteFM, which learns to characterize anonymous candidate models from behavioral context and infer their target-specific capabilities, rather than binding routing decisions to fixed model identities or a single environment. Through episodic pretraining across heterogeneous routing environments, this capability can be reused by a frozen router and adapted to new environments through context alone. Experiments demonstrate transfer across changes in domains, modalities, candidate pools, and context budgets, with the largest gains when behavioral evidence is limited. On MMR-Bench, which is excluded from pretraining, RouteFM outperforms the strongest baseline by 2.23 quality points with only eight observations per candidate. These results support moving LLM routing from repeated local fitting toward a pretrain once, route anywhere paradigm. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/RouteFM.
Original Article
View Cached Full Text

Cached at: 10/02/26, 04:26 AM

Paper page - Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

Source: https://huggingface.co/papers/2609.37362

Abstract

Largelanguagemodel(LLM)routingaimstoassigneachquerytothemostsuitablemodelfromaheterogeneouscandidatepool,improvingthequality--efficiencytrade-offofLLMinference.Existingroutersaretypicallylearnedthroughlocalfitting:arouterisoptimizedforaparticularqueryworkloadandcandidatepool,andoftenrequiresadditionalsupervisionorretrainingastheroutingenvironmentchanges.WeaskwhetherLLMroutingcaninsteadbeapproachedfromafoundation-modelperspective,learningareusableroutingcapabilitythatgeneralizesacrosstasks,candidatemodels,anddeploymentconditions.Tothisend,weintroduceRouteFM,whichlearnstocharacterizeanonymouscandidatemodelsfrombehavioralcontextandinfertheirtarget-specificcapabilities,ratherthanbindingroutingdecisionstofixedmodelidentitiesorasingleenvironment.Throughepisodicpretrainingacrossheterogeneousroutingenvironments,thiscapabilitycanbereusedbyafrozenrouterandadaptedtonewenvironmentsthroughcontextalone.Experimentsdemonstratetransferacrosschangesindomains,modalities,candidatepools,andcontextbudgets,withthelargestgainswhenbehavioralevidenceislimited.OnMMR-Bench,whichisexcludedfrompretraining,RouteFMoutperformsthestrongestbaselineby2.23qualitypointswithonlyeightobservationspercandidate.TheseresultssupportmovingLLMroutingfromrepeatedlocalfittingtowardapretrainonce,routeanywhereparadigm.Ourcodeispubliclyavailableathttps://github.com/LAMDA-Model-Reuse/RouteFM.

View arXiv pageView PDFGitHub3Add to collection

Get this paper in your agent:

hf papers read 2609\.37362

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### AIGNLAI/RouteFM

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37362 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37362 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Learning Agent Routing From Early Experience

arXiv cs.CL

This paper introduces BoundaryRouter, a training-free framework that optimizes LLM agent usage by routing queries to either lightweight inference or full agent execution based on early experience. It also presents RouteBench, a benchmark for evaluating routing performance, showing significant improvements in speed and accuracy.

FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

Hugging Face Daily Papers

FlexRouter proposes a coverage-oriented LLM routing framework that uses Determinantal Point Processes to model complementarity among models, maximizing the probability that at least one selected model answers correctly while avoiding redundant selections and fixed budgets.