A learned LLM router scored 0.84 AUC. Shuffling the labels within each task still scored 0.838 [R]
Summary
A study revealed that an LLM router trained to select between models learned task recognition instead of difficulty, causing poor generalization on held-out tasks, but deferral based on the cheap model's output yielded better performance.
Similar Articles
@LiorOnAI: A ~10K parameter router can beat every individual open model on MMLU by learning which model should answer which questi…
A tiny ~10K parameter router called tinyrouter learns which open model to use per question on MMLU, outperforming individual models by optimizing allocation.
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
LLMRouter presents a unified formulation of LLM routing as a sequential decision process, along with an open-source infrastructure and benchmark (xRouteBench) for developing, evaluating, and deploying LLM routers. Empirical results show learned routers achieve 14.6% relative improvement over the strongest fixed-model baseline.
LLM-as-judge anchored on one confidence value in 10 of 16 evals. Asking for a label fixed it.
An LLM judge consistently returned a fixed confidence score of 0.72 in evaluations, but switching to categorical labels improved score distribution, showing that models are better at classification than numerical estimation for assessments.
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
This paper evaluates four open-source LLM routers using a common protocol across four benchmarks, finding that performance gains are more closely tied to model tier composition than task-specific targeting.
@omarsar0: Great paper from DeepMind on effective model routing strategies.
Google DeepMind released a paper on effective model routing strategies, discussing how LLM routers are judged on accuracy and cost but can be meaningless if models respond identically.