A learned LLM router scored 0.84 AUC. Shuffling the labels within each task still scored 0.838 [R]

Reddit r/MachineLearning Papers

Summary

A study revealed that an LLM router trained to select between models learned task recognition instead of difficulty, causing poor generalization on held-out tasks, but deferral based on the cheap model's output yielded better performance.

I trained a router to choose between a cheap and an expensive model. It scored 0.84 AUC held out. Then I shuffled the labels within each task — preserving each task's escalation rate, destroying all per-item signal — retrained, and got 0.838. It had learned to recognize the task, not the difficulty. Eliminated in turn: too few labels (110k from RouteLLM's released set) the architecture (linear probe, similarity-weighted ranking, fine-tuned encoder) the representation (a probe recovers human difficulty from the same embeddings) label noise (test-retest kappa 0.88–0.97; the labels support AUC 0.91) On held-out tasks every router falls to ~0.55 and loses to prompt length. What does work: deferral on the cheap model's own output, 0.75 vs 0.60 on the same split, at no extra inference cost. https://brianfeeny.com/posts/replicating-routellm-on-amazon-bedrock/?utm_source=reddit&utm_medium=social&utm_campaign=routing-paper-2026-09 Curious whether anyone has run the within-group shuffle on their own router — my ordering is an engineering assessment, not a benchmark, and I would be glad to be corrected.
Original Article

Similar Articles