Tag
A study revealed that an LLM router trained to select between models learned task recognition instead of difficulty, causing poor generalization on held-out tasks, but deferral based on the cheap model's output yielded better performance.