engineering-assessment

Tag

Cards List
#engineering-assessment

A learned LLM router scored 0.84 AUC. Shuffling the labels within each task still scored 0.838 [R]

Reddit r/MachineLearning ↗ · 3d ago

A study revealed that an LLM router trained to select between models learned task recognition instead of difficulty, causing poor generalization on held-out tasks, but deferral based on the cheap model's output yielded better performance.

0 favorites 0 likes
← Back to home

Submit Feedback