limited-feedback

Tag

Cards List
#limited-feedback

Online Learning with LLM Experts from Limited Feedback

arXiv cs.LG · 6d ago Cached

This paper proposes algorithms for adaptively routing prompts to LLM experts in an online setting with limited feedback, formulated as a bandit problem to minimize regret and maximize response quality.

0 favorites 0 likes
#limited-feedback

Online Learning with LLM Experts from Limited Feedback

Hugging Face Daily Papers · 2026-09-05 Cached

This paper formulates the adaptive routing of prompts to large language model experts as a contextual bandit problem with limited feedback, proposing algorithms that achieve sublinear regret and demonstrate efficient learning of high-quality routing strategies.

0 favorites 0 likes
← Back to home

Submit Feedback