rl-guided

Tag

Cards List
#rl-guided

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling

Hugging Face Daily Papers · 2026-06-02 Cached

This paper formulates adaptive sampling for large language models as a Markov decision process and trains a lightweight RL controller to balance correctness, latency, and computational cost, achieving improved trade-offs.

0 favorites 0 likes
← Back to home

Submit Feedback