Tag
This paper studies KL-regularized contextual bandits and shows that greedy sampling can achieve logarithmic regret without explicit eluder-dimension dependence for both reward and preference feedback.
Dream-RSI is a framework for scalable recursive self-improvement in AI agents that uses historical discovery trees to create a replay simulator for offline policy evaluation, reducing online costs and improving discovery efficiency.