exploration-strategies

Tag

Cards List
#exploration-strategies

When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence

arXiv cs.LG · yesterday Cached

This paper studies KL-regularized contextual bandits and shows that greedy sampling can achieve logarithmic regret without explicit eluder-dimension dependence for both reward and preference feedback.

0 favorites 0 likes
#exploration-strategies

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Hugging Face Daily Papers · 2d ago Cached

Dream-RSI is a framework for scalable recursive self-improvement in AI agents that uses historical discovery trees to create a replay simulator for offline policy evaluation, reducing online costs and improving discovery efficiency.

0 favorites 0 likes
← Back to home

Submit Feedback