Tag
This paper proposes a t-step lookahead threshold policy for Whittle index in partially observable restless bandits, proving geometric convergence to the exact index and enhancing numerical accuracy.
This paper studies the sample complexity of learning in average-reward weakly-coupled MDPs and restless bandits, establishing finite-sample PAC guarantees with polynomial complexity using a novel Lyapunov-based analysis framework.
This paper studies restless bandits with binary latent states and imperfect binary feedback, developing a partial conservation laws (PCL)-based framework for establishing indexability and computing the Whittle index, with applications to opportunistic spectrum access.