stochastic-linear-bandits

Tag

Cards List
#stochastic-linear-bandits

Stochastic Linear Bandits with Partially Observed Actions

arXiv cs.LG · 2026-07-13 Cached

This paper studies stochastic linear bandits where the agent only observes a random subset of action coordinates, proving that sublinear regret is possible when actions have low intrinsic dimension, and proposes the TOFU-POV algorithm with theoretical guarantees.

0 favorites 0 likes
#stochastic-linear-bandits

Randomized Exploration for Linear Bandits via Absolute Perturbations

arXiv cs.LG · 2026-06-30 Cached

This paper proposes Absolute Thompson Sampling (ATS), a modification of Thompson Sampling that ensures optimism in expectation by using absolute exploration noise, enabling a simpler UCB-style regret analysis while maintaining computational efficiency. It achieves regret matching existing TS bounds, and introduces an ensemble variant that converges to UCB behavior.

0 favorites 0 likes
← Back to home

Submit Feedback