Tag
This paper investigates how historical A/B test data can inform adaptive experiments using contextual bandits, providing a practical methodology for deciding when and how to deploy adaptive policies based on offline policy evaluation.