RLDM 2017
Optimal sample size for A/B tests using cumulative regret
Abstract
This work presents theoretical results for determining optimal sample size for A/B tests in the online setting. Our explorations differ from previous studies on three accounts. The first is that we rec- ommend a sample size on the basis of cumulative regret, as opposed to the typical statistical significance tests commonly used in A/B tests. The second is that we seek to optimize expected cumulative regret rather than the upper bound on cumulative regret. The third and most important contribution is that, similar to the Bayesian framework, we model the theoretical means of the alternatives as a random variable, which enables us to go beyond the gap dependent and gap independent results that is typical in bandit studies. We study Gaussian and binary reward distributions, with corresponding Gaussian and Uniform distribution of means. Specifically, in the case which explores a Gaussian reward distribution (noise) along with a dif- ferent Gaussian distribution representing the theoretical means of the alternatives, we derive a closed form solution for the optimal sample size which is only a function of the trial horizon and a ratio of the standard deviations of these two distributions. Our results are compared to settings where an equivalent fixed gap is assumed between the means of the alternatives. Our results indicate that when the gap between alternatives is modeled as random variable, the optimal sample sizes deviate significantly from the corresponding fixed gap settings.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 84506270893827609