Arrow Research search
Back to RLDM

RLDM 2017

Optimal sample size for A/B tests using cumulative regret

Conference Abstract Accepted abstract Artificial Intelligence · Decision Making · Machine Learning · Reinforcement Learning

Abstract

This work presents theoretical results for determining optimal sample size for A/B tests in the online setting. Our explorations differ from previous studies on three accounts. The first is that we rec- ommend a sample size on the basis of cumulative regret, as opposed to the typical statistical significance tests commonly used in A/B tests. The second is that we seek to optimize expected cumulative regret rather than the upper bound on cumulative regret. The third and most important contribution is that, similar to the Bayesian framework, we model the theoretical means of the alternatives as a random variable, which enables us to go beyond the gap dependent and gap independent results that is typical in bandit studies. We study Gaussian and binary reward distributions, with corresponding Gaussian and Uniform distribution of means. Specifically, in the case which explores a Gaussian reward distribution (noise) along with a dif- ferent Gaussian distribution representing the theoretical means of the alternatives, we derive a closed form solution for the optimal sample size which is only a function of the trial horizon and a ratio of the standard deviations of these two distributions. Our results are compared to settings where an equivalent fixed gap is assumed between the means of the alternatives. Our results indicate that when the gap between alternatives is modeled as random variable, the optimal sample sizes deviate significantly from the corresponding fixed gap settings.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Multidisciplinary Conference on Reinforcement Learning and Decision Making
Archive span
2013-2025
Indexed papers
1004
Paper id
84506270893827609
v2026.09.13