Arrow Research search
Back to JMLR

JMLR 2018

Refining the Confidence Level for Optimistic Bandit Strategies

Journal Article Articles Artificial Intelligence ยท Machine Learning

Abstract

This paper introduces the first strategy for stochastic bandits with unit variance Gaussian noise that is simultaneously minimax optimal up to constant factors, asymptotically optimal, and never worse than the classical upper confidence bound strategy up to universal constant factors. Preliminary empirical evidence is also promising. Besides this, a conjecture on the optimal form of the regret is shown to be false and a finite-time lower bound on the regret of any strategy is presented that very nearly matches the finite-time upper bound of the newly proposed strategy. [abs] [ pdf ][ bib ] &copy JMLR 2018. ( edit, beta )

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Journal of Machine Learning Research
Archive span
2000-2026
Indexed papers
4180
Paper id
150323167570796030
v2026.09.13