NeurIPS 1996
Local Bandit Approximation for Optimal Learning Problems
Abstract
In general, procedures for determining Bayes-optimal adaptive controls for Markov decision processes (MDP's) require a pro(cid: 173) hibitive amount of computation-the optimal learning problem is intractable. This paper proposes an approximate approach in which bandit processes are used to model, in a certain "local" sense, a given MDP. Bandit processes constitute an important subclass of MDP's, and have optimal learning strategies (defined in terms of Gittins indices) that can be computed relatively efficiently. Thus, one scheme for achieving approximately-optimal learning for gen(cid: 173) eral MDP's proceeds by taking actions suggested by strategies that are optimal with respect to local bandit models.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Annual Conference on Neural Information Processing Systems
- Archive span
- 1987-2025
- Indexed papers
- 30776
- Paper id
- 578868591713009108