Arrow Research search
Back to NeurIPS

NeurIPS 1996

Local Bandit Approximation for Optimal Learning Problems

Conference Paper Artificial Intelligence ยท Machine Learning

Abstract

In general, procedures for determining Bayes-optimal adaptive controls for Markov decision processes (MDP's) require a pro(cid: 173) hibitive amount of computation-the optimal learning problem is intractable. This paper proposes an approximate approach in which bandit processes are used to model, in a certain "local" sense, a given MDP. Bandit processes constitute an important subclass of MDP's, and have optimal learning strategies (defined in terms of Gittins indices) that can be computed relatively efficiently. Thus, one scheme for achieving approximately-optimal learning for gen(cid: 173) eral MDP's proceeds by taking actions suggested by strategies that are optimal with respect to local bandit models.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Annual Conference on Neural Information Processing Systems
Archive span
1987-2025
Indexed papers
30776
Paper id
578868591713009108
v2026.09.13