NeurIPS 1994
Generalization in Reinforcement Learning: Safely Approximating the Value Function
Abstract
A straightforward approach to the curse of dimensionality in re(cid: 173) inforcement learning and dynamic programming is to replace the lookup table with a generalizing function approximator such as a neu(cid: 173) ral net. Although this has been successful in the domain of backgam(cid: 173) mon, there is no guarantee of convergence. In this paper, we show that the combination of dynamic programming and function approx(cid: 173) imation is not robust, and in even very benign cases, may produce an entirely wrong policy. We then introduce Grow-Support, a new algorithm which is safe from divergence yet can still reap the benefits of successful generalization.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Annual Conference on Neural Information Processing Systems
- Archive span
- 1987-2025
- Indexed papers
- 30776
- Paper id
- 194694709687510486