Arrow Research search
Back to NeurIPS

NeurIPS 1994

Generalization in Reinforcement Learning: Safely Approximating the Value Function

Conference Paper Artificial Intelligence ยท Machine Learning

Abstract

A straightforward approach to the curse of dimensionality in re(cid: 173) inforcement learning and dynamic programming is to replace the lookup table with a generalizing function approximator such as a neu(cid: 173) ral net. Although this has been successful in the domain of backgam(cid: 173) mon, there is no guarantee of convergence. In this paper, we show that the combination of dynamic programming and function approx(cid: 173) imation is not robust, and in even very benign cases, may produce an entirely wrong policy. We then introduce Grow-Support, a new algorithm which is safe from divergence yet can still reap the benefits of successful generalization.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Annual Conference on Neural Information Processing Systems
Archive span
1987-2025
Indexed papers
30776
Paper id
194694709687510486
v2026.09.13