Arrow Research search
Back to NeurIPS

NeurIPS 2001

A Natural Policy Gradient

Conference Paper Artificial Intelligence ยท Machine Learning

Abstract

We provide a natural gradient method that represents the steepest descent direction based on the underlying structure of the param(cid: 173) eter space. Although gradient methods cannot make large changes in the values of the parameters, we show that the natural gradi(cid: 173) ent is moving toward choosing a greedy optimal action rather than just a better action. These greedy optimal actions are those that would be chosen under one improvement step of policy iteration with approximate, compatible value functions, as defined by Sut(cid: 173) ton et al. [9]. We then show drastic performance improvements in simple MDPs and in the more challenging MDP of Tetris.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Annual Conference on Neural Information Processing Systems
Archive span
1987-2025
Indexed papers
30776
Paper id
571445430136829221
v2026.09.13