Arrow Research search

Author name cluster

Steven Bradtke

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

NeurIPS Conference 1994 Conference Paper

Reinforcement Learning Methods for Continuous-Time Markov Decision Problems

  • Steven Bradtke
  • Michael Duff

Semi-Markov Decision Problems are continuous time generaliza(cid: 173) tions of discrete time Markov Decision Problems. A number of reinforcement learning algorithms have been developed recently for the solution of Markov Decision Problems, based on the ideas of asynchronous dynamic programming and stochastic approxima(cid: 173) tion. Among these are TD(, x), Q-Iearning, and Real-time Dynamic Programming. After reviewing semi-Markov Decision Problems and Bellman's optimality equation in that context, we propose al(cid: 173) gorithms similar to those named above, adapted to the solution of semi-Markov Decision Problems. We demonstrate these algorithms by applying them to the problem of determining the optimal con(cid: 173) trol for a simple queueing system. We conclude with a discussion of circumstances under which these algorithms may be usefully ap(cid: 173) plied.

NeurIPS Conference 1992 Conference Paper

Reinforcement Learning Applied to Linear Quadratic Regulation

  • Steven Bradtke

Recent research on reinforcement learning has focused on algo(cid: 173) rithms based on the principles of Dynamic Programming (DP). One of the most promising areas of application for these algo(cid: 173) rithms is the control of dynamical systems, and some impressive results have been achieved. However, there are significant gaps between practice and theory. In particular, there are no con ver(cid: 173) gence proofs for problems with continuous state and action spaces, or for systems involving non-linear function approximators (such as multilayer perceptrons). This paper presents research applying DP-based reinforcement learning theory to Linear Quadratic Reg(cid: 173) ulation (LQR), an important class of control problems involving continuous state and action spaces and requiring a simple type of non-linear function approximator. We describe an algorithm based on Q-Iearning that is proven to converge to the optimal controller for a large class of LQR problems. We also describe a slightly different algorithm that is only locally convergent to the optimal Q-function, demonstrating one of the possible pitfalls of using a non-linear function approximator with DP-based learning.

v2026.09.13