Arrow Research search

Author name cluster

Robert Crites

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

NeurIPS Conference 1995 Conference Paper

Improving Elevator Performance Using Reinforcement Learning

  • Robert Crites
  • Andrew Barto

This paper describes the application of reinforcement learning (RL) to the difficult real world problem of elevator dispatching. The el(cid: 173) evator domain poses a combination of challenges not seen in most RL research to date. Elevator systems operate in continuous state spaces and in continuous time as discrete event dynamic systems. Their states are not fully observable and they are nonstationary due to changing passenger arrival rates. In addition, we use a team of RL agents, each of which is responsible for controlling one ele(cid: 173) vator car. The team receives a global reinforcement signal which appears noisy to each agent due to the effects of the actions of the other agents, the random nature of the arrivals and the incomplete observation of the state. In spite of these complications, we show results that in simulation surpass the best of the heuristic elevator control algorithms of which we are aware. These results demon(cid: 173) strate the power of RL on a very large scale stochastic dynamic optimization problem of practical utility.

NeurIPS Conference 1994 Conference Paper

An Actor/Critic Algorithm that is Equivalent to Q-Learning

  • Robert Crites
  • Andrew Barto

We prove the convergence of an actor/critic algorithm that is equiv(cid: 173) alent to Q-Iearning by construction. Its equivalence is achieved by encoding Q-values within the policy and value function of the ac(cid: 173) tor and critic. The resultant actor/critic algorithm is novel in two ways: it updates the critic only when the most probable action is executed from any given state, and it rewards the actor using cri(cid: 173) teria that depend on the relative probability of the action that was executed.

v2026.09.13