Arrow Research search
Back to NeurIPS

NeurIPS 1994

An Actor/Critic Algorithm that is Equivalent to Q-Learning

Conference Paper Artificial Intelligence ยท Machine Learning

Abstract

We prove the convergence of an actor/critic algorithm that is equiv(cid: 173) alent to Q-Iearning by construction. Its equivalence is achieved by encoding Q-values within the policy and value function of the ac(cid: 173) tor and critic. The resultant actor/critic algorithm is novel in two ways: it updates the critic only when the most probable action is executed from any given state, and it rewards the actor using cri(cid: 173) teria that depend on the relative probability of the action that was executed.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Annual Conference on Neural Information Processing Systems
Archive span
1987-2025
Indexed papers
30776
Paper id
1085134687944083407
v2026.09.13