Arrow Research search
Back to ICML

ICML 2020

Estimating Q(s, s') with Deep Deterministic Dynamics Gradients

Conference Paper Accepted Paper Artificial Intelligence · Machine Learning

Abstract

In this paper, we introduce a novel form of value function, $Q(s, s’)$, that expresses the utility of transitioning from a state $s$ to a neighboring state $s’$ and then acting optimally thereafter. In order to derive an optimal policy, we develop a forward dynamics model that learns to make next-state predictions that maximize this value. This formulation decouples actions from values while still learning off-policy. We highlight the benefits of this approach in terms of value function transfer, learning within redundant action spaces, and learning off-policy from state observations generated by sub-optimal or completely random policies. Code and videos are available at http: //sites. google. com/view/qss-paper.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
International Conference on Machine Learning
Archive span
1993-2025
Indexed papers
16471
Paper id
650251857048157184
v2026.09.13