Arrow Research search
Back to RLJ

RLJ 2024

PID Accelerated Temporal Difference Algorithms

Journal Article Articles Artificial Intelligence · Machine Learning · Reinforcement Learning

Abstract

Long-horizon tasks, which have a large discount factor, pose a challenge for most conventional reinforcement learning (RL) algorithms. Algorithms such as Value Iteration and Temporal Difference (TD) learning have a slow convergence rate and become inefficient in these tasks. When the transition distributions are given, PID~VI was recently introduced to accelerate the convergence of Value Iteration using ideas from control theory. Inspired by this, we introduce PID TD Learning and PID Q-Learning algorithms for the RL setting, in which only samples from the environment are available. We give a theoretical analysis of the convergence of PID TD Learning and its acceleration compared to the conventional TD Learning. We also introduce a method for adapting PID gains in the presence of noise and empirically verify its effectiveness.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Reinforcement Learning Journal
Archive span
2024-2025
Indexed papers
228
Paper id
928317356672585141
v2026.09.13