Arrow Research search
Back to IROS

IROS 2021

Memory-based Deep Reinforcement Learning for POMDPs

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

A promising characteristic of Deep Reinforcement Learning (DRL) is its capability to learn optimal policy in an end-to-end manner without relying on feature engineering. However, most approaches assume a fully observable state space, i. e. fully observable Markov Decision Processes (MDPs). In real-world robotics, this assumption is unpractical, because of issues such as sensor sensitivity limitations and sensor noise, and the lack of knowledge about whether the observation design is complete or not. These scenarios lead to Partially Observable MDPs (POMDPs). In this paper, we propose Long-Short-Term-Memory-based Twin Delayed Deep Deterministic Policy Gradient (LSTM-TD3) by introducing a memory component to TD3, and compare its performance with other DRL algorithms in both MDPs and POMDPs. Our results demonstrate the significant advantages of the memory component in addressing POMDPs, including the ability to handle missing and noisy observation data.

Authors

Keywords

  • Sensitivity
  • Reinforcement learning
  • Markov processes
  • Robot sensing systems
  • Noise measurement
  • Intelligent robots
  • Deep Reinforcement Learning
  • State Space
  • Markov Decision Process
  • Feature Engineering
  • Sensor Noise
  • Observation Space
  • Memory Component
  • Deep Reinforcement Learning Algorithm
  • Deterministic Policy Gradient
  • Neural Network
  • Hyperparameters
  • Value Function
  • Minimalist
  • Discrete-time
  • Real Applications
  • Recurrent Neural Network
  • Decrease In Performance
  • Design Space
  • Current Observations
  • Robot Control
  • History Length
  • State St
  • Action Pair
  • Dramatic Performance
  • Past Observations
  • Belief State
  • Policy Learning
  • Deep Q-network

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
314905584862537448
v2026.09.13