Arrow Research search
Back to IROS

IROS 2023

Dynamic Decision Frequency with Continuous Options

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

In classic reinforcement learning algorithms, agents make decisions at discrete and fixed time intervals. The duration between decisions becomes a crucial hyperparameter, as setting it too short may increase the problem's difficulty by requiring the agent to make numerous decisions to achieve its goal while setting it too long can result in the agent losing control over the system. However, physical systems do not necessarily require a constant control frequency, and for learning agents, it is often preferable to operate with a low frequency when possible and a high frequency when necessary. We propose a framework called Continuous-Time Continuous-Options (CTCO), where the agent chooses options as sub-policies of variable durations. These options are time-continuous and can interact with the system at any desired frequency providing a smooth change of actions. We demonstrate the effectiveness of CTCO by comparing its performance to classical RL and temporal-abstraction RL methods on simulated continuous control tasks with various action-cycle times. We show that our algorithm's performance is not affected by the choice of environment interaction frequency. Furthermore, we demonstrate the efficacy of CTCO in facilitating exploration in a real-world visual reaching task for a 7 DOF robotic arm with sparse rewards.

Authors

Keywords

  • Visualization
  • Time-frequency analysis
  • Heuristic algorithms
  • Reinforcement learning
  • Manipulators
  • Control systems
  • High frequency
  • Learning Algorithms
  • Frequent Interactions
  • Control Task
  • Reinforcement Learning Methods
  • Learning Agent
  • Fixed Time Interval
  • Neural Network
  • Discretion
  • Empirical Analysis
  • Discrete-time
  • Number Of Steps
  • Radial Basis Function
  • Visual Input
  • Pendulum
  • Markov Decision Process
  • Robot Manipulator
  • Policy Targets
  • Continuous Sequence
  • Policy Gradient
  • Real-world Tasks
  • Cheetah
  • Open-loop Control
  • High Frequency Of Interactions
  • Policy Agencies
  • Optional Parameters
  • Reinforcement Learning Agent
  • Regularization Term
  • Number Of Time Steps
  • Real-world Scenarios

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
400763754339233612
v2026.09.13