Arrow Research search
Back to IROS

IROS 2003

Using policy gradient reinforcement learning on autonomous robot controllers

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Robot programmers can often quickly program a robot to approximately execute a task under specific environment conditions. However, achieving robust performance under more general conditions is significantly more difficult. We propose a framework that starts with an existing control system and uses reinforcement feedback from the environment to autonomously improve the controller's performance. We use the policy gradient reinforcement learning (PGRL) framework, which estimates a gradient (in controller space) of improved reward, allowing the controller parameters to be incrementally updated to autonomously achieve locally optimal performance. Our approach is experimentally verified on a Cye robot executing a room entry and observation task, showing significant reduction in task execution time and robustness with respect to un-modelled changes in the environment.

Authors

Keywords

  • Learning
  • Robot control
  • Orbital robotics
  • Feedback
  • Switches
  • Control systems
  • Optimal control
  • State-space methods
  • Computer science
  • Information science
  • Policy Gradient
  • Policy Gradient Reinforcement Learning
  • Control Parameters
  • Local Optimum
  • Control Performance
  • Reinforcement Learning Framework
  • Task Execution Time
  • Robot Programming
  • State Space
  • Potential Moderators
  • Wildfire
  • Optimal Policy
  • Goal State
  • Field Gradient
  • Reward Function
  • Mobile Robot
  • Human Operator
  • Mode Switching
  • Sensor Readings
  • Goal Position
  • Simulated Robot
  • Real Robot
  • Discrete Modes
  • Hybrid Control
  • Policy Gradient Algorithm
  • Optimal Control Policy
  • Discrete State Space
  • Robot Dynamics
  • Robotic System

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
909327149489102502
v2026.09.13