Arrow Research search
Back to ICRA

ICRA 2024

Reinforcement Learning in a Safety-Embedded MDP with Trajectory Optimization

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Safe Reinforcement Learning (RL) plays an important role in applying RL algorithms to safety-critical real-world applications, addressing the trade-off between maximizing rewards and adhering to safety constraints. This work introduces a novel approach that combines RL with trajectory optimization to manage this trade-off effectively. Our approach embeds safety constraints within the action space of a modified Markov Decision Process (MDP). The RL agent produces a sequence of actions that are transformed into safe trajectories by a trajectory optimizer, thereby effectively ensuring safety and increasing training stability. This novel approach excels in its performance on challenging Safety Gym tasks, achieving significantly higher rewards and near-zero safety violations during inference. The method’s real-world applicability is demonstrated through a safe and effective deployment in a real robot task of box-pushing around obstacles. Further insights are available from the videos and appendix on our website: https://sites.google.com/view/safemdp.

Authors

Keywords

  • Training
  • Markov decision processes
  • Reinforcement learning
  • Safety
  • Task analysis
  • Trajectory optimization
  • Robots
  • Markov Decision Process
  • High Reward
  • Real Robot
  • Reinforcement Learning Agent
  • Safety Constraints
  • Safety-critical Applications
  • Time Step
  • Hierarchical Structure
  • Challenging Task
  • Root Node
  • Path Planning
  • Model Predictive Control
  • Constrained Optimization
  • Reward Function
  • Constrained Optimization Problem
  • Transition Function
  • Obstacle Avoidance
  • Constraint Satisfaction
  • Reinforcement Learning Policy

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
267814542825206810
v2026.09.13