Arrow Research search
Back to IROS

IROS 2015

Nonparametric Bayesian reward segmentation for skill discovery using inverse reinforcement learning

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

We present a method for segmenting a set of unstructured demonstration trajectories to discover reusable skills using inverse reinforcement learning (IRL). Each skill is characterised by a latent reward function which the demonstrator is assumed to be optimizing. The skill boundaries and the number of skills making up each demonstration are unknown. We use a Bayesian nonparametric approach to propose skill segmentations and maximum entropy inverse reinforcement learning to infer reward functions from the segments. This method produces a set of Markov Decision Processes (MDPs) that best describe the input trajectories. We evaluate this approach in a car driving domain and a simulated quadcopter obstacle course, showing that it is able to recover demonstrated skills more effectively than existing methods.

Authors

Keywords

  • Trajectory
  • Hidden Markov models
  • Learning (artificial intelligence)
  • Bayes methods
  • Markov processes
  • Heuristic algorithms
  • Context
  • Inverse Reinforcement Learning
  • Maximum Entropy
  • Reward Function
  • Markov Decision Process
  • Number Of Skills
  • Obstacle Course
  • Time Series
  • Environmental Changes
  • Graphical Representation
  • Linear System
  • Markov Chain Monte Carlo
  • Hidden Markov Model
  • Transition Probabilities
  • Single State
  • Average Reward
  • True Sequence
  • Linear Dynamical System
  • Multiple Skills
  • Bernoulli Process
  • Left Lane

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
458277832935224959
v2026.09.13