Arrow Research search
Back to ICRA

ICRA 2013

Reinforcement learning with misspecified model classes

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Real-world robots commonly have to act in complex, poorly understood environments where the true world dynamics are unknown. To compensate for the unknown world dynamics, we often provide a class of models to a learner so it may select a model, typically using a minimum prediction error metric over a set of training data. Often in real-world domains the model class is unable to capture the true dynamics, due to either limited domain knowledge or a desire to use a small model. In these cases we call the model class misspecified, and an unfortunate consequence of misspecification is that even with unlimited data and computation there is no guarantee the model with minimum prediction error leads to the best performing policy. In this work, our approach improves upon the standard maximum likelihood model selection metric by explicitly selecting the model which achieves the highest expected reward, rather than the most likely model. We present an algorithm for which the highest performing model from the model class is guaranteed to be found given unlimited data and computation. Empirically, we demonstrate that our algorithm is often superior to the maximum likelihood learner in a batch learning setting for two common RL benchmark problems and a third real-world system, the hydrodynamic cart-pole, a domain whose complex dynamics cannot be known exactly.

Authors

Keywords

  • Mathematical model
  • Data models
  • Computational modeling
  • Equations
  • Standards
  • Measurement
  • Training data
  • Classification Model
  • Model Misspecification
  • Maximum Likelihood
  • Standard Model
  • Prediction Error
  • Complex Dynamics
  • Minimum Error
  • Highest Performance
  • Real-world Systems
  • Unknown Dynamics
  • True Dynamics
  • Dynamic Model
  • Value Function
  • Step In This Direction
  • Real-world Problems
  • Amount Of Training Data
  • Reward Function
  • Markov Decision Process
  • Minimum Square Error
  • Model-based Reinforcement Learning
  • Learning Bias
  • Gradient Ascent
  • Maximum Likelihood Solution
  • Order Of Magnitude Reduction
  • Irrelevant Data
  • Policy Search
  • Revolute Joints
  • Model-free Reinforcement Learning
  • Continuous State Space

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
712968385030407574
v2026.09.13