Arrow Research search
Back to ICRA

ICRA 2021

Continual Model-Based Reinforcement Learning with Hypernetworks

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Effective planning in model-based reinforcement learning (MBRL) and model-predictive control (MPC) relies on the accuracy of the learned dynamics model. In many instances of MBRL and MPC, this model is assumed to be stationary and is periodically re-trained from scratch on state transition experience collected from the beginning of environment interactions. This implies that the time required to train the dynamics model - and the pause required between plan executions - grows linearly with the size of the collected experience. We argue that this is too slow for lifelong robot learning and propose HyperCRL, a method that continually learns the encountered dynamics in a sequence of tasks using task-conditional hypernetworks. Our method has three main attributes: first, it includes dynamics learning sessions that do not revisit training data from previous tasks, so it only needs to store the most recent fixed-size portion of the state transition experience; second, it uses fixed-capacity hypernetworks to represent non-stationary and task-aware dynamics; third, it outperforms existing continual learning alternatives that rely on fixed-capacity networks, and does competitively with baselines that remember an ever increasing coreset of past experience. We show that HyperCRL is effective in continual model-based reinforcement learning in robot locomotion and manipulation scenarios, such as tasks involving pushing and door opening. Our project website with videos is at this link http://rvl.cs.toronto.edu/blog/2020/hypercrl/

Authors

Keywords

  • Training
  • Conferences
  • Training data
  • Reinforcement learning
  • Switches
  • Robot learning
  • Planning
  • Incremental Learning
  • Continuous Reinforcement
  • Model-based Reinforcement Learning
  • Learning Models
  • Dynamic Model
  • Transition State
  • Model Predictive Control
  • Open Door
  • Previous Tasks
  • Neural Network
  • Friction
  • Artificial Neural Network
  • Dynamic Network
  • Sequence Of Actions
  • Multilayer Perceptron
  • Adaptive Control
  • Feed-forward Network
  • Catastrophic Forgetting
  • Target Network
  • Reward Function
  • Multi-task Learning
  • Network Weights
  • Replay Buffer
  • Problem Setting
  • Task Switching
  • Laplace Approximation
  • Task Data

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
176802674116647552
v2026.09.13