Arrow Research search
Back to IROS

IROS 2023

Hierarchical Imitation Learning for Stochastic Environments

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Many applications of imitation learning require the agent to generate the full distribution of behaviour observed in the training data. For example, to evaluate the safety of autonomous vehicles in simulation, accurate and diverse behaviour models of other road users are paramount. Existing methods that improve this distributional realism typically rely on hierarchical policies. These condition the policy on types such as goals or personas that give rise to multi-modal behaviour. However, such methods are often inappropriate for stochastic environments where the agent must also react to external factors: because agent types are inferred from the observed future trajectory during training, these environments require that the contributions of internal and external factors to the agent behaviour are disentangled and only internal factors, i. e. , those under the agent's control, are encoded in the type. Encoding future information about external factors leads to inappropriate agent reactions during testing, when the future is unknown and types must be drawn independently from the actual future. We formalize this challenge as distribution shift in the conditional distribution of agent types under environmental stochasticity. We propose Robust Type Conditioning (RTC), which eliminates this shift with adversarial training under randomly sampled types. Experiments on two domains, including the large-scale Waymo Open Motion Dataset, show improved distributional realism while maintaining or improving task performance compared to state-of-the-art baselines.

Authors

Keywords

  • Training
  • Roads
  • Stochastic processes
  • Training data
  • Encoding
  • Trajectory
  • Safety
  • Imitation Learning
  • Stochastic Environment
  • Task Performance
  • Conditional Distribution
  • Internal Factors
  • Control Agents
  • Domain Shift
  • Autonomous Vehicles
  • Types Of Agents
  • Future Trajectories
  • Road Users
  • Improve Task Performance
  • Environmental Stochasticity
  • Collision
  • Bimodal
  • Complex Environment
  • Information Theory
  • Mutual Information
  • Marginal Distribution
  • External Events
  • Traffic Light
  • Reconstruction Loss
  • Autoencoder Framework
  • Covariate Shift
  • Additional Objective
  • Data Modalities
  • Training Objective
  • Hierarchical Method
  • Policy Learning

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
13569736609951344
v2026.09.13