Arrow Research search
Back to AAMAS

AAMAS 2026

Modeling Dynamics under Random Delays in Reinforcement Learning

Conference Paper Research Paper Track Autonomous Agents and Multiagent Systems

Abstract

Reinforcement learning in real-world systems often encounters delays in sensing and actuation, violating the standard Markov decision process (MDP) assumptions of immediate and fully observed states. While world models offer a promising potential to solve such random-delayed MDPs by imagining undelayed environment dynamics, random actuation delays introduce uncertainty that hinders their direct application. Specifically, world models require the executed actions, rather than the issued ones, to make accurate imaginations of the current state. We present a novel analysis that distinguishes the effects of observation and action delays on world models, revealing an asymmetry that can be exploited to improve the learning process. To address the uncertainty caused by stochastic action execution delays, we propose representing imagined latent states as expected latent states-probability-weighted averages over all possible action-execution trajectories. Compared to sampling the latent state via a single possible action execution trajectory, the expected latent reduces variance in training targets and captures multiple plausible futures at inference. We instantiate our approach using DreamerV3 and validate it on the DeepMind Control Suite with visual inputs. Experimental results show that our method achieves significantly higher returns, more accurate dynamics predictions, and improved training stability across a wide range of delay settings compared to strong baselines.

Authors

Keywords

  • Random Delay
  • Model-Based Reinforcement Learning

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
313783921876398248
v2026.09.13