AAMAS 2026
Modeling Dynamics under Random Delays in Reinforcement Learning
Abstract
Reinforcement learning in real-world systems often encounters delays in sensing and actuation, violating the standard Markov decision process (MDP) assumptions of immediate and fully observed states. While world models offer a promising potential to solve such random-delayed MDPs by imagining undelayed environment dynamics, random actuation delays introduce uncertainty that hinders their direct application. Specifically, world models require the executed actions, rather than the issued ones, to make accurate imaginations of the current state. We present a novel analysis that distinguishes the effects of observation and action delays on world models, revealing an asymmetry that can be exploited to improve the learning process. To address the uncertainty caused by stochastic action execution delays, we propose representing imagined latent states as expected latent states-probability-weighted averages over all possible action-execution trajectories. Compared to sampling the latent state via a single possible action execution trajectory, the expected latent reduces variance in training targets and captures multiple plausible futures at inference. We instantiate our approach using DreamerV3 and validate it on the DeepMind Control Suite with visual inputs. Experimental results show that our method achieves significantly higher returns, more accurate dynamics predictions, and improved training stability across a wide range of delay settings compared to strong baselines.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 313783921876398248