AAMAS 2026
Safe Offline Reinforcement Learning using Diffusion Policies
Abstract
DiffusionmodelshaveshowngreatpromiseforofflineRLbycapturing complex data distributions, yet their standard formulations lack explicit safety mechanisms. We propose Safe Diffusion Q-learning, which extends Diffusion-QL by integrating a cost critic and a direct penalty term into the diffusion policy objective to enforce constraint satisfaction during action generation. Evaluated on the DSRL benchmark, our method achieves near-zero constraint violation on challenging BulletSafetyGym and SafetyGym tasks while maintainingcompetitiverewardperformance. Theseresultsdemonstrate that expressive diffusion policies can be robustly constrained, enabling their use in safety-critical offline RL.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 573043937360284781