Arrow Research search
Back to AAMAS

AAMAS 2026

Safe Offline Reinforcement Learning using Diffusion Policies

Conference Paper Extended Abstracts Autonomous Agents and Multiagent Systems

Abstract

DiffusionmodelshaveshowngreatpromiseforofflineRLbycapturing complex data distributions, yet their standard formulations lack explicit safety mechanisms. We propose Safe Diffusion Q-learning, which extends Diffusion-QL by integrating a cost critic and a direct penalty term into the diffusion policy objective to enforce constraint satisfaction during action generation. Evaluated on the DSRL benchmark, our method achieves near-zero constraint violation on challenging BulletSafetyGym and SafetyGym tasks while maintainingcompetitiverewardperformance. Theseresultsdemonstrate that expressive diffusion policies can be robustly constrained, enabling their use in safety-critical offline RL.

Authors

Keywords

  • Safe Offline RL
  • Diffusion Policies
  • Q-Learnig
  • Single Agent

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
573043937360284781
v2026.09.13