AAMAS 2026
Surrogate-Augmented Deception in Reinforcement Learning (SAD-RL)
Abstract
Reinforcement learning (RL) agents in adversarial environments risk being modeled and exploited by opponents that infer their goals or policies. We introduce Surrogate-Augmented Deception in Reinforcement Learning (SAD-RL), a framework that trains agents to resist such modeling by embedding a surrogate predictor into the learning loop and penalizing its accuracy. Rather than emphasizing mereunpredictability, SAD-RLpromotesstrategicopacity—learning behaviors that remain effective while defying opponent inference. We evaluate SAD-RL in two representative domains: a discrete Adversarial Grid World (AGW) and a continuous Sharks and Minnows (SaM) pursuit-evasion task. Across both settings, SAD-RL agents maintain high task performance while exhibiting measurabledeceptionagainstsurrogatemodels, achievingabettertrade-off between effectiveness and opacity than conventional RL agents. We further analyze the trade-off between goal achievement and opacity, identifying distinct modes of balanced and over-deceptive behavior. Together, these results establish SAD-RL as a general and domain-agnostic approach for inducing emergent deception in reinforcement learning.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 280240081072691922