AAMAS Conference 2026 Conference Paper
Surrogate-Augmented Deception in Reinforcement Learning (SAD-RL)
- Joe Shymanski
- Scott Nivison
- Sandip Sen
Reinforcement learning (RL) agents in adversarial environments risk being modeled and exploited by opponents that infer their goals or policies. We introduce Surrogate-Augmented Deception in Reinforcement Learning (SAD-RL), a framework that trains agents to resist such modeling by embedding a surrogate predictor into the learning loop and penalizing its accuracy. Rather than emphasizing mereunpredictability, SAD-RLpromotesstrategicopacity—learning behaviors that remain effective while defying opponent inference. We evaluate SAD-RL in two representative domains: a discrete Adversarial Grid World (AGW) and a continuous Sharks and Minnows (SaM) pursuit-evasion task. Across both settings, SAD-RL agents maintain high task performance while exhibiting measurabledeceptionagainstsurrogatemodels, achievingabettertrade-off between effectiveness and opacity than conventional RL agents. We further analyze the trade-off between goal achievement and opacity, identifying distinct modes of balanced and over-deceptive behavior. Together, these results establish SAD-RL as a general and domain-agnostic approach for inducing emergent deception in reinforcement learning.