Arrow Research search
Back to AAMAS

AAMAS 2026

Surrogate-Augmented Deception in Reinforcement Learning (SAD-RL)

Conference Paper Research Paper Track Autonomous Agents and Multiagent Systems

Abstract

Reinforcement learning (RL) agents in adversarial environments risk being modeled and exploited by opponents that infer their goals or policies. We introduce Surrogate-Augmented Deception in Reinforcement Learning (SAD-RL), a framework that trains agents to resist such modeling by embedding a surrogate predictor into the learning loop and penalizing its accuracy. Rather than emphasizing mereunpredictability, SAD-RLpromotesstrategicopacity—learning behaviors that remain effective while defying opponent inference. We evaluate SAD-RL in two representative domains: a discrete Adversarial Grid World (AGW) and a continuous Sharks and Minnows (SaM) pursuit-evasion task. Across both settings, SAD-RL agents maintain high task performance while exhibiting measurabledeceptionagainstsurrogatemodels, achievingabettertrade-off between effectiveness and opacity than conventional RL agents. We further analyze the trade-off between goal achievement and opacity, identifying distinct modes of balanced and over-deceptive behavior. Together, these results establish SAD-RL as a general and domain-agnostic approach for inducing emergent deception in reinforcement learning.

Authors

Keywords

  • ReinforcementLearning
  • MultiagentSystems
  • AdversarialLearning
  • Deception
  • Opponent Modeling
  • Surrogate Models

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
280240081072691922
v2026.09.13