EAAI Journal 2026 Journal Article
Trustworthy distributed mirror learning for secure and private multi-agent coordination
- Suhang Wei
- Jinfang Jia
- Xiang Feng
- Huiqun Yu
Real-world deployment of Multi-Agent Reinforcement Learning (MARL) in Internet of Things (IoT) systems requires convergence, verifiability, and privacy to be jointly guaranteed, a capability absent in current approaches. While Multi-Agent Trust Region Learning (MATRL) ensures Nash equilibrium convergence, it risks data exposure and malicious attacks. To address this, we propose Trustworthy Distributed Mirror Learning (TDML), the first method to unify convergence, verifiability, and privacy in MARL. TDML theoretically breaks MATRL’s centralized architecture into agent-local learning and inter-agent communication. This allows key data to be secured with advanced techniques without compromising the theoretical properties of trust-region learning. Specifically, TDML introduces three core innovations: (1) an information functional that unifies all communication behaviors in distributed MATRL and enables flexible integration of security mechanisms; (2) split advantage computation, which decouples raw inputs from global advantages via intermediate representations to protect local data privacy; and (3) a security scheme that ensures verifiable message exchange by attaching zero-knowledge proofs to inter-agent communications. We prove TDML converges to a Nash equilibrium while providing verifiability and privacy guarantees. More importantly, it constructs a mirror space for trustworthy MARL, where derivative algorithms inherit these theoretical guarantees. Experiments show TDML outperforms state-of-the-art methods, improving attack resilience by up to 76% (sign-flipping attack) and achieving a 90+% privacy reconstruction error, while reducing communication overhead by up to 99% compared to homomorphic encryption. TDML establishes a foundational framework for trustworthy MARL, from which derivative algorithms inherit core guarantees for secure, real-world IoT deployment.