Arrow Research search

Author name cluster

Woojun Kim

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

23 papers
2 author rows

Possible papers

23

AAMAS Conference 2026 Conference Paper

B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning

  • Woojun Kim
  • Katia P. Sycara

Overestimation arising from selecting unseen actions during policy evaluation is a major challenge in offline reinforcement learning (RL). A minimalist approach in the single-agent setting—adding behaviorcloning(BC)regularizationtoexistingonlineRLalgorithms— has been shown to be effective in terms of achieving competitive performance with minimal modification to online algorithms; however, thisapproachisunderstudiedinmulti-agentsettings. Inparticular, overestimation becomes worse in multi-agent settings due to the presence of multiple actions, resulting in the BC regularizationbased approach easily suffering from either over-regularization or critic divergence. To address this, we propose a simple yet effective method, BehaviorCloningregularizationwithCriticClipping(B3C), which clips the target critic value in policy evaluation based on the maximum return in the dataset and pushes the limit of the weight on the RL objective over BC regularization, thereby demonstrating superiorperformanceacrossbenchmarks. Additionally, weleverage existing value factorization techniques, particularly non-linear factorization, which is understudied in offline settings. Integrated with non-linear value factorization, B3C outperforms state-of-the-art algorithms on various offline multi-agent benchmarks.

TMLR Journal 2026 Journal Article

Disentangled Concept-Residual Models: Bridging the Interpretability–Performance Gap for Incomplete Concept Sets

  • Renos Zabounidis
  • Ini Oguntola
  • Konghao Zhao
  • Joseph Campbell
  • Woojun Kim
  • Simon Stepputtis
  • Katia P. Sycara

Deploying AI in high-stakes settings requires models that are not only accurate but also interpretable and amenable to human oversight. Concept Bottleneck Models (CBMs) support these goals by structuring predictions around human-understandable concepts, enabling interpretability and post-hoc human intervenability. However, CBMs rely on a ‘complete’ concept set, requiring practitioners to define and label enough concepts to match the predictive power of black-box models. To relax this requirement, prior work introduced residual connections that bypass the concept layer and recover information missing from an incomplete concept set. While effective in bridging the performance gap, these residuals can redundantly encode concept information, a phenomenon we term \textbf{concept-residual overlap}. In this work, we investigate the effects of concept-residual overlap and evaluate strategies to mitigate it. We (1) define metrics to quantify the extent of concept-residual overlap in CRMs; (2) introduce complementary metrics to evaluate how this overlap impacts interpretability, concept importance, and the effectiveness of concept-based interventions; and (3) present \textbf{Disentangled Concept-Residual Models (D-CRMs)}, a general class of CRMs designed to mitigate this issue. Within this class, we propose a novel disentanglement approach based on minimizing mutual information (MI). Using CelebA, CIFAR100, AA2, CUB, and OAI, we show that standard CRMs exhibit significant concept-residual overlap, and that reducing this overlap with MI-based D-CRMs restores key properties of CBMs, including interpretability, functional reliance on concepts, and intervention robustness, without sacrificing predictive performance.

AAMAS Conference 2026 Conference Paper

Generalized Per-Agent Advantage Estimation for Multi-Agent Policy Optimization

  • Seongmin Kim
  • Giseung Park
  • Woojun Kim
  • Jiwon Jeon
  • Seungyul Han
  • Youngchul Sung

In this paper, we propose a novel framework for multi-agent reinforcement learning that enhances sample efficiency and coordination through accurate per-agent advantage estimation. The core of our approach is Generalized Per-Agent Advantage Estimator (GPAE), which employs a per-agent value iteration operator to compute precise per-agent advantages. This operator enables stable off-policy learning by indirectly estimating values via action probabilities, eliminating the need for direct 𝑄-function estimation. To further refine estimation, we introduce a double-truncated importance sampling ratio scheme. This scheme improves credit assignment for off-policy trajectories by balancing sensitivity to the agent’s own policy changes with robustness to non-stationarity from other agents. Experiments on benchmarks demonstrate that our approach outperforms existing approaches, excelling in coordination and sample efficiency for complex scenarios.

AAMAS Conference 2026 Conference Paper

Theory of Mind Guided Strategy Adaptation for Zero-Shot Coordination

  • Andrew Ni
  • Simon Stepputtis
  • Stefanos Nikolaidis
  • Michael Lewis
  • Katia P. Sycara
  • Woojun Kim

A central challenge in multi-agent reinforcement learning is enabling agents to adapt to previously unseen teammates in a zeroshot fashion. Prior work in zero-shot coordination often follows a two-stageprocess, firstgeneratingadiversetrainingpoolofpartner agents, and then training a best-response agent to collaborate effectively with the entire training pool. While many previous works have achieved strong performance by devising better ways to diversify the partner agent pool, there has been less emphasis on how to leverage this pool to build an adaptive agent. One limitation is that the best-response agent may converge to a static, generalist policy that performs reasonably well across diverse teammates, rather than learning a more adaptive, specialist policy that can better adapt to teammates and achieve higher synergy. To address this, we propose an adaptive ensemble agent that uses Theory-of- Mind-based best-response selection to first infer its teammate’s intentions and then select the most suitable policy from a policy ensemble. WeconductexperimentsintheOvercookedenvironment to evaluate zero-shot coordination performance under both fully and partially observable settings. The empirical results demonstrate the superiority of our method over a single best-response baseline.

NeurIPS Conference 2025 Conference Paper

Adaptively Coordinating with Novel Partners via Learned Latent Strategies

  • Benjamin Li
  • Shuyang Shi
  • Lucia Romero
  • Huao Li
  • Yaqi Xie
  • Woojun Kim
  • Stefanos Nikolaidis
  • Charles Lewis

Adaptation is the cornerstone of effective collaboration among heterogeneous team members. In human-agent teams, artificial agents need to adapt to their human partners in real time, as individuals often have unique preferences and policies that may change dynamically throughout interactions. This becomes particularly challenging in tasks with time pressure and complex strategic spaces, where identifying partner behaviors and selecting suitable responses is difficult. In this work, we introduce a strategy-conditioned cooperator framework that learns to represent, categorize, and adapt to a broad range of potential partner strategies in real-time. Our approach encodes strategies with a variational autoencoder to learn a latent strategy space from agent trajectory data, identifies distinct strategy types through clustering, and trains a cooperator agent conditioned on these clusters by generating partners of each strategy type. For online adaptation to novel partners, we leverage a fixed-share regret minimization algorithm that dynamically infers and adjusts the partner's strategy estimation during interaction. We evaluate our method in a modified version of the Overcooked domain, a complex collaborative cooking environment that requires effective coordination among two players with a diverse potential strategy space. Through these experiments and an online user study, we demonstrate that our proposed agent achieves state of the art performance compared to existing baselines when paired with novel human, and agent teammates.

YNICL Journal 2025 Journal Article

Association of iron deposition in MS lesion with remyelination capacity using susceptibility source separation MRI

  • Hyeong-Geol Shin
  • Woojun Kim
  • Jung Hwan Lee
  • Hyun-soo Lee
  • Yoonho Nam
  • Jiwoong Kim
  • Xu Li
  • Peter C.M. van Zijl

OBJECTIVES: signals within MS lesions using χ-separation and evaluate the association between lesional iron and remyelination capability. METHODS: signals. RESULTS: myelin signals (P < 0.001). After adjustment, lesions with early HPS demonstrated an annual loss in myelin signal (-1.94 ppb/year), whereas those without early HPS exhibited annual recovery (+0.66 ppb/year). Participants with confirmed disability improvement (CDI) had fewer HPS-positive lesions at baseline than those without CDI (P < 0.001). CONCLUSION: The presence of HPS is associated with impaired remyelination capacity and a lack of disease improvement in pwMS. Identifying HPS may help demarcate lesions more amenable to myelin repair therapies.

ICRA Conference 2025 Conference Paper

Distributed Multi-Robot Source Seeking in Unknown Environments with Unknown Number of Sources

  • Lingpeng Chen
  • Siva Kailas
  • Srujan Deolasee
  • Wenhao Luo
  • Katia P. Sycara
  • Woojun Kim

We introduce a novel distributed source seeking framework, DIAS, designed for multi-robot systems in scenarios where the number of sources is unknown and potentially exceeds the number of robots. Traditional robotic source seeking methods typically focused on directing each robot to a specific strong source and may fall short in comprehensively identifying all potential sources. DIAS addresses this gap by introducing a hybrid controller that identifies the presence of sources and then alternates between exploration for data gathering and exploitation for guiding robots to identified sources. It further enhances search efficiency by dividing the environment into Voronoi cells and approximating source density functions based on Gaussian process regression. Additionally, DIAS can be integrated with existing source seeking algorithms. We compare DIAS with existing algorithms, including DOSS and GMES in simulated gas leakage scenarios where the number of sources outnumbers or is equal to the number of robots. The numerical results show that DIAS outperforms the baseline methods in both the efficiency of source identification by the robots and the accuracy of the estimated environmental density function.

ICRA Conference 2025 Conference Paper

Escaping Local Minima: Hybrid Artificial Potential Field with Wall-Follower for Decentralized Multi-Robot Navigation

  • Joonkyung Kim
  • Sangjin Park
  • Wonjong Lee
  • Woojun Kim
  • Hyunga Choi
  • Nakju Lett Doh
  • Changjoo Nam

We tackle the challenges of decentralized multi-robot navigation in environments with nonconvex obstacles, where complete environmental knowledge is unavailable. While reactive methods like Artificial Potential Field (APF) offer simplicity and efficiency, they suffer from local minima, causing robots to become trapped due to their lack of global environmental awareness. Other existing solutions either rely on inter-robot communication, are limited to single-robot scenarios, or struggle to overcome nonconvex obstacles effectively. Our proposed methods enable collision-free navigation using only local sensor and state information without a map. By incorporating a wall-following (WF) behavior into the APF approach, our method allows robots to escape local minima, even in the presence of nonconvex and dynamic obstacles including other robots. We introduce two algorithms for switching between APF and WF: a rule-based system and an encoder network trained on expert demonstrations. Experimental results show that our approach achieves substantially higher success rates compared to state-of-the-art methods, highlighting its ability to overcome the limitations of local minima in complex environments.

NeurIPS Conference 2025 Conference Paper

Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment

  • Woojun Kim
  • Katia Sycara

Multi-agent reinforcement learning in mixed-motive settings presents a fundamental challenge: agents must balance individual interests with collective goals, which are neither fully aligned nor strictly opposed. To address this, reward restructuring methods such as gifting and intrinsic motivation have been proposed. However, these approaches primarily focus on promoting cooperation by managing the trade-off between individual and collective returns, without explicitly addressing fairness with respect to agents’ task-specific rewards. In this paper, we propose an adaptive conflict-aware gradient adjustment method that promotes cooperation while ensuring fairness in individual rewards. The proposed method dynamically balances policy gradients derived from individual and collective objectives in situations where the two objectives are in conflict. By explicitly resolving such conflicts, our method improves collective performance while preserving fairness across agents. We provide theoretical results that guarantee monotonic non-decreasing improvement in both the collective and individual objectives and ensure fairness. Empirical results in sequential social dilemma environments demonstrate that our approach outperforms baselines in terms of social welfare, while maintaining fairness.

ICRA Conference 2025 Conference Paper

Integrating Multi-Robot Adaptive Sampling and Informative Path Planning for Spatiotemporal Natural Environment Prediction

  • Siva Kailas
  • Srujan Deolasee
  • Wenhao Luo
  • Woojun Kim
  • Katia P. Sycara

Learning to predict spatiotemporal (ST) environmental processes from a sparse set of samples collected autonomously is a difficult task from both a sampling perspective (collecting the best sparse samples) and from a learning perspective (predicting the next timestep). In this work, we focus on investigating the sample collection process via multirobot informative path planning. We present an approach for incorporating multi-robot informative path planning into a spatiotemporal adaptive sampling framework while considering path length constraints for sampling location selection. We also incorporate informative path planning to determine the best path to collect samples along while en route to collecting the desired sample. We achieve this in a decentralized manner by decoupling the process into two stages: the first stage uses our spatiotemporal mixture of Gaussian Processes (STMGP) model to determine the most informative sampling location via a mutual information lower bound heuristic and the second stage plans an informative path to collect the desired sample and other additional informative samples via submodular function optimization. Moreover, we effectively leverage peer-to-peer communication to enable coordination. Simulation results on real-world spatiotemporal data are provided to validate the effectiveness of our proposed approach.

NeurIPS Conference 2024 Conference Paper

Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning

  • Jeonghye Kim
  • Suyoung Lee
  • Woojun Kim
  • Youngchul Sung

Offline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce $Q$-Aided Conditional Supervised Learning (QCS), which effectively combines the stability of RCSL with the stitching capability of $Q$-functions. By analyzing $Q$-function over-generalization, which impairs stable stitching, QCS adaptively integrates $Q$-aid into RCSL's loss function based on trajectory return. Empirical results show that QCS significantly outperforms RCSL and value-based methods, consistently achieving or exceeding the highest trajectory returns across diverse offline RL benchmarks. QCS represents a breakthrough in offline RL, pushing the limits of what can be achieved and fostering further innovations.

ICLR Conference 2024 Conference Paper

Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision Making

  • Jeonghye Kim
  • Suyoung Lee
  • Woojun Kim
  • Youngchul Sung

The recent success of Transformer in natural language processing has sparked its use in various domains. In offline reinforcement learning (RL), Decision Transformer (DT) is emerging as a promising model based on Transformer. However, we discovered that the attention module of DT is not appropriate to capture the inherent local dependence pattern in trajectories of RL modeled as a Markov decision process. To overcome the limitations of DT, we propose a novel action sequence predictor, named Decision ConvFormer (DC), based on the architecture of MetaFormer, which is a general structure to process multiple entities in parallel and understand the interrelationship among the multiple entities. DC employs local convolution filtering as the token mixer and can effectively capture the inherent local associations of the RL dataset. In extensive experiments, DC achieved state-of-the-art performance across various standard RL benchmarks while requiring fewer resources. Furthermore, we show that DC better understands the underlying meaning in data and exhibits enhanced generalization capability.

IROS Conference 2024 Conference Paper

ShapeGrasp: Zero-Shot Task-Oriented Grasping with Large Language Models through Geometric Decomposition

  • Samuel Li
  • Sarthak Bhagat
  • Joseph Campbell
  • Yaqi Xie 0001
  • Woojun Kim
  • Katia P. Sycara
  • Simon Stepputtis

Task-oriented grasping of unfamiliar objects is a necessary skill for robots in dynamic in-home environments. Inspired by the human capability to grasp such objects through intuition about their shape and structure, we present a novel zero-shot task-oriented grasping method leveraging a geometric decomposition of the target object into simple, convex shapes that we represent in a graph structure, including geometric attributes and spatial relationships. Our approach employs minimal essential information – the object’s name and the intended task – to facilitate zero-shot task-oriented grasping. We utilize the commonsense reasoning capabilities of large language models to dynamically assign semantic meaning to each decomposed part and subsequently reason over the utility of each part for the intended task. Through extensive experiments on a real-world robotics platform, we demonstrate that our grasping approach’s decomposition and reasoning pipeline is capable of selecting the correct part in 92% of the cases and successfully grasping the object in 82% of the tasks we evaluate. Additional videos, experiments, code, and data are available on our project website: https://shapegrasp.github.io/.

AAMAS Conference 2023 Conference Paper

A Variational Approach to Mutual Information-Based Coordination for Multi-Agent Reinforcement Learning

  • Woojun Kim
  • Whiyoung Jung
  • Myungsik Cho
  • Youngchul Sung

In this paper, we propose a new mutual information (MMI) framework for multi-agent reinforcement learning (MARL) to enable multiple agents to learn coordinated behaviors by regularizing the accumulated return with the simultaneous mutual information between multi-agent actions. By introducing a latent variable to induce nonzero mutual information between multi-agent actions and applying a variational bound, we derive a tractable lower bound on the considered MMI-regularized objective function. The derived tractable objective can be interpreted as maximum entropy reinforcement learning combined with uncertainty reduction of other agents’ actions. Applying policy iteration to maximize the derived lower bound, we propose a practical algorithm named variational maximum mutual information multi-agent actor-critic (VM3-AC), which follows centralized learning with decentralized execution (CTDE). We evaluated VM3-AC for several games requiring coordination, and numerical results show that VM3-AC outperforms other MARL algorithms in multi-agent tasks requiring high-quality coordination.

ICML Conference 2023 Conference Paper

An Adaptive Entropy-Regularization Framework for Multi-Agent Reinforcement Learning

  • Woojun Kim
  • Youngchul Sung

In this paper, we propose an adaptive entropy-regularization framework (ADER) for multi-agent reinforcement learning (RL) to learn the adequate amount of exploration of each agent for entropy-based exploration. In order to derive a metric for the proper level of exploration entropy for each agent, we disentangle the soft value function into two types: one for pure return and the other for entropy. By applying multi-agent value factorization to the disentangled value function of pure return, we obtain a metric to determine the relevant level of exploration entropy for each agent, given by the partial derivative of the pure-return value function with respect to (w. r. t.) the policy entropy of each agent. Based on this metric, we propose the ADER algorithm based on maximum entropy RL, which controls the necessary level of exploration across agents over time by learning the proper target entropy for each agent. Experimental results show that the proposed scheme significantly outperforms current state-of-the-art multi-agent RL algorithms.

NeurIPS Conference 2023 Conference Paper

Domain Adaptive Imitation Learning with Visual Observation

  • Sungho Choi
  • Seungyul Han
  • Woojun Kim
  • Jongseong Chae
  • Whiyoung Jung
  • Youngchul Sung

In this paper, we consider domain-adaptive imitation learning with visual observation, where an agent in a target domain learns to perform a task by observing expert demonstrations in a source domain. Domain adaptive imitation learning arises in practical scenarios where a robot, receiving visual sensory data, needs to mimic movements by visually observing other robots from different angles or observing robots of different shapes. To overcome the domain shift in cross-domain imitation learning with visual observation, we propose a novel framework for extracting domain-independent behavioral features from input observations that can be used to train the learner, based on dual feature extraction and image reconstruction. Empirical results demonstrate that our approach outperforms previous algorithms for imitation learning from visual observation with domain shift.

ICML Conference 2023 Conference Paper

LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework

  • Woojun Kim
  • Jeonghye Kim
  • Youngchul Sung

In this paper, a unified framework for exploration in reinforcement learning (RL) is proposed based on an option-critic architecture. The proposed framework learns to integrate a set of diverse exploration strategies so that the agent can adaptively select the most effective exploration strategy to realize an effective exploration-exploitation trade-off for each given task. The effectiveness of the proposed exploration framework is demonstrated by various experiments in the MiniGrid and Atari environments.

AAMAS Conference 2023 Conference Paper

Parameter Sharing with Network Pruning for Scalable Multi-Agent Deep Reinforcement Learning

  • Woojun Kim
  • Youngchul Sung

Handling the problem of scalability is one of the essential issues for multi-agent reinforcement learning (MARL) algorithms to be applied to real-world problems typically involving massively many agents. For this, parameter sharing across multiple agents has widely been used since it reduces the training time by decreasing the number of parameters and increasing the sample efficiency. However, using the same parameters across agents limits the representational capacity of the joint policy and consequently, the performance can be degraded in multi-agent tasks that require different behaviors for different agents. In this paper, we propose a simple method that adopts structured pruning for a deep neural network to increase the representational capacity of the joint policy without introducing additional parameters. We evaluate the proposed method on several benchmark tasks, and numerical results show that the proposed method significantly outperforms other parameter-sharing methods.

NeurIPS Conference 2023 Conference Paper

Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble Agents

  • Woojun Kim
  • Yongjae Shin
  • Jongeui Park
  • Youngchul Sung

Deep reinforcement learning (RL) has achieved remarkable success in solving complex tasks through its integration with deep neural networks (DNNs) as function approximators. However, the reliance on DNNs has introduced a new challenge called primacy bias, whereby these function approximators tend to prioritize early experiences, leading to overfitting. To alleviate this bias, a reset method has been proposed, which involves periodic resets of a portion or the entirety of a deep RL agent while preserving the replay buffer. However, the use of this method can result in performance collapses after executing the reset, raising concerns from the perspective of safe RL and regret minimization. In this paper, we propose a novel reset-based method that leverages deep ensemble learning to address the limitations of the vanilla reset method and enhance sample efficiency. The effectiveness of the proposed method is validated through various experiments including those in the domain of safe RL. Numerical results demonstrate its potential for real-world applications requiring high sample efficiency and safety considerations.

ICML Conference 2022 Conference Paper

MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay Buffer

  • Jeewon Jeon
  • Woojun Kim
  • Whiyoung Jung
  • Youngchul Sung

In this paper, we consider cooperative multi-agent reinforcement learning (MARL) with sparse reward. To tackle this problem, we propose a novel method named MASER: MARL with subgoals generated from experience replay buffer. Under the widely-used assumption of centralized training with decentralized execution and consistent Q-value decomposition for MARL, MASER automatically generates proper subgoals for multiple agents from the experience replay buffer by considering both individual Q-value and total Q-value. Then, MASER designs individual intrinsic reward for each agent based on actionable representation relevant to Q-learning so that the agents reach their subgoals while maximizing the joint action value. Numerical results show that MASER significantly outperforms StarCraft II micromanagement benchmark compared to other state-of-the-art MARL algorithms.

ICLR Conference 2021 Conference Paper

Communication in Multi-Agent Reinforcement Learning: Intention Sharing

  • Woojun Kim
  • Jongeui Park
  • Youngchul Sung

Communication is one of the core components for learning coordinated behavior in multi-agent systems. In this paper, we propose a new communication scheme named Intention Sharing (IS) for multi-agent reinforcement learning in order to enhance the coordination among agents. In the proposed IS scheme, each agent generates an imagined trajectory by modeling the environment dynamics and other agents' actions. The imagined trajectory is the simulated future trajectory of each agent based on the learned model of the environment dynamics and other agents and represents each agent's future action plan. Each agent compresses this imagined trajectory capturing its future action plan to generate its intention message for communication by applying an attention mechanism to learn the relative importance of the components in the imagined trajectory based on the received message from other agents. Numeral results show that the proposed IS scheme outperforms other communication schemes in multi-agent reinforcement learning.

YNIMG Journal 2021 Journal Article

χ-separation: Magnetic susceptibility source separation toward iron and myelin mapping in the brain

  • Hyeong-Geol Shin
  • Jingu Lee
  • Young Hyun Yun
  • Seong Ho Yoo
  • Jinhee Jang
  • Se-Hong Oh
  • Yoonho Nam
  • Sehoon Jung

Obtaining a histological fingerprint from the in-vivo brain has been a long-standing target of magnetic resonance imaging (MRI). In particular, non-invasive imaging of iron and myelin, which are involved in normal brain functions and are histopathological hallmarks in neurodegenerative diseases, has practical utilities in neuroscience and medicine. Here, we propose a biophysical model that describes the individual contribution of paramagnetic (e.g., iron) and diamagnetic (e.g., myelin) susceptibility sources to the frequency shift and transverse relaxation of MRI signals. Using this model, we develop a method, χ-separation, that generates the voxel-wise distributions of the two sources. The method is validated using computer simulation and phantom experiments, and applied to ex-vivo and in-vivo brains. The results delineate the well-known histological features of iron and myelin in the specimen, healthy volunteers, and multiple sclerosis patients. This new technology may serve as a practical tool for exploring the microstructural information of the brain.

AAAI Conference 2019 Conference Paper

Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning

  • Woojun Kim
  • Myungsik Cho
  • Youngchul Sung

In this paper, we propose a new learning technique named message-dropout to improve the performance for multi-agent deep reinforcement learning under two application scenarios: 1) classical multi-agent reinforcement learning with direct message communication among agents and 2) centralized training with decentralized execution. In the first application scenario of multi-agent systems in which direct message communication among agents is allowed, the messagedropout technique drops out the received messages from other agents in a block-wise manner with a certain probability in the training phase and compensates for this effect by multiplying the weights of the dropped-out block units with a correction probability. The applied message-dropout technique effectively handles the increased input dimension in multi-agent reinforcement learning with communication and makes learning robust against communication errors in the execution phase. In the second application scenario of centralized training with decentralized execution, we particularly consider the application of the proposed messagedropout to Multi-Agent Deep Deterministic Policy Gradient (MADDPG), which uses a centralized critic to train a decentralized actor for each agent. We evaluate the proposed message-dropout technique for several games, and numerical results show that the proposed message-dropout technique with proper dropout rate improves the reinforcement learning performance significantly in terms of the training speed and the steady-state performance in the execution phase.

v2026.09.13