Arrow Research search

Author name cluster

Qingxin Xia

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

1 paper
1 author row

Possible papers

1

AAMAS Conference 2026 Conference Paper

OM 2 P: Offline Multi-Agent Mean-Flow Policy

  • Zhuoran Li
  • Xun Wang
  • Hai Zhong
  • Qingxin Xia
  • Lihua Zhang
  • Longbo Huang

Generative models, especially diffusion and flow-based models, have been promising in offline multi-agent reinforcement learning. However, integrating powerful generative models into this framework poses unique challenges. In particular, diffusion and flow-based policies suffer from low sampling efficiency due to their iterative generation processes, making them impractical in timesensitive or resource-constrained settings. To tackle these difficulties, we propose Offline Multi-Agent Mean-Flow Policy (OM2P), a novel offline MARL algorithm to achieve efficient one-step action generation. To address the misalignment between generative objectives and reward maximization, we introduce a reward-aware optimizationschemethatintegratesacarefully-designedmean-flow matching loss with Q-function supervision. Additionally, we design a generalized timestep distribution and a derivative-free estimation strategy to reduce memory overhead and improve training stability. Empirical evaluations on Multi-Agent Particle and MuJoCo benchmarks demonstrate that OM2P achieves superior performance, with up to a 3. 8× reduction in GPU memory usage and up to a 10. 1× speed-up in training time. Our approach represents the first to successfully integrate mean-flow model into offline MARL, paving the way for practical and scalable generative policies in cooperative multi-agent settings.

v2026.09.13