Arrow Research search
Back to AAMAS

AAMAS 2026

Efficient Device-Cloud Collaborative Offline-to-Online Reinforcement Learning

Conference Paper Extended Abstracts Autonomous Agents and Multiagent Systems

Abstract

Federated reinforcement learning in device-cloud architectures suffers from high interaction costs, low sample efficiency, and slow convergence due to device constraints. We propose a data-centric device-cloud collaborative training method that pre-trains a global model on the cloud to provide a high-quality initial policy and selects high-value samples to guide safe and efficient local policy finetuning on devices. Experiments on public benchmarks show our method achieves significantly faster convergence, higher training efficiency, and improved policy stability, outperforming baselines.

Authors

Keywords

  • Device-cloudcollaborative
  • Federatedreinforcementlearning
  • Offlineto-online reinforcement learning
  • Prioritized experience replay

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
993905138122596036
v2026.09.13