AAMAS 2026
Efficient Device-Cloud Collaborative Offline-to-Online Reinforcement Learning
Abstract
Federated reinforcement learning in device-cloud architectures suffers from high interaction costs, low sample efficiency, and slow convergence due to device constraints. We propose a data-centric device-cloud collaborative training method that pre-trains a global model on the cloud to provide a high-quality initial policy and selects high-value samples to guide safe and efficient local policy finetuning on devices. Experiments on public benchmarks show our method achieves significantly faster convergence, higher training efficiency, and improved policy stability, outperforming baselines.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 993905138122596036