Arrow Research search
Back to ECAI

ECAI 2025

Training Robotic Self-Evolving with GRPO

Conference Paper Accepted Paper Artificial Intelligence

Abstract

Current embodied robots heavily depend on pre-trained models, whose capabilities are inherently constrained by the data they were originally trained on. However, truly intelligent robots are expected to improve themselves autonomously when encountering novel environments where these pre-trained models fall short. This is the capability we define as self-evolving ability. In this paper, we investigate the self-evolving capacity of robotic vision models. Specifically, we simulate this process using the R3ED dataset and propose a training framework in which a policy learns to navigate through unfamiliar environments to collect informative data that can be used to refine the vision model. Our training pipeline is built upon the GRPO algorithm and incorporates historical states into the policy design to enhance contextual awareness. Furthermore, we introduce a novel reward mechanism based on supervision discrepancy to guide effective data collection. Experimental results validate the effectiveness of our proposed reinforcement training strategy. Our work highlights the potential of designing intelligent robots that can improve themselves without the intervene of human beings. Nevertheless, we acknowledge that robotic self-evolving remains a nascent and underexplored area, with significant room for further future research and the discovery of more optimal approaches.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
European Conference on Artificial Intelligence
Archive span
1982-2025
Indexed papers
5223
Paper id
843849846268488577
v2026.09.13