Arrow Research search
Back to IROS

IROS 2025

Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Reinforcement learning (RL) is a promising approach for robotic navigation, allowing robots to learn through trial and error. However, real-world robotic tasks often suffer from sparse rewards, leading to inefficient exploration and suboptimal policies due to sample inefficiency of RL. In this work, we introduce Confidence-Controlled Exploration (CCE), a novel method that improves sample efficiency in RL-based robotic navigation without modifying the reward function. Unlike existing approaches, such as entropy regularization and reward shaping, which can introduce instability by altering rewards, CCE dynamically adjusts trajectory length based on policy entropy. Specifically, it shortens trajectories when uncertainty is high to enhance exploration and extends them when confidence is high to prioritize exploitation. CCE is a principled and practical solution inspired by a theoretical connection between policy entropy and gradient estimation. It integrates seamlessly with on-policy and off-policy RL methods and requires minimal modifications. We validate CCE across REINFORCE, PPO, and SAC in both simulated and real-world navigation tasks. CCE outperforms fixed-trajectory and entropy-regularized baselines, achieving an 18% higher success rate, 20-38% shorter paths, and 9. 32% lower elevation costs under a fixed training sample budget. Finally, we deploy CCE on a Clearpath Husky robot, demonstrating its effectiveness in complex outdoor environments.

Authors

Keywords

  • Training
  • Uncertainty
  • Costs
  • Navigation
  • Estimation
  • Reinforcement learning
  • Entropy
  • Trajectory
  • Robots
  • Intelligent robots
  • Policy Learning
  • Robot Navigation
  • Gradient Approximation
  • Sampling Efficiency
  • Reward Function
  • Trajectory Length
  • Navigation Task
  • Policy Gradient
  • Robotic Tasks
  • Theoretical Connections
  • Entropy Regularization
  • Path Length
  • State Space
  • Feed-forward Network
  • Path Planning
  • Mixing Time
  • Real-world Environments
  • Markov Decision Process
  • Reinforcement Learning Algorithm
  • Uneven Terrain
  • Robotic Agents
  • Real Robot
  • Intrinsic Rewards
  • Gradient Update
  • Average Path Length
  • Elevation Gain
  • Navigation Performance
  • Reward Structure
  • Elevational Gradient

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
742276651213832065
v2026.09.13