EAAI Journal 2026 Journal Article
A trend-aware reinforcement learning approach for adaptive motion planning of robotic manipulators in dynamic environments
- Dexian Wang
- Peng Zhang
- Pengfei Ding
- Junliang Wang
- Jie Zhang
In engineering applications such as service robotics, human–robot collaboration, and industrial automation, robotic manipulators frequently operate in highly dynamic and partially observable environments. The presence of dynamic obstacles, unpredictable human behaviors, and rapidly changing task demands intensify the inherent conflicts between positioning accuracy, obstacle avoidance, and operational safety. At the same time, these systems must rely on incomplete and noisy sensor data to perceive and interpret their surroundings. Such dynamic scenarios significantly increase system complexity and require adaptive, real-time control strategies capable of making reliable decisions under uncertainty. To address these challenges, this paper proposes Trend Learning – Adaptive Reward Shaping – Temporal Difference Knowledge Distillation of Q-value(Action-Value Function), collectively referred to as TL-ARS-TDKDQ, a reinforcement learning framework designed to enable robotic manipulators to adaptively perform precise positioning and dynamic obstacle avoidance in dynamic environments. Trend Learning(TL) alleviates environmental uncertainty by extracting temporal dependencies from sequential data of the manipulator. Adaptive Reward Shaping(ARS) dynamically balances positioning accuracy and obstacle avoidance for the robotic manipulator while progressively increasing task difficulty via curriculum learning. To enhance stability during reward fluctuations caused by ARS, Temporal Difference Knowledge Distillation Q-value (TDKDQ) employs a dynamic teacher network and Temporal Difference(TD) error-based balancing, ensuring stable policy convergence in non-stationary scenarios involving robotic manipulator control. Experiments with a KUKA arm in CoppeliaSim demonstrate that TL-ARS-TDKDQ significantly improves convergence speed, control stability, and task success when integrated into mainstream continuous control reinforcement learning algorithms.