EAAI Journal 2026 Journal Article
Cognitive entropy proximal policy optimization for autonomous ship collision avoidance based on deep reinforcement learning
- Weijun Wang
- Mingjie Li
- Guoquan Chen
- Shenhua Yang
- Yongfeng Suo
- Xiuqin Guo
- Kunjie Wang
Owing to the complexity of maritime environments and the high-dimensional characteristics of ship collision avoidance decisions, existing deep reinforcement learning (DRL) algorithms exhibit substantial limitations in balancing the exploration–exploitation trade-off in policy learning. Underexplored preliminary stages of training result in premature convergence to suboptimal solutions, which directly affects the convergence rate and the decision-making performance of the intelligent agent. To address this challenge, a cognitive entropy proximal policy optimization (CEPPO) algorithm is proposed to quantify the internal learning uncertainty of the agent and adjust the intensity of exploration accordingly, facilitating a more informed and adaptive behavior throughout training. A three-phase adjustment mechanism is integrated: exploration is encouraged early to avoid local optima, adaptively moderated in the mid-phase to accelerate convergence, and reduced in the final phase to enhance policy stability and generalization. In this study, artificial intelligence (AI) is implemented via the proposed cognitive entropy proximal policy optimization algorithm, and the implemented artificial intelligence is subsequently applied to autonomous ship collision avoidance in complex maritime environments, thereby underscoring both the implementation and the application of artificial intelligence within autonomous maritime navigation. A high-fidelity simulation platform based on unreal engine is developed to emulate complex maritime scenarios. A multilayered reward mechanism is implemented within this environment to enhance the effectiveness of collision avoidance. The experimental results demonstrate that CEPPO achieves superior convergence and robustness compared to conventional baselines.