EAAI Journal 2026 Journal Article
Clinician-informed offline reinforcement learning for vasopressor administration optimization in shock management
- Feier Qiu
- Ying Chen
- Xiuxian Wang
- Na Geng
- Zhitao Yang
Vasopressor treatment strategies are essential for managing shock patients, yet determining optimal type, dosage, and timing of vasopressors remains challenging given variable clinician expertise and patient conditions. Applying existing reinforcement learning (RL) algorithms in treatment decision making risks Q-value overestimation and ignores the gap between artificial intelligence (AI)-driven recommendations and established clinical practices. This work introduces an offline RL algorithm called Safe Conservative Q-learning (SafeCQL), integrating conservative regularization and clinician-informed safety constraints. Patient trajectories from a large real-world intensive care database are modeled as a Markov decision process (MDP) to optimize vasopressor administration. Off-policy evaluations with model-based and model-free estimation methods demonstrate the superior performance of SafeCQL over clinician policies and existing RL models in improving survival rates and policy robustness. Findings show that SafeCQL improves the survival rate by 5. 1% and raises expected returns from 45. 78 to 69. 15 (model-based) and 44. 33 to 61. 24 (model-free). SafeCQL can derive a treatment policy that aligns closely with clinician preferences while surpassing clinician expertise in decision quality. This work offers a deployable solution for personalized vasopressor management in dynamic critical care environments.