Arrow Research search
Back to JBHI

JBHI 2026

Value Decomposition-Based Multi-Agent Learning for Anesthetics Collaborative Control

Journal Article journal-article Artificial Intelligence ยท Biomedical and Health Informatics

Abstract

Automated control of personalized multiple anesthetics in clinical Total Intravenous Anesthesia (TIVA) is crucial yet challenging. Current systems, including target-controlled infusion (TCI) and closed-loop systems, either rely on relatively static pharmacokinetic/pharmacodynamic (PK/PD) models or focus on single anesthetic control. So they limit both personalization and collaborative control. To address these issues, we propose a novel V alue D ecomposition M ulti- A gent D eep R einforcement L earning (VD-MADRL) framework based on Markov Game (MG) for P ersonalized M ultiple A nesthetics C ontrol in a C losed- L oop system (PMAC-CL). VD-MADRL optimizes the collaboration between two anesthetics propofol (Agent I) and remifentanil (Agent II) by leveraging a MG to identify optimal actions among heterogeneous agents. We employ various value function decomposition methods to resolve the credit allocation problem and enhance collaborative control. We also introduce a multivariate environment model based on random forest (RF) for anesthesia state simulation. To ensure data validity, we design a data resampling and alignment technique to synchronize trajectory data from different devices, avoiding gradient explosion and maintaining conformity to Markov property. Extensive experiments on general and thoracic surgery datasets demonstrate that VD-MADRL provides more refined dose adjustments and maintains multiple anesthesia state indicators more stably at target levels compared to human experience. Especially, the best-performing algorithm, VDN in general surgery with online training, achieved a 16. 4% increase in cumulative reward (CR) and a 58. 0% reduction in mean MDPE compared to human experience. This demonstrates its great clinical value.

Authors

Keywords

  • Anesthesia
  • Collaboration
  • Closed loop systems
  • Data models
  • Trajectory
  • Surgery
  • Brain modeling
  • Real-time systems
  • Monitoring
  • Deep reinforcement learning
  • Multi-agent Reinforcement Learning
  • Value Function
  • Random Forest
  • Human Experience
  • Cardiac Surgery
  • Multiple Indicators
  • Closed-loop System
  • Dose Adjustment
  • General Surgery
  • Trajectory Data
  • Optimal Action
  • Resampled Data
  • Gradient Explosion
  • Heterogeneous Agents
  • Total Intravenous Anesthesia
  • Target-controlled Infusion
  • Clinical Anesthesia
  • Deep Learning
  • Time Step
  • Mean Arterial Pressure
  • BIS Values
  • Effects Of Anesthesia
  • Effects Of Different Modes
  • Machine Learning Models
  • Vital Sign Data
  • Markov Decision Process
  • Depth Of Anesthesia
  • Anesthetic Agents
  • Effect-site Concentration
  • Multi-agent deep reinforcement learning
  • value function decomposition
  • multiple anesthesia states
  • personalized anesthesia
  • Humans
  • Propofol
  • Markov Chains
  • Remifentanil
  • Algorithms
  • Anesthesia, Intravenous
  • Anesthetics
  • Anesthetics, Intravenous

Context

Venue
IEEE Journal of Biomedical and Health Informatics
Archive span
2013-2026
Indexed papers
6337
Paper id
208034176986074606
v2026.09.13