Arrow Research search
Back to IROS

IROS 2021

Cooperative Assistance in Robotic Surgery through Multi-Agent Reinforcement Learning

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Cognitive cooperative assistance in robot-assisted surgery holds the potential to increase quality of care in minimally invasive interventions. Automation of surgical tasks promises to reduce the mental exertion and fatigue of surgeons. In this work, multi-agent reinforcement learning is demonstrated to be robust to the distribution shift introduced by pairing a learned policy with a human team member. Multi-agent policies are trained directly from images in simulation to control multiple instruments in a sub task of the minimally invasive removal of the gallbladder. These agents are evaluated individually and in cooperation with humans to demonstrate their suitability as autonomous assistants. Compared to human teams, the hybrid teams with artificial agents perform better considering completion time (44. 4% to 71. 2% shorter) as well as number of collisions (44. 7% to 98. 0% fewer). Path lengths, however, increase under control of an artificial agent (11. 4% to 33. 5% longer). A multi-agent formulation of the learning problem was favored over a single-agent formulation on this surgical sub task, due to the sequential learning of the two instruments. This approach may be extended to other tasks that are difficult to formulate within the standard reinforcement learning framework. Multi-agent reinforcement learning may shift the paradigm of cognitive robotic surgery towards seamless cooperation between surgeons and assistive technologies.

Authors

Keywords

  • Laparoscopes
  • Minimally invasive surgery
  • Medical robotics
  • Instruments
  • Reinforcement learning
  • Gallbladder
  • Fatigue
  • Multi-agent Reinforcement Learning
  • Path Length
  • Minimally Invasive
  • Completion Time
  • Assistive Technology
  • Intelligence Agencies
  • Policy Learning
  • Reinforcement Learning Framework
  • Intelligent Robots
  • Surgical Tasks
  • Degrees Of Freedom
  • Time Step
  • Number Of Steps
  • Learning Curve
  • Human Interaction
  • Long Short-term Memory
  • Simulation Environment
  • Learning Phase
  • Laparoscopic Surgery
  • Reward Function
  • Proximal Policy Optimization
  • Single Surgeon
  • Laparoscopic Cholecystectomy
  • Laparoscopic Instruments
  • Surgical Instruments
  • Evaluation Team
  • Team Composition
  • Number Of Instruments
  • Increase In Path Length
  • Policy Model

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
867685431441667463
v2026.09.13