PRL Workshop 2024 Workshop Paper
Online Planning in MDPs with Stochastic Durative Actions
- Tal Berman
- Ronen Brafman
- Erez Karpas
5 10 15 20 25 30 35 40 Markov Decision Processes (MDPs) are a popular model for probabilistic planning. Actions in MDPs are applied sequentially, and their effects are instantaneous. Yet, real-world scenarios often involve actions with duration and parallel action execution. This paper considers CoMDPs, a model that extends MDPs with durative, concurrent actions, and describes TP-MCTS, an online algorithm for solving CoMDPs that combines Monte Carlo Tree Search (MCTS) with techniques used in classical temporal planning. TP-MCTS uses a compilation of durative actions to Start and End actions and enhances each tree node with a Simple Temporal Network to maintain temporal consistency and schedule the plan’s action. Our empirical evaluation demonstrates the efficacy of the TPMCTS algorithm in tackling CoMDPs.