UAI Conference 2016 Conference Paper
MDPs with Unawareness in Robotics
- Nan Rong
- Joseph Y. Halpern
- Ashutosh Saxena
while most actions result in the robot losing control and falling down. We formalize decision-making problems in robotics and automated control using continuous MDPs and actions that take place over continuous time intervals. We then approximate the continuous MDP using finer and finer discretizations. Doing this results in a family of systems, each of which has an extremely large action space, although only a few actions are “interesting”. We can view the decision maker as being unaware of which actions are “interesting”. We an model this using MDPUs, MDPs with unawareness, where the action space is much smaller. As we show, MDPUs can be used as a general framework for learning tasks in robotic problems. We prove results on the difficulty of learning a near-optimal policy in an an MDPU for a continuous task. We apply these ideas to the problem of having a humanoid robot learn on its own how to walk. Halpern, Rong, and Saxena [2010] (HRS from now on) defined MDPs with unawareness (MDPUs), where a decision-maker (DM) can be unaware of the actions in an MDP. In the robotics applications in which we are interested, we can think of the DM (e. g. , a humanoid robot) as being unaware of which actions are the useful actions, and thus can model what is going on using an MDPU.