Arrow Research search
Back to AIJ

AIJ 2015

POMDPs under probabilistic semantics

Journal Article journal-article Artificial Intelligence

Abstract

We consider partially observable Markov decision processes (POMDPs) with limit-average payoff, where a reward value in the interval [ 0, 1 ] is associated with every transition, and the payoff of an infinite path is the long-run average of the rewards. We consider two types of path constraints: (i) a quantitative constraint defines the set of paths where the payoff is at least a given threshold λ 1 ∈ ( 0, 1 ]; and (ii) a qualitative constraint which is a special case of the quantitative constraint with λ 1 = 1. We consider the computation of the almost-sure winning set, where the controller needs to ensure that the path constraint is satisfied with probability 1. Our main results for qualitative path constraints are as follows: (i) the problem of deciding the existence of a finite-memory controller is EXPTIME-complete; and (ii) the problem of deciding the existence of an infinite-memory controller is undecidable. For quantitative path constraints we show that the problem of deciding the existence of a finite-memory controller is undecidable. We also present a prototype implementation of our EXPTIME algorithm and experimental results on several examples.

Authors

Keywords

  • POMDPs
  • Limit-average objectives
  • Almost-sure winning

Context

Venue
Artificial Intelligence
Archive span
1970-2026
Indexed papers
3976
Paper id
142030935502333462
v2026.09.13