日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
arXiv:2007.00085

Enforcing Almost-Sure Reachability in POMDPs

Enforcing Almost-Sure Reachability in POMDPs

シェア:XThreadsFacebookLINEはてブBluesky

著者: Sebastian Junges, Nils Jansen, Sanjit A. Seshia

分類: cs.AI, cs.RO, cs.SY, eess.SY

原文アブストラクト

Partially-Observable Markov Decision Processes (POMDPs) are a well-known stochastic model for sequential decision making under limited information. We consider the EXPTIME-hard problem of synthesising policies that almost-surely reach some goal state without ever visiting a bad state. In particular, we are interested in computing the winning region, that is, the set of system configurations from which a policy exists that satisfies the reachability specification. A direct application of such a winning region is the safe exploration of POMDPs by, for instance, restricting the behavior of a reinforcement learning agent to the region. We present two algorithms: A novel SAT-based iterative approach and a decision-diagram based alternative. The empirical evaluation demonstrates the feasibility and efficacy of the approaches.