Optimizing Local Satisfaction of Long-Run Average Objectives in Markov Decision Processes

1
Citations
#820
in AAAI 2024
of 2289 papers
5
Authors
1
Data Points

Abstract

Long-run average optimization problems for Markov decision processes (MDPs) require constructing policies with optimal steady-state behavior, i.e., optimal limit frequency of visits to the states. However, such policies may suffer from local instability, i.e., the frequency of states visited in a bounded time horizon along a run differs significantly from the limit frequency. In this work, we propose an efficient algorithmic solution to this problem.

Citation History

Jan 28, 2026
1