Modified monotone policy iteration for interpretable policies in Markov decision processes and the impact of state ordering rules.

Saved in:
Bibliographic Details
Title: Modified monotone policy iteration for interpretable policies in Markov decision processes and the impact of state ordering rules.
Authors: Lee, Sun Ju1 (AUTHOR) julee@gatech.edu, Gong, Xingyu1 (AUTHOR) xgong75@gatech.edu, Garcia, Gian-Gabriel1 (AUTHOR) giangarcia@gatech.edu
Source: Annals of Operations Research. Apr2025, Vol. 347 Issue 2, p783-841. 59p.
Subjects: Markov processes, Random sets, Dynamic programming, Algorithms, Integers
Abstract: Optimizing interpretable policies for Markov decision processes (MDPs) can be computationally intractable for large-scale MDPs, e.g., for monotone policies, the optimal interpretable policy depends on the initial state distribution, precluding standard dynamic programming techniques. Previous work has proposed monotone policy iteration (MPI) to produce a feasible solution for warm starting a mixed integer linear program that finds an optimal monotone policy. However, this prior work did not investigate the convergence and optimality of this algorithm, nor did they investigate the impact of state ordering rules, i.e., the order in which policy improvement steps are performed in MPI. In this study, we analytically characterize the convergence and optimality of the MPI algorithm, introduce a modified MPI (MMPI) algorithm, and show that our algorithm improves upon the MPI algorithm. To test MMPI numerically, we conduct experiments in two settings: (1) perturbations of a machine maintenance problem wherein the optimal policy is guaranteed to be monotone or near-monotone and (2) randomly generated MDPs. We propose and investigate 19 state ordering rules for MMPI based on each state's value function, initial probability, and stationary distribution. Computational results reveal a trade-off between computational time and optimality gap; in the structured machine maintenance setting, the fastest state ordering rules still yield high quality policies while the trade-off is more pronounced in the random MDP setting. Across both settings, the random state ordering rule performs the best in terms of optimality gap (less than approximately 5% on average) at the expense of computational time. [ABSTRACT FROM AUTHOR]
Copyright of Annals of Operations Research is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Optimizing interpretable policies for Markov decision processes (MDPs) can be computationally intractable for large-scale MDPs, e.g., for monotone policies, the optimal interpretable policy depends on the initial state distribution, precluding standard dynamic programming techniques. Previous work has proposed monotone policy iteration (MPI) to produce a feasible solution for warm starting a mixed integer linear program that finds an optimal monotone policy. However, this prior work did not investigate the convergence and optimality of this algorithm, nor did they investigate the impact of state ordering rules, i.e., the order in which policy improvement steps are performed in MPI. In this study, we analytically characterize the convergence and optimality of the MPI algorithm, introduce a modified MPI (MMPI) algorithm, and show that our algorithm improves upon the MPI algorithm. To test MMPI numerically, we conduct experiments in two settings: (1) perturbations of a machine maintenance problem wherein the optimal policy is guaranteed to be monotone or near-monotone and (2) randomly generated MDPs. We propose and investigate 19 state ordering rules for MMPI based on each state's value function, initial probability, and stationary distribution. Computational results reveal a trade-off between computational time and optimality gap; in the structured machine maintenance setting, the fastest state ordering rules still yield high quality policies while the trade-off is more pronounced in the random MDP setting. Across both settings, the random state ordering rule performs the best in terms of optimality gap (less than approximately 5% on average) at the expense of computational time. [ABSTRACT FROM AUTHOR]
ISSN:02545330
DOI:10.1007/s10479-024-06158-3