arXiv · 1804.02884
Policy Gradient With Value Function Approximation For Collective Multiagent Planning
Abstract
Decentralized (PO)MDPs provide an expressive framework for sequential decision making in a multiagent system. Given their computational complexity, recent research has focused on tractable yet practical subclasses of Dec-POMDPs. We address such a subclass called CDEC-POMDP where the collective behavior of a population of agents affects the joint-reward and environment dynamics. Our main contribution is an actor-critic (AC) reinforcement learning method for optimizing CDEC-POMDP policies. Vanilla AC has slow convergence for larger problems. To address this, we show how a particular decomposition of the approximate action-value function over agents leads to effective updates, and also derive a new way to train the critic based on local reward signals. Comparisons on a synthetic benchmark and a real-world taxi fleet optimization problem show that our new AC approach provides better quality solutions than previous best approaches.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Duc Thien Nguyen, Akshat Kumar, Hoong Chuin Lau. 2018-04-09. Policy Gradient With Value Function Approximation For Collective Multiagent Planning. https://arxiv.org/abs/1804.02884
Cite the original work for its findings. Save a collection to share your selection of sources.