arXiv · 2402.07182
Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning
Abstract
An important challenge in multi-objective reinforcement learning is obtaining a Pareto front of policies to attain optimal performance under different preferences. We introduce Iterated Pareto Referent Optimisation (IPRO), which decomposes finding the Pareto front into a sequence of constrained single-objective problems. This enables us to guarantee convergence while providing an upper bound on the distance to undiscovered Pareto optimal solutions at each step. We evaluate IPRO using utility-based metrics and its hypervolume and find that it matches or outperforms methods that require additional assumptions. By leveraging problem-specific single-objective solvers, our approach also holds promise for applications beyond multi-objective reinforcement learning, such as planning and pathfinding.
Explore related subjects
Keep this discovery
Willem Röpke, Mathieu Reymond, Patrick Mannion, Diederik M. Roijers, Ann Nowé, Roxana Rădulescu. 2024-02-11. Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning. https://arxiv.org/abs/2402.07182
Cite the original work for its findings. Save a collection to share your selection of sources.