arXiv · 1908.08773
Opponent Aware Reinforcement Learning
Abstract
In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit. We introduce Threatened Markov Decision Processes (TMDPs) as a framework to support an agent against potential opponents in an RL context as well as schemes resulting in novel learning approaches to deal with TMDPs. After introducing our framework and deriving theoretical results, empirical evidence is given via extensive experiments, showing the importance for an RL agent of acknowledging adversarial awareness.
Explore related subjects
Keep this discovery
Victor Gallego, Roi Naveiro, David Rios Insua, David Gomez-Ullate Oteiza. 2019-08-22. Opponent Aware Reinforcement Learning. https://doi.org/10.1016/j.ejor.2026.08.031
Cite the original work for its findings. Save a collection to share your selection of sources.