arXiv · 2609.18829
Epsilon-Nash Equilibria in History-Dependent SA-MDPs
Abstract
We study state-adversarial Markov decision processes (SA-MDP) as a game of observation-space attacks: at each step, an agent selects an action from a received observation while an adversary$\unicode{x2014}$who knows the true state the agent is in$\unicode{x2014}$chooses a perturbed observation within a state-dependent proximity set. While existing work focuses on Markovian policies, we develop a solution concept and computational approach for SA-MDPs under history dependence. This is motivated by results showing that history dependence can materially change equilibrium outcomes and can force both the agent and the adversary to adapt their strategies. First, we prove the non-existence of universal (agnostic of the initial state distribution) history-dependent equilibrium policies. In response to this finding, our main result presents the first algorithmic route to computing $ε$-approximations of initial-state dependent equilibria. We do so by reducing SA-MDPs to a strategically equivalent constrained zero-sum one-sided partially observable stochastic game. We conclude by testing our algorithm on small analytically verifiable games and showing it scales to larger, more realistic benchmarks, including Atari Freeway rollouts with a 12-period ahead horizon.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Brandon Gary Kaplowitz, Dominik Bohnet Zurcher, Akash Agrawal, Tala Jafari, Christian Schroeder de Witt, Paul W. Goldberg. 2026-09-16. Epsilon-Nash Equilibria in History-Dependent SA-MDPs. https://arxiv.org/abs/2609.18829
Cite the original work for its findings. Save a collection to share your selection of sources.