arXiv Science⌕ Search

arXiv · 2610.05428

Learning in Continuous Games from Pairwise Preference Feedback

Abstract

We study learning in continuous games when players receive only pairwise preference feedback, revealing which of two actions is preferred but neither payoff values nor preference magnitudes. We first show that standard external regret and coarse correlated equilibria (CCE) are not identifiable from this ordinal information: the same sequence of play can incur zero and linear regret in two ordinally equivalent games, while the distributions that remain CCE across all cardinal representations consistent with the same preferences are exactly those supported on pure Nash equilibria. Motivated by this gap, we develop a first-order ordinal theory based on normalized unilateral preference directions, introducing an ordinal directional regret benchmark and corresponding equilibrium notions. We show that block-normalized pseudogradient dynamics achieve sublinear ordinal regret and, under additional structure, Nash-convergence guarantees. We then use a single-comparison estimator to implement these dynamics from finite pairwise comparisons. With one comparison per player and round, the resulting algorithm achieves sublinear finite-resolution ordinal regret against arbitrary opponent behavior and, in ordinal potential games, almost-sure last-iterate convergence to the Nash set. Our results provide regret, dynamics, and equilibrium guarantees directly from preference feedback without reconstructing cardinal utilities.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Anas Barakat. 2026-10-04. Learning in Continuous Games from Pairwise Preference Feedback. https://arxiv.org/abs/2610.05428

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Co-opetition Equilibrium in Adversarial Team Games with Heterogeneous Utilities

The United Nations' 2030 Agenda for Sustainable Development requires all countries to collaborate against adversarial factors, a scenario that can be formalized as an adversarial team game. However, existing solution concepts assume team players share identical utility functions--an assumption inconsistent with real-world settings where countries have divergent objectives. This paper argues that studying adversarial team games must account for heterogeneous utilities among team players. We show that ignoring utility differences can render computed equilibria unstable in the original game, and we formalize this degradation via the \emph{Price of Ignoring Heterogeneity}, which can be unbounded. To address this, we introduce the Co-opetition Equilibrium (CoE), where team players with heterogeneous utilities correlate their strategies (cooperation) against an adversary (competition). We establish existence via reduction from Nash equilibrium, and prove that finding a CoE is PPAD-complete while computing a Team-Maximizing CoE (TMCoE) is NP-hard. Nevertheless, we identify a broad class of zero-sum adversarial team games--those satisfying a consistent-constraint condition that generalizes identical utilities--where TMCoEs are exchangeable and computable in polynomial time via linear programming. We outline directions for algorithm design, MARL, and extensions to extensive-form and multi-team games.

cs.GT↗

Last-Iterate Convergence of Policy Dynamics in Zero-Sum Networked Separable Markov Games

Solving Nash equilibria for general multi-player Markov games is computationally intractable, while two-player zero-sum Markov games admit fast last-iterate policy-optimization methods. Zero-sum networked separable Markov games occupy an important middle ground: they retain global multi-player competition structure through pairwise interactions, while preserving computational tractability of Nash equilibria (NE) in the finite-horizon setting. Existing algorithms for this class either proceed through equilibrium-collapse arguments for a simplified setting where a single controller determines the transition probability, or backward dynamic programming that relies on equilibrium solvers at each stage. However, the design and analysis of direct policy-update approaches remain inadequate. To address this issue, we propose the entropy-regularized optimistic multiplicative weights update (ER-OMWU), a complementary single-loop policy dynamic that updates players' policies symmetrically and returns an approximate NE in the last iteration. We provide the first last-iterate convergence analysis of policy dynamics in the games of interest: after $\widetilde O\left(1/ε\right)$ iterations, the returned policy is an $ε$-approximate Nash equilibrium. The result preserves the near-linear convergence rate achieved by policy optimization in two-player zero-sum Markov games, but extends the policy-dynamics viewpoint to a more complicated but structured multi-player setting.

cs.GT↗

Auction Design with ROI-Constrained Bidders: Truthfulness and Revenue Maximization

The return-on-investment (ROI) constraint is central to many auctions, particularly in online advertising, where a bidder is unwilling to pay more than a fixed fraction of the value obtained. We study truthful and revenue-maximizing auctions for ROI-constrained bidders. We first characterize truthful auctions when both valuations and ROI constraints are private, showing that the allocation rule uniquely determines the payment rule. Building on this characterization, for multiple bidders we introduce $σ$-increment mechanisms that resemble Myerson's optimal mechanism~\cite{journals/mor/Myerson81}; as $σ$ vanishes, these mechanisms become asymptotically optimal among deterministic truthful mechanisms, and their revenue approaches at least a $1/\bar r$ fraction of the optimal expected revenue over all truthful mechanisms, where $\bar r$ is the largest possible ROI constraint. In the single-bidder setting, we prove that every truthful auction can be replaced by a convex pricing function with weakly higher payments for every type, and we derive the optimal pricing functions when either the valuation or the ROI constraint is public.

cs.GT↗