arXiv Science⌕ Search

arXiv · 2609.38162

Why Last-Iterate Scale-Invariant Regret Matching Converges Linearly?

Abstract

IREG-PRM+ normalizes the cumulative regret vector by its own norm and attains optimal regret without knowledge of the payoff scale. Run unmodified on zero-sum matrix games, it converges linearly in the last iterate, and no analysis explains why. The obstacle is that the algorithm has no fixed step size to analyze: the step size is a state variable, the inverse of a regret norm that the trajectory itself moves. Every proved linear rate for regret-matching dynamics comes from restarting or modifying the update. We identify the mechanism as norm saturation: the regret norm rises to a finite limit and freezes the step size. We prove that it always does, with an explicit bound, and that saturation forces the last-iterate Nash gap to vanish on every matrix game; pointwise convergence follows whenever the equilibrium is unique. Near a unique strictly complementary equilibrium the active support freezes in one step, and the one-round Jacobian on that support has a closed form. The last-iterate then converges linearly at a closed-form rate, provided one scale-invariant quantity stays below one: the saturated step size times the largest singular value of the value-centered payoff submatrix on the support. On the $216$-instance testbed, the $184$ instances with a resolvable limit all satisfy it. The same analysis gives a ratio certificate: observable norm-increment ratios bound the unobservable Nash-gap ratio up to a constant that enters once and does not accumulate with the iteration count. Its slope-two law holds on $96.1\%$ of the instances where the slope is measurable, and the same increment monitors progress in extensive-form games, where best-response passes can be scheduled sparsely. The code is available at https://github.com/lbn187/NormCert.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Boning Li, Longbo Huang. 2026-09-29. Why Last-Iterate Scale-Invariant Regret Matching Converges Linearly?. https://arxiv.org/abs/2609.38162

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Throttling Equilibria in Auction Markets

Throttling is a popular method of budget management for online ad auctions in which the platform modulates the participation probability of an advertiser in order to smoothly spend her budget across many auctions. In this work, we investigate the setting in which all of the advertisers simultaneously employ throttling to manage their budgets, and we do so for both first-price and second-price auctions. We analyze the structural and computational properties of the resulting equilibria. For first-price auctions, we show that a unique equilibrium always exists, is well-behaved and can be computed efficiently via tatonnement-style decentralized dynamics. In contrast, for second-price auctions, we prove that even though an equilibrium always exists, the problem of finding even an approximate equilibrium is PPAD-complete, there can be multiple equilibria, and it is NP-hard to find the revenue maximizing one. We also compare the equilibrium outcomes of throttling to those of multiplicative pacing, which is the other most popular and well-studied method of budget management. Finally, we characterize the Price of Anarchy of these equilibria for liquid welfare by showing that it is at most 2 for both first-price and second-price auctions, and demonstrating that our bound is tight.

cs.GT↗

Efficiency of Generalized Proportional First-Price Auctions Under Auto-bidding

Auto-bidding is now widely adopted in online advertising platforms, allowing advertisers to specify high-level campaign objectives--such as maximizing total value subject to a return-on-spend (ROS) constraint--rather than manual per-query bids. A central question in algorithmic mechanism design is characterizing the worst-case efficiency loss, or Price of Anarchy (PoA), across auction formats in this prior-free setting. While randomized auctions are known to strictly improve efficiency over deterministic mechanisms for two bidders, two fundamental questions have remained open: (1) what is the optimal PoA for two bidders, and (2) can any mechanism beat the barrier of 2 for general $n \ge 3$ bidders? We resolve both questions using the family of $r$-proportional first-price auctions ($\text{pFPA}_r$), in which each bidder wins with probability proportional to their bid raised to an exponent $r > 0$ and pays their bid upon winning. First, for two bidders, we prove that the standard proportional first-price auction ($r = 1$) achieves a tight $\text{PoA} \le 1.5$, complemented by a matching lower bound showing that no anonymous, monotone mechanism can do better. Second, for general $n \ge 2$ bidders, setting $r = 2n$ achieves $\text{PoA} \le 2 - \frac{1}{4n+1} = 2 - Ω(1/n)$ across all undominated bid profiles, breaking the deterministic barrier of 2 for every finite $n$ and asymptotically matching the known $2 - Θ(1/n)$ lower bound.

cs.GT↗

Honest Reporting in Scored Oversight: True-KL0 Property via the Prekopa Principle

We prove the True-KL$_0$ property for a parametric family of heterogeneous scoring rules arising in scored elicitation mechanisms (AI oversight, forecasting, expert surveys). An agent with private type $M>1$, scored through a $d$-dimensional outcome interface, reports to a principal who evaluates via a power-$p$ pseudospherical scoring rule, $p \in (d,d+1)$; $M$ captures the agent's information quality relative to a reference. Honest reporting is dominant-strategy optimal for every $d$ and every $p>1$, without a prior over the agent's type: a consequence of strict properness and identifiability, with a quadratic misreport-loss rate. True-KL$_0$, the property $R(M,p,d)<1$ for all $M>1$, $d \in \{2,3,4\}$, $p \in (d,d+1)$, is the quantitative core: $R$ is the Rayleigh quotient of the radial misreport channel of an annular oversight model, and True-KL$_0$ certifies a uniform curvature-domination margin for that channel: $1-R \ge 0.26$ ($R \le 0.7324$, semi-rigorous numerical certificate). Two structural tools drive the proof: (i) a substitution $y=(x+1)/(x-1)$ rewrites the loss integral $I_L$ as $\int_1^M F(y)(M^2-y^2)^{d/2} dy$ with $M$-independent weight $F(y)>0$; (ii) log-concavity of $I_L$ in $M$: algebraic for $d=2$ up to a small certified compact verification, via Prekopa's theorem plus semi-rigorous certificates for $d \in \{3,4\}$. True-KL$_0$ then follows from elementary tail bounds plus a certified bound on $M \in [1.001, 20]$. We also characterise the dimensional boundary: True-KL$_0$ holds for all $p \in (d,d+1)$ when $d \le 4$; $d=5$ is the unique transition, with $p_{crit}(5) \in [5.5718, 5.5750]$ (mpmath, not interval-certified); for $d=6,7$ (and conjecturally all $d \ge 6$) no threshold exists: the bound fails at every sampled $p \in (d,d+1)$.

cs.GT↗