arXiv Science⌕ Search

arXiv · 2610.05516

Distributed Algorithms for $α$-Potential Functions in General-Sum Games

Abstract

We study the problem of computing the tightest \(α\)-potential approximation of a general-sum game over continuous action spaces, within a prescribed class of potential functions and when each player has access only to its own utility function. The difficulty is twofold: the approximation error involves a worst-case search over an infinite set of unilateral deviations, and the required utility information is distributed across players. For a linear-in-parameters potential class, we use an exact finite-tuple reformulation that separates the problem into a global outer search over deviation tuples and distributed convex inner problems. We develop a primal--dual inner oracle tailored to this structure and establish a uniform one-sided accuracy guarantee. This oracle can be combined with global outer search to obtain an end-to-end guarantee on the outer optimization error. We also develop a projected zeroth-order outer method as a computationally lighter alternative for higher-dimensional problems. Numerical experiments illustrate the accuracy--computation tradeoff between the two outer-search methods and show that the proposed optimization framework can improve upon analytical \(α\)-potential constructions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yifei Chen, Chinmay Maheshwari. 2026-10-04. Distributed Algorithms for $α$-Potential Functions in General-Sum Games. https://arxiv.org/abs/2610.05516

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards Strategy-Level RSI for Skill-Augmented Agents: Learning When to Reuse Skills from Execution Feedback

Long-running agents accumulate reusable Skills, but a Skill that is semantically relevant to a task is not necessarily worth loading in the current state. We study the applicability question that arises once a candidate Skill is known: should it be loaded in the current state? We propose SkillApt, which uses matched WITH/WITHOUT Skill executions on the same task state as persistent evidence, estimates the conditional marginal utility of the Skill, and chooses LOAD or ABSTAIN accordingly. The base model, agent architecture, and Skill contents stay fixed; only the external deployment policy changes. We call this constrained setting strategy-level recursive self-improvement (Strategy-Level RSI). On 20 Skills and 160 held-out states, as paired evidence accumulates, SkillApt's task success rises from 81.9% under a cold start to 91.3%, matching a strong zero-shot LLM controller; yet SkillApt activates Skills on only 26.3% of states, versus 98.8% for the zero-shot controller. A hard-candidate study shows that non-optimal Skills mostly leave correctness unchanged while raising execution cost, and occasionally cause correctness harm. An ablation shows that a history recording only WITH success makes the policy load almost everywhere, whereas paired evidence substantially improves selectivity. These results indicate that relevance is not applicability: the main effect of execution evidence is not to make the model stronger but to change how existing Skills are deployed, moving the system from near-always loading to selective reuse.

cs.MA↗

Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions

Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing

cs.MA↗

ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration

Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.

cs.MA↗