arXiv ScienceSearch

SEARCH · arXiv Science

Results for “math.GT”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,025 records · Page 3Linked to original sources

Turn-Based Combat Arena: A New Framework for Multiagent Training and Game Balancing

This paper is the first in a series on Turn-Based Combat Arena, a configurable framework for turn-based strategy games designed to support the efficient training and evaluation of machine learning agents. The proposed framework enables flexible modification of game rules and parameters, allowing rapid experimentation across diverse scenarios. Its architecture is optimized for high-throughput simulation, supporting tens of thousands of games per second and enabling the storage and processing of billions of gameplay records on a single machine. The problem of balancing the game, and particularly the parameters of game units, is investigated in detail. We evaluate several optimization approaches and show that multiple methods converge to comparable solutions, suggesting robustness in identifying balanced game configurations. These results indicate that the framework can serve as a practical platform for both game design analysis and agent training.

cs.GT

Scale-robust Auctions

We study auctions that are robust at any scale, i.e., they can be applied to sell both expensive and cheap items and achieve the best multiplicative approximation of the optimal revenue in the worst case. We first show that it is without loss of optimality to restrict attention to scale-invariant mechanisms whenever the family of possible distributions is closed under every positive rescaling. This conclusion uses no regularity or other distributional shape restriction. We then solve the two-agent, single-item problem with values drawn i.i.d. from an unknown regular distribution when only a high value bidder can receive a positive allocation. The robustly optimal mechanism in this class randomizes between the second-price auction, with probability approximately 0.806, and a markup auction that offers the item to the highest-valued bidder at a price equal to 2.447 times the second-highest value. Its worst-case approximation ratio is approximately 1.907.

cs.GT

The Art of Calling the Winner by Asking Just Enough Questions: Competitive Preference Elicitation with Next-Best Queries

We study active elicitation of agent preferences for collectively choosing among $m$ alternatives using prominent voting rules. We focus on the next-best query model, in which an agent responds to a query by revealing their next favorite alternative, and measure the competitive ratio, which is the worst-case ratio between the number of queries made by the active elicitation algorithm and the minimum number of queries needed to reveal the winning alternative(s) in hindsight. We show that sublinear competitive ratios are achievable for many positional scoring rules, whereas every Condorcet-consistent rule has competitive ratio linear in $m$. For Borda count, we develop two complementary techniques: level-wise pruning, whose analysis extends to general concave scoring rules, and multi-scale score thresholding, which gives an $O(\sqrt m)$ worst-case guarantee for Borda. We also demonstrate strong empirical performance of level-wise pruning on real data.

cs.GT

Batched Pandora's Box

Motivated by numerous parallelizable stochastic search problems, most notable and timely among them being LLM inference-time scaling, we propose and study batched versions of the Pandora's Box problem of Weitzman. In particular, boxes are opened in capacity-constrained batches, each batch has a setup cost, and all rewards in a batch are revealed together. We consider two different variants, motivated by different application environments: one where boxes are reusable (i.e., can provide multiple i.i.d.~samples) and another where they are not. For both variants we rule out most ``simple'' natural heuristics, and also formally prove NP-hardness of approximation in the traditional sense. We then relax the problem to allow bi-criteria approximations, with respect to both rewards and setup costs, where we exhibit constant approximation algorithms for both the reusable and non-reusable settings. This is obtained through a linear-programming relaxation of Pandora's Box problem, followed by randomized or Pipage rounding.

cs.DS

Sparse Disapproval Guarantees a Nonempty Hare Core

An approval committee is Hare-core stable if no coalition meeting the Hare quota can strictly improve by moving to another candidate set. Whether every approval election has such a committee remains open. We prove nonemptiness when each voter disapproves at most two candidates, with no bounds on the numbers of candidates, seats, or voter types. The result also permits arbitrary positive rational voter weights. Our deterministic rule represents a committee by its missing set. It first maximizes weighted coverage of two-candidate disapproval sets and then maximizes total disapproval incidence. An exact coverage inequality excludes targets one seat below the committee. The incidence objective excludes unanimous equal-size targets, while targets of size at most $k-2$ cannot improve any voter. Two implementation-level independent verifiers audit overlapping finite grids. The symbolic proof, not this bounded enumeration, establishes the theorem's unbounded quantifiers. The argument identifies complement-side coverage as a tractable mechanism for a broad parameter range within a sharply defined preference domain.

cs.GT

The Complexity of Justified Representation with Additive Utilities

We study the computational complexity of satisfying proportional representation -- in particular proportional, extended, and fully justified representation (PJR, EJR, and FJR) -- in participatory budgeting and committee elections with additive utilities. First, we give a complete picture of the complexity of the axioms for a constant number of voters or voter types. Second, we show that even for committee elections with integer utilities bounded above by a small constant, satisfying FJR is intractable, giving the first strong NP-hardness result for a justified representation axiom. Third, we extend the Expanding Approvals Rule to committee elections with additive utilities and show that it satisfies PJR. Lastly, we show that no sequential voting rule can improve on the known positive result, thus proving that novel, substantially different voting rules are needed to surpass these boundaries. Beyond their theoretical merit, our results carry practical importance, as multi-winner voting with additive utilities has recently been gaining prominence in online deliberation and real-world participatory budgeting.

cs.GT

Fair Stable Matching: A Nash Social Welfare Approach

While traditional stable matching algorithms, such as the Gale-Shapley algorithm, prioritize stability, they may fall short of achieving equitable outcomes among participants. We study the role of \emph{Nash social welfare} (NSW) as a fairness objective in the classic \emph{stable marriage problem}. We develop \texttt{SNSW-Alg} that finds a stable matching that maximizes Nash social welfare under rank-induced utilities in $\tilde{\mathcal{O}}(n^4)$ time, where $n$ is the number of men or women. We demonstrate that \texttt{SNSW-Alg} balances equity while preserving stability. We empirically evaluate our methods across diverse preference distributions, demonstrating significant gains in fairness without substantial losses in other key measures such as regret, egalitarian criterion, and sex equality. Our findings suggest that the stable matching produced by \texttt{SNSW-Alg} is statistically Pareto-undominated by stable matchings based on other fairness measures - regret, egalitarian, and sex equality. This study offers compelling insights for designing fair-stable matching.

cs.GT

Test-time Reinforcement Learning in Imperfect Information Games

Test-time reasoning has significantly improved performance in domains ranging from games to language models. However, test-time policy changes with formal guarantees on the performance of the resulting strategy remain a challenge in two-player zero-sum imperfect-information games. Existing solutions are limited to tabular methods or single gradient step updates. In this work, we investigate policy-gradient algorithms as a method for scalable test-time reasoning. We extend the concept of gadget game, tabular technique for test-time search, to the reinforcement learning setting. Unlike prior approaches, we represent the gadget game implicitly by modified sampling and neural policy rather then explicitly by constructing it, thereby removing constraints on subgame size. Furthermore, we formally prove that, unlike prior tabular algorithms, regularized policy-gradient algorithms limit possible strategy degradation caused by test-time reasoning, even without the gadget games. Our evaluation across small- and large-scale games confirms that additional test-time training often substantially improves performance relative to the blueprint strategy.

cs.GT

Utility-Driven Spatial Data Sampling for UAV-Assisted Scientific Smart Farming

Large smart-farming deployments generate continuous scientific data from spatially distributed sensors, including soil, humidity, temperature, crop-health, and pest-related measurements. In vast agricultural fields, however, an energy-constrained unmanned aerial vehicle (UAV) often cannot collect data from every sensor during each mission. Existing UAV-assisted collection methods typically optimize coverage, route length, data volume, or freshness, but they do not always distinguish between data that is merely available and data that is scientifically valuable. This poster introduces a utility-driven spatial sampling framework for UAV-assisted smart farming. The field is partitioned into grid cells, each sized according to the UAV ground coverage range. After an initial exploration phase, each cell receives a scientific utility score based on freshness, redundancy, anomaly likelihood, and model uncertainty. The UAV then selects and visits a subset of high-utility cells under battery and return-to-base constraints. The proposed framework reframes UAV-based collection as adaptive scientific data management rather than exhaustive sensing.

cs.GT

Individualized Algorithmic Advice as a Strategic Signal on Competitive Markets

As algorithms increasingly mediate competitive decision-making, their influence extends beyond individual outcomes to shaping strategic market dynamics. In our experiment, we examined how algorithmic advice affects human behavior in a classic economic game with a unique, non-collusive, and analytically traceable equilibrium. Participants (N = 129) played a Cournot quantity competition with equilibrium-aligned or strategically biased algorithmic recommendations. While individualized equilibrium advice supported stable convergence, collusively downward-biased advice led to sustained underproduction and supracompetitive profits - hallmarks of tacit collusion. Participants' quantities converged faster and more consistently toward individualized than collective equilibrium advice, potentially due to an objective quality advantage or greater perceived ownership of the former. These findings demonstrate that algorithmic advice can function as a strategic signal, shaping coordination even without explicit communication. The results echo real-world concerns about algorithmic collusion and underscore the need for careful design and oversight of algorithmic decision-support systems in competitive environments.

cs.HC

Construction of a DFA for Computing Grundy Numbers in the Successful Derivation Games on Right-Linear Grammars

Inoue et al. have introduced the successful derivation game (SDG) on context-free grammars (CFGs), which is a generalization of classic heap-based games including subtraction games and Keyles, and shown that the least upper bound of the Grundy numbers in the SDG on a given CFG G is undecidable in general even when we restrict G to be a linear CFG. This paper shows that for the SDG on a right-linear grammar (RLG), we can construct a DFA for computing the Grundy number of a given position. In other words, for the SDG on an RLG, the set of positions with a given Grundy number c is regular. As a corollary, the least upper bound of the Grundy numbers in the SDG on a given RLG is decidable. We also investigate the complexity of computing the least upper bound of the Grundy numbers in the SDG on a given RLG, and it is shown to be PSPACE-complete.

cs.FL

Make an Offer They Can't Refuse: Grounding Bayesian Persuasion in Real-World Dialogues without Pre-Commitment

Large language models (LLMs) still struggle with strategic persuasion, largely because existing approaches either neglect information asymmetry or rely on unrealistic pre-commitment assumptions. We introduce a type-induced commitment-communication mechanism that grounds Bayesian Persuasion (BP) in natural language dialogue without pre-commitment: the persuader narrates their potential types (e.g., honest vs. dishonest) to dynamically construct an information schema, enabling the persuadee to perform Bayesian belief updates within the conversation itself. We implement two variants: Semi-Formal-Natural-Language (SFNL) and Fully-Natural-Language (FNL), evaluating them against strong baselines across multiple LLMs and human judges. BP strategies consistently outperform baselines: SFNL excels in logical credibility, while FNL shows superior robustness and emotional resonance. We verify that gains stem from genuine Bayesian reasoning rather than superficial formatting, and we further show that supervised fine-tuning enables small models to match the persuasive performance of much larger ones.

cs.CL

Refundable Deposits: How to Restore Cooperation in Finitely Repeated Games

While infinitely repeated games admit a rich set of Nash equilibria, finitely repeated games typically have a much smaller and often inefficient one. We show how to enlarge this set using deposits: in each period a player may place a refundable sum with a neutral intermediary, returned when the game ends and forfeited following a deviation. Paying these deposits is voluntary and incentive compatible at every stage, so no commitment by the players is assumed, the only commitment required being that of the intermediary to a refund rule fixed before play begins. The mechanism sustains payoff profiles more efficient than those of the standard equilibria, without altering the underlying game and without transfers between players. We demonstrate it on the prisoner's dilemma, a congestion game, and a public goods game, all settings where cooperation cannot emerge in the standard finitely repeated version. We also apply it to a dynamic common-pool resource, suggesting that the construction extends beyond repeated stage-games.

cs.GT

Approval-Based Apportionment: Like Portioning, Approximately like Committee Voting

We study approval-based apportionment, a variant of committee elections in which candidates ("parties") can be selected several times. We show that the proportionality axioms EJR, EJR+, and FJR coincide, and so do PJR, PJR+, and FPJR; that Lindahl priceability, an axiom implying core stability, is equivalent to a notion of approximate optimality with respect to the proportional approval voting (PAV) score; and that locally PAV-optimal committees are priceable. Approval-based apportionment (where a candidate receives an integer number of seats) lies between committee elections (zero or one seat) and portioning (a fractional number of seats). We formally connect portioning and apportionment by giving a construction that lifts axioms from apportionment to portioning and preserves implications between them. Several of our new implications between apportionment axioms are natural from a portioning perspective, leading us to believe that apportionment sits closer to portioning than to committee elections. None of them holds in committee elections, but several extend approximately, which makes apportionment a fruitful setting for conjecturing approximate relationships in approval-based committee elections.

cs.GT

Distributed Places and Safe Net Reduction

Being able to find small Petri nets with the same behaviour as formal specifications of concurrent systems benefits both effective verification and practical implementation of such systems. This paper considers specifications given in the form of compositionally defined safe nets. The paper discusses a novel concept of distributed place which implements the behaviour of an individual net place. It is shown that if distributed places cover a safe Petri net, then it is possible to delete some places without changing the behaviour. Crucially, the reduction is both static and local, making it computationally feasible in practice. The resulting reduction technique is then applied to an algebra of safe Petri nets (boxes) derived compositionally from process (box) expressions. Though the original derivation can yield exponentially large boxes, prior research demonstrated that if a box expression does not involve cyclic behaviours, the exponential number of places can be reduced down to polynomial (quadratic). In this paper, using distributed places, it is show that similar optimisation can also be achieved in the case of process expressions with iteration.

cs.GT

Constant regret in general games via higher-order optimism

We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary $N$-player normal form game with up to $K$ actions per player, guarantees $O(N^3\log^2 K)$ individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting (HOOD) is a variant of optimistic follow-the-regularized-leader (OptFTRL) that combines a discounted $(N+1)$-th order predictor with entropic regularization over a suitable "lifting" of the game's strategy space. This combination of ingredients is purposefully designed to dampen large oscillations of the induced sequence of play in a controlled manner, removing in this way a key stumbling block of previous attempts to achieve constant regret in general games. Our approach bears several striking similarities to the concurrent - and completely independent - work of Liu, Farina, and Ozdaglar (arXiv:2608.31166), who very recently derived an $O(N^{21}\log^{4} K)$ regret bound through the use of higher-order optimism and an exponential moving average estimator.

cs.LG

Coalition Formation with Limited Information Sharing for Local Energy Management

Distributed energy systems with prosumers require new methods for coordinating energy exchange among agents. Coalitional control provides a framework in which agents form groups to cooperatively reduce costs; however, existing bottom-up coalition-formation methods typically require full information sharing, raising privacy concerns and imposing significant computational overhead. In this work, we propose a limited information coalition-formation algorithm that requires only limited aggregate information exchange among agents. By constructing an upper bound on the value of candidate coalitions, we eliminate the need to solve optimisation problems for each potential merge, significantly reducing computational complexity while limiting information exchange. We prove that the proposed method guarantees cost no greater than that of decentralised operation. Coalition strategies are optimised using a distributed approach based on the Alternating Direction Method of Multipliers (ADMM), further limiting information sharing within coalitions. We embed the framework within a model predictive control scheme and evaluate it on real-world data, demonstrating improved economic performance over decentralised control with substantially lower computational cost than full-information approaches.

cs.GT

Networked Multi-Resource Defense Capabilities in a General Lotto Game

Ensuring the security of complex systems involves the strategic allocation of defensive resources to prevent various types of attacks from succeeding. A defender often has multiple types of defensive assets at its disposal, where it must decide how to optimally deploy their heterogeneous capabilities across different attack types. In this paper, we formulate a multi-resource allocation problem in the form of a General Lotto game where a defender possesses various types of resources. A feature that we introduce is that their individual effectiveness against different types of attacks is characterized by a network weight matrix. In our analysis, we derive upper and lower bounds on the performance of the defender, and provide numerical evidence suggesting that they are tight. For the case of two attack types, we analytically prove that the bounds coincide, establishing an exact equilibrium characterization. We then numerically compare our proposed networked multi-resource architecture to an independent-defense benchmark from the existing literature. These results highlight fundamental and tractable structures underlying multi-attack-type defense problems.

cs.GT