arXiv Science⌕ Search

arXiv · 2610.12028

Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints

Abstract

Consider a finite population of agents with decoupled Markov transition dynamics and empirical-density feedback, subject to the following constraints: with probability at least $1-δ_r$, at least a fraction $α_r$ of agents must reach a target region at some time $t^*$, while, at each time up to $t^*$, the unsafe population fraction must remain below $β_u$ with probability at least $1-δ_u$. However, standard mean-field methods enforce these constraints only in expectation, which fails to account for stochastic fluctuations at finite fleet size $N$. To address this control problem, we propagate the second-order moment (variance) of the empirical density alongside the mean-field trajectory via a discrete-time Lyapunov recursion, and apply the Cantelli inequality to convert chance constraints into tractable deterministic conditions on the moments of the empirical density. We then incorporate these moment-based surrogate constraints into a gradient-based sequential convex approximation procedure for density-feedback policy synthesis. We further introduce additional moment-error bounds to construct a rigorous finite-$N$ certificate. The method is evaluated on a gridworld environment and a power-system EV-charging aggregation problem and compared with a standard deterministic population-level LP baseline.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jie Fu, Anamika Dubey. 2026-10-08. Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints. https://arxiv.org/abs/2610.12028

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Compositional and Equilibrium-Free Stability Certification for Power Systems--Part II: Algorithms and Applications

This two-part study develops a compositional and equilibrium-free framework for power system stability analysis. Building on the theoretical foundation established in Part I, i.e., local delta dissipativity (LDD), this second part translates the theory into a practical assessment tool. We address two core implementation challenges: 1) how to certify the LDD conditions for heterogeneous power device models, and 2) how to verify the network-wide coupling condition in a scalable manner. To this end, we present a method that employs Krasovskii-type storage functions to verify LDD conditions for heterogeneous power devices and applies to a wide range of nonlinear power device models. For coupling conditions, we propose an ADMM-based distributed algorithm, enhanced with a $p$-check subroutine, for interconnection verification. Our method enables three primary application scenarios that have been hindered by equilibrium-point-oriented and centralized methods: stability assessment for systems with multiple equilibria, rapid stability screening under varying operating conditions, and a privacy-preserving distributed assessment architecture for large-scale systems. Case studies on modified IEEE 9-bus, 39-bus, and 118-bus test systems demonstrate the effectiveness of the proposed methods across different network configurations and scales.

eess.SY↗

Towards Ultra-Reliable 6G in-X Subnetworks: Dynamic Link Adaptation by Deep Reinforcement Learning

6G networks are composed of subnetworks expected to meet ultra-reliable low-latency communication (URLLC) requirements for mission-critical applications such as industrial control and automation. An often-ignored aspect in URLLC is consecutive packet outages, which can destabilize control loops and compromise safety in in-factory environments. Hence, the current work proposes a link adaptation framework to support extreme reliability requirements using the soft actor-critic (SAC)-based deep reinforcement learning (DRL) algorithm that jointly optimizes energy efficiency (EE) and reliability under dynamic channel and interference conditions. Unlike prior work focusing on average reliability, our method explicitly targets reducing burst/consecutive outages through adaptive control of transmit power and blocklength based solely on the observed signal-to-interference-plus-noise ratio (SINR). The joint optimization problem is formulated under finite blocklength and quality of service constraints, balancing reliability and EE. Simulation results show that the proposed method significantly outperforms the baseline algorithms, reducing outage bursts while consuming only 18\% of the transmission cost required by a full/maximum resource allocation policy in the evaluated scenario. The framework also supports flexible trade-off tuning between EE and reliability by adjusting reward weights, making it adaptable to diverse industrial requirements.

eess.SY↗

An Information Theory of Finite Abstractions and their Fundamental Scalability Limits

Finite abstractions are discrete models of dynamical systems, such that the set of abstraction trajectories contains all system trajectories. There is a consensus that abstractions suffer from a scalability bottleneck: to obtain a sufficient ``accuracy" (how closely the abstraction models the system), the required abstraction size often becomes very large, rendering computation infeasible. This is accentuated for complex, high-dimensional systems, due to the curse of dimensionality. Yet, after decades of research, there are no formal results on the size-accuracy tradeoff of abstractions. Here, we derive a statistical, quantitative theory of the size-accuracy tradeoff of abstractions of autonomous deterministic systems and uncover fundamental limits on their scalability, through rate-distortion theory---the information theory of lossy compression. Abstractions are viewed as encoder-decoder pairs, encoding trajectories of dynamical systems. Rate measures abstraction size, while distortion describes accuracy, defined as the spatial average deviation between abstract trajectories and system ones. We obtain a fundamental lower bound on the minimum achievable abstraction distortion, given the system dynamics and the abstraction size; and vice-versa a lower bound on the minimum required size, for given distortion. The bound depends on the complexity of the dynamics, through trajectory entropy and a constant depending on the trajectory manifold's geometry. We demonstrate its tightness on some dynamical systems. Finally, we showcase how this new theory enables constructing minimal abstractions, optimizing the size-accuracy tradeoff, through an example on a chaotic system.

eess.SY↗