arXiv ScienceSearch

arXiv subjects

Austin Cooper

Publications and source records attributed to Austin Cooper.

6 recordsLinked to original sources

Reinforcement Learning for Optimal Stopping in POMDPs with Application to Quickest Change Detection

The field of quickest change detection (QCD) focuses on the design and analysis of online algorithms that estimate the time at which a significant event occurs. In this paper, design and analysis are cast in a Bayesian framework, where QCD is formulated as an optimal stopping problem with partial observations. An approximately optimal detection algorithm is sought using techniques from reinforcement learning. The contributions of the paper are summarized as follows: (i) A Q-learning algorithm is proposed for the general partially observed optimal stopping problem. It is shown to converge under linear function approximation, given suitable assumptions on the basis functions. An example is provided to demonstrate that these assumptions are necessary to ensure algorithmic stability. (ii) Prior theory motivates a particular choice of features in applying Q-learning to QCD. It is shown that, in several scenarios and under ideal conditions, the resulting class of policies contains one that is approximately optimal. (iii) Numerical experiments show that Q-learning consistently produces policies that perform close to the best achievable within the chosen function class.

math.OC

Quickest Change Detection Using Mismatched CUSUM

Quickest change detection concerns estimation of an unknown change time \(\tau_a\) from a sequence of partial observations \(\{Y_k:k\ge 0\}\). We consider stopping rules of CUSUM form, \[ \mathcal{X}_{n+1} = \max\{0,\mathcal{X}_n+F(Y_{n+1})\}, \quad \tau_s=\min\{n\ge 0:\mathcal{X}_n\ge \textrm{H}\}, \] where the function \(F\) and threshold \(\textrm{H}\) are design parameters. The observations and change time are modeled jointly through a hidden Markov model, and \( F\) is selected from a prescribed function class \(\mathcal{G}\) to minimize the weighted criterion \[ \textsf{E}\bigl[ (\tau_s-\tau_a)_+ + \kappa(\tau_s-\tau_a)_- \bigr]. \] When \(\mathcal{G}\) is a linear function class, the optimizer \(F^*\) is characterized by a convex program, whose dual yields extensions of classical likelihood-ratio constructions. This conclusion is based on analysis that is asymptotic in the regime \(\kappa\to\infty\). We show that the hidden Markov model admits an asymptotically equivalent conditionally independent approximation of the type commonly used in the quickest change detection literature. We then develop the design and asymptotic theory for a substantially broader class of conditionally independent models, so that the resulting conclusions are not tied to the particular POMDP reduction. Combining renewal theory and large deviations for reflected random walks, we obtain for each $F\in\mathcal{G}$ asymptotically accurate approximations of the optimal threshold and average cost, with error vanishing as \(\kappa\to\infty\). It is found in numerical experiments that the resulting approximations are accurate for moderate values of \(\kappa\).

math.ST

Reinforcement Learning Design for Quickest Change Detection

The field of quickest change detection (QCD) concerns design and analysis of algorithms to estimate in real time the time at which an important event takes place, and identify properties of the post-change behavior. It is shown in this paper that approaches based on reinforcement learning (RL) can be adapted based on any "surrogate information state" that is adapted to the observations. Hence we are left to choose both the surrogate information state process and the algorithm. For the former, it is argued that there are many choices available, based on a rich theory of asymptotic statistics for QCD. Two approaches to RL design are considered: (i) Stochastic gradient descent based on an actor-critic formulation. Theory is largely complete for this approach: the algorithm is unbiased, and will converge to a local minimum. However, it is shown that variance of stochastic gradients can be very large, necessitating the need for commensurately long run times; (ii) Q-learning algorithms based on a version of the projected Bellman equation. It is shown that the algorithm is stable, in the sense of bounded sample paths, and that a solution to the projected Bellman equation exists under mild conditions. Numerical experiments illustrate these findings, and provide a roadmap for algorithm design in more general settings.

math.OC

Uncertainty Error Modeling for Non-Linear State Estimation With Unsynchronized SCADA and $\mu$PMU Measurements

Distribution systems of the future smart grid require enhancements to the reliability of distribution system state estimation (DSSE) in the face of low measurement redundancy, unsynchronized measurements, and dynamic load profiles. Micro phasor measurement units ($\mu$PMUs) facilitate co-synchronized measurements with high granularity, albeit at an often prohibitively expensive installation cost. Supervisory control and data acquisition (SCADA) measurements can supplement $\mu$PMU data, although they are received at a slower sampling rate. Further complicating matters is the uncertainty associated with load dynamics and unsynchronized measurements-not only are the SCADA and $\mu$PMU measurements not synchronized with each other, but the SCADA measurements themselves are received at different time intervals with respect to one another. This paper proposes a non-linear state estimation framework which models dynamic load uncertainty error by updating the variances of the unsynchronized measurements, leading to a time-varying system of weights in the weighted least squares state estimator. Case studies are performed on the 33-Bus Distribution System in MATPOWER, using Ornstein-Uhlenbeck stochastic processes to simulate dynamic load conditions.

eess.SY

High Impedance Fault Detection Through Quasi-Static State Estimation: A Parameter Error Modeling Approach

This paper presents a model for detecting high-impedance faults (HIFs) using parameter error modeling and a two-step per-phase weighted least squares state estimation (SE) process. The proposed scheme leverages the use of phasor measurement units and synthetic measurements to identify per-phase power flow and injection measurements which indicate a parameter error through $\chi^2$ Hypothesis Testing applied to the composed measurement error (CME). Although current and voltage waveforms are commonly analyzed for high-impedance fault detection, wide-area power flow and injection measurements, which are already inherent to the SE process, also show promise for real-world high-impedance fault detection applications. The error distributions after detection share the measurement function error spread observed in proven parameter error diagnostics and can be applied to HIF identification. Further, this error spread related to the HIF will be clearly discerned from measurement error. Case studies are performed on the 33-Bus Distribution System in Simulink.

eess.SY

Distributed Software-Defined Network Architecture for Smart Grid Resilience to Denial-of-Service Attacks

An important challenge for smart grid security is designing a secure and robust smart grid communications architecture to protect against cyber-threats, such as Denial-of-Service (DoS) attacks, that can adversely impact the operation of the power grid. Researchers have proposed using Software Defined Network frameworks to enhance cybersecurity of the smart grid, but there is a lack of benchmarking and comparative analyses among the many techniques. In this work, a distributed three-controller software-defined networking (D3-SDN) architecture, benchmarking and comparative analysis with other techniques is presented. The selected distributed flat SDN architecture divides the network horizontally into multiple areas or clusters, where each cluster is handled by a single Open Network Operating System (ONOS) controller. A case study using the IEEE 118-bus system is provided to compare the performance of the presented ONOS-managed D3-SDN, against the POX controller. In addition, the proposed architecture outperforms a single SDN controller framework by a tenfold increase in throughput; a reduction in latency of $>20\%$; and an increase in throughput of approximately $11\%$ during the DoS attack scenarios.

eess.SY