arXiv ScienceSearch

arXiv subjects

Kejun Chen

Publications and source records attributed to Kejun Chen.

16 recordsLinked to original sources

Personalized Federated Learning for Tensor Regression

The growing availability of tensor-valued data across multiple institutions creates opportunities for collaborative analysis, but also raises challenges related to data privacy, high dimensionality, and client heterogeneity. This paper introduces a personalized federated tensor regression framework that addresses all three simultaneously. Each client's coefficient tensor is decomposed into a globally shared low-Tucker-rank component and a locally sparse deviation, estimated via a two-stage privacy-preserving procedure. We establish finite-sample upper bounds and minimax lower bounds that quantify the privacy-accuracy trade-off, and prove the consistency of the supporting initialization and rank-selection steps. Simulation studies confirm that the federated approach improves estimation and prediction over purely local methods, especially when per-client data are scarce, and an MRI-based ADHD study illustrates its strong performance under real privacy constraints.

stat.ME

Private Federated Learning for High-dimensional Time Series

In the era of big data, leveraging information from multiple clients while preserving data privacy has emerged as a critical challenge in modern statistical modeling and forecasting. This paper introduces a privacy-preserving federated learning framework for high-dimensional vector autoregressive models, where each client's dynamics are characterized by a common low-rank structure augmented with sparse client-specific deviations. We develop a two-stage estimation procedure that integrates differentially private representation learning for the shared component with local personalization for client-specific adjustments, enabling effective information pooling under selective privacy constraints. Non-asymptotic error bounds are established for both the single-client and federated estimators to characterize the inherent privacy-utility trade-off, and consistency of a ridge-type rank selection criterion is proved. Simulation studies demonstrate that federation substantially improves estimation accuracy when local sample sizes are limited. Two empirical applications to analyzing electricity-economy linkages across U.S. states and conducting multi-task macroeconomic forecasting across countries, highlight the superior predictive accuracy of the proposed method over existing single-client benchmarks.

stat.ME

A hard-constrained NN learning framework for rapidly restoring AC-OPF from DC-OPF

This paper proposes a hard-constrained unsupervised learning framework for rapidly solving the non-linear and non-convex AC optimal power flow (AC-OPF) problem in real-time operation. Without requiring ground-truth AC-OPF solutions, feasibility and optimality are ensured through a properly designed learning environment and training loss. Inspired by residual learning, the neural network (NN) learns the correction mapping from the DC-OPF solution to the active power setpoints of the generators through re-dispatch. A subsequent optimization model is utilized to restore the optimal AC-OPF solution, and the resulting projection difference is employed as the training loss. A replay buffer is utilized to enhance learning efficiency by fully leveraging past data pairs. The optimization model is cast as a differentiable optimization layer, where the gradient is derived by applying the implicit function theorem to the KKT conditions at the optimal solution. Tested on IEEE-118 and PEGASE-9241 bus systems, numerical results demonstrate that the proposed NN can obtain strictly feasible and near-optimal solutions with reduced computational time compared to conventional optimization solvers. In addition, aided by the updated DC-OPF solution under varying topologies, the trained NN, together with the PF solver, can rapidly find the corresponding AC solution. The proposed method achieves a $40\times$ time speedup, while maintaining an average constraint violation on the order of $10^{-4}$ and an optimization gap below $1\%$.

eess.SY

A robust and scalable estimation for high-dimensional volatility models

This paper introduces a robust and computationally efficient estimation framework for high-dimensional volatility models in the BEKK-ARCH class. The proposed approach employs data truncation to ensure robustness against heavy-tailed distributions and utilizes a regularized least squares method for efficient optimization in high-dimensional settings. Non-asymptotic error bounds are established for the resulting estimators under heavy-tailed regimes, and the minimax optimal convergence rate is derived. Moreover, a robust BIC and a Ridge-type estimator are introduced for selecting the model order and the number of BEKK components, respectively, with their selection consistency established under heavy-tailed settings. Simulation studies demonstrate finite-sample performance of the proposed method, and two empirical applications illustrate its practical utility. The results show that the new framework outperforms existing alternatives in both computational speed and forecasting accuracy.

math.ST

Physics-Informed Gradient Estimation for Accelerating Deep Learning based AC-OPF

The optimal power flow (OPF) problem can be rapidly and reliably solved by employing responsive online solvers based on neural networks. The dynamic nature of renewable energy generation and the variability of power grid conditions necessitate frequent neural network updates with new data instances. To address this need and reduce the time required for data preparation time, we propose a semi-supervised learning framework aided by data augmentation. In this context, ridge regression replaces the traditional solver, facilitating swift prediction of optimal solutions for the given input load demands. Additionally, to accelerate the backpropagation during training, we develop novel batch-mean gradient estimation approaches along with a reduced branch set to alleviate the complexity of gradient computation. Numerical simulations demonstrate that our neural network, equipped with the proposed gradient estimators, consistently achieves feasible and near-optimal solutions. These results underline the effectiveness of our approach for practical implementation in real-time OPF applications.

eess.SY

Adversarial Multi-Agent Reinforcement Learning for Proactive False Data Injection Detection

Smart inverters are instrumental in the integration of distributed energy resources into the electric grid. Such inverters rely on communication layers for continuous control and monitoring, potentially exposing them to cyber-physical attacks such as false data injection attacks (FDIAs). We propose to construct a defense strategy against a priori unknown FDIAs with a multi-agent reinforcement learning (MARL) framework. The first agent is an adversary that simulates and discovers various FDIA strategies, while the second agent is a defender in charge of detecting and locating FDIAs. This approach enables the defender to be trained against new FDIAs continuously generated by the adversary. In addition, we show that the detection skills of an MARL defender can be combined with those of a supervised offline defender through a transfer learning approach. Numerical experiments conducted on a distribution and transmission system demonstrate that: a) the proposed MARL defender outperforms the offline defender against adversarial attacks; b) the transfer learning approach makes the MARL defender capable against both synthetic and unseen FDIAs.

eess.SY

Continual Adversarial Reinforcement Learning (CARL) of False Data Injection detection: forgetting and explainability

False data injection attacks (FDIAs) on smart inverters are a growing concern linked to increased renewable energy production. While data-based FDIA detection methods are also actively developed, we show that they remain vulnerable to impactful and stealthy adversarial examples that can be crafted using Reinforcement Learning (RL). We propose to include such adversarial examples in data-based detection training procedure via a continual adversarial RL (CARL) approach. This way, one can pinpoint the deficiencies of data-based detection, thereby offering explainability during their incremental improvement. We show that a continual learning implementation is subject to catastrophic forgetting, and additionally show that forgetting can be addressed by employing a joint training strategy on all generated FDIA scenarios.

cs.LG

Detection of False Data Injection Attacks (FDIA) on Power Dynamical Systems With a State Prediction Method

With the deeper penetration of inverter-based resources in power systems, false data injection attacks (FDIA) are a growing cyber-security concern. They have the potential to disrupt the system's stability like frequency stability, thereby leading to catastrophic failures. Therefore, an FDIA detection method would be valuable to protect power systems. FDIAs typically induce a discrepancy between the desired and the effective behavior of the power system dynamics. A suitable detection method can leverage power dynamics predictions to identify whether such a discrepancy was induced by an FDIA. This work investigates the efficacy of temporal and spatio-temporal state prediction models, such as Long Short-Term Memory (LSTM) and a combination of Graph Neural Networks (GNN) with LSTM, for predicting frequency dynamics in the absence of an FDIA but with noisy measurements, and thereby identify FDIA events. For demonstration purposes, the IEEE 39 New England Kron-reduced model simulated with a swing equation is considered. It is shown that the proposed state prediction models can be used as a building block for developing an effective FDIA detection method that can maintain high detection accuracy across various attack and deployment settings. It is also shown how the FDIA detection should be deployed to limit its exposure to detection inaccuracies and mitigate its computational burden.

cs.CR

Presolving Convexified Optimal Power Flow with Mixtures of Gradient Experts

Convex relaxations and approximations of the optimal power flow (OPF) problem have gained significant research and industrial interest for planning and operations in electric power networks. One approach for reducing their solve times is presolving which eliminates constraints from the problem definition, thereby reducing the burden of the underlying optimization algorithm. To this end, we propose a presolving framework for convexified optimal power flow (C-OPF) problems, which uses a novel deep learning-based architecture called MoGE (Mixture of Gradient Experts). In this framework, problem size is reduced by learning the mapping between C-OPF parameters and optimal dual variables (the latter being representable as gradients), which is then used to screen constraints that are non-binding at optimum. The validity of using this presolve framework across arbitrary families of C-OPF problems is theoretically demonstrated. We characterize generalization in MoGE and develop a post-solve recovery procedure to mitigate possible constraint classification errors. Using two different C-OPF models, we show via simulations that our framework reduces solve times by upto 34% across multiple PGLIB and MATPOWER test cases, while providing an identical solution as the full problem.

math.OC

On LinDistFlow Model Congestion Pricing: Bounding the Changes in Power Tariffs

The optimal power flow (OPF) problem is an important mathematical program that aims at obtaining the best operating point of an electric power grid. The optimization problem typically minimizes the total generation cost subject to certain physical constraints of the system. The so-called linearized distribution flow (LinDistFlow) model leverages a set of linear equations to approximate the nonlinear AC power flows. In this paper, we consider an OPF problem based on the LinDistFlow model for a single-phase radial power network. We derive closed-form solutions to the marginal values of both real and reactive power demands. We also derive upper bounds on the congestion price (a.k.a. `shadow price'), which denotes the change in marginal demand prices when the apparent power flow limits of certain lines are binding at optimum. Various cases of our result are discussed while simulations are carried out on a $141$-bus radial power network.

math.OC

Physics-guided Residual Learning for Probabilistic Power Flow Analysis

Probabilistic power flow (PPF) analysis is critical to power system operation and planning. PPF aims at obtaining probabilistic descriptions of the state of the system with stochastic power injections (e.g., renewable power generation and load demands). Given power injection samples, numerical methods repeatedly run classic power flow (PF) solvers to find the voltage phasors. However, the computational burden is heavy due to many PF simulations. Recently, many data-driven based PF solvers have been proposed due to the availability of sufficient measurements. This paper proposes a novel neural network (NN) framework which can accurately approximate the non-linear AC-PF equations. The trained NN works as a rapid PF solver, significantly reducing the heavy computational burden in classic PPF analysis. Inspired by residual learning, we develop a fully connected linear layer between the input and output in the multilayer perceptron (MLP). To improve the NN training convergence, we propose three schemes to initialize the NN weights of the shortcut connection layer based on the physical characteristics of AC-PF equations. Specifically, two model-based methods require the knowledge of system topology and line parameters, while the purely data-driven method can work without power grid parameters. Numerical tests on five benchmark systems show that our proposed approaches achieve higher accuracy in estimating voltage phasors than existing methods. In addition, three meticulously designed initialization schemes help the NN training process converge faster, which is appealing under limited training time.

eess.SY

Unsupervised Deep Learning for AC Optimal Power Flow via Lagrangian Duality

Non-convex AC optimal power flow (AC-OPF) is a fundamental optimization problem in power system analysis. The computational complexity of conventional solvers is typically high and not suitable for large-scale networks in real-time operation. Hence, deep learning based approaches have gained intensive attention to conduct the time-consuming training process offline. Supervised learning methods may yield a feasible AC-OPF solution with a small optimality gap. However, they often need conventional solvers to generate the training dataset. This paper proposes an end-to-end unsupervised learning based framework for AC-OPF. We develop a deep neural network to output a partial set of decision variables while the remaining variables are recovered by solving AC power flow equations. The fast decoupled power flow solver is adopted to further reduce the computational time. In addition, we propose using a modified augmented Lagrangian function as the training loss. The multipliers are adjusted dynamically based on the degree of constraint violation. Extensive numerical test results corroborate the advantages of our proposed approach over some existing methods.

eess.SY

Variation-cognizant Probabilistic Power Flow Analysis via Multi-task Learning

With an increasing high penetration of solar photovoltaic generation in electric power grids, voltage phasors and branch power flows experience more severe fluctuations. In this context, probabilistic power flow (PPF) study aims at characterizing the statistical properties of the state of the system with respect to the random power injections. To avoid repeated power flow calculations involved in PPF study, the present paper leverages regression algorithms and neural networks to improve the estimation performance and speed up the computation. Specifically, based on the variation level of the voltage magnitude at each bus, we develop either a linear regression or a fully connected neural network to approximate the inverse AC power flow mappings. The proposed multi-task learning technique further improves the accuracy of branch flow estimation by incorporating the errors of voltage angle differences into the loss function design. Tested on IEEE-300 and IEEE-1354 bus systems with real data, the proposed methods achieve better performance in estimating voltage phasors and branch flows.

eess.SY

Paired Ru-O-Mo ensemble for efficient and stable alkaline hydrogen evolution reaction

Electrocatalytic hydrogen evolution reaction (HER) in alkaline media is a promising electrochemical energy conversion strategy. Ruthenium (Ru) is an efficient catalyst with a desirable cost for HER, however, the sluggish H2O dissociation process, due to the low H2O adsorption on its surface, currently hampers the performances of this catalyst in alkaline HER. Herein, we demonstrate that the H2O adsorption improves significantly by the construction of Ru-O-Mo sites. We prepared Ru/MoO2 catalysts with Ru-O-Mo sites through a facile thermal treatment process and assessed the creation of Ru-O-Mo interfaces by transmission electron microscope (TEM) and extended X-ray absorption fine structure (EXAFS). By using Fourier-transform infrared spectroscopy (FTIR) and H2O adsorption tests, we proved Ru-O-Mo sites have tenfold stronger H2O adsorption ability than that of Ru catalyst. The catalysts with Ru-O-Mo sites exhibited a state-of-the-art overpotential of 16 mV at 10 mA cm-2 in 1 M KOH electrolyte, demonstrating a threefold reduction than the previous bests of Ru (59 mV) and commercial Pt (31 mV) catalysts. We proved the stability of these performances over 40 hours without decline. These results could open a new path for designing efficient and stable catalysts.

physics.chem-ph

A family of multimagic squares based on large sets of orthogonal arrays

Large set of orthogonal arrays (LOA) were introduced by D. R. Stinson, and it is also used to construct multimagic squares recently. In this paper, multimagic squares based on strong double LOA are further investigated. It is proved that there exists an MS$(q^{2t-1},t)$ for any prime power $q\geq 2t-1$ with $t\geq3$, which provided a new family of multimagic squares.

math.CO

A generalization of product construction of multimagic squares

In this paper, constructions of multimagic squares are investigated. Diagonal Latin squares and Kronecker products are used to get some constructions of multimagic squares. Consequently, some new families of compound multimagic squares are obtained.

math.CO