arXiv ScienceSearch

arXiv subjects

Xun Li

Publications and source records attributed to Xun Li.

At least 19 recordsLinked to original sources

Closed-loop $\alpha$-Potential Stochastic Differential Games via a BSDE Approach

In this paper, we study the closed-loop $\alpha $-potential stochastic differential game (SDG) problem as a continuation of our prior research on open-loop control (see \cite{GLZ2025}). By utilizing the backward stochastic differential equation (BSDE) approach, we derive a precise estimate for the parameter $\alpha $. Compared to our earlier work \cite{GLZ2025}, this study incorporates both first- and second-order sensitivity state processes, as well as the sensitivity of the control process. A distinguishing feature of this work is that, in the context of $N$-player heterogeneous agent games involving mean-field type interactions, we derive an $N$-uniform upper bound for the minimal potential approximation error. In contrast to the corresponding open-loop estimates, the closed-loop bound contains feedback-induced contributions that need not vanish with $N$. Consequently, our present estimate does not in general guarantee $\alpha \rightarrow 0$.

math.OC

Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution

Effective research ideation requires moving beyond a static understanding of prior work to trace how research problems and solutions evolve across the literature. Existing methods either treat papers as unstructured context or model scholarly evolution as isolated citation chains, overlooking interactions among research trajectories. We propose Tree-of-Ideas (ToI), a two-stage framework. EvoTrace reconstructs branching scholarly trajectories from citations, tracking evolving methods, resolved problems, and gaps. EvoAgent then reasons across trajectories to identify convergent problems and complementary solutions, generating grounded research ideas. Across six AI research topics, ToI achieves the highest score among automatic methods (6.27 vs. 5.36 for the strongest baseline on a 10-point scale), with strong Novelty (6.36) and Groundedness (7.00). Also, its score approaches that of human-paper references (6.29), demonstrating the value of cross-path evolutionary reasoning.

cs.AI

Inverse reinforcement learning for indefinite mean-field social optimization with multiplicative noise

This paper studies the inverse reinforcement learning (RL) problem for linear-quadratic mean-field (MF) social optimization. The considered system features multiplicative noise and indefinite cost weights, which violate standard convexity assumptions and pose analytical challenges. The goal is to recover unknown social cost weights from expert demonstrations and reproduce the optimal control policies. This requires solving coupled stochastic algebraic Riccati equations and Lyapunov equations with unknown system dynamics. To this end, we first propose a model-based inverse RL algorithm with two sequential loops that separately handle individual and MF dynamics, and we prove its convergence and closed-loop stabilizability. Moreover, we characterize the non-uniqueness of the recovered cost weights. To eliminate reliance on system dynamics, we develop a model-free inverse RL algorithm using integral RL and least-squares identification, which requires only measured trajectory data satisfying mild rank conditions. Finally, numerical simulations validate the effectiveness of the proposed approaches.

math.OC

Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching

This paper addresses infinite-horizon continuous-time stochastic linear quadratic optimal control problems with regime switching. We propose a paradigm shift from model-based design by adopting an adaptive dynamic programming approach, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data. The theoretical core of our work consists of a complete proof of the equivalence between the on- and off-policy architectures, alongside a rigorous analysis establishing the stability of the closed-loop system and the convergence of the algorithms to the optimal solution. For computational tractability, we implement these algorithms using vectorization and Kronecker product algebra. The theoretical results are corroborated by numerical case studies that clearly demonstrate the operational effectiveness and practical feasibility of the proposed model-free control strategy.

math.OC

Dynamic mean-variance portfolio selection with no-shorting constraints and unknown investment opportunity sets

We study continuous-time mean-variance portfolio selection with no-shorting constraints and unknown investment opportunity sets from a reinforcement learning (RL) perspective. The problem is a constrained stochastic linear -- quadratic control problem for which the entropy-regularized exploratory formulation of Wang et al. (2020) leads to difficulty in theoretical analysis, because enforcing the constraint on the support of randomized policies nullifies the tractable Gaussian exploration. To tackle this challenge, we introduce an auxiliary exploratory problem without entropy in which exploratory policies are still Gaussian whose samples may violate the no-shorting requirement but their means satisfy it. We then prove that, for a suitable choice of exploration variance, the mean of the optimal Gaussian policy of the auxiliary problem coincides with the optimal policy of the original problem. Motivated by this theoretical result, we develop a model-free RL algorithm that learns the optimal policy of the auxiliary (and hence the original) problem directly from trajectory data without estimating the investment opportunity set. A numerical example demonstrates the performance of the proposed algorithm.

math.OC

Inverse Optimal Control for Linear Quadratic Problem with Poisson Jumps: Model-Free Inverse Reinforcement Learning Approaches

This paper addresses the inverse optimal control (IOC) problem for stochastic linear systems subject to both Brownian motion and Poisson jumps, using an inverse reinforcement learning (IRL) framework. Given a target feedback gain from an expert, the objective is to identify an equivalent cost functional-specifically, the set of all cost weights-that yields this same gain. To solve this problem when system dynamics are unknown, we propose two model-free, off-policy IRL algorithms that operate entirely from data, circumventing the need to solve the generalized algebraic Riccati equation or compute the cost weights analytically. The first is an inverse Q-learning algorithm that constructs data-driven equations from expert demonstrations to compute the Q-function matrix, with equivalent cost weights updated algebraically and without requiring additional trajectory data. The second is a model-free off-policy inverse policy iteration algorithm that leverages data collected under an initial stabilizing policy, offering a complementary approach suited to different data availability scenarios. Crucially, by decoupling the data-collection behavior policies from the policies being iteratively updated, both algorithms can learn equivalent cost weights from sufficiently excited trajectories without identifying the system dynamics or jump intensity. Numerical simulations validate the effectiveness of the proposed methods.

math.OC

Acoustic Cloning

Cloning refers to producing identical copies of existing objects. Here, we experimentally show how to clone acoustic scattering objects. We acquire a digital twin and bring it back to life - a simple two-step process. First, we use broadband speakers to illuminate the scattering object within a closed receiver aperture. From these recorded reverberative data, we retrieve the object's scattering Green's functions using multidimensional deconvolution. In the second step, the acoustic scatterer is holographically reconstructed using the acquired scattering Green's functions. The hologram scatters any wavefield in real-time exactly like the original object would. Low-latency feedback reproduces all orders of interactions between the physical wavefield and the numerically defined hologram. This two-step process is demonstrated by cloning and modifying several rigid scatterers in a two-dimensional acoustic waveguide. Applications range from fully realistic digital scattering models to efficient metamaterial experimentation.

physics.app-ph

Explicit Convergence Regions of PID-Damped Accelerated Gradient Methods in Nonconvex Optimization

Momentum-based accelerated gradient methods are widely adopted to expedite convergence in nonconvex optimization, but are prone to overshooting and oscillatory behavior. A class of PID-damped accelerated gradient methods mitigates this issue by augmenting classical momentum methods with a discrete-time derivative damping term. However, the coupling among the step size, momentum, and derivative gain renders the explicit characterization of their convergence regions analytically intractable, leaving explicit theoretical convergence boundaries unexplored. In this paper, we model this class of algorithms as a third-order nonlinear feedback dynamical system and establish explicit three-dimensional convergence regions for the step size, momentum, and derivative gain via a robust control-theoretic analysis based on the Kalman-Yakubovich-Popov (KYP) lemma, formally guaranteeing linear convergence under the Regularity Condition. Furthermore, we reveal a strict geometric upper bound on the derivative gain dictated by the nonconvex curvature, beyond which over-damping severely contracts the feasible step-size region, providing a rigorous theoretical explanation for the overdamped stagnation phenomenon. Numerical experiments corroborate the theoretical boundaries and illustrate practical parameter selection guidelines for the derivative gain.

math.OC

Relaxed Control with Entropy Regularization for It\^o Stochastic Systems with Input Delay

This paper investigates the infinite-horizon classical stochastic optimal control problem with input delay under an entropy-regularized relaxed control framework. In particular, by constructing a relaxed system and introducing an entropy regularization term, we reformulate the classical optimal control problem into an entropy regularized formulation, and derive the optimal controller that follows a Gaussian distribution. Furthermore, we show that the optimal Gaussian control distribution converges to the optimal Dirac measure as the exploration weight tends to zero. Numerical simulation is provided to validate the effectiveness of the proposed method.

math.OC

Quantum master equation approach for the multiphonon up-pumping model

A fully quantum multiphonon up-pumping model is proposed to characterize coherent energy transfer in energetic materials (EMs) subjected to external shock. After eliminating the degrees of freedom of the phonon bath within a mean-field approximation, we derive a quantum master equation governing the energy transfer among vibrational modes. Our analysis reveals that doorway modes of different frequencies undergo distinct levels of effective coherent driving and dissipation, induced by the shocked phonon environment. This not only clarifies the microscopic origin of coherent phonon generation, but also reveals the possibility of modulating such coherent driving and dissipation. Based on numerical simulations of a simplified model using the master equation, we demonstrate how doorway modes extract energy from the phonon environment and subsequently excite higher-frequency molecular vibrational modes. This work offers a renewed perspective for understanding the mechanisms of energy transfer in energetic materials.

quant-ph

Distributed Load Frequency Control of Multi-Area Smart Grid

In this paper, we investigate the distributed load frequency control problem in a multi-area smart grid under external load disturbances and measurement noise. The novelty lies in that the information privacy is fully taken into account, that is, the internal structural parameters and operational states of each area are not shared with non-neighboring areas, which makes traditional distributed optimal control methods ineffective. The main contribution is to propose a distributed algorithm for the global optimal power regulation command under information privacy constraints via distributed approximation of the control Riccati equation, the estimation Riccati equation, and the state estimation. Simulation results show that the proposed algorithm can approximate the performance of centralized optimal control, and the performance index under the proposed distributed controller is smaller than that under the commonly used distributed control.

math.OC

Distributed Algorithm for the Global Optimal Controller of Nonlinear Multi-Agent Systems

In this paper, we investigate the distributed optimal control problem for a kind of nonlinear multi-agent systems. In particular,both the state and the system dynamic structures of each agent are private and can only be shared among communicating agents.This type of information structure is inevitable in fields such as collaborative control for industrial confidentiality, and renders traditional distributed control methods using all systems' dynamic structures ineffective. The primary contribution is the proposal of a distributed algorithm for the global optimal controller under such practical information structure via distributed approximation of the Hamilton-Jacobi-Bellman equation. Practical numerical simulation demonstrates the effectiveness of the proposed algorithm.

math.OC

Knowledge Priors for Identity-Disentangled Open-Set Privacy-Preserving Video FER

Facial expression recognition relies on facial data that inherently expose identity and thus raise significant privacy concerns. Current privacy-preserving methods typically fail in realistic open-set video settings where identities are unknown, and identity labels are unavailable. We propose a two-stage framework for video-based privacy-preserving FER in challenging open-set settings that requires no identity labels at any stage. To decouple privacy and utility, we first train an identity-suppression network using intra- and inter-video knowledge priors derived from real-world videos without identity labels. This network anonymizes identity while preserving expressive cues. A subsequent denoising module restores expression-related information and helps recover FER performance. Furthermore, we introduce a falsification-based validation method that uses recognition priors to rigorously evaluate privacy robustness without requiring annotated identity labels. Experiments on three video datasets demonstrate that our method effectively protects privacy while maintaining FER accuracy comparable to identity-supervised baselines.

cs.CV

State-dependent temperature control in Langevin diffusions using numerical exploratory Hamiltonian-Jacobi-Bellman equations

Choosing how much noise to add in Langevin dynamics is essential for making these algorithms effective in challenging optimization problems. One promising approach is to determine this noise by solving Hamilton-Jacobi-Bellman (HJB) equations and their exploratory variants. Though these ideas have been demonstrated to work well in one dimension, extension to high-dimensional minimization has been limited by two unresolved numerical challenges: setting reliable control bounds and stably computing the second-order information (Hessians) required by the equations. These issues and the broader impact of HJB parameters have not been systematically examined. This work provides the first such investigation. We introduce principled control bounds and develop a physics-informed neural network framework that embeds the structure of exploratory HJB equations directly into training, stabilizing computation, and enabling accurate estimation of state-dependent noise in high-dimensional problems. Numerical experiments demonstrate that the resulting method remains robust and effective well beyond low-dimensional test cases.

math.NA

Constrained Zero-Sum Stochastic Linear-Quadratic Differential Game for Jump-Diffusion Systems with Random Coefficients

This paper studies a two-player zero-sum stochastic linear-quadratic (SLQ) differential game for controlled jump-diffusion systems with random coefficients, where the controls of both players are constrained to nonempty closed convex cones. Under a uniform convexity--concavity condition, we establish the existence and uniqueness of an open-loop saddle point and characterize it by a forward--backward stochastic differential equation with jumps (FBSDEJ) together with cone-type variational inequalities. Assuming the existence of positive bounded solutions to the associated system of indefinite extended stochastic Riccati equations with jumps (IESREJs), we derive a feedback-form representation of the unique open-loop saddle point by constructing predictable minimax selectors and combining the Meyer--It\^o formula with jumps, and the FBSDEJ characterization. Finally, under additional structural conditions, we prove the existence of positive bounded solutions to the IESREJs by a double-truncation approximation and a multidimensional BSDEJ comparison theorem.

math.OC

Continuous-time q-learning for Markov regime switching system under Tsallis entropy

This paper studies continuous-time q-learning (the continuous-time counterpart of Q-learning) for a Markov regime-switching system under Tsallis entropy regularization. The Tsallis entropy regularization yields an optimal policy distribution that may not necessarily be a Gibbs measure, thereby complicating algorithm design. Furthermore, to address the limited universality of current continuous-time regime-switching reinforcement learning algorithms (often restricted to the exploratory mean-variance framework), this study focuses on continuous-time q-learning for Markov regime-switching systems based on Tsallis entropy, aiming for a more universally applicable continuous-time reinforcement learning method. We establish the martingale characterization of the q-function under Tsallis entropy for continuous-time Markov regime-switching systems. We further design two q-learning algorithms that differ based on whether the Lagrange multiplier can be explicitly derived. We apply these algorithms to the continuous-time exploratory mean-variance portfolio optimization problem in a regime-switching market. Numerical experiments demonstrate the satisfactory performance of our q-learning algorithms.

math.OC

Fluctuation-induced quenching of chaos in quantum optics

Recent studies have extensively explored chaotic dynamics in quantum optical systems through the mean-field approximation, which corresponds to an ideal, fluctuation-free scenario. However, the inherent sensitivity of chaos to initial conditions implies that even minute fluctuations can be amplified, thereby questioning the applicability of this approximation. Here, we analyze these chaotic effects using stochastic Langevin equations or the Lindblad master equation. For systems operating at frequencies of $10^5$ to $10^7$ Hz, we demonstrate that room-temperature thermal fluctuations are sufficient to suppress chaos at the level of expectation values, even under weak nonlinearity. Furthermore, nonlinearity induces deviations from Gaussian phase-space distributions of the quantum state, revealing attractor-like features in the Wigner function. With increasing nonlinearity, the noise threshold for chaos suppression decreases, approaching the scale of vacuum fluctuations. These results provide a bidirectional validation of the quantum mechanical suppression of chaos.

quant-ph

Personalized Federated Distillation Assisted Vehicle Edge Caching Strategy

Vehicle edge caching is a promising technology that can significantly reduce the latency for vehicle users (VUs) to access content by pre-caching user-interested content at edge nodes. It is crucial to accurately predict the content that VUs are interested in without exposing their privacy. Traditional federated learning (FL) can protect user privacy by sharing models rather than raw data. However, the training of FL requires frequent model transmission, which can result in significant communication overhead. Additionally, vehicles may leave the road side unit (RSU) coverage area before training is completed, leading to training failures. To address these issues, in this paper, we propose a personalized federated distillation assisted vehicle edge caching strategy. The simulation results demonstrate that the proposed vehicle edge caching strategy has good robustness to variations in vehicle speed, significantly reducing communication overhead.

cs.LG