arXiv ScienceSearch

arXiv subjects

Chenglong Bao

Publications and source records attributed to Chenglong Bao.

At least 19 recordsLinked to original sources

Enhancing Low-resolution Image Representation Through Normalizing Flows

Low-resolution image representation can be regarded as a special form of sparse representation that retains only low-frequency information while discarding high-frequency components. This property reduces storage and transmission costs and benefits various image processing tasks. However, a key challenge is to preserve essential visual content while maintaining the ability to accurately reconstruct the original images. This work proposes LR2Flow, a nonlinear framework that learns low-resolution image representations by integrating wavelet tight frame blocks with normalizing flows. We conduct a reconstruction error analysis of the proposed network, which demonstrates the necessity of designing invertible neural networks in the wavelet tight frame domain. Experimental results on various tasks, including image rescaling, compression, and denoising, demonstrate the effectiveness of the learned representations and the robustness of the proposed framework. Code is available at https://github.com/Evanescentlove/LR2Flow-main.

cs.CV

Diffusion Based Unpaired Data Learning for Inverse Problems

Data is important in many deep learning-based inverse problem solvers. However, obtaining sufficient paired data in many scenarios remains highly challenging, while unpaired data is cheap. To maximize data utilization, this paper proposes LUD-DIF, a diffusion-based approach for solving inverse problems with unpaired data. Starting from the evidence lower bound (ELBO) of the joint distribution, we decouple it into two independent diffusion processes under the weak-coupling assumption. The method provides theoretical support from a variational inference perspective, derives the loss function, quantitatively analyzes the error bound introduced by the assumption, and offers a theorem-motivated heuristic for hyperparameter selection. Experimental results demonstrate that LUD-DIF achieves outstanding performance on multiple image inverse problems, validating its effectiveness and generalization capability in unpaired inverse problem settings.

cs.CV

A Structure-Exploiting Implicit-Explicit Trust Region Method for Computing Second-Order Stationary Points of the Landau-Brazovskii Model

This work focuses on the reliable computation of second-order stationary points in the high-dimensional nonconvex energy landscape of the Landau-Brazovskii (LB) model, a fundamental model for studying phases and phase transitions. For this purpose, we develop an efficient implicit-explicit trust region (IMEX-TR) method. Trust region (TR) methods can avoid saddle-point stagnation and guarantee convergence to second-order stationary points under appropriate conditions. However, their direct application to the LB model has been impractical because the Hessian is dense if treated directly. The proposed IMEX-TR method overcomes this difficulty by exploiting the Hessian's special structure: the linear interaction part is diagonal in reciprocal space, whereas the nonlinear bulk-energy part is diagonal in physical space. Based on this structure, we design an efficient solver for the TR subproblem that with globally convergent guarantee and enjoys FFT-based acceleration, with $\mathcal O(N \log N)$ complexity per iteration. Existing first-order gradient-based methods for the LB model only guarantee convergence to first-order stationary points and may stagnate at saddle points. In contrast, the proposed IMEX-TR method inherits the theoretical guarantee of converging to second-order stationary points while remaining computationally practical. Numerical experiments verify the theoretical properties of the algorithm and demonstrate its robustness in locating stable phases from different initial conditions. Numerical results also show that IMEX-TR can escape unstable stationary states reached by first-order schemes and converge to physically meaningful second-order stationary points. These results suggest that targeting second-order stationary points provides an effective computational paradigm for exploring complex free-energy landscapes and identifying stable or metastable states.

math.NA

Projected Sobolev Natural Gradient Descent for Efficient Neural Network Solution of the Gross--Pitaevskii Equation

This paper introduces a projected Sobolev natural gradient descent (NGD) method for computing ground states of the Gross--Pitaevskii equation. By projecting a continuous Riemannian Sobolev gradient flow onto the normalized neural network tangent space, we derive a discrete NGD algorithm that preserves the normalization constraint. The numerical implementation employs variational Monte Carlo with a hybrid sampling strategy to accurately account for the normalization constant. To enhance computational efficiency, a matrix-free Nyström-preconditioned conjugate gradient solver is adopted to approximate the NGD operator without explicit matrix assembly. Numerical experiments demonstrate that the proposed method converges significantly faster than physics-informed neural network approaches and exhibits near-constant wall-clock time and memory over the tested high-dimensional range when the hidden architecture is fixed and the implicit Gram operator is used. Moreover, the resulting neural-network solutions provide high-quality initial guesses that substantially accelerate subsequent refinement by traditional high-precision solvers.

math.NA

Unpaired Joint Distribution Modeling via Multi-Scale Image Representations

This paper studies the problem of learning a joint distribution from marginal observations, which is inherently ill-posed due to the ambiguity of feasible couplings. We propose LUD-MSR, a latent-variable probabilistic framework that models the joint distribution via auxiliary representations and optimizes evidence lower bounds using only marginal data. Under mild assumptions, we establish an upper bound on the distribution approximation error. This analysis reveals a trade-off in representation learning between domain consistency and information preservation. To address this trade-off, we introduce a Multi-Scale image Representation (MSR) mapping that exploits structural similarity at coarse scales while suppressing domain-specific variations. We show that MSR achieves a more favorable balance of this trade-off compared to existing approaches. Experiments on real-world denoising benchmarks, including cryo-electron microscopy (cryo-EM), demonstrate the effectiveness of the proposed framework.

cs.CV

Well-Posedness and Efficient Algorithms for Inverse Optimal Transport with Bregman Regularization

This work analyzes the inverse optimal transport (IOT) problem under Bregman regularization. We establish well-posedness results, including existence, uniqueness (up to equivalence classes of solutions), and stability, under several structural assumptions on the cost matrix. On the computational side, we investigate the existence of solutions to the optimization problem with general constraints on the cost matrix and provide a sufficient condition guaranteeing existence. In addition, we propose an inexact block coordinate descent (BCD) method for the problem with a strongly convex penalty term. In particular, when the penalty is quadratic, the subproblems admit a diagonal Hessian structure, which enables highly efficient element-wise Newton updates. We establish a linear convergence rate for the algorithm and demonstrate its practical performance through numerical experiments, including the validation of stability bounds, the investigation of regularization effects, and the application to a marriage matching dataset.

math.OC

Stratification for Nonlinear Semidefinite Programming

This paper introduces a stratification framework for nonlinear semidefinite programming (NLSDP) that reveals and utilizes the geometry behind the nonsmooth KKT system. Based on the \emph{index stratification} of $\mathbb{S}^n$ and its lift to the primal--dual space, a stratified variational analysis is developed. Specifically, we define the stratum-restricted regularity property, characterize it by the verifiable weak second order condition (W-SOC) and weak strict Robinson constraint qualification (W-SRCQ), and interpret the W-SRCQ geometrically via transversality, which provides its genericity over ambient space and stability along strata. The interactions of these properties across neighboring strata are further examined, leading to the conclusion that classical strong-form regularity conditions correspond to the local uniform validity of stratum-restricted counterparts. On the algorithmic side, a stratified Gauss--Newton method with normal steps and a correction mechanism is proposed for globally solving the KKT equation through a least-squares merit function. We demonstrate that the algorithm converges globally to directional stationary points. Moreover, under the W-SOC and the strict Robinson constraint qualification (SRCQ), it achieves local quadratic convergence to KKT pairs and eventually identifies the active stratum.

math.OC

On the Convergence Analysis of an Inexact Preconditioned Stochastic Model-Based Algorithm

This paper focuses on investigating an inexact stochastic model-based optimization algorithm that integrates preconditioning techniques for solving stochastic composite optimization problems. The proposed framework unifies and extends the fixed-metric stochastic model-based algorithm to its preconditioned and inexact variants. Convergence guarantees are established under mild assumptions for both weakly convex and convex settings, without requiring smoothness or global Lipschitz continuity of the objective function. By assuming a local Lipschitz condition, we derive nonasymptotic and asymptotic convergence rates measured by the gradient of the Moreau envelope. Furthermore, convergence rates in terms of the distance to the optimal solution set are obtained under an additional quadratic growth condition on the objective function. Numerical experiment results demonstrate the theoretical findings for the proposed algorithm.

math.OC

A Regularized Newton Method for Nonconvex Optimization with Global and Local Complexity Guarantees

Finding an $ε$-stationary point of a nonconvex function with a Lipschitz continuous Hessian is a central problem in optimization. Regularized Newton methods are a classical tool and have been studied extensively, yet they still face a trade-off between global and local convergence. Whether a parameter-free algorithm of this type can simultaneously achieve optimal global complexity and quadratic local convergence remains an open question. To bridge this long-standing gap, we propose a new class of regularizers constructed from the current and previous gradients, and leverage the conjugate gradient approach with a negative curvature monitor to solve the regularized Newton equation. The proposed algorithm is adaptive, requiring no prior knowledge of the Hessian Lipschitz constant, and achieves a global complexity of $O(ε^{-3/2})$ in terms of the second-order oracle calls, and $\tilde{O}(ε^{-7/4})$ for Hessian-vector products, respectively. When the iterates converge to a point where the Hessian is positive definite, the method exhibits quadratic local convergence. Preliminary numerical results, including training the physics-informed neural networks, illustrate the competitiveness of our algorithm.

math.OC

A Stochastic Gradient Descent Method for Globally Minimizing Nearly Convex Functions

This paper proposes a stochastic gradient descent method with an adaptive Gaussian noise term for the global minimization of nearly convex functions, which are nonconvex and possess multiple strict local minimizers. The noise term, independent of the gradient, is determined by the difference between the current function value and a lower bound estimate of the optimal value. In both probability space and state space, we show that the proposed algorithm converges linearly to a neighborhood of the global optimal solution. The size of this neighborhood depends on the variance of the gradient and the deviation between the estimated lower bound and the optimal value. In particular, when full gradient information is available and a sharp lower bound of the objective function is provided, the algorithm achieves linear convergence to the global optimum. Furthermore, we introduce a double-loop scheme that alternately updates the lower bound estimate and the optimization sequence, enabling convergence to a neighborhood of the global optimum that depends solely on the gradient variance. Numerical experiments on several benchmark problems demonstrate the effectiveness of the proposed algorithm.

math.OC

Interface Laplace Learning: Learnable Interface Term Helps Semi-Supervised Learning

We introduce a novel framework, called Interface Laplace learning, for graph-based semi-supervised learning. Motivated by the observation that an interface should exist between different classes where the function value is non-smooth, we introduce a Laplace learning model that incorporates an interface term. This model challenges the long-standing assumption that functions are smooth at all unlabeled points. In the proposed approach, we add an interface term to the Laplace learning model at the interface positions. We provide a practical algorithm to approximate the interface positions using k-hop neighborhood indices, and to learn the interface term from labeled data without artificial design. Our method is efficient and effective, and we present extensive experiments demonstrating that Interface Laplace learning achieves better performance than other recent semi-supervised learning approaches at extremely low label rates on the MNIST, FashionMNIST, and CIFAR-10 datasets.

cs.LG

The Global R-linear Convergence of Nesterov's Accelerated Gradient Method with Unknown Strongly Convex Parameter

The Nesterov accelerated gradient (NAG) method is an important extrapolation-based numerical algorithm that accelerates the convergence of the gradient descent method in convex optimization. When dealing with an objective function that is $μ$-strongly convex, selecting extrapolation coefficients dependent on $μ$ enables global R-linear convergence. In cases where $μ$ is unknown, a commonly adopted approach is to set the extrapolation coefficient using the original NAG method. This choice allows for achieving the optimal iteration complexity among first-order methods for general convex problems. However, it remains unknown whether the NAG method with an unknown strongly convex parameter exhibits global R-linear convergence for strongly convex problems. In this work, we answer this question positively by establishing the Q-linear convergence of certain constructed Lyapunov sequences. Furthermore, we extend our result to the global R-linear convergence of the accelerated proximal gradient method, which is employed for solving strongly convex composite optimization problems. Interestingly, these results contradict the findings of the continuous counterpart of the NAG method in [Su, Boyd, and Candés, J. Mach. Learn. Res., 2016, 17(153), 1-43], where the convergence rate by the suggested ordinary differential equation cannot exceed the $O(1/{\tt poly}(k))$ for strongly convex functions.

math.OC

Accelerated Gradient Methods with Gradient Restart: Global Linear Convergence

Gradient restarting has been shown to improve the numerical performance of accelerated gradient methods. This paper provides a mathematical analysis to understand these advantages. First, we establish global linear convergence guarantees for both the original and gradient restarted accelerated proximal gradient method when solving strongly convex composite optimization problems. Second, through analysis of the corresponding ordinary differential equation model, we prove the continuous trajectory of the gradient restarted Nesterov's accelerated gradient method exhibits global linear convergence for quadratic convex objectives, while the non-restarted version provably lacks this property by [Su, Boyd, and Candés, \textit{J. Mach. Learn. Res.}, 2016, 17(153), 1-43].

math.OC

Globalized distributionally robust chance-constrained support vector machine based on core sets

Support vector machine (SVM) is a well known binary linear classification model in supervised learning. This paper proposes a globalized distributionally robust chance-constrained (GDRC) SVM model based on core sets to address uncertainties in the dataset and provide a robust classifier. The globalization means that we focus on the uncertainty in the sample population rather than the small perturbations around each sample point. The uncertainty is mainly specified by the confidence region of the first- and second-order moments. The core sets are constructed to capture some small regions near the potential classification hyperplane, which helps improve the classification quality via the expected distance constraint of the random vector to core sets. We obtain the equivalent semi-definite programming reformulation of the GDRC SVM model under some appropriate assumptions. To deal with the large-scale problem, an approximation approach based on principal component analysis is applied to the GDRC SVM. The numerical experiments are presented to illustrate the effectiveness and advantage of our model.

math.OC

On the robust isolated calmness of a class of nonsmooth optimizations on Riemannian manifolds and its applications

This paper studies the robust isolated calmness property of the KKT solution mapping of a class of nonsmooth optimization problems on Riemannian manifolds. The manifold versions of the Robinson constraint qualification, the strict Robinson constraint qualification, and the second order conditions are defined and discussed. We show that the robust isolated calmness of the KKT solution mapping is equivalent to satisfying the M-SRCQ and M-SOSC conditions. Furthermore, under the above two conditions, we show that the Riemannian augmented Lagrangian method achieves a local linear convergence rate. Finally, we verify the proposed conditions and demonstrate the convergence rate on two minimization problems over the sphere and the manifold of fixed rank matrices.

math.OC

Reconstruction of dynamical systems from data without time labels

In this paper, we study the method to reconstruct dynamical systems from data without time labels. Data without time labels appear in many applications, such as molecular dynamics, single-cell RNA sequencing etc. Reconstruction of dynamical system from time sequence data has been studied extensively. However, these methods do not apply if time labels are unknown. Without time labels, sequence data becomes distribution data. Based on this observation, we propose to treat the data as samples from a probability distribution and try to reconstruct the underlying dynamical system by minimizing the distribution loss, sliced Wasserstein distance more specifically. Extensive experiment results demonstrate the effectiveness of the proposed method.

cs.LG

A Neural Network Framework for High-Dimensional Dynamic Unbalanced Optimal Transport

In this paper, we introduce a neural network-based method to address the high-dimensional dynamic unbalanced optimal transport (UOT) problem. Dynamic UOT focuses on the optimal transportation between two densities with unequal total mass, however, it introduces additional complexities compared to the traditional dynamic optimal transport (OT) problem. To efficiently solve the dynamic UOT problem in high-dimensional space, we first relax the original problem by using the generalized Kullback-Leibler (GKL) divergence to constrain the terminal density. Next, we adopt the Lagrangian discretization to address the unbalanced continuity equation and apply the Monte Carlo method to approximate the high-dimensional spatial integrals. Moreover, a carefully designed neural network is introduced for modeling the velocity field and source function. Numerous experiments demonstrate that the proposed framework performs excellently in high-dimensional cases. Additionally, this method can be easily extended to more general applications, such as crowd motion problem.

math.OC

Revealing the molecular structures of a-Al2O3(0001)-water interface by machine learning based computational vibrational spectroscopy

Solid-water interfaces are crucial to many physical and chemical processes and are extensively studied using surface-specific sum-frequency generation (SFG) spectroscopy. To establish clear correlations between specific spectral signatures and distinct interfacial water structures, theoretical calculations using molecular dynamics (MD) simulations are required. These MD simulations typically need relatively long trajectories (a few nanoseconds) to achieve reliable SFG response function calculations via the dipole-polarizability time correlation function. However, the requirement for long trajectories limits the use of computationally expensive techniques such as ab initio MD (AIMD) simulations, particularly for complex solid-water interfaces. In this work, we present a pathway for calculating vibrational spectra (IR, Raman, SFG) of solid-water interfaces using machine learning (ML)-accelerated methods. We employ both the dipole moment-polarizability correlation function and the surface-specific velocity-velocity correlation function approaches to calculate SFG spectra. Our results demonstrate the successful acceleration of AIMD simulations and the calculation of SFG spectra using ML methods. This advancement provides an opportunity to calculate SFG spectra for the complicated solid-water systems more rapidly and at a lower computational cost with the aid of ML.

cond-mat.mtrl-sci