arXiv ScienceSearch

arXiv subjects

Mengmeng Song

Publications and source records attributed to Mengmeng Song.

11 recordsLinked to original sources

On the local and global minimizers of the smooth stress function in Euclidean Distance Matrix problems

We consider the nonconvex minimization problem, with quartic objective function, that arises in the exact recovery of a configuration matrix $P\in \R^{nd}$ of $n$ points when a Euclidean distance matrix, \EDMp, is given with embedding dimension $d$. It is an open question in the literature whether there are conditions such that the minimization problem admits a local nonglobal minimizer, \lngmp. We prove that all second-order stationary points are global minimizers whenever $n \leq d + 1$. {And, for $d=1$ and $n\geq 7>d+1$, we present an example where we can analytically exhibit a local nonglobal minimizer. For more general cases,} we numerically find a second-order stationary point and then prove that there indeed exists a nearby \lngm for the quartic nonconvex minimization problem. Thus, we answer the previously open question about their existence in the affirmative. Our approach to finding the \lngm is novel in that we first exploit the translation and rotation invariance to remove the singularities of the Hessian, and reduce the size of the problem from $nd$ variables in $P$ to $(n-1)d - d(d-1)/2$ variables. This allows for stabilizing Newton's method, and for finding examples that satisfy the strict second order sufficient optimality conditions. The motivation for being able to find global minima is to obtain \emph{exact recovery} of the configuration matrix, even in the cases where the data is noisy and/or incomplete, without resorting to approximating solutions from convex (semidefinite programming) relaxations. In the process of our work we present new insights into when \lngmp s of the smooth stress function do and do not exist.

math.OC

Split-Merge: A Difference-based Approach for Dominant Eigenvalue Problem

The computation of the dominant eigenpair for symmetric positive semidefinite matrices is fundamental in numerical optimization. This work shifts the paradigm from the classical Rayleigh quotient to an unconstrained difference formulation, whose global optimum recovers the dominant eigenpair. Within this framework, we prove that gradient descent with a constant step-size $\alpha \in (0, 1)$ converges almost surely to the global optimum at a local linear rate. This analysis thereby reinterprets the classical power method as the conservative special case $\alpha=1/2$ and rigorously establishes its asymptotic sub-optimality. To advance this first-order scheme, we propose the Split-Merge algorithm based on the majorization-minimization principle. After splitting the matrix, we introduce auxiliary vectors to effectively merge the decomposition factors, resulting in a matrix-free and parameter-free iteration that captures tighter curvature information. We establish that Split-Merge converges almost surely to a global minimizer, and show that the iteration exhibits a spectral peeling mechanism that suppresses the targeted eigenspace, potentially surpassing the static linear rate of power iterations. Numerical evaluations across synthetic and real-world datasets confirm that our method has scalable efficiency, achieving speed-ups exceeding $10\times$ over the power method, with performance comparable to subspace iterations.

math.OC

Enhancing Deep Neural Network Training Efficiency and Performance through Linear Prediction

Deep neural networks (DNN) have achieved remarkable success in various fields, including computer vision and natural language processing. However, training an effective DNN model still poses challenges. This paper aims to propose a method to optimize the training effectiveness of DNN, with the goal of improving model performance. Firstly, based on the observation that the DNN parameters change in certain laws during training process, the potential of parameter prediction for improving model training efficiency and performance is discovered. Secondly, considering the magnitude of DNN model parameters, hardware limitations and characteristics of Stochastic Gradient Descent (SGD) for noise tolerance, a Parameter Linear Prediction (PLP) method is exploit to perform DNN parameter prediction. Finally, validations are carried out on some representative backbones. Experiment results show that compare to the normal training ways, under the same training conditions and epochs, by employing proposed PLP method, the optimal model is able to obtain average about 1% accuracy improvement and 0.01 top-1/top-5 error reduction for Vgg16, Resnet18 and GoogLeNet based on CIFAR-100 dataset, which shown the effectiveness of the proposed method on different DNN structures, and validated its capacity in enhancing DNN training efficiency and performance.

cs.LG

Linear programming on the Stiefel manifold

Linear programming on the Stiefel manifold (LPS) is studied for the first time. It aims at minimizing a linear objective function over the set of all $p$-tuples of orthonormal vectors in ${\mathbb R}^n$ satisfying $k$ additional linear constraints. Despite the classical polynomial-time solvable case $k=0$, general (LPS) is NP-hard. According to the Shapiro-Barvinok-Pataki theorem, (LPS) admits an exact semidefinite programming (SDP) relaxation when $p(p+1)/2\le n-k$, which is tight when $p=1$. Surprisingly, we can greatly strengthen this sufficient exactness condition to $p\le n-k$, which covers the classical case $p\le n$ and $k=0$. Regarding (LPS) as a smooth nonlinear programming problem, we reveal a nice property that under the linear independence constraint qualification, the standard first- and second-order {\it local} necessary optimality conditions are sufficient for {\it global} optimality when $p+1\le n-k$.

math.OC

On globally solving nonconvex trust region subproblem via projected gradient method

The trust region subproblem (TRS) is to minimize a possibly nonconvex quadratic function over a Euclidean ball. There are typically two cases for (TRS), the so-called ``easy case'' and ``hard case''. Even in the ``easy case'', the sequence generated by the classical projected gradient method (PG) may converge to a saddle point at a sublinear local rate, when the initial point is arbitrarily selected from a nonzero measure feasible set. To our surprise, when applying (PG) to solve a cheap and possibly nonconvex reformulation of (TRS), the generated sequence initialized with {\it any} feasible point almost always converges to its global minimizer. The local convergence rate is at least linear for the ``easy case'', without assuming that we have possessed the information that the ``easy case'' holds. We also consider how to use (PG) to globally solve equality-constrained (TRS).

math.OC

On local minimizers of generalized trust-region subproblem

Generalized trust-region subproblem (GT) is a nonconvex quadratic optimization with a single quadratic constraint. It reduces to the classical trust-region subproblem (T) if the constraint set is a Euclidean ball. (GT) is polynomially solvable based on its inherent hidden convexity. In this paper, we study local minimizers of (GT). Unlike (T) with at most one local nonglobal minimizer, we can prove that two-dimensional (GT) has at most two local nonglobal minimizers, which are shown by example to be attainable. The main contribution of this paper is to prove that, at any local nonglobal minimizer of (GT), not only the strict complementarity condition holds, but also the standard second-order sufficient optimality condition remains necessary. As a corollary, finding all local nonglobal minimizers of (GT) or proving the nonexistence can be done in polynomial time. Finally, for (GT) in complex domain, we prove that there is no local nonglobal minimizer, which demonstrates that real-valued optimization problem may be more difficult to solve than its complex version.

math.OC

Local Optimality Conditions for a Class of Hidden Convex Optimization

Hidden convex optimization is such a class of nonconvex optimization problems that can be globally solved in polynomial time via equivalent convex programming reformulations. In this paper, we focus on checking local optimality in hidden convex optimization. We first introduce a class of hidden convex optimization problems by jointing the classical nonconvex trust-region subproblem (TRS) with convex optimization (CO), and then present a comprehensive study on local optimality conditions. In order to guarantee the existence of a necessary and sufficient condition for local optimality, we need more restrictive assumptions. To our surprise, while (TRS) has at most one local non-global minimizer and (CO) has no local non-global minimizer, their joint problem could have more than one local non-global minimizer.

math.OC

Polyak's convexity theorem, Yuan's lemma and S-lemma: extensions and applications

We extend Polyak's theorem on the convexity of joint numerical range from three to any number of quadratic forms on condition that they can be generated by three quadratic forms with a positive definite linear combination. Our new result covers the fundamental Dines's theorem. As applications, we further extend Yuan's lemma and S-lemma, respectively. Our extended Yuan's lemma is used to build a more generalized assumption than that of Haeser (J. Optim. Theory Appl. 174(3): 641-649, 2017), under which the standard second-order necessary optimality condition holds at local minimizer. The extended S-lemma reveals strong duality of homogeneous quadratic optimization problem with two bilateral quadratic constraints.

math.OC

Trust-region and $p$-regularized subproblems: local nonglobal minimum is the second smallest objective function value among all first-order stationary points

The local nonglobal minimizer of trust-region subproblem, if it exists, is shown to have the second smallest objective function value among all KKT points. This new property is extended to $p$-regularized subproblem. As a corollary, we show for the first time that finding the local nonglobal minimizer of Nesterov-Polyak subproblem corresponds to a generalized eigenvalue problem.

math.OC

On Local Minimizers of Quadratically Constrained Nonconvex Homogeneous Quadratic Optimization with at Most Two Constraints

We study nonconvex homogeneous quadratically constrained quadratic optimization with one or two constraints, denoted by (QQ1) and (QQ2), respectively. (QQ2) contains (QQ1), trust region subproblem (TRS) and ellipsoid regularized total least squares problem as special cases. It is known that there is a necessary and sufficient optimality condition for the global minimizer of (QQ2). In this paper, we first show that any local minimizer of (QQ1) is globally optimal. Unlike its special case (TRS) with at most one local non-global minimizer, (QQ2) may have infinitely many local non-global minimizers. At any local non-global minimizer of (QQ2), both linearly independent constraint qualification and strict complementary condition hold, and the Hessian of the Lagrangian has exactly one negative eigenvalue. As a main contribution, we prove that the standard second-order sufficient optimality condition for any strict local non-global minimizer of (QQ2) remains necessary. Applications and the impossibility of further extension are discussed.

math.OC

A unified approach for projections onto the intersection of $\ell_1$ and $\ell_2$ balls or spheres

This paper focuses on designing a unified approach for computing the projection onto the intersection of an $\ell_1$ ball/sphere and an $\ell_2$ ball/sphere. We show that the major computational efforts of solving these problems all rely on finding the root of the same piecewisely quadratic function, and then propose a unified numerical method to compute the root. In particular, we design breakpoint search methods with/without sorting incorporated with bisection, secant and Newton methods to find the interval containing the root, on which the root has a closed form. It can be shown that our proposed algorithms without sorting possess $O(n log n)$ worst-case complexity and $O(n)$ in practice. The efficiency of our proposed algorithms are demonstrated in numerical experiments.

math.OC