arXiv ScienceSearch

arXiv subjects

Lingzi Jin

Publications and source records attributed to Lingzi Jin.

2 recordsLinked to original sources

A sequential regularized piecewise affine algorithm for nonconvex nonsmooth multicomposite optimization in RNN training

This paper focuses on a class of nonconvex nonsmooth multicomposite optimization problems for training RNNs (Recurrent Neural Networks). We first establish easily verifiable conditions under which an approximate first-order d-stationary point of the problem is guaranteed to be an approximate second-order d-stationary point. Subsequently, we propose a sequential regularized piecewise affine algorithm, SRPA, that minimizes a regularized piecewise affine function at each iteration, where the idea of active-set strategy is incorporated to reduce the computational cost per iteration. Leveraging the aforementioned conditions for second-order d-stationarity, we establish the global convergence of SRPA to second-order d-stationary points and the complexity bound of $ \mathcal{O}(ε^{-2}) $ for obtaining an $ ε$-approximate second-order d-stationary point. Finally, numerical experiments for training RNNs on synthetic and real-world datasets demonstrate the promising performance of SRPA compared with state-of-the-art algorithms.

math.OC

Nonconvex Nonsmooth Multicomposite Optimization and Its Applications to Recurrent Neural Networks

We consider a class of nonconvex nonsmooth multicomposite optimization problems where the objective function consists of a Tikhonov regularizer and a composition of multiple nonconvex nonsmooth component functions. Such optimization problems arise from tangible applications in machine learning and beyond. To define and compute its first-order and second-order d(irectional)-stationary points effectively, we first derive the closed-form expression of the tangent cone for the feasible region of its constrained reformulation. Building on this, we establish its equivalence with the corresponding constrained and $\ell_1$-penalty reformulations in terms of global optimality and d-stationarity. The equivalence offers indirect methods to attain the first-order and second-order d-stationary points of the original problem in certain cases. We apply our results to the training process of recurrent neural networks (RNNs).

math.OC