NEPv Approach for Optimization on Stiefel Manifold with the $(2,1)$-norm Regularization
Row-sparse projection provides a useful tool in machine learning (ML) when it comes to, for example, feature selection, aiming to choose most relevant features for various ML objectives. One way to seek a high quality row-sparse projection is to combine an ML objective, such as the ones for PCA, LDA, and OCCA, with the matrix $(2,1)$-norm regularization which is nonsmooth. Such combinations result in challenging optimization problems on the Stiefel manifold that need to be solved efficiently. In this paper, a unifying NEPv framework is established to efficiently deal with optimization on the Stiefel manifold with the $(2,1)$-norm regularization. The effect of the $(2,1)$-norm regularization is also investigated. The wide applicability of the framework is demonstrated through the combinations of common learning objectives in today's data science applications with the $(2,1)$-norm regularization. Numerical experiments are presented to illustrate the use of the NEPv approach and to gain insights as to what a proper regularizing parameter should have in real-world applications.