arXiv ScienceSearch

arXiv subjects

Hanbaek Lyu

Publications and source records attributed to Hanbaek Lyu.

At least 19 recordsLinked to original sources

Finding Koopman Invariant Subspaces via Personalized PageRank

Selecting a finite dictionary of observables whose span is Koopman-invariant is a central challenge in data-driven Koopman operator approximation. We address this problem by exploiting zero-block structure in Extended Dynamic Mode Decomposition (EDMD) matrices. We show that any sub-dictionary whose span is Koopman-invariant induces an exact zero block in the EDMD matrix, even for finite data. We then show that such blocks can be detected by applying PageRank to a row-normalized EDMD matrix constructed from a large initial dictionary. The theory extends to approximately invariant subspaces and yields stronger guarantees for personalized PageRank (PPR) when the seed observables lie inside the target block and reach all observables in that block. Combining EDMD concentration bounds with PageRank perturbation theory gives end-to-end detection guarantees with $O(1/\sqrt{M})$ finite-sample scaling and explicit constants. More generally, without assuming an invariant subspace exists, high PPR mass on a sub-dictionary controls discounted multi-step leakage from the seed observables. Numerical experiments on the Duffing oscillator, Van der Pol oscillator, Lorenz system, and a three-well Ramachandran potential suggest that the method identifies compact, interpretable dictionaries with accurate predictions.

math.DS

Likelihood landscape of binary latent model on a tree

We investigate the optimization landscape of maximum likelihood estimation (MLE) for the Cavender-Farris-Neyman (CFN) model, a two-state latent tree model fundamental to statistical phylogenetics and the ferromagnetic Ising model. Although the log-likelihood function is non-concave and may admit many critical points, simple coordinate maximization algorithms are remarkably effective in practice. We provide the first theoretical justification for this success. We prove that sufficiently deep inside the reconstruction regime, the population log-likelihood is strongly concave and smooth within a box around the true parameter, whose size is independent of tree topology and number of leaves. This fundamental result implies that the empirical landscape shares these regularity properties with high probability given polynomial sample complexity and also that coordinate maximization converges exponentially fast to an $O(1/\sqrt{m})$-consistent MLE. Our analysis centers on a novel decay property of the population Hessian: diagonal entries remain large while off-diagonal entries decay exponentially with graph distance. These results provide rigorous theoretical evidence for the efficacy of likelihood-based tree inference and suggest broader principles for latent variable models.

math.ST

Scaling limit of Sinkhorn-rescaled Random Matrices via Stability of Static Schrödinger Bridges

We analyze the asymptotic behavior and scaling limits of large random matrices rescaled via the Sinkhorn algorithm to match prescribed row and column margins. For a random matrix with independent sub-exponential entries, we show that its Sinkhorn rescaling concentrates around the rescaling of its mean matrix, both at the level of the Schrödinger potentials and as random measures on the unit square, with explicit non-asymptotic rates. As the dimensions grow, the rescaled random matrix converges to the continuous static Schrödinger bridge (SSB) determined by the limiting margins and reference density. Around this scaling limit we develop a fluctuation theory: bulk rigidity for the empirical spectral distribution of the associated sample covariance matrix, and a central limit theorem for the empirical Schrödinger potentials of the rescaled empirical mean. Our analysis is driven by a new quantitative stability theory for the SSB, developed in three forms: Lipschitz continuity in the Hellinger distance under perturbations of the reference measure (kernel stability); Hölder-$1/2$ continuity in the Hellinger distance under $L^1$ perturbations of the margins (margin stability); and $L^\infty$ stability of the discrete Schrödinger potentials under margin perturbation (potential stability). Translated to the discrete random-matrix setting, these bounds yield the concentration and scaling-limit results, while a local law for random Gram matrices with a non-uniform variance profile drives the bulk rigidity. Our SSB stability theory may be of independent interest.

math.PR

Diffusive Scaling limit of stochastic Box-Ball systems and PushTASEP

We introduce the Stochastic Box-Ball System (SBBS), a probabilistic cellular automaton that generalizes the classic Takahashi-Satsuma Box-Ball System. In SBBS, particles are transported by a carrier with a fixed capacity that may fail to pick up any given particle with a fixed probability $ε$. This model interpolates between two known integrable systems: the Box-Ball System (as $ε\rightarrow 0$) and the PushTASEP (as $ε\rightarrow 1$). We show that the long-term behavior of SBBS is governed by isolated particles and the occasional emergence of short solitons, which can form longer solitons but are more likely to fall apart. More precisely, we first show that all particles are isolated except for a $1/\sqrt{n}$-fraction of times in any given $n$ steps, and solitons keep forming for this fraction of times. We then show that under diffusive scaling, both SBBS (for any carrier capacity) and PushTASEP converge weakly to semimartingale reflecting Brownian Motions (SRBMs) on the Weyl chamber with explicit covariance and reflection matrices, which are consistent with the microscale relations between these systems. The reflection matrix for SBBS is determined by how 2-solitons behave and exhibit ``solitonic bias'' visible in the diffusive scale. Our proof relies on a new, extended SRBM invariance principle that we develop in this work. This principle can handle processes with complex boundary behavior that can be written as "overdetermined" Skorokhod decompositions, which is crucial for analyzing the complex solitonic interaction in SBBS. We believe this tool may be of independent interest.

math.PR

Regularized Overestimated Newton

We propose Regularized Overestimated Newton (RON), a Newton-type method with low per-iteration cost and strong global and local convergence guarantees for smooth convex optimization. RON interpolates between gradient descent and globally regularized Newton, with behavior determined by the largest Hessian overestimation error. Globally, when the optimality gap of the objective is large, RON achieves an accelerated $O(n^{-2})$ convergence rate; when small, its rate becomes $O(n^{-1})$. Locally, RON converges superlinearly and linearly when the overestimation is exact and inexact, respectively, toward possibly non-isolated minima under the local Quadratic Growth (QG) condition. The linear rate is governed by an improved effective condition number depending on the overestimation error. Leveraging a recent randomized rank-$k$ Hessian approximation algorithm, we obtain a practical variant with $O(\text{dim}\cdot k^2)$ cost per iteration. When the Hessian rank is uniformly below $k$, RON achieves a per-iteration cost comparable to that of first-order methods while retaining the superior convergence rates even in degenerate local landscapes. We validate our theoretical findings through experiments on entropic optimal transport and inverse problems.

math.OC

Sobolev acceleration for neural networks

Sobolev training, which integrates target derivatives into the loss functions, has been shown to accelerate convergence and improve generalization compared to conventional $L^2$ training. However, the underlying mechanisms of this training method remain only partially understood. In this work, we present the first rigorous theoretical framework proving that Sobolev training accelerates the convergence of Rectified Linear Unit (ReLU) networks. Under a student-teacher framework with Gaussian inputs and shallow architectures, we derive exact formulas for population gradients and Hessians, and quantify the improvements in conditioning of the loss landscape and gradient-flow convergence rates. Extensive numerical experiments validate our theoretical findings and show that the benefits of Sobolev training extend to modern deep learning tasks.

cs.LG

Sample Complexity of Branch-length Estimation by Maximum Likelihood

We consider the branch-length estimation problem on a bifurcating tree: a character evolves along the edges of a binary tree according to a two-state symmetric Markov process, and we seek to recover the edge transition probabilities from repeated observations at the leaves. This problem arises in phylogenetics, and is related to latent tree graphical model inference. In general, the log-likelihood function is non-concave and may admit many critical points. Nevertheless, simple coordinate maximization has been known to perform well in practice, defying the complexity of the likelihood landscape. In this work, we provide the first theoretical guarantee as to why this might be the case. We show that deep inside the Kesten-Stigum reconstruction regime, provided with polynomially many $m$ samples (assuming the tree is balanced), there exists a universal parameter regime (independent of the size of the tree) where the log-likelihood function is strongly concave and smooth with high probability. On this high-probability likelihood landscape event, we show that the standard coordinate maximization algorithm converges exponentially fast to the maximum likelihood estimator, which is within $O(1/\sqrt{m})$ from the true parameter, provided a sufficiently close initial point.

stat.CO

Large random matrices with given margins

We study large random matrices with i.i.d. entries conditioned to have prescribed row and column sums (margins), a problem connected to relative entropy minimization, Schrödinger bridges, contingency tables, and random graphs with given degree sequences. Our central result is a `transference principle': the complex margin-conditioned matrix can be closely approximated by a simpler matrix whose entries are independent and drawn from an exponential tilting of the original model. The tilt parameters are determined by the sum of two potentials. We establish phase diagrams for `tame margins', where these potentials are uniformly bounded. This framework resolves a 2011 conjecture by Chatterjee, Diaconis, and Sly on $δ$-tame degree sequences and generalizes a sharp phase transition in contingency tables obtained by Dittmer, Lyu, and Pak in 2020. For tame margins, we show that a generalized Sinkhorn algorithm can compute the potentials at a dimension-free exponential rate. Our limit theory further establishes that for a convergent sequence of tame margins, the potentials converge as fast as the margins converge. We apply this framework and obtain several key results for the conditioned matrix: The marginal distribution of any single entry is asymptotically an exponential tilting of the base measure, resolving a 2010 conjecture by Barvinok on contingency tables. The conditioned matrix concentrates in cut norm around a `typical table' (the expectation of the tilted model), which acts as a static Schrödinger bridge between the margins. The empirical singular value distribution of the rescaled matrix converges to an explicit law determined by the variance profile of the tilted model. In particular, we confirm the universality of the Marchenko-Pastur law for constant linear margins.

math.PR

Likelihood-Based Root State Reconstruction on a Tree: Sensitivity to Parameters and Applications

We consider a broadcasting problem on a tree where a binary digit (e.g., a spin or a nucleotide's purine/pyrimidine type) is propagated from the root to the leaves through symmetric noisy channels on the edges that randomly flip the state with edge-dependent probabilities. The goal of the reconstruction problem is to infer the root state given the observations at the leaves only. Specifically, we study the sensitivity of maximum likelihood estimation (MLE) to uncertainty in the edge parameters under this model, which is also known as the Cavender-Farris-Neyman (CFN) model. Our main result shows that when the true flip probabilities are sufficiently small, the posterior root mean (or magnetization of the root) under estimated parameters (within a constant factor) agrees with the root spin with high probability and deviates significantly from it with negligible probability. This provides theoretical justification for the practical use of MLE in ancestral sequence reconstruction in phylogenetics, where branch lengths (i.e., the edge parameters) must be estimated. As a separate application, we derive an approximation for the gradient of the population log-likelihood of the leaf states under the CFN model, with implications for branch length estimation via coordinate maximization.

math.PR

Block majorization-minimization with diminishing radius for constrained nonsmooth nonconvex optimization

Block majorization-minimization (BMM) is a simple iterative algorithm for constrained nonconvex optimization that sequentially minimizes majorizing surrogates of the objective function in each block while the others are held fixed. BMM entails a large class of optimization algorithms such as block coordinate descent and its proximal-point variant, expectation-minimization, and block projected gradient descent. We first establish that for general constrained nonsmooth nonconvex optimization, BMM with $ρ$-strongly convex and $L_g$-smooth surrogates can produce an $ε$-approximate first-order optimal point within $\widetilde{O}((1+L_g+ρ^{-1})ε^{-2})$ iterations and asymptotically converges to the set of first-order optimal points. Next, we show that BMM combined with trust-region methods with diminishing radius has an improved complexity of $\widetilde{O}((1+L_g) ε^{-2})$, independent of the inverse strong convexity parameter $ρ^{-1}$, allowing improved theoretical and practical performance with `flat' surrogates. Our results hold robustly even when the convex sub-problems are solved as long as the optimality gaps are summable. Central to our analysis is a novel continuous first-order optimality measure, by which we bound the worst-case sub-optimality in each iteration by the first-order improvement the algorithm makes. We apply our general framework to obtain new results on various algorithms such as the celebrated multiplicative update algorithm for nonnegative matrix factorization by Lee and Seung, regularized nonnegative tensor decomposition, and the classical block projected gradient descent algorithm. Lastly, we numerically demonstrate that the additional use of diminishing radius can improve the convergence rate of BMM in many instances.

math.OC

Convergence and complexity of block majorization-minimization for constrained block-Riemannian optimization

Block majorization-minimization (BMM) is a simple iterative algorithm for nonconvex optimization that sequentially minimizes a majorizing surrogate of the objective function in each block coordinate while the other block coordinates are held fixed. We consider a family of BMM algorithms for minimizing smooth nonconvex objectives, where each parameter block is constrained within a subset of a Riemannian manifold. We establish that this algorithm converges asymptotically to the set of stationary points, and attains an $ε$-stationary point within $\widetilde{O}(ε^{-2})$ iterations. In particular, the assumptions for our complexity results are completely Euclidean when the underlying manifold is a product of Euclidean or Stiefel manifolds, although our analysis makes explicit use of the Riemannian geometry. Our general analysis applies to a wide range of algorithms with Riemannian constraints: Riemannian MM, block projected gradient descent, optimistic likelihood estimation, geodesically constrained subspace tracking, robust PCA, and Riemannian CP-dictionary-learning. We experimentally validate that our algorithm converges faster than standard Euclidean algorithms applied to the Riemannian setting.

math.OC

Stochastic optimization with arbitrary recurrent data sampling

For obtaining optimal first-order convergence guarantee for stochastic optimization, it is necessary to use a recurrent data sampling algorithm that samples every data point with sufficient frequency. Most commonly used data sampling algorithms (e.g., i.i.d., MCMC, random reshuffling) are indeed recurrent under mild assumptions. In this work, we show that for a particular class of stochastic optimization algorithms, we do not need any other property (e.g., independence, exponential mixing, and reshuffling) than recurrence in data sampling algorithms to guarantee the optimal rate of first-order convergence. Namely, using regularized versions of Minimization by Incremental Surrogate Optimization (MISO), we show that for non-convex and possibly non-smooth objective functions, the expected optimality gap converges at an optimal rate $O(n^{-1/2})$ under general recurrent sampling schemes. Furthermore, the implied constant depends explicitly on the `speed of recurrence', measured by the expected amount of time to visit a given data point either averaged (`target time') or supremized (`hitting time') over the current location. We demonstrate theoretically and empirically that convergence can be accelerated by selecting sampling algorithms that cover the data set most effectively. We discuss applications of our general framework to decentralized optimization and distributed non-negative matrix factorization.

math.OC

Phase transition in one-dimensional excitable media with variable interaction range

We investigate two discrete models of excitable media on a one-dimensional integer lattice $\mathbb{Z}$: the $κ$-color Cyclic Cellular Automaton (CCA) and the $κ$-color Firefly Cellular Automaton (FCA). In both models, sites are assigned uniformly random colors from $\mathbb{Z}/κ\mathbb{Z}$. Neighboring sites with colors within a specified interaction range $r$ tend to synchronize their colors upon a particular local event of 'excitation'. We establish that there are three phases of CCA/FCA on $\mathbb{Z}$ as we vary the interaction range $r$. First, if $r$ is too small (undercoupled), there are too many non-interacting pairs of colors, and the whole graph $\mathbb{Z}$ will be partitioned into non-interacting intervals of sites with no excitation within each interval. If $r$ is within a sweet spot (critical), then we show the system clusters into ever-growing monochromatic intervals. For the critical interaction range $r=\lfloor κ/2 \rfloor$, we show the density of edges of differing colors at time $t$ is $Θ(t^{-1/2})$ and each site excites $Θ(t^{1/2})$ times up to time $t$. Lastly, if $r$ is too large (overcoupled), then neighboring sites can excite each other and such 'defects' will generate waves of excitation at a constant rate so that each site will get excited at least at a linear rate. For the special case of FCA with $r=\lfloor 2/κ\rfloor+1$, we show that every site will become $(κ+1)$-periodic eventually.

math.PR

Supervised low-rank semi-nonnegative matrix factorization with frequency regularization for forecasting spatio-temporal data

We propose a novel methodology for forecasting spatio-temporal data using supervised semi-nonnegative matrix factorization (SSNMF) with frequency regularization. Matrix factorization is employed to decompose spatio-temporal data into spatial and temporal components. To improve clarity in the temporal patterns, we introduce a nonnegativity constraint on the time domain along with regularization in the frequency domain. Specifically, regularization in the frequency domain involves selecting features in the frequency space, making an interpretation in the frequency domain more convenient. We propose two methods in the frequency domain: soft and hard regularizations, and provide convergence guarantees to first-order stationary points of the corresponding constrained optimization problem. While our primary motivation stems from geophysical data analysis based on GRACE (Gravity Recovery and Climate Experiment) data, our methodology has the potential for wider application. Consequently, when applying our methodology to GRACE data, we find that the results with the proposed methodology are comparable to previous research in the field of geophysical sciences but offer clearer interpretability.

stat.ML

Scaling limit of soliton lengths in a multicolor box-ball system

The box-ball systems are integrable cellular automata whose long-time behavior is characterized by soliton solutions, with rich connections to other integrable systems such as the Korteweg-de Vries equation. In this paper, we consider a multicolor box-ball system with two types of random initial configurations and obtain sharp scaling limits of the soliton lengths as the system size tends to infinity. We obtain a sharp scaling limit of soliton lengths that turns out to be more delicate than that in the single color case established in [Levine, Lyu, Pike '20]. A large part of our analysis is devoted to studying the associated carrier process, which is a multi-dimensional Markov chain on the orthant, whose excursions and running maxima are closely related to soliton lengths. We establish the sharp scaling of its ruin probabilities, Skorokhod decomposition, strong law of large numbers, and weak diffusive scaling limit to a semimartingale reflecting Brownian motion with explicit parameters. We also establish and utilize complementary descriptions of the soliton lengths and numbers in terms of modified Greene-Kleitman invariants for the box-ball systems and associated circular exclusion processes.

math.PR

Convergence and Complexity Guarantee for Inexact First-order Riemannian Optimization Algorithms

We analyze inexact Riemannian gradient descent (RGD) where Riemannian gradients and retractions are inexactly (and cheaply) computed. Our focus is on understanding when inexact RGD converges and what is the complexity in the general nonconvex and constrained setting. We answer these questions in a general framework of tangential Block Majorization-Minimization (tBMM). We establish that tBMM converges to an $ε$-stationary point within $O(ε^{-2})$ iterations. Under a mild assumption, the results still hold when the subproblem is solved inexactly in each iteration provided the total optimality gap is bounded. Our general analysis applies to a wide range of classical algorithms with Riemannian constraints including inexact RGD and proximal gradient method on Stiefel manifolds. We numerically validate that tBMM shows improved performance over existing methods when applied to various problems, including nonnegative tensor decomposition with Riemannian constraints, regularized nonnegative matrix factorization, and low-rank matrix recovery problems.

math.OC

New directions in algebraic statistics: Three challenges from 2023

In the last quarter of a century, algebraic statistics has established itself as an expanding field which uses multilinear algebra, commutative algebra, computational algebra, geometry, and combinatorics to tackle problems in mathematical statistics. These developments have found applications in a growing number of areas, including biology, neuroscience, economics, and social sciences. Naturally, new connections continue to be made with other areas of mathematics and statistics. This paper outlines three such connections: to statistical models used in educational testing, to a classification problem for a family of nonparametric regression models, and to phase transition phenomena under uniform sampling of contingency tables. We illustrate the motivating problems, each of which is for algebraic statistics a new direction, and demonstrate an enhancement of related methodologies.

math.ST

On the Complexity of First-Order Methods in Stochastic Bilevel Optimization

We consider the problem of finding stationary points in Bilevel optimization when the lower-level problem is unconstrained and strongly convex. The problem has been extensively studied in recent years; the main technical challenge is to keep track of lower-level solutions $y^*(x)$ in response to the changes in the upper-level variables $x$. Subsequently, all existing approaches tie their analyses to a genie algorithm that knows lower-level solutions and, therefore, need not query any points far from them. We consider a dual question to such approaches: suppose we have an oracle, which we call $y^*$-aware, that returns an $O(ε)$-estimate of the lower-level solution, in addition to first-order gradient estimators {\it locally unbiased} within the $Θ(ε)$-ball around $y^*(x)$. We study the complexity of finding stationary points with such an $y^*$-aware oracle: we propose a simple first-order method that converges to an $ε$ stationary point using $O(ε^{-6}), O(ε^{-4})$ access to first-order $y^*$-aware oracles. Our upper bounds also apply to standard unbiased first-order oracles, improving the best-known complexity of first-order methods by $O(ε)$ with minimal assumptions. We then provide the matching $Ω(ε^{-6})$, $Ω(ε^{-4})$ lower bounds without and with an additional smoothness assumption on $y^*$-aware oracles, respectively. Our results imply that any approach that simulates an algorithm with an $y^*$-aware oracle must suffer the same lower bounds.

math.OC