arXiv ScienceSearch

arXiv subjects

Tianmin Yu

Publications and source records attributed to Tianmin Yu.

7 recordsLinked to original sources

DiffATS: Diffusion in Aligned Tensor Space

Direct diffusion modeling of high-resolution spatiotemporal fields is computationally challenging. Parameter-efficient primitives address this by representing high-dimensional data with a compact set of parameters. In this paper, we construct data-dependent tensor primitives without pretrained compression autoencoders. Our construction starts from Tucker decomposition, which captures low-rank multilinear structure through a core tensor and mode-wise factors. However, Tucker factors are non-unique: the same tensor can be represented by different rotated factors, which complicates generative modeling. We address this issue with orthogonal Procrustes (OP) alignment. Specifically, we select medoid anchor matrices from the data and align the factor matrices to resolve the gauge ambiguity. This yields matrix Grassmannian primitives and tensor Grassmannian primitives that are compact, data-adaptive, and directly decodable by explicit multilinear reconstruction. Theoretically, we prove that the proposed primitive maps are homeomorphisms between low-rank tensors and their corresponding primitive spaces, certifying that the representations are non-degenerate and topologically faithful. Building on these primitives, we propose *Diffusion in Aligned Tensor Space* (DiffATS), a generative framework that trains diffusion models directly on aligned tensor primitives. Across images, videos, and PDE solutions, DiffATS achieves strong unconditional and conditional generation performance while compressing original data by $3.9\times$ to $210\times$, without relying on any pretrained deep compression autoencoders.

cs.LG

Mixing times of Langevin dynamics for spiked matrix models

We investigate the Langevin dynamics for Wigner matrices with a spherical spike, in the regime where the signal-to-noise ratio $\theta$ is large, but order one. For large, order-$1$, signal-to-noise, the (worst-case) mixing time undergoes a sharp transition around the critical inverse temperature $\beta_c(\theta) = \frac{1}{\theta}$. Namely, if $\beta = \alpha/\theta$, and $\alpha<1$ then at large $\theta$ the mixing time is $O(\log N)$, and if $\alpha>1$ it is exponential in $N$. We show that initialized from the uniform-at-random spherical prior, however, the mixing time in the low-temperature $\alpha>1$ regime circumvents the exponential bottleneck and the mixing time is $O(\log N)$. In fact, this fast mixing holds for any initialization that is symmetric with respect to the top eigenvector of the spiked matrix. Using this, we are able to show a low-temperature metastability picture, pinning down the exact exponential rate of the (worst-case initialization) mixing time for low temperatures, showing it is given by the difference of the free energies of the spiked and null models.

math.PR

Scalable Mean-Field Variational Inference via Preconditioned Primal-Dual Optimization

In this work, we investigate the large-scale mean-field variational inference (MFVI) problem from a mini-batch primal-dual perspective. By reformulating MFVI as a constrained finite-sum problem, we develop a novel primal-dual algorithm based on an augmented Lagrangian formulation, termed primal-dual variational inference (PD-VI). PD-VI jointly updates global and local variational parameters in the evidence lower bound in a scalable manner. To further account for heterogeneous loss geometry across different variational parameter blocks, we introduce a block-preconditioned extension, P$^2$D-VI, which adapts the primal-dual updates to the geometry of each parameter block and improves both numerical robustness and practical efficiency. We establish convergence guarantees for both PD-VI and P$^2$D-VI under properly chosen constant step size, without relying on conjugacy assumptions or explicit bounded-variance conditions. In particular, we prove $O(1/T)$ convergence to a stationary point in general settings and linear convergence under strong convexity. Numerical experiments on synthetic data and a real large-scale spatial transcriptomics dataset demonstrate that our methods consistently outperform existing stochastic variational inference approaches in terms of convergence speed and solution quality.

stat.ML

An entropy formula for the Deep Linear Network

We study the Riemannian geometry of the Deep Linear Network (DLN) as a foundation for a thermodynamic description of the learning process. The main tools are the use of group actions to analyze overparametrization and the use of Riemannian submersion from the space of parameters to the space of observables. The foliation of the balanced manifold in the parameter space by group orbits is used to define and compute a Boltzmann entropy. We also show that the Riemannian geometry on the space of observables defined in [2] is obtained by Riemannian submersion of the balanced manifold. The main technical step is an explicit construction of an orthonormal basis for the tangent space of the balanced manifold using the theory of Jacobi matrices.

cs.LG

Riemannian Langevin Monte Carlo schemes for sampling PSD matrices with fixed rank

This paper introduces two explicit schemes to sample matrices from Gibbs distributions on $\mathcal S^{n,p}_+$, the manifold of real positive semi-definite (PSD) matrices of size $n\times n$ and rank $p$. Given an energy function $\mathcal E:\mathcal S^{n,p}_+\to \mathbb{R}$ and certain Riemannian metrics $g$ on $\mathcal S^{n,p}_+$, these schemes rely on an Euler-Maruyama discretization of the Riemannian Langevin equation (RLE) with Brownian motion on the manifold. We present numerical schemes for RLE under two fundamental metrics on $\mathcal S^{n,p}_+$: (a) the metric obtained from the embedding of $\mathcal S^{n,p}_+ \subset \mathbb{R}^{n\times n} $; and (b) the Bures-Wasserstein metric corresponding to quotient geometry. We also provide examples of energy functions with explicit Gibbs distributions that allow numerical validation of these schemes.

math.NA

Siegel Brownian motion

We construct an analogue of Dyson Brownian motion in the Siegel half-space H that we term Siegel Brownian motion. Given \beta in (0,\infty], a stochastic flow for Z_t in H is introduced so that the law of the eigenvalues \lambda_t of the cross ratio matrix R(Z_t,iI_n) is determined by the Ito differential equation corresponds to stochastic gradient ascent of a function S. S turns out to be the log volume of isospectral orbit in H and can be understood as a Boltzmann entropy. In the limit \beta=\infty, the group orbits evolve by motion by minus a half times mean curvature.

math.PR

The Riemannian Langevin equation and conic programs

Diffusion limits provide a framework for the asymptotic analysis of stochastic gradient descent (SGD) schemes used in machine learning. We consider an alternative framework, the Riemannian Langevin equation (RLE), that generalizes the classical paradigm of equilibration in R^n to a Riemannian manifold (M^n, g). The most subtle part of this equation is the description of Brownian motion on (M^n, g). Explicit formulas are presented for some fundamental cones.

math.PR