arXiv ScienceSearch

arXiv subjects

Lek-Heng Lim

Publications and source records attributed to Lek-Heng Lim.

At least 19 recordsLinked to original sources

Singular value decomposition of unbounded operators

The singular value decomposition has been established for matrices, Hilbert--Schmidt operators, trace-class operator, compact operators, and bounded operators, but surprisingly not for unbounded operators. Unfortunately, most interesting operators in applied math are unbounded, as any operators involving some form of derivatives --- gradient, exterior derivatives, Laplacians, Fourier and other transforms of derivatives, Hamiltonians, etc. --- are likely unbounded. In this article, we fill in this last missing piece by establishing the existence of singular value decompositions for unbounded operators in three natural forms: multiplication-operator, direct-integral, and operator-valued-measure. We show it inherits classical properties of finite-dimensional singular value decomposition including approximation results, relationships with fundamental subspaces, and the Moore--Penrose inverse. This discovery opens the door to the singular value decompositions of a myriad of well-known unbounded operators in mathematics, physics, statistics, and finnance --- gradients on Euclidean spaces and manifolds, Petrov--Galerkin method, finite-difference operators, Hilbert--Schmidt operators, Hilbert complexes, supersymmetric quantum mechanics, Sturm--Liouville theory, nonparametric density estimation, and the Black--Scholes equation. The resulting decompositions reveal a number of novel insights, among many others: bosonic and fermionic states in supersymmetric quantum mechanics arise as left and right singular vectors of generalized ladder operators; the Riesz transform appears as the left singular operator of the Euclidean gradient; and the Hodge decomposition follows directly from the singular value decompositions of the exterior derivatives.

math.SP

Pierce-Birkhoff conjecture is false

We provide a counterexample to the Pierce-Birkhoff conjecture: a continuous function that is piecewise quadratic on a finite collection of semialgebraic sets that partition $\mathbb{R}^n$ but that cannot be expressed as a finite lattice combination of polynomials. The counterexample was found with the assistance of our multi-agent, multi-model setup that chains together GPT 5.6, GPT 6, Claude Opus 5, and Claude Fable 5.1.

math.AG

The Grassmannian of indefinite subspaces

The homogeneous space $\operatorname{O}_{m,n}(\mathbb{R})/(\operatorname{O}_{p,q}(\mathbb{R}) \times \operatorname{O}_{m-p,n-q}(\mathbb{R}))$ is an object that has received scant attention, with just two brief mentions in existing literature, and christened the indefinite Grassmannian in one of them. In this article, we develop some of its basic properties, building it from ground up. We will see that, aside from its homogeneous space description, the indefinite Grassmannian may be characterized in several other ways: set-theoretically, it is the manifold of indefinite $(p+q)$-dimensional subspaces in $(m +n)$-dimensional space; it is an adjoint orbit of a Lie group; a semialgebraic smooth manifold of matrices; a base space of a principal bundle whose total space is the indefinite Stiefel manifold, a natural corresponding notion. It may also be naturally equipped with various structures, turning the indefinite Grassmannian into a pseudo-Riemannian manifold; a Einstein manifold; a symplectic manifold; and a pseudo-Kähler manifold (last two only over $\mathbb{C}$). As a centerpiece of this article, we establish two attributes of the indefinite Grassmannian in relation to the standard Grassmannian: (i) any Grassmannian has a Whitney stratification whose highest dimensional strata are indefinite Grassmannians; (ii) the indefinite Grassmannian is a strong deformation retract of a product of two Grassmannians, thereby allowing us to completely ascertain the topology of the former. We will also determine some of the indefinite Grassmannian's features that are within reach --- both geometric (Riemann, Ricci, sectional, and scalar curvatures; second fundamental form) and topological (cohomology ring, homotopy groups, characteristic classes).

math.DG

Low-dimensional topology of deep neural networks

We study layered models, including feedforward networks, ResNets, and transformers, by limiting each layer to a width of $d = 3$, i.e., $\mathbb{R}^3$ as representation space. This allows us to track how a neural network changes low-dimensional topological invariants through its layers. Just about any topological structure may be simplified or even trivialized by simply increasing dimension; e.g., any knot is equivalent to an unknot in $\mathbb{R}^4$. By restricting to $\mathbb{R}^3$, we not only isolate the effects of activation and depth from that of width, we work in a space that lends itself to easy visualization. We focus on linking number here, deferring other invariants like link groups, Milnor's $\barμ$-invariants, knot types, ambient cobordisms, to a sequel. We provide full proofs and empirical experiments to justify the following insights: When measured by their power to effect changes in linking numbers, the layer-skipping feature in ResNets is as powerful as the attention mechanism in transformers; both ResNets and transformers are strictly more powerful than feedforward neural networks with monotonic activations, which are in turn more powerful than invertible and flow-based models; but replacing monotonic activation with a nonmonotonic one elevates a feedforward network into the same expressivity class as ResNets and transformers. These results suggest that low-dimensional topology can be a useful tool to guide designs of AI architectures. We also generalize our results from $d = 3$ to arbitrary $d > 3$.

cs.LG

Generalized matrix nearness problems II

Given a matrix $A$, a matrix nearness problem seeks an $X$ that most closely approximates $A$ in the sense of minimizing $\lVert A - X\rVert$ under a variety of constraints on $X$. A generalized matrix nearness problem seeks the same but with three given matrices $A,B,C$ and $\lVert A - BXC\rVert$ in place of $\lVert A - X\rVert$. We extend previous studies of the latter problem in three directions: incorporating an affine term, replacing matrix product by Kronecker product in various manners, and generalizing Frobenius norm to any orthogonally invariant norm. We will solve several of these in closed form. For the rest, we develop an iterative algorithm that works for any Schatten norm, proving that it converges to a global minimizer regardless of the initial point. In addition, the algorithm relies purely on numerical linear algebra, and notably does not compute any explicit gradients or subgradients. Along the way, we will also show that there is no Mirsky-type theorem for rank constrained generalized matrix nearness problems.

math.NA

Linear representations of manifolds

A finite-dimensional linear representation of a group or an algebra may be regarded as a map into a space of matrices, endowing abstract elements with coordinates, and encoding algebraic operations as matrix products. With this in mind, we define a linear representation of a $\mathsf{G}$-manifold $\mathcal{M} $ as a map into a space of matrices, representing points as matrices and the $\mathsf{G}$-action as matrix products. We show that this generalizes group representations to any $\mathsf{G}$-manifold that may not have a group structure, with homogeneous spaces $\mathsf{G}/\mathsf{H}$ an important special case; and in this case it also generalizes Cartan embeddings of symmetric spaces to more general $\mathsf{G}/\mathsf{H}$. To demonstrate the utility of such manifold representations, we use them to provide effective bounds for Mostow-Palais $\mathsf{G}$-equivariant embeddings of $\mathsf{G}$-manifolds into $\mathsf{G}$-modules $\mathbb{V}$. Unlike Whitney and Nash embeddings, Mostow-Palais embeddings have no known effective bounds; before our work, it was only known that $\dim \mathbb{V} < \infty$ if $\mathsf{G}$ is compact. We will give explicit values for $\dim \mathbb{V}$ and show that our bounds are sharp. Furthermore, our method is constructive, giving explicit expressions for these minimal-dimensional Mostow-Palais embeddings.

math.DG

Eigen, singular, cosine-sine, and Autonne--Takagi vectors distributions of random matrix ensembles

We show that some of the best-known matrix decompositions of some of the best-known random matrix ensembles give us the unique $G$-invariant uniform distributions on some of the best-known manifolds. The eigenvectors distributions of the Gaussian, Laguerre, and Jacobi ensembles are all given by the uniform distribution on the complete flag manifold. The singular vectors distributions of Ginibre ensembles are given by the uniform distribution on a product of the complete flag manifold with a Stiefel manifold. Circular ensembles split into two types: The cosine-sine vectors distributions of circular real, unitary, and quaternionic ensembles are given by the uniform distributions on products of a (partial) flag manifold with copies of the orthogonal, unitary, or compact symplectic groups. The Autonne--Takagi vectors distributions of circular orthogonal, Lagrangian, and symplectic ensembles are given by the uniform distributions on Lagrangian Grassmannians.

math.PR

Stiefel optimization is NP-hard

We show that linearly constrained linear optimization over a Stiefel or Grassmann manifold is NP-hard in general. We show that the same is true for unconstrained quadratic optimization over a Stiefel manifold. We will show that unless $\mathrm{P}=\mathrm{NP}$, these optimization problems over a Stiefel manifold do not have $\mathrm{FPTAS}$. As an aside we extend our results to flag manifolds. Combined with earlier findings, this shows that manifold optimization is a difficult endeavor -- even the simplest problems like LP and unconstrained QP are already NP-hard on the most common manifolds.

math.OC

Special orthogonal, special unitary, and symplectic groups as products of Grassmannians

We describe a curious structure of the special orthogonal, special unitary, and symplectic groups that has not been observed, namely, they can be expressed as matrix products of their corresponding Grassmannians realized as involution matrices. We will show that $\operatorname{SO}(n)$ is a product of two real Grassmannians, $\operatorname{SU}(n)$ a product of four complex Grassmannians, and $\operatorname{Sp}(2n, \mathbb{R})$ or $\operatorname{Sp}(2n, \mathbb{C})$ a product of four symplectic Grassmannians over $\mathbb{R}$ or $\mathbb{C}$ respectively.

math.RT

Degree of the Grassmannian as an affine variety

The degree of the Grassmannian with respect to the Plücker embedding is well-known. However, the Plücker embedding, while ubiquitous in pure mathematics, is almost never used in applied mathematics. In applied mathematics, the Grassmannian is usually embedded as projection matrices $\operatorname{Gr}(k,\mathbb{R}^n) \cong \{P \in \mathbb{R}^{n \times n} : P^{\scriptscriptstyle\mathsf{T}} = P = P^2,\; \operatorname{tr}(P) = k\}$ or as involution matrices $\operatorname{Gr}(k,\mathbb{R}^n) \cong \{X \in \mathbb{R}^{n \times n} : X^{\scriptscriptstyle\mathsf{T}} = X,\; X^2 = I,\; \operatorname{tr}(X)=2k - n\}$. We will determine an explicit expression for the degree of the Grassmannian with respect to these embeddings. In so doing, we resolved a conjecture of Devriendt, Friedman, Reinke, and Sturmfels about the degree of $\operatorname{Gr}(2, \mathbb{R}^n)$ and in fact generalized it to $\operatorname{Gr}(k, \mathbb{R}^n)$. We also proved a set theoretic variant of another conjecture of theirs about the limit of $\operatorname{Gr}(k,\mathbb{R}^n)$ in the sense of Gröbner degneration.

math.AG

Simple matrix expressions for the curvatures of Grassmannian

We show that modeling a Grassmannian as symmetric orthogonal matrices $\operatorname{Gr}(k,\mathbb{R}^n) \cong\{Q \in \mathbb{R}^{n \times n} : Q^{\scriptscriptstyle\mathsf{T}} Q = I, \; Q^{\scriptscriptstyle\mathsf{T}} = Q,\; \operatorname{tr}(Q)=2k - n\}$ yields exceedingly simple matrix formulas for various curvatures and curvature-related quantities, both intrinsic and extrinsic. These include Riemann, Ricci, Jacobi, sectional, scalar, mean, principal, and Gaussian curvatures; Schouten, Weyl, Cotton, Bach, Plebański, cocurvature, nonmetricity, and torsion tensors; first, second, and third fundamental forms; Gauss and Weingarten maps; and upper and lower delta invariants. We will derive explicit, simple expressions for the aforementioned quantities in terms of standard matrix operations that are stably computable with numerical linear algebra. Many of these aforementioned quantities have never before been presented for the Grassmannian.

math.DG

Pierce-Birkhoff conjecture is true for splines

We prove the Pierce--Birkhoff conjecture for splines, i.e., continuous piecewise polynomials of degree $d$ in $n$ variables on a hyperplane partition of $\mathbb{R}^n$, can be written as a finite lattice combination of polynomials. We will provide a purely existential proof, followed by a more in-depth analysis that yields effective bounds.

math.AG

Geometric Programming for 3D Circuits

With the soaring demand for high-performing integrated circuits, 3D integrated circuits (ICs) have emerged as a promising alternative to traditional planar structures. Unlike existing 3D ICs that stack 2D layers, a full 3D IC features cubic circuit elements unrestricted by layers, offering greater design freedom. Design problems such as floorplanning, transistor sizing, and interconnect sizing are highly complex due to the 3D nature of the circuits and unavoidably require systematic approaches. We introduce geometric programming to solve these design optimization problems systematically and efficiently.

math.OC

Glivenko-Cantelli for $f$-divergence

We extend the celebrated Glivenko-Cantelli theorem, sometimes called the fundamental theorem of statistics, from its standard setting of total variation distance to all $f$-divergences. A key obstacle in this endeavor is to define $f$-divergence on a subcollection of a $σ$-algebra that forms a $π$-system but not a $σ$-subalgebra. This is a side contribution of our work. We will show that this notion of $f$-divergence on the $π$-system of rays preserves nearly all known properties of standard $f$-divergence, yields a novel integral representation of the Kolmogorov-Smirnov distance, and has a Glivenko-Cantelli theorem. We will also discuss the prospects of a Vapnik-Chervonenkis theory for $f$-divergence.

math.ST

Euclidean distance degree in manifold optimization

We determine the Euclidean distance degrees of the three most common manifolds arising in manifold optimization: flag, Grassmann, and Stiefel manifolds. For the Grassmannian, we will also determine the Euclidean distance degree of an important class of Schubert varieties that often appear in applications. Our technique goes further than furnishing the value of the Euclidean distance degree; it will also yield closed-form expressions for all stationary points of the Euclidean distance function in each instance. We will discuss the implications of these results on the tractability of manifold optimization problems.

math.OC

Summing divergent matrix series

We extend several celebrated methods in classical analysis for summing series of complex numbers to series of complex matrices. These include the summation methods of Abel, Borel, Cesáro, Euler, Lambert, Nörlund, and Mittag-Leffler, which are frequently used to sum scalar series that are divergent in the conventional sense. One feature of our matrix extensions is that they are fully noncommutative generalizations of their scalar counterparts -- not only is the scalar series replaced by a matrix series, positive weights are replaced by positive definite matrix weights, order on $\mathbb{R}$ replaced by Loewner order, exponential function replaced by matrix exponential function, etc. We will establish the regularity of our matrix summation methods, i.e., when applied to a matrix series convergent in the conventional sense, we obtain the same value for the sum. Our second goal is to provide numerical algorithms that work in conjunction with these summation methods. We discuss how the block and mixed-block summation algorithms, the Kahan compensated summation algorithm, may be applied to matrix sums with similar roundoff error bounds. These summation methods and algorithms apply not only to power or Taylor series of matrices but to any general matrix series including matrix Fourier and Dirichlet series. We will demonstrate the utility of these summation methods: establishing a Fejér's theorem and alleviating the Gibbs phenomenon for matrix Fourier series; extending the domains of matrix functions and accurately evaluating them; enhancing the matrix Padé approximation and Schur--Parlett algorithms; and more.

math.NA

Grassmannian optimization is NP-hard

We show that unconstrained quadratic optimization over a Grassmannian $\operatorname{Gr}(k,n)$ is NP-hard. Our results cover all scenarios: (i) when $k$ and $n$ are both allowed to grow; (ii) when $k$ is arbitrary but fixed; (iii) when $k$ is fixed at its lowest possible value $1$. We then deduce the NP-hardness of unconstrained cubic optimization over the Stiefel manifold $\operatorname{V}(k,n)$ and the orthogonal group $\operatorname{O}(n)$. As an addendum we demonstrate the NP-hardness of unconstrained quadratic optimization over the Cartan manifold, i.e., the positive definite cone $\mathbb{S}^n_{\scriptscriptstyle++}$ regarded as a Riemannian manifold, another popular example in manifold optimization. We will also establish the nonexistence of $\mathrm{FPTAS}$ in all cases.

math.OC

Attention is a smoothed cubic spline

We highlight a perhaps important but hitherto unobserved insight: The attention module in a transformer is a smoothed cubic spline. Viewed in this manner, this mysterious but critical component of a transformer becomes a natural development of an old notion deeply entrenched in classical approximation theory. More precisely, we show that with ReLU-activation, attention, masked attention, encoder-decoder attention are all cubic splines. As every component in a transformer is constructed out of compositions of various attention modules (= cubic splines) and feed forward neural networks (= linear splines), all its components -- encoder, decoder, and encoder-decoder blocks; multilayered encoders and decoders; the transformer itself -- are cubic or higher-order splines. If we assume the Pierce-Birkhoff conjecture, then the converse also holds, i.e., every spline is a ReLU-activated encoder. Since a spline is generally just $C^2$, one way to obtain a smoothed $C^\infty$-version is by replacing ReLU with a smooth activation; and if this activation is chosen to be SoftMax, we recover the original transformer as proposed by Vaswani et al. This insight sheds light on the nature of the transformer by casting it entirely in terms of splines, one of the best known and thoroughly understood objects in applied mathematics.

cs.AI