arXiv ScienceSearch

arXiv subjects

Zhaiming Shen

Publications and source records attributed to Zhaiming Shen.

17 recordsLinked to original sources

Reduced order model for parametric Boltzmann equation and its application to inverse problems

The Boltzmann equation plays an important role in modeling mesoscopic behavior in a wide range of scientific and engineering applications. However, its numerical solution is computationally expensive due to the high dimensionality of the model and the nonlinear nonlocal collision operator, especially for steady-state problems that require iterative solvers. This cost becomes prohibitive for inverse problems, where the induced optimization problem requires repeated forward solves. In this work, we propose a reduced-order model (ROM) for the parametric Boltzmann equation to address this computational challenge. The ROM constructs a low-dimensional approximation space for the parameter-induced solution manifold through a residual-based greedy strategy, and the reduced solution is then obtained via residual minimization over the reduced space, subject to mass conservation. The overall efficiency of the ROM is achieved by exploiting the quadratic structure of the collision operator and a precomputed separable approximation of the collision kernel. The resulting ROM is further applied to a thermally-driven inverse problem for reconstructing collision parameters from the observed macroscopic temperature data. This is accomplished either by directly replacing the PDE constraint with the ROM, leading to a bilevel optimization formulation, or by reformulating the task as a single-level optimization problem through the Karush--Kuhn--Tucker (KKT) conditions. Numerical experiments in both collision-dominated and transport-dominated cases are performed to demonstrate the efficiency and accuracy of the proposed ROM and its effectiveness in inverse problems. In particular, the resulting inverse problem is computationally much more tractable, achieving speedups of several orders of magnitude over that based on the full-order model while maintaining comparable accuracy.

math.NA

Scalable Graph Coreset Selection via Greedy Sampling

Sampling representative nodes from large graphs is fundamental to graph signal processing and network analysis, yet existing methods require access to the full graph Laplacian, making them impractical at scale. We propose a simple and effective column-selective graph sampling algorithm based on a minimum inner product greedy selection rule. At each iteration, the algorithm accesses only a small random subset of Laplacian columns, requiring no eigendecomposition or global graph traversal, making it well-suited for large-scale graphs where the full Laplacian cannot be stored in memory. We analyze the algorithm under the stochastic block model and show that, when the degree distribution is balanced across nodes, the algorithm achieves sampling proportional to cluster size, and that the resulting mean estimate is controlled for band-limited graph signals in the Paley-Wiener space, with the error decaying as inter-cluster connectivity weakens. Numerical experiments on both synthetic and real-world data validate the effectiveness of the proposed method.

math.NA

Structure-Preserving Reduced-Order Modeling via Low-Rank Transport Signatures

Parametrized PDEs with density-valued solutions are often difficult to approximate with classical linear reduced-order models, especially in transport-dominated regimes. We introduce an optimal-transport-based reduced-order modeling that represents each density by the Kantorovich potential transporting a fixed reference density to the target density, and then maps these potentials to transport signatures using a weighted Laplacian associated with the reference measure. This embeds the density-valued solution map in a Hilbert space while preserving control of the induced transport maps and Wasserstein error. We treat the signature map as a continuous matrix indexed by parameters and space, construct a low-rank skeleton decomposition using a maximal-volume criterion, and learn the parameter-to-coefficient map with a neural network for efficient non-intrusive online evaluation. The reconstructed solution is obtained by pushing forward the reference density, so mass preservation is built into the method. We prove a mean-squared Wasserstein error bound separating low-rank approximation, discretization, sampling, and learning errors, and demonstrate the method on a two-dimensional continuity equation, where transport signatures yield substantially lower-rank structure than the original density snapshots.

math.NA

Advancing Local Clustering on Graphs via Compressive Sensing: Semi-supervised and Unsupervised Methods

Local clustering aims to identify specific substructures within a large graph without any additional structural information of the graph. These substructures are typically small compared to the overall graph, enabling the problem to be approached by finding a sparse solution to a linear system associated with the graph Laplacian. In this work, we first propose a method for identifying specific local clusters when very few labeled data are given, which we term semi-supervised local clustering. We then extend this approach to the unsupervised setting when no prior information on labels is available. The proposed methods involve randomly sampling the graph, applying diffusion through local cluster extraction, then examining the overlap among the results to find each cluster. We establish the co-membership conditions for any pair of nodes, and rigorously prove the correctness of our methods. Additionally, we conduct extensive experiments to demonstrate that the proposed methods achieve state of the art results in the low-label rates regime.

cs.LG

Transformers for Learning on Noisy and Task-Level Manifolds: Approximation and Generalization Insights

Transformers serve as the foundational architecture for large language and video generation models, such as GPT, BERT, SORA and their successors. Empirical studies have demonstrated that real-world data and learning tasks exhibit low-dimensional structures, along with some noise or measurement error. The performance of transformers tends to depend on the intrinsic dimension of the data/tasks, though theoretical understandings remain largely unexplored for transformers. This work establishes a theoretical foundation by analyzing the performance of transformers for regression tasks involving noisy input data near a manifold. Specifically, the input data are in a tubular neighborhood of a manifold, while the ground truth function depends on the projection of the noisy data onto this manifold, referred to as the task-level manifold. We prove approximation and generalization errors which crucially depend on the intrinsic dimension of the task-level manifold. Our results demonstrate that transformers can leverage low-complexity structures in learning task even when the input data are perturbed by high-dimensional noise. Our novel proof technique constructs representations of basic arithmetic operations by transformers, which may hold independent interest.

cs.LG

Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods

While in-context learning (ICL) has achieved remarkable success in natural language and vision domains, its theoretical understanding-particularly in the context of structured geometric data-remains unexplored. This paper initiates a theoretical study of ICL for regression of Hölder functions on manifolds. We establish a novel connection between the attention mechanism and classical kernel methods, demonstrating that transformers effectively perform kernel-based prediction at a new query through its interaction with the prompt. This connection is validated by numerical experiments, revealing that the learned query-prompt scores for Hölder functions are highly correlated with the Gaussian kernel. Building on this insight, we derive generalization error bounds in terms of the prompt length and the number of training tasks. When a sufficient number of training tasks are observed, transformers give rise to the minimax regression rate of Hölder functions on manifolds, which scales exponentially with respect to the prompt length with the exponent depending on the intrinsic dimension of the manifold, rather than the ambient space dimension. Our result also characterizes how the generalization error scales with the number of training tasks, shedding light on the complexity of transformers as in-context kernel algorithm learners. Our findings provide foundational insights into the role of geometry in ICL and novels tools to study ICL of nonlinear models.

cs.LG

Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Through the interaction mechanism in attention, we explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, we establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features. Our theory provides finite-sample generalization error bounds in terms of context length and training set size. We numerically validate the theory on synthetic regression tasks.

cs.LG

The Kolmogorov Superposition Theorem can Break the Curse of Dimensionality When Approximating High Dimensional Functions

We explain how to use Kolmogorov Superposition Theorem (KST) to break the curse of dimensionality when approximating a dense class of multivariate continuous functions. We first show that there is a class of functions called Kolmogorov-Lipschitz (KL) continuous in $C([0,1]^d)$ which can be approximated by a special ReLU neural network of two hidden layers with a dimension independent approximation rate $O(1/n)$ with approximation constant increasing quadratically in $d$. The number of parameters used in such neural network approximation equals to $(6d+2)n$. Next we introduce KB-splines by using linear B-splines to replace the outer function and smooth the KB-splines to have the so-called LKB-splines as the basis for approximation. Our numerical evidence shows that the curse of dimensionality is broken in the following sense: When using the standard discrete least squares (DLS) method to approximate a continuous function, there exists a pivotal set of points in $[0,1]^d$ with size at most $O(nd)$ such that the rooted mean squares error (RMSE) from the DLS based on the pivotal set is similar to the RMSE of the DLS based on the original set with size $O(n^d)$. The pivotal point set is chosen by using matrix cross approximation technique and the number of LKB-splines used for approximation is the same as the size of the pivotal data set. Therefore, we do not need too many basis functions nor too many function values to approximate a high dimensional continuous function $f$. Hence, the study in this paper provides an approach for dimension reduction problems.

math.NA

Local Clustering for Lung Cancer Image Classification via Sparse Solution Technique

In this work, we propose to use a local clustering approach based on the sparse solution technique to study the medical image, especially the lung cancer image classification task. We view images as the vertices in a weighted graph and the similarity between a pair of images as the edges in the graph. The vertices within the same cluster can be assumed to share similar features and properties, thus making the applications of graph clustering techniques very useful for image classification. Recently, the approach based on the sparse solutions of linear systems for graph clustering has been found to identify clusters more efficiently than traditional clustering methods such as spectral clustering. We propose to use the two newly developed local clustering methods based on sparse solution of linear system for image classification. In addition, we employ a box spline-based tight-wavelet-framelet method to clean these images and help build a better adjacency matrix before clustering. The performance of our methods is shown to be very effective in classifying images. Our approach is significantly more efficient and either favorable or equally effective compared with other state-of-the-art approaches. Finally, we shall make a remark by pointing out two image deformation methods to build up more artificial image data to increase the number of labeled images.

cs.CV

The Optimal Linear B-splines Approximation via Kolmogorov Superposition Theorem and its Application

We propose a new approach for approximating functions in $C([0,1]^d)$ via Kolmogorov superposition theorem (KST) based on the linear spline interpolation of the outer function in the Kolmogorov representation. We improve the results in \cite{LaiShenKST21} by showing that the optimal rate of approximation based on our proposed approach is $O(\frac{1}{n^2})$, where $n$ denotes the number of knots over $[0,1]$. Furthermore, the approximation constant scales linearly with the dimension $d$. We show that there exists a dense subclass in $C([0,1]^d)$ whose approximation can achieve such optimal rate, and the number of parameters needed in such approximation is at most $O(nd)$. Thus, there is no curse of dimensionality when approximating functions in this subclass. Moreover, for $d\geq 4$, we apply tensor product spline denoising technique to denoise KB-splines and get the smooth LKB-splines. We use LKB-splines as basis to approximate functions for the cases when $d=4$ and $d=6$, which extends the results in \cite{LaiShenKST21}. In addition, we validate via numerical experiments that fewer than $O(nd)$ function values are needed to achieve the rate $O(\frac{1}{n^β})$ for some $β>0$ based on the smoothness of the outer function. Finally, we demonstrate that our approach can be applied to numerically solving partial differential equation such as the Poisson equation with accurate results.

math.NA

Maximal Volume Matrix Cross Approximation for Image Compression and Least Squares Solution

We study the classic matrix cross approximation based on the maximal volume submatrices. Our main results consist of an improvement of the classic estimate for matrix cross approximation and a greedy approach for finding the maximal volume submatrices. More precisely, we present a new proof of the classic estimate of the inequality with an improved constant. Also, we present a family of greedy maximal volume algorithms to improve the computational efficiency of matrix cross approximation. The proposed algorithms are shown to have theoretical guarantees of convergence. Finally, we present two applications: image compression and the least squares approximation of continuous functions. Our numerical results at the end of the paper demonstrate the effective performance of our approach.

math.NA

Graph-based Semi-supervised Local Clustering with Few Labeled Nodes

Local clustering aims at extracting a local structure inside a graph without the necessity of knowing the entire graph structure. As the local structure is usually small in size compared to the entire graph, one can think of it as a compressive sensing problem where the indices of target cluster can be thought as a sparse solution to a linear system. In this paper, we apply this idea based on two pioneering works under the same framework and propose a new semi-supervised local clustering approach using only few labeled nodes. Our approach improves the existing works by making the initial cut to be the entire graph and hence overcomes a major limitation of the existing works, which is the low quality of initial cut. Extensive experimental results on various datasets demonstrate the effectiveness of our approach.

cs.LG

Tree-based Ensemble Learning for Out-of-distribution Detection

Being able to successfully determine whether the testing samples has similar distribution as the training samples is a fundamental question to address before we can safely deploy most of the machine learning models into practice. In this paper, we propose TOOD detection, a simple yet effective tree-based out-of-distribution (TOOD) detection mechanism to determine if a set of unseen samples will have similar distribution as of the training samples. The TOOD detection mechanism is based on computing pairwise hamming distance of testing samples' tree embeddings, which are obtained by fitting a tree-based ensemble model through in-distribution training samples. Our approach is interpretable and robust for its tree-based nature. Furthermore, our approach is efficient, flexible to various machine learning tasks, and can be easily generalized to unsupervised setting. Extensive experiments are conducted to show the proposed method outperforms other state-of-the-art out-of-distribution detection methods in distinguishing the in-distribution from out-of-distribution on various tabular, image, and text data.

cs.LG

A Compressed Sensing Based Least Squares Approach to Semi-supervised Local Cluster Extraction

A least squares semi-supervised local clustering algorithm based on the idea of compressed sensing is proposed to extract clusters from a graph with known adjacency matrix. The algorithm is based on a two-stage approach similar to the one in \cite{LaiMckenzie2020}. However, under a weaker assumption and with less computational complexity than the one in \cite{LaiMckenzie2020}, the algorithm is shown to be able to find a desired cluster with high probability. The ``one cluster at a time" feature of our method distinguishes it from other global clustering methods. Several numerical experiments are conducted on the synthetic data such as stochastic block model and real data such as MNIST, political blogs network, AT\&T and YaleB human faces data sets to demonstrate the effectiveness and efficiency of our algorithm.

cs.LG

A Quasi-Orthogonal Matching Pursuit Algorithm for Compressive Sensing

In this paper, we propose a new orthogonal matching pursuit algorithm called quasi-OMP algorithm which greatly enhances the performance of classical orthogonal matching pursuit (OMP) algorithm, at some cost of computational complexity. We are able to show that under some sufficient conditions of mutual coherence of the sensing matrix, the QOMP Algorithm succeeds in recovering the s-sparse signal vector x within s iterations where a total number of 2s columns are selected under the both noiseless and noisy settings. In addition, we show that for Gaussian sensing matrix, the norm of the residual of each iteration will go to zero linearly depends on the size of the matrix with high probability. The numerical experiments are demonstrated to show the effectiveness of QOMP algorithm in recovering sparse solutions which outperforms the classic OMP and GOMP algorithm.

math.NA

The exponential map is chaotic: An invitation to transcendental dynamics

We present an elementary and conceptual proof that the complex exponential map is "chaotic" when considered as a dynamical system on the complex plane. (This result was conjectured by Fatou in 1926 and first proved by Misiurewicz 55 years later.) The only background required is a first undergraduate course in complex analysis.

math.DS

Group Testing with Pools of Fixed Size

In the classical combinatorial (adaptive) group testing problem, one is given two integers \(d\) and \(n\), where \(0\le d\le n\), and a population of \(n\) items, exactly \(d\) of which are known to be defective. The question is to devise an optimal sequential algorithm that, at each step, tests a subset of the population and determines whether such subset is contaminated (i.e. contains defective items) or otherwise. The problem is solved only when the \(d\) defective items are identified. The minimum number of steps that an optimal sequential algorithm takes in general (i.e. in the worst case) to solve the problem is denoted by \(M(d, n)\). The computation of \(M(d, n)\) appears to be very difficult and a general formula is known only for \(d = 1\). We consider here a variant of the original problem, where the size of the subsets to be tested is restricted to be a fixed positive integer \(k\). The corresponding minimum number of tests by a sequential optimal algorithm is denoted by \(M^{\lbrack k\rbrack}(d, n)\). In this paper we start the investigation of the function \(M^{\lbrack k\rbrack}(d, n)\).

math.CO