arXiv ScienceSearch

arXiv subjects

ShengLi Tzeng

Publications and source records attributed to ShengLi Tzeng.

12 recordsLinked to original sources

Structured Covariate-Informed Empirical Orthogonal Functions for Spatio-Temporal Environmental Fields

Low-rank representations such as empirical orthogonal function (EOF) decompositions are widely used for analyzing large spatio-temporal environmental fields. However, conventional EOF identifies latent modes solely from covariance structure and does not utilize observed environmental covariates, limiting its ability to incorporate external information into low-rank representations. This study introduces Structured Covariate-Informed EOF (SCIEOF), a covariate-informed extension of EOF that bridges low-rank dimension reduction and prediction-oriented spatio-temporal modeling. SCIEOF embeds spatial and temporal covariates into the latent bases while incorporating spatio-temporal covariates through an additive component, yielding low-rank representations with latent modes informed by observed covariates. Estimation procedures are developed and evaluated through simulation studies and an application to global near-surface air temperature from the MERRA-2 reanalysis. Simulation studies demonstrate that incorporating informative covariates improves latent structure recovery and predictive accuracy, particularly when the spatial basis is appropriately specified. The advantage is more pronounced at moderate-to-large sample sizes, while methods with stronger structural assumptions remain competitive when data are limited. In the MERRA-2 application, SCIEOF achieves competitive or improved predictive performance relative to commonly used methods while providing a compact and physically interpretable low-rank representation. Overall, SCIEOF provides a flexible and computationally scalable framework for integrating structural covariate information into low-rank spatio-temporal representations, extending EOF toward predictive environmental modeling.

stat.ME

Sturm-Liouville-Type Parity and Oscillation of a Cubic Spline Eigenbasis

We study the eigen-structure of the penalty matrix arising from cubic smoothing splines on equally spaced knots. Using purely matrix-theoretic arguments, we show that its positive eigenvalues are simple, that the associated eigenvectors alternate between even and odd, and that the eigenvector for the $k$th largest eigenvalue has exactly $k+1$ sign changes. The approach provides a direct and transparent alternative to existing variational proofs of the oscillation property. These results show that equally spaced knots support a spline basis with both a parity structure and an oscillation pattern.

math.ST

Calibrated Predictive Distributions from Sample-Based Generators

Conditional generative models, including diffusion models and ensemble forecasters, often produce predictive samples without a tractable likelihood representation. Such sample-based predictive distributions can be systematically biased and poorly calibrated. We propose bias-corrected conformal probability integral transform (PIT) calibration, a split-sample post-processing framework that outputs a calibrated predictive distribution rather than a single fixed-level prediction interval. The method first estimates an affine location-scale correction on a held-out bias split, then calibrates randomized PIT values using a conformal calibrator. The resulting predictive law is represented as a weighted empirical distribution on the generator order statistics, enabling the direct computation of threshold-coherent exceedance probabilities, arbitrary quantiles, highest-density intervals, expected tail losses, and calibrated resamples. In contrast, standard conformal prediction primarily provides fixed-level prediction sets or threshold decisions and does not directly estimate predictive probabilities or high-density regions. We establish finite-sample calibration in probability under exchangeability and show how an optional split-conformal wrapper based on a PIT-centrality score gives nested prediction intervals with finite-sample marginal coverage at user-specified levels. Simulation studies with controlled misspecification and a WeatherBench-2 precipitation-forecasting application demonstrate substantial improvements in probabilistic calibration and downstream distributional summaries relative to uncalibrated sample-based forecasts and interval-only conformal baselines.

stat.ME

cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data

High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables. Existing methods either scale poorly, rely heavily on low-dimensional displays detached from the original data matrix, or prioritize predictive accuracy over interpretability. To address this gap, we introduce categorical Generalized Association Plots (cGAP), a visualization framework for nominal, ordinal, and binary data that preserves the original data matrix while augmenting it with interpretable geometric structure. cGAP uses Homogeneity Analysis (HOMALS) to embed subjects and category levels in a three-dimensional Euclidean space and maps the embedding to red-green-blue coordinates so that similar patterns receive similar colors. The framework integrates three coordinated views: a HOMALS-guided heatmap of the raw data matrix, a subject proximity matrix, and a variable proximity matrix. Seriation algorithms are then used to reorder rows and columns to reveal coherent clusters, outliers, and local-to-global structure. We also derive barycentric traceability, projection-distortion, and contrast-preservation properties that clarify how embedding geometry is transferred to the display. We demonstrate the versatility of cGAP through applications to student-animal classification data, mammalian dentition profiles, mushroom records from the UCI Machine Learning Repository, and the Clusters of Orthologous Genes database. These examples show that cGAP supports transparent exploratory analysis by maintaining traceability between derived visual structure and the original categorical observations. cGAP provides a full-matrix, heatmap-based visualization environment for investigating complex categorical datasets across scientific domains.

stat.ML

Assessing Spatial Stationarity and Segmenting Spatial Processes into Stationary Components

In this research, we propose a novel technique for visualizing nonstationarity in geostatistics, particularly when confronted with a single realization of data at irregularly spaced locations. Our method hinges on formulating a statistic that tracks a stable microergodic parameter of the exponential covariance function, allowing us to address the intricate challenges of nonstationary processes that lack repeated measurements. We implement the fused lasso technique to elucidate nonstationary patterns at various resolutions. For prediction purposes, we segment the spatial domain into stationary sub-regions via Voronoi tessellations. Additionally, we devise a robust test for stationarity based on contrasting the sample means of our proposed statistics between two selected Voronoi subregions. The effectiveness of our method is demonstrated through simulation studies and its application to a precipitation dataset in Colorado.

stat.ME

The R Package HCV for Hierarchical Clustering from Vertex-links

The HCV package implements the hierarchical clustering for spatial data. It requires clustering results not only homogeneous in non-geographical features among samples but also geographically close to each other within a cluster. We modified typically used hierarchical agglomerative clustering algorithms to introduce the spatial homogeneity, by considering geographical locations as vertices and converting spatial adjacency into whether a shared edge exists between a pair of vertices. The main function HCV obeying constraints of the vertex links automatically enforces the spatial contiguity property at each step of iterations. In addition, two methods to find an appropriate number of clusters and to report cluster members are also provided.

stat.CO

Interpretable, predictive spatio-temporal models via enhanced Pairwise Directions Estimation

This article concerns the predictive modeling for spatio-temporal data as well as model interpretation using data information in space and time. We develop a novel approach based on supervised dimension reduction for such data in order to capture nonlinear mean structures without requiring a prespecified parametric model. In addition to prediction as a common interest, this approach emphasizes the exploration of geometric information from the data. The method of Pairwise Directions Estimation (PDE; Lue, 2019) is implemented in our approach as a data-driven function searching for spatial patterns and temporal trends. The benefit of using geometric information from the method of PDE is highlighted, which aids effectively in exploring data structures. We further enhance PDE, referring to it as PDE+, by incorporating kriging to estimate the random effects not explained in the mean functions. Our proposal can not only increase prediction accuracy, but also improve the interpretation for modeling. Two simulation examples are conducted and comparisons are made with four existing methods. The results demonstrate that the proposed PDE+ method is very useful for exploring and interpreting the patterns and trends for spatio-temporal data. Illustrative applications to two real datasets are also presented.

stat.ME

Spatially Adaptive Calibrations of AirBox PM$_{2.5}$ Data

Two networks are available to monitor PM$_{2.5}$ in Taiwan, including the Taiwan Air Quality Monitoring Network (TAQMN) and the AirBox network. The TAQMN, managed by Taiwan's Environmental Protection Administration (EPA), provides high-quality PM$_{2.5}$ measurements at $77$ monitoring stations. More recently, the AirBox network was launched, consisting of low-cost, small internet-of-things (IoT) microsensors (i.e., AirBoxes) at thousands of locations. While the AirBox network provides broad spatial coverage, its measurements are not reliable and require calibrations. However, applying a universal calibration procedure to all AirBoxes does not work well because the calibration curves vary with several factors, including the chemical compositions of PM$_{2.5}$, which are not homogeneous in space. Therefore, different calibrations are needed at different locations with different local environments. Unfortunately, most AirBoxes are not close to EPA stations, making the calibration task challenging. In this article, we propose a spatial model with spatially varying coefficients to account for heteroscedasticity in the data. Our method gives adaptive calibrations of AirBoxes according to their local conditions and provides accurate PM$_{2.5}$ concentrations at any location in Taiwan, incorporating two types of measurements. In addition, the proposed method automatically calibrates measurements from a new AirBox once it is added to the network. We illustrate our approach using hourly PM$_{2.5}$ data in the year 2020. After the calibration, the results show that the PM$_{2.5}$ prediction improves about 37% to 67% in root mean-squared prediction error for matching EPA data. In particular, once the calibration curves are established, we can obtain reliable PM$_{2.5}$ values at any location in Taiwan, even if we ignore EPA data.

stat.ME

Spatial Process Decomposition for Quantitative Imaging Biomarkers Using Multiple Images of Varying Shapes

Quantitative imaging biomarkers (QIB) are extracted from medical images in radiomics for a variety of purposes including noninvasive disease detection, cancer monitoring, and precision medicine. The existing methods for QIB extraction tend to be ad-hoc and not reproducible. In this paper, a general and flexible statistical approach is proposed for handling up to three-dimensional medical images in an objective and principled way. In particular, a model-based spatial process decomposition is developed where the random weights are unique to individual patients for component functions common across patients. Model fitting and selection are based on maximum likelihood, while feature extractions are via optimal prediction of the underlying true image. A simulation study evaluates the properties of the proposed methodology and for illustration, a cancer image data set is analyzed and QIBs are extracted in association with a clinical endpoint.

stat.ME

Empirical Likelihood Based Summary ROC Curve for Meta-Analysis of Diagnostic Studies

Objectives: This study provides an effective model selection method based on the empirical likelihood approach for constructing summary receiver operating characteristic (sROC) curves from meta-analyses of diagnostic studies. Methods: We considered models from combinations of family indices and specific pairs of transformations, which cover several widely used methods for bivariate summary of sensitivity and specificity. Then a final model was selected using the proposed empirical likelihood method. Simulation scenarios were conducted based on different number of studies and different population distributions for the disease and non-disease cases. The performance of our proposal and other model selection criteria was also compared. Results: Although parametric likelihood-based methods are often applied in practice due to its asymptotic property, they fail to consistently choose appropriate models for summary under the limited number of studies. For these situations, our proposed method almost always performs better. Conclusion: When the number of studies is as small as 10 or 5, we recommend choosing a summary model via the proposed empirical likelihood method.

stat.ME

Distance for Functional Data Clustering Based on Smoothing Parameter Commutation

We propose a novel method to determine the dissimilarity between subjects for functional data clustering. Spline smoothing or interpolation is common to deal with data of such type. Instead of estimating the best-representing curve for each subject as fixed during clustering, we measure the dissimilarity between subjects based on varying curve estimates with commutation of smoothing parameters pair-by-pair (of subjects). The intuitions are that smoothing parameters of smoothing splines reflect inverse signal-to-noise ratios and that applying an identical smoothing parameter the smoothed curves for two similar subjects are expected to be close. The effectiveness of our proposal is shown through simulations comparing to other dissimilarity measures. It also has several pragmatic advantages. First, missing values or irregular time points can be handled directly, thanks to the nature of smoothing splines. Second, conventional clustering method based on dissimilarity can be employed straightforward, and the dissimilarity also serves as a useful tool for outlier detection. Third, the implementation is almost handy since subroutines for smoothing splines and numerical integration are widely available. Fourth, the computational complexity does not increase and is parallel with that in calculating Euclidean distance between curves estimated by smoothing splines.

stat.ME

Multi-Resolution Spatial Random-Effects Models for Irregularly Spaced Data

The spatial random-effects model is flexible in modeling spatial covariance functions, and is computationally efficient for spatial prediction via fixed rank kriging. However, the success of this model depends on an appropriate set of basis functions. In this research, we propose a class of basis functions extracted from thin-plate splines. These functions are ordered in terms of their degrees of smoothness with a higher-order function corresponding to larger-scale features and a lower-order one corresponding to smaller-scale details, leading to a parsimonious representation for a nonstationary spatial covariance function. Consequently, only a small to moderate number of functions are needed in a spatial random-effects model. The proposed class of basis functions has several advantages over commonly used ones. First, we do not need to concern about the allocation of the basis functions, but simply select the total number of functions corresponding to a resolution. Second, only a small number of basis functions is usually required, which facilitates computation. Third, estimation variability of model parameters can be considerably reduced, and hence more precise covariance function estimates can be obtained. Fourth, the proposed basis functions depend only on the data locations but not the measurements taken at those locations, and are applicable regardless of whether the data locations are sparse or irregularly spaced. In addition, we derive a simple close-form expression for the maximum likelihood estimates of model parameters in the spatial random-effects model. Some numerical examples are provided to demonstrate the effectiveness of the proposed method.

stat.ME