arXiv ScienceSearch

arXiv subjects

Chul Moon

Publications and source records attributed to Chul Moon.

12 recordsLinked to original sources

A survival analysis of glioma patients using topological features and locations of tumors

Tumor shape plays a critical role in influencing both growth and metastasis. We introduce a novel topological radiomic feature derived from persistent homology to characterize tumor shape, focusing on its association with time-to-event outcomes in gliomas. These features effectively capture diverse tumor shape patterns that are not represented by conventional radiomic measures. To incorporate these features into survival analysis, we employ a functional Cox regression model in which the topological features are represented in a functional space. We further include interaction terms between shape features and tumor location to capture lobe-specific effects. This approach enables interpretable assessment of how tumor morphology relates to survival risk. We evaluate the proposed method in two case studies using radiomic images of high-grade and low-grade gliomas. The findings suggest that the topological features serve as strong predictors of survival prognosis, remaining significant after adjusting for clinical variables, and provide additional clinically meaningful insights into tumor behavior.

stat.ME

Spatial Analysis for AI-segmented Histopathology Images: Methods and Implementation

Quantitative characterization of cellular spatial organization is critical for understanding tumor progression and immune response. Recent advances in artificial intelligence (AI) enable large-scale segmentation and classification of nuclei from digitized histopathology slides, producing massive point pattern and marked point pattern data. However, accessible and standardized tools for downstream spatial statistical analysis remain limited. We present SASHIMI (Spatial Analysis for Segmented Histopathology Images using Machine Intelligence), a browser-based platform for real-time spatial analysis of AI-segmented histopathology images. Rather than proposing new spatial methods, SASHIMI systematically organizes and operationalizes 27 widely used spatial summary statistics, areal indices, and topological features within a unified computational framework. The platform computes mathematically grounded descriptors including K-, L-, G-, F-, and J-functions, pair correlation and mark connection functions, spatial autocorrelation measures, similarity indices, and persistent homology-based topological summaries. Outputs include both functional curves and scalar feature tables suitable for downstream statistical modeling. We illustrate the framework using two cancer cohorts: oral potentially malignant disorders and non-small-cell lung cancer. Across datasets, cross-type spatial interactions and topological descriptors show associations with patient survival, demonstrating that complementary spatial features capture distinct aspects of tumor microenvironment architecture. SASHIMI provides an accessible, reproducible platform for single-cell-level spatial profiling of tumor tissue, enabling interactive visualization and standardized feature extraction without requiring programming expertise.

stat.AP

generalRSS: Sampling and Inference for Balanced and Unbalanced Ranked Set Sampling in R

Ranked set sampling (RSS) is a stratified sampling method that improves efficiency over simple random sampling (SRS) by utilizing auxiliary information for ranking and stratification. While balanced RSS (BRSS) assumes equal allocation across strata, unbalanced RSS (URSS) allows unequal allocation, making it particularly effective for skewed distributions. The generalRSS package provides extensive tools for both BRSS and URSS, addressing limitations in existing RSS software that primarily focus on balanced designs. It supports RSS data generation, efficient sample allocation strategies for URSS, and statistical inference for both balanced and unbalanced designs. This paper presents the RSS methodology and demonstrates the utility of generalRSS through two medical data applications: a one-sample mean inference and a two-sample area under the curve (AUC) comparison using NHANES datasets. These applications illustrate the practical implementation of URSS and show how generalRSS facilitates ranked set sampling and inference in real-world data analysis.

stat.ME

Enhancing Empathic Accuracy: Penalized Functional Alignment Method to Correct Temporal Misalignment in Real-time Emotional Perception

Empathic accuracy (EA) is the ability to accurately understand another person\textquotesingle s thoughts and feelings, which is crucial for social and psychological interactions. Traditionally, EA is assessed by comparing a perceiver\textquotesingle s moment-to-moment ratings of a target\textquotesingle s emotional state with the target\textquotesingle s own self-reported ratings at corresponding time points. However, misalignments between these two sequences are common due to the complexity of emotional interpretation and individual differences in behavioral responses. Conventional methods often ignore or oversimplify these misalignments, for instance, by assuming a fixed time lag, which can introduce bias into EA estimates. To address this, we propose a novel alignment approach that captures a wide range of misalignment patterns. Our method leverages the square-root velocity framework to decompose emotional rating trajectories into amplitude and phase components. To ensure realistic alignment, we introduce a regularization constraint that limits temporal shifts to ranges consistent with human perceptual capabilities. This alignment is efficiently implemented using a constrained dynamic programming algorithm. We validate our method through simulations and real-world applications involving video and music datasets, demonstrating its superior performance over traditional techniques.

stat.AP

Discovering Clinically Meaningful Shape Features for the Analysis of Tumor Pathology Images

With the advanced imaging technology, digital pathology imaging of tumor tissue slides is becoming a routine clinical procedure for cancer diagnosis. This process produces massive imaging data that capture histological details in high resolution. Recent developments in deep-learning methods have enabled us to automatically detect and characterize the tumor regions in pathology images at large scale. From each identified tumor region, we extracted 30 well-defined descriptors that quantify its shape, geometry, and topology. We demonstrated how those descriptor features were associated with patient survival outcome in lung adenocarcinoma patients from the National Lung Screening Trial (n=143). Besides, a descriptor-based prognostic model was developed and validated in an independent patient cohort from The Cancer Genome Atlas Program program (n=318). This study proposes new insights into the relationship between tumor shape, geometrical, and topological features and patient prognosis. We provide software in the form of R code on GitHub: https://github.com/estfernandez/Slide_Image_Segmentation_and_Extraction.

stat.AP

Using Persistent Homology Topological Features to Characterize Medical Images: Case Studies on Lung and Brain Cancers

Tumor shape is a key factor that affects tumor growth and metastasis. This paper proposes a topological feature computed by persistent homology to characterize tumor progression from digital pathology and radiology images and examines its effect on the time-to-event data. The proposed topological features are invariant to scale-preserving transformation and can summarize various tumor shape patterns. The topological features are represented in functional space and used as functional predictors in a functional Cox proportional hazards model. The proposed model enables interpretable inference about the association between topological shape features and survival risks. Two case studies are conducted using consecutive 133 lung cancer and 77 brain tumor patients. The results of both studies show that the topological features predict survival prognosis after adjusting clinical variables, and the predicted high-risk groups have worse survival outcomes than the low-risk groups. Also, the topological shape features found to be positively associated with survival hazards are irregular and heterogeneous shape patterns, which are known to be related to tumor progression.

cs.CV

Bayesian Landmark-based Shape Analysis of Tumor Pathology Images

Medical imaging is a form of technology that has revolutionized the medical field in the past century. In addition to radiology imaging of tumor tissues, digital pathology imaging, which captures histological details in high spatial resolution, is fast becoming a routine clinical procedure for cancer diagnosis support and treatment planning. Recent developments in deep-learning methods facilitate the segmentation of tumor regions at almost the cellular level from digital pathology images. The traditional shape features that were developed for characterizing tumor boundary roughness in radiology are not applicable. Reliable statistical approaches to modeling tumor shape in pathology images are in urgent need. In this paper, we consider the problem of modeling a tumor boundary with a closed polygonal chain. A Bayesian landmark-based shape analysis (BayesLASA) model is proposed to partition the polygonal chain into mutually exclusive segments to quantify the boundary roughness piecewise. Our fully Bayesian inference framework provides uncertainty estimates of both the number and locations of landmarks. The BayesLASA outperforms a recently developed landmark detection model for planar elastic curves in terms of accuracy and efficiency. We demonstrate how this model-based analysis can lead to sharper inferences than ordinary approaches through a case study on the 246 pathology images from 143 non-small cell lung cancer patients. The case study shows that the heterogeneity of tumor boundary roughness predicts patient prognosis (p-value < 0.001). This statistical methodology not only presents a new model for characterizing a digitized object's shape features by using its landmarks, but also provides a new perspective for understanding the role of tumor surface in cancer progression.

stat.AP

Empirical Likelihood Inference for Area under the ROC Curve using Ranked Set Samples

The area under a receiver operating characteristic curve (AUC) is a useful tool to assess the performance of continuous-scale diagnostic tests on binary classification. In this article, we propose an empirical likelihood (EL) method to construct confidence intervals for the AUC from data collected by ranked set sampling (RSS). The proposed EL-based method enables inferences without assumptions required in existing nonparametric methods and takes advantage of the sampling efficiency of RSS. We show that for both balanced and unbalanced RSS, the EL-based point estimate is the Mann-Whitney statistic, and confidence intervals can be obtained from a scaled chi-square distribution. Simulation studies and two case studies on diabetes and chronic kidney disease data suggest that using the proposed method and RSS enables more efficient inference on the AUC.

stat.ME

Bayesian Elastic Net based on Empirical Likelihood

We propose a Bayesian elastic net that uses empirical likelihood and develop an efficient tuning of Hamiltonian Monte Carlo for posterior sampling. The proposed model relaxes the assumptions on the identity of the error distribution, performs well when the variables are highly correlated, and enables more straightforward inference by providing posterior distributions of the regression coefficients. The Hamiltonian Monte Carlo method implemented in Bayesian empirical likelihood overcomes the challenges that the posterior distribution lacks a closed analytic form and its domain is nonconvex. We develop the leapfrog parameter tuning algorithm for Bayesian empirical likelihood. We also show that the posterior distributions of the regression coefficients are asymptotically normal. Simulation studies and real data analysis demonstrate the advantages of the proposed method in prediction accuracy.

stat.ME

Hypothesis Testing for Shapes using Vectorized Persistence Diagrams

Topological data analysis involves the statistical characterization of the shape of data. Persistent homology is a primary tool of topological data analysis, which can be used to analyze topological features and perform statistical inference. In this paper, we present a two-stage hypothesis test for vectorized persistence diagrams. The first stage filters vector elements in the vectorized persistence diagrams to enhance the power of the test. The second stage consists of multiple hypothesis tests, with false positives controlled by false discovery rates. We demonstrate the flexibility of our method by applying it to a variety of simulated and real-world data types. Our results show that the proposed hypothesis test enables accurate and informative inferences on the shape of data compared to the existing hypothesis testing methods for persistent homology.

stat.ME

Persistent homology machine learning for fingerprint classification

The fingerprint classification problem is to sort fingerprints into pre-determined groups, such as arch, loop, and whorl. It was asserted in the literature that minutiae points, which are commonly used for fingerprint matching, are not useful for classification. We show that, to the contrary, near state-of-the-art classification accuracy rates can be achieved when applying topological data analysis (TDA) to 3-dimensional point clouds of oriented minutiae points. We also apply TDA to fingerprint ink-roll images, which yields a lower accuracy rate but still shows promise, particularly since the only preprocessing is cropping; moreover, combining the two approaches outperforms each one individually. These methods use supervised learning applied to persistent homology and allow us to explore feature selection on barcodes, an important topic at the interface between TDA and machine learning. We test our classification algorithms on the NIST fingerprint database SD-27.

stat.ML

Persistence Terrace for Topological Inference of Point Cloud Data

Topological data analysis (TDA) is a rapidly developing collection of methods for studying the shape of point cloud and other data types. One popular approach, designed to be robust to noise and outliers, is to first use a smoothing function to convert the point cloud into a manifold and then apply persistent homology to a Morse filtration. A significant challenge is that this smoothing process involves the choice of a parameter and persistent homology is highly sensitive to that choice; moreover, important scale information is lost. We propose a novel topological summary plot, called a persistence terrace, that incorporates a wide range of smoothing parameters and is robust, multi-scale, and parameter-free. This plot allows one to isolate distinct topological signals that may have merged for any fixed value of the smoothing parameter, and it also allows one to infer the size and point density of the topological features. We illustrate our method in some simple settings where noise is a serious issue for existing frameworks and then we apply it to a real data set by counting muscle fibers in a cross-sectional image.

stat.ME