arXiv · 2311.16375
Testing for a difference in means of a single feature after clustering
Abstract
For many applications, it is critical to interpret and validate groups of observations obtained via clustering. A common validation approach involves testing differences in feature means between observations in two estimated clusters. In this setting, classical hypothesis tests lead to an inflated Type I error rate. To overcome this problem, we propose a new test for the difference in means in a single feature between a pair of clusters obtained using hierarchical or $k$-means clustering. The test based on the proposed $p$-value controls the selective Type I error rate in finite samples and can be efficiently computed. We further illustrate the validity and power of our proposal in simulation and demonstrate its use on single-cell RNA-sequencing data.
Explore related subjects
Keep this discovery
Yiqun T. Chen, Lucy L. Gao. 2023-11-27. Testing for a difference in means of a single feature after clustering. https://arxiv.org/abs/2311.16375
Cite the original work for its findings. Save a collection to share your selection of sources.