arXiv · 2610.05605
Handling Missing Data in Performance Portability Studies
Abstract
Missing data is a common challenge in many real-world datasets, often leading to biased results or reduced accuracy. Missing data is a pervasive problem in performance efficiency datasets, particularly in high-performance computing (HPC) performance portability studies. Performance portability scores rely on complete sets of performance efficiencies as inputs to their respective metrics; however, missing values can arise due to incomplete benchmarking, hardware constraints, or implementation gaps. Many established performance portability metrics lack mechanisms for handling missing inputs, which can result in incomplete analyses or the inability to compute scores altogether. In this study, we evaluate a range of imputation methods and algorithms for addressing missing performance efficiencies. We present illustrative examples for each approach and apply them to real-world HPC performance portability datasets. Our results highlight the trade-offs between accuracy, robustness, and computational cost, offering practical recommendations for researchers seeking to mitigate the impact of missing data in performance portability assessments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ami Marowka. 2026-10-04. Handling Missing Data in Performance Portability Studies. https://arxiv.org/abs/2610.05605
Cite the original work for its findings. Save a collection to share your selection of sources.