arXiv · 2609.34317
The Construction of an Empirical Dataset of Incomplete Software Changes from Open Source Projects
Abstract
During software development, a modification to a software component may propagate across the system, requiring precise identification and correct revision of all affected components. This is a complex task, and developers often (45.7%) miss related changes. To address this, several methods have been developed to extract co-change rules from files that are frequently changed together in the revision history. However, previous research evaluated the methods using artificially created incomplete changes, which may not be representative of real-world data. To solve this problem, we construct a dataset by mining incomplete changes from a collection of open-source software, using information about induced bugs and their respective fixes from an issue tracking platform. We also analyze the characteristics of incomplete changes using this constructed dataset and found that 89.4% of missed changes involved five or fewer files. Finally, we re-evaluate LCExtractor, an existing co-change rule extraction method, on our constructed dataset, and we identify the optimal sorting criterion and the impact of the number of used commits.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Savira Ramadhanty, Profir-Petru Pârţachi, Yoshiya Ishida, Takashi Kobayashi. 2026-09-28. The Construction of an Empirical Dataset of Incomplete Software Changes from Open Source Projects. https://arxiv.org/abs/2609.34317
Cite the original work for its findings. Save a collection to share your selection of sources.