arXiv · 2609.27936
Quality over Quantity: Semi-Supervised Detection of Illicit Bitcoin Flows via Feature Engineering
Abstract
Detecting illicit cryptocurrency transactions is hampered by extreme class imbalance, adversarial obfuscation, and a scarcity of reliable labels. While semi-supervised learning (SSL) offers a promising solution by leveraging unlabeled data, we show that its success is not guaranteed by data volume alone but is contingent on data quality. We introduce an SSL framework for detecting illicit Bitcoin flows in Shared Send Mixers (SSM) transactions, built on a comprehensive historical dataset comprising 163 million transactions. Our main conclusion is that the success of SSL depends on data quality rather than volume: high-fidelity features such as KeyLinker address clustering and Shared Send Untangling (SSU) complexity metrics achieve an F1 score of 0.84 on unlabeled data. Finally, we empirically show that common heuristics like One-Time Change (OTC), though abundant, introduce noise, while strategic reliance on higher-fidelity features like KeyLinker is essential. Our work establishes that in blockchain forensics, the path to better performance lies in smarter feature engineering for data quality, not just larger datasets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yekaterina Smolenkova, Nickolay Larionov, Nikolay Ivanov, Yury Yanovich. 2026-08-22. Quality over Quantity: Semi-Supervised Detection of Illicit Bitcoin Flows via Feature Engineering. https://arxiv.org/abs/2609.27936
Cite the original work for its findings. Save a collection to share your selection of sources.