arXiv Science⌕ Search

arXiv subjects

Jianwen Tian

Publications and source records attributed to Jianwen Tian.

2 recordsLinked to original sources

Weaponizing Ground Truth: Data Poisoning Attacks by Exploiting Boundary Misalignment Between Antivirus Software and Learning-Based Detectors

Machine-learning (ML)-based malware detectors are commonly trained using labels obtained from antivirus (AV) engines and aggregation services (e.g., VirusTotal). This practice assumes AV-generated labels provide reliable supervision. However, small byte-level modifications can substantially alter AV verdicts while leaving the representations perceived by downstream ML detectors largely unchanged, producing label-feature inconsistencies that can contaminate training datasets and create poisoning opportunities for ML-based malware detection. We present Bi-Iocane, a black-box poisoning framework that exploits the reliance of malware-labeling pipelines on AV-generated labels. Bi-Iocane identifies AV-sensitive bytes and modifies them to induce label changes. It rewrites such bytes in malware to obtain benign labels (evasion-oriented poisoning) and injects malware-associated byte patterns into benign software to obtain malicious labels (defamation-oriented poisoning). These poisoned samples and their lightly modified variants corrupt training data and cause selected targets to be misclassified. We evaluate Bi-Iocane with 13 AV engines simulating AV aggregation services and eight ML detectors. For 30 malware and 30 benign clean targets, Bi-Iocane combines AV-specific manipulations to generate malware-to-benign and benign-to-malware poisoned samples whose all tested AV-based labels are flipped. After these poisoned samples and variants are used for downstream training, the resulting ML models misclassify 92.08% of the original clean targets on average with only a 0.06\% poisoning budget per target. Meanwhile, the poisoned models largely preserve clean-set performance, and six evaluated poisoning defenses show only limited mitigation. VirusTotal evaluation further confirms practical defamation risk and reveals potential evasion risk in real-world AV-to-ML labeling supply chains.

cs.CR↗

Breaking the Black Box: Byte-Level Boundary Inference of Real-World Antivirus Systems

Existing approaches for understanding the detection logic of real-world antivirus (AV) software infer only binary malware/benign decisions from black-box queries, providing limited insight into the fine-grained decision-critical regions that govern AV detection. In this paper, we present \textbf{AVHunter}, the first framework for inferring byte-level decision-critical regions of real-world AV products under a black-box threat model. AVHunter constructs the first large-scale Byte-Level AV Boundary Dataset (BABD) by systematically probing 11 real-world AV products, revealing that modern AV detections are largely associated with a small number of compact decision-critical byte regions. Leveraging BABD, AVHunter trains AV-specific models that not only reproduce binary AV decisions, but also localize the decision-critical byte regions underlying these decisions, achieving an average boundary prediction recall of 85.07% while maintaining 97.43% detection agreement with the target AVs. We further validate that the predicted regions capture genuine AV decision knowledge through boundary-guided malware evasion, false-positive induction on benign executables, and a seven-month longitudinal study demonstrating that the inferred regions remain largely stable as AV products evolve. Overall, AVHunter moves beyond conventional binary-label AV modeling by enabling fine-grained boundary-region localization and revealing a new form of AV knowledge leakage with important implications for malware analysis, AV security, and boundary-aware defenses.

cs.CR↗