arXiv · 2609.20959
Don't Blame the Model, Verify the Data: An Evaluation of SMT-based Dataset Verification
Abstract
The EU AI Act mandates that datasets for high-risk machine learning (ML) systems meet strict quality criteria such as soundness and bias mitigation. While Satisfiability Modulo Theory (SMT) solving offers a formal approach to verifying these properties, its scalability in realistic ML settings remains unexplored. To bridge this gap, this work presents the first large-scale empirical study of SMT-based dataset verification on two real-world ML datasets. We systematically evaluate how solver performance is shaped by three key dimensions: the type of data-quality property, the specification style, and the dataset encoding strategy. Our findings demonstrate that SMT-based verification is feasible for practical scenarios, but each dimension shapes it: the property type sets the tractability limit, the specification style drives scalability (exceeding 2000x for aggregate properties), and the encoding strategy has a systematic effect, with extracted feature columns performing best.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sehee Park, Dominik Geißler, Andrei Aleksandrov, Kim Völlinger. 2026-09-17. Don't Blame the Model, Verify the Data: An Evaluation of SMT-based Dataset Verification. https://arxiv.org/abs/2609.20959
Cite the original work for its findings. Save a collection to share your selection of sources.