Don't Blame the Model, Verify the Data: An Evaluation of SMT-based Dataset Verification

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过SMT求解方法验证高风险ML系统的数据集质量,评估了数据属性类型、规范风格和编码策略对验证性能的影响。
📝 Abstract
The EU AI Act mandates that datasets for high-risk machine learning (ML) systems meet strict quality criteria such as soundness and bias mitigation. While Satisfiability Modulo Theory (SMT) solving offers a formal approach to verifying these properties, its scalability in realistic ML settings remains unexplored. To bridge this gap, this work presents the first large-scale empirical study of SMT-based dataset verification on two real-world ML datasets. We systematically evaluate how solver performance is shaped by three key dimensions: the type of data-quality property, the specification style, and the dataset encoding strategy. Our findings demonstrate that SMT-based verification is feasible for practical scenarios, but each dimension shapes it: the property type sets the tractability limit, the specification style drives scalability (exceeding 2000x for aggregate properties), and the encoding strategy has a systematic effect, with extracted feature columns performing best.
Problem

Research questions and friction points this paper is trying to address.

SMT
dataset verification
machine learning
data quality
scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

SMT-based dataset verification
data-quality property
specification style
dataset encoding strategy
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
S
Sehee Park
Technische Universität Berlin, Berlin, Germany
D
Dominik Geißler
Technische Universität Berlin, Berlin, Germany
A
Andrei Aleksandrov
Technische Universität Berlin, Berlin, Germany; Fraunhofer FOKUS, Berlin, Germany
K
Kim Völlinger
Technische Universität Berlin, Berlin, Germany