🤖 AI Summary
Addressing challenges in data engineering—including schema drift, difficulty handling heterogeneous data types, and insufficient interpretability in file, database, and query-result diffing—this paper introduces the first unified, scalable differential analysis framework. Our method integrates schema-aware mapping, type-specific comparators, and an LLM-enhanced retrieval-constrained multi-label explanation generator to enable high-accuracy comparison and root-cause localization across structured and semi-structured data. Evaluated on million-row datasets, the framework achieves >95% precision and recall, outperforms baselines by 30–40% in throughput, reduces memory consumption by 30–50%, and shortens root-cause analysis time from 10 hours to 12 minutes. These advances significantly improve reliability and interpretability in data migration validation, regression testing, and regulatory compliance auditing.
📝 Abstract
Data engineering workflows require reliable differencing across files, databases, and query outputs, yet existing tools falter under schema drift, heterogeneous types, and limited explainability. SmartDiff is a unified system that combines schema-aware mapping, type-specific comparators, and parallel execution. It aligns evolving schemas, compares structured and semi-structured data (strings, numbers, dates, JSON/XML), and clusters results with labels that explain how and why differences occur. On multi-million-row datasets, SmartDiff achieves over 95 percent precision and recall, runs 30 to 40 percent faster, and uses 30 to 50 percent less memory than baselines; in user studies, it reduces root-cause analysis time from 10 hours to 12 minutes. An LLM-assisted labeling pipeline produces deterministic, schema-valid multilabel explanations using retrieval augmentation and constrained decoding; ablations show further gains in label accuracy and time to diagnosis over rules-only baselines. These results indicate SmartDiff's utility for migration validation, regression testing, compliance auditing, and continuous data quality monitoring. Index Terms: data differencing, schema evolution, data quality, parallel processing, clustering, explainable validation, big data