Score
Designs and implements procedures, tests, and tooling to verify that query outputs meet specified correctness, completeness, consistency, and integrity requirements; this includes automated comparison of returned results against expected results, constraints, or ground truth and the construction of checks for schema, types, and value ranges. Builds validation pipelines and analyzes discrepancies and error patterns to diagnose faults in queries, data, or processing logic and to determine corrective actions.
This work addresses the challenge of verifying relational database management systems’ (RDBMS) compliance with SQL semantics at the standard specification level. We present the first executable Prolog reference implementation grounded in the complete formal SQL semantics defined by ISO/IEC 9075, integrated with differential fuzz testing for semantic-level black-box validation. Unlike prior approaches relying solely on crash detection or meta-transformation, our method enables end-to-end verifiable modeling of SQL standard semantics. Empirical evaluation across MySQL, TiDB, SQLite, and DuckDB uncovered 19 previously unknown vulnerabilities and 11 semantic inconsistencies—each traceable to explicit violations, omissions, or ambiguities in the SQL standard. Our approach significantly enhances the decidability and interpretability of SQL implementation correctness.
Ad hoc SQL development lacks engineering rigor, leading to data silos, logical redundancy, and ineffective data governance. Method: This paper proposes a DataOps-driven CI/CD framework for analytical SQL warehouses, featuring a novel five-stage automated pipeline—Lint, Optimize, Parse, Validate, Observe—that embeds quality assurance and enables end-to-end lifecycle governance. Contribution/Results: We introduce the DataOps Controls Scorecard and a requirements traceability matrix, explicitly mapping 12 governance criteria to CI/CD stages to ensure control completeness and scalability. The framework integrates Agile, Lean, and DevOps principles with static analysis, syntactic parsing, optimization recommendations, validation testing, and observability. Empirical evaluation demonstrates significant improvements in data quality, development transparency, and cross-functional collaboration, providing a sustainable, production-ready pathway for large-scale analytical systems.
This work addresses the propagation of label errors in data validation, which can severely compromise the reliability of downstream query results. To quantify the impact of such errors and identify high-risk tuples whose uncertainty may be exacerbated by validation, the authors propose Maximum Error Score (MES)—a data-distribution-agnostic metric. Building on MES, they design MESReduce, an interactive validation optimization algorithm that adaptively guides the verification process by efficiently computing MES and incorporating feedback from external validators. Experimental evaluation on both real-world and synthetic datasets demonstrates that MESReduce significantly reduces the maximum error score and effectively enhances validation accuracy.
In industrial settings, limited production data severely compromises the fidelity of test data for SQL generation services (e.g., NL2SQL), hindering simultaneous preservation of structural integrity and semantic coherence. Method: This paper proposes an LLM-driven high-fidelity test data generation method, integrating Gemini with schema-aware preprocessing, SQL-semantic alignment postprocessing, and constraint-guided sampling—supporting complex patterns including nested columns, multi-table JOINs, aggregations, and deep subqueries. Contribution/Results: The method jointly optimizes semantic consistency, syntactic correctness, and structural fidelity, significantly improving test coverage and defect detection rates. Evaluated on Google’s real-world NL2SQL workloads, it generates high-quality mock data out-of-the-box, effectively addressing the semantic incoherence prevalent in existing approaches under large-scale, complex database schemas.
Existing conformance checking approaches between process models and reference models suffer from limited semantic expressiveness and insufficient automation, hindering fine-grained compliance verification. This paper proposes a semantic consistency checking method grounded in causal dependency analysis of tasks and events, transcending traditional trajectory-based dependency modeling by formally encoding causal constraints at the semantic level. We establish a unified framework integrating causal dependency modeling, semantic representation, and formal verification, and design an automated conformance checking algorithm implemented in a prototype tool. Empirical evaluation demonstrates that our approach significantly outperforms state-of-the-art techniques in both accuracy and flexibility, achieving— for the first time—the fully automated, high-expressivity semantic conformance verification of process models against reference models.
论文针对LLM数据代理在结构化任务中可能产生无效推理路径的问题,提出通过执行合约来实现可审计的Trace Integrity方法,以评估输出背后的计算是否可靠。
本文提出了一种形式化的语义块模型和执行评判基准来独立评估规范质量,通过结构化表示和机器可验证条件解决规范确定性问题。
This study addresses the difficulty coding agents face in verifying whether programs satisfy specifications and their inability to effectively leverage expert diagnostic experience from verification failures. Building upon the executable semantics of the K framework, this work proposes a method that transforms expert diagnostics into reusable guidance, establishing a comprehensive verification pipeline encompassing specification generation, proof repair, and adequacy auditing. By designing paired clean and defective program packages, the effectiveness of the auditing mechanism is systematically evaluated. The proposed approach achieves a 164/164 pass rate on the HumanEval benchmark, with all defects accurately identified. Furthermore, experiments on KleverBench reveal optimization opportunities for guidance selection strategies under resource constraints. This research offers a novel paradigm for enhancing the formal verification capabilities of coding agents.
本文提出ERIQ方法,通过对比VIEW、CTE和TEMPT表示的中间查询结果一致性来检测DBMS逻辑错误,共发现64个bug。
This study addresses the inefficiencies and impeded knowledge transfer arising from fragmented verification and validation (V&V) practices at the Jet Propulsion Laboratory (JPL). To overcome these challenges, this work proposes a unified V&V architecture grounded in human-centered design. By decoupling methodologies while maintaining a common attribute set, the architecture achieves bidirectional traceability through relational design and platform-independent SysML modeling. Furthermore, it establishes a comprehensive toolchain by integrating the Jama platform, modular templates, and digital thread technologies. This research effectively balances engineering rigor with agility, facilitating process automation, pattern reuse, and efficient cross-project collaboration. Ultimately, it provides a scalable and unified paradigm for the V&V of complex systems.