data integrity

Designs, builds, or audits systems, processes, and controls that ensure data remain accurate, consistent, complete, and reliable across their lifecycle; this includes specifying and implementing validation rules, checksums/hashes, versioning and transactional guarantees, access controls, audit trails, and procedures for detecting, correcting, and preventing corruption or unauthorized modification.

dataintegrity

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

From What to How: A Taxonomy of Formalized Security Properties

May 20, 2025
IS
Imen Sayar
🏛️ University of Toulouse | IRIT

High-level security properties (e.g., confidentiality, integrity) in the Software Development Life Cycle (SDLC) lack systematic refinement mechanisms, leading to semantic disconnects between these properties and concrete artifacts such as threats, defenses, and assets. Method: We propose the first SDLC-wide security property refinement taxonomy, implemented as a formal, refinable, verifiable, and traceable classification framework in Event-B. The framework integrates principles from security engineering and adaptive systems theory. Contribution: It bridges the semantic gap between high-level security objectives and mid-to-low-level security models, enabling co-evolution of security properties with threat and defense models. Rigorously verified in Event-B, the framework ensures logical consistency and correctness. It provides both theoretically sound foundations and practically actionable guidance for security requirements–driven system development.

Aligning security properties refinement with attacks and defensesProposing an SDLC taxonomy for security propertiesVerifying taxonomy correctness using Event-B formal language

This study addresses critical challenges faced by regulated enterprises—including cross-system data inconsistencies, reconciliation difficulties, asset record drift, and overreliance on manual audits—by proposing the GERA framework. GERA innovatively integrates deterministic reconciliation, robust anomaly detection based on Z-Score and its variants, governance-driven semantic standardization, and NIST CSF 2.0 security controls within a four-layer architecture comprising ingestion, staging, core modeling, and semantic services. Empirical validation across banking, broadband service providers, and technology firms demonstrates that the framework significantly enhances reconciliation automation and audit readiness, effectively mitigating 39% of compliance deficiencies identified during PCAOB inspections.

Audit ReadinessCross-System ReconciliationData Fragmentation

Simplified integrity checking for an expressive class of denial constraints

Dec 30, 2024
DM
Davide Martinenghi
🏛️ Politecnico di Milano

Efficiently verifying complex (especially negation-based) integrity constraints in data-intensive systems remains challenging. Method: This paper proposes a lightweight, automated checking framework for extended denial constraints—more expressive than tgds and egds—by introducing program transformation into integrity verification. It employs formal constraint modeling, constraint-driven SQL equivalence simplification, and automatic SQL generation, enabling incremental validation under database consistency assumptions. Contribution/Results: The approach breaks from traditional syntax-bound verification paradigms, supporting direct deployment of standard SQL without custom extensions. Under strict semantic equivalence, it significantly reduces runtime overhead while maintaining compatibility with mainstream OLTP systems.

Automated Correctness CheckingData Integrity RulesData Quality Assurance

DataOps-driven CI/CD for analytics repositories

Nov 15, 2025
DV
Dmytro Valiaiev
🏛️ University of Arkansas Little Rock

Ad hoc SQL development lacks engineering rigor, leading to data silos, logical redundancy, and ineffective data governance. Method: This paper proposes a DataOps-driven CI/CD framework for analytical SQL warehouses, featuring a novel five-stage automated pipeline—Lint, Optimize, Parse, Validate, Observe—that embeds quality assurance and enables end-to-end lifecycle governance. Contribution/Results: We introduce the DataOps Controls Scorecard and a requirements traceability matrix, explicitly mapping 12 governance criteria to CI/CD stages to ensure control completeness and scalability. The framework integrates Agile, Lean, and DevOps principles with static analysis, syntactic parsing, optimization recommendations, validation testing, and observability. Empirical evaluation demonstrates significant improvements in data quality, development transparency, and cross-functional collaboration, providing a sustainable, production-ready pathway for large-scale analytical systems.

Addressing ad-hoc SQL development lacking software engineering rigorProviding standardized DataOps framework for analytics pipeline managementSolving data governance challenges and validation impossibility in analytics

Latest Papers

What's happening recently
View more

This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.

design conformancedistributed systemsimplementation drift

This study addresses the challenges of applying formal verification to production-grade software, where high modeling costs and consistency risks in fault handling hinder adoption. By integrating runtime execution traces with formal specifications, the authors verify a real-world restaurant point-of-sale (POS) payment workflow and leverage large language models (LLMs) to automatically generate these specifications. Their analysis reveals that the structural form—not the natural language phrasing—of specifications primarily governs LLM-generated correctness, and uncovers a shared “relevant oracle failure” issue between code and simulators. Extending fault models to include crash-recovery, stale reads, and retries, the team conducts simulation-based audits, verifying core protocol correctness, identifying and reproducing seven fault-handling vulnerabilities, and revalidating after fixes. They also expose how deviations in API response structures render recovery paths unreachable—a finding consistently replicated across seven LLMs from two vendors.

failure handlingformal verificationpayment workflow

Current AI systems rely heavily on manual auditing and documentation, which hinders scalable governance for automated services. This work proposes Ontological Knowledge Blocks (OKBs), a novel framework that formalizes regulatory obligations as quintuples comprising ontologies, SHACL rules, evidence requirements, and provenance links. By leveraging RDF/OWL modeling, PROV-O for provenance tracking, and an intermediate representation–driven deterministic compiler, the approach enables dynamic switching of governance configurations without modifying service code. Evaluation in an AI-assisted HPC scheduling scenario demonstrates that compliance checks are configuration-sensitive, violations accumulate strictly additively, SHACL validation incurs only 12.6–100.3 milliseconds of latency, and the Combined configuration provides the most comprehensive coverage.

AI governanceautomated verificationcompliance

This work addresses the limitations of existing AI trustworthiness assessment approaches, which are either too abstract to support full lifecycle monitoring or rely on single metrics insufficient for governance needs. The paper proposes a lightweight, auditable framework for dynamic trustworthiness management that integrates formal modeling with governance processes. By employing context-sensitive trustworthiness dimension protocols and interpretable rule learning based on decision trees, the framework enables end-to-end monitoring and documentation of AI systems—from design and deployment through re-evaluation. Novel diagnostic tools, including hierarchical transitions, margin-of-boundary analysis, and profile drift detection, are introduced alongside clearly accountable human-in-the-loop checkpoints. Experiments on synthetic AI lifecycle trajectories demonstrate the framework’s effectiveness in detecting performance degradation, abrupt perturbations, and impacts of system updates, thereby establishing a transparent, traceable, and contestable evidentiary basis for AI governance.

AI governanceauditableconformity documentation