Domain Bounds as a Silent-Fault Detector for AI-Ready Scientific Data

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of silent failures in AI scientific data reuse, where context loss can produce numerically valid yet semantically conflicting results. To overcome the limitations of traditional lineage and descriptive provenance tracking, this work proposes the concept of "domain boundaries," which explicitly encodes data generation conditions and binds them to standardized data containers, thereby enabling automated detection of invalid reuse scenarios. Experimental evaluation demonstrates that the proposed approach successfully identifies 17 out of 18 injected faults, significantly outperforming existing baseline methods. Furthermore, its detection performance remains robust regardless of dataset scale. These findings establish a new paradigm for ensuring safe cross-context data reuse in AI-driven scientific research.
📝 Abstract
We address an incipient type of data-based faults, driven by increasing amounts of data reuse in AI-enhanced computational science workflows. A dataset is created under specific conditions; its categories, measures, and labels depend on that context, but that context is usually unavailable to downstream users and is completely abandoned when datasets are used to train models. We propose the encoding and association of this context, called domain bounds, with the data to which it applies. By specifying the conditions for valid reuse, a domain bound flags apparently valid data from being reused in an incompatible context, even when every value falls within its expected range. We present cases from computational science and biomedicine where missing domain bounds allow silent reuse errors, and show why description and provenance approaches do not detect them. Our detector uses standard data container technologies, catches faults drawn from a taxonomy of published cases, irrespective of dataset size. Against injected faults it caught 17 of 18 domain-bound reuses that lineage and descriptive records accepted.
Problem

Research questions and friction points this paper is trying to address.

silent-fault detection
data reuse
domain bounds
AI-ready scientific data
context loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain Bounds
Silent-Fault Detection
Data Reuse
Provenance
Computational Science
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.