๐ค AI Summary
This work proposes a paradigm shift from validity-driven to value-driven data dependency discovery, addressing the limitations of traditional approaches that focus narrowly on statistical strength or validity while overlooking holistic value in real-world data governance tasksโsuch as relevance, redundancy, and lifecycle costs. Drawing on decision theory, the paper formally defines the utility and net value of dependencies and introduces a comprehensive, value-aware framework spanning discovery, validation, selection, and maintenance. Its key innovation lies in unifying task-specific loss reduction and lifecycle cost within a single evaluation framework, integrating budget-constrained optimization, lossโcost learning, and dynamic monitoring mechanisms. This establishes both the theoretical foundation and key technical pathways for value-driven dependency discovery, offering a new direction toward efficient and cost-effective data governance.
๐ Abstract
Data dependency discovery has traditionally focused on identifying dependencies that hold in the data or are statistically strong. Yet a dependency may be valid without being valuable: it may be irrelevant to the governance task, redundant given existing knowledge, or too costly to discover, validate, maintain, and apply. We call for a shift from validity-driven to value-driven dependency discovery. We define dependency use value decision-theoretically as the expected reduction in task-specific loss from incorporating a dependency into the governance process, and define net value by further accounting for lifecycle costs. Building on this framework, we outline principles for value-aware search, validation, dependency-set selection, and maintenance, and identify a research agenda spanning value estimation before full discovery, loss and cost learning, budgeted set selection, lifecycle monitoring, and benchmarking.