Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses

📅 2026-08-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a deterministic, evidence-driven approach to semantic recovery in production data warehouses lacking documentation and semantic annotations. By integrating a language model with a validation framework, the method leverages structured evidence—including value fingerprints, a library of 26 semantic patterns, and verification rules—to generate column-level semantic descriptions accompanied by confidence scores and provenance traces. A novel “capability detector” mechanism enables calibrated abstention when evidence is insufficient. Evaluated on 680 hidden columns, the approach achieves an accuracy of 0.475, substantially outperforming the baseline of 0.223. In blind tests on clinical data, it recovers 95.5% of ICD-9 codes, fully abstains on columns with no supporting evidence, and maintains 86% execution accuracy with 59% coverage under completely opaque conditions.
📝 Abstract
Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic identifiers, partial or absent documentation. We address the problem they pose first: recovering what columns and values mean from the data itself. Rosetta places a language model inside a verification harness: a deterministic profiler extracts structural evidence (value fingerprints, a 26-pattern library, checksum verdicts), the model proposes semantics conditioned on that evidence, and every fact carries provenance and a confidence bounded by its evidence class. Against human documentation on 680 paired columns across eleven BIRD databases, identifiers destroyed, the harness delivers metadata that is 0.475 accurate on the 42% of columns it commits to, against 0.223 on 94% for the same model used directly. Restricted to the 283 columns where both arms speak, the harness writes no better prose than the model alone; the gain is selection: deterministic evidence governs whether the system speaks (coverage +0.257 [0.128, 0.378]), not how well. The deterministic layer is a competence detector, not a competence amplifier. A backbone swap bounds the claim: the prose finding reproduces, but prompt-requested abstention does not transfer; a code-enforced commit gate (predictions registered first; measured on a third backbone and held-out databases) makes no-evidence coverage 0.000 on every backbone. On a blind i2b2 clinical warehouse Rosetta decodes 95.5% of 134 real ICD-9 codes from values alone and abstains on all 44 NDC drug codes. The catalog supports calibrated abstention at query time: under full schema opacity a naive translator falls from 0.92 to 0.42 execution accuracy while our gate answers at 86% accuracy over 59% coverage. Negative results are reported plainly, including that our own authority ladder is not the mechanism behind the headline.
Problem

Research questions and friction points this paper is trying to address.

metadata reconstruction
column semantics
undocumented warehouses
schema understanding
data profiling
Innovation

Methods, ideas, or system contributions that make the work stand out.

metadata reconstruction
deterministic verification harness
calibrated abstention
column semantics recovery
language model grounding
🔎 Similar Papers
No similar papers found.