Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
This study addresses the difficulty of attributing model behaviors to their origins due to the lack of generative provenance in synthetic speech data. We propose a compact provenance contract and auditing protocol, formally establishing provenance as a necessary but insufficient condition for behavioral attribution. Methodologically, we construct synthetic research objects that bind source specifications to content within a Japanese nursing care scenario, implementing audits through immutable manifests, disjoint versioning of scenario seeds, and multimodal asset linkage. Experimentally, we audit 1.55 hours of speech, revealing impediments to precise upstream attribution and establishing a candidate causal graph framework. This work provides a novel paradigm for enhancing the traceability and causal analysis of synthetic data.