🤖 AI Summary
This study addresses the difficulty of attributing model behaviors to their origins due to the lack of generative provenance in synthetic speech data. We propose a compact provenance contract and auditing protocol, formally establishing provenance as a necessary but insufficient condition for behavioral attribution. Methodologically, we construct synthetic research objects that bind source specifications to content within a Japanese nursing care scenario, implementing audits through immutable manifests, disjoint versioning of scenario seeds, and multimodal asset linkage. Experimentally, we audit 1.55 hours of speech, revealing impediments to precise upstream attribution and establishing a candidate causal graph framework. This work provides a novel paradigm for enhancing the traceability and causal analysis of synthetic data.
📝 Abstract
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.