π€ AI Summary
This study addresses the challenge of manually verifying residual target-speaker speech following the de-identification of psychiatric audio recordings. To this end, we propose a symmetric dual-pipeline automated verification system built upon Gemma and Nemotron series open-source audio and large language models. The framework introduces a novel symmetric dual-pipeline architecture combined with a multi-model view disjunctive ensemble strategy, significantly enhancing detection robustness through model complementarity. Experimental evaluation on a corpus of 48 dialogue recordings demonstrates that the proposed method achieves an overall F1 score of 0.478 and a recall of 0.870, substantially outperforming single-model baselines. These results indicate that the approach offers an efficient, automated quality-control solution for clinical audio privacy protection.
π Abstract
Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora. We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines: confirming that clinician speech has been removed from psychiatric interview recordings, and confirming that patient speech has been removed from the same recordings. Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions. We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting: Gemma-4-12B, Gemma-4-31B, Nemotron-3-Nano, and Nemotron-3-Nano-Omni. A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.