Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of manually verifying residual target-speaker speech following the de-identification of psychiatric audio recordings. To this end, we propose a symmetric dual-pipeline automated verification system built upon Gemma and Nemotron series open-source audio and large language models. The framework introduces a novel symmetric dual-pipeline architecture combined with a multi-model view disjunctive ensemble strategy, significantly enhancing detection robustness through model complementarity. Experimental evaluation on a corpus of 48 dialogue recordings demonstrates that the proposed method achieves an overall F1 score of 0.478 and a recall of 0.870, substantially outperforming single-model baselines. These results indicate that the approach offers an efficient, automated quality-control solution for clinical audio privacy protection.
πŸ“ Abstract
Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora. We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines: confirming that clinician speech has been removed from psychiatric interview recordings, and confirming that patient speech has been removed from the same recordings. Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions. We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting: Gemma-4-12B, Gemma-4-31B, Nemotron-3-Nano, and Nemotron-3-Nano-Omni. A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.
Problem

Research questions and friction points this paper is trying to address.

Speaker Deletion Verification
Clinical Psychiatry
Audio Language Models
Dyadic Dialogue
Speech Redaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speaker Deletion Verification
Audio Language Models
Role-guided Redaction
Clinical Psychiatry Speech
Ensemble Strategy
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
J
Joseph T Colonel
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
D
Daniel Katzman
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
K
Kelsey Kirker
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
A
Adam N Davidson
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
S
Shalaila S Haas
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
C
Cheryl Corcoran
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
R
RenΓ© S Kahn
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
G
Guillermo Checci
Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA
Baihan Lin
Baihan Lin
Tenure-Track Professor, Mount Sinai, Harvard University
Speech / NLPML / RL / BanditsComputational PsychiatryTheoretical NeuroscienceBio-Inspired AI