The Impact of Automatic Speech Transcription on Speaker Attribution

📅 2025-07-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the effectiveness and robustness of automatic speech recognition (ASR) transcripts for speaker attribution, particularly examining the impact of transcription errors. Contrary to the conventional assumption that ASR errors degrade performance, experiments reveal that word-level errors do not significantly impair attribution accuracy—and may even introduce speaker-discriminative linguistic cues; in certain settings, ASR-generated transcripts yield superior attribution compared to human transcriptions. The work systematically evaluates how ASR system characteristics—including recognition accuracy, vocabulary coverage, and contextual modeling capability—affect attribution performance, corroborating findings through linguistic pattern analysis and speaker特征 modeling. Results demonstrate that speaker attribution based on ASR output is highly resilient and practically viable, challenging the prevailing assumption that high-fidelity transcription is a prerequisite. This establishes a novel paradigm for speaker identification in audio-deprived scenarios, where only ASR transcripts are available.

Technology Category

Natural Language Processing: SpeechCognitive Modeling & Cognitive Systems: Affective ComputingPlanning, Routing, and Scheduling: Activity and Plan Recognition

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsSecurity and Privacy: Large-scale security measurements
📝 Abstract
Speaker attribution from speech transcripts is the task of identifying a speaker from the transcript of their speech based on patterns in their language use. This task is especially useful when the audio is unavailable (e.g. deleted) or unreliable (e.g. anonymized speech). Prior work in this area has primarily focused on the feasibility of attributing speakers using transcripts produced by human annotators. However, in real-world settings, one often only has more errorful transcripts produced by automatic speech recognition (ASR) systems. In this paper, we conduct what is, to our knowledge, the first comprehensive study of the impact of automatic transcription on speaker attribution performance. In particular, we study the extent to which speaker attribution performance degrades in the face of transcription errors, as well as how properties of the ASR system impact attribution. We find that attribution is surprisingly resilient to word-level transcription errors and that the objective of recovering the true transcript is minimally correlated with attribution performance. Overall, our findings suggest that speaker attribution on more errorful transcripts produced by ASR is as good, if not better, than attribution based on human-transcribed data, possibly because ASR transcription errors can capture speaker-specific features revealing of speaker identity.
Problem

Research questions and friction points this paper is trying to address.

Impact of ASR transcription errors on speaker attribution
Comparison of speaker attribution using human vs ASR transcripts
Resilience of attribution to word-level transcription errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses automatic speech recognition for transcription
Analyzes speaker attribution with errorful transcripts
Resilient to word-level transcription errors
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.