SPACER: A Parallel Dataset of Speech Production And Comprehension of Error Repairs

📅 2025-03-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates systematic asymmetries between speaker self-repairs and listener text corrections in handling single-word substitution errors during spoken interaction. We construct SPACER—the first parallel corpus synchronously annotated for both speaker repair and listener editing behaviors—based on the Switchboard corpus, enabling the first joint annotation of production and comprehension systems. Using offline editing experiments, combined semantic similarity measures (WordNet and BERT embeddings), and phonological distance metrics (Levenshtein edit distance augmented with phoneme-level alignment), we identify a consistent functional asymmetry: speakers primarily respond to error severity (deviation strength), whereas listeners rely more heavily on contextual fit and phonological substitutability. We publicly release the SPACER dataset as an open resource, providing a benchmark for cognitively grounded, integrated modeling of speech production and comprehension, and advancing unified theories of language processing.

Technology Category

Natural Language Processing: SpeechCognitive Modeling & Cognitive Systems: Social Cognition And InteractionData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Assisted, interactive, and conversational search
📝 Abstract
Speech errors are a natural part of communication, yet they rarely lead to complete communicative failure because both speakers and comprehenders can detect and correct errors. Although prior research has examined error monitoring and correction in production and comprehension separately, integrated investigation of both systems has been impeded by the scarcity of parallel data. In this study, we present SPACER, a parallel dataset that captures how naturalistic speech errors are corrected by both speakers and comprehenders. We focus on single-word substitution errors extracted from the Switchboard corpus, accompanied by speaker's self-repairs and comprehenders' responses from an offline text-editing experiment. Our exploratory analysis suggests asymmetries in error correction strategies: speakers are more likely to repair errors that introduce greater semantic and phonemic deviations, whereas comprehenders tend to correct errors that are phonemically similar to more plausible alternatives or do not fit into prior contexts. Our dataset enables future research on integrated approaches toward studying language production and comprehension.
Problem

Research questions and friction points this paper is trying to address.

Lack of parallel data on speech error correction in production and comprehension.
Asymmetries in error correction strategies between speakers and comprehenders.
Need for integrated study of language production and comprehension systems.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel dataset for speech error correction
Focus on single-word substitution errors
Analyzes speaker and comprehender correction strategies
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shiva Upadhye
Department of Language Science, University of California, Irvine
J
Jiaxuan Li
Department of Language Science, University of California, Irvine
Richard Futrell
Richard Futrell
Associate Professor, UC Irvine Language Science
Computational linguisticspsycholinguisticsinformation theory