InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of speech recognition distortion in VR interviews and the poor transferability of NLP models to spoken data by constructing a German multimodal corpus. By integrating VR avatar-based interviews with eye and facial tracking, this work achieves, for the first time, high-precision synchronization between multimodal behavioral signals, manually corrected transcripts, and fine-grained linguistic annotations such as negation scope. Experiments quantify the ASR error rate on short responses, demonstrating that the proposed approach significantly enhances downstream NLP model performance in spoken-language scenarios. Ultimately, this research provides a high-quality benchmark resource for multimodal language analysis within VR environments.
📝 Abstract
We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated (α=0.87 for cues; α=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.
Problem

Research questions and friction points this paper is trying to address.

multimodal corpus
virtual reality
automatic speech recognition
domain transfer
negation detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Corpus
Virtual Reality
Synchronized Behavioral Data
Negation Annotation
ASR Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Patrick Schrottenbacher
Goethe University, Frankfurt am Main, Germany
L
Leon Hammerla
Goethe University, Frankfurt am Main, Germany
L
Lydia Kleine
Leibniz Institute for Educational Trajectories (LIfBi), Bamberg, Germany
D
Doris Stingl
Leibniz Institute for Educational Trajectories (LIfBi), Bamberg, Germany
Alexander Mehler
Alexander Mehler
Professor of Computer Science, Goethe University Frankfurt am Main
Computational HumanitiesText-technology