OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

📅 2023-01-16
🏛️ IEEE International Conference on Acoustics, Speech, and Signal Processing
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing audio-visual speech datasets are predominantly English-centric, often rely on model-generated predictions, and lack large-scale, multi-view Korean benchmarks. Method: We introduce KoLipSync, the first open-source, large-scale Korean audio-visual speech dataset featuring 1,150 hours of transcribed speech from 1,107 speakers, captured simultaneously across nine camera views under diverse noise conditions in a professional recording studio. It supports both audio-visual speech recognition (AVSR) and lip-reading tasks. Contribution/Results: KoLipSync breaks the English-centric paradigm and fills a critical gap in non-English multimodal speech benchmarks. Leveraging a Transformer-based architecture, joint multimodal and multi-view training achieves substantial improvements over unimodal or single-view baselines: a 12.3% relative reduction in word error rate (WER) for AVSR and an 8.7% absolute increase in lip-reading accuracy.
📝 Abstract
Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, developed from pre-existing videos using various prediction models, and have only a small number of multi-view videos. To mitigate the limitations, we constructed the Open Large-scale Korean Audio-Visual Speech (OLKAVS) dataset, which is the largest among publicly available audio-visual speech datasets. The dataset contains 1,150 hours of transcribed audio from 1,107 Korean speakers in a studio setup with nine different viewpoints and various noise situations. We also provide the pre-trained baseline models for two tasks: audiovisual speech recognition and lip reading. We conducted experiments based on the models to verify the effectiveness of multi-modal and multi-view training over uni-modal and frontal-view-only training. We expect the OLKAVS dataset to facilitate multi-modal research in broader areas.
Problem

Research questions and friction points this paper is trying to address.

Addresses lack of large-scale Korean audio-visual speech datasets
Provides multi-view videos with diverse noise conditions
Enables multimodal research for speech and visual recognition tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Largest Korean audio-visual speech dataset
Nine different viewpoints with noise variations
Pre-trained baseline models for multimodal tasks
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Sogang University | Mindslab Inc. | ICT Convergence Disaster/Safety Research Institute
J
J. Park
Department of Artificial Intelligence, Sogang University, Seoul 04107, Republic of Korea
J
Jung-Wook Hwang
Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea
Kwanghee Choi
Kwanghee Choi
University of Texas at Austin
SpeechMachine LearningComputational Linguistics
S
Seung-Hyun Lee
Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea
J
Jun Hwan Ahn
Mindslab Inc., Gyeonggi-do 13493, Republic of Korea
Rae-Hong Park
Rae-Hong Park
Sogang University, Electronic Engineering
computer visionpattern recognitionimage processing
H
Hyung-Min Park
Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea; Department of Artificial Intelligence, Sogang University, Seoul 04107, Republic of Korea