Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决日语语音识别中词汇读音区分问题,提出Ruby-ASR方法,通过结合正字法和词汇读音序列提高识别准确性。
📝 Abstract
Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and cannot be recovered reliably by post-hoc text-only grapheme-to-phoneme conversion. We present Ruby-ASR, which refines the conventional target into a span-bound orthographic--lexical-reading sequence. Unlike separate full-sentence orthographic and phonological outputs, the ruby representation locally binds each written span to its realized reading and permits deterministic recovery of both views. We instantiate the target under subtitle-style and verbatim-style transcription conventions using a Qwen3-ASR backbone; a mora-level CTC objective provides auxiliary monotonic reading supervision. The experimental results across five Japanese benchmarks show that refining the recognition target can improve lexical-reading recovery without sacrificing readable orthographic transcription. We release the checkpoints and inference code.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
Orthographic Transcript
Lexical Reading
Supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ruby-ASR
Orthographic-Lexical-Reading Sequence
Evidence-Preserving Supervision
Mora-Level CTC Objective
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hao Shi
Independent Researcher
Y
Yun Liu
Independent Researcher
X
Xuehao Yang
Independent Researcher
J
Jun Liu
Independent Researcher
Chuanbo Hua
Chuanbo Hua
Postdoctoral Researcher @ KAIST
Reinforcement LearningCombination OptimizationLLM for Algorithm Design
Xuanjun Chen
Xuanjun Chen
National Taiwan University
Speech ProcessingMachine LearningGenerative AIDeepfakes
L
Lianbo Liu
Independent Researcher
S
Shiao Zhu
Independent Researcher
Zixiong Su
Zixiong Su
The University of Tokyo
Human-Computer InteractionSilent Speech InterfaceHuman-AI Interaction