🤖 AI Summary
ASR systems struggle to accurately recognize domain-specific terminology in professional settings such as academic lectures. To address this, we introduce the SlideASR task—leveraging slide visual content to enhance speech transcription. Existing pipeline approaches are cumbersome and inefficient, while multimodal large language models (MLLMs) often degenerate into pure OCR systems. We propose Visually-Anchored Policy Optimization (VAPO): an MLLM-based framework integrating Chain-of-Thought reasoning and visual anchoring to jointly model OCR, ASR, and visual–acoustic consistency, augmented with a quadruple-reward reinforcement learning objective for end-to-end post-training. Experiments demonstrate substantial improvements in domain-term recognition accuracy, achieving state-of-the-art performance on both synthetic and real-world academic lecture data. Furthermore, we release SlideASR-Bench—the first benchmark dataset for slide-augmented ASR—featuring rich academic entities and fine-grained multimodal alignment annotations, advancing research in domain-specialized speech understanding.
📝 Abstract
Automatic speech recognition (ASR) systems often struggle with domain-specific terminology, especially in specialized settings such as academic lectures. To address this, we define the SlideASR task, which leverages the rich visual information from presentation slides to improve transcription accuracy. Existing pipeline methods for this task tend to be complex and underperform. Although omni-modal large language models (OLLMs) provide a promising end-to-end framework, they frequently fail in practice by degenerating into simple optical character recognition (OCR) systems. To overcome this, we propose Visually-Anchored Policy Optimization (VAPO), a novel post-training method designed to control the model's reasoning process. Drawing on the Chain-of-Thought reasoning paradigm, VAPO enforces a structured "Look before Transcription" procedure using a <think><answer> format. Specifically, the model first performs OCR on the slide content within the think step, then generates the transcription by referencing this recognized visual information in the answer step. This reasoning process is optimized via reinforcement learning with four distinct rewards targeting format compliance, OCR accuracy, ASR quality, and visual anchoring consistency. To support further research, we construct SlideASR-Bench, a new entity-rich benchmark consisting of a synthetic dataset for training and testing, and a challenging real-world set for evaluation. Extensive experiments demonstrate that VAPO significantly improves recognition of domain-specific terms, establishing an effective end-to-end paradigm for SlideASR.