Institution profile

Unisound

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Unified Target-Speaker ASR with Text and Enrollment Speech Cues

Sep 27, 2026

This study addresses the insufficient complementarity caused by the separation of textual and enrollment speech cues in multi-speaker target speech recognition (TS-ASR). To this end, we propose a unified dual-cue TS-ASR framework that introduces a pioneering single-model unification mechanism. Built upon a Conformer architecture, the proposed method deeply fuses lexical and speaker information through a shared cross-attention module and incorporates a negative sampling strategy to enhance the supervision of cue effectiveness. Experimental results on 30,000 mixed-speech utterances demonstrate that combining five-character text prompts with enrollment speech reduces the character error rate (CER) to 8.80%, significantly outperforming the text-only (17.32%) and enrollment-only (29.06%) baselines.

0 citationsRead paper

ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning

Sep 25, 2026

This study addresses the challenges of disconnected open-ended questioning from Bayesian networks and insufficient decision interpretability in automated medical consultation by proposing an uncertainty-aware framework. Methodologically, it introduces a novel automated pipeline that integrates heterogeneous clinical narratives to construct a Disease-Symptom Bayesian Network (DSBN). By combining large language model-assisted knowledge extraction with uncertainty-quantified reasoning, the framework drives adaptive inquiry and diagnostic decision-making through dynamic posterior probability updates. Experimental results demonstrate that the proposed approach improves Top-1 and Top-3 diagnostic accuracy by over 20%. Furthermore, physician evaluations confirm that its explanation quality and diagnostic plausibility significantly surpass those of existing baselines.

0 citationsRead paper

Look before Transcription: End-to-End SlideASR with Visually-Anchored Policy Optimization

Oct 08, 2025

ASR systems struggle to accurately recognize domain-specific terminology in professional settings such as academic lectures. To address this, we introduce the SlideASR task—leveraging slide visual content to enhance speech transcription. Existing pipeline approaches are cumbersome and inefficient, while multimodal large language models (MLLMs) often degenerate into pure OCR systems. We propose Visually-Anchored Policy Optimization (VAPO): an MLLM-based framework integrating Chain-of-Thought reasoning and visual anchoring to jointly model OCR, ASR, and visual–acoustic consistency, augmented with a quadruple-reward reinforcement learning objective for end-to-end post-training. Experiments demonstrate substantial improvements in domain-term recognition accuracy, achieving state-of-the-art performance on both synthetic and real-world academic lecture data. Furthermore, we release SlideASR-Bench—the first benchmark dataset for slide-augmented ASR—featuring rich academic entities and fine-grained multimodal alignment annotations, advancing research in domain-specialized speech understanding.

0 citationsRead paper
Recent publications

Latest Papers

Unified Target-Speaker ASR with Text and Enrollment Speech Cues

Sep 27, 2026

This study addresses the insufficient complementarity caused by the separation of textual and enrollment speech cues in multi-speaker target speech recognition (TS-ASR). To this end, we propose a unified dual-cue TS-ASR framework that introduces a pioneering single-model unification mechanism. Built upon a Conformer architecture, the proposed method deeply fuses lexical and speaker information through a shared cross-attention module and incorporates a negative sampling strategy to enhance the supervision of cue effectiveness. Experimental results on 30,000 mixed-speech utterances demonstrate that combining five-character text prompts with enrollment speech reduces the character error rate (CER) to 8.80%, significantly outperforming the text-only (17.32%) and enrollment-only (29.06%) baselines.

0 citationsRead paper

ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning

Sep 25, 2026

This study addresses the challenges of disconnected open-ended questioning from Bayesian networks and insufficient decision interpretability in automated medical consultation by proposing an uncertainty-aware framework. Methodologically, it introduces a novel automated pipeline that integrates heterogeneous clinical narratives to construct a Disease-Symptom Bayesian Network (DSBN). By combining large language model-assisted knowledge extraction with uncertainty-quantified reasoning, the framework drives adaptive inquiry and diagnostic decision-making through dynamic posterior probability updates. Experimental results demonstrate that the proposed approach improves Top-1 and Top-3 diagnostic accuracy by over 20%. Furthermore, physician evaluations confirm that its explanation quality and diagnostic plausibility significantly surpass those of existing baselines.

0 citationsRead paper

Look before Transcription: End-to-End SlideASR with Visually-Anchored Policy Optimization

Oct 08, 2025

ASR systems struggle to accurately recognize domain-specific terminology in professional settings such as academic lectures. To address this, we introduce the SlideASR task—leveraging slide visual content to enhance speech transcription. Existing pipeline approaches are cumbersome and inefficient, while multimodal large language models (MLLMs) often degenerate into pure OCR systems. We propose Visually-Anchored Policy Optimization (VAPO): an MLLM-based framework integrating Chain-of-Thought reasoning and visual anchoring to jointly model OCR, ASR, and visual–acoustic consistency, augmented with a quadruple-reward reinforcement learning objective for end-to-end post-training. Experiments demonstrate substantial improvements in domain-term recognition accuracy, achieving state-of-the-art performance on both synthetic and real-world academic lecture data. Furthermore, we release SlideASR-Bench—the first benchmark dataset for slide-augmented ASR—featuring rich academic entities and fine-grained multimodal alignment annotations, advancing research in domain-specialized speech understanding.

0 citationsRead paper