Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过重用语音识别编码器状态与文本嵌入融合的方法,提高长文稿中基于查询的话题定位准确性,尤其在结构化或半结构化语音上效果显著。
📝 Abstract
Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.
Problem

Research questions and friction points this paper is trying to address.

query-conditioned topic localization
transcripts
sentence span
speech representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Speech Representations
Query-Conditioned Topic Localization
ASR Encoder States
Sentence-Level Representations
💼 Related Jobs
No related jobs found.
S
Steffen Freisinger
Technische Hochschule Nürnberg Georg Simon Ohm
P
Philipp Seeberger
Technische Hochschule Nürnberg Georg Simon Ohm
T
Thomas Ranzenberger
Technische Hochschule Nürnberg Georg Simon Ohm
Tobias Bocklet
Tobias Bocklet
Technische Hochschule Nürnberg & Intel Labs
Automatic Speech ProcessingMachine LearningDeep LearningArtificial Intelligence
Korbinian Riedhammer
Korbinian Riedhammer
Technische Hochschule Nürnberg
speech recognitionatypical speechpathologic speech