Speech Block Influence: Component-Specific Layer Scoring for Pruning Speech LLMs

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the poor transferability of existing scoring metrics in layer pruning for large speech models, which stems from architectural heterogeneity and multimodal sequences. To this end, we propose a component-specific layer importance scoring framework (SBI) that introduces the first dedicated evaluation system tailored for speech LLMs. Specifically, SBI-Enc assesses the impact of encoder adapter outputs, while SBI-Dec computes inter-layer similarity within the decoder using exclusively text tokens to optimize the pruning strategy. Experimental results demonstrate that SBI significantly enhances model robustness under high pruning ratios. Furthermore, our findings confirm that purely textual calibration serves as an efficient and cost-effective alternative to multimodal approaches.
📝 Abstract
Speech LLMs are costly to deploy in resource-constrained settings. Layer pruning can cut this cost, but existing scoring metrics transfer poorly to speech LLMs: they assume a decoder-only architecture with homogeneous token sequences, whereas speech LLMs add encoder and adapter components and process multimodal sequences. We propose Speech Block Influence (SBI), the first layer-importance scoring framework designed for speech LLM pruning that consists of two component-specific scores: SBI-Enc measures the effect of encoder-layer removal at the adapter's output to better reflect downstream impact; SBI-Dec measures layer-wise input-output similarity over text-token positions only to avoid audio-token dominance. Across three speech LLMs, SBI improves pruning robustness, with stronger encoder performance at higher pruning rates and more reliable decoder layer selection by scoring text tokens rather than the audio-dominated full sequence. We further find that text-only calibration yields decoder rankings highly correlated with those from speech-text calibration, suggesting a cheaper alternative to measure decoder layer importance.
Problem

Research questions and friction points this paper is trying to address.

Speech LLMs
Layer Pruning
Scoring Metrics
Multimodal Sequences
Model Compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech Block Influence
Layer Pruning
Speech LLMs
Component-Specific Scoring
Model Compression
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.