Steering Speech-Language Models: Training-Free Task Specialization via Contrastive Activation Addition

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of training-free inference-time task control methods for large speech models by introducing a pioneering training-free steering mechanism. Specifically, it proposes a contrastive activation addition protocol that extracts steering vectors from limited labeled data. By integrating vector arithmetic with script-normalized direction techniques, these vectors are injected into the representation space during inference to modulate model behavior. Experimental results demonstrate that, when combined with prompt engineering, the proposed approach significantly outperforms pure prompting baselines on tasks such as automatic speech recognition and emotion recognition. Furthermore, the method exhibits strong out-of-domain transferability and effectively enforces alignment with target language scripts while optimizing transcription quality.
📝 Abstract
Activation steering has proven effective for controlling the behavior of Large Language Models (LLMs) at inference time, but its application to SpeechLLMs remains new, and training-free steering approaches for such models are still largely unexplored. We propose a training-free Contrastive Activation Addition (CAA) protocol that derives steering vectors for common speech tasks (e.g. transcription) in SpeechLLMs from a small number of labeled utterances. We showcase that adding these vectors in the representation space, at inference time, enforces better the targeted speech task. We further show that, when combined with prompting, these vectors yield to consistent improvement over prompting alone on most evaluated tasks such as Automatic Speech Recognition (ASR) or Emotion Recognition (ER), and transfer to out-of-domain data. We additionally demonstrate the usefulness of script-normalization directions to enforce the target script of a specific language.
Problem

Research questions and friction points this paper is trying to address.

SpeechLLMs
Activation Steering
Training-Free
Task Specialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Activation Addition
Training-Free Steering
Speech-Language Models
Activation Steering
Task Specialization
🔎 Similar Papers
No similar papers found.