speech-to-llm projection

Designs and trains projection modules (e.g., linear projectors, adapter modules, or mapper networks) that map speech-encoder outputs into the embedding space expected by a large language model so the LLM can accept and reason over speech-derived representations. Analyzes and optimizes the alignment by measuring and minimizing embedding-space distance, reducing modality representation gap, and validating cross-modal integration via downstream-task performance.

speech-to-llmprojection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high computational cost, susceptibility to overfitting, and limited generalization of existing speech large language models (LLMs), which typically rely on expensive end-to-end training. The authors propose an efficient and scalable method for mapping speech embeddings into LLM-compatible representations through a novel task-agnostic pretraining followed by lightweight fine-tuning paradigm. Specifically, a speech-to-text embedding projector is first pretrained without any involvement of the target LLM—enabling training on inexpensive hardware—and subsequently adapted to the target LLM via only approximately 1,000 steps of instruction tuning. This approach drastically reduces computational requirements while supporting both task-agnostic and task-specific configurations. In evaluations on speech translation and spoken question answering, the task-agnostic variant matches the performance of the best IWSLT25 model, while the task-specific variant surpasses existing methods using less data and computational resources.

generalizationinstruction tuningoverfitting

Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models

Oct 02, 2025
TÒ
Tolúld{o}pé Ògúnrèmí
🏛️ Stanford University | Toyota Technological Institute at Chicago

How do modality adapters (MAs) in speech large models (SLMs) transform encoder outputs into language-model-compatible representations? Method: We systematically analyze MA intermediate representations across prominent SLMs—including SALMONN, Qwen2-Audio, and Phi-4-Multimodal-Instruct—using Whisper-family encoders and a nearest-neighbor decoding token analysis approach. Contribution/Results: We identify, for the first time, two fundamentally distinct representation strategies employed by MAs: (1) semantic representations mediated through English, and (2) phonetic representations encoded directly in English tokens. These strategies are determined by the pretraining objective of the underlying speech encoder (e.g., ASR versus self-supervised learning). This dichotomy provides a unified explanation for observed cross-lingual generalization disparities across SLMs and establishes a theoretical foundation for MA architecture design and multilingual speech understanding modeling.

Compares phonetic versus semantic representation strategies across three SLMsExplores how speech encoder training affects multilingual processing capabilitiesInvestigates how modality adapters transform speech representations in language models

This work addresses the structural misalignment between speech encoders and large language models caused by inconsistencies between language-specific and language-agnostic representations. To resolve this, the authors propose incorporating a translation objective into speech encoder pretraining, leveraging multilingual speech translation tasks to encourage the learning of language-independent speech representations. This study presents the first systematic demonstration of the critical role of translation-based pretraining in building effective speech-augmented large language models, establishing a new paradigm wherein translation-driven alignment enhances cross-lingual representation learning. Experimental results show that the proposed approach significantly improves performance across multiple downstream speech-language tasks, effectively strengthening both cross-modal fusion and linguistic generalization capabilities.

language-agnostic spacelanguage-specific representationsLarge Language Model

OneLLM: One Framework to Align All Modalities with Language

Dec 06, 2023
JH
Jiaming Han
🏛️ The Chinese University of Hong Kong | Shanghai Artificial Intelligence Laboratory

Current multimodal large language models (MLLMs) exhibit strong modality-specificity and weak generalization, hindering unified processing of heterogeneous non-linguistic data. To address this, we propose OneLLM—a unified framework enabling end-to-end alignment from eight distinct modalities—images, audio, video, point clouds, depth maps, normal maps, IMU signals, and fMRI—to language. Our method introduces: (1) modality-agnostic representations via a unified multimodal encoder and a universal projection module (UPM); (2) the first progressive cross-modal alignment paradigm spanning all eight modalities; and (3) the first large-scale multimodal instruction dataset covering seven non-text modalities and containing 2 million samples. Evaluated on 25 cross-modal benchmarks—including description, question answering, and reasoning tasks—OneLLM achieves state-of-the-art or near-state-of-the-art performance. The code, models, dataset, and interactive demo are fully open-sourced.

Intermodality GeneralizationLarge Language ModelsMultimodal

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Aug 21, 2024
YY
Yuanyang Yin
🏛️ University of Science and Technology of China | Peking University | Kuaishou Technology

Current multimodal large language models (MLLMs) struggle to achieve token-level semantic alignment between vision and language under image-level alignment, limiting the visual understanding and reasoning capabilities of small-scale LLMs. To address this, we propose the first supervised token-level embedding alignment mechanism: leveraging vision–language priors from pre-trained models (e.g., CLIP), we distill cross-modal alignment knowledge via contrastive learning and design a lightweight adapter to precisely map visual tokens into the LLM’s embedding space. Our method requires no additional training data or inference overhead. Evaluated on multiple visual question answering and reasoning benchmarks, it improves performance of small-scale MLLMs by 3.2–5.7 points, while significantly enhancing model interpretability and generalization.

Addressing suboptimal modality integration in multimodal systemsEnhancing cross-modal understanding for smaller language modelsImproving token-level visual-textual alignment in MLLMs

Latest Papers

What's happening recently
View more

This study investigates the root causes of the performance gap between speech and text inputs in end-to-end spoken large language models. Through cross-layer centered kernel alignment (CKA) analysis and speech–text token alignment, combined with evaluations on SpeechMMLU and VoiceBench BBH across four open-source models, the authors find that speech representations form broad alignment bands across layers. This suggests that the modality gap stems primarily from the difficulty of compressing redundant acoustic information into stable high-level semantic representations, rather than mere distributional shift. Furthermore, statistical calibration at the input layer proves ineffective or even detrimental, reinforcing the structural stability of the observed alignment patterns. These findings provide theoretical grounding for future modeling approaches operating at the token or temporal granularity.

end-to-end modelsmodality gapspeech representation

研究探讨了仅训练投影器而非整个多模态大语言模型骨干以适应新模ality的方法,证明此方法有效且能保持原有能力,同时提高训练效率。

fine-tuningmultimodal large language modelprojector

This study addresses the “modality gap”—the significant performance disparity between speech and text inputs—in Large Speech-Language Models (LSLMs). To overcome the limitation of existing alignment mechanisms, which lack fine-grained characterization, we conduct the first systematic empirical analysis and propose a representation-level quantitative metric, the Alignment Path Score. We further design a token-level representation intervention method based on angular projection and length normalization. Using cosine similarity, Euclidean distance, hierarchical representation analysis, and alignment path modeling, we reveal a strong correlation between the modality gap and representation similarity. Experiments demonstrate that our approach substantially narrows the performance gap under speech input, improving model accuracy across multiple downstream tasks. Our work establishes an interpretable and intervenable paradigm for speech–text cross-modal alignment, advancing both diagnostic understanding and controllable representation learning in LSLMs.

Analyzing performance gap between speech and text inputs in large language modelsDeveloping interventions to improve speech input correctness through alignment mechanismsInvestigating representation alignment patterns across different model layers

This work addresses the inefficiency of cross-modal fusion caused by representational discrepancies among audio, visual, and textual modalities by proposing a semantic alignment method based on optimal transport (OT). For the first time, OT is integrated into a large language model–driven audio-visual speech recognition (LLM-AVSR) framework. Leveraging LLaMA3.2-3B language embeddings as anchors, the approach employs OT-derived coupling matrices to generate soft pseudo-labels that guide contrastive learning, thereby explicitly aligning features from Whisper’s audio encoder and AV-HuBERT’s visual encoder in a language-anchored semantic space. Evaluated on the LRS3-TED benchmark, the method significantly outperforms strong existing baselines and achieves state-of-the-art performance across varying signal-to-noise ratios, substantially enhancing recognition robustness and semantic consistency in challenging acoustic environments.

audio-visual speech recognitioncross-modal integrationlarge language model

LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures

Sep 10, 2025
HH
Hai Huang
🏛️ Atlassian | NYU | Brown University

In vision, embedding-space prediction (e.g., Joint Embedding Predictive Architecture, JEPA) substantially outperforms input-space reconstruction—yet current large language models (LLMs) still rely predominantly on token-level reconstruction for pretraining and fine-tuning, leading to weak generalization and susceptibility to overfitting. This work introduces LLM-JEPA, the first JEPA-based framework for language modeling: it jointly trains an encoder and a predictor in the embedding space, eliminating token-level reconstruction objectives. LLM-JEPA is architecture-agnostic, successfully integrated with Llama3, Gemma2, OpenELM, and Olmo. Empirical evaluation across diverse benchmarks—including NL-RX, GSM8K, Spider, and RottenTomatoes—demonstrates consistent superiority over standard training objectives, with marked improvements in generalization and robustness against overfitting. This study bridges a critical gap by extending JEPA to language modeling, establishing a novel, principled paradigm for LLM training grounded in predictive representation learning.

Combining Joint Embedding Predictive Architectures with large language modelsDeveloping robust pretraining and finetuning methods resistant to overfittingImproving language model training using vision-inspired embedding-space objectives

Hot Scholars

YH

Yeongwoo Hwang

Graduate Student, Harvard University
Quantum Complexity Theory
LA

Li An

Solon & Martha Dixon Endowed Professor, Auburn University
human-environmentlandscape ecologyGISciencecomplex systems
SZ

Shaofeng Zhang

Southern University of Science and Technology
Learn to Optimize
JB

John Bostanci

PhD Student, Columbia University
Quantum ComputingComputer Science
QF

Qi Fan

Nanjing University