Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing multimodal embedding models struggle to uniformly support text, images, video, and audio, often omitting audio or covering only a subset of modalities. This work proposes a unified four-modality embedding space built upon a frozen vision–language foundation model, augmented with a lightweight audio tower connector and modality-gated deep adapters, requiring no updates to the base model parameters. By aligning only audio with text, the approach enables cross-modal retrieval across all pairs—including audio–image—while fully preserving the original performance on pre-existing modalities. Both generations of the proposed model are trained within hours on a single GPU and achieve strong results in audio–text and audio–image retrieval. The code, model weights, and evaluation tools are publicly released.
📝 Abstract
A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio entirely, while audio-text retrieval is led by specialist systems that serve no other modality. We present the Fusion Embedding family, which adds audio to a frozen vision-language embedding base whose parameters are never updated: generation 1 (fusion-embedding-1) trains only a 16.4M-parameter connector between a frozen audio tower and the frozen base, and generation 2 (fusion-embedding-2) adds modality-gated deep adapters (44.2M parameters) whose branch never executes on text, image, or video inputs: their outputs are bit-for-bit those of the released base, verified after every training run. Because the base already binds text, images, and video, aligning audio to text alone makes audio-image retrieval emerge, with zero paired audio-visual training data. Alongside the recipe we map its design space with controlled negative results (rewriting training captions with an LLM, substituting a leaderboard-stronger audio tower, and widening the connector each reduce retrieval) and with training-protocol findings that we expect to transfer to any frozen decoder-LM embedding backbone. Both generations train in hours on a single GPU. Weights, code, and the evaluation harness are openly released.
Problem

Research questions and friction points this paper is trying to address.

multimodal embedding
audio-text retrieval
unified embedding space
cross-modal retrieval
frozen backbone
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fusion Embedding
unified embedding space
frozen backbone
modality-gated adapters
zero-shot audio-visual retrieval
🔎 Similar Papers
No similar papers found.