Projector Is All You Train

📅 2026-08-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了仅训练投影器而非整个多模态大语言模型骨干以适应新模ality的方法,证明此方法有效且能保持原有能力,同时提高训练效率。
📝 Abstract
The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it to a new modality. Through experiments on 3D MLLMs, we find that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and our jointly trained MLLMs with the same encoder and backbone. We also show that joint training leads to undesirable drift in existing capabilities of the language model, which projector-only training avoids by definition. Furthermore, projector-only training has approximately twice the training sample throughput of joint training. We validate our findings across different language model backbones via 3D classification and captioning benchmarks as well as standard benchmarks evaluating language, vision, and spatial reasoning capabilities.
Problem

Research questions and friction points this paper is trying to address.

multimodal large language model
fine-tuning
projector
Innovation

Methods, ideas, or system contributions that make the work stand out.

Projector-Only Training
Multimodal Large Language Model (MLLM)
3D MLLMs
Sample Throughput
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Nyx Iskandar
Ramen VR
S
Saathvik Selvan
University of California, Berkeley
S
Slater Victoroff
iph.so