Virtual Encoders in Multimodal Transformers

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在没有专用感知编码器的情况下,多模态Transformer如何通过内部早期到中期层构建任务可用的感知表示,提出了一种称为虚拟编码器的新结构。
📝 Abstract
Multimodal language models traditionally rely on dedicated perceptual encoders to construct task-usable representations. More integrated architectures have recently emerged, which instead expose the shared transformer to lightly projected patches, audio frames, or discrete visual tokens. Where does this encoding happen when such representations are not provided? We find that the transformer can internalize this missing computation, constructing task-usable perceptual representations within its own early-to-middle layers before the downstream language model. We call this computational structure a Virtual Encoder. Across linear probing, similarities to perceptual encoders, and causal analyses, we identify signatures of this structure in models that receive perceptual tokens without continuous encoder-derived features. These analyses also suggest that the boundary between perception and language processing need not coincide within an architectural module. Instead, encoder-like computation can emerge as a functional regime within a shared transformer, providing a new perspective for understanding where and how multimodal models process perception.
Problem

Research questions and friction points this paper is trying to address.

Virtual Encoder
Multimodal Transformers
Perceptual Representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Virtual Encoder
Multimodal Transformers
Perceptual Representations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.