How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of scaling law research for encoder-free multimodal models, which hinders assessing their viability as alternatives to conventional vision encoders. We systematically compare the scaling laws and compute-optimal allocation strategies of both architectures, incorporating end-to-end pixel learning and expert routing techniques. Our analysis reveals that as computational resources increase, the advantages of pretrained visual priors diminish, enabling language models to adaptively assume visual tasks through early-layer processing and expert routing. Extrapolations indicate that at a scale of 10²² FLOPs, encoder-free models are projected to match the performance of encoder-based counterparts. These findings establish encoder-free architectures as a highly promising direction for future multimodal modeling.
📝 Abstract
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Encoder-Free
Scaling Laws
Visual Encoder
Multimodal Pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

Encoder-Free MLLMs
Scaling Laws
Multimodal Pretraining
Compute-Optimal Allocation
Vision-Specific Adaptation
🔎 Similar Papers
2024-08-29arXiv.orgCitations: 7
💼 Related Jobs
No related jobs found.
Lin Chen
Lin Chen
CASIA
Computer Vision
B
Bolin Ni
Foundation Model Department, Tencent
Q
Qi Yang
Foundation Model Department, Tencent
L
Lan Jiang
Foundation Model Department, Tencent
Kun Ding
Kun Ding
CASIA
CVMultimodal
Xiaoran Fan
Xiaoran Fan
Fudan University
H
Hower Yang
Foundation Model Department, Tencent
Y
Ying Wang
CASIA; UCAS
Shiming Xiang
Shiming Xiang
National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences
Distance Metric LearningSemi-supervised LearningManifold LearningRegressionFeature Selection