Fitting Vision Adapters at Frontier Scales

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear limits and scaling laws governing visual capability injection via lightweight projectors when large language model (LLM) weights remain frozen. Employing GLM—a text-only LLM—as the base model and Kimi as the vision encoder, this work proposes a reproducible training recipe utilizing a frontier-scale visual adapter comprising only 50M parameters. It systematically investigates how multimodal capabilities scale with LLM size under frozen pretrained weights. The primary contribution lies in successfully endowing a purely linguistic model with visual functionality while revealing the boundaries of such capabilities at scale. Comprehensive evaluations on the MMMU-Pro and BLINK benchmarks validate both overall performance and task-specific visual proficiency, thereby establishing a new paradigm for efficiently constructing multimodal large language models.
📝 Abstract
Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.
Problem

Research questions and friction points this paper is trying to address.

Vision Adapters
Multimodal Learning
Frontier Scales
Language Model Scaling
Visual Capabilities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Adapters
Multimodal Learning
Frontier Scales
Projector Training
Scaling Laws