Procedural Core: A Compact Recurrent Initialization for Vision Transformers

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high optimization costs associated with random initialization in Transformers and the limited reusability of existing programmatic pretraining approaches. To overcome these challenges, this work proposes a reusable, compact cyclic weight initialization strategy that encapsulates general inductive biases into an independent core. By integrating a cyclic Transformer architecture with programmatic data generation, the method leverages weight expansion techniques to enable parameter reuse across models of arbitrary scale and facilitate low-cost cross-domain transfer. Experimental results demonstrate that the proposed approach improves classification accuracy by 2.2% on ImageNet while significantly enhancing performance in image segmentation, object localization, and depth estimation. These findings establish a new paradigm for the efficient initialization of large-scale models.
πŸ“ Abstract
Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. We train a minimal recurrent transformer on procedural data, then expand its weights to initialize transformers of arbitrary width and depth. The resulting initialization improves performance on image classification, self-supervised visual learning (DINO), and modeling natural language (FineWeb-Edu) and code (CodeParrot). For image classification, expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization. Our analysis identifies recurrence as essential for learning compact weights that transfer across models. In ViTs, we localize a key benefit in the suppression of high-norm tokens that produces substantial improvements in zero-shot segmentation (ImageNet-S mAP 32.3 to 42.9), object localization (VOC07 CorLoc 9.9 to 18.4), and depth estimation (NYUv2 RMSE 1.104 to 0.998). This demonstrates that transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
Model Initialization
Procedural Data
Inductive Bias
Transferability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Procedural Core
Recurrent Initialization
Vision Transformers
Weight Expansion
High-norm Token Suppression
πŸ”Ž Similar Papers
No similar papers found.