🤖 AI Summary
This study addresses the challenge of adapter placement in continual learning, which constrains the trade-off between storage efficiency and model performance. To this end, this work introduces hierarchical plasticity mechanisms from the human visual cortex into AI architecture design for the first time, proposing an fMRI readout-based LS-B method to guide layer-wise specialization in Vision Transformers (ViTs) and enable backpropagation-free adaptive capacity allocation. The experiments employ a ViT-B/16 backbone, BiLoRA adapters, and weight spectral analysis techniques. Validated on Split ImageNet-R, the proposed approach demonstrates the optimality of intermediate-depth adapter placement, reducing storage requirements by 60% while maintaining accuracy degradation below 1.5 percentage points. These findings establish a novel paradigm for neuroscience-inspired parameter-efficient learning.
📝 Abstract
Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspectives. Algorithmically, training all contiguous four-block placements yields an inverted U: final accuracy peaks at intermediate depth and varies by up to 3.5 percentage points (pp), while inexpensive criteria based on weight spectra or activation statistics favor the deepest blocks. From neuroscience, the hierarchical organization and intermediate-stage plasticity of the visual cortex motivate us to ask whether a measurement taken outside the learner can guide layer specialization without placement search. LS-B observes the first tasks through a frozen fMRI encoding model of twelve human visual areas and commits task-specific capacity once to the blocks whose readouts vary most across tasks relative to their stable structure. Across three ViT-B/16 backbones, LS-B yields stable, backbone-specific allocations. On the two backbones with placement search, AugReg and iBOT, the selected blocks overlap the intermediate-depth region identified by search. Under matched storage and observation budgets, the selected blocks outperform the shallowest and deepest four-block configurations. On Split ImageNet-R, LS-B uses 60% of full-BiLoRA adapter storage while remaining within 1.5 pp of its final accuracy. The allocation requires no labels or backpropagation, adds under 0.6% runtime, and exhibits backbone-specific cortical signatures.