VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing visual backbones, whose memory and learning rules remain fixed after training and lack the capacity for dynamic self-modification conditioned on image context. This work proposes the first general-purpose self-modifying visual backbone, achieving co-evolution of memory and learning through nested learning. Methodologically, it introduces a novel self-referential update mechanism coupling five memories, and derives a stability-matched step-size control scheme to ensure non-expansive dynamics. Furthermore, 2D feature processing is optimized via soft upper-bound injection, spectral clamping, and a four-directional scanning chunking strategy. Experiments demonstrate that this architecture achieves competitive performance on ImageNet-1K, COCO, and ADE20K, validating the practical value of the self-modifying paradigm.
📝 Abstract
Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an image. Across this progression, visual computation has become increasingly adaptive to each input, yet the rules governing that adaptation remain largely prescribed by the trained backbone. We introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system, in which what the model remembers and how it learns co-evolve within an image. Building on the self-referential construction of Nested Learning (NL), VisionHOPE realizes this co-evolution through five coupled memories that store content, generate key and value representations, and govern learning rate and retention. These memories evolve jointly as visual context accumulates along each scan. However, directly applying the unconstrained self-referential update to a visual backbone leads to instability. We therefore derive a stability-matched step-size control scheme that combines a soft cap on self-referential injection with a spectral clamp on the retained memory transition, and prove that the resulting memory dynamics are non-expansive along each scan. For two-dimensional feature maps, we adapt NL's chunk formulation by aligning chunks with image rows and columns across four directional scans. The proposed VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K, establishing self-modifying learning systems as a practical foundation for general-purpose visual backbones. The code is available at https://github.com/PSRben/VisionHOPE.
Problem

Research questions and friction points this paper is trying to address.

visual backbone
self-modifying learning system
training instability
adaptive computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Modifying Learning System
Coupled Memories
Stability-Matched Step-Size Control
Nested Learning
Visual Backbone
🔎 Similar Papers