🤖 AI Summary
Existing studies lack systematic methods to dissect representational commonalities and idiosyncrasies across diverse large vision models—especially those differing in architecture and training paradigms.
Method: We propose a biologically inspired, multi-dimensional representational analysis framework integrating Representational Similarity Analysis (RSA), Soft Matching, and Linear Predictivity, augmented by an improved Similarity Network Fusion (SNF) technique for cross-architectural representational similarity modeling.
Contribution/Results: Our framework uncovers, for the first time, cross-architectural convergence of self-supervised models in geometric structure, unit tuning properties, and linear decodability—revealing pronounced representational alignment between hybrid architectures and masked autoencoders. The resulting robust “representation fingerprints” significantly improve model-family discrimination accuracy and expose previously unrecognized inter-model associations. Collectively, these findings establish a new paradigm for understanding how architectural inductive biases and training objectives jointly shape computational strategies in vision models.
📝 Abstract
Large vision models differ widely in architecture and training paradigm, yet we lack principled methods to determine which aspects of their representations are shared across families and which reflect distinctive computational strategies. We leverage a suite of representational similarity metrics, each capturing a different facet-geometry, unit tuning, or linear decodability-and assess family separability using multiple complementary measures. Metrics preserving geometry or tuning (e.g., RSA, Soft Matching) yield strong family discrimination, whereas flexible mappings such as Linear Predictivity show weaker separation. These findings indicate that geometry and tuning carry family-specific signatures, while linearly decodable information is more broadly shared. To integrate these complementary facets, we adapt Similarity Network Fusion (SNF), a method inspired by multi-omics integration. SNF achieves substantially sharper family separation than any individual metric and produces robust composite signatures. Clustering of the fused similarity matrix recovers both expected and surprising patterns: supervised ResNets and ViTs form distinct clusters, yet all self-supervised models group together across architectural boundaries. Hybrid architectures (ConvNeXt, Swin) cluster with masked autoencoders, suggesting convergence between architectural modernization and reconstruction-based training. This biology-inspired framework provides a principled typology of vision models, showing that emergent computational strategies-shaped jointly by architecture and training objective-define representational structure beyond surface design categories.