A Data-driven Typology of Vision Models from Integrated Representational Metrics

📅 2025-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing studies lack systematic methods to dissect representational commonalities and idiosyncrasies across diverse large vision models—especially those differing in architecture and training paradigms. Method: We propose a biologically inspired, multi-dimensional representational analysis framework integrating Representational Similarity Analysis (RSA), Soft Matching, and Linear Predictivity, augmented by an improved Similarity Network Fusion (SNF) technique for cross-architectural representational similarity modeling. Contribution/Results: Our framework uncovers, for the first time, cross-architectural convergence of self-supervised models in geometric structure, unit tuning properties, and linear decodability—revealing pronounced representational alignment between hybrid architectures and masked autoencoders. The resulting robust “representation fingerprints” significantly improve model-family discrimination accuracy and expose previously unrecognized inter-model associations. Collectively, these findings establish a new paradigm for understanding how architectural inductive biases and training objectives jointly shape computational strategies in vision models.

Technology Category

Computer Vision: Representation Learning for VisionMachine Learning: Deep Neural Architectures and Foundation ModelsCognitive Modeling & Cognitive Systems: (Computational) Cognitive Architectures

Application Category

Search and Retrieval-Augmented AI: Web query analysis, representation and understandingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Large vision models differ widely in architecture and training paradigm, yet we lack principled methods to determine which aspects of their representations are shared across families and which reflect distinctive computational strategies. We leverage a suite of representational similarity metrics, each capturing a different facet-geometry, unit tuning, or linear decodability-and assess family separability using multiple complementary measures. Metrics preserving geometry or tuning (e.g., RSA, Soft Matching) yield strong family discrimination, whereas flexible mappings such as Linear Predictivity show weaker separation. These findings indicate that geometry and tuning carry family-specific signatures, while linearly decodable information is more broadly shared. To integrate these complementary facets, we adapt Similarity Network Fusion (SNF), a method inspired by multi-omics integration. SNF achieves substantially sharper family separation than any individual metric and produces robust composite signatures. Clustering of the fused similarity matrix recovers both expected and surprising patterns: supervised ResNets and ViTs form distinct clusters, yet all self-supervised models group together across architectural boundaries. Hybrid architectures (ConvNeXt, Swin) cluster with masked autoencoders, suggesting convergence between architectural modernization and reconstruction-based training. This biology-inspired framework provides a principled typology of vision models, showing that emergent computational strategies-shaped jointly by architecture and training objective-define representational structure beyond surface design categories.
Problem

Research questions and friction points this paper is trying to address.

Developing a principled typology to classify diverse vision models
Identifying shared and distinctive representational features across model families
Integrating multiple similarity metrics to reveal computational strategy signatures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrated multiple representational similarity metrics for analysis
Applied Similarity Network Fusion for enhanced model separation
Revealed computational strategies beyond surface architecture categories
🔎 Similar Papers
No similar papers found.
J
Jialin Wu
Department of Computer Science and Engineering, UC San Diego
S
Shreya Saha
Department of Electrical and Computer Engineering, UC San Diego
Y
Yiqing Bo
Department of Computer Science and Engineering, UC San Diego
Meenakshi Khosla
Meenakshi Khosla
UC San Diego
Computational NeuroscienceArtificial IntelligenceVisionAuditionLanguage