CI-JEPA: A Counterfactual Analysis of Latent Representations in Joint-Embedding Predictive Architectures for Self-Supervised Learning

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of visual intervention response modeling and counterfactual robustness analysis in I-JEPA by proposing CI-JEPA. The method introduces a counterfactual intervention-aware mechanism that quantifies semantic sensitivity and nuisance invariance by predicting representational shifts between original and modified images, thereby establishing a representation robustness evaluation framework grounded in selective sensitivity. Experimental evaluations employ joint embedding prediction, image masking interventions, and frozen-encoder linear probing techniques. Achieving 78.14% accuracy on the Flowers102 dataset, the results demonstrate that CI-JEPA significantly outperforms baseline models in relative semantic selectivity.
📝 Abstract
Self-supervised visual representation learning learns useful features without manual annotations during representation training. The image-based joint-embedding predictive architecture (I-JEPA) predicts latent representations of masked image regions, but its objective does not explicitly model responses to specified visual interventions. We introduce CI-JEPA, a counterfactual intervention-aware extension that learns to predict the representation change $\Delta Z = Z_{\mathrm{CF}} - Z$ between an original image and a modified counterpart. We assess representation robustness through selective sensitivity: stronger responses to task-relevant semantic changes than to nuisance changes. Experiments on Flowers102 use flower-center occlusion as a candidate semantic intervention and background blur and tint as candidate nuisance interventions. With frozen-encoder linear probing, CI-JEPA achieves a best validation accuracy of 78.14\%, compared with 77.55\% for both the pretrained ViT-B/16 and the I-JEPA baseline, a gain of 0.59 percentage points. The reported mean $L_2$ representation changes are 4.42 for center occlusion, 3.48 for background tint, and 2.83 for background blur. This ordering is consistent with relative semantic selectivity for the evaluated interventions, rather than complete nuisance invariance. The accuracy comparison is complementary and does not establish improved robustness over the baselines. These controlled image modifications provide a framework for studying intervention-induced changes in JEPA representations; they do not establish causal feature discovery or robustness to all visual changes.
Problem

Research questions and friction points this paper is trying to address.

Self-supervised learning
Joint-embedding predictive architecture
Counterfactual intervention
Representation robustness
Visual representation learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Intervention
Self-Supervised Learning
Joint-Embedding Predictive Architecture
Latent Representation
Selective Sensitivity
🔎 Similar Papers
M
Mintu Dutta
Department of Information and Communication Technology, Pandit Deendayal Energy University, Gandhinagar, Gujarat, India
R
Ritesh Vyas
Department of Electrical and Electronics Engineering, Birla Institute of Technology & Science, Pilani, Rajasthan, India
Mohendra Roy
Mohendra Roy
Pandit Deendayal Energy University (PDEU), India
AIMachine LearningBio-PhotonicsBio-ImagingBioSensors