Does Mechanistic Interpretability Transfer Across Data Modalities? A Cross-Domain Causal Circuit Analysis of Variational Autoencoders

📅 2026-03-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether causal mechanisms from variational autoencoders (VAEs) in the image domain can be effectively transferred to tabular data to enhance model interpretability. By extending a four-layer causal intervention framework and introducing three novel components—posterior-calibrated Causal Effect Strength (CES), path-specific activation patching, and Feature Group Disentanglement (FGD)—the authors systematically evaluate cross-modal transferability of causal structures across multiple VAE architectures and multimodal datasets. Results reveal that tabular VAEs exhibit approximately 50% lower modularity than their image counterparts; β-VAE shows a marked decline in CES on heterogeneous tabular data (0.043 vs. 0.133); and high-specificity interventions yield the best downstream prediction performance (r = 0.460, p < 0.001). This work provides the first quantitative evidence of modality-dependent differences in the causal interpretability of VAEs.

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyMachine Learning: Causal LearningReasoning under Uncertainty: Causality

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Although mechanism-based interpretability has generated an abundance of insight for discriminative network analysis, generative models are less understood -- particularly outside of image-related applications. We investigate how much of the causal circuitry found within image-related variational autoencoders (VAEs) will generalize to tabular data, as VAEs are increasingly used for imputation, anomaly detection, and synthetic data generation. In addition to extending a four-level causal intervention framework to four tabular and one image benchmark across five different VAE architectures (with 75 individual training runs per architecture and three random seed values for each run), this paper introduces three new techniques: posterior-calibration of Causal Effect Strength (CES), path-specific activation patching, and Feature-Group Disentanglement (FGD). The results from our experiments demonstrate that: (i) Tabular VAEs have circuits with modularity that is approximately 50% lower than their image counterparts. (ii) $β$-VAE experiences nearly complete collapse in CES scores when applied to heterogeneous tabular features (0.043 CES score for tabular data compared to 0.133 CES score for images), which can be directly attributed to reconstruction quality degradation (r = -0.886 correlation coefficient between CES and MSE). (iii) CES successfully captures nine of eleven statistically significant architecture differences using Holm--Šidák corrections. (iv) Interventions with high specificity predict the highest downstream AUC values (r = 0.460, p < .001). This study challenges the common assumption that architectural guidance from image-related studies can be transferred to tabular datasets.
Problem

Research questions and friction points this paper is trying to address.

Mechanistic Interpretability
Data Modalities
Variational Autoencoders
Causal Circuit
Tabular Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Circuit Analysis
Variational Autoencoders
Mechanistic Interpretability
Tabular Data
Causal Effect Strength
Dip Roy
Dip Roy
Indian Institute of Technology, Patna
Explainable AIMechanistic Interpretibility
Rajiv Misra
Rajiv Misra
Professor of Computer Science at IIT Patna
wireless networks
S
Sanjay Kumar Singh
Department of Computer Science, Rajarshi School of Management Technology, Varanasi, India
A
Anisha Roy
Department of Electronics and Communication Engineering, Jaypee Institute of Information Technology, Noida