Improving Interpretability and Accuracy in Neuro-Symbolic Rule Extraction Using Class-Specific Sparse Filters

📅 2025-01-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In neural-symbolic image classification, extracting symbolic rules from a pre-trained CNN leads to significant accuracy degradation due to information loss during post-hoc rule extraction. Method: This paper proposes a class-specific sparse loss function that embeds differentiable, class-aware filter activation binarization into the end-to-end training pipeline, mitigating information loss at its source. The method jointly integrates CNN-based feature extraction, class-conditional sparse regularization, differentiable binarization approximation, and neural-symbolic rule distillation. Contribution/Results: Compared to state-of-the-art approaches, our method achieves a 9% average accuracy gain, compresses the rule set size by 53%, and attains 97% of the original CNN’s accuracy—marking the first unified modeling paradigm that simultaneously delivers high accuracy and strong interpretability in neural-symbolic classification.

Technology Category

Computer Vision: Visual Reasoning & Symbolic RepresentationsMachine Learning: Neuro-Symbolic LearningCognitive Modeling & Cognitive Systems: Symbolic Representations

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
There has been significant focus on creating neuro-symbolic models for interpretable image classification using Convolutional Neural Networks (CNNs). These methods aim to replace the CNN with a neuro-symbolic model consisting of the CNN, which is used as a feature extractor, and an interpretable rule-set extracted from the CNN itself. While these approaches provide interpretability through the extracted rule-set, they often compromise accuracy compared to the original CNN model. In this paper, we identify the root cause of this accuracy loss as the post-training binarization of filter activations to extract the rule-set. To address this, we propose a novel sparsity loss function that enables class-specific filter binarization during CNN training, thus minimizing information loss when extracting the rule-set. We evaluate several training strategies with our novel sparsity loss, analyzing their effectiveness and providing guidance on their appropriate use. Notably, we set a new benchmark, achieving a 9% improvement in accuracy and a 53% reduction in rule-set size on average, compared to the previous SOTA, while coming within 3% of the original CNN's accuracy. This highlights the significant potential of interpretable neuro-symbolic models as viable alternatives to black-box CNNs.
Problem

Research questions and friction points this paper is trying to address.

Neuro-Symbolic Models
CNN Rule Extraction
Accuracy Degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interpretable Neural-Symbolic Models
Filter Simplification in CNNs
Accuracy-Interpretability Trade-off