🤖 AI Summary
In neural-symbolic image classification, extracting symbolic rules from a pre-trained CNN leads to significant accuracy degradation due to information loss during post-hoc rule extraction.
Method: This paper proposes a class-specific sparse loss function that embeds differentiable, class-aware filter activation binarization into the end-to-end training pipeline, mitigating information loss at its source. The method jointly integrates CNN-based feature extraction, class-conditional sparse regularization, differentiable binarization approximation, and neural-symbolic rule distillation.
Contribution/Results: Compared to state-of-the-art approaches, our method achieves a 9% average accuracy gain, compresses the rule set size by 53%, and attains 97% of the original CNN’s accuracy—marking the first unified modeling paradigm that simultaneously delivers high accuracy and strong interpretability in neural-symbolic classification.
📝 Abstract
There has been significant focus on creating neuro-symbolic models for interpretable image classification using Convolutional Neural Networks (CNNs). These methods aim to replace the CNN with a neuro-symbolic model consisting of the CNN, which is used as a feature extractor, and an interpretable rule-set extracted from the CNN itself. While these approaches provide interpretability through the extracted rule-set, they often compromise accuracy compared to the original CNN model. In this paper, we identify the root cause of this accuracy loss as the post-training binarization of filter activations to extract the rule-set. To address this, we propose a novel sparsity loss function that enables class-specific filter binarization during CNN training, thus minimizing information loss when extracting the rule-set. We evaluate several training strategies with our novel sparsity loss, analyzing their effectiveness and providing guidance on their appropriate use. Notably, we set a new benchmark, achieving a 9% improvement in accuracy and a 53% reduction in rule-set size on average, compared to the previous SOTA, while coming within 3% of the original CNN's accuracy. This highlights the significant potential of interpretable neuro-symbolic models as viable alternatives to black-box CNNs.