🤖 AI Summary
To address the limited expressive power and interpretability of MLP-Mixer in fine-grained image feature extraction, this paper proposes KAN-Mixers—the first pure-MLP vision model integrating Kolmogorov–Arnold Networks (KANs) into the Mixer architecture. Our method replaces conventional fully connected layers with differentiable spline-parameterized KAN layers and incorporates channel- and spatial-mixing mechanisms, yielding an end-to-end trainable, highly interpretable hybrid architecture. The key contribution is the first deep integration of KANs with the Mixer paradigm, enhancing feature modeling capacity while preserving the conceptual simplicity of MLP-based models. Extensive experiments demonstrate that KAN-Mixers achieve 90.30% and 69.80% test accuracy on Fashion-MNIST and CIFAR-10, respectively—substantially outperforming standard MLPs, MLP-Mixers, and baseline KAN models. These results validate both the effectiveness and superior generalization capability of our approach.
📝 Abstract
Due to their effective performance, Convolutional Neural Network (CNN) and Vision Transformer (ViT) architectures have become the standard for solving computer vision tasks. Such architectures require large data sets and rely on convolution and self-attention operations. In 2021, MLP-Mixer emerged, an architecture that relies only on Multilayer Perceptron (MLP) and achieves extremely competitive results when compared to CNNs and ViTs. Despite its good performance in computer vision tasks, the MLP-Mixer architecture may not be suitable for refined feature extraction in images. Recently, the Kolmogorov-Arnold Network (KAN) was proposed as a promising alternative to MLP models. KANs promise to improve accuracy and interpretability when compared to MLPs. Therefore, the present work aims to design a new mixer-based architecture, called KAN-Mixers, using KANs as main layers and evaluate its performance, in terms of several performance metrics, in the image classification task. As main results obtained, the KAN-Mixers model was superior to the MLP, MLP-Mixer and KAN models in the Fashion-MNIST and CIFAR-10 datasets, with 0.9030 and 0.6980 of average accuracy, respectively.