🤖 AI Summary
ReLU-KAN suffers from limited feature representation capacity in multi-input scenarios due to ReLU’s hard thresholding of negative inputs. To address this negative-value sensitivity, we propose Activation Function-driven Kolmogorov–Arnold Networks (AF-KAN), which replace fixed basis functions (e.g., B-splines or ReLU) with learnable, diverse activation functions—including multivariate compositions and piecewise modeling—thereby enhancing expressivity and robustness. Methodologically, AF-KAN integrates attention-guided parameter sparsification, adaptive grid optimization, and batch normalization to achieve efficient parameter compression. Experiments demonstrate that AF-KAN significantly outperforms comparably sized MLPs, ReLU-KAN, and other KAN variants on image classification tasks; it attains comparable accuracy using only 1/6–1/10 the parameters of baseline KANs. The implementation is publicly available.
📝 Abstract
Kolmogorov-Arnold Networks (KANs) have inspired numerous works exploring their applications across a wide range of scientific problems, with the potential to replace Multilayer Perceptrons (MLPs). While many KANs are designed using basis and polynomial functions, such as B-splines, ReLU-KAN utilizes a combination of ReLU functions to mimic the structure of B-splines and take advantage of ReLU's speed. However, ReLU-KAN is not built for multiple inputs, and its limitations stem from ReLU's handling of negative values, which can restrict feature extraction. To address these issues, we introduce Activation Function-Based Kolmogorov-Arnold Networks (AF-KAN), expanding ReLU-KAN with various activations and their function combinations. This novel KAN also incorporates parameter reduction methods, primarily attention mechanisms and data normalization, to enhance performance on image classification datasets. We explore different activation functions, function combinations, grid sizes, and spline orders to validate the effectiveness of AF-KAN and determine its optimal configuration. In the experiments, AF-KAN significantly outperforms MLP, ReLU-KAN, and other KANs with the same parameter count. It also remains competitive even when using fewer than 6 to 10 times the parameters while maintaining the same network structure. However, AF-KAN requires a longer training time and consumes more FLOPs. The repository for this work is available at https://github.com/hoangthangta/All-KAN.