Score
Design and implement feedforward neural modules using the KAN (kan feedforward / kan-ffn) architecture as a replacement for standard MLP blocks, constructing layers that model nonlinear mappings from latent representations. Build, tune, and analyze these KAN-based feedforward transforms to produce fine-grained heterogeneous mappings and to adapt feedforward computations to behavior- or task-specific signals.
This work addresses the poor interpretability and low parameter efficiency of traditional multilayer perceptrons (MLPs) in nonlinear modeling. Methodologically, grounded in the Kolmogorov–Arnold representation theorem, it replaces fixed activation functions with learnable piecewise spline functions, introducing a novel “learnable activations-as-weights” paradigm, and develops a symbol-numerical co-optimization framework with differentiable grid refinement. Contributions include: (i) the first comprehensive survey of Kolmogorov–Arnold Networks (KANs); (ii) theoretical and empirical identification of their structural advantages in function approximation, intrinsic interpretability, and parameter efficiency; and (iii) experimental validation showing KANs significantly outperform MLPs in fitting accuracy, generalization, and few-shot learning. To foster reproducibility and adoption, we publicly release a unified training framework, advancing the field of interpretable neural modeling.
This work theoretically compares Kolmogorov–Arnold Networks (KANs) and multilayer perceptrons (MLPs) in terms of expressive power and spectral bias, specifically assessing KANs’ potential as MLP alternatives with improved efficiency in modeling high-frequency components. Method: Leveraging tools from approximation theory, spectral analysis, and spline interpolation, the study establishes rigorous theoretical characterizations of both architectures. Contribution/Results: We prove for the first time that any MLP can be exactly represented by a KAN with strictly lower parametric complexity. We reveal that KANs’ learnable spline grids inherently mitigate low-frequency bias—enhancing fidelity to high-frequency signal components. Theoretically, KANs match or exceed MLPs in expressivity; moreover, large-grid KANs achieve substantial parameter savings for specific target functions. Empirical validation confirms weaker spectral bias and significantly improved high-frequency approximation accuracy compared to MLPs.
Multi-channel MLPs suffer from low training efficiency. Method: This work establishes, for the first time, a rigorous geometric and algebraic equivalence between free-knot B-spline Kolmogorov–Arnold networks (KANs) and a specific class of multi-channel MLPs. Leveraging this correspondence, we propose a hierarchical geometric refinement strategy along the channel dimension and design an end-to-end preconditioned training framework that jointly optimizes spline knot positions and network weights. The method synergistically integrates KAN’s localized basis function property with multi-channel MLPs’ parallel representational capacity. Contribution/Results: On regression and scientific machine learning benchmarks, our approach achieves up to 3.2× iteration speedup and reduces average test error by 18.7%, significantly improving both training efficiency and generalization accuracy.
To address the high computational overhead and poor hardware compatibility of Kolmogorov–Arnold Networks (KANs), this paper proposes SineKAN—a novel architecture that, for the first time, replaces conventional B-spline or Fourier basis functions in KANs with learnable sinusoidal activation functions at the edge level. This design preserves the theoretical expressivity guaranteed by the Kolmogorov–Arnold representation theorem while simultaneously enhancing periodic representation capability, gradient stability, and hardware efficiency. Experiments on visual benchmarks—including image classification—demonstrate that SineKAN matches or exceeds the accuracy of both B-spline and Fourier KANs, achieves significantly faster training convergence, and exhibits numerical precision scalability comparable to dense neural networks. By unifying interpretability, lightweight design, and hardware-aware computation, SineKAN establishes a new paradigm for efficient, theoretically grounded, and interpretable neural networks.
Kolmogorov–Arnold Networks (KANs) remain unexplored for 3D point cloud processing, despite their promise as differentiable, structured alternatives to conventional activation functions. Method: We propose PointNet-KAN, the first KAN-based architecture for point clouds, which replaces standard MLP layers in PointNet with learnable KAN layers while preserving permutation invariance. Specifically, we parameterize KAN’s activation functions using Jacobi polynomial families—including Lagrange, Chebyshev, and Gegenbauer polynomials—and introduce a shared-weight design coupled with symmetric aggregation to ensure equivariance and efficiency. Contribution/Results: On ModelNet40 classification and ShapeNet part segmentation, PointNet-KAN achieves performance on par with the original PointNet+MLP baseline using significantly shallower architectures. These results empirically validate KANs as effective, generalizable substitutes for handcrafted or learned activations in geometric deep learning, opening new avenues for structured, differentiable function approximation in point cloud analysis.
To address the limited expressive capacity of Kolmogorov–Arnold Networks (KANs), this paper proposes FC-KAN—a novel KAN architecture that explicitly integrates elementary mathematical functions—including B-splines, Difference-of-Gaussians (DoG), wavelets, radial basis functions (RBFs), and polynomials—via low-dimensional, element-wise operations. We introduce a flexible multi-function composition mechanism encompassing summation, multiplication, quadratic/cubic representations, concatenation, and linear projection. Notably, we present two novel variants: DoG-B-spline and quadratic-linear hybrid functions—the first such incorporation in KAN literature. Extensive evaluation on MNIST and Fashion-MNIST across five independent trials demonstrates that FC-KAN achieves significantly higher average accuracy than standard MLPs and state-of-the-art KAN variants, including BSRBF-KAN, EfficientKAN, and FastKAN. These results empirically validate that explicit, mathematically grounded function composition enhances both the interpretability and modeling capability of neural networks.
This work addresses the practical limitations of Kolmogorov–Arnold Networks (KANs)—notably their high computational cost and lack of a unified implementation framework—by introducing a modular, efficient, and scalable PyTorch-native KAN framework. For the first time, it integrates core features from PyKAN, EfficientKAN, and FastKAN within a unified architecture, supporting adaptive grid scaling, dynamic expansion, and fine-grained structural customization. The framework replaces conventional linear weights with learnable univariate functions and incorporates diverse basis functions, substantially improving computational efficiency. Experiments demonstrate that it accurately reproduces state-of-the-art KAN performance on the California Housing dataset with competitive computational overhead, while enabling flexible exploration of non-standard architectures with negligible performance degradation.
This work addresses the parameter redundancy and poor interpretability of MLPs in the DreamerV3 world model. We propose the first integration of Kolmogorov–Arnold Networks (KANs) and their efficient variant, FastKAN, into an online model-based reinforcement learning framework. Leveraging a JAX-native fully vectorized implementation and a lightweight grid management strategy, we systematically evaluate KAN-based approximators across three core subsystems: visual encoding, latent dynamics modeling, and reward/continuation prediction. Experimental results on the *walker_walk* task show that replacing only the reward and continuation prediction modules with FastKAN achieves performance, sample efficiency, and training speed comparable to the original MLP baseline—while substantially improving parameter efficiency and function-level interpretability. This work establishes a new paradigm for designing interpretable and parameter-efficient world models.
This work investigates the replacement of conventional MLP-based feedforward networks in small language models with Kolmogorov–Arnold Networks (KANs) to enhance interpretability and compression capabilities. By employing a KAN architecture grounded in learnable univariate edge functions—such as GR-KAN—the study explicitly models feedforward pathways and integrates functional principal component analysis (fPCA) for compression alongside edge-level pruning. Systematic evaluations are conducted on benchmarks including BabyLM and Wikitext-103. Experimental results demonstrate that sparse KAN variants enable high-ratio pruning and permit auditable scalar transformations; however, they fail to consistently outperform MLP baselines in language modeling and grammaticality judgment tasks, showing no reliable gains in either performance or inference latency.
This study systematically evaluates the practical utility of Kolmogorov–Arnold Networks (KANs) relative to Multilayer Perceptrons (MLPs) for structured data classification tasks. Under unified preprocessing, architectural design, and hyperparameter settings, a standardized comparison is conducted across twelve publicly available datasets encompassing binary, multiclass, multilabel, and ordinal classification scenarios. Performance is assessed using accuracy, F1 scores, hypothesis testing, and Cohen’s d effect sizes. Results demonstrate that KANs significantly outperform MLPs on most tasks, with a moderate average effect size (d = −0.46), albeit at the cost of higher computational overhead and parameter count. This work presents the first rigorous empirical comparison between KANs and MLPs across diverse structured classification tasks and offers practical guidelines for architecture selection based on accuracy requirements and resource constraints.
This work addresses the challenges of efficiently deploying both Multilayer Perceptrons (MLPs) and the emerging Kolmogorov–Arnold Networks (KANs) on edge devices, where MLPs incur high memory overhead and KANs lack dedicated hardware support, compounded by their divergent compute and memory access patterns that hinder unified acceleration. To bridge this gap, we propose VIKIN, a reconfigurable accelerator that, for the first time, enables efficient unified support for both KANs and MLPs. VIKIN employs a hybrid execution paradigm—pipeline-based for KANs and parallel for MLPs—augmented with two-level sparsity optimizations to accommodate their distinct computational characteristics. Experimental results on real-world datasets show that VIKIN achieves a 1.28× speedup for KAN inference over MLPs with 19.58% lower accuracy loss; compared to an edge GPU, it delivers a 1.25× speedup and 4.87× higher energy efficiency for KAN workloads.