train kolmogorov-arnold networks

Designs and trains neural-network architectures that implement the Kolmogorov–Arnold representation to approximate multivariate functions by composing sums of univariate mappings. Builds low-parameter surrogate models and interpolants that represent complex, potentially multiscale or discontinuous mappings by selecting appropriate layer structures, activation functions, and training procedures.

trainkolmogorov-arnoldnetworks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Sinusoidal Approximation Theorem for Kolmogorov-Arnold Networks

Jul 31, 2025
SG
Sergei Gleyzer
🏛️ University of Alabama

To address the limited expressive power of Kolmogorov–Arnold Networks (KANs), this paper proposes SinKAN—a novel KAN architecture employing learnable-frequency sinusoidal activation functions. Methodologically, SinKAN replaces both inner and outer spline activations in standard KANs with weighted sinusoidal basis functions, where frequencies are trainable parameters and phases are fixed to uniformly spaced constants; it further incorporates the Lorentz–Sprecher simplification framework to preserve theoretical soundness. The key contribution is the first systematic integration of learnable frequency parameters into KANs, substantially enhancing approximation capability for high-frequency and nonsmooth multivariate functions. Experiments demonstrate that SinKAN outperforms fixed-frequency Fourier networks across multiple benchmark function approximation tasks, achieves performance on par with multilayer perceptrons (MLPs), and retains KANs’ intrinsic advantages—namely, interpretability and parameter sparsity.

Compares performance with Fourier transforms and multilayer perceptrons (MLPs)Proposes sinusoidal-based Kolmogorov-Arnold Networks (KANs) for multivariable function approximationReplaces spline activations with learnable-frequency sinusoidal functions

This study addresses the longstanding trade-off between accuracy and efficiency in conventional neural networks by systematically comparing Kolmogorov–Arnold Networks (KANs) with Multilayer Perceptrons (MLPs). Built upon the Kolmogorov representation theorem, KANs employ learnable spline-based activation functions within a grid-structured architecture, achieving both high accuracy and low computational cost. Experimental results across diverse tasks—including nonlinear function approximation, time series forecasting, and multivariate classification—demonstrate that KANs consistently outperform MLPs in predictive performance while significantly reducing floating-point operations (FLOPs). These findings position KANs as an interpretable, efficient, and accurate alternative architecture, particularly well-suited for resource-constrained and real-time applications.

computational efficiencyKolmogorov-Arnold NetworksMulti-Layer Perceptrons

KAN: Kolmogorov-Arnold Networks

Apr 30, 2024
ZL
Ziming Liu
🏛️ Massachusetts Institute of Technology | California Institute of Technology | Northeastern University

This paper addresses the poor interpretability and limited accuracy of traditional multilayer perceptrons (MLPs) by proposing Kolmogorov–Arnold Networks (KANs), a novel neural architecture grounded in the Kolmogorov–Arnold representation theorem. Unlike MLPs—which employ fixed activation functions and linear weight layers—KANs parameterize learnable B-spline functions on **edges**, enabling flexible nonlinear modeling; nodes perform only summation, with no weights or activations. This design yields three key contributions: (1) Both theoretical analysis and empirical evaluation demonstrate superior neural scaling laws: small KANs significantly outperform large MLPs in data fitting and partial differential equation solving. (2) KANs enable direct parameter visualization and semantic interpretability, facilitating human–machine collaborative scientific discovery. (3) KANs constitute the first general-purpose neural network paradigm featuring edge-level learnable activations.

Enhancing accuracy and interpretability in neural networks.Facilitating discovery in mathematics and physics.Proposing Kolmogorov-Arnold Networks as MLP alternatives.

FC-KAN: Function Combinations in Kolmogorov-Arnold Networks

Sep 03, 2024
HT
Hoang Thang Ta
🏛️ Dalat University | University of Nebraska Omaha | Instituto Politécnico Nacional

To address the limited expressive capacity of Kolmogorov–Arnold Networks (KANs), this paper proposes FC-KAN—a novel KAN architecture that explicitly integrates elementary mathematical functions—including B-splines, Difference-of-Gaussians (DoG), wavelets, radial basis functions (RBFs), and polynomials—via low-dimensional, element-wise operations. We introduce a flexible multi-function composition mechanism encompassing summation, multiplication, quadratic/cubic representations, concatenation, and linear projection. Notably, we present two novel variants: DoG-B-spline and quadratic-linear hybrid functions—the first such incorporation in KAN literature. Extensive evaluation on MNIST and Fashion-MNIST across five independent trials demonstrates that FC-KAN achieves significantly higher average accuracy than standard MLPs and state-of-the-art KAN variants, including BSRBF-KAN, EfficientKAN, and FastKAN. These results empirically validate that explicit, mathematically grounded function composition enhances both the interpretability and modeling capability of neural networks.

K-A NetworkMathematical FunctionsPerformance Improvement

SineKAN: Kolmogorov-Arnold Networks Using Sinusoidal Activation Functions

Jul 04, 2024
EA
Eric A. F. Reinhardt
🏛️ The University of Alabama

To address the high computational overhead and poor hardware compatibility of Kolmogorov–Arnold Networks (KANs), this paper proposes SineKAN—a novel architecture that, for the first time, replaces conventional B-spline or Fourier basis functions in KANs with learnable sinusoidal activation functions at the edge level. This design preserves the theoretical expressivity guaranteed by the Kolmogorov–Arnold representation theorem while simultaneously enhancing periodic representation capability, gradient stability, and hardware efficiency. Experiments on visual benchmarks—including image classification—demonstrate that SineKAN matches or exceeds the accuracy of both B-spline and Fourier KANs, achieves significantly faster training convergence, and exhibits numerical precision scalability comparable to dense neural networks. By unifying interpretability, lightweight design, and hardware-aware computation, SineKAN establishes a new paradigm for efficient, theoretically grounded, and interpretable neural networks.

Image RecognitionKolmogorov-Arnold NetworkPerformance Evaluation

Latest Papers

What's happening recently
View more

This work addresses the inefficiency of conventional neural network training in complex tasks such as physics-informed neural networks (PINNs), where the lack of exploitable structure hinders optimization. For the first time, a multigrid algorithm is introduced into the training of Kolmogorov–Arnold Networks (KANs). By establishing an equivalence between KANs and multi-channel MLPs through basis transformation, the authors design a multilevel optimization strategy that progressively refines spline knots across layers. Leveraging the compact support of spline basis functions and analytical geometric interpolation operators, the method enables lossless transfer and complementary optimization from coarse to fine models. Experiments demonstrate that this approach achieves several orders of magnitude improvement in accuracy over standard KAN or MLP training, significantly accelerating convergence and enhancing generalization across multiple tasks.

Kolmogorov-Arnold Networksmultilevel trainingneural network structure

This work addresses the high parameter redundancy, poor scalability, and low training efficiency of Kolmogorov–Arnold Networks (KANs) by introducing hyperbolic geometry into the KAN framework for the first time. The proposed method embeds inputs into the bounded hyperbolic latent space of the Poincaré ball, performs KAN-style updates in the tangent space, and incorporates a low-rank prototype module to share function transformations across hidden dimensions. By integrating spline-based function learning, a radial coordinate structure, and a radius control mechanism, the approach enhances model interpretability and training stability. Empirical evaluations on eight benchmark datasets demonstrate that the method achieves predictive performance comparable to or better than existing approaches while substantially improving parameter efficiency.

efficiencyfunction approximationKolmogorov-Arnold Networks

This work proposes Feature-Enhanced Kolmogorov–Arnold Networks (FEKAN) to address the limitations of existing Kolmogorov–Arnold Networks (KANs), which suffer from high computational cost and slow convergence, hindering their scalability. FEKAN enhances model expressivity and accelerates convergence through a feature augmentation mechanism without introducing additional trainable parameters, while preserving the interpretability inherent to KANs. Grounded in the Kolmogorov–Arnold representation theorem, FEKAN is well-suited for tasks such as function approximation, physics-informed neural networks, and neural operators. Experimental results demonstrate that FEKAN consistently outperforms current KAN variants across multiple benchmarks, achieving both higher accuracy and faster convergence.

computational costKolmogorov-Arnold Networkspractical applicability

This work addresses the practical limitations of Kolmogorov–Arnold Networks (KANs)—notably their high computational cost and lack of a unified implementation framework—by introducing a modular, efficient, and scalable PyTorch-native KAN framework. For the first time, it integrates core features from PyKAN, EfficientKAN, and FastKAN within a unified architecture, supporting adaptive grid scaling, dynamic expansion, and fine-grained structural customization. The framework replaces conventional linear weights with learnable univariate functions and incorporates diverse basis functions, substantially improving computational efficiency. Experiments demonstrate that it accurately reproduces state-of-the-art KAN performance on the California Housing dataset with competitive computational overhead, while enabling flexible exploration of non-standard architectures with negligible performance degradation.

computational costframework inconsistencyKAN implementations

This study addresses the challenge of efficiently and accurately predicting surface pressure distributions on subsonic and transonic airfoils by introducing Kolmogorov–Arnold Networks (KANs)—based on the Kolmogorov–Arnold representation theorem—into aerodynamic surrogate modeling for the first time. The authors systematically compare KANs against Multilayer Perceptrons (MLPs) and Graph Neural Networks (GNNs). Experimental results demonstrate that KANs achieve effective prediction of pressure coefficient distributions with lower model complexity and faster training speeds, albeit with slightly lower accuracy than hyperparameter-optimized MLPs. While GNNs attain the highest predictive accuracy, they incur substantially greater computational costs. The work also highlights challenges associated with KANs, particularly regarding training stability and sensitivity to hyperparameters, thereby offering a novel architectural alternative and empirical benchmark for surrogate modeling in fluid dynamics.

aerodynamic predictionfluid dynamicsKolmogorov Arnold networks

Hot Scholars

AH

Amaury Habrard

Professor of Computer Science, University Jean Monnet of Saint-Etienne (France)
machine learning
DR

Dara Rahmati

Faculty of Computer Science and Engineering, Shahid Beheshti University
Computer ArchitectureNetworks on ChipHardware AcceleratorsScientific Computing
ND

Nilanjan Dey

Professor, Techno International New, Town, Kolkata
Medical ImagingBiomedical TechnologiesMachine LearningHeuristic Algorithm Applications
HP

Hemant Purohit

Associate Professor, School of Computing, George Mason University
Human-centered AISemantic WebHuman-AI CollaborationCrisis Informatics
BK

Bikram Keshari Parida

Sun Moon University, South Korea
Artificial IntelligenceGeneral RelativityQuantum Field theoryAstrophysics