neural networks

Designs, implements, and evaluates artificial neural network models and their components—such as layer architectures, activation functions, loss functions, optimization and training procedures, and regularization—used to map inputs to outputs for prediction, representation, or control tasks. Analyzes model behavior including learning dynamics, generalization, efficiency, robustness, and scalability.

neuralnetworks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.8
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$208K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

This work addresses the limited interpretability of artificial neural networks despite their high predictive accuracy. To bridge this gap, the authors propose a novel approach that partitions ReLU-based neural networks into individual units and represents them as three-dimensional bit tensors. For the first time, ternary formal concept analysis is introduced to extract symbolic logical rules from these tensors, which preserve the original classification performance. These rules are then organized into a human-readable logical decision tree. The method provides a transparent representation of internal attribute interactions within the network, significantly enhancing model interpretability without compromising accuracy. This study thus offers a new pathway toward symbolic interpretation of neural networks, combining the strengths of connectionist models with the clarity of symbolic reasoning.

artificial neural networkconcept analysisinterpretability

This study addresses the challenge classical frequentist statisticians face in understanding neural networks by proposing a reconstruction of neural networks through the lens of linear regression. By simplifying network architecture and integrating statistical interpretability techniques, the approach reformulates deep learning models into a modeling paradigm familiar to statisticians. The method preserves the expressive power of neural networks while offering intuitive parameter interpretations and customizable pathways, thereby significantly lowering the cognitive barrier for statisticians entering the field of deep learning. The resulting framework balances theoretical rigor with practical usability, fostering meaningful integration and methodological exchange between traditional statistics and modern deep learning.

barrier to entryfrequentist perspectivelinear regression

Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Tensorflow Pretrained Models

Sep 20, 2024
KC
Keyu Chen
🏛️ Georgia Institute of Technology | Indiana University | Kyoto University | AppCubic | Rutgers University | Purdue University | University of Wisconsin-Madison | National Taiwan Normal University

High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.

Comparing linear probing versus fine-tuning approaches in transfer learningExploring TensorFlow pre-trained models for image classification tasksProviding practical guidance and code examples for deep learning implementation

SAGRAD: A Program for Neural Network Training with Simulated Annealing and the Conjugate Gradient Method

Jun 17, 2015
JB
J. Bernal
🏛️ National Institute of Standards and Technology | CINVESTAV-Tamaulipas

To address the non-convex optimization challenge in neural network classification—specifically, susceptibility to poor local minima and flat regions—this paper proposes SAGRAD, a batch-training algorithm integrating Simulated Annealing (SA) with Møller’s Scaled Conjugate Gradient (SCG) method. Its core innovation lies in the first incorporation of SA into the SCG framework, enabling a dynamic restart and escape mechanism that synergistically balances global exploration and local acceleration. Implemented in Fortran 77, SAGRAD incorporates efficient Hessian-vector multiplication, optimized gradient computation, and an adaptive SA weight initialization strategy. Empirical evaluation across multiple classification benchmarks demonstrates significantly improved convergence robustness and generalization performance, while markedly reducing the probability of converging to suboptimal local minima. These results validate SAGRAD’s effectiveness and practicality for non-convex optimization in neural network training.

ClassificationLocal OptimaNeural Network Training

Latest Papers

What's happening recently
View more

This study addresses the limited interpretability, reliability, and controllability of deep learning models—particularly large generative models—stemming from their “black-box” nature. From the perspective of representation learning, the work integrates optimization theory, information theory, linear algebra, and calculus to construct a unified mathematical framework that systematically elucidates the internal mechanisms of neural networks. This framework transforms network architecture design from an empirical, alchemy-like practice into a principled, analytically tractable process, substantially enhancing model interpretability and controllability. The proposed approach achieves performance on par with or exceeding that of existing black-box models across multiple tasks, thereby unifying theoretical rigor with practical efficacy.

black boxdeep learninggenerative models

This work addresses the lack of formal correctness guarantees for neural networks in safety-critical applications by proposing a unified formal verification framework applicable to diverse architectures, including feedforward networks, recurrent networks, and Transformers. The framework integrates expressive specification languages—such as linear temporal logic—with advanced algorithmic techniques, including abstract interpretation and SMT solving, to uniformly capture both semantic representations and verification methodologies. By generalizing existing verification theories to broader classes of models, this study not only extends the theoretical foundations of neural network verification but also establishes a scalable and rigorous basis for safety analysis of complex architectures. Consequently, it advances both the theoretical understanding and practical deployment of trustworthy artificial intelligence systems.

Formal VerificationNeural Network VerificationRecurrent Neural Networks

This work addresses the lack of publicly available, diverse benchmark datasets for systematically evaluating neural network code verification, refactoring, and migration tools. To bridge this gap, the authors propose a novel approach that leverages large language models to automatically generate neural network code spanning a wide range of architectural components, input types, and tasks. The generated samples are rigorously validated through static analysis and symbolic tracing to ensure both structural and semantic adherence to precise design specifications. The resulting benchmark comprises 608 correct and diverse neural network implementations, constituting the first publicly reusable dataset of its kind. This resource significantly advances reproducibility and enables systematic evaluation in research on neural network reliability and maintainability.

adaptabilitydatasetneural networks

Standard artificial neural networks rely on simplified point-neuron models that fail to capture the complex computational properties of biological neurons. This work proposes, for the first time, integrating biologically plausible dynamical models—derived from cutting-edge neuroscience and closely aligned with cortical neuron physiology—directly into deep network architectures as drop-in replacements for conventional units, without increasing parameter count. Theoretical analysis and empirical experiments demonstrate that this approach substantially enhances model expressivity, accelerates learning, and improves robustness, while simultaneously reducing overfitting and decreasing reliance on large training datasets.

artificial neural networkscortical cellsneural modeling