deep learning

Designs, builds, and evaluates models based on multi‑layer neural networks and their training pipelines, including architectures, loss functions, optimization algorithms, regularization, and representation learning for supervised, unsupervised, and self‑supervised objectives. Analyzes model behavior, scaling, generalization, robustness, and efficiency, and implements training and inference systems and techniques for deployment.

deeplearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Tensorflow Pretrained Models

Sep 20, 2024
KC
Keyu Chen
🏛️ Georgia Institute of Technology | Indiana University | Kyoto University | AppCubic | Rutgers University | Purdue University | University of Wisconsin-Madison | National Taiwan Normal University

High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.

Comparing linear probing versus fine-tuning approaches in transfer learningExploring TensorFlow pre-trained models for image classification tasksProviding practical guidance and code examples for deep learning implementation

SAGRAD: A Program for Neural Network Training with Simulated Annealing and the Conjugate Gradient Method

Jun 17, 2015
JB
J. Bernal
🏛️ National Institute of Standards and Technology | CINVESTAV-Tamaulipas

To address the non-convex optimization challenge in neural network classification—specifically, susceptibility to poor local minima and flat regions—this paper proposes SAGRAD, a batch-training algorithm integrating Simulated Annealing (SA) with Møller’s Scaled Conjugate Gradient (SCG) method. Its core innovation lies in the first incorporation of SA into the SCG framework, enabling a dynamic restart and escape mechanism that synergistically balances global exploration and local acceleration. Implemented in Fortran 77, SAGRAD incorporates efficient Hessian-vector multiplication, optimized gradient computation, and an adaptive SA weight initialization strategy. Empirical evaluation across multiple classification benchmarks demonstrates significantly improved convergence robustness and generalization performance, while markedly reducing the probability of converging to suboptimal local minima. These results validate SAGRAD’s effectiveness and practicality for non-convex optimization in neural network training.

ClassificationLocal OptimaNeural Network Training

This paper addresses non-convex optimization in shallow neural networks across three fundamental tasks: exact representation, function approximation, and regression. We propose a unified convexification framework grounded in mean-field theory. Theoretically, we rigorously prove that the convexified problem admits no relaxation gap and derive an interpretable, closed-form generalization bound that explicitly characterizes hyperparameter influence and provides principled guidelines for optimal selection. Algorithmically, we design a task-adaptive solver: for low-dimensional settings, we employ the simplex method with theoretical guarantees of exact recovery; for high-dimensional settings, we combine sparsification with gradient descent to achieve efficient approximation. Empirical results demonstrate substantial improvements in test performance over standard training heuristics. To our knowledge, this is the first work achieving unified modeling, gap-free convexification, and joint optimization of generalization and algorithmic efficiency across all three tasks.

Convexify non-convex optimization in shallow neural networksDevelop efficient algorithms for high-dimensional dataset discretizationEstablish generalization bounds for neural network solutions

Three Mechanisms of Feature Learning in a Linear Network

Jan 13, 2024
YX
Yizhou Xu
🏛️ Abdus Salam International Center for Theoretical Physics | Massachusetts Institute of Technology | NTT Research

This work investigates how neural network width governs training dynamics. For single-hidden-layer linear networks, we derive the first exact analytical solution of learning dynamics at arbitrary finite width, unifying the characterization of the two-phase evolution—kernel learning and feature learning—and establishing a complete phase diagram parameterized by width, layer-wise learning rates, and initialization scale. Methodologically, we integrate analytical dynamical systems analysis, phase-diagram modeling, and empirical validation on nonlinear networks. Crucially, we identify three novel mechanisms operative during the feature-learning phase: alignment learning, de-alignment learning, and rescaling learning—each transcending the conventional kernel-method paradigm. These theoretical insights are empirically reproduced in realistic deep networks, offering a new conceptual framework for understanding training dynamics and designing adaptive optimization algorithms. (138 words)

Analyzes learning dynamics in neural networksExplores hyperparameter impact on training trajectoriesIdentifies feature learning mechanisms in networks

ReLU Neural Networks with Linear Layers are Biased Towards Single- and Multi-Index Models

May 24, 2023
SP
Suzanna Parkinson
🏛️ University of Chicago | Marquette University

This paper investigates the structural inductive bias of overparameterized deep ReLU networks (depth > 2) in interpolation learning. Focusing on architectures that prepend a linear layer to shallow subnetworks, we develop an analytical framework based on representation cost—defined as the minimal squared ℓ²-norm of network weights—and show that this design substantially reduces the *mixed variation* of learned functions. Consequently, the network implicitly favors structured functions supported on low-dimensional subspaces: those exhibiting restricted variation along orthogonal directions and admitting exact characterization via single- or multi-index models. Theoretically and empirically, we demonstrate that this mechanism aligns learned weights closely with the true latent low-dimensional subspace, achieves near-optimal low-dimensional approximation on multi-index model-generated data, and improves generalization. To our knowledge, this is the first work to rigorously characterize the critical role of input-side linear layers in inducing implicit regularization for deep networks through representation bias.

Analyzes representation cost favoring low mixed variation.Explores function properties in deep ReLU networks.Shows linear layers improve generalization in networks.

Latest Papers

What's happening recently
View more

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

This study investigates how to assign optimal layer-wise learning rates in the early phase of training deep linear neural networks to minimize test loss. By deriving an exact closed-form solution after two steps of gradient descent, the authors characterize the initial training dynamics and propose a gradient-based inter-layer learning rate scaling strategy, constructing an analytically tractable surrogate loss function. Theoretical analysis reveals that employing unequal learning rates across layers at initialization enhances performance, whereas uniform learning rates become preferable in subsequent steps. This work presents the first precise theoretical characterization of the dynamics during the first two training iterations, and numerical experiments confirm that the proposed strategy significantly reduces test loss compared to baseline approaches.

early training dynamicslayer-wise optimizationlearning rate

This work proposes a biologically plausible multilayer neuronal network model that moves beyond the simplified weighted-sum neuron paradigm of conventional artificial neural networks, which struggles to capture the true learning mechanisms of biological systems. The proposed architecture employs cascaded adaptive combiners to enable efficient online, streaming learning without relying on backpropagation. By integrating neuron-level computations that more closely mirror biological reality with practical machine learning principles, the model establishes a concise and scalable learning framework. Empirical evaluation on image classification tasks demonstrates competitive performance, confirming both the efficacy and practicality of the approach.

alternative to backpropagationbiological neuron modelscascaded adaptive combiners

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Hot Scholars

YJ

Yuling Jiao

University of Wuhan
Deep learningScientific and statistical computingInverse problem
FK

Freddie Kalaitzis

Senior Research Fellow, University of Oxford
Machine LearningComputational StatisticsEarth ObservationAI for Social Good
SM

Shuangge Ma

Yale University
Genetic epidemiologySurvival analysisCancerHealth economics
BK

Berkcan Kapusuzoglu

Capital One
Machine LearningOptimization under UncertaintyUncertainty QuantificationComputational
PJ

Pierre Jannin

MediCIS, LTSI, Inserm, Université de Rennes
Surgical data scienceComputer Assisted Surgery