benchmark pretrained deep models

Designs and implements standardized evaluation suites and experimental pipelines for pretrained deep neural networks, specifying data splits, transfer-learning and fine-tuning protocols, and a set of quantitative performance and robustness metrics. Builds comparative analyses and leaderboards across architecture families (e.g., CNNs, transformers, hybrids) to quantify performance gaps, ranking, and transferability between models.

benchmarkpretraineddeepmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Tensorflow Pretrained Models

Sep 20, 2024
KC
Keyu Chen
🏛️ Georgia Institute of Technology | Indiana University | Kyoto University | AppCubic | Rutgers University | Purdue University | University of Wisconsin-Madison | National Taiwan Normal University

High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.

Comparing linear probing versus fine-tuning approaches in transfer learningExploring TensorFlow pre-trained models for image classification tasksProviding practical guidance and code examples for deep learning implementation

Benchmarking Transferability: A Framework for Fair and Robust Evaluation

Apr 28, 2025
AK
Alireza Kazemi
🏛️ The University of Queensland

This work addresses the fundamental lack of fairness and robustness in evaluating model transferability across domains. We propose the first systematic, standardized benchmarking framework for assessing cross-domain transfer capability. Our method introduces a unified multi-source-domain–target-domain evaluation protocol, encompassing diverse transfer tasks and perturbation-robustness analysis, and adopts head-training (i.e., linear-probe fine-tuning) as the consistent evaluation paradigm. Empirical analysis reveals significant performance discrepancies among existing transferability metrics under varying experimental settings, undermining their reliability. Our framework substantially improves assessment fidelity, yielding an average 3.5% gain in transfer performance under standard head-training configurations. To foster reproducibility and rigorous comparison, we fully open-source all code, datasets, and evaluation pipelines—establishing a new, standardized paradigm for transferability measurement.

Addressing inconsistencies in transferability measurement methodsEvaluating reliability of transferability scores across domainsProposing standardized framework for robust transferability assessment

CTBENCH: A Library and Benchmark for Certified Training

Jun 07, 2024
YM
Yuhao Mao
🏛️ ETH Zurich | INSAIT | Sofia University

Existing certified training algorithms suffer from inconsistent evaluation protocols and suboptimal hyperparameter tuning, leading to incomparable performance claims and unreliable SOTA conclusions. Method: We introduce CTBENCH—the first unified benchmark for certified training—enabling fair, cross-algorithm evaluation of mainstream methods (e.g., IBP, CROWN-IBP, DeepPoly) under a standardized training pipeline, consistent ℓ∞/ℓ2 certification framework, and systematic hyperparameter optimization (grid search + Bayesian optimization). Contribution/Results: Our evaluation reveals that most recently proposed algorithms are substantially overestimated in prior work; after baseline enhancement, their relative improvements drop by over 40% on average. Crucially, all methods achieve significantly higher certified accuracy on CTBENCH than reported in their original papers. This work establishes a reproducible, extensible standard for evaluating certified training, redefining both the robustness training baseline and the SOTA landscape.

Method EvaluationNeural Network ReliabilityStandardization

Benchmarking Neural Network Training Algorithms

Jun 12, 2023
GE
George E. Dahl
🏛️ Google | University of Tübingen | Vector Institute | Dalhousie University | University of Toronto | Stanford University | Meta AI | Dell Technologies

Fair evaluation of deep learning training algorithms faces three key challenges: inconsistent termination criteria, high workload sensitivity, and difficulty isolating hyperparameter tuning. This paper introduces AlgoPerf—the first time-oriented, multi-workload training algorithm benchmark—featuring robustness-aware workload variant design and a standardized termination protocol, with hyperparameter tuning rigorously isolated. Evaluated on a unified hardware platform, AlgoPerf employs a diverse multi-task workload suite and a systematic optimizer comparison methodology to enable latency-accuracy co-evaluation across models, datasets, and hardware. Experiments reveal substantial latency disparities among mainstream optimizers, establish reproducible state-of-the-art baselines, and deliver the first quantitative, fair, and engineering-practical evaluation standard for training algorithm improvement.

Compare hyperparameter-tuned algorithms fairlyIdentify state-of-the-art training algorithms reliablyMeasure training time accurately and decide completion

Latest Papers

What's happening recently
View more

This work addresses the lack of publicly available, diverse benchmark datasets for systematically evaluating neural network code verification, refactoring, and migration tools. To bridge this gap, the authors propose a novel approach that leverages large language models to automatically generate neural network code spanning a wide range of architectural components, input types, and tasks. The generated samples are rigorously validated through static analysis and symbolic tracing to ensure both structural and semantic adherence to precise design specifications. The resulting benchmark comprises 608 correct and diverse neural network implementations, constituting the first publicly reusable dataset of its kind. This resource significantly advances reproducibility and enables systematic evaluation in research on neural network reliability and maintainability.

adaptabilitydatasetneural networks

Traditional small-scale datasets such as MNIST struggle to effectively differentiate the performance of advanced neural network architectures, particularly due to their limited capacity to capture inductive biases for sequential data. This work presents the first systematic evaluation of diverse models—including ResNet, Temporal Convolutional Networks (TCNs), and Dilated Convolutional Neural Networks (DCNNs)—on the lightweight, structured sequential dataset MNIST-1D, benchmarking them against baselines such as logistic regression, MLPs, CNNs, and GRUs. Experimental results demonstrate that TCNs and DCNNs substantially outperform conventional approaches, achieving near-human accuracy, while ResNet also exhibits strong performance. These findings validate MNIST-1D as an efficient and effective benchmark and underscore the critical role of inductive bias in resource-constrained settings.

architectural comparisoninductive biasesMNIST-1D

This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.

AI evaluationbenchmarkingdeployment conditions

This study addresses the challenge of selecting suitable ImageNet-pretrained models for image classification tasks in target domains by introducing a multidimensional evaluation framework. The authors systematically fine-tune the output layers and general parameters of eleven pretrained models across five diverse datasets, evaluating their performance under both single-run and multiple-run training settings. Through comprehensive assessment of accuracy, accuracy density, training time, and model size, the work quantifies the cross-domain transferability differences among pretrained models, revealing consistent patterns in how model characteristics align with task-specific requirements. These findings provide empirical evidence and practical guidelines for informed model selection in real-world applications.

image classificationmodel selectionpre-trained models

Hot Scholars

MH

Mehrtash Harandi

Department of Electrical and Computer Systems Engineering, Monash University
Machine LearningComputer Vision
VK

Vicky Kouni

Postdoctoral Fellow, Cambridge University
Learning TheoryMachine LearningDeep UnfoldingSignal Processing
QL

Qian Lin

Research Engineer, ByteDance
DatabaseDistributed SystemData Streams
WL

Weichen Liu

College of Computing and Data Science, Nanyang Technological University
Embedded SystemsMultiprocessor SystemsNetwork-on-Chip